Task urgency determination based on acoustic features of audio data

CN116724353BActive Publication Date: 2026-09-08MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202280010473.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-01-29
Filing Date
2022-01-13
Publication Date
2026-09-08
Estimated Expiration
2042-01-13

AI Technical Summary

Benefits of technology

[0005]The disclosed techniques include determining audio substream data of received audio input associated with a task command, and generating values ​​of acoustic features for embedded vector data and/or the corresponding audio substream data. The disclosed techniques also include machine learning (ML) systems for determining the importance and urgency of the task based on the values ​​of the embedded vector data and/or acoustic features. The ML system can use a regression model for sequential analysis of the acoustic features, or use an embedding model for parallel processing. The ML model can use neural networks, such as recurrent neural networks. As a result, the disclosed techniques efficiently and accurately determine the level of urgency and importance of user utterances or audio data by focusing on classifying the acoustic features of the data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116724353B_ABST
    Figure CN116724353B_ABST
Patent Text Reader

Abstract

Systems and methods are provided for determining importance and urgency of a task based on acoustic features of audio input associated with the task. The determination includes classifying the task into one or more categories associated with importance, urgency, and priority of the task. The classification can use a trained machine learning model of acoustic features and embeddings for a neural network. A task classifier uses feature acoustics of one or both of foreground and background audio. The feature acoustics include pitch, tone, and volume over a duration of the audio input. A combination of the acoustic features determines the category associated with the task. The machine learning model includes a regression model of acoustic features over time and a model that utilizes embeddings for a neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Various calendar and task management apps are widely available for conveniently creating and managing tasks and applications. Some apps reside on mobile devices (e.g., smartphones and tablets), while others operate on smart devices (e.g., smart speakers). These apps interactively receive commands from users to create tasks (e.g., schedule calendar events, make phone calls, etc.). As users appreciate the benefits of app-based task management, it is desirable to develop technologies that better meet users' needs for managing tasks in various scenarios of daily life without adding to their burden.

[0002] It is with regard to these and other general considerations that the aspects disclosed herein have been made. Furthermore, while relatively specific problems may be discussed, it should be understood that the examples are not limited to solving specific problems identified in the background or elsewhere in this disclosure. Summary of the Invention

[0003] According to this disclosure, the above and other problems are addressed by determining the importance and priority of a task based on the acoustic features of audio data associated with the task.

[0004] While prior methods exist for identifying tasks based on received voice commands, this disclosure relates to determining the importance and urgency of a task based on acoustic features of audio data associated with the task. The task may include initiating communication (e.g., a telephone call) or an alarm and performing procedures (e.g., scheduling or updating calendar appointments and events). Audio data may include voice data and ambient sounds. Specifically, task-related audio data includes background sounds such as sirens, engine sounds, people talking, and ambient noise. The disclosed techniques address this problem by analyzing voice data or audio data to identify its acoustic features, classifying the data into categories that characterize situations based on levels of urgency and importance. For example, voice data may include commands to be performed, while the acoustic features of the voice data may indicate the level of stress by the speaker, and ambient noise may indicate the level of urgency of the situation. Machine learning processes determine the importance and urgency of a task by analyzing the acoustic features of received audio data.

[0005] The disclosed techniques include determining audio substream data of received audio input associated with a task command, and generating values ​​of acoustic features for embedded vector data and / or the corresponding audio substream data. The disclosed techniques also include machine learning (ML) systems for determining the importance and urgency of the task based on the values ​​of the embedded vector data and / or acoustic features. The ML system can use a regression model for sequential analysis of the acoustic features, or use an embedding model for parallel processing. The ML model can use neural networks, such as recurrent neural networks. As a result, the disclosed techniques efficiently and accurately determine the level of urgency and importance of user utterances or audio data by focusing on classifying the acoustic features of the data.

[0006] The present invention is provided as an option to present the concepts in a simplified form, which are further described in the detailed description below. The present invention is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Additional aspects, features, and / or advantages of the examples will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this disclosure. Attached Figure Description

[0007] Non-restrictive and non-exhaustive examples are described with reference to the following figures.

[0008] Figure 1 An overview of an example system for determining the importance and urgency of a task based on acoustic features of audio data using a machine learning model, according to aspects of this disclosure, is shown.

[0009] Figure 2 An example of a data structure for determining the importance and urgency of a task based on the acoustic features of audio data according to aspects of this disclosure is shown.

[0010] Figure 3 An example of a data structure is shown that, according to aspects of this disclosure, determines the importance and urgency of a task by extracting context from speech and audio data.

[0011] Figure 4 An example of a data structure for determining the importance and urgency of a task according to aspects of this disclosure is shown.

[0012] Figure 5 An example of a method, according to aspects of this disclosure, for using machine learning models to determine the importance and urgency of a task based on acoustic features of audio data associated with the task.

[0013] Figure 6Examples of methods for determining the importance and urgency of a task based on acoustic features of audio data associated with the task, according to aspects of this disclosure, are shown.

[0014] Figure 7 This is a block diagram illustrating an example physical component of a computing device in which aspects of this disclosure can be utilized.

[0015] Figure 8A This is a simplified illustration of aspects of this disclosure that can be utilized in practice with mobile computing devices.

[0016] Figure 8B This is another simplified block diagram of a mobile computing device in which aspects of this disclosure can be utilized. Detailed Implementation

[0017] Various aspects of this disclosure are described more fully below with reference to the accompanying drawings, which are taken in part and illustrate specific example aspects. However, different aspects of this disclosure may be implemented in many different ways and should not be construed as limited to the aspects set forth herein; rather, these aspects are provided so that this disclosure will be thorough and complete, and will fully convey the scope of these aspects to those skilled in the art. The aspects may be practiced as methods, systems, or apparatuses. Thus, aspects may take the form of hardware implementations, entirely software implementations, or implementations combining software and hardware aspects. Therefore, the following detailed description is not intended to be limiting.

[0018] Traditional systems generate tasks and determine their importance and urgency based on user interactions with the device (such as calendar apps and digital assistants) and their relationship to other tasks. These systems generate tasks or items in a to-do list when users interact with the device via graphical user interfaces, voice commands, and other methods as input commands. Some other systems automatically generate tasks based on predefined sets of rules or patterns of user activity determined by the system.

[0019] The system can prioritize tasks in various ways. In some cases, it can prioritize tasks based on commands received from the user. In others, it can automatically prioritize tasks based on predefined prioritization rules and pattern matching between activities and those rules. Some applications prioritize tasks relative to existing tasks based on the context of the corresponding task. However, problems arise when tasks based on a given command indicate the same or similar context, which users expect the system to automatically determine the urgency and priority of tasks. For example, tasks with the command "Call Mom" ​​verbally entered by the user at 7 p.m. on two separate Saturdays could indicate the same level of priority for the "Call Mom" ​​task. However, the user might expect one task to have a higher priority than any other task at that time, since the user was told to "Call Mom" ​​when a car accident occurred at 7 p.m. on the second Saturday.

[0020] Traditional task management relies on receiving commands or events based on predetermined rules to generate tasks and prioritize them. This invention determines the urgency and priority of tasks based on sound / audio / acoustic features, assigning importance, urgency, and / or priority to tasks. In various aspects, tasks can be explicit entries in any to-do list or task management application. Alternatively, tasks can be implicit entries extracted from emails or a set of actions requested by a voice assistant. Task entries can be captured via multimodal channels. Machine learning (ML) models use acoustic features (not limited to voice in foreground audio, but also including background audio). The acoustic features determine the importance of tasks based on feedback from the model. Additional actions categorize tasks based on importance and provide users with improved task execution.

[0021] As discussed in more detail below, this disclosure relates to using machine learning models to determine the importance and urgency of a task based on acoustic features of audio data associated with that task. Specifically, the acoustic features of the audio data may include pitch, volume, prosody, etc. Audio data may be associated with a task when the audio data contains commands for that task or when the task can be inferred from the audio data. The audio data may include one or more audio substreams. Accordingly, substreams may be associated with different types of audio data and context.

[0022] Figure 1An overview of an example system for determining the importance and urgency of a task based on acoustic features of audio data using a machine learning model, according to aspects of this disclosure, is shown. System 100 represents a system for a machine learning model or neural network to generate tasks and determine their importance, urgency, and priority using a model generator and a task manager. System 100 includes a client device 110, a task manager 130, and a model generator 160. A user 102 interacts with the client device 110. The task manager 130 and the model generator 160 may be implemented as one or more servers connected to the client device 110 via a network (not shown). Additionally or alternatively, the task manager 130 and the model generator 160 may include one or more sets of instructions to be executed at least partially as an application on the client device 110. The client device 110 communicates with the task manager 130. Task Manager 130 includes an audio receiver 132, an audio substream generator 134, an audio type / feature determiner 136, a storage device for storing trained audio type models 138, a task generator 140, a task classifier 142, and a task processor 146. Task Manager 130 communicates with Model Generator 160. Model Generator 160 generates trained models by training models for determining audio types and task categories. Model Generator 160 includes an audio type model trainer 162, a task classification model trainer 164, a storage device for storing training audio type data 166, a storage device for storing training task category data 168, and a model deployer 170.

[0023] Client device 110 may be a smart speaker, smartphone, personal computing device, or general-purpose computing device that provides input and output capabilities. Client device 110 includes input components (e.g., one or more microphones, cameras, touch sensors, keyboards, etc.) and output components (e.g., one or more speakers, displays, etc.). In various respects, client device 110 may communicate with task manager 130 via a network (not shown).

[0024] In various aspects, the audio receiver 132 may receive audio data 112 from an audio input portion (e.g., a microphone) in the client device 110. The audio receiver 132 receives the audio data 112 as an audio signal stream. The audio substream generator 134 generates one or more audio substreams based on the received audio data 112. In various aspects, the audio data stream 112 may include one or more substreams correspondingly representing foreground sound 114 and background sound 122. In various aspects, the audio substream generator 134 uses techniques similar to, but not limited to, identifying and removing background noise data from the audio data to separate the audio data stream 112 into one or more substreams of background sound 122.

[0025] The audio type / feature determiner 136 uses a machine learning model or neural network to determine the acoustic type and features in the received audio data 112. The audio type or data type (114A and 122A) indicates the type of the corresponding substream of the audio data. The type 114A of the foreground sound 114 in the audio data 112 can be speech data. For example, the type 122A of the background sound 122 can be a siren from a police car or ambulance. In various respects, the audio type / feature determiner 136 can use a machine learning (ML) model or neural network to determine or predict the audio type and features. The ML model or neural network used to determine the audio type can use a trained audio type model 138. Acoustic features include pitch, tone, intensity, and volume. The foreground speech data can include values ​​for pitch 116A, tone 118A, and intensity 120A. The background siren sound can include values ​​for pitch 116B, tone 118B, and intensity 120B. The audio type / feature determiner 136 uses a trained machine learning (ML) model to determine the type of the corresponding substream of audio data 112. A storage device for the trained audio type model 138 stores the trained ML model used by the audio type / feature determiner 136. In each respect, the trained ML model determines the type of audio data based on a set of predetermined rules to analyze the substreams of received audio stream data.

[0026] In various respects, the audio type / feature determiner 136 can determine the audio type and acoustic features of the audio stream data based on analysis of the audio stream data without splitting the audio stream data into audio substreams. This makes the processing for determining the audio type and acoustic features less intensive than generating audio substreams and determining the audio type and acoustic features for each corresponding audio substream. There may be a trade-off in accuracy when determining audio type and acoustic features based on audio stream data, because analyzing the corresponding audio substreams can utilize more detailed analysis than using the audio stream data, enabling a more accurate determination of the audio type and acoustic features.

[0027] Task generator 140 generates a task based on the received audio input and a determined substream of audio data 112. In various aspects, task generator 140 generates a task using a foreground sound 114 having a data type 114A (e.g., speech in the foreground). For example, when the foreground sound 114 includes the data type 114A of speech and the speech produced by the speech is "calling mom," the task could be "calling mom."

[0028] Task classifier 142 classifies tasks into one or more predefined categories using acoustic features of substreams of received audio data 112. These predefined categories are associated with the importance and urgency of the tasks (e.g., "important," "not important," "urgent," "not urgent," etc.). Task classifier 142 can use a machine learning model to classify tasks. The machine learning model, which can be stored in a storage device for a trained model 126, can use the values ​​of the acoustic features of the audio data to determine the importance and urgency of a given task. Alternatively or additionally, task classifier 142 can use a neural network with a set of trained parameters to classify tasks. Storage for a trained task classification model 144 includes a set of trained parameters used by task classifier 142 to classify tasks. The neural network can use embeddings that at least represent the acoustic features of the received audio data 112. In each respect, an embedding represents a mapping from a set of multidimensional vectors or acoustic features to vectors of continuous numbers. Audio substream generator 134 can convert the audio signals of substreams of received audio data 112 into embeddings as input to the neural network in task classifier 142. Task classifier 142 can use a neural network to determine the likelihood of a classification by determining a probability distribution for the corresponding categories associated with the importance and urgency of the task. For example, when the probability distribution indicates that the category “important” is more likely than the category “unimportant,” task classifier 142 will classify the task as “important” rather than “unimportant.”

[0029] In various aspects, the task classifier 142 can use a regression data model to determine the importance, urgency, and priority of a category. Specifically, the regression data model refers to the acoustic characteristics of the audio substream over previous times. For example, a task with a background sound of an siren that becomes louder over time can be classified as important and / or urgent. The task classifier 142 can regressively analyze the acoustic characteristics of the siren background audio substream over the duration of the audio input. In various aspects, the pitch and volume of the audio substream change over time to determine the source of the siren (e.g., a fire truck) that is approaching.

[0030] Task processor 146 processes and executes generated tasks at a determined level of importance and urgency. In some aspects, task processor 146 interacts with a telephone application to make phone calls based on the task. Additionally or alternatively, task processor 146 may interface with external applications, including but not limited to calendar applications, task management applications (not shown), and telecommunications applications, and send tasks according to a determined category based on their importance and urgency. For example, when the task category indicates that the task is “urgent,” task processor 146 may initiate the execution of a task with urgency. In some other aspects, when the determined task category indicates that the task is “important,” task processor 146 may cause a calendar application to highlight or emphasize the task.

[0031] Model generator 160 generates a trained audio type model 138 and a trained task classification model 144, and deploys the trained data to a storage device in task manager 130 for the trained model 137. The audio type model trainer trains the audio type model using a training set of correct audio data and types. The training set may include, for example, a pair of audio stream data with voice and type "voice". Another pair of training data may be another pair of audio stream data, including an alarm and type "alarm". Audio type model trainer 162 may store the trained audio type data in a storage device for training audio type data 166.

[0032] The task classification model trainer 164 uses the received training data to train a task classification model. The training data for task classification may include a set of acoustic features of one or more audio data substreams and correct classifications for importance, urgency, and priority. The classification for priority may include more than two priority categories for ranking. The trained data for classification may include a set of rules or conditions (i.e., a trained model) for the values ​​of acoustic features and specified categories. Alternatively or additionally, the trained data for classification may include a set of trained parameters (or a trained model) for a neural network that receives the embeddings of the audio data as input and determines one or more categories of the audio data for importance, urgency, and priority.

[0033] By storing trained audio type data in a storage device for a trained audio type model 138 and storing trained data for task classification in a storage device for a trained task classification model 144, the model deployer 170 deploys the trained audio type data and the trained data for classification to the task manager 130.

[0034] Will understand, about Figure 1The various methods, devices, applications, features, etc., described are not intended to limit System 100 to being performed by the specific applications and features described. Therefore, additional configurations may be used to practice the methods and systems disclosed herein and / or the features and applications described may be excluded without departing from the methods and systems disclosed herein.

[0035] Figure 2 An example of a data structure for determining the importance and urgency of a task based on acoustic features of audio data according to aspects of this disclosure is shown. The example data structure 200 includes three types of data structures: an audio stream input 210, a set of audio types 214 and acoustic features 216, and a task 220 and its task category 222.

[0036] Audio stream input 210 represents audio input data from client device 110, received by audio receiver 132 of task manager 130. Audio stream input 210 includes multiple audio substreams (e.g., audio substream A 212A, audio substream B 212B, and audio substream C 212C, for example). Each audio substream includes a set of audio signals during the duration of audio stream input 210. For example, this set of audio signals may include amplitude values ​​for corresponding audio frequencies. In various respects, each audio substream represents an audio data stream of a different sound. For example, different sounds may be voice, fire truck siren, indistinguishable human voice, clapping, gunshots, etc. Sound processing techniques (e.g., noise cancellation, speaker recognition) can identify various types of audio and isolate them into different audio substreams.

[0037] Audio type 214 includes one or more corresponding audio substreams of different types. For audio substream A 212A, audio type 214 includes a foreground voice sound with the spoken text “Calling Mom.” The acoustic features 216 of audio substream A 212A include a tempo of 10, a beat of 10, a prosody of 10, and a volume of 10. In each respect, a larger value for each acoustic feature indicates a larger magnitude of that feature. Audio substream B 212B is converted to the audio type of an alarm in the background, with acoustic features 216 of a tempo of 3, a beat of 10, a step of 78, a prosody of 5, and a volume of 7. Audio substream C 212C is converted to the audio type of human speech in the background (indistinguishable), with acoustic features of a tempo of 10, a beat of 5, a step of 3, a prosody of 3, and a volume of 5.

[0038] Based on the audio type 214, a combination of audio types 214 across multiple substreams of audio stream input 210, the present invention generates, for example, a task 220 of “Call Mom,” which involves calling Mom as the recipient (or the destination of the call). The present invention also determines the importance of the task as a task category: “Important,” based on a set of acoustic features 216 across multiple audio substreams of audio stream input 210. Alternatively or additionally, the present invention determines a category associated with the urgency and priority of the task based on the acoustic features 216. For example, the task 220 of “Call Mom” can be classified as “Urgent” when one of the background sound substreams includes a fire truck siren. A task classifier (ML or NN) 142 can be trained to determine that calling Mom is urgent when the received audio input includes a fire truck siren in the background. In some other aspects, the task 220 of “Call Mom” can be “Non-Urgent” when the received audio stream input 210 does not include a siren from any of the audio substreams. In some other respects, in addition to acoustic features, classification can also take into account the context of task 220 (e.g., calling mom).

[0039] Figure 3 An example of a data structure is shown that determines the importance and urgency of a task by extracting context from speech data and audio data according to aspects of this disclosure. Additionally or alternatively, this disclosure extracts speech data 312A and sound context 312B from audio stream input 310. In various aspects, speech data 312A may be based on foreground sound from audio stream input 310. Sound context 312B may be based on acoustics from a background audio substream. The invention generates task 330 with command 332, such as “Call Mom” (e.g., initiating a phone call). Concurrently, this disclosure identifies task category 334 as important.

[0040] In all respects, this disclosure transforms the audio stream input 310 into a set of embeddings (i.e., multidimensional vectors). A task classifier (e.g., using a neural network with a trained model) 142 classifies the sound context into one or more categories indicating the importance, urgency, and priority of the task. Figure 3 In the example, task category 334 indicates "important". Alternatively, task category 334 may indicate other values, including "not important", "urgent", "not urgent", and priority level (e.g., values ​​ranging from 1 to 10 according to task priority).

[0041] Figure 4 An example of a data structure for determining the importance and urgency of a task according to aspects of this disclosure is shown. Figure 4 The values ​​for the corresponding categories of the audio stream input are shown: task category 410, task ranking 420, and category rule 430.

[0042] Task classification 410 may include categories based on importance 412 and urgency 414. The importance category 412 may be either "important" or "not important". The urgency category 414 may be either "urgent" or "not urgent".

[0043] Task ranking (420) can include a ranking score to determine task priority. Urgency score (422) takes a value between 1 and 5, where 5 represents the most urgent task. Importance score (424) takes a value between 1 and 5, where 5 represents the most important task.

[0044] Classification rule 430 includes a set of rules for classifying tasks. In various aspects, this disclosure determines classification based on a set of conditions. As an example, classification rule 430 instructs three acoustic features as a set of conditions to determine a category. Exemplary classification rule 430 instructs that the foreground pitch is greater than 7, the foreground volume is greater than 9, and the background audio type is "alarm". This rule specifies that a set of conditions is converted into categories indicating "urgent" and "important". In various aspects, classification rule 430 can be used to train a machine learning model for task classifier 142.

[0045] Figure 5 An example of a method, according to aspects of this disclosure, for using machine learning models to determine the importance and urgency of a task based on acoustic features of audio data associated with the task.

[0046] The general order of operations for method 500 is as follows: Figure 5 The method is shown in the diagram. Typically, method 500 begins with start operation 502 and ends with end operation 522. Method 500 may include more or fewer steps or may be combined with... Figure 5 The steps shown are arranged differently in order. Method 500 can be executed as a set of computer-executable instructions, which are executed by a computer system and encoded or stored on a computer-readable medium. Furthermore, method 500 can be executed by gates or circuits associated with a processor, ASIC, FPGA, SOC, or other hardware device. In the following, method 500 will be referred to in conjunction with… Figure 1 , 2 The systems, components, devices, modules, software, data structures, data feature representations, signaling diagrams, and methods described in sections 3, 4, 6, 7, and 8A to 8B are explained.

[0047] After initiating operation 502, method 500 begins with receiving operation 504, which receives audio training data comprising substreams of audio data and acoustic features with correct audio types and task categories. For example, the audio training data may include audio data with a person's voice shouting the phrase "Call Mom" ​​in the foreground audio and indistinguishable human voices from a fire truck siren in the background audio.

[0048] Training operation 506 uses the received audio training data to train a machine learning model. Specifically, the machine learning model for determining the audio type can be trained using example training data to train the sound of shouting "Call Mom." Another machine learning model can be trained to classify the task into importance, urgency, and priority. A set of acoustic features from the training data can be used to train the machine learning model to determine the category of the task "Call Mom" ​​as "important" when the acoustic features of the foreground speech include high volume and high pitch (i.e., "shout or scream"). Additionally, training operation 506 trains the machine learning model to classify the task "Call Mom" ​​as "urgent" when the audio input includes the sound of a fire truck siren in the background. In various aspects, steps 504 through 506 are used to train models for determining the audio type and the category of the task. In some aspects, training operation 506 uses training data to train a machine learning model for task classification, where the audio stream data includes foreground audio types with the word "stop," and the acoustic features of the foreground audio include high pitch, short rhythm, and high intensity. The training data can specify a set of categories: "important" and "urgent." Therefore, the use of a trained task classification model can cause the task executor to immediately stop the ongoing task (e.g., playing music) without requiring confirmation from the user.

[0049] The receiving operation 508 receives audio stream input. In various aspects, the audio receiver 132 of the task manager 130 can receive audio stream input. The audio stream input may include multiple audio substreams, each representing a different type of audio data (e.g., foreground voice, background alarm, etc.).

[0050] Generation operation 510 generates an audio substream based on the received audio stream input. Generation operation 510 can split the received audio stream input into substreams by identifying different types of audio data and splitting the audio stream into a set of substreams. In various aspects, generation operation 510 uses techniques similar to those used in noise cancellation to identify noise or sound in the audio stream to identify and generate the substreams of the audio stream. In some other aspects, the client device may include multiple audio input devices (e.g., directional microphones) to receive audio data and identify different sound sources based on the direction of the audio input devices. Based on the audio sources identified by the directional audio inputs, the audio stream data can be separated into substreams.

[0051] In all aspects, generation operation 510 splits the received audio stream input into segments of fixed time length (e.g., 10 seconds). Log-mel filter bank (LMFB) features can be used as audio features from the received audio stream input. The input audio waveform can be analyzed using a sparse fast Fourier transform (SFFT) with a fixed window size and frameshift (e.g., 2048 SFFT points, a window size of 2048 samples, and a frameshift of 1024 samples). In all aspects, the use of some software packages or libraries for audio analysis (e.g., Librosa, see B. McFee, C. Raffel, D. Liang, DPelis, M. McVicar, E. Battenberg, and O. Nieto, “librosa: Audio and musicsignal analysis in python,” Proceedings of the 14th Python in Science Conference, Vol. 8, 2015) can determine the LMFB features. The Hidden Markov Toolkit (HTK) formula definition for the Mel scale can be adopted to determine the pitch of the audio stream input. The sampling rate can be dependent on model size constraints. For example, there could be two tasks classifying urgency and importance. The final spectrogram could have 431 time bins, with 128 frequency bins for each task. Unpadded Log-mel delta and delta-deltas can also be computed, reducing the number of time samples to 423. This results in a final tensor size of 423 × 128 × 3. In all respects, the generation operation 510 can normalize or scale each feature value to a value between zero and one before feeding the speech feature tensor to the convolutional neural network (CNN) classifier.

[0052] The input tensors are fed into different classifiers using a fully convolutional neural network (FCNN). For example, an FCNN could include 910 stacked convolutional layers with small kernel sizes (e.g., the VGG architecture, see K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition”, arXiv preprint arXiv:1409.1556, 2014). Each convolutional layer is followed by batch normalization and a rectified linear activation function (e.g., a rectified linear unit (ReLU) activation layer). Random dropout is used in convolutional layers 5 through 10 to mitigate potential overfitting issues, with 2×2 max pooling layers attached to the fourth and eighth ReLU layers. Additionally, attention is applied to each output channel of the last convolutional layer. A global pooling layer, followed by a softmax layer, generates the final classification decision for each task.

[0053] For example, an audio substream could represent foreground speech with a voice input of “calling Mom.” Another audio substream could represent background sounds from a fire truck siren. Yet another audio substream could represent background sounds of people talking, but words spoken by one person in the background are indistinguishable from words spoken by another person in the background.

[0054] Operation 512 determines the acoustic features of the corresponding audio substream. These acoustic features may include, but are not limited to, the rhythm, tempo, prosody, step, and volume of the corresponding audio substream. In some aspects, audio with a speech voice type may include the pitch and speed of the speech as acoustic features to determine whether the speaker is shouting, screaming, or talking casually.

[0055] Operation 514 determines the audio type of the corresponding audio substream. The audio type can include whether the audio substream is foreground or background audio. Audio types can also include voice, sirens from fire trucks or ambulances, indistinguishable voices of people in the background, sounds of objects hitting objects, etc. An audio stream input can include more than one audio substream. Each audio substream can be associated with an audio type different from the others. Operation 514 can use a machine learning model to determine the audio type. The machine learning model can have already been trained in operation 506.

[0056] Generation operation 516 generates a task based on audio input. For example, when the audio input includes foreground audio of a spoken voice and the voice text is "Call Mom," generation operation 516 can generate a task for initiating a phone call to Mom using a pre-defined phone book. In some other aspects, generation operation 516 can generate a task based on background audio of the audio input. For example, a digital assistant with a client device is operating in an vacant house. However, the audio input may include background audio of a high-volume siren from a fire truck and the sound of someone knocking on the house door. Generation operation 516 can automatically generate a task to call the owner of the vacant house to notify the owner of a possible emergency at the vacant house. Steps 508 to 516 correspond to generating task 532. Alternatively or additionally, generation operation 516 can receive an existing task to modify its category. Generation operation 516 can interact with external applications that manage tasks (e.g., calendar applications, to-do list applications, and appointment scheduler applications).

[0057] Classification operation 518 categorizes tasks based on their importance, urgency, and priority according to acoustic characteristics. In each respect, classification operation 518 can determine the category of the task. Classification operation 518 can use a learned machine learning model to classify tasks based on acoustic features. For example, the task of calling mom could include a "call mom" voice with high volume and high pitch (e.g., screaming) acoustic features in the foreground sound, and background sounds from a fire truck siren based on acoustic features. Classification operation 518 classifies the task as "important" and "urgent" based on this combination of acoustic features. In each respect, the combination of acoustic features of the foreground voice "call 911" with background noise indicating alarms, sirens, explosions, etc., would be more important than the same foreground voice without background noise. Classification operation 518 classifies the task accordingly by analyzing both the acoustic characteristics of the foreground audio (e.g., voice) data and the background audio data. In some other respects, when the foreground speech sound is a “stop that” speech sound with acoustic features indicating a high level of stress on the speaker (e.g., high pitch, unstable tone, velocity above a predetermined threshold, etc.), classification operation 518 classifies the task of stopping the music playing (e.g., on a smart speaker) as “urgent.” Thus, classification operation 518 can classify, for example, speech sound with certain acoustic features as a shout under stress, and therefore the task is urgent. The level of stress can be a speech sound or an attribute of the speech data determined based on the acoustic features of the speech data.

[0058] In some respects, there could be a classifier for urgency (i.e., class C). urg Task (where the urgency category result is F) urg Another classifier can classify importance (i.e., C). impThe task importance category (Class) result label is F. imp The results of the two classifiers can be combined using the following equation (1):

[0059]

[0060] In various aspects, training the split audio stream input can be based on recommended training data with correct answers. For example, there could be 14,000 audio clips for training the model, labeled with category tags for "urgency" and "importance." There could be 2,500 test audio clips. Data augmentation techniques can be used to improve the model. Stochastic gradient descent (SGD) with a cosine decaying restart learning rate scheduler can be used to train the model. The maximum and minimum learning rates can be 0.1 and 1e-5, respectively. In the absence of validation data, the average output of the model when the learning rate is near its minimum can be used. For example, Pyrotch or Keras can be used to implement CNN-based models for classification. After training is complete, the network can be used for inference by applying the network to small audio segments (e.g., 200 milliseconds).

[0061] In some other respects, there may be an audio stream with no foreground sound but only a set of background sounds. Generation operation 517 generates a task to call the homeowner. This set of background sounds may include the siren of a fire truck and another background sound of someone knocking on the door. Classification operation 518 may classify the task as "important" and "urgent." In some respects, classification operation 518 may determine the priority for performing the task based on the acoustic characteristics of the audio stream.

[0062] Action 520 performs tasks based on their categorized importance and / or urgency. Action 520 can interact with external servers and devices to perform tasks based on the levels of importance, urgency, and priority categorized by action 518. For example, action 520 can interact with an external telephone server to initiate a phone call at a time reflecting the level of importance and the level of urgency determined by the caller. Action 520 performs tasks that are both important and urgent, starting as early as possible before performing other tasks. In some other respects, action 520 can interact with an email server to send emails. In yet another respect, action 520 can save tasks in a calendar application or task management application.

[0063] Figure 6 Examples of methods for determining the importance and urgency of a task based on acoustic features of audio data associated with the task, according to various aspects of this disclosure, include embedding audio data and using neural networks.

[0064] Figure 6The general sequence of operations for method 600 is shown. Typically, method 600 begins with start operation 602 and ends with end operation 620. Method 600 may include more or fewer steps or may be combined with... Figure 6 The steps shown are arranged differently in different order. Method 600 can be executed as a set of computer-executable instructions that are executed by a computer system and encoded or stored on a computer-readable medium. Furthermore, method 600 can be executed by gates or circuits associated with a processor, ASIC, FPGA, SOC, or other hardware device. In the following, method 600 will be referred to in conjunction with... Figure 1 , 2 The systems, components, devices, modules, software, data structures, data characteristic representations, signaling diagrams, and methods described in sections 3, 4, 5, 7, and 8A-B are explained.

[0065] After initiating operation 602, method 600 begins with receiving operation 604, which receives training data. The training data includes a substream of audio data and acoustic features with the correct audio type and task category. For example, the audio training data may include audio data of a person's voice shouting the phrase "Call Mom" ​​in the foreground audio and indistinguishable human voices from a fire truck siren in the background audio.

[0066] Training operation 606 uses the received audio training data to train a task classifier neural network. Specifically, the neural network used to determine the audio type can be trained using example training data to train the sound of shouting "Call Mommy". Another neural network can be trained to classify the task as important, urgent, and priority. The set of acoustic features of the training data can be used to train a machine learning model to determine the category of the task "Call Mommy" when the acoustic features of the foreground voice include a high pitch and a large volume (i.e., "shout or scream"). Additionally, training operation 606 trains the neural network to classify the task "Call Mommy" as urgent when the audio input includes the sound of a fire truck siren in the background. In all respects, steps 604 to 606 are used to train the model for determining the audio type and classifying the task.

[0067] The receiving operation 608 receives an audio stream input. In various respects, the audio receiver 132 of the task manager 130 can receive audio stream input. The audio stream input may include multiple audio substreams, each representing a different type of audio data (e.g., foreground voice, background alarm, etc.).

[0068] The generation operation 610 generates an audio substream based on the received audio stream input. The generation operation 610 can divide the received audio stream input into substreams by identifying different types of audio data and splitting the audio stream into a set of substreams. For example, an audio substream could represent foreground speech with a voice input of "calling mom." Another audio substream could represent background sounds from a fire truck siren. Yet another audio substream could represent background sounds of people talking, but words spoken by one person in the background are indistinguishable from words spoken by another person in the background.

[0069] The generation operation 612 generates embeddings of the corresponding audio substreams as input to a classifier (i.e., a neural network). In all respects, the embeddings represent multidimensional vectors. Each dimension of the vector is associated with acoustic features and other parameters characterizing the audio stream and its audio substreams.

[0070] Generation operation 614 generates a task based on audio input. For example, when the audio input includes foreground audio of a spoken voice and the voice text is “Call Mom,” generation operation 614 can generate a task for initiating a phone call to Mom using a pre-defined phone book. In some other aspects, generation operation 614 can generate a task based on background audio of the audio input. For example, a digital assistant with a client device is placed in an vacant house. However, the audio input may include background audio of a high-volume siren from a fire truck and the sound of someone knocking on the house door. Generation operation 614 can automatically generate a task to call the owner of the vacant house to notify the owner of a possible emergency at the vacant house. Steps 608 to 616 correspond to generating task 632. Alternatively or alternatively, generation operation 614 can receive an existing task to modify its category. Generation operation 614 can interact with external applications that manage tasks (e.g., calendar applications, to-do list applications, and appointment scheduler applications).

[0071] The classification operation 616 categorizes tasks based on their acoustic importance, urgency, and priority. In various aspects, the classification operation 616 can determine the category of the task. The classification operation 616 can use a neural network (e.g., a multi-layered recurrent neural network) that uses embeddings as input to determine the task category based on acoustic features. For example, a task to call mom might include the voiced "call mom" in the foreground sounds, high volume and high pitch (e.g., screaming) acoustic features, and background sounds from a fire truck siren based on acoustic features. The classification operation 616 categorizes the task as "important" and "urgent" based on a combination of acoustic features. The neural network can use a trained set of parameters trained by the training operation 606.

[0072] Execution operation 618 performs a task according to a defined category. Execution operation 618 can interact with external servers and devices to perform tasks based on the level of importance, urgency, and priority categorized by classification operation 616. For example, execution operation 618 can interact with an external telephone server to initiate a phone call at a time reflecting the level of importance and the level of urgency determined. Execution operation 618 performs a task that is both important and urgent, allowing it to begin as early as possible before other tasks. In some other aspects, execution operation 618 can interact with an email server to send an email. In yet another aspect, execution operation 618 can save the task in a calendar application or task management application. In each aspect, method 600 can end with termination operation 620. In each aspect, execution operation 618 can update the position of the task on the to-do list in a to-do list application based on the category. The position of the task can indicate the ranking level of the task's importance and / or urgency on the to-do list.

[0073] It should be understood that operations 602 to 620 are described for the purpose of illustrating the method and system and are not intended to limit the disclosure to a particular sequence of steps. For example, the steps may be performed in a different order, additional steps may be performed, and the disclosed steps may be excluded without departing from the disclosure.

[0074] Figure 7 This is a block diagram illustrating the physical components (e.g., hardware) of a computing device 700 in which aspects of this disclosure can be implemented. The computing device components described below are applicable to the aforementioned computing device. In a basic configuration, computing device 700 may include at least one processing unit 702 and system memory 704. Depending on the configuration and type of the computing device, system memory 704 may include, but is not limited to, volatile storage devices (e.g., random access memory), non-volatile storage devices (e.g., read-only memory), flash memory, or any combination of such memories. System memory 704 may include an operating system 705 and one or more program tools 706 suitable for executing the aspects disclosed herein. For example, operating system 705 may be suitable for controlling the operation of computing device 700. Furthermore, aspects of this disclosure may be implemented in conjunction with graphics libraries, other operating systems, or any other applications, and are not limited to any particular application or system. This basic configuration in Figure 7 The components within the dashed line 708 are shown. The computing device 700 may have additional features or functions. For example, the computing device 700 may also include additional data storage devices (removable and / or non-removable), such as, for example, a disk, optical disc, or magnetic tape. Such additional storage... Figure 7 The image shows a removable storage device 709 and a non-removable storage device 710.

[0075] As described above, multiple program tools and data files can be stored in system memory 704. While executing on at least one processing unit 702, program tool 706 (e.g., application 720) can perform processes including, but not limited to, those described herein. Application 720 includes an audio receiver 722, an acoustic type / feature determiner 724, a task generator 726, a task classifier 728, and a task processor 730, as described above. Figure 1 More detailed descriptions are available. Other program tools that may be used under various aspects of this disclosure may include email and contact applications, task management applications, calendar applications, word processing applications, spreadsheet applications, database applications, presentation applications, drawing or computer-aided applications, etc.

[0076] Furthermore, aspects of this disclosure can be implemented in discrete electronic components, packaged or integrated electronic chips containing logic gates, circuits utilizing microprocessors, or circuits on a single chip containing electronic components or microprocessors. For example, aspects of this disclosure can be implemented via a system-on-a-chip (SOC), wherein... Figure 7 Each or many of the components shown can be integrated onto a single integrated circuit. Such a SoC device may include one or more processing units, graphics units, communication units, system virtualization units, and various application functions, all integrated (or “burned in”) onto a chip substrate as a single integrated circuit. When operating via the SoC, the capabilities described herein regarding the client switching protocol can be operated via dedicated logic integrated onto a single integrated circuit (chip) along with other components of the computing device 700. Aspects of this disclosure can also be practiced using other techniques capable of performing logical operations (such as, for example, AND, OR, and NOT), including but not limited to mechanical, optical, fluid, and quantum technologies. Furthermore, aspects of this disclosure can be practiced in general-purpose computers or any other circuit or system.

[0077] The computing device 700 may also have one or more input devices 712, such as a keyboard, mouse, pen, voice or audio input device, touch or swipe input device, etc. Multiple output devices 714 (such as a monitor, speaker, printer, etc.) may also be included. The above devices are examples and other devices may be used. The computing device 700 may include one or more communication connections 716 that allow communication with other computing devices 750. Examples of suitable communication connections 716 include, but are not limited to, radio frequency (RF) transmitters, receivers, and / or transceiver circuitry; universal serial buses (USB), parallel and / or serial ports.

[0078] As used herein, the term computer-readable medium may include computer storage media. Computer storage media may include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information, such as computer-readable instructions, data structures, or program tools. System memory 704, removable storage device 709, and non-removable storage device 710 are examples of computer storage media (e.g., memory storage devices). Computer storage media may include RAM, ROM, electrically erasable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical storage devices, magnetic tape, magnetic tape, disk storage or other magnetic storage devices, or any other article of manufacture that can be used to store information and is accessible by computing device 700. Any such computer storage medium may be part of computing device 700. Computer storage media does not include carrier waves or other propagated or modulated data signals.

[0079] Communication media can embody computer-readable instructions, data structures, program tools, or other data in modulated data signals (such as carrier waves or other transmission mechanisms), and include any information transmission medium. The term "modulated data signal" can describe a signal having one or more sets of characteristics that are set or altered in a manner that encodes information in the signal. By way of example and not limitation, communication media can include wired media (such as wired networks or direct wired connections) and wireless media (such as acoustic, radio frequency (RF), infrared, and other wireless media).

[0080] Figure 8A and Figure 8B The aspects of this disclosure are illustrated using computing devices or mobile computing devices 800 from which they can be implemented, such as mobile phones, smartphones, wearable computers (such as smartwatches), tablet computers, laptop computers, etc. In some aspects, the client utilized by the user may be a mobile computing device. Reference Figure 8AOne aspect of a mobile computing device 800 used to implement these aspects is shown. In a basic configuration, the mobile computing device 800 is a handheld computer with both input and output elements. The mobile computing device 800 typically includes a display 805 and one or more input buttons 810 that allow users to enter information into the mobile computing device 800. The display 805 of the mobile computing device 800 can also be used as an input device (e.g., a touchscreen display). If included as an optional input element, a side input element 815 allows additional user input. The side input element 815 can be a rotary switch, a button, or any other type of manual input element. Alternatively, the mobile computing device 800 can incorporate more or fewer input elements. For example, the display 805 may not be a touchscreen in some respects. In yet another alternative, the mobile computing device 800 is a portable telephone system, such as a cellular phone. The mobile computing device 800 may also include an optional keypad 835. The optional keypad 835 can be a physical keypad or a “soft” keypad generated on a touchscreen display. In various aspects, output elements include a display 805 for displaying a graphical user interface (GUI), a visual indicator 820 (e.g., a light-emitting diode), and / or an audio transducer 825 (e.g., a speaker). In some aspects, the mobile computing device 800 incorporates a vibration transducer for providing haptic feedback to the user. In another aspect, the mobile computing device 800 incorporates input and / or output ports (such as audio input (e.g., a microphone jack), audio output (e.g., a headphone jack), and video output (e.g., an HDMI port)) for sending signals to or receiving signals from external devices.

[0081] Figure 8B This indicates computing devices, servers (e.g., Figure 1 The diagram illustrates the architecture of one aspect of a mobile computing device, such as a task manager 130 and a model generator 160. Specifically, the mobile computing device 800 can be incorporated into system 802 (e.g., a system architecture) to implement certain aspects. System 802 can be implemented as a "smartphone" capable of running one or more applications (e.g., a browser, email, calendar, contact manager, messaging client, game, and media client / player). In some aspects, system 802 is integrated as a computing device, such as an integrated digital assistant (PDA) and a wireless phone.

[0082] One or more applications 866 may be loaded into memory 862 and run on or associated with operating system 864. Examples of applications include telephone dialers, email programs, PIM (Personal Information Management) programs, word processing programs, spreadsheet programs, internet browser programs, messaging programs, etc. System 802 also includes a non-volatile storage area 868 within memory 862. The non-volatile storage area 868 can be used to store persistent information that should not be lost in the event of a power outage of system 802. Applications 866 may use and store information, such as emails or other messages used by email applications, in the non-volatile storage area 868. A synchronization application (not shown) also resides on system 802 and is programmed to interact with a corresponding synchronization application residing on a host computer to keep the information stored in the non-volatile storage area 868 synchronized with the corresponding information stored on the host computer. It should be understood that other applications may be loaded into memory 862 and run on the mobile computing device 800 described herein.

[0083] System 802 has a power supply 870, which can be implemented as one or more batteries. The power supply 870 may also include an external power source, such as an AC adapter or a power docking station for replenishing or recharging the batteries.

[0084] System 802 may also include a radio interface layer 872, which performs the functions of transmitting and receiving radio frequency communications. Radio interface layer 872 facilitates wireless connections between system 802 and the "external world" via a communications operator or service provider. Transmissions to and from radio interface layer 872 are conducted under the control of operating system 864. In other words, communications received by radio interface layer 872 can be propagated to application program 866 via operating system 864, and vice versa.

[0085] A visual indicator 820 (e.g., an LED) can be used to provide visual notifications, and / or an audio interface 874 can be used to generate audible notifications via an audio transducer 825. In the illustrated configuration, the visual indicator 820 is a light-emitting diode (LED) and the audio transducer 825 is a speaker. These devices can be directly coupled to a power supply 870 such that when activated, they remain on for a duration specified by the notification mechanism, even if the processor 860 and other components may be turned off to conserve battery power. The LED can be programmed to remain on indefinitely until the user takes an action to indicate the device's power-on status. The audio interface 874 is used to provide and receive audible signals from the user. For example, in addition to being coupled to the audio transducer 825, the audio interface 874 can also be coupled to a microphone to receive audible input, such as to facilitate telephone conversations. According to various aspects of this disclosure, the microphone can also be used as an audio sensor to facilitate control of notifications, as described below. System 802 may also include a video interface 876, which enables the operation of the onboard camera 830 to record still images, video streams, etc.

[0086] The mobile computing device 800 implementing system 802 may have additional features or functions. For example, the mobile computing device 800 may also include additional data storage devices (removable and / or non-removable), such as disks, optical discs, or magnetic tapes. Such additional storage... Figure 8B The non-volatile storage region 868 is shown in the middle.

[0087] Data / information generated or captured by mobile computing device 800 and stored via system 802 can be locally stored on mobile computing device 800, as described above, or the data can be stored on any number of storage media that can be accessed by the device via radio interface layer 872 or via a wired connection between mobile computing device 800 and a separate computing device associated with mobile computing device 800, such as a server computer in a distributed computing network (such as the Internet). It should be understood that such data / information can be accessed via radio interface layer 872 or via a distributed computing network through mobile computing device 800. Similarly, such data / information can be easily transferred between computing devices for storage and use according to well-known data / information transmission and storage components, including email and collaborative data / information sharing systems.

[0088] The descriptions and illustrations of one or more aspects provided in this application are not intended to limit or restrict the scope of the claimed disclosure in any way. The aspects, examples, and details provided in this application are considered sufficient to convey ownership and enable others to make and use the best model of the claimed disclosure. The claimed disclosure should not be construed as limited to any aspect, such as, or the details provided in this application. Whether shown and described in combination or separately, various features (both structural and methodological) are intended to be selectively included or omitted to produce embodiments with a particular set of features. Having provided the descriptions and illustrations of this application, those skilled in the art can contemplate variations, modifications, and alternatives falling within the spirit of the broader aspects of the overall inventive concept embodied in this application without departing from the broader scope of the claimed disclosure.

[0089] As will be understood from the foregoing disclosure, one aspect of this technology relates to a computer-implemented method for determining the category of a task based on acoustic features of audio data associated with the task. The method includes receiving audio input, wherein the audio input is associated with a task; determining one or more acoustic features of the received audio input; determining the category of the task based on the one or more acoustic features, wherein the category of the task at least indicates importance or urgency; and performing the task according to the category of the task. The method also includes generating a plurality of audio substreams based on the received audio input; determining one or more acoustic features corresponding to one or more of the generated plurality of audio substreams; determining an audio type based on the plurality of audio substreams; generating a task based on the audio type; and determining the category of the task based on the one or more acoustic features using a trained machine learning model for classifying the task, wherein the category of the task is associated with the importance and urgency of the task. The method further includes receiving a set of audio type data for training, wherein the set of audio type data includes the correct audio types associated with the set of audio data; and training a machine learning model for determining the audio type of the audio input based on the received set of audio type data for training. The method further includes receiving a set of task classification data for training, wherein the set of task classification data for training includes at least one correct category of a task associated with a set of acoustic features, wherein the task category indicates one or more of importance, urgency, and priority, and wherein the set of acoustic features includes at least pitch, prosody, and intensity of the audio data; and training a machine learning model for determining categories based on acoustic features based on the received set of task classification data for training, wherein the categories include one or more of the following: important, unimportant, urgent, not urgent, and level of priority. The method also includes generating embeddings of audio data based on multiple generated audio substreams, wherein the embeddings are associated with multidimensional vector representations of one or more acoustic features of the multiple generated audio substreams, and wherein the trained machine learning model for classifying the task uses a neural network. The trained machine learning model is based on a regression model used to regressively classify the task based on the acoustic features of at least one audio substream of the multiple audio substreams over time. The one or more acoustic features include one or more of the following: volume, pitch, tone, or intensity. Tasks can be categorized into one or more of the following: important, unimportant, urgent, or not urgent.

[0090] Another aspect of this technology relates to a system. The system includes a processor and a memory storing computer-executable instructions that, when executed by the processor, cause the system to: receive audio input; receive a task, wherein the task is associated with the received audio input; determine one or more acoustic features of the received audio input; determine a category of the task based on the one or more acoustic features, wherein the category of the task at least indicates importance or urgency; and perform the task according to the category of the task. The computer-executable instructions also cause the system to: generate multiple audio substreams based on the received audio input; determine one or more acoustic features corresponding to one or more of the generated multiple audio substreams; determine an audio type of the received audio input based on the multiple audio substreams; generate a task based on the audio type; and determine a category of the task based on the one or more acoustic features using a trained machine learning model for classifying the task, wherein the category of the task is associated with the importance and urgency of the task. The computer-executable instructions further enable the system to: receive a set of audio type data for training, wherein the set of audio type data includes correct audio types associated with the set of audio data; and train a machine learning model to determine the audio type of the audio input based on the received set of audio type data for training. The computer-executable instructions further enable the system to: receive a set of task classification data for training, wherein the set of task classification data for training includes at least one correct category of the task associated with a set of acoustic features, wherein the category indicates one or more of the following: importance, urgency, and priority, and wherein the set of acoustic features includes at least pitch, prosody, and intensity of the audio data; and train a machine learning model for determining categories based on acoustic features based on the received set of task classification data for training, wherein the categories include levels of importance, unimportance, urgency, non-urgency, and priority. The computer-executable instructions further enable the system to: generate embeddings of audio data based on multiple generated audio substreams, wherein the embeddings are associated with multidimensional vector representations of one or more acoustic features of the received audio substreams, and wherein the trained machine learning model for the classification task is based on a neural network. The trained machine learning model is based on a regression model, which is used to classify tasks based on acoustic features that change over time.

[0091] In another aspect of this technology, a computer-readable recording medium is disclosed. This computer-readable recording medium stores computer-executable instructions that, when executed by a processor, cause a computer system to: receive audio input; receive a task, wherein the task is associated with the received audio input; determine one or more acoustic features of the received audio input; determine a category of the task based on the one or more acoustic features, wherein the category of the task at least indicates importance or urgency; and update the position of the task in a to-do list according to the task category, wherein the to-do list comprises multiple tasks ordered at least based on the importance or urgency of corresponding tasks among a plurality of tasks. The computer-executable instructions, when executed by a processor, also cause the computer system to: terminate the execution of an ongoing task based on the determined task category without interactively confirming the task's command. The computer-executable instructions, when executed by a processor, also cause the computer system to: send the task and the determined task category to a calendar application server at a time according to the determined task category, causing the calendar application to insert the task as a new task item according to the level of importance and urgency specified by the determined task category. When executed by the processor, the computer-executable instructions also cause the computer system to: initiate remote communication to a destination specified by the task, based on the level of importance and urgency indicated by the determined task category; send tasks with task categories, causing the receiving application to interactively display tasks with emphasis according to the task categories; and execute the task before another task, based on the task category.

[0092] Any one of the above aspects combined with any other aspect of the above aspects. Any one of the one or more aspects described herein.

Claims

1. A computer-implemented method for determining the category of a task based on acoustic features of audio data associated with the task, the method comprising: Receive audio input data, wherein the audio input data is associated with a task, and wherein the audio input data includes foreground audio and background audio; Receive the task, wherein the task is associated with the received audio input data; The processor determines multiple acoustic features of the received audio input data based on a first trained machine learning model, the first trained machine learning model predicting acoustic features from the audio signal data by embedding the audio signal data as input in vector form, wherein the multiple acoustic features include a first acoustic feature associated with the foreground audio and a second acoustic feature associated with the background audio. The processor determines the category of the task based on a combination of the first acoustic feature and the second acoustic feature, the first acoustic feature and the second acoustic feature being input to a second trained machine learning model, wherein the category of the task at least indicates urgency, and wherein the second trained machine learning model at least predicts the importance or urgency of the task as output, the importance or urgency being inferred at least from the background audio. as well as The processor automatically updates the position of the task in the to-do list according to the category of the task, wherein the to-do list includes multiple tasks, the multiple tasks being sorted at least based on the importance or urgency of the respective tasks among the multiple tasks, wherein the category specifies the level of urgency for performing the task.

2. The computer-implemented method according to claim 1, further comprising: Multiple audio sub-streams are generated based on the received audio input data; Determine one or more acoustic features corresponding to one or more audio sub-streams among the generated plurality of audio sub-streams; The audio type is determined based on the multiple audio sub-streams; The task is generated based on the audio type; as well as Based on the one or more acoustic features, a trained machine learning model for classifying tasks is used to determine the category of the task, wherein the category of the task is associated with the importance and urgency of the task.

3. The computer-implemented method according to claim 1, further comprising: Receive a set of audio type data for training, wherein the set of audio type data includes the correct audio type associated with the set of audio data; as well as Based on the received set of audio type data used for training, a machine learning model is trained to determine the audio type of the audio input data.

4. The computer-implemented method according to claim 1, further comprising: Receive a set of task classification data for training, wherein the set of task classification data for training includes at least one correct category of the task associated with a set of acoustic features, wherein the category of the task indicates one or more of importance, urgency and priority, and wherein the set of acoustic features includes at least the pitch, rhythm and intensity of the audio input data. as well as Based on the received set of task classification data for training, a machine learning model is trained to determine the category based on the plurality of acoustic features, wherein the category includes one or more of the following: important, unimportant, urgent, not urgent, and priority level.

5. The computer-implemented method according to claim 2, further comprising: An embedding of the audio input data is generated based on the generated plurality of audio substreams, wherein the embedding is associated with a multidimensional vector representation of one or more acoustic features of the generated plurality of audio substreams, and wherein the trained machine learning model for the classification task uses a neural network, and wherein the trained machine learning model is based on a regression model for regressively classifying the task based on the acoustic features of at least one of the plurality of audio substreams over time.

6. The computer-implemented method of claim 2, wherein the plurality of audio substreams include foreground sound and background sound, and the category of the one or more acoustic features having higher values ​​in the background sound indicates higher importance compared to another task having lower values ​​in the background sound.

7. The computer-implemented method of claim 6, wherein the one or more acoustic features include one or more of the following: volume, pitch, pitch, or strength, The received audio input data includes speech data, which has a pressure level as an attribute of the speech data, and The pressure level is associated with one or more of the acoustic features.

8. The computer-implemented method of claim 6, wherein the category of the task includes one or more of the following: important, unimportant, urgent, and not urgent.

9. A system for determining the category of a task using acoustic features of audio data associated with the task, the system comprising: processor; as well as The memory stores computer-executable instructions that, when executed by the processor, cause the system to perform operations including: Receive audio input data, wherein the audio input data includes foreground audio and background audio; Receive the task, wherein the task is associated with the received audio input data; The processor determines multiple acoustic features of the received audio input data based on a first trained machine learning model, the first trained machine learning model predicting acoustic features from the audio signal data by embedding the audio signal data as input in vector form, wherein the multiple acoustic features include a first acoustic feature associated with the foreground audio and a second acoustic feature associated with the background audio. The processor determines the category of a task based on a combination of the first acoustic feature and the second acoustic feature, the first acoustic feature and the second acoustic feature being used as input to a second trained machine learning model, wherein the category of the task at least indicates importance or urgency, and wherein the second trained machine learning model at least predicts the urgency of the task as output, the urgency being inferred at least from the background audio. as well as The processor automatically updates the position of the task in the to-do list according to the category of the task, wherein the to-do list includes multiple tasks, the multiple tasks being sorted at least based on the importance or urgency of the respective tasks among the multiple tasks, wherein the category specifies the level of urgency for performing the task.

10. The system of claim 9, wherein the computer-executable instructions further cause the system to: Multiple audio sub-streams are generated based on the received audio input data; Determine one or more acoustic features corresponding to one or more audio sub-streams among the generated plurality of audio sub-streams; The audio type of the received audio input data is determined based on the multiple audio sub-streams; The task is generated based on the audio type; as well as Based on the one or more acoustic features, a trained machine learning model for classifying tasks is used to determine the category of the task, wherein the category of the task is associated with the importance and urgency of the task.

11. The system of claim 9, wherein the computer-executable instructions further cause the system to: Receive a set of audio type data for training, wherein the set of audio type data includes the correct audio type associated with the set of audio data; and Based on the received set of audio type data used for training, a machine learning model is trained to determine the audio type of the audio input data.

12. The system of claim 9, wherein the computer-executable instructions further cause the system to: Receive a set of task classification data for training, wherein the set of task classification data for training includes at least one correct category of the task associated with a set of acoustic features, wherein the category indicates one or more of importance, urgency, and priority, and wherein the set of acoustic features includes at least the pitch, prosody, and intensity of the audio data; and Based on the received set of task classification data for training, a machine learning model is trained to determine the categories based on the plurality of acoustic features, wherein the categories include important, unimportant, urgent, non-urgent, and priority level.

13. The system of claim 10, wherein the computer-executable instructions further cause the system to: An embedding of audio data is generated based on the generated plurality of audio substreams, wherein the embedding is associated with a multidimensional vector representation of one or more acoustic features of the received audio substreams, and wherein the trained machine learning model for the classification task is based on a neural network.

14. The system of claim 10, wherein the trained machine learning model is based on a regression model used to regressively classify a task based on acoustic features that change over time.

15. A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, cause a computer system to: Receive audio input data, wherein the audio input data includes foreground audio and background audio; Receive a task, wherein the task is associated with the received audio input data; The processor determines one or more acoustic features of the received audio input data based on a first trained machine learning model, the first trained machine learning model predicting acoustic features from the audio signal data by embedding the audio signal data as input into a vector form, wherein the plurality of acoustic features includes a first acoustic feature associated with the foreground audio and a second acoustic feature associated with the background audio. The processor determines the category of a task based on a combination of the first acoustic feature and the second acoustic feature, the first acoustic feature and the second acoustic feature being input to a second trained machine learning model, wherein the category of the task at least indicates importance or urgency, and wherein the second trained machine learning model at least predicts the importance or urgency of the task as output, the importance or urgency being inferred at least from the background audio. as well as The processor automatically updates the position of the task in the to-do list according to the category of the task, wherein the to-do list includes multiple tasks, the multiple tasks being sorted at least based on the importance or urgency of the respective tasks among the multiple tasks, wherein the category specifies the level of urgency for performing the task.

16. The computer-readable storage medium of claim 15, wherein the computer-executable instructions, when executed by the processor, further cause the computer system to: Based on the determined category of the task, the execution of the ongoing task is terminated without interactive confirmation of the task's command.

17. The computer-readable storage medium of claim 15, wherein the computer-executable instructions, when executed by the processor, further cause the computer system to: At the time specified according to the determined category of the task, the task and the determined category of the task are sent to a calendar application, causing the calendar application to insert the task as a new task item based on the level of importance and urgency specified by the determined category of the task.

18. The computer-readable storage medium of claim 15, wherein the computer-executable instructions, when executed by the processor, further cause the computer system to: Based on the level of importance and urgency specified by the category of the determined task, initiate remote communication to the destination specified by the task.

19. The computer-readable storage medium of claim 15, wherein the computer-executable instructions, when executed by the processor, further cause the computer system to: Send the task with the category of the task, causing the receiving application to interactively display the task with emphasis according to the category of the task.

20. The computer-readable storage medium of claim 15, wherein the computer-executable instructions, when executed by the processor, further cause the computer system to: Based on the category of the task, the task is executed before another task.

Citation Information

Patent Citations

  • Categorizationing and prioritization of managing tasks

    CN108475365A

  • Call priority based on audio stream analysis

    US20080310398A1

  • Improving natural language interactions using emotional modulation

    US20160210985A1

  • Methods and systems for voice profiling as a service

    US20210020191A1