Methods and systems for generating actionable insights for multi-speaker audio conversation

By processing both verbal and non-verbal audio signals to determine conversation characteristics and predict relationships, the system generates actionable insights that enhance the accuracy and completeness of multi-speaker audio conversation analysis.

WO2026105986A1PCT designated stage Publication Date: 2026-05-21SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2025-04-21
Publication Date
2026-05-21

Smart Images

  • Figure KR2025005362_21052026_PF_FP_ABST
    Figure KR2025005362_21052026_PF_FP_ABST
Patent Text Reader

Abstract

A method (800) and a system (102) for generating actionable insights for multi-speaker audio conversation are disclosed. The method (800) for generating actionable insights for multi-speaker audio conversation is disclosed. The method (800) comprises receiving an audio input, associated with an audio conversation, having verbal audio signal and non-verbal audio signal. Further, the method (800) comprises determining a plurality of conversation characteristics corresponding to each speaker based on the verbal audio signal and the non-verbal audio signal. The method (800) comprises determining conversation tonal attributes corresponding to each speaker based on the determined conversation characteristics. The method (800) comprises predicting at least one relationship among the speakers based on the determined conversation tonal attributes. Further, the method (800) comprises generating actionable insights, for each speaker, based on the determined relationship.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND SYSTEMS FOR GENERATING ACTIONABLE INSIGHTS FOR MULTI-SPEAKER AUDIO CONVERSATION

[0001] The present disclosure relates to audio processing, and in particular, relates to a method and a system for generating actionable insights for multi-speaker audio conversation.

[0002] With the advancement in technology, various electronic devices, such as smartphones, are provided with speech-text conversion functions that generally convert an audio conversation into a textual data. For instance, such speech-text conversion functions are usually provided to capture the conversation audio and convert such audio into a textual summary or insights related to the audio conversation. Generally, during capturing of the conversation audio, both verbal audio and non-verbal audio, such as echoes, static, or background noises, are captured by the speech-text conversion functions.

[0003] However, currently, the speech-text conversion functions are not capable of effectively analyzing non-verbal audio in order to properly process the entire audio conversation for generating textual summaries or actionable insights. In particular, the non-verbal audio is usually discarded by the speech-text conversion functions as noise. This might result in an incomplete analysis of the audio conversation as some of the verbal audio overlapping with the non-verbal audio might also be discarded as noise, and therefore this further leads to inaccurate generation of summary or insights related to the audio conversation. Further, currently, the summaries or insights generated are merely based on the utterances in the audio conversation, and conversational relationships between speakers are neglected. Therefore, currently, the speech-text conversion functions fail to provide an optimal user experience.

[0004] Therefore, it is desirable to provide a system and a method that can eliminate one or more of the above-mentioned problems associated with the generation of actionable insights for multi-speaker audio conversation.

[0005] This summary is provided to introduce a selection of concepts, in a simplified format, that are further described in the detailed description of the invention. This summary is neither intended to identify key or essential inventive concepts of the invention and nor is it intended for determining the scope of the invention.

[0006] In an embodiment, a method for generating actionable insights for multi-speaker audio conversation is disclosed. The method comprises receiving an audio input, associated with an audio conversation, having verbal audio signal and non-verbal audio signal. Further, the method comprises determining a plurality of conversation characteristics corresponding to each speaker based on the verbal audio signal and the non-verbal audio signal. The method comprises determining conversation tonal attributes corresponding to each speaker based on the determined conversation characteristics. The method comprises predicting at least one relationship among the speakers based on the determined conversation tonal attributes. Further, the method comprises generating actionable insights, for each speaker, based on the determined relationship.

[0007] In another embodiment, a system for generating actionable insights for multi-speaker audio conversation is disclosed. The system comprises at least one processor configured to receive an audio input, associated with an audio conversation, having verbal audio signal and non-verbal audio signal. Further, the at least one processor is configured to determine a plurality of conversation characteristics corresponding to each speaker based on the verbal audio signal and the non-verbal audio signal. The at least one processor is configured to determine conversation tonal attributes corresponding to each speaker based on the determined conversation characteristics. Further, the at least one processor is configured to predict at least one relationship among the speakers based on the determined conversation tonal attributes. Furthermore, the at least one processor is configured to generate actionable insights, for each speaker, based on the determined relationship.

[0008] To further clarify the advantages and features of the methods, systems, and apparatuses / devices, a more particular description of the methods, systems, and apparatuses / devices will be rendered by reference to specific embodiments thereof, which are illustrated in the appended drawings. It is appreciated that these drawings depict only typical embodiments of the disclosure and are therefore not to be considered limiting of its scope. The disclosure will be described and explained with additional specificity and detail with the accompanying drawings.

[0009] These and other features, aspects, and advantages of the present invention will become better understood when the following detailed description is read with reference to the accompanying drawings in which like characters represent like parts throughout the drawings, wherein:

[0010] Figures 1a-1b illustrates an environment depicting an electronic device having a system for generating actionable insights for multi-speaker audio conversation, according to various embodiments of the present disclosure;

[0011] Figure 2 illustrates a block diagram of a system for generating actionable insights for multi-speaker audio conversation, according to an embodiment of the present disclosure;

[0012] Figure 3 illustrates a block diagram depicting clustering of an audio input associated with the audio conversation, according to an embodiment of the present disclosure;

[0013] Figure 4 illustrates a block diagram depicting determining of plurality of conversation characteristics corresponding to each speaker, according to an embodiment of the present disclosure;

[0014] Figure 5 illustrates a block diagram depicting determining of conversation tonal attributes corresponding to each speaker, according to an embodiment of the present disclosure;

[0015] Figure 6 illustrates a block diagram depicting determining of at least one relationship among the speakers, according to an embodiment of the present disclosure;

[0016] Figure 7 illustrates a block diagram depicting generation of actionable insights, according to an embodiment of the present disclosure;

[0017] Figure 8 illustrates a flowchart depicting a method for generating actionable insights for multi-speaker audio conversation, according to an embodiment of the present disclosure;

[0018] Figure 9 illustrates a method of providing actionable insights, according to an embodiment of the present disclosure; and

[0019] Figure 10 illustrates a block diagram of an electronic device according to another embodiment of the present disclosure.

[0020] Further, skilled artisans will appreciate that elements in the drawings are illustrated for simplicity and may not have necessarily been drawn to scale. For example, the flow charts illustrate the method in terms of the most prominent steps involved to help to improve understanding of aspects of the disclosure. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the embodiments of the disclosure so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.

[0021] For the purpose of promoting an understanding of the principles of the invention, reference will now be made to the embodiment illustrated in the drawings and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the invention is thereby intended, such alterations and further modifications in the illustrated system, and such further applications of the principles of the invention as illustrated therein being contemplated as would normally occur to one skilled in the art to which the invention relates. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skilled in the art to which this invention belongs. The system, methods, and examples provided herein are illustrative only and not intended to be limiting.

[0022] Embodiments of the present invention will be described below in detail with reference to the accompanying drawings.

[0023] Figures 1a and 1b illustrate environments 100-1, 100-2 depicting an electronic device 101 having a system 102 for generating actionable insights 104, 104-1, 104-2 for multi-speaker audio conversation, according to various embodiments of the present disclosure. In an embodiment, the electronic device 101 may be embodied as a smartphone, a tablet, a personal computer, a smart television, a smart speaker, a wearable device, or any other electronic device having audio processing capabilities, without departing from the scope of the present disclosure. In the illustrated embodiments, the electronic device 101 may include, but is not limited to, the system 102 for generating actionable insights 104, 104-1, 104-2 for multi-speaker audio conversation.

[0024] Referring to Figure 1a, in one embodiment, the system 102 of the electronic device 101 may be configured to receive an audio input corresponding to an audio conversation between a speaker S1 and a speaker S2. In the illustrated embodiment, the speaker S1 and the speaker S2 may be located in a same environment during the audio conversation, and the electronic device 101 may also be located in the same environment as of the speaker S1 and the speaker S2. Based on the received audio input, the system 102 may be configured to generate actionable insights 104 for each of the speaker S1 and the speaker S2. The actionable insights 104 may include a contextual information / element which is determined based on the audio conversation between the speaker S1 and the speaker S2. For example, the actionable insights 104 may include reminders, tasks, summaries, objectives, descriptive entries, action items, constructive and / or concerning elements, service descriptors (in example where speakers are related as client and customer), etc.

[0025] Referring to Figure 1b, in another embodiment, a speaker S1 and a speaker S2 may be located at different locations. For example, the speaker S1 and the speaker S2 may be on a call with each other. In such an embodiment, the speaker S1 and the speaker S2 may be engaged in an audio conversation through an electronic device 101-1 and an electronic device 101-2, respectively. Each of the electronic devices 101-1, 101-2 may be provided with the system 102 for generating actionable insights 104-1, 104-2 for multi-speaker audio conversation. In the present embodiment, the system 102 of each of the electronic devices 101-1, 101-2 may receive an audio input corresponding to the audio conversation between the speaker S1 and the speaker S2. The audio input received by the system 102 of each of the electronic devices 101-1, 101-2 may have verbal audio signal and non-verbal audio signal. In the illustrated embodiment, considering that the speaker S1 and the speaker S2 are located at different locations, the non-verbal audio signal may be distinct for both the speaker S1 and the speaker S2. Therefore, the system 102 may use both the non-verbal audio signal and the verbal audio signal for precise and effective audio processing, such as audio diarization, prior to generating the actionable insights 104-1, 104-2. Based on the received audio input, the system 102 of the electronic device 101-1 may be configured to generate the actionable insights 104-1 for the speaker S1, and the system 102 of the electronic device 101-2 may be configured to generate the actionable insights 104-2 for the speaker S2. The actionable insights 104-1, 104-2 may include a contextual information / element which is determined based on the audio conversation between the speaker S1 and the speaker S2. For example, the actionable insights 104-1, 104-2 may include reminders, tasks, summaries, objectives, descriptive entries, etc.

[0026] Constructional and operational details of the system 102 are explained in detail in the description of Figure 2.

[0027] Figure 2 illustrates a block diagram of the system 102 for generating the actionable insights 104, 104-1, 104-2 for multi-speaker audio conversation, according to an embodiment of the present disclosure. In the illustrated embodiment, the system 102 may be implemented in the electronic device 101 for generating the actionable insights 104, 104-1, 104-2 for multi-speaker audio conversation. In an embodiment, the system 102 may include a processor 202, memory 204, module(s) 206, and data 208. The module(s) 206 and the memory 204 are coupled to the processor 202.

[0028] The processor 202 can be a single processing unit or a number of units, all of which could include multiple computing units. The processor 202 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuitries, and / or any devices that manipulate signals based on operational instructions. Among other capabilities, the processor 202 is configured to fetch and execute computer-readable instructions and data stored in the memory 204.The processor 202 may include one or a plurality of processors. At this time, one or a plurality of processors may be a general-purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU).

[0029] The processor may include various processing circuitry and / or multiple processors.  For example, as used herein, including the claims, the term "processor" may include various processing circuitry, including at least one processor, wherein one or more of at least one processor, individually and / or collectively in a distributed manner, may be configured to perform various functions described herein. As used herein, when "a processor", "at least one processor", and "one or more processors" are described as being configured to perform numerous functions, these terms cover situations, for example and without limitation, in which one processor performs some of recited functions and another processor(s) performs other of recited functions, and also situations in which a single processor may perform all recited functions. Additionally, the at least one processor may include a combination of processors performing various of the recited / disclosed functions, e.g., in a distributed manner.  At least one processor may execute program instructions to achieve or perform various functions.

[0030] The one or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or artificial intelligence (AI) model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning.

[0031] Here, being provided through learning means that, by applying a learning technique to a plurality of learning data, a predefined operating rule or AI model of a desired characteristic is made. The learning may be performed in a device itself in which AI according to an embodiment is performed, and / or may be implemented through a separate server / system.

[0032] The AI model may consist of a plurality of neural network layers. Each layer has a plurality of weight values, and performs a layer operation through calculation of a previous layer and an operation of a plurality of weights. Examples of neural networks include, but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann Machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), generative adversarial networks (GAN), and deep Q-networks.

[0033] The artificial intelligence models, explained in the disclosure, may be obtained by training. Here, "obtained by training" means that a predefined operation rule or artificial intelligence model configured to perform a desired feature (or purpose) is obtained by training a basic artificial intelligence model with multiple pieces of training data by a training technique. The training technique is a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of learning techniques include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0034] The memory 204 may include any non-transitory computer-readable medium known in the art including, for example, volatile memory, such as static random-access memory (SRAM) and dynamic random-access memory (DRAM), and / or non-volatile memory, such as read-only memory (ROM), erasable programmable ROM, flash memories, hard disks, optical disks, and magnetic tapes.

[0035] The module(s) 206, amongst other things, include routines, programs, objects, components, data structures, etc., which perform particular tasks or implement data types. The module(s) 206 may also be implemented as, signal processor(s), state machine(s), logic circuitries, and / or any other device or component that manipulate signals based on operational instructions.

[0036] Further, the module(s) 206 may be implemented in hardware, instructions executed by at least one processing unit, for e.g., the processor 202, or by a combination thereof. The processing unit may comprise a computer, a processor, a state machine, a logic array and / or any other suitable devices capable of processing instructions. The processing unit may be a general-purpose processor which executes instructions to cause the general-purpose processor to perform operations or, the processing unit may be dedicated to performing the required functions. In some example embodiments, the module(s) 206 may be machine-readable instructions (software, such as web-application, mobile application, program, etc.) which, when executed by a processor / processing unit, perform any of the described functionalities.

[0037] In an implementation, the module(s) 206 may include a receiving module 210, an audio separation module 212, a speaker diarization module 214, a feature extraction module 216, a tone determination module 218, a relationship determination module 220, a prompt generation module 222, and an insight generation module 224. The receiving module 210, the audio separation module 212, the speaker diarization module 214, the feature extraction module 216, the tone determination module 218, the relationship determination module 220, the prompt generation module 222, and the insight generation module 224 are in communication with each other. At least one of the modules 206 may be implemented through an AI model. A function associated with AI may be performed through the non-volatile memory, the volatile memory, and the processor. The data 208 serves, amongst other things, as a repository for storing data processed, received, and generated by one or more of the modules 206.

[0038] In an embodiment of the present disclosure, the module(s) 206 may be implemented as part of the processor 202. In another embodiment of the present disclosure, the module(s) 206 may be external to the processor 202. In yet another embodiment of the present disclosure, the module(s) 206 may be part of the memory 204. In another embodiment of the present disclosure, the module(s) 206 may be part of hardware, separate from the processor 202.

[0039] In an embodiment, the processor 202 may be configured to receive an audio input, associated with an audio conversation, having verbal audio signal and non-verbal audio signal. In the illustrated embodiment, the receiving module 210 may be configured to receive the audio input associated with the audio conversation.

[0040] In an embodiment, the processor 202 may be configured to separate the verbal audio signal and the non-verbal audio signal from the received audio input. The verbal audio signal may be associated with the use of spoken words and language to convey information in the audio conversation between the multiple speakers. Further, the non-verbal audio may be associated with elements of the audio input that do not rely on the spoken words but instead convey context through other sound-related factors, such as echoes, static, background noises, dropped calls or surrounding sounds. In the illustrated embodiment, the audio separation module 212 may be configured to separate the verbal audio signal and the non-verbal audio signal from the received audio input. In an exemplary embodiment, the audio separation module 212 may be deployed with a Wiener-Kolmogorov filter to separate the non-verbal audio signal, representing noise, from the received audio input, without departing from the scope of the present disclosure.

[0041] In an embodiment, at least one module of the modules 206 may be implemented on a server. For example, the receiving module 210, the audio separation module 212, the speaker diarization module 214 and the feature extraction module 216 may be implemented on the electronic device 102, and the tone determination module 218, the relationship determination module 220, the prompt generation module 222, and the insight generation module 224 may be implemented on the server. In this case, the electronic device 102 may transmit textual data of the audio input and a plurality of conversation characteristics for the speaker S1 and the speaker S2 to the server, receive the actionable insights 104, 104-1, 104-2 from the server.

[0042] Figure 3 illustrates a block diagram depicting clustering of the audio input associated with the audio conversation, according to an embodiment of the present disclosure. Referring to Figures 2 and 3, the processor 202 may be configured to generate verbal audio clusters corresponding to each speaker based on the verbal audio signal and the non-verbal audio signal. In the illustrated embodiment, the speaker diarization module 214 may be configured to generate the verbal audio clusters corresponding to each speaker. The speaker diarization module 214 may be configured to receive the separated verbal audio signal and the separated non-verbal audio signal from the audio separation module 212.

[0043] Referring to Figure 3, at block 301, the speaker diarization module 214 may be configured to determine verbal Melfrequency Cepstra Confficients (MFCCs) based on the separated verbal audio signal. In an exemplary embodiment, the speaker diarization module 214 may amplify higher frequencies to balance spectrum associated with the separated verbal audio signal. Further, the separated verbal audio signal may be divided into small and overlapping frames. Thereafter, a Hamming window may be applied on each frame to soften edges of each frame. Further, each frame may be converted from a time domain to a frequency domain. Furthermore, Mel-filterbank may be used for processing each frame. The speaker diarization module 214 may calculate logarithm of the output from the Mel-filterbank. Thereafter, the speaker diarization module 214 may apply Discrete Cosine Transform (DCT) to the logarithm of the output from the Mel-filterbank to obtain verbal MFCCs. Similarly, at block 302, the speaker diarization module 214 may be configured to determine non-verbal Cepstral Coefficients (CCs) using the similar process, as explained for the verbal MFCCS, except that the Mel-filterbank may not be applied for processing the non-verbal audio signal.

[0044] In an embodiment, the speaker diarization module 214 may be configured to segment each of the separated verbal audio signal and the separated non-verbal audio signal. Referring to Figure 3, at block 303, the speaker diarization module 214 may be configured to segment the separated verbal audio signal based on the verbal MFFCs. In an embodiment, the speaker diarization module 214 may be configured to perform Voice Activity Detection (VAD), at block 305 in Figure 3, prior to segmenting the separated verbal audio signal. Similarly, at block 304, the speaker diarization module 214 may be configured to segment the separated non-verbal audio signal using the non-verbal CCs.

[0045] Further, at blocks 306, the speaker diarization module 214 may be configured to extract X-vector based on the segmented verbal audio signal. Similarly, at 307, the speaker diarization module 214 may be configured to extract X-vector based on the segmented non-verbal audio signal. In an embodiment, the speaker diarization module 214 may be deployed with a Time Delay Neural Network (TDNN) and a full connect NN model to extract X-vector for each of the segmented verbal audio signal and the segmented non-verbal audio signal.

[0046] Referring to Figure 3, at blocks 308 and 309, the speaker diarization module 214 may be configured to generate the verbal audio clusters and non-verbal audio clusters by clustering each of the segmented verbal audio signal and the segmented non-verbal audio signal. Each cluster of the verbal audio clusters may be indicative of an audio utterance of one of the speakers and each cluster of the non-verbal audio clusters is indicative of a noise associated with one of the speakers. In an embodiment, at each block 308 and 309, the speaker diarization module 214 may be configured to perform a Probabilistic Linear Discriminant Analysis (PLDA) for clustering each of the segmented verbal audio signal and the segmented non-verbal audio signal. In the PLDA, the X-vector may be decomposed into multiple factors as represented by (1) below:

[0047] ..........(1)

[0048] Where, m: Speaker independent mean

[0049] Fhi: Speaker inherent feature

[0050] Gwij: Utterance dependent feature

[0051] ni,j: Random noise

[0052] i: Speaker index

[0053] j: Utterance index

[0054] m, F, G, and : Model parameters pre-determined using training data

[0055] Based on the decomposition of the X-vector, the speaker diarization module 214 may be configured to determine cluster likelihood, as represented by (2), for each of the segmented verbal audio signal and the non-verbal audio signal.

[0056] ....(2)

[0057] Further, at block 310, the speaker diarization module 214 may be configured to perform segmentation refinement 310 by mapping the verbal audio clusters with the non-verbal audio clusters. In the segmentation refinement 310, based on the mapping of the clustered verbal audio signal and the clustered non-verbal audio signal, the speaker diarization module 214 may be configured to determine non-conflicting overlapping clusters and conflicting overlapping clusters. The non-conflicting overlapping clusters may refer to the overlapping clusters where the same speaker may be identified in each of the overlapped verbal audio cluster and the overlapped non-verbal audio cluster. The conflicting overlapping clusters may refer to the overlapping clusters where different speakers may be identified in each of the overlapped verbal audio cluster and the overlapped non-verbal audio cluster.

[0058] Further, the speaker diarization module 214 may be configured to verify speaker corresponding to each conflicting overlapping cluster and each non-conflicting overlapping cluster. In the preferred embodiment, non-verbal audio content in the non-verbal audio signal for each speaker may be distinct and therefore, the non-verbal audio signal may be used in combination with the verbal audio signal for verifying the speaker at each time interval in the verbal audio cluster. In an embodiment, the speaker diarization module 214 may be deployed with an Neural Network trained to verify the speaker at each time interval in the conflicting overlapping clusters and the non-conflicting overlapping clusters. At block 311, the speaker diarization module may be configured to generate modified verbal audio clusters, hereinafter referred to as the verbal audio clusters 311 or the generated verbal audio clusters 311, based on the verified speakers.

[0059] The processor 202 may be configured to generate textual data based on the generated verbal audio cluster. Machine learning models may be trained on a large dataset of audio recordings and corresponding transcripts to make the system 102 understand the patterns and variations in human speech. The trained model may then be used to recognize the speech patterns in the generated verbal audio cluster. By comparing the extracted features with those in the training data, the system 102 may determine the most likely sequence of words spoken. In the illustrated embodiment, the processor 202 may deploy an Automatic Speech Recognition (ASR) system with one or more Artificial Intelligence (AI)-based language models and AI-based acoustic models to generate textual data 401 based on the generated verbal audio cluster, without departing from the scope of the present disclosure. In such an embodiment, the textual data 401 may be embodied as a transcribed text representing spoken words associated with the audio conversation between the multi-speakers.

[0060] Figure 4 illustrates a block diagram depicting determining of a plurality of conversation characteristics corresponding to each speaker, according to an embodiment of the present disclosure. The processor 202 may be configured to determine the plurality of conversation characteristics corresponding to each speaker based on the generated verbal audio clusters 311. Referring to Figure 4, in the illustrated embodiment, the feature extraction module 216 may be configured to determine the plurality of conversation characteristics corresponding to each speaker based on the generated verbal audio clusters and the generated textual data. The plurality of conversation characteristics may include, but is not limited to, frequency spectrum 402, loudness 404, conversation speed 408, characteristic word vectors 406, and sentiment 410. In an embodiment, the characteristic word vectors 406 may interchangeably be referred to as start / end word vectors 406 or a start word vector and an end word vector 406, without departing from the scope of the present disclosure.

[0061] In an exemplary embodiment, the feature extraction module 216 may be configured to determine the frequency spectrum 402, corresponding to the audio utterance of each speaker, by applying frequency transformation technique on the generated verbal audio clusters 311. In such an embodiment, the feature extraction module 216 may apply frequency transformation technique, such as Fourier Transform, to determine the frequency spectrum 402 for each speaker from the generated verbal audio clusters 311. The Fourier Transform may be used to decompose the time-domain audio signal into its frequency-domain representation, effectively separating the signal into its constituent frequency components. This is represented as a Fourier spectrum, where the amplitude Ak, as represented by (1) below, corresponds to the intensity of the k-th frequency component in the signal. By performing this transformation, a detailed view of the frequency content of the speaker's speech may be obtained.

[0062] .......(1)

[0063] Further, in an exemplary embodiment, the feature extraction module 216 may be configured to determine the loudness 404, corresponding to the audio utterance of each speaker, by calculating an average intensity associated with the generated verbal audio clusters 311. To calculate the average intensity associated with the generated verbal audio clusters 311, energy of the generated verbal audio clusters 311 may be approximated by summing a square of an amplitude at each sample point over a time segment. To find the average intensity, the sum of the square roots of the sample values over the segment and divide by the total time duration associated with the generated verbal audio clusters 311. Finally, the loudness 404 in decibels (dB) can be derived from the average intensity, typically using a logarithmic scale to represent the perceived loudness as represented by (2) given below:

[0064] β = 10 log(I / I0) ......(2)

[0065] In an exemplary embodiment, the feature extraction module 216 may be configured to determine the conversation speed 408 by calculating a number of words uttered by each speaker based on the textual data 401 of the verbal audio clusters for each speaker, and calculating a time duration in which the number of words is uttered. The conversation speed 408 may be calculated by dividing the number of words uttered with the time duration in which the number of words is uttered.

[0066] In an exemplary embodiment, the feature extraction module 216 may be configured to determine a start word vector and an end word vector 406 by evaluating a start word phrase and an end word phrase in the textual data 401 of the generated verbal audio clusters 311 for each speaker. In an exemplary embodiment, the feature extraction module 216 may be deployed with a pre-trained embedding model to evaluate the start word phrase and the end word phrase to determine the start word vector and the end word vector 406. In one or more exemplary embodiments, the start word vector and the end word vector 406 may be extracted from a pre-calculated database such as fastText, GloVe, or word2vec, without departing from the scope of the present disclosure.

[0067] In an exemplary embodiment, the feature extraction module 216 may be configured to determine the sentiment 410 based on the generated verbal audio clusters 311 of each speaker and the textual data 401 of the generated verbal audio clusters 311 for each speaker. In such an embodiment, firstly, the sentiment 410 may be determined based on the generated verbal audio clusters 311 using a Convolutional Neural Network (CNN), without departing from the scope of the present disclosure. The generated verbal audio clusters 311 may be inputted in the CNN to extract features associated with the generated verbal audio clusters 311 of each speaker. Subsequently, the CNN may classify the extracted features to generate output indicative of the sentiment 410 associated with the generated verbal audio clusters 311 of each speaker. Secondly, the sentiment 410 may be determined based on the textual data 401 of the generated verbal audio clusters 311 for each speaker. In one or more embodiments, the sentiment 410 may be determined based on the textual data 401 using one of lexicon-based and machine learning-based approaches, such as implementing deep learning AI models. Thereafter, the sentiment 410 based on the generated verbal audio clusters 311 and the sentiment 410 based on the textual data 401 may be combined by weighted average to obtain final sentiment 410 corresponding to each speaker.

[0068] Figure 5 illustrates a block diagram depicting determining of conversation tonal attributes corresponding to each speaker, according to an embodiment of the present disclosure. In an embodiment, the processor 202 may be configured to determine conversation tonal attributes corresponding to each speaker. In an embodiment, the plurality of conversation tonal attributes may include, but is not limited to, formal tone 502, informal tone 504, engaged tone 506, disengaged tone 508, concerned tone 510, unconcerned tone 512, dominant tone 514, and subordinate tone 516. Referring to Figure 5, in the illustrated embodiment, the tone determination module 216 may be configured to determine the conversation tonal attributes corresponding to each speaker.

[0069] In an embodiment, the tone determination module 216 may be deployed with a multi-output Neural Network (NN) 518 to determine the conversation tonal attributes corresponding to each speaker. The multi-output NN 518 may be a type of neural network designed to predict multiple outcomes or targets simultaneously from a single input. The multi-output NN 518 may be trained to generate multiple predictions (or regressions) at the same time. The plurality of conversation characteristics corresponding to each speaker may be inputted in the multi-output NN 518 to determine the conversation tonal attributes for the respective speaker. In an embodiment, the multi-output NN 518 may have multiple output layers, each responsible to predict / determine a different output. Each output may be associated with distinct activation function and loss function, depending on a nature of the prediction / determination, such as classification or regression. Further, the multi-output NN 518 may have underlying hidden layers, representing learned features of the plurality of conversation characteristics corresponding to each speaker, that are shared across all the outputs. Such shared layers may extract features from the plurality of conversation characteristics corresponding to each speaker, which are then used by different output layers to determine the conversation tonal attributes corresponding to each speaker. Further, each output in the multi-output NN 518 may have a distinct associated loss function depending on the type of output.

[0070] Figure 6 illustrates a block diagram depicting determining of at least one relationship among the speakers, according to an embodiment of the present disclosure. In an embodiment, the processor 202 may be configured to determine at least one relationship 602 among the speakers of the audio conversation. The at least one relationship 602 among the speakers may include, but is not limited to, family relationship, professional relationship, formal relationship, romantic relationship, and friendly relationship. Referring to Figure 6, in the illustrated embodiment, the relationship determination module 220 may be configured to determine the at least one relationship 602 among the speakers of the audio conversation.

[0071] In an embodiment, the relationship determination module 220 may be deployed with a multi-class Machine Learning (ML) model 604 to determine the at least one relationship 602 among the speakers of the audio conversation based on the determined conversation tonal attributes. The multi-class ML model 604 may be a type of ML model designed for multi-class classification tasks, where the objective may be to classify an input into one of multiple categories. A preferred embodiment may implement a multi-class Neural Network (NN) as the multi-class ML model. The multi-class classification may involve two or more than two classes, without departing from the scope of the present disclosure.

[0072] Further, in an embodiment, the processor 202 may be configured to generate actionable insights, such as the actionable insights 104, 104-1, 104-2, for each speaker. In the illustrated embodiment, the insight generation module 224 may be configured to generate the actionable insights 104, 104-1, 104-2 for each speaker. In order to generate the actionable insights 104, 104-1, 104-2, firstly, the prompt generation module 222 may receive the relationship determined by the relationship determination module 220 and the conversation tonal attributes determined by the tone determination module 218. Further, the prompt generation module 222 may be configured to generate prompts based on a combination of the relationship and the conversation tonal attributes for each speaker. In one embodiment, the prompt generation module 222 may be deployed with a rule-based system to generate prompts based on the combination of the relationship and the conversation tonal attributes. In another embodiment, the prompt generation module 222 may be deployed with a generative AI engine which is trained to generate prompts based on the combination of the relationship and conversation tonal attributes.

[0073] Figure 7 illustrates a block diagram depicting generation of the actionable insights 104, 104-1, 104-2, according to an embodiment of the present disclosure. Referring to Figure 7, in the illustrated embodiment, the insight generation module 224 may be configured to generate the actionable insights 104, 104-1, 104-2, using the prompts as input, based on the determined relationship and the determined conversation. In an embodiment, the prompts along with the textual data associated with the verbal audio cluster for each speaker may be inputted into the insight generation module 224 to generate the actionable insights 104, 104-1, 104-2. In the illustrated embodiment, the insight generation module 224 may be deployed with a generative AI model, to generate the actionable insights 104, 104-1, 104-2.

[0074] In an embodiment, each actionable insight may include a contextual element indicative of a context associated with one or more insights for each speaker. Further, in various embodiments, each actionable insight may be embodied as an audio-based insight, a visual-based insight, a text-based insight, or a combination thereof, without departing from the scope of the present disclosure. For instance, in one example, the actionable insight may be embodied as a textual summary / description containing the context associated with each action to be performed by the speaker. In another example, the actionable insight may be embodied as an audio reminder to perform a specific task, where such a reminder is combined with a textual summary of the specific task to be performed by the speaker.

[0075] In one of the exemplary implementations, the audio input related to an audio conversation, as depicted in Table 1, between a speaker 1 and a speaker 2 may be received by the processor 202.

[0076] Speaker 1Speaker 2Good morning, how are you today?Great, thanks! So, I wanted to talk to you about our upcoming project. We have a lot of work to do, and we need everyone to pitch in. Can you handle the workload?We need to finish the design phase by next week, and then move on to implementation. Can you take charge of the design part?Excellent! Let me know if there's anything else I can help with.Hi, I'm doing well. How about yourself?Yes, definitely. What exactly needs to be done?Sure, no problem. I'll get started right away.Thank you, I will. Have a great day!

[0077] In such an exemplary implementation, the processer 202 may generate verbal audio clusters corresponding to the speaker 1 and the speaker 2 based on the audio input. Further, the processor 202 may determine the plurality of conversation characteristics for the speaker 1 and the speaker 2 based on the verbal audio clusters. Thereafter, the processor 202 may determine the conversation tonal attributes for the speaker 1 as "formal" and "dominant", and the conversation tonal attributes for the speaker 2 as "formal" and "subordinate" based on the plurality of conversation characteristics for the speaker 1 and the speaker 2. Further, the processor 202 may determine the relationship between the speaker 1 and the speaker 2 as "professional". Based on the conversation tonal attributes and the relationship, the processor 202 may generate the actionable insights for the speaker 2 as mentioned in the Table 2 below:

[0078] Speaker 2 - Actionable InsightsAction Items:1.Finish the design phase of the project by next week.2.Start working on the design part of the project immediately.3.Keep the manager updated on progress and any challenges encountered during the design process.Reminders:1.Don't forget to prioritize the design task and allocate sufficient time to complete it before the deadline.2.Stay organized and maintain clear communication with the team throughout the project3.Remember to ask for help from the manager or other colleagues if needed.

[0079] In an embodiment, the processor 202 may summarize the audio conversation as "The manager asked about your availability and capacity to handle the workload for the upcoming project. They mentioned that there is a lot of work to be done and expressed the importance of everyone pitching in. Specifically, the manager requested that you take charge of the design part of the project, ensuring that it is completed by next week. The manager offered support if needed and encouraged you to let them know if there's anything else they can assist with."

[0080] In one of the exemplary implementations, the audio input related to a phone conversation, as depicted in Table 3, between a speaker 1 and a speaker 2 may be received by the processor 202.

[0081] Speaker 1Speaker 2Hi Son! How are you doing today?Just wanted to check in and see how your day is going.That's great! Did you remember to take your vitamins?Good job! Now, let me ask you something. Have you thought about what you want to be when you grow up?That's okay. You have plenty of time to figure it out. But it's important to explore different options and follow your passions.Excellent attitude! Remember, I'm always here to support you no matter what path you choose.Love you too, Son. Have a great rest of your day!I'm good, Mom.What's up?It's going well. I just finished my homework and now I'm taking a break.Yes, I did. And I had a healthy snack too.Not really. I have a lot of interests, but I haven't decided yet.Yeah, I know. I'll keep trying new things and see where it takes me.Thanks, Mom. Love you!

[0082] In such an exemplary implementation, the processer 202 may generate verbal audio clusters corresponding to the speaker 1 and the speaker 2 based on the audio input. Further, the processor 202 may determine the plurality of conversation characteristics for the speaker 1 and the speaker 2 based on the verbal audio clusters. Thereafter, the processor 202 may determine the conversation tonal attributes for the speaker 1 as "informal" and "concerned", and the conversation tonal attributes for the speaker 2 as "informal" and "subordinate" based on the plurality of conversation characteristics for the speaker 1 and the speaker 2. Further, the processor 202 may determine the relationship between the speaker 1 and the speaker 2 as "family".

[0083] In an embodiment, the processor 202 may summarize the audio conversation as "Mom called to check in on me and make sure I'm doing okay. She asked if I took my vitamins and had a healthy snack, which I did. We talked about my future plans, and even though I don't know exactly what I want to be, she reminded me to explore different options and follow my passions. She said she would support me no matter what path I choose. It was a nice talk, and we ended it by telling each other "I love you."

[0084] Based on the conversation tonal attributes and the relationship, the processor 202 may generate the actionable insights for the speaker 2 as mentioned in the Table 4 below:

[0085] Speaker 2 - Actionable InsightsAction Items:1.Continue exploring different options and passions to determine future career choices.2.Take vitamins and have healthy snacks daily.3.Keep in touch with mom regularly to update her on progress and discuss future plans.Reminders:1.Remember to prioritize health and well-being.2.Stay open-minded and flexible in considering career paths.3.Appreciate the support and guidance provided by mom.

[0086] In one of the exemplary implementations, the audio input related to an audio conversation, as depicted in Table 5, between a speaker 1 and a speaker 2 may be received by the processor 202.

[0087] Speaker 1Speaker 2Hello, I'm interested in your software services. Can you please tell me more about them?That sounds great! Can you give me some examples of projects you've worked on?Wow, those sound like impressive projects! How much would it cost to develop a website similar to the ones you mentioned?Okay, sure. Let me gather my thoughts and provide you with the necessary information. Thank you!Sure! We offer a wide range of software services including web development, mobile app development, software testing, and software maintenance. We have experienced professionals who are experts in different technologies and platforms.Yes, definitely! We have worked on various projects ranging from small business websites to complex enterprise applications. For example, we recently developed a custom CRM system for a leading healthcare company. They were able to streamline their patient management process and increase efficiency using our solution. Another project we worked on was developing a mobile application for a food delivery startup. The app allowed customers to order food online and track their orders in real-time.The cost depends on several factors such as the complexity of the website, the number of features required, and the timeline. However, we can provide you with a detailed quote once we understand your specific requirements. Please let us know more about your project and we'll be happy to assist you further.No problem! Looking forward to hearing from you soon.

[0088] In such an exemplary implementation, the processer 202 may generate verbal audio clusters corresponding to the speaker 1 and the speaker 2 based on the audio input. Further, the processor 202 may determine the plurality of conversation characteristics for the speaker 1 and the speaker 2 based on the verbal audio clusters. Thereafter, the processor 202 may determine the conversation tonal attributes for the speaker 1 as "formal", "Engaged" and "professional", and the conversation tonal attributes for the speaker 2 as "formal", "Engaged" and "Dominant" based on the plurality of conversation characteristics for the speaker 1 and the speaker 2. Further, the processor 202 may determine the relationship between the speaker 1 and the speaker 2 as "client".

[0089] In an embodiment, the processor 202 may summarize the audio conversation as "The client expressed interest in the software services offered by the provider. The provider explained their expertise in web development, mobile app development, software testing, and maintenance. The client asked for examples of past projects, and the provider shared successful cases such as developing a custom CRM system for a healthcare company and a mobile app for a food delivery startup. The client then enquired about the cost of developing a similar website, and the provider emphasized the importance of understanding specific requirements before providing a detailed quote."

[0090] Based on the conversation tonal attributes and the relationship, the processor 202 may generate the actionable insights for the speaker 2 as mentioned in the Table 6 below:

[0091] Speaker 2 - Actionable InsightsAction Items:1.Gather more information about the client's specific requirements for the website development project.2.Prepare a detailed quote for the client, taking into account the complexity of the website, the number of features required, and the desired timeline.3.Follow up with the client to schedule a meeting or call to discuss the project further and present the quote.Reminders:1.Remember to emphasize the importance of understanding the client's specific requirements before providing a quote.2.Ensure that the quote includes all the necessary details and is tailored to meet the client's needs.3.Schedule a follow-up meeting or call with the client to present the quote and discuss the next steps of the project.

[0092] Figure 8 illustrates a flowchart depicting a method 800 for generating actionable insights for multi-speaker audio conversation, according to an embodiment of the present disclosure. The method 800 may be implemented in the system 102 using components thereof, as described above. In an embodiment, the method 800 may be executed by the processor 202 of the system 102, as described above. Further, for the sake of brevity, details of the present disclosure that are explained in detail in the description of Figures 1a-7 are not explained in detail in the description of Figure 8.

[0093] At step 802, the method 800 may include receiving the audio input, associated with the audio conversation, having the verbal audio signal and the non-verbal audio signal. For example, the system 102 may receive the audio input, associated with call conversation between a user of the electronic device 101 and the other party, from the electronic device 101.

[0094] At step 806, the method 800 may include determining the plurality of conversation characteristics corresponding to each speaker based on the verbal audio signal and the non-verbal audio signal.

[0095] The method 800 may include generating the verbal audio clusters corresponding to each speaker based on the verbal audio signal and the non-verbal audio signal. In an embodiment, the method 800 may include separating the verbal audio signal and the non-verbal audio signal from the received audio signal. Further, the method 800 may include segmenting each of the separated verbal audio signal and the separated non-verbal audio signal. The method may include generating the verbal audio clusters and the non-verbal audio clusters by clustering each of the segmented verbal audio signal and the segmented non-verbal audio signal. Each cluster of the verbal audio clusters may be indicative of the audio utterance of one of the speakers and each cluster of the non-verbal audio clusters may be indicative of the noise associated with one of the speakers.

[0096] Further, the method 800 may be include mapping the verbal audio clusters with the non-verbal audio clusters. The method 800 may include determining non-conflicting overlapping clusters and conflicting overlapping clusters based on the mapping of the clustered verbal audio signal and the clustered non-verbal audio signal. Further, the method 800 may include verifying speaker corresponding to each conflicting overlapping cluster and each non-conflicting overlapping cluster. Furthermore, the method 800 may include generating modified verbal audio clusters based on the verified speakers.

[0097] Further, the method 800 may include determining the plurality of conversation characteristics corresponding to each speaker based on the generated verbal audio clusters. In an embodiment, the method 800 may include generating textual data based on the generated verbal audio clusters. Further, the method 800 may include determining the plurality of conversation characteristics corresponding to each speaker based on the generated verbal audio clusters and the generated textual data. The plurality of conversation characteristics includes at least one of frequency spectrum, loudness, conversation speed, characteristic word vectors, or sentiment.

[0098] At step 808, the method 800 may include determining the conversation tonal attributes corresponding to each speaker based on the determined conversation characteristics. In an embodiment, the method 800 may include determining, using the multi-output NN, conversation tonal attributes corresponding to each speaker. The conversation tonal attributes includes at least one of formal tone, informal tone, engaged tone, disengaged tone, concerned tone, unconcerned tone, dominant tone, or subordinate tone. When the determined conversation characteristics is input to the multi-output NN, the multi-output NN may output the conversation tonal attributes.

[0099] At step 810, the method may include predicting at least one relationship among the speakers based on the determined conversation tonal attributes. In an embodiment, the method 800 may include determining, using the multi-class ML model, the at least one relationship among the speakers of the audio conversation based on the determined conversation tonal attributes. The at least one relationship includes family relationship, professional relationship, formal relationship, romantic relationship or friendly relationship. When the determined conversation tonal attributes is input to the multi-class ML model, the multi-class ML model may output the at least one relationship among the speakers.

[0100] At step 812, the method 800 may include generating the actionable insights, for each speaker, based on the determined relationship. In an embodiment, the method 800 may include generating, using the generative AI model, the actionable insights based on the determined relationship and the determined conversation tonal attributes. Each actionable insight may include the contextual element indicative of the context associated with one or more insights for each speaker. In an embodiment, the method 800 may include generating, using a prompt engine, a prompt to be input to the generative AI model based on the determined relationship and the determined conversation tonal attributes. When the generated prompt is input to the generative AI model, the generative AI model may output the actionable insights.

[0101] If the determined relationship is professional, the actionable insights may include action items along with constructive and / or concerning elements. If the determined relationship is family / friend / romantic, the actionable insights may include important reminders that may be automatically set for the user in a suitable calendar / clock / alarm application. If the determined relationship is client / customer, the actionable insights may include service descriptors which describe satisfaction level of the client / customer as well as ways to improve upon it.

[0102] The method 800 may include matching the user of the electronic device 101 with one of the speakers by comparing the plurality of conversation characteristics corresponding to each speaker with the conversation characteristics of the user. The method 800 may include obtaining the conversation characteristics of the user separately from the call conversation. For example, the system 102 may receive the conversation characteristics of the user from the electronic device 101 either before or during the phone call.

[0103] The method 800 may include providing the actionable insights corresponding to the user. For example, the system 102 may display the actionable insights corresponding to the user through a display of the electronic device 101. The system 102 may display the predicted at least one relationship with the other party on the call conversation through the display of the electronic device 101.

[0104] The system 102 may include summarizing the call conversation and displaying the summarized call conversation through a display of the electronic device 101.

[0105] As would be gathered, the present disclosure offers a comprehensive approach of generating actionable insights for multi-speaker audio conversation. The system 102 and the method 800 of the present disclosure determines various parameters, such as the plurality of conversation characteristics, the relationship among speakers, and the conversation tonal attributes. Further, the system 102 and the method 800 may generate the actionable insights based on the plurality of conversation characteristics, the relationship among speakers, and the conversation tonal attributes. Therefore, the actionable insights generated by the system 102 and the method 800 are highly relevant and precise, enabling speakers to take effective and informed actions. This further enhances the user experience and also substantially reduces cognitive load for the speakers to maintain the context of the audio conversation.

[0106] Additionally, the system 102 and the method 800 are capable of effectively analyzing non-verbal audio in order to properly process the entire audio conversation for generating actionable insights. In particular, the non-verbal audio is not discarded by the system 108 and the method 800, rather the non-verbal audio is processed similar as the verbal audio to enhance the effectiveness of diarization process for audio clustering. This results in a complete analysis of the audio conversation as the non-verbal audio is not discarded as noise, and therefore this further increases accuracy while generating the actionable insights related to the audio conversation.

[0107] Figure 9 illustrates a method of providing actionable insights, according to an embodiment of the present disclosure.

[0108] Referring to FIG. 9, the electronic device 101 may display the actionable insights of the Table 6 in Fig. 7 through a display. The actionable insights includes action items 940 and reminders 950. The electronic device 101 may display the actionable insights during a call, and may alternatively display the actionable insights after the call ends.

[0109] The electronic device 101 may display the relationship 920 with the other party on the call. Additionally, the electronic device 101 can display the subject 930 of the call conversation.

[0110] The electronic device 101 may delete, edit, and save at least one of the displayed actionable insights. For example, the electronic device 101 may receive a user input of modifying the displayed actionable insights. The electronic device 101 may receive a user input of storing the displayed actionable insights in a calendar application or a To-Do application via button interfaces 960 or 970. The electronic device 101 may store and display the actionable insights corresponding to identification information (e.g., a name in a contact list or a phone number) of the other party on the call. Accordingly, when a call from the other party comes in, the electronic device 101 may display the actionable insights stored corresponding to the other party.

[0111] Figure 10 illustrates a block diagram of an electronic device according to another embodiment of the present disclosure.

[0112] Referring to FIG. 10, the electronic device 101 may include a microphone 1200, a communication module 1300, a memory 204, an input interface 1500, an output module 1600, a sensor 1700, and a processor 202. The same components as those shown in FIG. 2 are assigned like reference numerals.

[0113] All the shown components may not be essential components of the electronic device 101. The electronic device 101 may be configured with more components than those shown in FIG. 10 or with less components than those shown in FIG. 10.

[0114] The output module 1600 may include a sound output module 1620 and a display 1610.

[0115] The sound output module 1620 may output a sound signal to outside of the electronic device 101. The sound output module 1620 may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback.

[0116] The display 1610 may output image data subject to image-processing by an image processor (not shown) through a display panel (not shown), according to a control by the processor 202. The display panel (not shown) may include at least one of a liquid crystal display, a thin film transistor-liquid crystal display, an organic light-emitting diode, a flexible display, a 3-Dimensional (3D) display, or an electrophoretic display.

[0117] The input interface 1500 may receive a user input for controlling the electronic device 101. The input interface 1500 may receive the user input and transfer the user input to the processor 202.

[0118] The input interface 1500 may include a user input device including a touch panel that detects a user's touch, a button that receives a user's push operation, a wheel that receives a user's rotation operation, a keyboard, and a dome switch, although not limited thereto.

[0119] Also, the input interface 1500 may include a voice recognition device (not shown) for voice recognition. For example, the voice recognition device may be a microphone (1200), and the voice recognition device may receive a user's voice command or a user's voice request. Accordingly, the processor 202 may perform a control of performing an operation corresponding to a voice command or a voice request.

[0120] The memory 204 may store various information, data, instructions, programs, etc. required for operations of the electronic device 101. The memory 204 may include at least one of a volatile memory or a non-volatile memory, or a combination thereof. The memory 204 may include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, card type memory (for example, Secure Digital (SD) memory or eXtreme Digital (XD) memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, a magnetic disk, or an optical disk. Also, the electronic device 101 may manage a web storage or a cloud server that performs a storage function on the Internet.

[0121] The communication module 1300 may transmit / receive information to / from an external apparatus or an external server according to a protocol under a control by the processor 202. The communication module 1300 may include at least one communication module and at least one port that transmits / receives data to / from an external apparatus (not shown).

[0122] Also, the communication module 1300 may perform communication with an external apparatus through at least one wired or wireless communication network. The communication module 1300 may include at least one of a short-range communication module 1310 or a mobile communication module 1320 or a combination thereof. The communication module 1300 may include at least one antenna for communicating with another apparatus in a wireless manner.

[0123] The short-range communication module 1310 may include at least one communication module (not shown) that performs communication according to communication standards, such as Bluetooth, Wireless Fidelity (Wi-Fi), BLE, NFC / Radio Frequency Identification (RFID), Wifi Direct, UWB, or ZIGBEE. Also, the mobile communication module 1320 may include a communication module that performs communication through a network for internet communication. Also, the mobile communication module 1320 may include a mobile communication module that performs communication according to communication standards, such as 3-Generation (3G), 4-Generation (4G), 5-Generation (5G), and / or 6-Generation (6G).

[0124] Also, the communication module 1300 may include a communication module capable of receiving a control command from a remote controller (not shown) located within a short distance, for example, an infrared (IR) communication module.

[0125] The sensor 1700 may include various kinds of sensors. For example, the sensor 1700 may include various kinds of sensors, such as an image sensor, an infrared sensor, an ultrasonic sensor, a LIDAR sensor, a human detection sensor, a motion detection sensor, a proximity sensor, and an illumination sensor. Because a function of each sensor may be intuitively deduced by a person skilled in the art from the name, detailed descriptions thereof will be omitted.

[0126] The processor 202 may control overall operations of the electronic device 101. The processor 202 may execute a program stored in the memory 204 to control components of the electronic device 101.

[0127] In an embodiment, the processor 202 includes at least one processor.

[0128] The at least one processor may receive an audio input, associated with an audio conversation, having verbal audio signal and non-verbal audio signal.

[0129] The at least one processor may determine a plurality of conversation characteristics corresponding to each speaker based on the verbal audio signal and the non-verbal audio signal.

[0130] The at least one processor may determine conversation tonal attributes corresponding to each speaker based on the determined conversation characteristics.

[0131] The at least one processor may determine relationship among the speakers based on the determined conversation tonal attributes.

[0132] The at least one processor may generate actionable insights, for each speaker, based on the determined relationship.

[0133] The at least one processor may separate the verbal audio signal and the non-verbal audio signal from the received audio signal.

[0134] The at least one processor may segment each of the separated verbal audio signal and the separated non-verbal audio signal.

[0135] The at least one processor may generate verbal audio clusters and non-verbal audio clusters by clustering each of the segmented verbal audio signal and the segmented non-verbal audio signal. Each cluster of the verbal audio clusters is indicative of an audio utterance of one of the speakers and each cluster of the non-verbal audio clusters is indicative of a noise associated with one of the speakers.

[0136] The at least one processor may map the verbal audio clusters with the non-verbal audio clusters.

[0137] The at least one processor may determine non-conflicting overlapping clusters and conflicting overlapping clusters based on the mapping of the clustered verbal audio signal and the clustered non-verbal audio signal.

[0138] The at least one processor may verify speaker corresponding to each conflicting overlapping cluster and each non-conflicting overlapping cluster.

[0139] The at least one processor may generate modified verbal audio clusters based on the verified speakers.

[0140] The at least one processor may generate textual data based on the generated verbal audio clusters.

[0141] The at least one processor may determine a plurality of conversation characteristics corresponding to each speaker based on the generated verbal audio clusters and the generated textual data. The plurality of conversation characteristics comprises frequency spectrum, loudness, conversation speed, characteristic word vectors, and sentiment.

[0142] The at least one processor may determine, using a multi-output Neural Network (NN), conversation tonal attributes corresponding to each speaker based on the determined conversation characteristics. The plurality of conversation tonal attributes comprises formal tone, informal tone, engaged tone, disengaged tone, concerned tone, unconcerned tone, dominant tone, and subordinate tone.

[0143] The at least one processor may determine, using a multi-class Machine Learning (ML) model, the relationship among the speakers of the audio conversation based on the determined conversation tonal attributes. The relationship among the speakers comprises family relationship, professional relationship, formal relationship, romantic relationship, and friendly relationship.

[0144] The at least one processor may generate, using generative AI model, actionable insights based on the determined relationship and the determined conversation tonal attributes. Each actionable insights comprises a contextual element indicative of a context associated with one or more insights for each speaker.

[0145] In an embodiment, the audio conversation includes a call conversation between a user of an electronic device (101) and the other party in a phone call. The at least one processor control to the display 1610 to display the generated actionable insights and the determined relationship between the user of the electronic device (101) and the other party.

[0146] Machine-readable storage media may be provided in the form of non-transitory storage media. Herein, 'non-transitory storage media' means that the storage media do not include a signal and current and are tangible, without meaning that data is semi-permanently or temporarily stored in the storage media.

[0147] According to an embodiment, the method according to various embodiments disclosed in the present document may be included in a computer program product and provided. The computer program product may be traded between a seller and a purchaser. The computer program product may be distributed in the form of a machine-readable storage medium (for example, compact disc read only memory (CD-ROM)), or be distributed (for example, downloadable or uploadable) online via an application store or between two user devices (for example, smart phones) directly. When distributed online, at least part of the computer program product (for example, downloadable app) may be temporarily generated or at least temporarily stored in the machine-readable storage medium, such as a memory of the manufacturer's server, a server of the application store, or a relay server.

[0148] The term "module" or "unit (or portion)" used in various embodiments herein may include a unit implemented by hardware, software, or firmware, and may be used interchangeably with a term such as logic, a logic block, a part, or a circuit. The module may be an integrated part or be a minimum unit or portion of the part, which performs one or more functions. For example, according to an embodiment of the disclosure, the module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0149] While specific language has been used to describe the present disclosure, any limitations arising on account thereto, are not intended. As would be apparent to a person in the art, various working modifications may be made to the method in order to implement the inventive concept as taught herein. The drawings and the foregoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment.

Claims

1.A method (800) for generating actionable insights for multi-speaker audio conversation, the method (800) comprising:receiving an audio input, associated with an audio conversation, having verbal audio signal and non-verbal audio signal;determining a plurality of conversation characteristics corresponding to each speaker based on the verbal audio signal and the non-verbal audio signal;determining conversation tonal attributes corresponding to each speaker based on the determined conversation characteristics;determining relationship among the speakers based on the determined conversation tonal attributes; andgenerating actionable insights, for each speaker, based on the determined relationship.2.The method (800) according to claim 1 further comprising:separating the verbal audio signal and the non-verbal audio signal from the received audio signal;segmenting each of the separated verbal audio signal and the separated non-verbal audio signal; andgenerating verbal audio clusters and non-verbal audio clusters by clustering each of the segmented verbal audio signal and the segmented non-verbal audio signal, wherein each cluster of the verbal audio clusters is indicative of an audio utterance of one of the speakers and each cluster of the non-verbal audio clusters is indicative of a noise associated with one of the speakers.3.The method (800) according to claim 2 further comprising:mapping the verbal audio clusters with the non-verbal audio clusters;determining non-conflicting overlapping clusters and conflicting overlapping clusters based on the mapping of the clustered verbal audio signal and the clustered non-verbal audio signal; andverifying speaker corresponding to each conflicting overlapping cluster and each non-conflicting overlapping cluster; andgenerating modified verbal audio clusters based on the verified speakers.4.The method (800) of any one of claims 1 to 3, wherein determining the plurality of conversation characteristics corresponding to each speaker comprises:generating textual data based on the generated verbal audio clusters; anddetermining a plurality of conversation characteristics corresponding to each speaker based on the generated verbal audio clusters and the generated textual data,wherein the plurality of conversation characteristics comprises frequency spectrum, loudness, conversation speed, characteristic word vectors, and sentiment.5.The method (800) of any one of claims 1 to 4, wherein determining conversation tonal attributes corresponding to each speaker comprises:determining, using a multi-output Neural Network (NN), conversation tonal attributes corresponding to each speaker based on the determined conversation characteristics,wherein the plurality of conversation tonal attributes comprises formal tone, informal tone, engaged tone, disengaged tone, concerned tone, unconcerned tone, dominant tone, and subordinate tone.6.The method (800) of any one of claims 1 to 5, wherein determining the relationship among the speakers of the audio conversation comprises:determining, using a multi-class Machine Leaning (ML) model, the relationship among the speakers of the audio conversation based on the determined conversation tonal attributes,wherein the relationship among the speakers comprises family relationship, professional relationship, formal relationship, romantic relationship, and friendly relationship.7.The method (800) of any one of claims 1 to 6, wherein generating actionable insights for each speaker comprises:generating, using generative AI model, actionable insights based on the determined relationship and the determined conversation tonal attributes,wherein each actionable insight comprises a contextual element indicative of a context associated with one or more insights for each speaker.8.A system (102) for generating actionable insights for multi-speaker audio conversation, the system (102) comprising:at least one processor configured to:receive an audio input, associated with an audio conversation, having verbal audio signal and non-verbal audio signal;determine a plurality of conversation characteristics corresponding to each speaker based on the verbal audio signal and the non-verbal audio signal;determine conversation tonal attributes corresponding to each speaker based on the determined conversation characteristics;determine relationship among the speakers based on the determined conversation tonal attributes; andgenerate actionable insights, for each speaker, based on the determined relationship.9.The system (102) according to claim 8, the at least one processor is configured to:separate the verbal audio signal and the non-verbal audio signal from the received audio signal;segment each of the separated verbal audio signal and the separated non-verbal audio signal; andgenerate verbal audio clusters and non-verbal audio clusters by clustering each of the segmented verbal audio signal and the segmented non-verbal audio signal, wherein each cluster of the verbal audio clusters is indicative of an audio utterance of one of the speakers and each cluster of the non-verbal audio clusters is indicative of a noise associated with one of the speakers.10.The system (102) according to claim 9, wherein the at least one processor is configured to:map the verbal audio clusters with the non-verbal audio clusters;determine non-conflicting overlapping clusters and conflicting overlapping clusters based on the mapping of the clustered verbal audio signal and the clustered non-verbal audio signal; andverify speaker corresponding to each conflicting overlapping cluster and each non-conflicting overlapping cluster; andgenerate modified verbal audio clusters based on the verified speakers.11.The system (102) of any one of claims 8 to 10, wherein to determine the plurality of conversation characteristics corresponding to each speaker, the at least one processor is configured to:generate textual data based on the generated verbal audio clusters; anddetermine a plurality of conversation characteristics corresponding to each speaker based on the generated verbal audio clusters and the generated textual data,wherein the plurality of conversation characteristics comprises frequency spectrum, loudness, conversation speed, characteristic word vectors, and sentiment.12.The system (102) of any one of claims 8 to 11, wherein to determine conversation tonal attributes corresponding to each speaker, the at least one processor is configured to:determine, using a multi-output Neural Network (NN), conversation tonal attributes corresponding to each speaker based on the determined conversation characteristics,wherein the plurality of conversation tonal attributes comprises formal tone, informal tone, engaged tone, disengaged tone, concerned tone, unconcerned tone, dominant tone, and subordinate tone.13.The system (102) of any one of claims 8 to 12, wherein to determine the relationship among the speakers of the audio conversation, the at least one processor is configured to:determine, using a multi-class Machine Learning (ML) model, the relationship among the speakers of the audio conversation based on the determined conversation tonal attributes,wherein the relationship among the speakers comprises family relationship, professional relationship, formal relationship, romantic relationship, and friendly relationship.14.The system (102) of any one of claims 8 to 13, wherein to generate actionable insights for each speaker, the at least one processor is configured to:generate, using generative AI model, actionable insights based on the determined relationship and the determined conversation tonal attributes,wherein each actionable insights comprises a contextual element indicative of a context associated with one or more insights for each speaker.15.The method (800) of any one of claims 1 to 7, wherein the audio conversation includes a call conversation between a user of an electronic device (101) and the other party in a phone call, further comprising:displaying, through a display of the electronic device (101), the generated actionable insights and the determined relationship between the user of the electronic device (101) and the other party.