System and method for processing text, image, and audio signals using artificial intelligence

The system processes audio and image signals to generate analytical data, addressing the limitations of conventional methods by providing accurate sentiment measurement and real-time guidance for improving conversational behavior.

JP2026516008APending Publication Date: 2026-05-19KAI CONVERSATIONS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
KAI CONVERSATIONS LTD
Filing Date
2024-05-02
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Conventional methods for analyzing remote conversations lack depth and specificity, are subjective, and fail to combine emotional text analysis with facial emotion analysis to provide accurate sentiment measurement and feedback on interpersonal interactions.

Method used

A system and method for simultaneously processing audio and image signals using artificial intelligence algorithms to generate analytical data, including sentiment measurement, by analyzing facial expressions, body language, and speech patterns to identify decision-making points and emotional sincerity.

Benefits of technology

Provides accurate and comprehensive emotional analysis of conversations, enabling real-time guidance for improving conversational behavior and identifying critical moments that influence discussion outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026516008000001_ABST
    Figure 2026516008000001_ABST
Patent Text Reader

Abstract

A system (100,200) is disclosed that processes at least simultaneously acquired audio signals (AS) and image signals (IS) to generate corresponding analytical data, including emotion measurement. The system (100,200) comprises a computing configuration (102) configured to include at least an audio processing module (APM,104) and an image processing module (IPM,106) for processing the audio signals (AS) and image signals (IS). Each module (104,106) employs an artificial intelligence (AI) algorithm. The image processing module (106) is configured to process facial image information contained in the image signal (IS) to identify key facial image points indicating facial expressions and generate temporal facial state data (TFSD), and to identify multiple key body image points indicating body language and generate temporal body language state data (TBLSD). The speech processing module (104) analyzes utterances in the speech signal (AS), compares the utterances with a word database (108) to generate text data (TD), processes the utterances to determine temporal speech frequency information (TSFI), and is configured to temporally associate the TSFI with the text data (TD). The computing configuration (102) further comprises an analysis module (110) that uses an AI algorithm, processes TFSD, TBLSD, TD and TSFI using an emotion model, generates interpretations of the speech signal (AS) and image signal (IS), and generates analysis data including emotion measurement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a system that processes text, image, and audio signals to generate corresponding processed data, where the processed data includes analysis data including emotion measurement. Further, the present disclosure relates to a method of processing text, image, and audio signals using the system to generate corresponding processed data, where the processed data includes analysis data including emotion measurement. Moreover, the present disclosure relates to a software product executable on computing hardware for implementing the system, where the software product uses an algorithm for implementing the method when executed on computing hardware.

Background Art

[0002] In today's world, communication has become an essential element in our daily lives. Usually, when working from home, communicating with friends and family, or conducting business around the world, we rely on remote conversations to maintain connections. In particular, these remote conversations contain valuable insights such as the language used during the conversation, the tone of voice, and non-verbal cues.

[0003] The analysis of recorded speech is well known. Conventional methods include processing the speech signal using a parser to generate corresponding text. In the process of converting the speech signal to corresponding text, certain information contained in the speech signal is removed and not reflected in the corresponding text. For example, information indicating the pitch of the voice, speech rate, and other intonation and nuances is not transmitted to the corresponding text via the parser. The generated text can then be analyzed using AI engines based on deep neural networks, hidden Markov models (HMMs), etc., to extract its substantive meaning. Furthermore, the analysis of images obtained from such communication is also well known and is used, for example, in facial recognition. However, parts of the face change when making facial expressions. These changes are known to be easily detected by facial recognition. In addition, especially since the COVID-19 pandemic, the frequency of use of video conferencing tools has increased as people have been forced to work from home and communicate with others through video calls. Such video calls allow people to not only hear the voices of others but also see them through video. Furthermore, wearable devices such as VR headsets, goggles, and glasses can also provide video and audio streams in face-to-face environments.

[0004] However, when a specific observer manually analyzes a conversation, the sample size is limited, and the results depend on the observer's skills and interpretation. This means the observer may overlook important aspects of the conversation, and bias may arise from applying subjective interpretations to specific expressions that occur during the conversation. Appropriately, emotional intelligence is a crucial element for effective communication and obtaining positive outcomes from human interaction. Understanding the flow of emotions among interacting participants over time allows for an understanding of the factors that cause positive and negative interactions. Visualizing these emotional flows allows participants to better recognize the positive and negative behaviors and patterns that give rise to them. However, conventional conversation analysis methods rely on a specific person observing the conversation and drawing their own opinions and conclusions. Such methods are labor-intensive, expensive, time-consuming, and often subjective and inconsistent.

[0005] Conventional AI-based conversational intelligence solutions have two main limitations: (i) While it is possible to point out general issues (i.e., emotional tendencies) or search for occurrences of specific words or phrases (e.g., words indicating sadness), (ii) The conversation cannot be analyzed with sufficient depth or specificity to be actually useful for a particular individual (i.e., accurate sentiment measurement cannot be performed).

[0006] Furthermore, existing findings are general and lack accuracy, specificity, and practicality. None of the known solutions can combine emotional text analysis and facial emotion analysis to provide feedback on the impact an evaluator has on others. Nor can any known solution provide deeper conversational analysis that measures how emotions develop over time and what positive outcomes can be obtained based on such interpersonal interaction patterns.

[0007] Therefore, based on the above discussion, it is necessary to overcome the aforementioned shortcomings that arise when analyzing verbal and visual interactions that occur between multiple individuals configured to interact with one another. [Overview of the project] [Means for solving the problem]

[0008] This disclosure provides a system for processing at least simultaneously acquired audio and image signals to generate corresponding analytical data, including sentiment measurement. Furthermore, this disclosure provides a method for processing at least audio and image signals using the system to generate corresponding analytical data, including sentiment measurement. This disclosure presents a solution to the problems of the prior art, namely, a system for efficiently processing at least simultaneously acquired audio and image signals and a method for using the system. The object of this disclosure is to provide an improved system for processing at least simultaneously acquired audio and image signals, at least partially overcoming the aforementioned problems in the prior art. Furthermore, the object of this disclosure is to provide an improved method for processing at least audio and image signals using the system to generate corresponding analytical data, including sentiment measurement.

[0009] One or more of the objects of this disclosure are achieved by the solutions described in the appended independent claims. Advantageous embodiments of this disclosure are further defined in the dependent claims.

[0010] In a first embodiment, the Disclosure provides a system for processing at least simultaneously acquired audio and image signals to generate corresponding analytical data, including emotion measurement, comprising a computing configuration configured to include at least an audio processing module for processing audio signals and an image processing module for processing image signals. Each module is configured to use one or more artificial intelligence algorithms to process their respective signals. The image processing module is configured to process facial image and body language information contained in the image signal to identify multiple key facial image points indicating facial expressions and generate temporal facial state data, as well as to identify multiple key body image points indicating body language and generate temporal body language state data. The speech processing module is configured to process speech contained in a speech signal, analyze the speech and match it with a word database to generate corresponding text data, and process the speech to determine temporal speech frequency information indicating at least one of emphasis, hesitation, and speech rate, and to temporally associate this temporal speech frequency information with the text data. Furthermore, the computing configuration includes an analysis module that uses one or more artificial intelligence algorithms, which is configured to process temporal facial state data, temporal body language state data, text data, and temporal speech frequency information using an emotion model, generate interpretations of speech signals and image signals, and generate analysis data including emotion measurement.

[0011] This system uses a computing architecture to process at least simultaneously acquired audio and video signals and generate corresponding analytical data, including sentiment measurement. By analyzing audio and video signals, this system makes it possible to uncover latent human insights that arise during a particular conversation. Furthermore, the depth and accuracy of the analysis, when combined with a conversation model, can derive more specific insights. In addition, by eliminating noise contained in simultaneously acquired audio and video signals, this system provides highly relevant analytical results for multiple participants in a particular conversation. Moreover, this system can be configured to perform real-time analysis, allowing participants engaged in conversations with other participants to obtain guidance for adjusting their conversational behavior in real time, for example, by providing real-time guidance for gaining a better negotiating position in an argument.

[0012] In one embodiment, the image processing module is configured to process an image signal that includes video data acquired simultaneously with an audio signal.

[0013] In this embodiment, the system is designed to process image signals, particularly those containing video data acquired simultaneously with audio signals. Specifically, it can perform various image processing operations on video data of a particular individual to extract primary emotional information, and perform various processing operations on simultaneously acquired audio data corresponding to the same individual to extract secondary emotional information. By synchronizing the corresponding audio and video signals and the primary and secondary emotional information associated with them, more accurate and effective analysis can be achieved. For example, if the audio signal contains emotional information indicating sadness while the video signal contains emotional information indicating joy, analyzing this discrepancy between the primary and secondary emotional information can provide an index for evaluating the emotional sincerity of the individual.

[0014] In yet another optional embodiment, the system is configured to process a text signal in addition to at least audio and image signals. In this case, the information contained in the text signal is used in combination with the information contained in the audio and image signals to generate corresponding analytical data, including sentiment measurement.

[0015] This system is configured to process text signals together with audio and video signals, integrate this information, and generate analytical data, including sentiment measurement. This makes it possible to analyze information about the communication being analyzed in a more comprehensive way, including subtle nuances. For example, in video calls where audio and video signals are transmitted and received, it is common for documents to be discussed to be distributed among participants before the call, and it is useful to analyze these documents in advance before the video call takes place.

[0016] In yet another embodiment, the one or more artificial intelligence algorithms include at least one of a neural network, a deep neural network, a Boltzmann machine, or a hidden Markov model for processing at least audio and image signals.

[0017] In such any embodiment, the one or more artificial intelligence algorithms have the ability to process large amounts of data and extract meaningful information from audio and image signals. This allows them to be used in a variety of applications, such as emotion recognition, speech analysis, facial expression analysis, and body language analysis. Furthermore, these algorithms can adaptively improve their performance over time through learning with additional data, possessing high flexibility and enabling them to handle different types of audio and image signals. The technical benefits of using these algorithms include improved accuracy and efficiency in processing audio and image signals, and improved system performance and reliability when generating emotion measurement and analysis data from input signals.

[0018] In yet another embodiment, the analysis module is configured to identify decision-making points that occur during a video discussion generating audio and video signals, using at least sentiment measurement.

[0019] In such embodiments, the analysis module uses sentiment measurements derived from audio and video signals to identify decision-making points that occur during a video discussion. Preferably, the decision-making points, i.e., abrupt changes in sentiment measurements over time, allow the system to gain insight into critical moments during a discussion that may influence the outcome or direction of the discussion.

[0020] In yet another optional embodiment, the analysis module is configured to identify decision points occurring during a video discussion generating audio and video signals, using at least sentiment measurement, which are identified by the analysis module based on at least one of temporal abruptions in sentiment measurement or temporal abruptions in the content of the audio signals.

[0021] The analysis module is configured to identify decision-making points that occur during video discussions by using sentiment measurement and detecting sudden changes in sentiment measurement or speech content in the audio signal. Through such analysis, the analysis module can gain insights into critical moments in video discussions that may influence the outcome or direction of the conversation.

[0022] In a second embodiment, the Disclosure provides a method for training a system of the first embodiment, the method comprising the following steps: (i) The process of constructing a first corpus of training material related to training values ​​for sentiment measurement on samples of audio signals containing speech information, (ii) A step of constructing a second corpus of training material related to training values ​​for emotion measurement on a sample of image signal containing facial information, and (iii) The step of applying the first and second corpora of training materials to one or more artificial intelligence algorithms to construct analytical characteristics for processing at least audio and video signals.

[0023] In such embodiments, the performance and accuracy of one or more algorithms that process audio and video signals and generate sentiment measurements based on the characteristics and features of the input signals are improved, for example, by learning from training materials and being optimized.

[0024] In a third aspect, the present disclosure provides a method for processing at least an audio signal and an image signal using the system of the first aspect to generate corresponding analysis data including emotion measurement. The system includes a computing structure configured to include an audio processing module for processing at least an audio signal and an image processing module for processing an image signal, and each module is configured to use one or more artificial intelligence algorithms to process its respective signal. The method includes the following steps. Using the image processing module, processing the face image information included in the image signal, identifying a plurality of main face image points indicating expressions to generate temporal face state data, and further identifying a plurality of main body image points indicating body gauges to generate temporal body gauge data; Using the audio processing module, analyzing the utterances included in the audio signal, collating with a word database to generate corresponding text data, processing the audio to determine temporal audio frequency information indicating at least one of emphasis, hesitation, and speech rate, and temporally associating the temporal audio frequency information with the text data; and A step of using the analysis module of the computing structure, the analysis module including one or more artificial intelligence algorithms, using an emotion model to process the temporal face state data, the temporal body gauge state data, the text data, and the temporal audio frequency information, generating an interpretation of the audio signal and the image signal, and generating analysis data including emotion measurement. The method realizes all the advantages and technical effects of the system of the present disclosure.

[0025] In a fourth aspect, the present disclosure provides a software product executable on computing hardware for implementing the methods of the second and third aspects.

[0026] It should be understood that all embodiments described above can be used in combination with one another. It should be understood that all devices, elements, circuits, units, and means described herein can be implemented by software elements, hardware elements, or any combination thereof. Furthermore, the processes and functions performed by each component described herein mean that the component is adapted or configured to perform the respective process and function. In addition, even if a particular function or process performed by a particular external component is not explicitly shown in the detailed description of that component in the description of the specific embodiments below, it will be obvious to those skilled in the art that these methods and functions can be implemented by the corresponding software elements, hardware elements, or any combination thereof. It should be understood that the features of this disclosure are applicable in various combinations without departing from the scope of this disclosure as defined by the appended claims.

[0027] Other aspects, advantages, features, and objectives of this disclosure will become apparent from the accompanying drawings and the detailed description of the embodiments, which will be interpreted in connection with the claims set forth below. [Brief explanation of the drawing]

[0028] The above summary and the detailed description of the embodiments shown below will be better understood when read in conjunction with the accompanying drawings. For the purpose of illustrating this disclosure, the drawings show exemplary configurations of this disclosure. However, this disclosure is not limited to the specific methods or means disclosed herein. Furthermore, those skilled in the art will understand that the drawings are not necessarily drawn to scale. Also, identical elements are numbered together where possible.

[0029] The embodiments of this disclosure will be described below with reference to the following drawings, which are merely illustrative examples. [Figure 1-2]Figures 1A, 1B, and 2 are schematic diagrams of a system according to one embodiment of the present disclosure that processes text, image, and audio signals (e.g., at least an image signal and an audio signal) to generate corresponding analytical information, including sentiment measurement. [Figure 3] Figure 3 is a diagram showing an example of a flowchart according to one embodiment of the present disclosure, which includes steps for training the system shown in Figures 1 and 2. [Figure 4] Figure 4 is a diagram showing an example of a flowchart according to one embodiment of the present disclosure, which includes steps for using the system shown in Figures 1 and 2.

[0030] In the attached drawings, underlined numbers are used to indicate the item to which that number is located, or an item adjacent to that number. Ununderlined numbers are associated with the item identified by the line connecting the number and the item. Furthermore, if an arrow is attached to an ununderlined number, that number is used to identify the general item pointed to by the arrow. [Modes for carrying out the invention]

[0031] The following detailed description illustrates embodiments of the Disclosure and forms for carrying them out. While several embodiments of the Disclosure are disclosed, those skilled in the art will understand that other embodiments for carrying out or applying the Disclosure are also possible.

[0032] Figures 1A, 1B, and 2 are schematic diagrams of a system 100 according to one embodiment of the present disclosure, which processes at least an audio signal and an image signal to generate corresponding analytical data, including emotion measurement. Referring to Figures 1A and 1B, the system 100 comprises a computing configuration 102, an audio processing module 104, an image processing module 106, a database 108, and an analysis module 110.

[0033] System 100 includes a computing configuration 102 that corresponds to a form of configuring and managing computing resources to efficiently perform a specific task. Optionally, the computing resources include hardware, software, and network resources that work together to process and analyze the participant's voice and image signals and generate corresponding analysis data. Furthermore, the participant may be a person (i.e., a human) or a virtual program (e.g., an autonomous program or bot) associated with or operating a user device. The user device is an electronic device associated with (or used by) a participant that has the functionality to enable the participant to engage in conversation. The term "user device" is also interpreted broadly to include any electronic device capable of voice and video communication over a wired or wireless network. Examples of user devices include, but are not limited to, mobile phones, personal digital assistants (PDAs), handheld devices, wireless modems, notebook computers, personal computers, and wearable devices.

[0034] It should be understood that the computing configuration 102 consists of three layers: the application layer, the middleware layer, and the hardware layer. Typically, the application layer includes software such as word processing software and email applications. The middleware layer provides a bridge between the hardware layer and the application layer, managing communication between different software and enabling multiple software to work together. The hardware layer includes components such as servers, storage devices, and network infrastructure, which are used to run software and process data.

[0035] Appropriately, with respect to the audio signal, the computing configuration 102 includes software for speech recognition, speech noise reduction, and speech enhancement, as well as hardware such as a microphone and a sound card. The middleware layer may include speech-to-text functionality that converts the audio signal into transcribed text, and natural language processing algorithms that extract meaning from the transcribed text. The hardware layer may include servers or groups of servers capable of processing large amounts of audio data in real time. These servers or groups of servers may be configured as array processors that perform large-scale parallel computing.

[0036] Similarly, with respect to image signals, the computing configuration 102 includes software for performing image recognition, image enhancement, and computer vision, as well as hardware such as cameras and graphics processing units (GPUs). The middleware layer may include algorithms for performing object detection, face recognition, or optical character recognition (OCR), which can identify and analyze features within images. The hardware layer may also include servers or groups of servers capable of processing large amounts of image data in real time. These servers or groups of servers may be configured as array processors for performing large-scale parallel computing.

[0037] System 100 includes an audio processing module 104. "Audio processing module" refers to software or hardware components used to manipulate audio signals. The audio processing module 104 is used for audio editing, speech recognition, noise reduction, and speech enhancement. Optionally, the audio processing module 104 may include a filter module. This filter module is used to remove unwanted noise or interference from the audio signal, such as hum or hiss. Examples of filter modules may include one or more high-pass filters, low-pass filters, and notch filters. These filters may also be configured to adapt to the time-varying spectral characteristics of the audio signal, for example, to address voice fatigue that occurs during video calls.

[0038] In one embodiment, the speech processing module 104 further includes a tone analyzer. The tone analyzer operates by analyzing the tone of the speech signal based on the use of specific words or phrases and the structure of sentences. It should be understood that machine learning algorithms are further used to determine the tone of the entire text, i.e., whether it is positive, negative, or neutral, and to identify specific emotions or sentiments contained therein. In another embodiment, the speech processing module 104 further includes a sentiment analyzer. The sentiment analyzer operates by analyzing the emotional tendencies of the speech signal. Overall, the speech processing module 104 contributes to achieving desired sound quality and clarity by manipulating and emphasizing the speech signal.

[0039] System 100 includes an image processing module 106. “Image processing module” refers to a module used to acquire images of participants. Optionally, the image processing module 106 can acquire video of participants. Furthermore, the image processing module 106 can acquire each frame of video acquired by an imaging device. Optionally, the imaging device may be a camera mounted on a device used for conversation. For example, in a video conference with two participants using laptops, the audio signal is recorded by a microphone, and the video signal is acquired by the laptop's camera. Notably, the image processing module 106 is designed to process a video signal consisting of a sequence of images played continuously at a constant frame rate. The video signal may include visual content such as scenes, objects, or events acquired by a camera or other imaging device. As shown in the diagram, the system simultaneously captures both audio and video signals associated with the participants.

[0040] In one embodiment, the image processing module 106 is configured to process an image signal that includes video data acquired simultaneously with an audio signal. Notably, the image processing module 106 is designed to process the image signal acquired by the image processing module 106 and the audio signal recorded by the audio processing module 104 simultaneously. "Acquired simultaneously" means that the image signal and the audio signal are recorded or acquired at the same time, typically from the same device. Such simultaneous acquisition refers to situations in which audio and video are recorded simultaneously, such as video recording, live streaming, or video conferencing.

[0041] By processing audio and image signals simultaneously, it becomes possible to analyze and manipulate both modalities synchronously, thereby obtaining highly accurate processing results. For example, in video conferencing, the image processing module 106 can analyze the image signal to detect facial expressions, body language, gestures, and other visual cues, while simultaneously processing the audio signal to detect speech and other auditory features. Through such synchronous processing, the system 100 can accurately acquire and analyze both audio and image information, improving overall quality and effectiveness.

[0042] System 100 includes a database 108, which refers to a storage medium for storing words. In one embodiment, the database 108 includes a dictionary containing multiple words. The speech processing module 104 may be configured to analyze utterances in a speech signal, match the utterances with words stored in the database 108 to generate corresponding text data, and process the utterances to determine temporal speech frequency information indicating at least one of emphasis, hesitation, and speech rate, and to temporally associate the temporal speech frequency information with the text data. Optionally, the database 108 may include, but is not limited to, internal storage devices, external storage devices, Universal Serial Bus (USB), hard disk drives (HDDs), flash memory, Secure Digital (SD) cards, solid-state drives (SSDs), computer-readable storage media, or a suitable combination thereof.

[0043] "Temporal speech frequency information" refers to the change or fluctuation of speech frequency over time. More precisely, temporal speech frequency information includes analyzing and characterizing the frequency components of speech as they change or transition over different time intervals. Speech is composed of multiple frequency components corresponding to different phonemes or utterances, and these frequency components change rapidly over time as an utterance is produced and transitions to the next sound occur. Temporal speech frequency information captures the dynamic characteristics of speech and provides insights into the temporally fluctuating spectral characteristics of speech. Temporal speech frequency information can provide important information about speech signals, such as prosody, intonation, and speech rate. For example, temporal speech frequency information can reveal pitch or frequency fluctuations in a speech signal, thereby indicating information about a participant's emotion, emphasis, or linguistic meaning.

[0044] System 100 includes an analysis module 110. The "analysis module" 110 refers to a module that analyzes and interprets audio and image signals using an artificial intelligence algorithm. Specifically, the analysis module 110 is configured to analyze temporal facial state data, temporal body language data, text data, and temporal audio frequency information using an emotion model. The analysis module 110 is configured to generate analysis data, including emotion measurements.

[0045] Appropriately, “temporal facial state data” refers to information about facial expressions, movements, or other changes acquired over time. Similarly, “temporal body language state data” refers to information about body language, movements, postures, gestures, or other changes in the body acquired over time. Optionally, the analysis module 110 can use computer vision technology to track and analyze facial features and movements, as well as body language features and movements, thereby providing insights into the emotional state of the participant being analyzed. For example, the analysis module 110 can detect facial expressions of participants indicating different emotions, such as smiles, frowns, and eyebrow raising and lowering. Similarly, text data includes transcripts of spoken words, captions, or other text data. Optionally, the analysis module 110 can analyze text data using natural language processing technology. Furthermore, the aforementioned temporal speech frequency information is also used to generate analysis data, including sentiment measurement. “Sentiment measurement” refers to quantitative or qualitative measurements or indicators of emotion extracted or derived from the analysis of speech and image signals.

[0046] Please understand that the analysis module 110 is configured as follows: (i) Analyzing the audio and video signals of participants during a video conference, and (ii) Tracking temporal facial state data, text data, and temporal voice frequency information.

[0047] By tracking participants' audio and video signals on user devices, such tracking helps monitor and analyze interactions between at least two participants during a video conference. This tracking allows for the identification of patterns and tendencies in each participant's conversation and the detection of potential problems and misunderstandings that may arise during the conversation. In particular, this analysis focuses primarily on each participant's expressive performance.

[0048] Furthermore, by tracking participants on user devices, the analytics module 110 can monitor and analyze interactions when participants share their screens during video conferences, and identify patterns and trends. For example, the analytics module 110 can be configured to analyze specific customers (first participant) and specific advisors (second participant) during a meeting. "Screen sharing" typically refers to the process of one participant sharing the contents of their screen with other participants during a meeting (e.g., a video conference) for the purpose of presenting information. When screen sharing occurs during a video conference, it becomes possible to analyze valuable insights such as satisfaction, interaction time, and overall engagement from the audio and video signals of participants in the conversation. For example, if a customer concentrates on a shared presentation for an extended period, it may suggest that the customer is particularly interested in the shared content.

[0049] In one embodiment, the system 100 is configured to process text signals in addition to at least audio and image signals, and the information contained in the text signals is used in combination with the information contained in the audio and image signals to generate corresponding analytical data, including sentiment measurement. In this regard, “text signals” means any character or textual information that is part of the communication or data processed by the system 100. Text signals may include speech transcripts, text-based chat messages, captions, or other forms of text data associated with the audio and image signals being analyzed.

[0050] System 100 is configured to process text signals in addition to audio and image signals in order to generate analytical data. The information contained in the text signals is used in combination with the information contained in the audio and image signals to generate analytical data, including sentiment measurement. By integrating data from different sources, System 100 achieves accurate and comprehensive analysis of sentiment. Preferably, by integrating information obtained from text, audio, and image signals, System 100 can generate analytical data that allows for a more holistic and deeper understanding of the emotional aspects of communication.

[0051] In one embodiment, the one or more artificial intelligence algorithms include at least one of a neural network, a deep neural network, a Boltzmann machine, or a hidden Markov model for processing at least audio and image signals. In this regard, the analysis module 110 processes temporal facial state data, text data, and temporal speech frequency information using one or more artificial intelligence algorithms. Optionally, the one or more artificial intelligence algorithms may include at least one of a neural network, a deep neural network, a Boltzmann machine, or a hidden Markov model trained or designed to generate interpretations of audio and image signals. Further optionally, the analysis module 110 may be pre-trained or customized to specific emotional contexts such as happiness, sadness, anger, or surprise, and may use pattern recognition, statistical analysis, or other techniques to identify emotional cues from the temporal facial state data, text data, and temporal speech frequency information. Based on the output generated by the analysis module 110, the analysis results represent the emotional state or intensity of the participant in the audio and image signals, and can be used in a variety of applications such as emotion recognition, affective computing, virtual reality (VR), or human-computer interaction (HCI).

[0052] In one embodiment, the analysis module 110 is configured to identify decision-making points that occur during a video discussion generating audio and video signals, using at least sentiment measurement. In this regard, the analysis module 110 uses at least sentiment measurement of participants (such as emotional state, emotional intensity, or emotional expression) to identify decision-making points that occur during the video discussion. A “decision-making point” refers to a specific moment or event during the video discussion in which an important decision is made or an important action is taken. Typically, decision-making points include moments in which participants express strong emotions, provide important information, make important statements, or engage in important interactions that influence the overall outcome or direction of the conversation. For example, if sentiment measurement indicates heightened excitement, frustration, or disagreement among participants, it may indicate a decision-making point where emotions are heightened and an important decision is being made. Optionally, by identifying decision-making points during a video discussion, the analysis module 110 can provide insights or clues that require attention within the discussion.

[0053] In one embodiment, the analysis module 110 is configured to identify decision points occurring during a video discussion generating audio and video signals, using at least sentiment measurement, which are identified by the analysis module 110 based on at least one of a temporal abruption in sentiment measurement or a temporal abruption in the content of the audio signal utterances. In this regard, the analysis module 110 uses sentiment measurement to identify decision points occurring during a video discussion. Typically, decision points are determined based on abrupt changes in sentiment measurement or the content of the audio signal utterances. The analysis module 110 identifies decision points by detecting temporal abruptions in sentiment measurement, using at least sentiment measurement. This may include processing to detect sudden and significant changes or fluctuations in the emotional state, intensity of emotion, or emotional expression of participants during a video conference. For example, if sentiment measurement shows a sudden rise, such as a change from a calm state to an excited state, it may suggest that it is a decision point where emotions are heightened and an important discussion or action is taking place.

[0054] Referring to Figure 2, a system 200 is shown that processes text, image, and audio signals (e.g., image and audio signals) to generate corresponding analytical information, including sentiment measurement, in real time. In the illustrated system 200, the data flow is in real time. The application layout includes multiple icons and visual elements representing different analytical components of the conversation analysis system. These icons are strategically placed within a graphical user interface (GUI) and are visually associated with corresponding parts of the conversation transcript. A meeting is initiated by participants, and participants are invited to the meeting. Appropriately, the conversation is presented along with the video and audio signals. The transcript is color-coded and highlights different stages of the conversation, automatically identified by one or more AI algorithms based on the topic of the sentences. The stages of the conversation are represented by different labels such as "Introduction," "Information Gathering," and "Resolution," and are visually associated with corresponding parts of the transcript. One or more AI algorithms are configured to perform sentiment analysis, linguistic analysis (tone, confidence level, speech style, etc.), engagement analysis, and outcome analysis. The scoring and analysis components provide insights into how the conversation took place, how it was received, and what results were obtained.

[0055] As shown in the diagram, system 200 captures the participant's audio and video signals. These audio and video signals provide system 200 with information including voice tone, facial expressions, and body language. By combining and analyzing this information, a deeper understanding of the conversation can be achieved.

[0056] The participants' audio and video signals are provided for time-series data generation. "Time-series data generation" refers to the process of generating a series of data points arranged in chronological order. This system utilizes one or more AI algorithms, such as acoustic analysis, word analysis, and visual analysis, integrating them seamlessly. For example, the system uses the audio processing module 104 to acquire the tone, pitch, and volume of the participant's voice, performs real-time transcription and analysis of the speech content using word analysis, and further acquires facial expressions and body language using the image processing module 106. Combining these elements enables a holistic and multidimensional analysis of the conversation, providing more accurate insights and revealing subtle nuances that might be overlooked in individual analyses. In this regard, the application monitors the quality of the audio and video signals, as well as screen-sharing content. Furthermore, this data is sent to a server for real-time or near-real-time monitoring.

[0057] System 200 further enables the generation of specific conversation models by identifying structured subsets of information and interaction states from conversations. Such pre-configured conversation models allow for targeted analysis based on predefined parameters and objectives, leading to more actionable insights. For example, the system can be programmed to identify specific keywords or phrases, affect cues, or conversation patterns, which can then be used to generate real-time responses tailored to the conversation model. System 200 utilizes real-time or near-real-time data to construct "real-time moments" and "real-time prompts." A "real-time moment" refers to a moment when participants are simultaneously engaged with the same content. These real-time moments stimulate activity on the platform and increase the likelihood of users actively engaging with the content. A "real-time prompt," on the other hand, refers to a function designed to encourage participants to engage with the content in real time.

[0058] Real-time responses generated by the system are displayed to participants and suggest actions they can take to improve the outcome of the conversation. These actionable suggestions can take various forms, from subtle prompts encouraging adjustments to tone and speaking speed to clearer recommendations on how to address specific issues or achieve desired outcomes. Real-time prompts enable participants to actively engage in the conversation and make informed decisions, thereby leading to more effective communication and better conversation outcomes. System 200 includes a knowledge base database, which is a repository of information and data available for decision support, problem solving, learning, and content delivery to participants. This knowledge base database works in conjunction with the real-time prompts. The knowledge base database includes text documents, images, videos, etc. Furthermore, the knowledge base database is systematically organized in an easily viewable and searchable structure to provide participants with up-to-date, accurate, and relevant information tailored to their needs.

[0059] This system also enables the introduction of human labeling and feedback loops. These feedback loops allow for continuous improvement of the conversational model through machine learning. Human evaluators can provide feedback on the accuracy of the system's responses and identify areas for improvement. This feedback is used to train one or more AI algorithms, refining the system's performance over time and increasing the accuracy and effectiveness of real-time responses. Furthermore, the system includes filter prompts to select available content based on participant feedback and preferences.

[0060] System 200 identifies key points in a conversation, known as "Moments-that-Matter (MTM)," where specific actions or interventions can significantly impact the outcome. Using algorithms and data analysis techniques, System 200 automatically identifies these key points, enabling timely and accurate responses with reproducibility and scalability. Furthermore, System 200 allows for the analysis of specific types of conversations based on a defined data structure, facilitating collective analysis to identify the most effective interaction patterns that can lead to the best possible conversation outcomes. Such conversation analysis enables data-driven insights and evidence-based decision-making, contributing to improved conversation outcomes.

[0061] Furthermore, System 200 utilizes sentiment analysis to perform the conversation analysis process. By acquiring and analyzing participants' emotions, System 200 enables participants to become more aware of their own emotions and their impact on others. This mode of operation promotes improved emotional intelligence, enhancing participants' ability to better control their emotions and helping to achieve more empathetic and effective communication. Preferably, the meeting concludes after feedback has been received. System 200 also enables post-conversation analysis in a consistent and systematic manner. Such post-conversation analysis allows participants to reflect on their expressive performance and identify areas for improvement. System 200 can provide detailed insights into conversation patterns, sentiment cues, and the overall effectiveness of the conversation. Moreover, such post-conversation analysis enables participants to enhance self-awareness and consciously adjust their communication style to achieve better emotional outcomes in future conversations. The conversation data is also stored in a conversation database, and the application is configured to display the stored conversations in the conversation view database tab.

[0062] Figure 3 is a flowchart showing the steps of method 300 for training system 200. The method includes steps 302 to 306. Step 302 includes the step of method 300 constructing a first corpus of training material related to training values ​​of sentiment measurements for samples of speech signals containing speech information. In this regard, “first corpus (training material)” refers to the set of data used as input data for generating training values ​​of sentiment measurements. The training material of the first corpus generally consists of combinations of samples of speech signals containing utterance data and their corresponding sentiment measurements. Sentiment measurements may include at least one of the following: a quantitative measure of emotion, a qualitative measure of emotion (e.g., emotion labels such as “happiness,” “sadness,” or “anger”), or an emotion intensity score. “Training values ​​of sentiment measurements” refers to known sentiment measurements corresponding to speech signal samples in the training material. These training values ​​are used as reference data for learning the relationship between speech signals and their corresponding sentiment measurements.

[0063] Samples of speech signals containing speech information refer to specific examples or instances of speech signals containing speech data that constitute part of the training material. These samples may optionally be other forms of speech signals containing speech information, such as recordings of participants' speech, or transcripts, conversation data, or other speech-related data. The “construction” process of the first corpus (training material) involves collecting, selecting, and preparing samples of speech signals and their corresponding sentiment measures to create a comprehensive dataset that can be used to train System 200. This construction process may include collecting speech data from various sources, annotating or labeling sentiment measures, and organizing the data into a format suitable for model training. Once the training material of the first corpus is constructed, it functions to represent the relationship between speech signals and sentiment measures based on the provided training values. In one embodiment, the model thus trained can be used to analyze novel or unknown speech signals in real time and to automatically estimate or predict the sentiment content of speech data.

[0064] Step 304 of Method 300 includes the step of constructing a second corpus of training material related to training values ​​of sentiment measurements for samples of image signals containing facial expression information. Constructing this second corpus of training material refers to the process of collecting and organizing training data that associates sentiment measurements with samples of image signals, and in particular aims to create a second corpus of training material that includes data related to facial expression information. This training material is used to train an artificial intelligence algorithm or model to analyze and interpret facial expression information in image signals and generate sentiment measurements. The second corpus of training material functions as a dataset providing examples of image signals associated with known sentiment measurements, enabling the algorithm or model to learn and generalize from this data to accurately interpret facial expression information in image signals and generate sentiment measurements.

[0065] Step 306 includes the step of applying the first and second corpora of training material to one or more artificial intelligence algorithms to configure analytical characteristics for processing at least audio and video signals. The first corpus of training material includes training values ​​of sentiment measurements associated with audio signal samples, and the second corpus of training material includes training values ​​of sentiment measurements associated with video signal samples. By applying the first and second corpora of training material to one or more artificial intelligence algorithms, the algorithms are configured to adaptively adjust or configure analytical characteristics for processing audio and video signals.

[0066] Training one or more artificial intelligence algorithms enables pattern recognition, learning from relationships, and inference based on sentiment measurements in training data, allowing for accurate interpretation of audio and video signals and the generation of corresponding sentiment measurements in the analyzed data. Advantageously, the performance and accuracy of one or more artificial intelligence algorithms are optimized using first and second corpora of training materials, improving their ability to process audio and video signals based on the characteristics and features of the input signals and generate sentiment measurements.

[0067] Steps 302 to 306 are illustrative examples only, and other alternative configurations, i.e., adding, removing, or performing one or more steps in a different order, are possible without departing from the claims herein.

[0068] Figure 4 shows a flowchart illustrating a method for processing at least audio and image signals using System 200 and generating corresponding analytical data, including emotion measurement. The method includes steps 402 to 406. Step 402 includes the step of using an image processing module 106 to process facial image information contained in an image signal and identify a plurality of key facial image points indicating facial expressions to generate temporal facial state data. Preferably, the image processing module 106 identifies a plurality of key facial image points indicating facial expressions, such as the corners of the eyes, corners of the mouth, and eyebrow positions, and generates temporal facial state data. This data captures changes in facial expressions over time and provides insights into the emotional state of an individual in an image. The facial state data can be obtained using techniques such as facial landmark detection, facial feature extraction, or facial expression recognition.

[0069] Step 404 of Method 400 includes the step of using a speech processing module 104 to analyze utterances contained in a speech signal, compare the utterances with a word database to generate corresponding text data, process the utterances to determine temporal speech frequency information indicating at least one of emphasis, hesitation, and speech rate, and temporally associate the temporal speech frequency information with the text data. The speech processing module 104 analyzes the utterances using techniques such as speech recognition and natural language processing and generates corresponding text data by comparing them with a word database. Furthermore, the speech processing module 104 can analyze the speech to determine temporal speech frequency information indicating emphasis, hesitation, speech rate, etc. The temporal speech frequency information is temporally associated with the text data to provide a synchronous representation of the utterance content. In one embodiment, the image signal includes video data acquired simultaneously with the speech signal.

[0070] In step 406, method 400 includes a step using an analysis module 110 of the computing configuration 102, the analysis module 110 including one or more artificial intelligence algorithms, processing temporal facial state data, text data and temporal speech frequency information using an emotion model, generating interpretations of speech signals and image signals, and generating analysis data including emotion measurement. The analysis module 110 utilizes one or more artificial intelligence algorithms to process temporal facial state data, text data and temporal speech frequency information using an emotion model. The one or more artificial intelligence algorithms include at least one of a neural network, a deep neural network, a Boltzmann machine and a hidden Markov model for processing at least speech signals and image signals.

[0071] The emotion model is designed to interpret audio and video signals and generate emotion measurements, thereby providing insights into the emotional content of the processed signals. The emotion model may be trained using a corpus of training material related to training values ​​of emotion measurements for samples of audio signals containing speech information. Emotion measurements may include quantitative or qualitative indicators of emotion, such as emotion labels (e.g., "happy," "sad," "angry"), emotion intensity scores, or other relevant emotion parameters. The generated analysis data, including emotion measurements, can be used to comprehensively understand the emotional aspects of audio and video signals. For example, the analysis module 110 can use emotion measurements to identify decision points that occur during a video discussion, such as moments of heightened emotion intensity or sudden shifts in emotion. Such decision points provide useful insights for further analysis, such as sentiment analysis, emotion recognition, or behavioral analysis.

[0072] Steps 402 to 406 are illustrative examples only, and other alternative configurations, i.e., adding, removing, or performing one or more steps in a different order, are possible without departing from the claims herein.

[0073] A software product executable on computing hardware for carrying out the methods relating to the second and third embodiments.

[0074] The software product may be implemented as an algorithm when executed on computing hardware, stored in a non-temporary machine-readable data storage medium, and configured to perform methods 300 and 400. The software may be stored in a non-temporary machine-readable data storage medium, which includes, but is not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or appropriate combinations thereof. Examples of computer-readable media include, but are not limited to, electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), read-only memory (ROM), hard disk drives (HDDs), flash memory, secure digital (SD) cards, solid state drives (SSDs), computer-readable storage media, and CPU cache memory.

[0075] The foregoing regarding embodiments of this disclosure can be modified in various ways without departing from the scope of this disclosure as defined in the appended claims. Expressions such as “including,” “comprising,” “incorporating,” “have,” and “is” used in the description and claims of this disclosure should be interpreted as non-restrictive, allowing for the existence of items, components, or elements not explicitly mentioned, rather than as exclusive. Furthermore, singular descriptions should be interpreted as including plural forms depending on the context. The term “exemplary” as used herein means “as an example, embodiment, or description.” Therefore, an “exemplary” embodiment is not said to be preferable or advantageous to other embodiments, nor does it preclude combining features of other embodiments. Furthermore, the term “optionally” means “provided in one embodiment but not in other embodiments.” Certain features of this disclosure described in the context of a separate embodiment for clarity may also be provided in combination in a single embodiment. Conversely, various features of the present invention described in the context of a single embodiment for the sake of brevity may also be provided separately, in any suitable combination, or applied in any of the other embodiments of the present disclosure.

Claims

1. A system (100, 200) for processing at least simultaneously acquired audio and image signals to generate corresponding analytical data including emotion measurement, wherein the system comprises a computing configuration (102) configured to include at least an audio processing module (104) for processing audio signals and an image processing module (106) for processing image signals, and each of the modules (104, 106) is configured to use one or more artificial intelligence algorithms to process their respective signals. The image processing module (106) is configured to process facial image information contained in the image signal to identify a plurality of main facial image points indicating facial expressions and generate temporal facial state data, and to identify a plurality of main body image points indicating body language and generate temporal body language state data. The aforementioned speech processing module (104) is configured to analyze utterances in a speech signal, compare those utterances with a word database (108) to generate corresponding text data, process the utterances to determine temporal speech frequency information indicating at least one of emphasis, hesitation, and speech rate, and to temporally associate said temporal speech frequency information with the text data. Furthermore, the computing configuration is characterized by comprising an analysis module (110) that uses one or more artificial intelligence algorithms to process the temporal facial state data, temporal body language state data, text data, and temporal voice frequency information based on an emotion model, generates interpretations of voice signals and image signals, and generates analysis data including emotion measurement.

2. The system (100, 200) according to claim 1, wherein the image processing module (106) is configured to process an image signal including video data acquired simultaneously with an audio signal.

3. The system (100, 200) according to claim 1 or 2, wherein the system (100, 200) is configured to process a text signal in addition to an audio signal and an image signal, and the information contained in the text signal is used in combination with the information contained in the audio signal and the image signal to generate corresponding analytical data including emotion measurement.

4. The system according to claim 1, 2, or 3 (100, 200), wherein the one or more artificial intelligence algorithms used to process at least the audio signal and image signal include at least one of a neural network, a deep neural network, a Boltzmann machine, and a hidden Markov model.

5. The system (100, 200) according to any one of claims 1 to 4, wherein the analysis module (110) is configured to use at least sentiment measurement to identify decision-making points that occur during a video discussion generating the aforementioned audio and image signals.

6. The analysis module (110) is configured to use at least sentiment measurement to identify decision-making points that occur during a video discussion that generates the aforementioned audio and image signals. The system according to any one of claims 1 to 4 (100, 200), wherein the decision-making point is identified by the analysis module based on at least one of a temporal abruption in emotion measurement or a temporal abruption in the content of speech signals.

7. A method (300) for training a system (100, 200) according to any one of claims 1 to 6, (i) A step of constructing a first corpus of training material related to training values ​​for sentiment measurement on samples of speech signals containing speech information, (ii) A step of constructing a second corpus of training material related to training values ​​for emotion measurement on a sample of image signal containing facial expression information, (iii) A step of applying the first and second corpora of training material to one or more artificial intelligence algorithms to configure analytical characteristics for processing at least the audio signal and video signal, A method (300) characterized by including the following.

8. A method (400) of using a system (100, 200) to process at least audio and image signals to generate corresponding analytical data including emotion measurement, wherein the system (100, 200) comprises a computing configuration configured to include at least an audio processing module (104) for processing audio signals and an image processing module (106) for processing image signals, and each of the modules (104, 106) is configured to use one or more artificial intelligence algorithms to process their respective signals. The above method is characterized by including the following steps (400): The process involves using the image processing module (106) to process the facial image information contained in the image signal to identify a plurality of main facial image points indicating facial expressions and generate temporal facial state data, and to identify a plurality of main body image points indicating body language and generate temporal body language state data. The process involves using the speech processing module (104) to analyze the utterances contained in the aforementioned speech signal, to compare the utterances with a word database to generate corresponding text data, to process the utterances to determine temporal speech frequency information indicating at least one of emphasis, hesitation, and speech rate, and to temporally associate the temporal speech frequency information with the text data. The process involves using an analysis module (110) of the computing configuration, which includes one or more artificial intelligence algorithms that process temporal facial state data, temporal body language state data, text data, and temporal speech frequency information using an emotion model, and which generates an interpretation of the speech signal and image signal to generate the analysis data, including the emotion measurement.

9. The method (400) according to claim 8, wherein the method (400) includes the step of configuring an image processing module (106) to process the image signal, which includes video data acquired simultaneously with the audio signal.

10. The method (400) according to claim 8 or 9, wherein the method (400) is configured to process a text signal in addition to the audio signal and the image signal, and the information contained in the text signal is used in combination with the information contained in the audio signal and the image signal to generate the corresponding analysis data, including the emotion measurement.

11. The method (400) according to claim 8, 9, or 10, wherein the method (400) includes the step of applying one or more artificial intelligence algorithms, which include at least one of a neural network, a deep neural network, a Boltzmann machine, and a hidden Markov model, to process at least the audio signal and the image signal.

12. The method (400) according to any one of claims 8 to 11, wherein the method (400) includes the step of configuring the analysis module (110) to use at least the sentiment measurement to identify decision points occurring during a video discussion that generates the audio and image signals.

13. The method (400) according to any one of claims 8 to 12, wherein the method (400) includes the step of configuring an analysis module (110) to use at least the sentiment measurement to identify decision points occurring during a video discussion that generates the audio and video signals, the decision points being identified by the analysis module based on at least one of a temporal abruption in the sentiment measurement or a temporal abruption in the content of the utterances in the audio signals.

14. A software product executable on computing hardware for carrying out the method (300, 400) according to any one of claims 7 to 13.