System for analyzing the Emotional Quotient using multimodal deep learning on audio-video inputs

DE202025104181U1Active Publication Date: 2025-10-16DESAI SHARMISHTA DR PUNE +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE202025104181
Authority / Receiving Office
DE · DE
Patent Type
Utility models
Current Assignee / Owner
Filing Date
2025-07-20
Publication Date
2025-10-16
Estimated Expiration
2035-07-31

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A system (100) for analyzing emotional quotient (EQ) using multimodal deep learning on audio-video inputs, comprising: a computing device (102) having a processor (104) and a memory (106) configured to store one or more instructions executable by the processor (104); wherein the processor (104) is configured to execute a plurality of modules (108) to process synchronized video and audio data from natural conversations, extract multimodal behavioral characteristics with deep learning models trained with curriculum learning strategies, and calculate standardized EQ component scores and a composite EQ score; wherein the computing device (102) is connected to a database (122), user devices (124), cameras (126) and microphones (128) via a network (120); wherein the plurality of modules (108) comprises: a data acquisition module (110) configured to acquire and synchronize real-time video data via the cameras (126) and audio data via the microphones (128); a preprocessing module (112) configured to transcribe the captured audio data using a speech recognition engine including the Whisper model and to match word-based timestamps with the captured video data; a feature extraction module (114) configured to extract multimodal behavioral features, including facial expressions, speech patterns, gaze behavior, nodding frequency, emotional correctness, emotional entropy, and intensity variation, by applying deep learning models with curriculum learning strategies; an assessment module (116) configured to apply quantifiable and scalable assessment logics in real time, including emotional entropy assessment and synchronization of emotions between listener and speaker, to calculate standardized EQ component scores corresponding to the domains of self-awareness, social awareness, relationship management, and self-management; an output module (118) configured to map the EQ component values ​​to a standardized scale of 1 to 5, aggregate the component values ​​into a composite EQ value in the range of 4 to 20, store the results in the database (122), and transmit the values ​​to the user devices (124) via a user interface (130).
Need to check novelty before this filing date? Find Prior Art

Description

Field of the invention

[0001] The present disclosure relates generally to the technical field of automated emotional intelligence assessment, and more specifically to a system that analyzes the Emotional Quotient (EQ) using multimodal deep learning on audio-video inputs. The system implements quantifiable and scalable real-time assessment logic, audio-video fusion, emotional entropy assessment, and speaker-listener synchronization to objectively measure emotional intelligence. Background of the invention

[0002] Emotional intelligence (EQ) significantly influences success in personal relationships, professional growth, and leadership skills. Conventional EQ measurements are based on subjective questionnaires—but they can be manipulated and often do not reflect actual behavior in real emotional situations. Furthermore, they rarely capture multimodal behavioral cues such as facial expressions, speech patterns, eye contact, or nodding.

[0003] Existing computer-based systems for automatic EQ prediction are mostly theoretical in nature. While the state of the art suggests the use of eye tracking or pitch detection, they lack technical evaluation logic, formulas for entropy or correctness, and real-time fusion pipelines that deliver reproducible EQ values.

[0004] Furthermore, there is a lack of dynamic synchronization between the emotional timelines of speaker and listener, models for aligning these emotions, eye contact assessments, or nodding detection with computer vision. The state of the art is based solely on proposed weightings, without scalable and validated evaluation approaches.

[0005] Therefore, no modular software solutions exist that capture and process real-time audio-video data from natural conversations to generate standardized EQ scores. The lack of implementation of automation-capable deep learning models with curriculum learning and validated assessment logic limits scalability, accuracy, and practical relevance.

[0006] There is therefore a need for a system that analyzes EQ using multimodal deep learning on audio-video data. It must be able to capture measurable behavioral traits—e.g., facial expressions, speech patterns, eye contact, nodding frequency, emotional accuracy, intensity variation, and reaction time—and convert these into standardized component scores and an overall EQ using clinically validated weightings.

[0007] Equally required is a system that processes video and audio data from conversations using commercially available cameras and microphones in real time, uses deep learning models with curriculum learning, and validates the computational logic against established psychological frameworks – applicable across industries in recruitment, education, therapy, and human resource development. Summary of the invention

[0008] The present disclosure proposes a system for analyzing the Emotional Quotient (EQ) using multimodal deep learning on audio-video data. This summary is intended to provide an overview and serve as an introduction to the detailed description.

[0009] The invention solves existing problems through quantifiable, scalable real-time scoring logic, audio-video fusion, emotional entropy assessment, and speaker-listener synchronization.

[0010] One embodiment describes a system for EQ analysis based on multimodal deep learning processing of audio and video data. It includes a computing device, a network, a database, end-user devices, and cameras and microphones.

[0011] The computing device has a processor and memory with executable instructions. It is connected to a database, cameras, microphones, and end devices via the network.

[0012] The processor executes modules for synchronizing video and audio streams, extracting multimodal features using curriculum learning deep learning models, and calculating standardized EQ component scores and a composite EQ.

[0013] The data acquisition module synchronizes live video (cameras) and audio (microphones). The preprocessing module transcribes speech using Whisper technology and aligns word timestamps with video frames.

[0014] The feature extraction module uses deep learning to capture facial expressions, speech patterns, eye contact, nodding frequency, emotional accuracy, entropy, and intensity variation. It detects emotions in speakers and listeners and quantifies trust and intensity.

[0015] The eye contact score measures the percentage of gaze directed toward the microphone / camera (eye / iris landmarks). Nod detection is based on vertical head movement across connected frames.

[0016] The scoring module calculates real-time scores via emotional entropy, synchronization, and correctness assessment, standardized on components such as self-awareness, social awareness, relationship management, and self-regulation.

[0017] Entropy measurement (positive vs. negative emotions) results in a quantitative value, combined with clinical weightings to produce standardized component scores.

[0018] The output module scales the components (1-5), sums them to the EQ (4-20), stores the data in the database, and sends the results to user devices. The scores are validated against psychological reference tests (acceptable mean deviation).

[0019] The advantages and objects of the invention will become apparent from the following statements, claims and drawings. DETAILED DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings, which form a part of this specification, illustrate an embodiment of the invention and, together with the description, explain the principles of the invention. Fig. 1 shows a block diagram of a system for analyzing emotional quotient (EQ) using multimodal deep learning on audio and video inputs according to an exemplary embodiment of the invention. Fig. 2 shows a flowchart of a workflow for analyzing the emotional quotient (EQ) using multimodal deep learning on audio-video inputs by the system according to an exemplary embodiment of the invention. DETAILED DISCLOSURE OF THE INVENTION

[0021] Various embodiments of the present invention will be described with reference to the accompanying drawings. Where possible, the same or similar reference numerals are used throughout the drawings and the description to refer to the same or similar parts or steps.

[0022] The present disclosure was developed with a view to solving the problems described in the prior art. The objective of the invention is to provide a system that analyzes the emotional quotient (EQ) using multimodal deep learning methods based on audio and video inputs. Quantifiable and scalable real-time scoring logic, audio-video fusion, emotional entropy calculation, and speaker-listener synchronization are used to objectively measure emotional intelligence.

[0023] Fig.Figure 1 shows a block diagram of a system 100 for analyzing EQ using multimodal deep learning on audio-video inputs. In one embodiment, the system 100 analyzes EQ using audio-video inputs, using quantifiable and scalable real-time scoring logic, audio-video fusion, entropy calculations, and speaker-listener synchronization to objectively assess emotional intelligence. The system 100 includes a computing device 102, a network 120, a database 122, user devices 124, cameras 126, and microphones 128.

[0024] In one embodiment, computing device 102 includes a processor 104 and memory 106 storing one or more instructions executable by processor 104. Upon executing these instructions, the processor causes system 100 to acquire and synchronize audio-video data, extract multimodal behavioral traits, apply deep learning models with curriculum learning strategies, and calculate EQ component scores and an overall EQ value. Processor 104 acts as the central processing unit (CPU) of system 100, coordinating data processing, model execution, and decision making by retrieving and executing instructions from memory 106.

[0025] In one embodiment, memory 106 serves as the storage unit of system 100, storing executable instructions as well as data required for EQ analysis. This data may include extracted features such as facial expressions, speech patterns, gaze behavior, nodding frequency, emotional accuracy, emotional entropy, intensity variation, and user-specific configuration or context data. The interaction between processor 104 and memory 106 enables system 100 to process synchronized audio-video inputs, execute deep learning models, and calculate EQ values ​​in real time.

[0026] In one embodiment, computing device 102 represents an electronic device used to deploy and operate system 100. Computing device 102 may be, among other things, a personal computer, server, laptop, tablet, or any suitable computing device capable of executing AI models and managing data input / output operations. Users can interact with system 100 via computing device 102 and user devices 124 to utilize EQ analysis features.

[0027] In one embodiment, user interface 130 is an integral part of computing device 102, allowing users to enter commands, configure settings, and display calculated EQ component scores and overall results. User interface 130 may include, among other things, a graphical interface, a touchscreen, a keyboard, a mouse, speech recognition modules, or other input and display means. This versatility allows users such as human resources professionals, educators, therapists, or researchers to intuitively interact with system 100 and gain real-time insights into emotional intelligence.

[0028] In one embodiment, users include administrators and analysts. Administrators use computing device 102 to manage and configure system 100 and to respond to analysts' analysis requests. Analysts, in turn, use user devices 124 to interact with system 100 and utilize emotional quotient (EQ) analysis features.

[0029] In one embodiment, computing device 102 is connected to cameras 126 and microphones 128 for capturing synchronized audio-video data, to database 122 for storing extracted features and calculated EQ values, and to user devices 124 for displaying the results—all via network 120. Network 120 enables seamless transmission of data, results, and user input, allowing system 100 to operate in a distributed and real-time manner.

[0030] In one embodiment, network 120 may include, but is not limited to, wireless networks, local area networks (LANs), wide area networks (WANs), cellular networks, virtual private networks (VPNs), or any other communications infrastructure that enables data exchange between computing device 102, database 122, user devices 124, cameras 126, and microphones 128. The flexibility of network 120 ensures that system 100 can operate in diverse environments and remains accessible across locations and devices.

[0031] In one embodiment, processor 104 is configured to execute a plurality of modules 108 to process synchronized video and audio data from natural conversations, extract multimodal behavioral features using deep learning models trained with curriculum learning strategies, and calculate standardized EQ component scores and a composite EQ total score. The plurality of modules 108 includes a data acquisition module 110, a preprocessing module 112, a feature extraction module 114, a scoring module 116, and an output module 118.

[0032] In one embodiment, the data acquisition module 110 is configured to capture and synchronize real-time video data using the cameras 126 and audio data using the microphones 128. In one embodiment, the preprocessing module 112 is configured to transcribe the captured audio data using a speech recognition engine such as the Whisper model and to match word timestamps with the captured video data at the word level.

[0033] In one embodiment, the feature extraction module 114 is configured to extract multimodal behavioral features including facial expressions, speech patterns, gaze behavior, nodding frequency, emotional accuracy, emotional entropy, and intensity variation using deep learning models trained with curriculum learning strategies. The feature extraction module 114 detects speaker and listener emotions, predicts confidence scores, and calculates intensity variation using the trained models.

[0034] The feature extraction module 114 calculates gaze behavior by tracking eye and iris landmarks, calculating normalized gaze coordinates, and determining the proportion of frames in which gaze is directed toward the microphone 128 or camera 126. Nodding is detected by tracking vertical head movements against facial landmarks, calculating the normalized vertical distance across consecutive frames, and counting the number of nods per minute based on a dynamic threshold.

[0035] The evaluation module 116 is configured to apply quantifiable and scalable real-time evaluation logic, including entropy assessments and synchronization of speaker and listener emotions, to calculate standardized EQ component scores in the areas of self-awareness, social awareness, relationship management, and self-management. Emotional correctness is calculated by comparing the detected speaker and listener emotions with context-dependent expected responses and calculating correctness scores based on timestamp alignment.

[0036] The evaluation module 116 calculates emotional entropy by classifying detected emotions into positive and negative classes and applying an entropy formula to their distribution. Furthermore, it applies clinically validated weights to convert the extracted features into standardized EQ component scores.

[0037] The output module 118 is configured to map the EQ component scores to a standardized scale of 1 to 5, aggregate the component scores into a composite EQ score ranging from 4 to 20, store the results in the database 122, and transmit the scores to the user devices 124 via the user interface 130. The EQ component scores, as well as the overall EQ score, are stored in the database 126 to enable validation, auditing, and historical analysis, and are output to user devices 130 or external psychological assessment systems. The composite EQ score and the component scores are compared to an established psychological test to confirm reliability and an acceptable mean deviation.

[0038] In one embodiment, the system 100 analyzes the emotional quotient (EQ) using multimodal deep learning on audio-video inputs. The system 100 assesses the EQ based on four main components: self-awareness, social awareness, relationship management, and self-management. Each component is rated on a scale of 1 to 5. The final EQ score is the sum of these four components and thus ranges from 4 to 20. The selection of the assessed factors and the determination of the threshold ranges are based on the recommendations of an experienced clinical psychologist. The proposed system 100 was evaluated using data from 31 participants who also completed the publicly available Global Emotional Intelligence Test based on Mayer and Salovey's four-branch model. The results generated by the system 100 were compared with those of the test.According to the clinical psychologist's assessment, System 100 showed an acceptable mean deviation, thus confirming its reliability and effectiveness.

[0039] In one embodiment, the computational framework for evaluating the EQ follows the formula: Emotional Quotient(EQ)=SA+SCA+RM+SM, where SA stands for self-awareness, SCA for social awareness, RM for relationship management and SM for self-management.

[0040] In one embodiment, the Self-Awareness (SA) component is measured using the following weighted formula: SA=0.1×self-confidence+0.1×patience+0.2×expressiveness+0.6×entropy, Self-confidence is trained using a curriculum learning method. Patience is calculated as the average of speaking rate, mean pause time, and reaction time to initial questions.

[0041] In one embodiment, speaking rate is extracted from the transcript generated by the Whisper model by calculating the number of words per minute. In one embodiment, Table 1 shows the mapping of word-per-minute (wpm) ranges to assessment scores. Table 1: wpm area Score 120-150 5 150-160 4 160-180 3 180-250 2 250-400 1

[0042] In one embodiment, the mean pause time is calculated by measuring the average duration between the timestamps of consecutive words in the transcript. Table 2 shows the correlation of the mean pause time (t in seconds) to the scores: Table 2: Average break time Score 0.2-0.3 s 5 0.15-0.2 s or 0.3-0.35 s 4 0.1-0.15 s or 0.35-0.5 s 3 0.05-0.1 s or 0.5-0.75 s 2 <0.05 s or >0.75 s 1

[0043] In one embodiment, the initial reaction time measures the delay between the end of the speaker's speech and the listener's first reaction (nod, word, or emotion), extracted via timestamps. In one embodiment, Table 3 shows the mapping of the initial reaction time t (in seconds) to the scores: Table 3: response time Score 0.1-0.3 s 5 0.08-0.1 s or 0.3-0.35 s 4 0.05-0.08 s or 0.35-0.5 s 3 0-0.05 s 2 >0.7 s 1

[0044] In one embodiment, the expressiveness (EOE) is the average of the speaker's expressiveness (SEoE) and the listener's expressiveness (LEoE).

[0045] In one embodiment, the SEoE is calculated as follows: SEoE=0.1×number of different emotions (De)+0.4×correctness rating (Cs)+0.3×smoothness of emotion transition (Ts)+0.2×variation of emotion intensity (Iv).

[0046] Table 4 shows the assignment of the normalized SEoE values ​​to a scale of 1 to 5: Table 4: Normalized SEOE Score interpretation 0.81-1.00 5 Extremely natural expression 0.61-0.80 4 Mostly, of course, with minor discrepancies 0.41-0.60 3 Moderate abrupt transitions 0.21-0.40 2 Noticeably unnatural expression 0.00-0.20 1 Highly enforced or robotic

[0047] In one embodiment herein, the correctness score (Cs) is calculated by first extracting emotions from the transcript context using the Facebook / Bart-Large-Mnli model and extracting emotions from the video using curriculum learning, as shown in Table 5. Table 5: Extracted context Mapped emotions achievement, compliment, success, praise Luck Unexpected beauty, wonder, inspiration awe Encouragement, support, reassurance sympathy Entertainment, playfulness, humor Luck Loss, failure, disappointment sadness Frustration, injustice, criticism, danger, threat, insecurity AngerFear Disgust, moral outrage, aversion disgust Exchange of information, explanations, facts Neutral Ask questions, seek clarity Neutral Unexpected result, shock Surprise Concern for the pain of others, empathy sympathy sympathy

[0048] In one embodiment, the extracted context words, such as "success," "praise," or "achievement," are mapped to emotions such as joy, awe, compassion, etc. according to a predefined mapping. The correctness score (Cs) is then calculated as the ratio of correctly mapped emotions to the total number of expressed emotions.

[0049] In one embodiment, the smoothness of the emotion transition (Ts) is calculated by extracting timestamps of the detected emotions and calculating the average time to transition between them (Tavg) using an exponential decay function: Ts=e(−k⋅Tavg), where k = 0.2 to derive the final smoothing value that reflects the naturalness of the emotional flow.

[0050] In one embodiment, the variation in emotion intensity (Iv) is calculated from the predicted emotion intensity values ​​obtained through deep learning. The mean (µi) and standard deviation (σi) are used to calculate the coefficient of variation (CV = σi / µi). The smoothness or stability of the emotion intensity is then given by: Iv=1−CV.

[0051] In one embodiment, the listener expressiveness (LEoE) is calculated as follows: LEoE=0.1×number of different emotions (De)+0.4×correctness value (Cs)+0.3×smoothness of emotion transition (Ts)+0.2×reaction time (Rt).

[0052] In one embodiment, Table 6 shows the assignment of the normalized LEoE values ​​to the scale from 1 to 5: Table 6: Normalized SEOE Score interpretation 0.81-1.00 5 Consistently coordinated, smooth and timely responses 0.61-0.80 4 Mostly natural with minor deviations 0.41-0.60 3 Some incorrect or delayed reactions 0.21-0.40 2 Frequent mismatches or unnatural shifts 0.00-0.20 1 Severely delayed or forced reactions

[0053] In one embodiment, the emotions are extracted from the video using curriculum learning. Count the distinct emotions (Ds) of the extracted emotions. In one embodiment, the speaker's emotions are extracted from the video using curriculum learning, and the corresponding timestamps are extracted. In one embodiment, the listener's emotions are extracted from the video using curriculum learning, and the corresponding timestamps are extracted. The timestamps of the speaker and listener videos are matched with a permissible delay of 5 seconds. The speaker and listener emotions are assigned using Table 7. Table 7: Speaker emotion Expected listener reaction Happy Happy, Awe, Neutral Sad Compassion, Sad, Neutral Angry Neutral, Sad, Compassionate Fearful Compassion, Neutral, Reverence Disgusted Neutral, Disgusted, Sad Surprised Awe, Happy, Neutral Neutral Neutral, Happy, Compassionate sympathy Compassion, Neutral, Sad awe Awe, Happy, Neutral

[0054] In one embodiment, the correctness value (Cs) is calculated as the ratio of correctly assigned emotions to the total number of expressed emotions.

[0055] In one embodiment, the smoothness of the emotion transition (Ts) is calculated by extracting timestamps of the detected emotions and calculating the average time (Tavg) for the transition between them using the exponential decay function: Ts=e(−k⋅Tavg), where k = 0.2. This formula is used to determine the final smoothing value that reflects the naturalness of the emotional transition.

[0056] In one embodiment, the speaker's emotions are extracted from the video using curriculum learning, as are the corresponding timestamps.

[0057] In one embodiment, the listener's emotions are also extracted from the video using curriculum learning, as well as the corresponding timestamps.

[0058] The delay between the timestamps of the speaker's and the listener's emotions is calculated. The average delay (ΔTavg) between the speaker's and the listener's timestamps is determined. In one embodiment, the listener's reaction time is calculated using the following exponential decay function: Rt=e(−α⋅(ΔTavg)) where ΔTavg represents the average delay between the speaker's emotion and the listener's response and α = 0.3.

[0059] In one embodiment, entropy is calculated by classifying the detected emotions into positive (e.g., joy, surprise) and negative (e.g., sadness, anger, fear, disgust, contempt) classes. Using the count values ​​P (positive) and N (negative), the entropy H is calculated as follows: H=−(P / (P+N))⋅log2(P / (P+N))−(N / (P+N))⋅log2(N / (P+N)).

[0060] In one embodiment, Table 8 shows the assignment of the entropy values ​​H to the evaluation range from 1 to 5: Table 8: Entropy H Score 0.75-0.85 5 0.70-0.75 or 0.85-0.9 4 0.65-0.70 or 0.9-0.95 3 0.6-0.65 or 0.95-1.0 2 <0.6 1

[0061] In one embodiment, the social awareness component (SCA) is calculated by combining emotional appropriateness (EA), listener reaction time (Rt), and gaze behavior (GE) using the following formula: SCA=0.4×EA+0.3×Rt+0.3×GE.

[0062] In one embodiment, emotional appropriateness (EA) is determined by calculating the cosine similarity between an aggregated emotion vector (Eagg) from the video and a speaker context vector (C) derived from the transcript using the facebook / bart-large-mnli model. EA=(Eagg⋅C) / (‖Eagg‖⋅‖C‖).

[0063] In one embodiment, Table 9 shows the assignment of cosine similarity values ​​EA to rating scales: Table 9: EA (cosine similarity) Score 0.81-1.00 5 0.61-0.80 4 0.41-0.60 3 0.21-0.40 2 0.00-0.20 1

[0064] In one embodiment, the listener response time Rt is calculated similarly to LEoE: Using the timestamps of the speaker and listener emotions, the average delay ΔTavg is determined and the listener response time is calculated using the exponential decay function as follows: Rt=e∧(−α⋅ΔTavg) where α = 0.3.

[0065] Table 10 shows the assignment of the normalized Rt to the scores: Table 10: RT area Score 0.81-1.00 5 0.61-0.80 4 0.41-0.60 3 0.21-0.40 2 0.00-0.20 1

[0066] RT ≥ 0.81, score 5: Very fast, context-appropriate reactions. RT ≤ 0.20, score 1: Delayed or weak reactions.

[0067] In one embodiment, gaze engagement (GE) is determined by detecting eye features using Mediapipe Face Mesh, calculating normalized gaze coordinates, and measuring the proportion of images in which the gaze remains within a threshold toward the camera. Table 11 shows the mapping of GE percentages to scores. Table 11: GE series Score 0.81-1.00 5 0.61-0.80 4 0.41-0.60 3 0.21-0.40 2 0.00-0.20 1

[0068] In one embodiment, the Relationship Management (RM) component is calculated using the following formula: RM=0.6×Active Listening+0.4×Speaker’s Expressiveness.

[0069] In one embodiment, active listening skills are derived as the average of two subcomponents: expression mirroring and nodding.

[0070] In one embodiment, expression mirroring is measured by matching the timestamps of the speaker's and listener's emotions (using curriculum learning) and calculating the percentage of correctly assigned emotion responses using a predefined speaker-listener mapping table (see Table 12). Table 12: Speaker emotion Expected listener reaction Happy Happy, Awe, Neutral Sad Compassion, Sad, Neutral Angry Neutral, Sad, Compassionate Fearful Compassion, Neutral, Reverence Disgusted Neutral, Disgusted, Sad Surprised Awe, Happy, Neutral Neutral Neutral, Happy, Compassionate sympathy Compassion, Neutral, Sad awe Awe, Happy, Neutral

[0071] In one embodiment herein, Table 13 represents the mapping of expression mirroring percentage P to scores: Table 13: P (%) Score 50-70 5 40-50 or 70-80 4 30-40 or 80-90 3 20-30 or 90-100 2 <20 1

[0072] In one embodiment, nodding is detected by tracking forehead and chin landmarks using Mediapipe Face Mesh, calculating normalized vertical motion, and counting valid nods per minute. Table 14 shows the mapping of nods per minute (n) to scores: Table 14: n Score 6-10 5 4-6 or 10-12 4 2-4 or 12-14 3 1-2 or 14-15 2 >1 or >15 1

[0073] In one embodiment, the speaker's expressiveness is calculated as the average of emotional variability, emotion intensity, and entropy. Emotional variability (Ev) is calculated by dividing the video into time segments, detecting unique emotions in each segment, and calculating: Ev=(number of unique emotions) / (total number of emotions across all segments).

[0074] Table 15 shows the assignment of Ev to ratings: Table 15: Ev Score interpretation 0.81-1.00 5 Very wide emotional range 0.61-0.80 4 Varied, but slightly repetitive 0.41-0.60 3 Moderate variability 0.21-0.40 2 Limited selection 0.00-0.20 1 Emotional monotony

[0075] In one embodiment, the emotion intensity (I) is calculated as the mean of the predicted intensity values, normalized, and mapped to a scale of 1 to 5.

[0076] In one embodiment, Table 16 shows the mapping of normalized I to scores: Table 16: I Score 0.4-0.6 5 0.3-0.4 or 0.6-0.7 4 0.2-0.3 or 0.7-0.8 3 0.1-0.2 or 0.8-0.9 2 <0.1 or >0.9 1

[0077] In one embodiment, the entropy of the speaker's expressiveness is calculated using the same method as previously described: emotions are classified into positive and negative, then the entropy H is calculated and assigned to the value. The self-management component (SM) is calculated as follows: SM=0.5×Optimism+0.5×Entropy

[0078] In one embodiment, optimism (R) is calculated as the proportion of positive emotions: R=P / (P+N)

[0079] Table 14 shows the assignment of R to the values: R Score 0.65-0.80 5 0.60-0.65 or 0.80-0.85 4 0.55-0.60 or 0.85-0.90 3 0.50-0.55 or 0.90-0.95 2 <0.50 or >0.95 1

[0080] In one embodiment, entropy is calculated by classifying the detected emotions into positive (e.g., joy, surprise) and negative (e.g., sadness, anger, fear, disgust, contempt) classes. Based on the count values ​​P (positive) and N (negative), the entropy H is calculated as follows: H=−(P / (P+N))⋅log2(P / (P+N))−(N / (P+N))⋅log2(N / (P+N)).

[0081] Table 17 shows the assignment of entropy H to the point range 1-5: Table 17: Entropy H Score 0.75-0.85 5 0.70-0.75 or 0.85-0.9 4 0.65-0.70 or 0.9-0.95 3 0.6-0.65 or 0.95-1.0 2 <0.6 1

[0082] In one embodiment, system 100 comprises a modular architecture designed to process synchronized audio and video inputs from natural conversations and calculate Emotional Quotient (EC) scores based on predefined scoring logic and psychological validation. System 100 comprises multiple functional modules that together enable complete EQ analysis based on synchronized audio-video data. System 100 is configured to capture and synchronize video recordings from cameras 126 with corresponding audio data from microphones 128, ensuring temporal consistency between the modalities.

[0083] In one embodiment, system 100 is further capable of transcribing recorded speech and matching word timestamps with synchronized video frames, enabling precise temporal mapping of spoken content to facial expressions and other nonverbal cues. System 100 extracts multimodal behavioral features, including (but not limited to) facial expressions, speech patterns, gaze behavior, nodding frequency, emotional correctness, emotional entropy, intensity variation, emotional appropriateness, synchronization between speaker and listener emotions, and reaction times.

[0084] In one embodiment, the system applies 100 validated weighted formulas developed in collaboration with an experienced clinical psychologist to calculate standardized EQ component scores for the four main dimensions: self-awareness (SA), social awareness (SCA), relationship management (RM), and self-management (SM). The calculated component scores are then aggregated into a composite EQ score scaled from 4 to 20, creating an interpretable and quantifiable measure of emotional intelligence.

[0085] The system 100 is based on a modular architecture that ingests synchronized audio-video data and calculates four interpretable EQ dimension scores: Self-Awareness (SA), Social Awareness (SCA), Relationship Management (RM), and Self-Management (SM). Each of these scores is calculated individually using weighted formulas from extracted multimodal behavioral traits and aggregated to a total EQ score ranging from 4 to 20.

[0086] In one embodiment, system 100 captures real-time video data via cameras 126 and audio data via microphones 128 during natural conversations. This setup does not require specialized sensors or wearable devices, enabling broad accessibility and ease of use in various environments using standard cameras and microphones. Preprocessing module 112 transcribes the audio input using a speech recognition engine such as the Whisper model and extracts word timestamps. A lightweight buffering method matches these timestamps with the corresponding video frames to ensure synchronized analysis of facial expressions, speech patterns, and other behavioral cues.

[0087] The feature extraction module 114 extracts multimodal behavioral features, including facial expressions, speech patterns, gaze behavior, nodding frequency, emotional correctness, entropy, intensity variation, emotional appropriateness, synchronization between speaker and listener emotions, and reaction times. These features are converted into four standardized EQ component scores using weighted formulas, which have been validated by an experienced clinical psychologist and aligned with psychological theories such as Mayer and Salovey's four-branch model of emotional intelligence.

[0088] For example, in one embodiment for self-awareness (SA), the focus (60%) is placed on emotional entropy, as it is central to understanding and regulating one's own emotional state. Self-confidence, patience, and expressiveness are considered as secondary factors. The assessment module 116 calculates the final composite EQ score by summing the four component scores (SA, SCA, RM, SM), resulting in a total score ranging from 4 to 20. The output module 118 assigns each component a scale of 1-5, stores the detailed and aggregated scores in the database 122 for review and historical analysis, and transmits them to the user devices 124 via the user interface 130.

[0089] The System 100 also supports the validation of the calculated scores against established psychological instruments, such as the publicly available Global Emotional Intelligence Test, and confirms reliability through acceptable mean deviation when evaluated by an experienced clinical psychologist.

[0090] In one embodiment, the system 100 is compared to the publication "May I Know My EQ? Factors to Automate EQ Prediction Using Technology." This publication describes four EQ dimensions—self-awareness, social awareness, relationship management, and self-management—and suggests measurable parameters from audio-video inputs. However, it remains conceptual and dispenses with technical evaluation logics such as entropy mapping, eye-tracking models, or nod detection methods.

[0091] Although the publication highlights aspects such as gaze behavior, nodding frequency, and intensity variation, it does not implement real-time detection or algorithmic processing. While weightings and mapping matrices are presented, detailed models, evaluation domains, and process flows are lacking. In contrast, the proposed system 100 fully implements all these aspects with defined formulas, evaluation logic, and real-time methods. It calculates entropy, synchronizes audio-video data, and analyzes behavioral features in real time. Gaze tracking is performed using Mediapipe and dynamic thresholding logic, and nodding is detected by analyzing vertical head movements.

[0092] System 100 assigns the extracted features to a standardized rating scale from 1 to 5, supported by explicit technical formulas for emotional entropy and transition smoothness. Instead of general ML references, it uses deep learning methods such as curriculum learning, Whisper for speech recognition, and BART-MNLI for analyzing emotions in context. It also uses timestamp matching, differential lag calculation, normalization, and weighted scoring, validated by a clinical psychologist, to provide a real-time, quantifiable EQ assessment—something missing from the aforementioned publication.

[0093] The novelty of System 100 lies in its provision of quantifiable and scalable real-time assessment logic, audio-video fusion, entropy assessment, and emotion synchronization between speaker and listener—aspects not disclosed in the publication. System 100 can be used, for example, in job interviews for objective assessment of emotional intelligence, in employee training for personalized EQ feedback and leadership development, in education and consulting to promote emotional awareness and interaction, in therapy and psychological assessment to capture emotional states, and in remote collaboration to analyze emotional cues and increase team bonding.

[0094] Fig.Figure 2 shows a flowchart 300 of a workflow for analyzing the emotional quotient (EQ) using multimodal deep learning on audio-video inputs by system 100 using steps 302, 304, 306, 308, and 310. In step 302, the data acquisition module 110 captures and synchronizes real-time video data via the cameras 122 and audio data via the microphones 124. In step 204, the preprocessing module 112 transcribes the captured audio data using the speech recognition engine (e.g., Whisper model) and matches word timestamps with the video data.

[0095] In step 306, the feature extraction module 114 extracts the multimodal behavioral features, including facial expressions, speech patterns, gaze behavior, nodding frequency, emotional correctness, emotional entropy, and intensity variation, by applying deep learning models trained with curriculum learning strategies. In step 208, the assessment module 116 calculates the standardized EQ component scores for self-awareness, social awareness, relationship management, and self-management using quantifiable and scalable real-time assessment logic, including entropy assessment and synchronization of speaker and listener emotions. In step 310, the output module 118 assigns the component scores, aggregates them into a composite EQ score ranging from 4 to 20, stores the results in the database 122, and transmits them to the user devices 124 via the user interface 130.

[0096] Numerous advantages of the present disclosure arise from the above. System 100 analyzes the emotional quotient (EQ) using multimodal deep learning on audio-video data, providing quantifiable and scalable real-time scoring logic, audio-video fusion, entropy scoring, and emotion synchronization between speaker and listener for objective assessment of emotional intelligence. System 100 extracts and processes measurable behavioral traits—such as facial expressions, speech patterns, gaze behavior, nod frequency, emotional correctness, intensity variation, and reaction times—and applies clinically sound weightings and mathematically defined models to calculate standardized EQ component scores and a composite EQ score.

[0097] It will be apparent that numerous changes and modifications may be made to the described methods without departing from the basic principles of the invention. All such changes and modifications are incorporated into this application. Reference number list 100 systems 102 Computer device 104 processor 106 Reminder 108 variety of modules 110 Data acquisition module 112 Preprocessing module 114 Feature Extraction Module 116 Scoring module 118 Output module 120 Network 122 Database 124 user devices 126 cameras 128 microphones 130 User Interface

Claims

[1] A system (100) for analyzing the emotional quotient (EQ) using multimodal deep learning on audio-video inputs, comprising: a computing device (102) with a processor (104) and a memory (106) configured to store one or more instructions that can be executed by the processor (104); wherein the processor (104) is configured to execute a variety of modules (108) to process synchronized video and audio data from natural conversations, to extract multimodal behavioral features using deep learning models trained with curriculum learning strategies, and to calculate standardized EQ component values ​​as well as a composite EQ value; wherein the computing device (102) is connected via a network (120) to a database (122), user devices (124), cameras (126) and microphones (128); the multitude of modules (108) includes the following: a data acquisition module (110) configured to capture and synchronize video data via the cameras (126) and audio data via the microphones (128) in real time; a preprocessing module (112) configured to transcribe the captured audio data using a speech recognition engine, including the Whisper model, and to match word-based timestamps with the recorded video data; a feature extraction module (114) configured to extract multimodal behavioral features, including facial expressions, speech patterns, gaze behavior, frequency of nodding, emotional correctness, emotional entropy and intensity variation, by applying deep learning models with curriculum learning strategies; an evaluation module (116) configured to apply quantifiable and scalable evaluation logics in real time, including emotional entropy evaluation and synchronization of emotions between listener and speaker to calculate standardized EQ component values ​​corresponding to the areas of self-awareness, social awareness, relationship management and self-management; an output module (118) configured to map the EQ component values ​​to a standardized scale of 1 to 5, aggregate the component values ​​to a composite EQ value in the range of 4 to 20, store the results in the database (122) and transmit the values ​​to the user devices (124) via a user interface (130). [2] The system (100) according to claim 1, wherein the feature extraction module (114) detects speaker and listener emotions, predicts confidence levels and calculates the variation in emotional intensity by applying deep learning models with curriculum learning strategies. [3] The system (100) according to claim 1, wherein the feature extraction module (114) calculates the gaze behavior by tracking eye and iris landmarks, calculating normalized gaze coordinates and determining the proportion of images in which the gaze is directed towards the microphones (128) and cameras (126). [4] The system (100) according to claim 1, wherein the feature extraction module (114) detects nodding by tracking vertical head movements using facial landmarks, calculating normalized vertical distances over successive images, and counting nods per minute based on a dynamic threshold. [5] The system (100) according to claim 1, wherein the evaluation module (116) calculates the emotional correctness by comparing recognized emotions of the speaker and listener with context-based expected reactions and calculating correctness values ​​based on timestamp matches. [6] The system (100) according to claim 1, wherein the evaluation module (116) calculates the emotional entropy by categorizing detected emotions into positive and negative classes and applying an entropy formula based on the distribution. [7] The system (100) according to claim 1, wherein the output module (118) stores the calculated EQ component values ​​and the composite EQ value in the database (122) for validation, testing and historical analysis and transmits the values ​​to user devices (124) or external psychological rating systems. [8] The system (100) according to claim 1, wherein the evaluation module (116) applies weights derived from clinical validations to convert extracted features into standardized EQ component values. [9] The system (100) according to claim 1, wherein the composite EQ score and the component scores are validated using an established psychological test to confirm reliability and acceptable mean deviations.