System for real-time analysis of emotional feedback during motivational presentations
Patent Information
- Application Number
- DE202025104706
- Authority / Receiving Office
- DE · DE
- Patent Type
- Utility models
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-11-06
- Estimated Expiration
- 2035-08-31
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical field of the invention
[0001] The present invention relates to the field of human-computer interaction and emotion analysis. More specifically, it relates to a system and a device configured for the real-time capture, analysis, and visualization of the emotional feedback of the audience during live motivational speeches. This involves the use of multimodal sensor arrays, embedded emotional AI, and adaptive feedback loops to improve speaker performance and audience engagement. Background of the invention
[0002] In motivational speeches, the effectiveness of communication depends not only on the content conveyed but also on the speaker's ability to connect with the audience's emotions. However, speakers often lack objective and timely feedback on their listeners' emotional state. Traditional surveys and feedback forms after the event are limited by subjectivity, processing delays, and a lack of detail. Furthermore, existing emotion recognition systems are either post-processed, do not operate in real time, or are not integrated into the live dynamics.
[0003] There is a need for a robust, non-invasive, real-time system capable of dynamically capturing the emotional states of the audience during live motivational sessions. The system must be able to collect data across various modalities, including facial expressions, tone of voice, physiological changes, and body language. The analyzed emotional responses must be correlated with the content and timing of the speech to generate meaningful insights. Such a system would enable motivational speakers to dynamically adapt their presentations and provide analytics for long-term performance optimization.
[0004] A motivational speaker's ability to effectively captivate and inspire an audience is inextricably linked to emotional connection. However, gauging the emotional response of a live audience is largely subjective and inferred from superficial cues such as applause, facial expressions, or post-event verbal feedback. Traditional methods for capturing audience engagement in motivational speeches include manual observation, feedback forms, post-event surveys, and interviewer-led focus groups. While these techniques provide valuable information, they are retrospective, time-consuming, and often fail to capture the nuances of audience emotions in real time. Furthermore, such approaches are affected by perceptual and memory biases, reducing the reliability and accuracy of the data collected.
[0005] In more technologically advanced environments, speakers and organizers attempt to gain insights into audience sentiment at various points during a presentation using audience response systems such as clickers or mobile app polls. However, these systems require active participation, disrupt the natural flow of the event, and offer only discrete, predefined feedback options (e.g., Likert scales or binary responses). Therefore, these tools are not well-suited to capturing the dynamic, continuous nature of emotional engagement, which fluctuates over time and varies considerably from person to person.
[0006] Another important category of emotion recognition systems are facial expression analysis tools. These systems, often developed using computer vision and machine learning, can classify basic emotions such as joy, sadness, anger, and surprise based on facial features. Commercial tools like the Microsoft Azure Emotion API or Affectiva's SDK have been integrated into marketing and retail experiences to measure consumer responses. However, in the context of motivational speeches, such systems have several limitations. First, they focus primarily on facial expressions and ignore other important emotional channels such as tone of voice, physiological arousal, or posture.Second, these tools are typically designed for static or semi-controlled environments such as individual face scans or focus groups, and perform poorly in large, dynamic environments with variable lighting, occlusion, and diverse ethnic facial representations. Third, most commercially available facial emotion systems are cloud-dependent, which introduces latency and raises privacy concerns when used in real time.
[0007] Several research prototypes are exploring multimodal emotion recognition, combining facial, voice, and physiological data. In academic studies, for example, participants are equipped with electroencephalography (EEG) caps, electrocardiogram (ECG) sensors, or galvanic skin response (GSR) electrodes to record their emotional state. While such setups can provide high-resolution emotional data, they are impractical for live presentations to large and mobile audiences. The required instrumentation is intrusive, expensive, and logistically complex, making it unsuitable for scalable use in real-world presentations. Furthermore, these systems are often limited to laboratory conditions and not integrated with speaker support tools, rendering them impractical for adaptive live feedback.
[0008] Attempts have also been made to use real-time sentiment analysis from social media posts or mobile live chat feedback in online presentations. These solutions are particularly common in virtual webinars or live streams where the audience can comment or react with emojis or text. Sentiment analysis techniques using natural language processing (NLP) are employed to classify these reactions as positive, negative, or neutral. While useful in digital environments, this form of emotion analysis is inherently sparse and imprecise. It relies heavily on uneven audience participation and the interpretation of short, context-free messages, limiting its emotional resolution and accuracy. Furthermore, such systems are useless in face-to-face sessions where live interaction on social media is minimal or even undesirable.
[0009] Biometric feedback systems, commonly used in the wellness and gaming industries, offer a different approach to capturing emotional states. Devices such as Fitbit, Apple Watch, and Empatica E4 allow for continuous monitoring of heart rate, electrodermal activity, and movement. However, these systems are designed for individual users and do not provide collective emotional insights suitable for addressing an audience as a group. More importantly, every member of the audience must be equipped with and synchronized to the system, which is both impractical and ethically complex. Furthermore, they do not offer real-time visualization for the speaker or contextual correlation between biometric trends and the speaker's spoken content.
[0010] Some high-end conference rooms or corporate experience centers are experimenting with custom-built emotion-tracking rooms equipped with ceiling-mounted infrared cameras, directional microphones, and custom control software. In theory, these systems can capture audience engagement. However, they are prohibitively expensive, location-bound, and often lack flexible software. They are neither mobile nor portable and cannot be adapted to the varying acoustics and seating arrangements of different venues. This severely limits their applicability for mobile motivational speakers or dynamic presentation environments such as workshops and seminars.
[0011] A major limitation of most existing solutions lies in the lack of contextual linking of emotional responses to the content of the speech. Even when emotional data is captured, current systems do not assign it to specific moments or topics within the speech. This leads to emotional interpretations that are detached from context. Speakers therefore struggle to understand which parts of their speech were effective, confusing, or uninteresting. Without timeline-based emotional annotation, post-event feedback is often vague, such as "the middle section was a bit slow," without precise details about where or why.
[0012] Real-time feedback for speakers during live sessions remains one of the least developed aspects of emotion analysis. With the few systems that offer live reporting, the feedback is either too complex to interpret in real time or it's presented on devices like tablets or laptops that aren't ergonomically designed for stage use. As a result, speakers can't use the data to adjust their presentation in the moment, and the opportunity for dynamically improving audience engagement is lost.
[0013] Data privacy and ethical concerns also pose significant challenges. Most existing solutions that process facial or physiological data lack robust anonymization, federated learning protocols, or consent management frameworks. This limits their use in compliant environments, particularly under regulatory frameworks such as the General Data Protection Regulation (GDPR). Without integrated ethical safeguards, event organizers risk violating personal data protection laws, and speakers cannot use such systems without legal risks.
[0014] Therefore, there remains an urgent and unmet need for a comprehensive, modular, and privacy-conscious system capable of capturing, processing, and visualizing the emotional feedback of the audience during live motivational speeches in real time. This system must support multimodal sensing, be easily deployable in a wide variety of venues, and offer real-time visualization mechanisms that are intuitive, non-distracting, and actionable for the speaker. It must also enable content synchronization to facilitate contextually relevant post-event analysis that is useful for long-term improvements. None of the current solutions on the market or in the research literature offer this combination of features, particularly not in a portable, non-intrusive, and speaker-centric form, highlighting the novelty and necessity of the proposed invention. Summary of the invention
[0015] The invention provides a system for real-time analysis of emotional feedback during motivational speeches. It comprises a distributed network of multimodal sensors integrated into the venue or audience wearable devices, a central edge AI processing unit for emotional state detection, a timestamping and content alignment engine, and a real-time feedback rendering module configured in a wearable or structural display device for the speaker. The system uses trained deep learning models for multimodal emotion recognition, combines input from visual, auditory, and physiological channels, and outputs emotional heatmaps indexed to the speech content both in real time and after the session.
[0016] The main objective of the present invention is to provide a system capable of capturing, analyzing, and relaying real-time emotional feedback from the audience during motivational speeches. This allows speakers to dynamically adapt their presentations to the emotional responses of the listeners. A further objective of the invention is to create a non-intrusive, privacy-compliant architecture that utilizes multimodal sensory inputs—including visual cues, vocal variations, and physiological signals—to infer emotional states without requiring active audience participation or intrusive data collection mechanisms. Another objective is the integration of a fusion-based emotion inference engine that combines these inputs using deep learning techniques optimized for real-time processing on edge computing hardware. This eliminates latency and reliance on external cloud infrastructure.
[0017] The invention also aims to provide a speaker-facing feedback mechanism in the form of a portable device or a structural visualization interface, such as a display integrated into a podium or a transparent head-up projection, which can unobtrusively represent aggregated emotional states in intuitive visual formats such as trend lines, emotional indices, or color-coded cues. A further objective is to enable precise timestamping and alignment of the audience's emotional responses with the content timeline of the speech through embedded speech-to-text conversion and semantic segmentation techniques, thereby allowing both real-time and retrospective correlation of emotional feedback with specific speech segments.Furthermore, the invention attempts to offer longitudinal analyses by storing session-by-session emotional attributions and performance metrics that can be compared over time to help speakers assess their progress and optimize future interactions.
[0018] A key objective of the invention is to ensure that the emotion analysis system is modular and portable, allowing it to be easily deployed in a wide variety of locations—from large auditoriums to small workshop spaces—without requiring permanent installations or specialized infrastructure. The invention also aims to integrate robust data anonymization, secure transmission protocols, and optional federated learning frameworks to comply with data protection regulations such as the GDPR, thereby facilitating ethical deployment at scale. Furthermore, it seeks to democratize access to real-time audience feedback by providing speakers—whether beginners or professionals—with actionable insights previously unavailable through traditional observation or feedback surveys.Ultimately, the invention aims to increase the effectiveness of motivational speeches by enabling a continuous, adaptive, and emotionally intelligent communication cycle between speaker and audience. BRIEF DESCRIPTION OF THE FIGURE
[0019] These and other features, aspects, and advantages of the present invention will be better understood if the following detailed description is read with reference to the accompanying drawing, in which the same symbols consistently represent the same parts. The following applies: Fig. Figure 1 shows a block diagram of a system for real-time analysis of emotional feedback during motivational speeches.
[0020] Experts will also recognize that the elements in the drawing are shown for the sake of simplicity and are not necessarily to scale. For example, the flowcharts illustrate the process by highlighting the main steps to enhance understanding of the aspects of this disclosure. Furthermore, with regard to the design of the device, one or more components of the device may be represented in the drawing by conventional symbols, and the drawing may show only the specific details relevant to understanding the embodiments of this disclosure, so as not to clutter the drawing with details that are readily apparent to those skilled in the art after reading this description. Detailed description of the invention
[0021] For a better understanding of the inventive principles, reference is made below to the embodiment shown in the drawing, which is described in specific terminology. However, this does not limit the scope of the invention. Changes and further modifications of the illustrated system, as well as further applications of the inventive principles, are possible, as would normally occur to a person skilled in the art in the field of invention.
[0022] It is clear to the person skilled in the art that the preceding general description and the following detailed description are exemplary and explanatory of the invention and are not intended as a limitation of it.
[0023] References in this specification to “an aspect”, “another aspect”, or similar expressions mean that a particular feature, structure, or property described in connection with the embodiment is included in at least one embodiment of the present disclosure. Therefore, occurrences of the expressions “in one embodiment”, “in another embodiment”, and similar expressions in this specification may all refer to the same embodiment, but need not.
[0024] The terms "includes," "include," or other variations thereof are intended to cover non-exclusive inclusion, such that a process or method that includes a list of steps may not only contain those steps but may also include other steps not expressly listed or inherent in such process or method. Likewise, the statement "includes..." in the case of one or more devices, subsystems, elements, structures, or components does not, without further limitations, preclude the existence of other devices, subsystems, elements, structures, components, or additional devices, subsystems, elements, structures, or components.
[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as understood by a person skilled in the art in the field of the invention. The system, methods, and examples provided here serve only for illustration and are not to be construed as a limitation.
[0026] Embodiments of the present disclosure are described in detail below with reference to the attached drawing.
[0027] In Fig.Figure 1 shows a block diagram of a system 100 for real-time analysis of emotional feedback during motivational speeches. The system 100 comprises: an array of multimodal sensors (102), including at least one visual sensor (102a) for capturing the facial expressions of spectators, at least one directional microphone (102b) for capturing the acoustic responses of the audience, and optionally one or more physiological sensors (102c) for capturing biometric signals from spectators; an edge-based processing unit (104) that is communicatively coupled to the array of multimodal sensors and comprises: (a) a feature extraction module (104a) for extracting visual features from captured facial images, acoustic features from voice responses, and physiological features from biometric signals;(b) an emotion inference engine (104b) for processing the features using a deep learning-based emotion recognition model comprising a convolutional neural network (CNN) for classifying facial expressions, a recurrent neural network (RNN) for classifying voice emotions, and a multimodal late fusion layer for computing a composite emotion state vector representing the aggregated emotions of the audience; (c) a timestamping and speech alignment module (104c) configured to correlate the computed composite emotion state vector with segmented portions of a live motivational speech based on real-time speech-to-text transcription and semantic analysis; and (d) a session-based storage unit (104d) configured to log time-indexed emotion state vectors and corresponding speech segments for post-event analysis;a speaker feedback interface (106) comprising a portable display or a podium-mounted visualization panel (106a), wherein the interface is configured to display visual indicators for emotional feedback in real time, the indicators being derived from the emotion state vector and comprising at least emotion trend graphs, threshold alerts, or engagement indices.
[0028] In one embodiment, the visual sensor (102a) comprises a plurality of high frame rate RGB and near-infrared cameras mounted at various vantage points of an event venue, each camera being calibrated with face recognition techniques using Haar cascade or HOG-SVM techniques, and the visual data being preprocessed prior to feature extraction using illumination correction and landmark stabilization.
[0029] In one embodiment, the directional microphone (102b) comprises a beamforming microphone array configured with real-time noise reduction and spatial filtering techniques to isolate the audience's vocal responses. The acoustic signal is converted into Mel spectrogram representations and categorized into emotional categories using a CNN-LSTM hybrid model trained on labeled vocal emotion datasets.
[0030] In one embodiment, the physiological sensors (102c) are attached to the wrist or integrated into a badge and comprise a photoplethysmography sensor and a galvanic skin response sensor, wherein the biometric data are transmitted via a Bluetooth Low Energy (BLE) mesh network with edge aggregation nodes that perform local encryption and data smoothing using adaptive window averaging.
[0031] In one embodiment, the emotion inference machine (104b) comprises an attention-based fusion mechanism configured to dynamically assign weights to input modalities based on signal reliability assessments and context vectors derived from audience temporal response windows, thereby improving robustness against signal dropouts or noise in each individual modality.
[0032] In one embodiment, the timestamping and speech alignment module (104c) includes a real-time automatic speech recognition engine that uses a transformer-based speech model to generate a continuous transcript of the motivational speech. The transcript is segmented into semantically coherent blocks using natural language analysis and lexical cohesion evaluation, and further aligned with emotion data using a synchronized clock source.
[0033] In one embodiment, the speaker feedback interface (106) comprises a flexible OLED display embedded in a wearable wristband. The display is controlled by a microcontroller that receives a live stream of emotional index data via a secure wireless protocol and is configured to show a radial color wheel representation of dominant emotions, a linear trend graph of emotional fluctuations, and warning symbols when persistent negative emotional states above a preset threshold are detected.
[0034] In one embodiment, the podium-mounted visualization panel (106a) is a transparent OLED screen integrated into a teleprompter surface. The panel is configured to overlay visualizations of emotional feedback in a non-obstructive manner, including speaker-facing engagement graphs, emotional zone shifts, and segmented audience response values synchronized with the speech.
[0035] In one embodiment, the session-based storage unit (104d) is implemented as a hybrid on-device and cloud-synchronized repository, wherein the on-device module stores immediate session data and the cloud module supports longitudinal analysis across sessions, with a differential privacy layer being applied during cloud synchronization to anonymize sensitive data.
[0036] In one embodiment, it also includes a federated learning framework configured to update the emotion recognition models using distributed training across audience devices or venue edge nodes, with the raw data remaining on the source device and only encrypted model gradients being aggregated, thus ensuring GDPR-compliant data processing and adaptive model improvement.
[0037] The invention described herein relates to a comprehensive system for the real-time analysis of emotional feedback during motivational speeches. It comprises a tightly integrated architecture that includes multimodal sensing, deep-learning-based emotional inference, semantic speech alignment, and dynamic feedback visualization. The core technical framework of the system is distributed across several components—a multimodal feature acquisition layer, a hierarchical emotion inference engine, a timestamping and content alignment module, and a feedback rendering subsystem—that work together in near real-time to provide the speaker with interpretable emotional insights during live sessions.
[0038] The process begins with the multimodal data acquisition pipeline. Visual data is captured by an array of high-frame-rate RGB and near-infrared cameras installed throughout the venue. Each camera is calibrated for consistent facial detection and expression analysis. These cameras continuously feed video images into a facial feature detection module, which uses either Haar cascade classifiers or HOG-SVM models to initially generate bounding frames. Once the facial features are located, the relevant areas are normalized and the lighting is adjusted. A convolutional neural network (CNN), pre-trained with facial expression datasets such as FER+ and AffectNet, extracts high-dimensional features indicative of microexpressions and macroaffective shifts.These characteristics are passed to a classification level that assigns the input to one of several emotional categories or creates continuous ratings for dimensions such as valence and arousal.
[0039] Simultaneously, the audience's audio signals are captured via a directional microphone array integrated into the venue. Each microphone uses beamforming technology to spatially isolate the sound emanating from the seating area. Real-time spectral subtraction and Wiener filtering are employed for noise reduction. The resulting signal is segmented into short-term windows and transformed into Mel spectrogram representations. These spectrograms are fed into a CNN-LSTM hybrid model trained to recognize affective vocal signals such as laughter, surprise, distance, or agreement. The temporal LSTM component captures the dynamic flow of vocal utterances over time, thus ensuring context-sensitive and temporally coherent emotion recognition.
[0040] Where available, optional physiological data from spectators wearing wrist-worn or badge-integrated sensors will be incorporated into the engineering pipeline. These sensors monitor photoplethysmography (PPG) and galvanic skin response (GSR) and transmit the data to a Bluetooth Low Energy (BLE) mesh network. The signals are preprocessed through window smoothing and motion artifact filtering, followed by feature extraction, including heart rate variability, skin conductance, and peak amplitude. These features are normalized and interpreted using a regression-based model trained on physiological-affective correlation datasets. This generates emotional estimates corresponding to levels of stress, arousal, or relaxation.
[0041] The Emotional Inference Machine is a late-fusion neural network architecture that synthesizes these heterogeneous data modalities. After independent, modality-specific processing, the outputs are passed to an attention-based fusion module. This module uses modality-level attention vectors derived from signal quality confidence scores and contextual weights. The fusion layer is implemented as a bidirectional gated recurrent unit (Bi-GRU) that combines the temporal dynamics of multimodal features.
[0042] To correlate these emotional states with the speaker's actual discourse, a timestamp and content matching module is activated in parallel. This module captures the speaker's live audio and processes it using a real-time automatic speech recognition (ASR) system based on a transformer-based language model. The ASR continuously transcribes the spoken content and applies semantic segmentation using cohesion analysis, discourse markers, and lexical boundary detection to divide the transcript into meaningful units such as statements, narratives, or transitions. A synchronized timestamp mechanism based on NTP or GPS timing ensures the correspondence between emotional state vectors and semantic language units.
[0043] The emotional vectors are mapped to the corresponding speech segments, creating a temporal-emotional matrix that is stored in a session-specific local database. Each row of this matrix contains emotion scores, entropy indicators (audience dispersion), and segment metadata such as topic, speaker tone, and pause duration. In parallel, emotional entropy metrics are calculated using Shannon entropy to assess the diversity of emotional responses within the audience. This helps identify emotionally polarized or fragmented speech segments.
[0044] The visualization layer translates this complex emotional data into real-time visual feedback for the speaker—unobtrusive and cognitively efficient. In wearable applications, a curved OLED display integrated into a wristband receives live emotion vectors and displays them using radial emotion diagrams, trend graphs, and dynamic threshold alerts. These visuals are updated with an adaptive refresh rate—speeding up during emotionally unstable moments and slowing down during stable phases—to avoid cognitive overload for the speaker. In podium-based displays, a transparent OLED screen integrated into the teleprompter surface overlays emotional feedback such as engagement graphs, audience mood tracings, and warning symbols when negative emotions persist for defined periods.
[0045] The system also stores all data generated during the session for post-event analysis. This data is indexed using semantic tags, timestamps, and speaker-defined keyframes. Optionally, an LDA-based theme modeling technique is applied to the transcribed speech content to identify evolving themes. This allows emotional feedback to be contextualized thematically rather than just temporally. Federated learning mechanisms are used to improve the performance of the emotion model over time without exporting raw audience data. Local devices regularly train on session data and only transmit model weight updates, in encrypted form, to a central server. This ensures GDPR-compliant and privacy-compliant learning.
[0046] By integrating this detailed technical pipeline, the system enables robust, real-time emotional feedback analysis specifically tailored to the dynamics of motivational speeches. It delivers actionable insights during and after the engagement, enhances the speaker-audience connection, and facilitates long-term improvement in speaker effectiveness through a scientifically sound and ethically safe technological framework.
[0047] The proposed system consists of three core components: a sensor architecture, a processing and inference unit, and a feedback visualization interface, all functionally integrated into a structured platform that can be used in live speaking rooms.
[0048] The sensor architecture comprises multiple optical sensors (RGB and infrared cameras) strategically placed to capture audience facial expressions. These cameras are calibrated for facial action unit detection, utilizing convolutional neural network (CNN) models trained on emotion datasets such as AffectNet or FER+. Complementing the visual input, directional microphone arrays are integrated throughout the hall to capture vocal responses such as laughter, sighs, or applause. These signals are processed using spectrogram-based emotion classification models trained to differentiate between emotional tones such as amusement, confusion, or excitement.
[0049] Furthermore, where permitted, spectators can voluntarily wear wristbands or smart badges with physiological sensors (e.g., photoplethysmography for measuring heart rate variability, galvanic skin response sensors). These sensors transmit data to the central unit via a secure Bluetooth Low Energy (BLE) network. The architecture is modular, allowing the system to operate with only camera and audio data in scenarios with minimal instrumentation.
[0050] All data streams are routed to an on-site Edge AI Processing Unit (EPU) to reduce latency. The EPU consists of a GPU-enabled compute board running pre-trained deep learning models for real-time emotion classification. The output of each modality is processed by a late-fusion emotion inference engine, which utilizes attention-based recurrent neural networks (RNNs) to temporally align and weight emotional signals. For each timeframe, the fusion engine derives a composite emotion state score vector for the audience, which is then stored in a time-indexed database.
[0051] The timestamping and content alignment engine synchronizes these emotional vectors with the speaker's live transcript. This is achieved through the integration of speech-to-text modules that transcribe the speech and segment it into semantically distinct parts. These segments are tagged and indexed based on the inferred emotions of the audience to enable post-session correlation analysis.
[0052] To provide real-time feedback to the speaker, the system features a feedback visualization interface integrated into a specially designed wearable device or a podium-mounted display. The wearable device is a curved OLED band worn by the speaker on their forearm or wrist. This band displays dynamic emotional metrics in the form of color-coded arcs, emotional trend lines, and warning indicators if negative emotions (e.g., disinterest, confusion) dominate for an extended period. Alternatively, a head-up display integrated into a transparent teleprompter lens can project emotional charts for the speaker's discreet visualization.
[0053] The system structure is contained in a mobile deployment kit, which includes the following: 1. A visual sensor array mounted on a telescopic tripod with PTZ (pan-tilt-zoom) function; 2. An audio recording and preprocessing module housed in a rack-mounted acoustic digitizer; 3. A robust EPU enclosure with thermal and EMI shielding; 4. A wireless interface gateway that supports Wi-Fi 6, BLE and optional LTE uplink for cloud analytics.
[0054] The post-session analytics module, which can be hosted either in the cloud or locally, generates heatmaps, engagement scores, emotional arcs, and speaker effectiveness metrics. These can be exported in standardized formats and support longitudinal tracking of speaker performance across sessions.
[0055] Security and data protection are ensured through the use of federated learning updates for the emotion models. This guarantees that no raw facial data leaves the edge device. Consent protocols and anonymization pipelines are embedded in the data collection logic and comply with the GDPR and related data protection frameworks.
[0056] The invention relates to the technical field of human-computer interaction, emotion analysis, and real-time audience response systems. In particular, it relates to systems and methods for capturing, analyzing, and visualizing real-time emotional feedback during live motivational speeches through the integration of multimodal sensor data, deep learning techniques for emotion recognition, edge-based computing infrastructures, and adaptive visualization devices for the speaker. The invention is applicable to public speaking, live performance monitoring, communication coaching, and intelligent presentation systems where emotional engagement and real-time feedback are crucial for optimizing speaker performance and the audience experience.
[0057] The drawing and the preceding description show examples of embodiments. Those skilled in the art will recognize that one or more of the described elements can be combined to form a single functional element. Alternatively, certain elements can be divided into several functional elements. Elements of one embodiment can be added to another embodiment. For example, the sequence of the processes described here can be changed and is not limited to the manner described here. Furthermore, the actions of a flowchart need not be implemented in the sequence shown; nor does it necessarily have to be performed by all actions. Actions that are not dependent on other actions can also be performed in parallel with the other actions. The scope of the embodiments is in no way limited by these specific examples.Numerous variations are possible, whether explicitly stated in the specification or not, such as differences in structure, dimensions, and material use. The range of embodiments is at least as broad as specified in the following claims.
[0058] Advantages, further benefits, and problem solutions have been described above with reference to specific embodiments. However, the advantages, benefits, problem solutions, and all components that can lead to an advantage, benefit, or solution occurring or becoming more apparent are not to be construed as critical, necessary, or essential features or components of individual or all claims. REFERENCES 100 A system for real-time analysis of emotional feedback during motivational presentations. 102 Array of Multimodal Sensors 102a Visual Sensor 102b Directional Microphone 102c Physiological Sensors 104 edge-based processing units 104a Module for Feature Extraction 104b Emotional Inference Engine 104c Timestamp and Language Alignment Module 104d Session-based storage unit 106 Speaker Feedback Interface 106a A portable display or a pedestal-mounted visualization panel
Claims
[1] A system for real-time analysis of emotional feedback in motivational speeches, consisting of: a series of multimodal sensors, including at least one visual sensor configured to capture facial expressions of spectators, at least one directional microphone configured to capture the audio responses of the audience, and optionally one or more physiological sensors configured to capture biometric signals from spectators; an edge-based processing unit that is communicatively coupled to the arrangement of multimodal sensors, wherein the edge-based processing unit comprises the following: (a) a feature extraction module configured to extract visual features from captured facial images, acoustic features from voice responses, and physiological features from biometric signals; (b) an emotion inference machine configured to process the features using a deep learning-based emotion recognition model comprising a convolutional neural network (CNN) for classifying facial expressions, a recurrent neural network (RNN) for classifying voice emotions, and a multimodal late fusion layer configured to compute a composite emotion state vector representing the aggregated emotions of the audience; (c) a timestamp and speech alignment module configured to correlate the calculated composite emotion state vector with segmented portions of a live motivational speech based on real-time speech-to-text transcription and semantic analysis; and (d) a session-based storage unit configured to log time-indexed emotional state vectors and corresponding speech segments for post-event analysis; A speaker feedback interface comprising a portable display or a podium-mounted visualization panel, wherein the interface is configured to display visual indicators of emotional feedback in real time, the indicators being derived from the emotional state vector and including at least emotional trend graphs, threshold alerts, or engagement indices. [2] System according to claim 1, wherein the visual sensor comprises multiple high frame rate RGB and near-infrared cameras mounted at different vantage points of an event venue, each camera being calibrated using facial recognition techniques using Haar cascade or HOG-SVM techniques, and wherein the visual data are preprocessed prior to feature extraction using illumination correction and landmark stabilization. [3] System according to claim 1, wherein the physiological sensors are attached to the wrist or integrated into the badge and comprise a photoplethysmography sensor and a galvanic skin response sensor, wherein the biometric data are transmitted via a Bluetooth Low Energy (BLE) mesh network with edge aggregation nodes that perform local encryption and data smoothing using adaptive window averaging. [4] System according to claim 1, wherein the emotional inference machine comprises an attention-based fusion mechanism configured to dynamically assign weights to input modalities based on signal reliability assessments and context vectors derived from audience temporal response windows, thereby improving robustness against signal dropouts or noise in each individual modality. [5] System according to claim 1, wherein the timestamping and speech alignment module comprises an automatic real-time speech recognition machine which uses a transformer-based speech model to generate a continuous transcript of the motivational speech, wherein the transcript is segmented into semantically coherent blocks using natural language analysis and evaluation of lexical cohesion and is further aligned with emotion data using a synchronized clock source. [6] System according to claim 1, wherein the interface for speaker feedback comprises a flexible OLED display embedded in a wearable wristband. The display is controlled by a microcontroller that receives a live stream of emotional index data via a secure wireless protocol and is configured to show a radial color wheel representation of dominant emotions, a linear trend graph of emotional fluctuations, and warning symbols when sustained negative emotional states above a preset threshold are detected. [7] System according to claim 1, wherein the visualization panel mounted on the podium is a transparent OLED screen integrated into a teleprompter surface, the panel being configured to overlay visualizations of emotional feedback in a non-obstructive manner, including speaker-facing engagement graphs, emotional zone shifts and segmented audience response values synchronized with the speech flow. [8] System according to claim 1, wherein the session-based storage unit is implemented as a hybrid on-device and cloud-synchronized repository, wherein the on-device module stores immediate session data and the cloud module supports longitudinal analysis across sessions, wherein a differential privacy layer is applied during cloud synchronization to anonymize sensitive data.
Citation Information
Cited By
Emotion recognition and intervention system based on facial micro-expression and physiological signal fusion
CN121456675A
Emotion recognition and intervention system based on facial micro-expression and physiological signal fusion
CN121456675B
Family accompanying spherical unmanned aerial vehicle family accompanying method and system
CN121812103A
A family companion spherical unmanned aerial vehicle family accompanying method and system
CN121812103B
Conference terminal interaction system and method based on large model
CN121887837A