Apparatus and method for extended reality perception of a user
The apparatus and method address the challenge of personalized stimulus management in XR environments by using physiological sensors and AR processing to filter out stress-inducing audio and visual stimuli, improving user comfort and immersion.
Patent Information
- Application Number
- PCT/EP2025/071521
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-01
- Filing Date
- 2025-07-25
- Publication Date
- 2026-02-05
AI Technical Summary
Traditional noise cancellation technologies fail to discriminate between different types of audio and visual stimuli, potentially blocking important sounds or visuals that may be beneficial or necessary for the user, and existing XR environments lack personalized stimulus management to address individual variability in user perception.
An apparatus and method that utilizes physiological sensors to measure user parameters, correlates these with audio and video signals, and adjusts or filters out stress-inducing stimuli in real-time through an AR processing circuit, employing machine learning for precise identification and mitigation of specific auditory and visual stressors.
Enhances user comfort and immersion in XR environments by accurately identifying and modifying stress-inducing stimuli, ensuring a personalized and effective reduction of negative emotional impact.
Smart Images

Figure EP2025071521_05022026_PF_FP_ABST
Abstract
Description
[0001] APPARATUS AND METHOD FOR EXTENDED REALITY PERCEPTION OF A USER
[0002] Field
[0003] The present disclosure generally relates to extended reality (XR) perception, and, more particularly to methods and apparatuses for identifying and mitigating disturbing audiovisual stimuli.
[0004] Background
[0005] Traditional noise cancellation technology functions by generating sound waves that are the inverse of the unwanted noise, effectively neutralizing it. This approach, while effective in creating a quieter environment, can inadvertently block all sounds, including those that might be important or necessary for the listener to hear. For example, in everyday life, missing crucial audio cues such as a baby's cry, an alarm, or an incoming phone call can pose significant challenges or even risks. This blanket approach to noise cancellation does not discriminate between different types of sounds, treating all noise as undesirable and removing it entirely.
[0006] In a similar vein, visual stimuli can also have a profound impact on individuals. Certain visual patterns, colors, or even the movement of objects might evoke stress, anxiety, or discomfort in some people but not in others. For instance, a specific color like a bright pink might be perceived as jarring and stressful to one person, while another might find it perfectly acceptable or even pleasant. The individual variability in visual perception means that a one- size-fits-all approach to visual stimulus management is inadequate.
[0007] Within XR (Extended Reality) environments, where immersive experiences may be created by blending the physical and virtual worlds, the ability to finely tune what users see and hear becomes even more crucial. XR environments can magnify the effects of both audio and visual stimuli, making the need for selective filtering mechanisms paramount.
[0008] Thus, there may be a demand for identifying and mitigating stimuli that are disturbing to an individual user, allowing beneficial or neutral stimuli to remain unobstructed. Hence, a sophisticated detection and response systems capable of distinguishing between different types of stimuli and their impact on users' emotional and psychological states may be needed. Summary
[0009] This demand is addressed by apparatuses and methods in accordance with the appended claims.
[0010] According to a first aspect, the present disclosure proposes an apparatus for extended reality (XR) perception of a user. The apparatus comprises an input interface configured to receive an audio signal of a scene and to receive a video signal of the scene. The apparatus further comprises at least one physiological sensor configured to measure at least one physiological parameter of the user while the user is experiencing the scene. The apparatus further comprises a processing circuit configured to correlate the measured at least one physiological parameter of the user with the audio and / or video signal, and based on the correlation, determine auditory and / or visual stimuli causing a stress level of the user above a threshold. This approach may allow for the identification and mitigation of specific stressinducing stimuli (e.g., in real-time). This means that the user can have a more comfortable and personalized experience in the XR environment, as the system can actively adjust or filter out elements that cause stress, enhancing overall user well-being and immersion.
[0011] In some embodiments, the at least physiological sensor is configured to determine, as the physiological parameter, at least one of a heart rate signal, a blood pressure signal, a cortisol level (a stress hormone) signal, a galvanic skin response (GSR, which measures skin conductivity) signal, an electroencephalography (EEG, which monitors brain activity) signal, an electromyography (EMG, which tracks muscle activity) signal, and an eye activity signal. An advantage of this measurement capability may be the ability to obtain a detailed and accurate assessment of the user's physiological state from one or even multiple angles. By monitoring a range of signals, the system can more reliably detect stress responses and provide more precise adjustments to the XR environment, ensuring a more personalized and effective reduction of stress-inducing stimuli.
[0012] In some embodiments, the processing circuit is configured to determine an auditory stimulus in the audio signal based on an entrainment of the at least one physiological parameter with a frequency and / or phase the auditory stimulus. That is, the processing circuit may be designed to identify an auditory stimulus within the audio signal by analyzing how the user's physiological parameters, such as heart rate or brain activity, synchronize with the frequency and / or phase of the auditory stimulus. This synchronization, known as entrainment, may help the system pinpoint which specific sounds are affecting the user's physiological state. An advantage of this approach is that it may allow for identification of the specific sounds that are causing stress or discomfort. By accurately targeting these auditory stimuli, the system can selectively filter or modify them, enhancing the user's comfort and experience in the XR environment without affecting other, non-stressful sounds.
[0013] In some embodiments, the processing circuit is configured to determine a visual stimulus in the video signal based on an entrainment of the at least one physiological parameter with an appearance of the visual stimulus within a visual field (e.g., including foveal and / or peripheral vision) of the user. Thus, the processing circuit may identify a visual stimulus within the video signal by examining how the user's physiological parameters, such as eye activity or brain signals, synchronize with the appearance of the visual stimulus within the user's visual field, which may include both central (foveal) and peripheral vision. This synchronization, or entrainment, may help the system determine which visual elements are influencing the user's physiological state. An advantage of this may be that it may enable precise detection of specific visual stimuli that cause stress or discomfort. By identifying and potentially altering or filtering these visual elements, the system can improve the user's experience in the XR environment, reducing visual stress and enhancing overall immersion and comfort.
[0014] In some embodiments, the at least physiological sensor is configured to determine magnetencephalography (MEG) or EEG data. The processing circuit may be configured to correlate the audio and / or video signal with brainwave responses using the MEG and / or EEG data. Thus, the physiological sensor may be designed to capture magnetoencephalography (MEG) or electroencephalography (EEG) data, which monitor brain activity. The processing circuit may then correlate the audio and / or video signals with the brainwave responses detected by the MEG and / or EEG data. This correlation may help identify which auditory or visual stimuli are affecting the user's brain activity. An advantage of this is the ability to obtain highly detailed and accurate insights into how specific stimuli impact the user's brain. This may allow for more precise adjustments to the XR environment, effectively reducing stress-inducing stimuli and enhancing the user's overall experience and comfort.
[0015] In some embodiments, the processing circuit is configured to decompose the MEG and / or EEG signal into auditory and visual brainwave responses, correlate the user’s auditory brainwave response with the audio signal, and correlate the user’s visual brainwave response with the video signal. Thus, the processing circuit may be designed to break down the MEG and / or EEG signals into separate auditory and visual brainwave responses. It may then correlate the user's auditory brainwave response with the audio signal and the visual brainwave response with the video signal. This process may help identify how specific sounds and visuals are influencing the user's brain activity. An advantage of this is that it may allow for a detailed and nuanced understanding of how different types of stimuli affect the user's mental state. By distinguishing between auditory and visual influences, the system can more effectively tailor the XR environment to minimize stress and enhance the user's experience.
[0016] In some embodiments, the apparatus further comprises a classification circuit configured to classify auditory stimuli of the audio signal and / or visual stimuli of video signal. The processing circuit may be configured to correlate the measured at least one physiological parameter with the one or more classified auditory and / or visual stimuli. That is the apparatus may include a classification circuit that categorizes auditory stimuli in the audio signal and visual stimuli in the video signal. The processing circuit then correlates the user's physiological measurements, such as heart rate or brain activity, with these categorized auditory and visual stimuli. This means the system can identify which specific types of sounds and visuals are causing changes in the user's physiological state. An advantage of this approach is that it may allow for a more sophisticated and precise identification of stress-inducing stimuli. By classifying and correlating specific auditory and visual inputs with physiological responses, the system can more accurately adjust the XR environment to reduce stress and enhance the user's comfort and overall experience.
[0017] In some embodiments, the apparatus further comprises a camera to capture the scene and output the video signal. The apparatus may an eye tracking circuit configured to determine the user’s foveal vision. The processing circuit may be configured to map the user’s foveal vision to the video signal to determine which part of the scene is within the user’s foveal vision. Thus, the apparatus may include a camera to capture the scene and produce the video signal. Additionally, it may feature an eye-tracking circuit that determines where the user's foveal vision, or central focus, is directed. The processing circuit then maps this foveal vision to the video signal to identify which part of the scene the user is directly looking at. An advantage of this system is that it may allow for precise identification of the visual elements the user is focusing on. This enables the apparatus to make targeted adjustments to reduce stress-inducing visuals within the user's direct line of sight, thereby enhancing comfort and immersion in the XR environment.
[0018] In some embodiments, the apparatus further comprises a classification circuit configured to classify one or more objects / features in the scene based on the video signal. The processing circuit may be configured to correlate the measured at least one physiological parameter with one or more classified objects / features within the user’s foveal vision. Thus, the apparatus may include a classification circuit that identifies and categorizes various objects or features in the scene based on the video signal. The processing circuit may then correlate the user's physiological measurements, such as stress levels or heart rate, with the classified objects or features that are within the user's foveal vision, or direct line of sight. An advantage of this approach is that it may enable the system to pinpoint which specific objects or features in the user's central vision are causing physiological stress. This precise identification may allow for more effective adjustments to the visual elements in the XR environment, thereby reducing stress and improving the user's overall experience.
[0019] In some embodiments, the classification circuit is configured to translate the video signal into text describing the one or more objects in the scene. Thus, the classification circuit may be designed to convert the video signal into text descriptions of the various objects present in the scene. This means the system can automatically generate textual representations of what the camera captures, detailing the characteristics and identities of these objects. An advantage of this feature is that it may allow for a clear and detailed understanding of the visual environment, which can be used to accurately correlate specific objects with the user's physiological responses. This textual information can enhance the system's ability to identify and mitigate stress-inducing elements in the XR environment, leading to a more comfortable and personalized user experience.
[0020] In some embodiments, the classification circuit comprises a trained machine learning network configured for video captioning. Thus, the classification circuit may include a trained machine learning network, such as a Convolutional Neural Network (CNN), specifically designed for video captioning. This means the network may analyze the video signal and generate descriptive text captions for the objects and features within the scene. An advantage of using a trained machine learning network for video captioning is its ability to accurately and efficiently process complex visual data and provide detailed descriptions. This may enhance the system's capability to identify and respond to stress-inducing visual stimuli, thereby improving the user's overall experience in the XR environment.
[0021] In some embodiments, the processing circuit is configured to detect the stress level caused by an object in the scene based on the classified objects, the user’s foveal vision (what they are directly looking at), and the at least one physiological parameter, such as heart rate or brain activity. This means the system can determine which specific objects in the user’s direct line of sight are causing stress. An advantage of this capability is that it may allow for identification of stress-inducing objects, enabling the system to make targeted adjustments to the visual environment. This can help reduce stress and enhance the user's comfort and experience in the XR environment.
[0022] In some embodiments, the apparatus further comprises an output interface configured to output information on the scene and / or auditory stimuli and / or visual stimuli causing the stress level above the threshold. This means the system can communicate details about which specific elements in the environment are inducing stress. An advantage of this feature is that it may allow users or other systems to be informed about the stress-inducing factors, enabling them to take appropriate actions to mitigate these elements and improve the user's experience and well-being in the XR environment.
[0023] In some embodiments, the apparatus further comprises an augmented reality (AR) processing circuit configured to modify the audio and / or video signal if the stress level is above the threshold. This means the system can actively adjust what the user hears and sees in real-time to reduce stress. An advantage of this feature is that it may allow for immediate and personalized modifications to the XR environment, enhancing the user's comfort and reducing stress by directly addressing and altering the stimuli causing discomfort.
[0024] In some embodiments, the AR processing circuit is configured to selectively modify the auditory and / or visual stimuli causing the stress level above the threshold. This means the system may target and modify only those elements that are identified as stress-inducing. An advantage of this feature is that it may allow for precise adjustments, ensuring that only the problematic stimuli are changed while leaving other, non-stressful elements unaffected. This may result in a more personalized and effective reduction of stress without compromising the overall XR experience.
[0025] In some embodiments, the AR processing circuit is configured to selectively cancel the auditory and / or visual stimuli causing the stress level above the threshold from the audio and / or video signal. Thus, the AR processing circuit may be designed to specifically remove the auditory and / or visual stimuli that are causing the user's stress level to exceed the threshold from the audio and / or video signal. This means the system can selectively eliminate the stress-inducing elements from what the user hears and sees. An advantage of this feature is that it may allows for the precise removal of only the problematic stimuli, effectively reducing stress without affecting other aspects of the audio and visual experience. This enhances the user's comfort and maintains the overall integrity of the XR environment. According to a further aspect, the present disclosure proposes an apparatus for XR perception of a user. The apparatus comprises an input interface configured to receive an audio signal of a scene and to receive a video signal of the scene. The apparatus further comprises a trained machine learning model configured to predict the user’s stress level based on the audio and / or video signal. The apparatus further comprises an AR processing circuit configured to modify the audio and / or video signal if the predicted stress level is above a threshold. An advantage of this may be that it proactively adjusts the XR environment to prevent stress, ensuring a more comfortable and personalized user experience by anticipating and mitigating stress-inducing stimuli before they cause significant discomfort.
[0026] In some embodiments, the AR processing circuit is configured to selectively modify auditory and / or visual stimuli causing the stress level above the threshold.
[0027] In some embodiments, the AR processing circuit is configured to selectively cancel the auditory and / or visual stimuli causing the stress level above the threshold from the audio and / or video signal.
[0028] In some embodiments, the apparatus further comprises a classification circuit configured to translate the audio and / or video signal into text describing the audio and / or video signal. The trained machine learning model may be configured to predict the user’s stress level based on the text describing the audio and / or video signal. The AR processing circuit may be configured to modify auditory and / or visual stimuli described in the audio and / or video signal and causing the stress level above the threshold.
[0029] According to a further aspect, the present disclosure proposes a method for XR perception of a user. The method includes receiving an audio signal of a scene, receiving a video signal of the scene, measuring at least one physiological parameter of the user while the user is experiencing the scene, correlating the measured at least one physiological parameter of the user with the audio and video signal, and, based on the correlation, determining auditory and / or visual stimuli causing a stress level of the user above a threshold.
[0030] In some embodiments, the method further includes training at least one machine learning model to predict the user’s stress level related to the audio and / or video signal based on the audio and / or video signal and based on the determined stress level.
[0031] In some embodiments, the method further includes translating, using at least one trained machine learning classification network, the audio and / or video signal into text describing the scene, and determining the auditory and / or visual stimuli causing the user’s stress level above the threshold based on the at least one physiological parameter and the text describing the scene.
[0032] In some embodiments, the method further includes translating, using the trained machine learning classification network, the video signal into text describing at least one object in the scene, determining that the user is gazing at the object, measuring the at least one physiological parameter of the user while the user is gazing at the object, and detecting the stress level related to the object using the physiological parameter.
[0033] In some embodiments, the method further includes modifying the audio and / or video signal in response to a stress level above the threshold.
[0034] The proposed solution proposes noise and visual stimulus cancellation in XR (Extended Reality) settings, specifically targeting stimuli that induce negative emotions such as stress or disgust.
[0035] Brief description of the Figures
[0036] Some examples of apparatuses and / or methods will be described in the following by way of example only, and with reference to the accompanying figures, in which
[0037] Fig. 1 shows a schematic block diagram of an apparatus for enhancing XR perception of a user;
[0038] Fig. 2 shows an example implementation of the apparatus including earphones in combination with a XR headset;
[0039] Fig. 3A, B show embodiments of EEG sensors;
[0040] Fig. 4 illustrates an example determining an auditory stimulus in an audio signal based on an entrainment of an EEG signal; and
[0041] Fig. 5 shows a schematic flowchart of a method for enhancing XR perception of a user according to an embodiment.
[0042] Detailed Description Some examples are now described in more detail with reference to the enclosed figures. However, other possible examples are not limited to the features of these embodiments described in detail. Other examples may include modifications of the features as well as equivalents and alternatives to the features. Furthermore, the terminology used herein to describe certain examples should not be restrictive of further possible examples.
[0043] Throughout the description of the figures same or similar reference numerals refer to same or similar elements and / or features, which may be identical or implemented in a modified form while providing the same or a similar function. The thickness of lines, layers and / or areas in the figures may also be exaggerated for clarification.
[0044] When two elements A and B are combined using an “or”, this is to be understood as disclosing all possible combinations, i.e. only A, only B as well as A and B, unless expressly defined otherwise in the individual case. As an alternative wording for the same combinations, "at least one of A and B" or "A and / or B" may be used. This applies equivalently to combinations of more than two elements.
[0045] If a singular form, such as “a”, “an” and “the” is used and the use of only a single element is not defined as mandatory either explicitly or implicitly, further examples may also use several elements to implement the same function. If a function is described below as implemented using multiple elements, further examples may implement the same function using a single element or a single processing entity. It is further understood that the terms "include", "including", "comprise" and / or "comprising", when used, describe the presence of the specified features, integers, steps, operations, processes, elements, components and / or a group thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, processes, elements, components and / or a group thereof.
[0046] Fig. 1 shows a schematic block diagram of an apparatus 10 for enhancing XR perception of a user. XR perception refers to a way users experience and interpret the virtual, augmented, or mixed reality environments created by extended reality (XR) technologies. This perception may be shaped by immersive qualities of apparatus 10, including visual, auditory, and potentially haptic feedback, which together may create a convincing and interactive digital experience. An effectiveness of XR perception relies on the apparatus' ability to seamlessly blend virtual elements with the real world, ensuring that users feel present and engaged in the simulated environment. Examples of apparatus 10 (XR system) include virtual reality (VR) headsets or glasses in combination with earphones / earbuds, which may immerse users in a digital environment by covering their field of vision and optionally incorporating hand controllers for interaction. Other examples of apparatus 10 may include smartphones, notebooks, laptops PCs, TVs, and the like. Augmented reality (AR) systems may overlay digital information onto the real world, enhancing the user's perception without completely replacing it. Mixed reality (MR) systems may blend both virtual and real worlds, allowing digital and physical objects to interact in real time. These systems may employ various sensors, cameras, and software to create and maintain the immersive experience.
[0047] Apparatus 10 comprises an input interface configured to receive an audio signal 11 of a scene or environment. An input interface for receiving the audio signal 11 may be a component of a device or system designed to capture and process sound from the environment. This interface may include one or more microphones that pick up audio signals from various sources in the scene. The microphones may convert sound waves into electrical signals, which are then processed by the device. The input interface may also include hardware and software components to filter, amplify, and digitize the audio signals for further analysis and processing. For example, in an XR headset, the input interface may comprise built-in microphones that capture ambient sounds, dialogue, and other auditory elements from the user's surroundings.
[0048] Input interface of apparatus 10 is further configured to receive a video signal 12 of the (same) scene or environment. For example, input interface may comprise a camera and / or a display. An input interface for receiving the video signal 12 may be a component designed to capture and process visual information from the environment. This interface may include a camera or an array of cameras that record the visual aspects of the scene. The captured video data may then be converted into a digital format that can be processed by the device. For instance, in an XR headset, the input interface for video might comprise of forwardfacing cameras that capture the user's surroundings. These cameras may record visual information, which is then digitized and processed by the device's internal systems. In a television (TV), the input interface for receiving video and audio signals may include several types of ports and connectors that allow the TV to receive signals from various external sources.
[0049] The scene, in the context of audio and visual media, may refer to a specific segment or sequence within a larger piece of content, such as a movie, video, game, or virtual environment. It may comprise of a series of actions, events, or interactions that take place in a par- ticular setting and time frame. In visual media, a scene may include the visual elements such as characters, objects, and background, as well as the lighting, camera angles, and movements. In audio media, it may include the sounds, dialogue, music, and effects that are designed to accompany and enhance the visual elements.
[0050] The audio signal 11 may be an electrical or digital representation of sound, which is created when sound waves are converted into a format that can be processed, transmitted, or recorded. This conversion may involve one or more microphones or other transducers that capture sound waves and convert their pressure variations into corresponding electrical voltage changes. Audio signals can be analog, where the signal is a continuous wave that directly corresponds to the sound wave, or digital, where the signal is represented by discrete values through sampling and quantization.
[0051] The video signal 12 may be an electrical signal that represents visual information for display on electronic devices such as televisions, monitors, or projectors. This signal can be transmitted in either analog or digital format. Analog video signals, like those used in traditional television broadcasting, represent images through continuous variations in voltage or current, corresponding to the brightness (luminance) and color (chrominance) information of the picture. Examples include composite video, component video, and S-Video. Digital video signals, used in modern devices and streaming technologies, represent images using binary data, consisting of discrete values. This format allows for higher quality and more efficient transmission and storage. Examples include HDMI, DVI, and DisplayPort.
[0052] Apparatus 10 further comprises one or more physiological sensors 13 configured to measure at least one physiological parameter 14 of the user while the user is experiencing the scene. Physiological sensor 13 may be an integral or external part of a VR headset, for example. Physiological sensor 13 is a device designed to detect and measure specific physiological parameters or signals from the human body. Such sensors can monitor a wide range of bodily functions and conditions, providing critical data for medical diagnosis, treatment, and research. Examples of physiological sensor 13 include electrocardiogram (ECG) sensors, which measure electrical activity of the heart to monitor heart rate and rhythm; electromyography (EMG) sensors, which detect electrical activity produced by skeletal muscles to assess muscle function; electroencephalogram (EEG) sensors, which record electrical activity of the brain to study brain function and diagnose neurological conditions; magnetoencephalography (MEG) sensors, which are used to measure the magnetic fields produced by neural activity in the brain; pulse oximeters, which measure oxygen saturation levels in the blood and heart rate; temperature sensors, which monitor body temperature to detect fever or hypothermia; blood pressure monitors, which measure the force of blood against the walls of arteries to assess cardiovascular health; glucose sensors, which monitor blood sugar levels, essential for managing diabetes; and respiratory sensors, which measure breathing rate and pattern to assess respiratory function.
[0053] At least some of the mentioned sensors can be implemented in wearables (e.g., smartwatches, wristbands, etc.) or VR headsets to enhance functionality and provide real-time monitoring of physiological parameters. For example, electrocardiogram (ECG) sensors can be integrated into wearable devices such as smartwatches or fitness trackers to continuously monitor heart rate and detect irregular heart rhythms. Electromyography (EMG) sensors can be embedded in smart clothing or accessories to track muscle activity and provide feedback for athletic training or rehabilitation. Electroencephalogram (EEG) sensors may be used in headsets to monitor brain activity, which can be particularly useful in VR applications for brain-computer interface (BCI) technologies. These can enable users to control virtual environments using their thoughts or assess cognitive load and mental states during VR experiences. Pulse oximeters can be included in wearables to monitor blood oxygen levels and heart rate, while temperature sensors can track body temperature to detect fever or other health changes. Blood pressure monitors may also be implemented in wearable devices, allowing for continuous cardiovascular monitoring. Glucose sensors, which are vital for diabetes management, are being miniaturized and integrated into wearable patches or devices that can continuously monitor blood sugar levels. Respiratory sensors can be included in wearables to measure breathing rate and patterns, providing valuable data for respiratory health and fitness tracking.
[0054] Apparatus 10 further comprises a processing circuit 15 which is configured to correlate the measured at least one physiological parameter 14 of the user with the audio signal 11 , the video signal 12, or both. Based on the correlation, processing circuit 15 is configured to determine auditory and / or visual stimuli 16 causing a stress level of the user above a (stress) threshold. Examples of processing circuit 15 include digital processing circuits which may perform a variety of functions from simple computations to complex data processing tasks. Microprocessors, for instance, are the central processing units (CPUs) found in computers, smartphones, and other electronic devices. They may perform a wide range of tasks by executing instructions from software programs, handling complex calculations, data processing, and control tasks. Digital Signal Processors (DSPs) are specialized microprocessors designed specifically for high-speed numerical calculations, which may be used in audio, video, and communications applications. They may process signals in real-time, performing tasks such as filtering, compression, and modulation. Field-Programmable Gate Arrays (FPGAs) are integrated circuits that can be configured by the user after manufacturing. They may be used in a variety of applications, from prototyping new digital circuits to performing specific processing tasks in high-performance computing systems. Application- Specific Integrated Circuits (ASICs) are custom-designed chips optimized for specific applications or tasks. Unlike general-purpose processors, ASICs are tailored for particular functions, making them more efficient for those tasks. They may be used in consumer electronics, automotive systems, and telecommunications. Graphics Processing Units (GPUs) are specialized processors designed to handle graphics rendering tasks. They are highly parallelized, making them suitable for processing large blocks of data simultaneously. GPUs may be used in gaming, graphics design, and increasingly in machine learning and artificial intelligence applications.
[0055] Correlating the measured at least one physiological parameter of the user with the audio signal 11 and / or the video signal 12 means to analyze and establish a relationship between a user's physiological data (such as heart rate, brain activity, or muscle movements, for example) and the audio or video content they are interacting with. This process may involve measuring the user's physiological responses and then comparing these measurements to the corresponding audio or video signals to understand how the user's body reacts to different stimuli. For example, if a person is watching a movie (video signal) while their heart rate (physiological parameter) is being monitored, correlating these two could help determine how certain scenes affect the viewer's emotional state. Similarly, if a person is listening to music (audio signal) while their brain activity (physiological parameter) is measured, this correlation could reveal how different types of music influence cognitive or emotional responses.
[0056] The processing circuit 15 may monitor physiological data 14 such as heart rate, skin conductance, or brain activity to detect signs of stress. It may then compare these stress indicators with the audio or video content being presented to the user. If the stress levels surpass a threshold, the processing circuit 15 may identify the particular stimuli responsible for this elevated stress. This information can be used to modify the (virtual) environment, provide feedback to the user, or make adjustments to the content to reduce stress and improve user experience.
[0057] Fig. 2 shows an example implementation of apparatus 10. A user 20 wears ear- phones / earbuds 22 in combination with a XR headset or glasses 21. XR headset 21 together with the earphones 22 may form apparatus 10. As shown in Fig. 3A and 3B, XR headset 21 or earphones 22 can comprise a physiological sensor 13 in form of an EEG sensor 33. For example, Fig. 3A depicts user 20 wearing an EEG (Electroencephalogram) cap with electrodes attached to their head, which may be used to measure brain activity. EEG can also be performed without a cap by individually placing electrodes on the scalp. XR headset 21 may include EEG electrodes acting as EEG sensor 33.
[0058] For another example, Fig. 3B depicts an EEG sensor 33 designed to be placed behind an ear. The shape of the EEG sensor 33 is contoured to fit the area behind the ear, ensuring secure and stable placement. The EEG sensor 33 includes several electrodes. These electrodes may be positioned to capture brain activity from the region behind the ear. On the right side of Fig. 3B, there are example EEG waveforms corresponding to each electrode. This setup may be useful for capturing brain activity from a less intrusive location compared to traditional scalp electrodes shown in Fig. 3A. Earphones 22 may include such an EEG sensor 33, for example. The skilled person having benefit form the present disclosure will appreciate that there are many other options for alternative or additional physiological sensors, such as heart rate sensors, blood pressure sensors, cortisol level sensors, GSR sensors, MEG sensors, EMG sensors, or eye activity sensors.
[0059] If desired, user 20 can select an intelligent noise / video cancellation modus of apparatus 10.
[0060] During the noise / video cancellation modus, processing circuit 15 may be configured to decompose EEG signal 14 into auditory and visual brainwave responses, for example. This process may involve analyzing the EEG signal 14 to isolate the neural activity related to different types of sensory stimuli. Firstly, the EEG signal 14, which represents the overall electrical activity of the brain, may be recorded while the user is exposed to various auditory and visual stimuli. This raw EEG data may contain overlapping brainwave patterns generated by the brain's response to these different stimuli. Processing circuit 15 may apply signal processing techniques to this EEG data to decompose it into its constituent parts. Techniques such as Independent Component Analysis (ICA) or Principal Component Analysis (PCA) may be used. These methods may identify and separate the independent sources of brain activity within the mixed EEG signals. For auditory stimuli, processing circuit 15 may look for brainwave patterns associated with hearing and processing sounds. This may involve identifying specific features in the EEG signal, such as the P300 wave, which is a known marker of attention to auditory stimuli. For visual stimuli, processing circuit 15 may focus on brainwave patterns related to visual processing. This can involve analyzing responses in the visual cortex, such as visual evoked potentials (VEPs), which are specific changes in the EEG signal triggered by visual events. By decomposing the EEG signal 14 into these separate components, processing circuit 15 can effectively differentiate between the brain's responses to audio signal 11 and video signal 12. This decomposition allows for a more detailed analysis of how the user reacts to different types of stimuli, enabling the system to tailor the environment in ways that minimize stress and enhance user comfort.
[0061] Processing circuit 15 may be configured to correlate the user’s auditory brainwave response with the audio signal 11. There may be several options for configuring processing circuit 15 to correlate the user's auditory brainwave response with the audio signal 11.
[0062] One approach may involve the use of time-frequency analysis, where the processing circuit 15 may apply techniques such as Short-Time Fourier Transform (STFT) or wavelet transform to both the EEG signals 14 and the audio signals 11. This allows processing circuit 15 to analyze how specific frequency components of the brainwave responses correspond to the frequency components of the audio signal over time, identifying patterns of synchronization or entrainment.
[0063] Another option may be the application of machine learning algorithms. Processing circuit 15 may comprise a trained machine learning model configured to predict the user’s stress level based on the audio signal 11, and an AR processing circuit configured to modify / cancel the audio stimuli in the audio signal if the predicted stress level is above a threshold. The processing circuit 15 can employ supervised learning techniques, where it is trained on a dataset of EEG recordings and corresponding audio stimuli. Neural networks or support vector machines can be used to learn the complex relationships between the auditory brainwave responses and the features of the audio signal, such as amplitude, frequency, and temporal structure. Once trained, the model can predict how new audio stimuli are likely to affect the user's brainwave patterns.
[0064] A further approach may involve using phase-locking value (PLV) analysis. The processing circuit 15 may calculate a PLV between the EEG signal 14 and the audio signal 11 to measure the degree of synchronization between the phase of the brainwaves and the phase of the auditory stimuli. High PLV values indicate strong phase synchronization, suggesting that the auditory brainwave responses are closely tied to the audio signal.
[0065] Further, coherence analysis may be employed. The processing circuit 15 may compute the coherence between the EEG signal 14 and the audio signal 11, which measures the linear relationship between their frequency components. High coherence values may indicate a strong correlation between specific frequency bands of the brainwave responses and the audio signal, helping to identify which aspects of the sound are influencing brain activity.
[0066] Cross-correlation analysis may be another method. The processing circuit 15 may calculate the cross-correlation function between the EEG signal 14 and the audio signal 11 to determine the time lag at which the correlation is maximized. This may identify a temporal relationship between the auditory brainwave responses and the audio stimuli, revealing how quickly the brain responds to different sounds.
[0067] Based on the frequency distribution and temporal predictability of the auditory stimulus, the disturbing noise may be cancelled actively by the headphones 22, for example.
[0068] Processing circuit 15 may be configured to correlate the user’s visual brainwave response with the video signal. There may be several options for configuring processing circuit 15 to correlate the user's visual brainwave response with the video signal 12.
[0069] One approach may involve using event-related potential (ERP) analysis, where the processing circuit 15 may identify specific time points in the video signal 12, such as the onset of visual stimuli. The EEG data may then be segmented around these time points, and the processing circuit 15 may analyze the corresponding brainwave responses, such as visual evoked potentials (VEPs), to determine how the visual stimuli affect brain activity.
[0070] Another option may be time-frequency analysis, similar to the method used for auditory signals. The processing circuit 15 can apply techniques like Short-Time Fourier Transform (STFT) or wavelet transform to both the EEG signals and the video signal. This allows the circuit to analyze how specific frequency components of the brainwave responses correspond to the visual elements in the video, identifying patterns of synchronization or entrainment.
[0071] Machine learning algorithms can also be employed. Processing circuit 15 may comprise a trained machine learning model configured to predict the user’s stress level based on the video signal 12, and an AR processing circuit configured to modify / cancel the visual stimuli in the video signal 12 if the predicted stress level is above a threshold The processing circuit 15 can use supervised learning techniques, training a model on a dataset of EEG recordings and corresponding video stimuli. Neural networks or support vector machines can learn the complex relationships between visual brainwave responses and features of the video signal, such as color, movement, and patterns. The trained model can then predict how new visual stimuli are likely to affect the user's brainwave patterns.
[0072] Phase-locking value (PLV) analysis may be another method. The processing circuit 15 may calculate the PLV between the EEG signal and the video signal 12 to measure the degree of synchronization between the phase of the brainwaves and the phase of the visual stimuli. High PLV values may indicate strong phase synchronization, suggesting that the visual brainwave responses are closely tied to the video signal.
[0073] Coherence analysis may also be used. The processing circuit 15 may compute the coherence between the EEG signal 14 and the video signal 12, measuring the linear relationship between their frequency components. High coherence values indicate a strong correlation between specific frequency bands of the brainwave responses and the video signal, helping to identify which visual aspects are influencing brain activity.
[0074] Cross-correlation analysis may be another technique. The processing circuit 15 may calculate the cross-correlation function between the EEG signal 14 and the video signal 12 to determine a time lag at which the correlation is maximized. This may identify the temporal relationship between the visual brainwave responses and the visual stimuli, revealing how quickly the brain responds to different elements in the video.
[0075] Fig. 4 illustrates an example, where processing circuit 15 is configured to determine an auditory stimulus in the audio signal 11 based on an entrainment of the EEG signal 14 with a frequency and / or phase the auditory stimulus. Processing circuit 15 may employ a technique known as entrainment, where it synchronizes or aligns the EEG signal 14 with the frequency and / or phase of the auditory stimuli in audio signal 11. Entrainment involves observing how the user's brain waves (EEG signals) exhibit patterns that align with the temporal characteristics (frequency and phase) of the sounds being heard. By identifying these correlations, the system can pinpoint which auditory stimuli are affecting the user's physiological state, such as causing stress or discomfort. Depending on the timing of the audio stimuli, EEG signal 14 may be segmented to correspond with the specific instances when the stimuli occur. This process may involve identifying exact moments, or onsets, at which the auditory stimuli begin. Once these onsets are determined, the EEG signal 14 may be "cut" or segmented around these points, creating segments of EEG recordings that are time-locked to the stimuli. Within these segments, a specific component of the EEG signal 14 known as the P300 may be analyzed. The P300 is a phase-locked brain response that occurs approximately 300 milliseconds after an individual pays attention to a stimulus. This response is characterized by a positive deflection in the EEG signal 14 and is considered an indicator of cognitive processes related to attention and the processing of relevant stimuli. By analyzing the P300 response, the system can determine which auditory stimuli the user is focusing on or paying attention to. This analysis may involve examining the amplitude and timing of the P300 wave within the segmented EEG data, providing insights into how the brain is reacting to the auditory stimuli. This information can then be used to understand the user's auditory preferences and stress triggers, allowing the system to intelligently filter or modify the audio input to enhance the user's experience.
[0076] As shown in Fig. 4, in some embodiments, incoming sounds from the external world may be classified (e.g., a train is passing by, there is a nail hammered into the wall, there are children playing, clock is ticking load....). For this purpose, the apparatus 10 may further comprise a classification circuit configured to classify auditory stimuli of the audio signal 11 and / or visual stimuli of video signal 12 that have been identified based on entrainment. The classification circuit may be part of processing circuit 15. Processing circuit 15 may be configured to correlate the measured physiological parameter(s) with the one or more classified auditory and / or visual stimuli to identify stressors.
[0077] The classification circuit may be designed to categorize incoming sounds from the external environment. This classification may involve identifying and labeling various auditory stimuli within the audio signal, such as the sound of a train passing by, a nail being hammered into a wall, children playing, or a clock ticking. The (trained) classification circuit may process the audio signal 11 to detect and distinguish these specific sounds based on their acoustic characteristics, such as frequency, amplitude, and temporal patterns.
[0078] In addition to auditory stimuli, the classification circuit may also be configured to categorize visual stimuli from the video signal 12. This may involve analyzing the video signal 12 to identify and label different elements within the scene, such as moving objects, color patterns, and textures. Processing circuit 15 may correlate the measured physiological parameters, such as EEG signals or other biometric data, with the classified auditory and visual stimuli. This correlation process may involve examining how the user's physiological responses align with the identified sounds and visual elements. By doing so, the processing circuit 15 can determine which specific stimuli are causing particular physiological reactions, such as increased stress or heightened attention. For example, if the classification circuit identifies the sound of a train passing by and the processing circuit 15 detects a corresponding spike in the user's stress levels as indicated by the EEG signals 14, it can conclude that this auditory stimulus is a stressor for the user 20. Similarly, if a visual stimulus, like a flashing light, is identified and correlated with a physiological response, the system can understand the impact of that visual input on the user 20.
[0079] A combination with an eye-tracking circuit may be used to detect negative emotion causing stimuli (e.g., avoiding looking at the part of the scene where the natural gaze would go over). The eye tracking circuit may be configured to determine the user’s foveal vision. For example, the eye-tracking circuit may employ infrared light or near-infrared LEDs to illuminate the user’s eyes. Infrared cameras or sensors may then capture high-resolution images of the eyes, focusing particularly on the pupils and the reflections of the infrared light, known as Purkinje images, on the cornea. The captured images may be processed using advanced image processing algorithms to detect the position and movements of the pupils. The algorithms may calculate the center of the pupil and the corneal reflections, determining the line of sight or gaze direction. By continuously tracking these parameters, the eyetracking circuit can accurately determine where the user is looking at any given moment. To identify the user's foveal vision, the circuit may map the gaze direction onto the visual scene or display the user is viewing. This may involve geometric calculations that consider the position of the eyes, the orientation of the head, and the distance to the screen or objects in the environment. The circuit may then determine the exact point of focus within the user's field of view, which corresponds to the foveal region.
[0080] The processing circuit 15 may be configured to map the user’s foveal vision to the video signal 12 to determine which part of the scene is within the user’s foveal vision. Processing circuit may determine gaze direction data and use it to identify the point of focus, which corresponds to the user's foveal vision. The video signal 12, which captures the visual scene, may be processed frame by frame to identify the spatial coordinates of visual elements within the scene. The processing circuit 15 may use geometric transformation algorithms to align the coordinate system of the gaze data with that of the video signal 12. This may involve taking into account the position and orientation of the user's head and the distance to the screen or objects in the environment. Once the gaze direction is mapped onto the coordinate system of the video signal 12, the processing circuit 15 may determine the specific area of the video frame that corresponds to the user's foveal vision. This area may be a small, central region of the frame where visual acuity is highest, allowing the system to identify which parts of the scene the user is focusing on.
[0081] Processing circuit 15 may be configured to classify one or more objects / features in the scene based on the video signal 12 and correlate the measured at least one physiological pa- rameter 14 (e.g., EEG signal) with one or more classified objects / features within the user’s foveal vision.
[0082] Based on the features of the visual input and the corresponding stress level, the visual scene may be adapted (e.g., the pink car will turn blue, blood from a wound may be overlayed with skin tissue in people who cannot see blood). That is, apparatus 10 or processing circuit 15 may further comprise an AR processing circuit configured to modify the audio and / or video signal if the stress level is above the threshold. The AR processing circuit may be configured to selectively modify or cancel the auditory and / or visual stimuli causing the stress level above the threshold.
[0083] In some embodiments, processing circuit 15 may be configured to translate the video signal 12 of a captured scene into text describing the objects within the scene, the colors of the object, in case there are irregularities in the objects color, structure and texture, those will be also mentioned within the text (e.g., a car with only three wheels or a flashing light source). Processing circuit 15 may be configured to analyze the video signal 12 frame by frame using computer vision algorithms. These algorithms may detect and identify objects within each frame by comparing the visual data to a database of known object models. Techniques such as object detection and recognition, which may be implemented through convolutional neural networks (CNNs), may be used to achieve this. Once objects are identified, the processing circuit 15 may assess their attributes, such as shape, size, and position within the scene. It may then evaluate the colors of these objects using color detection algorithms, which may analyze the pixel values to determine dominant colors. If the system detects any irregularities in the objects, such as unusual color patterns, structural anomalies, or texture differences, these may be noted. For example, if an object has an unexpected color gradient, or a surface texture that differs from the typical model, these details may be highlighted.
[0084] The next step may involve natural language processing (NLP) to generate descriptive text. The identified objects and their attributes, including any noted irregularities, may be converted into a structured format that an NLP algorithm can process. The processing circuit 15 may then use this information to construct coherent sentences that describe the scene. For instance, it might generate a description like, "A red car with a dented hood and three wheels is parked next to a blue mailbox." To enhance accuracy, the processing circuit 15 may also incorporate contextual information from previous frames and adjacent objects to ensure that descriptions are consistent and comprehensive. Advanced NLP techniques, such as sequence-to-sequence models and transformers, can be employed to improve the fluency and readability of the generated text.
[0085] Apparatus 10 may further comprise an output interface configured to output information on the scene and / or auditory stimuli and / or visual stimuli causing the stress level above the threshold. Due to the classification of the audio sound / visual input, feedback can be given to the user (e.g., “you were disturbed by a dropping water in your bathroom sink”, “you were disturbed by the combination of shape and color within the car”) in combination with recommendations (e.g., “you can fix this by doing XY”) or also given to manufacturers of technical devices (“the sound of your fridge KS34E is disturbing 75% of your clients” I “the design of your car is disturbing people who have impaired color vision”). This information, if shared, could also help to find a good flat (e.g., “the traffic in the morning is disturbing 36% of the people living in this area”, “the architecture of this new building is inducing negative emotions in your citizens”) and optimize town planning. Beside detection of disturbing sounds / visuals, the proposed system can also be used for creating pleasurable mu- sic / visuals. By understanding which elements of the music / of the video are triggering the most attention, lead to relaxation or create more engagement in the participant.
[0086] Apparatus 10 may implement a method 50 for XR perception of a user as shown in Fig. 5.
[0087] Method 50 includes receiving 51 an audio signal and a video signal of a scene, measuring 52 at least one physiological parameter of the user while the user is experiencing the scene, correlating 53 the measured at least one physiological parameter of the user with the audio and video signal, and, based on the correlation, determining 54 auditory and / or visual stimuli causing a stress level of the user above a threshold.
[0088] The present disclosure proposes an apparatus and method for enhancing Extended Reality (XR) perception by identifying and mitigating disturbing audio-visual stimuli. The apparatus comprises an input interface for receiving audio and video signals of a scene, and one or more physiological sensors that measure parameters such as heart rate, blood pressure, cortisol levels, galvanic skin response, EEG, MEG, EMG, and eye activity. A processing circuit correlates these physiological parameters with the audio and / or video signals to determine which stimuli cause the user’s stress level to exceed a threshold.
[0089] The apparatus may capture audio and visual data from the user's environment using microphones and cameras. Physiological sensors may monitor the user's responses, and the processing circuit may analyze these responses to identify stress-inducing stimuli. Tech- niques such as Independent Component Analysis (ICA) or Principal Component Analysis (PCA) may be used to decompose the EEG signals into auditory and visual brainwave responses. The processing circuit may also employ machine learning algorithms, timefrequency analysis, phase-locking value (PLV) analysis, coherence analysis, and crosscorrelation analysis to correlate brainwave responses with audio and visual stimuli.
[0090] Additionally, the apparatus may include a classification circuit that categorizes these stimuli and correlates them with physiological responses to identify stressors. An eye-tracking circuit may determine the user's foveal vision, mapping it onto the video signal to pinpoint which part of the scene the user is focusing on. The apparatus can selectively modify or cancel stress-inducing auditory and visual stimuli in real-time, enhancing the user’s comfort and immersion in the XR environment.
[0091] Moreover, the apparatus may translate the video signal into text descriptions of the objects within the scene, including any irregularities. The apparatus may output information on stress-inducing stimuli to users or manufacturers for further action, and it can also be used to create pleasurable audio-visual experiences by identifying elements that trigger relaxation or engagement.
[0092] The present disclosure also proposes methods for training machine learning models to predict user stress levels based on audio and video signals, and for modifying these signals to reduce stress. This technology aims to provide a personalized and adaptive XR experience, improving user well-being and engagement.
[0093] In the following, some examples of the proposed concept are presented:
[0094] An example (e.g., example 1) relates to an apparatus for extended reality, XR, perception of a user, the apparatus comprising an input interface configured to receive an audio signal of a scene, receive a video signal of the scene, at least one physiological sensor configured to measure at least one physiological parameter of the user while the user is experiencing the scene, a processing circuit configured to correlate the measured at least one physiological parameter of the user with the audio and / or video signal, and based on the correlation, determine auditory and / or visual stimuli causing a stress level of the user above a threshold.
[0095] Another example (e.g., example 2) relates to a previous example (e.g., example 1) or to any other example, further comprising that at least physiological sensor is configured to determine, as the physiological parameter, at least one of a heart rate signal, a blood pressure signal, a cortisol level signal, a galvanic skin response, GSR, signal, an electroencephalography, EEG, signal, a magnetencephalography, MEG, signal, an electromyography, EMG, signal, and an eye activity signal.
[0096] Another example (e.g., example 3) relates to a previous example (e.g., one of the examples 1 or 2) or to any other example, further comprising that the processing circuit is configured to determine an auditory stimulus in the audio signal based on an entrainment of the at least one physiological parameter with a frequency and / or phase the auditory stimulus.
[0097] Another example (e.g., example 4) relates to a previous example (e.g., one of the examples 1 to 3) or to any other example, further comprising that the processing circuit is configured to determine a visual stimulus in the video signal based on an entrainment of the at least one physiological parameter with an appearance of the visual stimulus within a visual field of the user.
[0098] Another example (e.g., example 5) relates to a previous example (e.g., one of the examples 1 to 4) or to any other example, further comprising that the at least physiological sensor is configured to determine MEG or EEG data, and wherein the processing circuit is configured to correlate the audio and / or video signal with brainwave responses using the MEG or EEG data.
[0099] Another example (e.g., example 6) relates to a previous example (e.g., example 5) or to any other example, further comprising that the processing circuit is configured to decompose the EEG signal into auditory and visual brainwave responses, and correlate the user’s auditory brainwave response with the audio signal, and correlate the user’s visual brainwave response with the video signal.
[0100] Another example (e.g., example 7) relates to a previous example (e.g., one of the examples 1 to 6) or to any other example, further comprising a classification circuit configured to classify auditory stimuli of the audio signal and / or visual stimuli of video signal, and wherein the processing circuit is configured to correlate the measured at least one physiological parameter with the one or more classified auditory and / or visual stimuli.
[0101] Another example (e.g., example 8) relates to a previous example (e.g., one of the examples 1 to 7) or to any other example, further comprising a camera to capture the scene and output the video signal, an eye tracking circuit configured to determine the user’s foveal vision, wherein the processing circuit is configured to map the user’s foveal vision to the video signal to determine which part of the scene is within the user’s foveal vision.
[0102] Another example (e.g., example 9) relates to a previous example (e.g., example 8) or to any other example, further comprising a classification circuit configured to classify one or more objects / features in the scene based on the video signal, and wherein the processing circuit is configured to correlate the measured at least one physiological parameter with one or more classified objects / features within the user’s foveal vision.
[0103] Another example (e.g., example 10) relates to a previous example (e.g., example 9) or to any other example, further comprising that the classification circuit is configured to translate the video signal into text describing the one or more objects in the scene.
[0104] Another example (e.g., example 11) relates to a previous example (e.g., one of the examples 9 or 10) or to any other example, further comprising that the classification circuit comprises a trained machine learning network (e.g., CNN) configured for video captioning.
[0105] Another example (e.g., example 12) relates to a previous example (e.g., one of the examples 9 to 11) or to any other example, further comprising that the processing circuit is configured to detect the stress level caused by an object in the scene based on the classified objects, the user’s foveal vision, and the at least one physiological parameter.
[0106] Another example (e.g., example 13) relates to a previous example (e.g., one of the examples 1 to 12) or to any other example, further comprising an output interface configured to output information on the scene and / or auditory stimuli and / or visual stimuli causing the stress level above the threshold.
[0107] Another example (e.g., example 14) relates to a previous example (e.g., one of the examples 1 to 13) or to any other example, further comprising an augmented reality, AR, processing circuit configured to modify the audio and / or video signal if the stress level is above the threshold.
[0108] Another example (e.g., example 15) relates to a previous example (e.g., example 14) or to any other example, further comprising that the AR processing circuit is configured to selectively modify the auditory and / or visual stimuli causing the stress level above the threshold. Another example (e.g., example 16) relates to a previous example (e.g., example 15) or to any other example, further comprising that the AR processing circuit is configured to selectively cancel the auditory and / or visual stimuli causing the stress level above the threshold from the audio and / or video signal.
[0109] An example (e.g., example 17) relates to an apparatus for XR perception of a user, the apparatus comprising an input interface configured to receive an audio signal of a scene, receive a video signal of the scene, and a trained machine learning model configured to predict the user’s stress level based on the audio and / or video signal, and an AR processing circuit configured to modify the audio and / or video signal if the predicted stress level is above a threshold.
[0110] Another example (e.g., example 18) relates to a previous example (e.g., example 17) or to any other example, further comprising that the AR processing circuit is configured to selectively modify auditory and / or visual stimuli causing the stress level above the threshold.
[0111] Another example (e.g., example 19) relates to a previous example (e.g., example 18) or to any other example, further comprising that the AR processing circuit is configured to selectively cancel the auditory and / or visual stimuli causing the stress level above the threshold from the audio and / or video signal.
[0112] Another example (e.g., example 20) relates to a previous example (e.g., one of the examples 16 to 18) or to any other example, further comprising a classification circuit configured to translate the audio and / or video signal into text describing the audio and / or video signal, wherein the trained machine learning model is configured to predict the user’s stress level based on the text describing the audio and / or video signal, wherein the AR processing circuit is configured to modify auditory and / or visual stimuli described in the audio and / or video signal and causing the stress level above the threshold.
[0113] An example (e.g., example 21) relates to a method for XR perception of a user, the method comprising receiving an audio signal of a scene, receiving a video signal of the scene, measuring at least one physiological parameter of the user while the user is experiencing the scene, correlating the measured at least one physiological parameter of the user with the audio and video signal, and based on the correlation, determining auditory and / or visual stimuli causing a stress level of the user above a threshold. Another example (e.g., example 22) relates to a previous example (e.g., example 21) or to any other example, further comprising training at least one machine learning model to predict the user’s stress level related to the audio and / or video signal based on the audio and / or video signal and based on the determined stress level.
[0114] Another example (e.g., example 23) relates to a previous example (e.g., one of the examples 21 or 22) or to any other example, further comprising translating, using at least one trained machine learning classification network, the audio and / or video signal into text describing the scene, determining the auditory and / or visual stimuli causing the user’s stress level above the threshold based on the at least one physiological parameter and the text describing the scene.
[0115] Another example (e.g., example 24) relates to a previous example (e.g., one of the examples 21 to 23) or to any other example, further comprising translating, using the trained machine learning classification network, the video signal into text describing at least one object in the scene, determining that the user is gazing at the object, measuring the at least one physiological parameter of the user while the user is gazing at the object, and detecting the stress level related to the object using the physiological parameter.
[0116] Another example (e.g., example 25) relates to a previous example (e.g., one of the examples 21 to 24) or to any other example, further comprising modifying the audio and / or video signal in response to a stress level above the threshold.
[0117] The aspects and features described in relation to a particular one of the previous examples may also be combined with one or more of the further examples to replace an identical or similar feature of that further example or to additionally introduce the features into the further example.
[0118] Examples may further be or relate to a (computer) program including a program code to execute one or more of the above methods when the program is executed on a computer, processor or other programmable hardware component. Thus, steps, operations or processes of different ones of the methods described above may also be executed by programmed computers, processors or other programmable hardware components. Examples may also cover program storage devices, such as digital data storage media, which are machine-, processor- or computer-readable and encode and / or contain machineexecutable, processor-executable or computer-executable programs and instructions. Program storage devices may include or be digital storage devices, magnetic storage media such as magnetic disks and magnetic tapes, hard disk drives, or optically readable digital data storage media, for example. Other examples may also include computers, processors, control units, (field) programmable logic arrays ((F)PLAs), (field) programmable gate arrays ((F)PGAs), graphics processor units (GPU), application-specific integrated circuits (ASICs), integrated circuits (ICs) or system-on-a-chip (SoCs) systems programmed to execute the steps of the methods described above.
[0119] It is further understood that the disclosure of several steps, processes, operations or functions disclosed in the description or claims shall not be construed to imply that these operations are necessarily dependent on the order described, unless explicitly stated in the individual case or necessary for technical reasons. Therefore, the previous description does not limit the execution of several steps or functions to a certain order. Furthermore, in further examples, a single step, function, process or operation may include and / or be broken up into several sub-steps, -functions, -processes or -operations.
[0120] If some aspects have been described in relation to a device or system, these aspects should also be understood as a description of the corresponding method. For example, a block, device or functional aspect of the device or system may correspond to a feature, such as a method step, of the corresponding method. Accordingly, aspects described in relation to a method shall also be understood as a description of a corresponding block, a corresponding element, a property or a functional feature of a corresponding device or a corresponding system.
[0121] The following claims are hereby incorporated in the detailed description, wherein each claim may stand on its own as a separate example. It should also be noted that although in the claims a dependent claim refers to a particular combination with one or more other claims, other examples may also include a combination of the dependent claim with the subject matter of any other dependent or independent claim. Such combinations are hereby explicitly proposed, unless it is stated in the individual case that a particular combination is not intended. Furthermore, features of a claim should also be included for any other independent claim, even if that claim is not directly defined as dependent on that other independent claim.
Claims
Claims1. An apparatus for extended reality, XR, perception of a user, the apparatus comprising an input interface configured to receive an audio signal of a scene; receive a video signal of the scene; at least one physiological sensor configured to measure at least one physiological parameter of the user while the user is experiencing the scene; a processing circuit configured to correlate the measured at least one physiological parameter of the user with the audio and / or video signal, and based on the correlation, determine auditory and / or visual stimuli causing a stress level of the user above a threshold.
2. The apparatus of claim 1 , wherein at least physiological sensor is configured to determine, as the physiological parameter, at least one of a heart rate signal, a blood pressure signal, a cortisol level signal, a galvanic skin response, GSR, signal, an electroencephalography, EEG, signal, a magnetencephalography, MEG, signal, an electromyography, EMG, signal, and an eye activity signal.
3. The apparatus of claim 1 , wherein the processing circuit is configured to determine an auditory stimulus in the audio signal based on an entrainment of the at least one physiological parameter with a frequency and / or phase the auditory stimulus.
4. The apparatus of claim 1 , wherein the processing circuit is configured to determine a visual stimulus in the video signal based on an entrainment of the at least one physiological parameter with an appearance of the visual stimulus within a visual field of the user.
5. The apparatus of claim 1 , wherein the at least physiological sensor is configured to determine MEG or EEG data, and wherein the processing circuit is configured to correlate the audio and / or video signal with brainwave responses using the MEG or EEG data.
6. The apparatus of claim 5, wherein the processing circuit is configured to decompose the EEG signal into auditory and visual brainwave responses, and correlate the user’s auditory brainwave response with the audio signal, and correlate the user’s visual brainwave response with the video signal.
7. The apparatus of claim 1 , further comprising a classification circuit configured to classify auditory stimuli of the audio signal and / or visual stimuli of video signal, and wherein the processing circuit is configured to correlate the measured at least one physiological parameter with the one or more classified auditory and / or visual stimuli.
8. The apparatus of claim 1 , further comprising a camera to capture the scene and output the video signal; an eye tracking circuit configured to determine the user’s foveal vision; wherein the processing circuit is configured to map the user’s foveal vision to the video signal to determine which part of the scene is within the user’s foveal vision.
9. The apparatus of claim 8, further comprising a classification circuit configured to classify one or more objects / features in the scene based on the video signal, and wherein the processing circuit is configured to correlate the measured at least one physiological parameter with one or more classified objects / features within the user’s foveal vision.
10. The apparatus of claim 9, wherein the classification circuit is configured to translate the video signal into text describing the one or more objects in the scene.11 . The apparatus of claim 9, wherein the classification circuit comprises a trained machine learning network configured for video captioning.
12. The apparatus of claim 9, wherein the processing circuit is configured to detect the stress level caused by an object in the scene based on the classified objects, the user’s foveal vision, and the at least one physiological parameter.
13. The apparatus of claim 1 , further comprising an output interface configured to output information on the scene and / or auditory stimuli and / or visual stimuli causing the stress level above the threshold.
14. The apparatus of claim 1 , further comprisingan augmented reality, AR, processing circuit configured to modify the audio and / or video signal if the stress level is above the threshold.
15. The apparatus of claim 14, wherein the AR processing circuit is configured to selectively modify the auditory and / or visual stimuli causing the stress level above the threshold.
16. The apparatus of claim 15, wherein the AR processing circuit is configured to selectively cancel the auditory and / or visual stimuli causing the stress level above the threshold from the audio and / or video signal.
17. An apparatus for XR perception of a user, the apparatus comprising an input interface configured to receive an audio signal of a scene; receive a video signal of the scene; and a trained machine learning model configured to predict the user’s stress level based on the audio and / or video signal; and an AR processing circuit configured to modify the audio and / or video signal if the predicted stress level is above a threshold.
18. The apparatus of claim 17, wherein the AR processing circuit is configured to selectively modify or cancel auditory and / or visual stimuli causing the stress level above the threshold.
19. The apparatus of claim 16, further comprising a classification circuit configured to translate the audio and / or video signal into text describing the audio and / or video signal, wherein the trained machine learning model is configured to predict the user’s stress level based on the text describing the audio and / or video signal; wherein the AR processing circuit is configured to modify auditory and / or visual stimuli described in the audio and / or video signal and causing the stress level above the threshold.
20. A method for XR perception of a user, the method comprising receiving an audio signal of a scene; receiving a video signal of the scene;measuring at least one physiological parameter of the user while the user is experiencing the scene; correlating the measured at least one physiological parameter of the user with the audio and video signal, and based on the correlation, determining auditory and / or visual stimuli causing a stress level of the user above a threshold.
Citation Information
Patent Citations
Methods and systems for immersive reality in a medical environment
US10976806B1
Device, system, and method for reducing coronasomnia to enhance immunity and immune response
US20210338973A1
System for extended reality visual contributions
US20210383912A1
Stress detection
US20240164672A1