Method, computer program and device for adjusting services consumed by human, robot, vehicle, enteraction system and smart home system
By using event-based optical sensors to capture and analyze optical event data, the problem of efficiently identifying and adjusting services based on human emotional states in existing technologies has been solved, enabling real-time adjustments and improved user experience.
Patent Information
- Application Number
- CN202480018565.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-21
- Filing Date
- 2024-03-15
- Publication Date
- 2025-11-07
AI Technical Summary
Existing technologies struggle to efficiently identify and adjust services based on human emotional states. Conventional cameras have difficulty capturing facial micro-expressions and suffer from high data processing redundancy, leading to delays and high costs.
It uses event-based optical sensors to capture optical event data, analyzes micro-movements to detect predefined emotional patterns, and enables real-time or near-real-time service adjustments.
It enables real-time or near-real-time service adjustments based on emotional state, reducing data processing latency and costs, and improving user experience satisfaction.
Smart Images

Figure CN120917404A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Examples relate to methods, computer programs and devices for adapting services consumed by humans, robots, vehicles, entertainment systems and smart home systems, and more specifically, but not exclusively, to the idea of adapting services consumed by humans based on their emotional state. BACKGROUND
[0002] Conventionally, human emotion recognition is a challenging task as it is difficult to reveal one’s inner emotions. However, with facial mimicry and body language, human emotion analysis can be conducted. For example, in our daily communication, human emotions are mainly expressed with facial expressions. Facial expressions are complex as they are based on a multitude of factors and facilitators. However, known models map certain human emotions to certain facial expressions. For example, it is well known that laughing or smiling is an expression of happiness and when laughing or smiling, one’s mouth corners are upturned. SUMMARY
[0003] Examples are based on the finding that micro-movements in a human face can be captured using event-based optical sensors. Based on the optical event data comprising information about the micro-movements, the emotional state of the human can be detected based on predefined emotional patterns. Then, the service can be adapted or adjusted based on the emotional state of the human.
[0004] Examples provide a method for adapting a service consumed by a human. The method comprises capturing optical event data of the human using an event-based sensor and analyzing the optical event data to detect a predefined emotional pattern. The method further comprises determining information about an emotional state of the human based on the emotional pattern detected in the optical event data and adapting the service based on the information about the emotional state of the human. In contrast to conventional video data, event-based optical data can have a higher frame rate or resolution, so that shorter movements can be detected and considered for the determination of the emotional state.
[0005] For example, the service is a game, a movie, a smart home service or a metaverse. As will be shown in the following description, there are a multitude of various services that can benefit from an adaptation based on the emotional state of the human. Furthermore, the capturing can comprise capturing optical event data from a plurality of event-based sensors. Using event data from a plurality of sensors can facilitate the resolution of the data and a more reliable detection of the emotional pattern. In some examples, the capturing can comprise capturing the optical event data in a three-dimensional (3D) data volume comprising a time component. Thereby, a more detailed analysis of the predefined pattern can be enabled, e.g., using the 3D data over time,
[0006] Depending on the specific implementation in the example, the optical event data can comprise pixel-level polarity data on luminance, color, or contrast changes. The sampling of the optical quantity changes allows for a reduced data rate while still maintaining a high frame rate or sampling rate. Transmitting only the differential data, e.g. the polarity, luminance, and / or color changes, allows for a reduced amount of data to be transmitted. For example, the optical event data has a sampling rate of at least 500 frames per second. Thereby, even very short optical changes can be tracked. The analysis can comprise determining micro-movements in the optical event data to detect a predefined emotional pattern. For example, micro-movements refer to one or more elements of the group of a human's mouth, eyes, eyebrows, eyelids, lips, jaw, skin, cheek, and nose. The example can implement tracking or acquiring optical event data of micro-movements, thereby allowing for a detailed analysis of, for example, a human's facial expression. Thus, the analysis can comprise analyzing a human's facial expression to detect an emotional state. Since a facial expression can not be the only indicator of an emotional state, the analysis can also comprise analyzing a human's body posture or gesture to detect an emotional state.
[0007] As outlined above in general, in the example, any conceivable service can be adjusted based on the emotional state. The adjustment of the service can comprise one or more elements of the group of increasing or decreasing agility of a game character, increasing or decreasing audio volume, increasing or decreasing speed of verbal or audio content, changing style of music, adjusting room temperature, adjusting brightness and / or color of light, providing special effects in the service, adjusting a virtual avatar of the human, and / or introducing other humans to the human. The service can be provided in any environment, examples being a smart home, a car, virtual or augmented reality, entertainment, etc.
[0008] The example also provides a computer program having a program code for performing the methods as described herein, when the computer program is executed in a computer, a processor, or a programmable hardware component. Another example is a device for adjusting a service consumed by a human, the device comprising circuitry configured to perform any of the methods described herein. For example, the device further comprises a plurality of event-based optical sensors having different fields of view of the user.
[0009] Other examples are a robot comprising the device, a vehicle comprising the device, an entertainment system comprising the device, and / or a smart home system comprising the device. BRIEF DESCRIPTION OF DRAWINGS
[0010] Some examples of devices and / or methods will be described below by way of example only, and with reference to the accompanying drawings, in which:
[0011] Figure 1 A block diagram of a flowchart illustrating an example of a method for adjusting a service consumed by a human is shown;
[0012] Figure 2 A block diagram illustrating an example of a device for adjusting a service consumed by a human is shown;
[0013] Figure 3 An example of a discrete emotion model is shown;
[0014] Figure 4 An example of a dimensional emotion model is shown;
[0015] Figure 5 Examples of facial expressions are depicted;
[0016] Figure 6 A block diagram illustrating a method for detecting an emotional state of a human in an example is shown;
[0017] Figure 7 A block diagram illustrating another example of a method for adjusting a service for a human based on an emotional state of the human is shown;
[0018] Figure 8 Emotion state detection based on predefined patterns is shown;
[0019] Figure 9 A block diagram illustrating another example of a method for adjusting a service for a human (game / movie or metaverse) based on an emotional state of the human is shown;
[0020] Figure 10 A block diagram illustrating an example of Figure 9 with additional information is shown;
[0021] Figure 11 A block diagram illustrating another example of a method for adjusting a robotic service for a human based on an emotional state of the human is shown; and
[0022] Figure 12 A block diagram illustrating an example method of Figure 11 with additional information is shown. DETAILED DESCRIPTION
[0023] Some examples will now be described in greater detail below with reference to the accompanying drawings. However, other possible examples are not limited to features of these detailed examples. Other examples can include modifications and / or enhancements to features, as well as equivalents and alternatives of features. Additionally, terminology used herein to describe certain examples should not limit other possible examples.
[0024] Throughout the description of the drawings, the same or similar reference numbers refer to the same or similar elements and / or features, which can be identical, or modified versions of identical, while providing the same or similar functions. The thicknesses of lines, layers, and / or regions can also be exaggerated for clarity.
[0025] When two elements A and B are used in "or" combination, it is to be understood that all possible combinations are disclosed, i.e. only A, only B and A and B, unless otherwise specifically defined in the specific case. As an alternative to the use of "or" for the same combination, "at least one of A and B" or "A and / or B" can be used. The same applies to combinations of more than two elements.
[0026] If the singular form is used (such as "a", "an" and "the") and it is not explicitly or implicitly defined that the use of only a single element is mandatory, other examples can also use several elements to implement the same function. If a certain function is described below as being implemented using a plurality of elements, other examples can implement the same function using a single element or a single processing entity. It is also to be understood that when the terms "comprise", "comprising", "include", "including", "comprises" and / or "comprising" are used, it is described that the specified features, integers, steps, operations, processes, elements, components and / or groups thereof are present, but not excluding the presence or addition of one or more other features, integers, steps, operations, processes, elements, components and / or groups thereof.
[0027] Figure 1 A block diagram showing a flowchart of an example of a method 10 for adjusting a service consumed by a human is shown. The method 10 comprises capturing 12 optical event data of a human using an event-based sensor and analyzing 14 the optical event data to detect a predefined mood pattern. The method 10 further comprises determining 16 information about a mood state of the human based on the mood pattern detected in the optical event data and adjusting 18 the service based on the information about the mood state of the human.
[0028] Figure 2 A block diagram showing an example of an apparatus 20 for adjusting a service consumed by a human is shown. The apparatus 20 comprises circuitry 22 configured to perform the method 10 or any method described herein. Figure 2 An optional entity 200 comprising the apparatus 20 is also shown. For example, the entity 200 can be a robot (industrial or household), a vehicle (car, bus, truck, train, airplane, etc.), an entertainment system (household, on-board, in-car, etc.), or a smart home system (light control, temperature control, etc.).
[0029] The circuitry 22 can be implemented using one or more processing units, one or more processing devices, any means for processing such as a processor, a computer or a programmable hardware component being operable with accordingly adapted software. In other words, the described functionality of the circuitry 22 can be implemented in software that is subsequently executed using one or more programmable hardware components. Such hardware components can include a general purpose processor, a Digital Signal Processor (DSP), a Graphics Processing Unit (GPU), a micro-controller, etc. Other examples are the apparatus or entity 200, e.g. an automation system, e.g. for entertainment, gaming, exercise, vehicles, comprising an example of the device 20.
[0030] As Figure 2 As further shown in the middle, the device 20 can optionally (shown in dashed lines) comprise one or more interfaces 24 coupled to the respective circuitry 22. The one or more interfaces 24 can correspond to one or more inputs and / or outputs for receiving and / or transmitting information within a module, between modules or between modules of different entities, which can be transmitted in digital (bit) values according to a specified code. For example, the interface 24 can comprise interface circuitry configured to receive and / or transmit information. In examples, the interface 24 can correspond to any means for obtaining, receiving, transmitting, or providing analog or digital signals or information, e.g. any connector, contact, pin, register, input port, output port, conductor, channel, etc. that allows for providing or obtaining a signal or information. The interface 24 can be configured to communicate (transmit, receive or both) in a wireless or wired manner and it can be configured to communicate, i.e. transmit and / or receive signals, information, with other internal or external components. The one or more interfaces 24 can comprise other components to enable communication in a (mobile) communication system, such components can include transceiver (transmitter and / or receiver) components such as one or more low noise amplifiers (LNA), one or more power amplifiers (PA), one or more duplexers, one or more duplex filters, one or more filters or filter circuitry, one or more converters, one or more mixers, accordingly adapted radio frequency components, etc. The interface 24 can be used to provide service adjustment information to a device that provides a service for a user or human.
[0031] Human emotion recognition is a challenging task as it is difficult to reveal a person's inner emotions. However, with facial mimicry and body language, human emotion analysis can be performed. Research reveals that human emotions are primarily (55%) expressed using facial expressions in our daily communication, see A. Mehrabian, Communicating Without Words, Psychol. Today. (1968) 53-55. Ordinary cameras can not be sufficient to capture facial microexpressions, which can be achieved using event-based vision (EVS) sensors. EVS sensors can detect sudden changes in expressions, such as eyebrow movements, swallowing, pulsating blood vessels, or sudden eye movements, for example, when people are lying, they tend to look down, and when they are recalling, they tend to look up. Due to the speed and low contrast of regular cameras, changes and movements of human faces are difficult to detect for these cameras. Moreover, frame-based data can have high redundancy in the acquired data for emotion analysis. Image-level camera sensing can cause large data storage and processing delays. Regular cameras can be difficult to capture facial microexpressions that occur in a short amount of time and can result in high costs when using high-speed cameras.
[0032] Examples utilize one or more event sensors. Event cameras are imaging sensors that respond to local changes in, for example, luminance. Event cameras do not use shutters to capture images like regular (frame) cameras. Instead, each pixel inside an event camera operates independently and asynchronously, reporting changes in luminance when they occur, and otherwise remaining silent. Event cameras are considered the next generation technology for computer vision tasks due to their high speed, high dynamic range, and data acquisition efficiency. Event cameras are used to capture motion in a scene, and static objects do not activate the event sensors.
[0033] In other examples, the device 20 can also include one or more (e.g., multiple) event-based optical sensors, which can have different fields of view of the user. The one or more sensors can be used to capture 12 optical event data of the human. For example, the sensors can include pixel-level sensors that operate independently and asynchronously. The pixel sensors can be activated upon detecting a change, i.e., if an optical property captured by the pixel changes, for example, when a motion or event occurs. The optical event data can include pixel-level polarity data about changes in luminance, color, and / or contrast.
[0034] This can allow for lower latency, lower power consumption, and lower data processing requirements compared to frame-based sensors that capture full frame information at each sampling step. Event-based sensors can have a higher dynamic range than frame-based sensors, even though they use high-speed vision (high frame rate). For example, optical event data has a sampling rate of at least 500 frames per second. Using event-based sensors, events can be recorded that would otherwise require a regular camera to run at 10,000 images / second and higher speeds. Event-based sensors can allow for latencies of 40 ps to 200 ps to be achieved at high bandwidth, e.g., 60 Mep (mega events per second) and above. Furthermore, 3D capture can even be enabled by event-based sensors. Capture 12 can include capturing optical event data in a 3D data volume that includes a temporal component, which can form a data basis for very efficient and accurate motion detection.
[0035] Emotion recognition and affective analysis with event-based vision sensors (face / body) can enable real-time or near real-time adjustments, 1 ms to 10 ms algorithmic adjustments. For example, high-speed cameras can operate at 100 frames / second, while event-based sensors can operate at 1,000 frames / second (events / second) and greater.
[0036] In an example, a predefined emotion pattern can be mapped to a human emotional state using an emotion model. Figure 3 An example of a discrete emotion model is shown. Two discrete emotion models for affective computing are shown. Figure 3 (a) on the left shows six basic emotions (anger, disgust, fear, joy, sadness, and surprise) shown as emoji types. Figure 3The right side (b) shows a more advanced emotional wheel model that makes finer distinctions among different emotions. Discrete emotion models define emotions as limited categories, two widely used emotion models are Ekman’s six basic emotions (anger, disgust, fear, joy, sadness, surprise), see P. Ekman, Universals and cultural differences in facial expressions of emotion, Nebr. Symp. Motiv. 19 (1971) 207-283; and Plutchik’s emotional wheel model, Plutchik Robert, Emotion and Life: perspective from psychology biology and evolution, Am. Physiol. Assoc. (2003), both of which are shown in Figure 3 .
[0037] The development of Ekman’s basic emotion model is based on the assumption that human emotions are shared across ethnic groups and cultures, see Wang Y, Song W, Tao W, et al. A systematic review on affective computing: Emotion models, databases, and recent advances. Information Fusion, 2022. From emojis, differences in facial expressions for different emotions can be recognized.
[0038] Unlike Ekman’s basic emotion model, Plutchik’s emotional wheel model involves eight emotions (joy, fear, surprise, sadness, anticipation, anger, and disgust) and the relationship between one emotion and another.
[0039] Figure 4 Examples of dimensional emotion models are shown. Continuous multi-dimensional models are introduced to describe fine-grained affect. As shown in Figure 4 , the two most recognized models are the Pleasure-Arousal-Dominance (PAD) model (a) on the left side and the Valence-Arousal (V-A) model (b) on the right side. Figure 4 Figure 4 on the right (a)). The two-dimensional PAD model can represent most different emotions, and the V-A model is used to represent complex emotions.
[0040] Examples can consider facial expressions and emotions. As we move different muscles and make different expressions on our faces, humans communicate their emotions and feelings to others, as the following examples will show: https: / / www.scienceofpeople.com / microexpressions / .
[0041] Figure 5 Examples of facial expressions are depicted. Figure 5 Examples of facial expressions of a pair of women and men are depicted for surprise 52 (at the top), sadness 54 (in the middle), and happiness 56 (at the bottom).
[0042] For the facial expression of surprise 52, the following facial actions (emotional patterns) help express the feeling of surprise:
[0043] • the eyebrows are raised and curved,
[0044] • the skin below the eyebrows is stretched,
[0045] • horizontal wrinkles are revealed across the forehead,
[0046] • the eyelids are open, and the whites of the eyes are revealed above and below, and
[0047] • the lower jaw is open and the teeth are parted, but the mouth is not tense or stretched.
[0048] For the facial expression of sadness 54, the following facial actions (emotional patterns) help express the feeling of sadness:
[0049] • the inner corners of the eyebrows are pulled in and up,
[0050] • the skin below the eyebrows is triangular, with the inner corners pointing up,
[0051] • the corners of the mouth are pulled down,
[0052] • the lower jaw is raised, and
[0053] • the lower lip is puffed out.
[0054] For the facial expression of happiness 56, the following facial actions (emotional patterns) help express the feeling of happiness:
[0055] • the corners of the mouth are pulled back and up,
[0056] • the mouth can or can not be open, revealing the teeth,
[0057] • wrinkles extend from the outside of the nose to the outside of the lips,
[0058] • the cheeks are raised,
[0059] • The lower eyelid can show wrinkles or be taut, and
[0060] • The crow's feet near the outer corner of the eye.
[0061] As part of the above analysis 14 steps, the example can determine micro-motions in the optical event data to detect a predefined emotional pattern. Micro-motions can refer to one or more elements in the group of a human's mouth, eyes, eyebrows, eyelids, lips, jaw, skin, cheeks, and nose. As outlined above and shown by the facial expressions in FIG. 13, the analysis 14 can include analyzing a human's facial expressions to detect an emotional state. In other examples, the analysis 14 can additionally include analyzing a human's body posture or gestures to detect an emotional state. Figure 5
[0062] As mentioned above, facial expressions can be used as a reliable feature for analyzing a human's emotions, as humans tend to make similar facial expressions when they feel similar emotions. Thus, as changes on a human's face, a person's emotions will be analyzed according to typical facial expression patterns. Event cameras can be used to capture motions and changes on a face in an efficient manner, specifically, facial micro-expressions can be captured. In comparison to conventional RGB (red, green, blue) cameras, event cameras have a higher frame rate to capture micro-expressions that can occur as fast as 1 / 15 to 1 / 25 of a second.
[0063] Figure 6 A block diagram of a method for detecting an emotional state of a human in an example is shown. Figure 6 A diagram of a process for facial emotion detection with event cameras is shown. Figure 6 A human face 61 is shown being recorded 62 using an event camera. Subsequently, optical event data from the event sensor 62 is analyzed, extracting 63 events with facial landmarks. In facial data 64, a contrast of old events 64a versus new events 64b (mouth starts to smile) is depicted. In step 65, facial motions are identified (detect emotional pattern). In emotion analysis 66, it is determined that the corners of the mouth are pulled back and up, which is an emotional pattern indicating an emotional state of "happy" 67. With EVS, people can detect sudden changes in expression, eyebrow motion, pulsating blood vessels, or sudden eye movements (people look down when they are lying, look up when they are remembering).
[0064] In other examples, event cameras or sensors can also be used for human emotion recognition with body gestures. In addition to facial expressions, the posture of the body also indicates different emotions of people. With event cameras, human body motions can also be recorded to assist in emotion recognition. Unsettling motions like leg fidgeting or twitching and tense movements like cramps can be detected to indicate unease.
[0065] For potential application areas of the examples, first, emotion recognition can be used as an auxiliary monitor for entertainment systems or interactive games. Second, a driving assistance monitor can also benefit from driver emotion recognition to avoid fatigued driving and to recognize extreme emotions (anger, sadness, etc.) or signs of illness (sadness, pain, etc.) to avoid traffic accidents and improve driving safety. Furthermore, human emotion recognition can help human-robot interaction, as robots (e.g., service robots) can understand the owner’s emotions, and they can take actions accordingly, which can improve user satisfaction with the robots. In a general example, the adjustment 18 of the service includes one or more elements from the following group: increasing or decreasing the agility of a game character, increasing or decreasing the volume of audio, increasing or decreasing the speed of verbal or audio content, changing the style of music, adjusting the room temperature, adjusting the brightness and / or color of the light, providing special effects in the service, adjusting a virtual avatar of a human, or introducing other humans to the human.
[0066] If the whole face of a person moves, the optical flow of most events can be detected and subtracted from the scene. In this way, only the parts of the face that move in a different way (e.g., the eyebrows) can be segmented and analyzed. The examples can improve emotion recognition by capturing 12 human facial micro-expressions. The examples can provide an affordable and accurate analysis method for emotion recognition. The examples can be implemented with a simple device (event sensor) and can not require high computational costs. The examples can be integrated into many areas to improve user satisfaction.
[0067] Figure 7 A block diagram of another example of a method for adjusting a service for a human based on the emotional state of the human is shown. Figure 7 A user 71 is shown, whose optical event data is captured, the optical events are detected and analyzed 72. Based on the determined emotional state of the user, real-time parameter adjustments are performed 73 and a refreshed stream is provided to the user 71 in the user device 74. Figure 7 A control loop for parameter adjustments on a user device based on the emotional state of the user is shown. An event camera is used for the emotion monitoring loop to allow real-time (or almost real-time) decision making and adjustments. To provide a better user experience, changes or adjustments of the device settings should be imperceptible to the user. In the examples, adjustments can be faster compared to conventional imaging sensors that capture data frame by frame with an exposure time. With conventional data capture and transmission, a delay can arise to refresh the device according to the new settings. However, event sensors have a data stream, which can record data continuously, and the burden of transmitting data can be much lower compared to full image transmission. Furthermore, event cameras can capture a 3D data volume that contains temporal information that can help to better predict the facial emotions of the user. EVS can allow refreshing within a duration that is (almost) imperceptible.
[0068] If the user observes that the adjustment to the device (game, movie, music, robot...) is slowing down or pausing, they can be annoyed by the discontinuity of the stream and this can decrease the user's satisfaction.
[0069] The following shows parameter adjustment in an example.
[0070]
[0071]
[0072] The following shows an example of parameter adjustment in a weighted emotion matrix.
[0073]
[0074] In an example, the face region can be detected by detecting blinks. Additionally or alternatively, event-based face detection can be performed, e.g., using an enhanced kernel correlation filter for event-based face detection.
[0075] Figure 8 Emotion state detection based on predefined patterns is shown. Figure 8 A detected face is shown at the top. After detecting the emotion pattern in the face, confidence or ranking information for different emotion states of the user can be determined. These are shown in a bar chart, which shows that happy has the highest confidence / ranking. Figure 8 Another face is shown at the bottom of Figure 1 1. Relevant keypoints are shown along the eyebrows, eyes, nose, and mouth. For example, a distance measurement between consecutive keypoints that mark the corners of the mouth can indicate whether the corners of the mouth are moving upwards (smiling (happy)) or downwards (sad).
[0076] For example, with the positive event rate in the relevant regions, the ranking (confidence) of happy can be obtained by evaluating the percentage of positive events in the mouth and eye regions. Determining the distance of keypoints in the face can simplify face landmark morphing or face pose alignment detection.
[0077] Figure 9 A block diagram of another example of a method of adjusting a service (game / movie or metaverse) for a human based on the emotion state of the human is shown. The service can be a virtual reality movie or game, and the user 98 can wear a virtual reality (VR) device 99 (e.g., glasses) with multiple event-based sensors 99a, 99b, 99c, each with a certain field of view angle. There can be more than one sensor 99a, 99b, 99c per eye, and the sensors 99a, 99b, 99c can be arranged in a way that the field of view of the sensors 99a, 99b, 99c overlaps. Figure 9The three sensors shown are not all sensors. For example, one sensor can be positioned in the corner of the glasses and used to detect movements in the mouth area. Another sensor can be positioned on the frame near the ear and used to collect movements in the user's cheeks. Another sensor can be positioned towards the forehead, and yet another sensor can be positioned towards the eyes.
[0078] Figure 9 The activation of the service is shown at the top 91, followed by the VR product 99 being launched or added to the service / metaverse 92. This example method, consistent with the above description, detects 93 the boredom on the user's 98 face. In a game / movie scenario, based on the level of boredom, 94 the agility of the game character can be increased, or 94 the volume of the music can be increased. The user 98 may then feel invigorated and in a better mood 95. Steps 93, 94, and 95 run in a feedback loop for real-time updates or adjustments. In the case of a metaverse service, special effects can be provided on the virtual avatar, or new friends can be introduced 96. The user 98 may then show interest and be happier 97. Steps 93, 96, and 97 also run in a feedback loop for real-time updates or adjustments.
[0079] Figure 10 It shows Figure 9 A block diagram of an example. Figure 10 The document further instructs game developers to collect real-time and realistic feedback on recently launched features, and that games can be personalized or customized, such as adjusting background music or game scenes (e.g., on a beach or in the mountains). For movie / movie scenes, multiple EVSs can be used to develop systems for adjusting engagement levels (e.g., fear, laughter, immersion, boredom) to generate automatic movie ratings. Additionally, in some examples, virtual avatars in the metaverse can be optimized in real-time with similar speech rates based on individual user databases. Figure 9 The loop shown allows the virtual avatar to be continuously refreshed, providing smoother movements such as smiling or frowning based on the real-world user. The advantage of this example is real-time information gathering, rather than the traditional steps of data collection, processing, and signal transmission. The event flow is continuous, so all steps are parallel, and all data collection and parameter adjustments can be completed in real time.
[0080] Figure 11 A block diagram is shown as another example of a method for adjusting robotic services for humans based on their emotional states. Figure 11On the left side, the robot 110 is shown interacting with a user who is first shown in a sad emotional state 111a, but then feels uplifted and becomes a happier emotional state 111b. The corresponding flowchart is shown on the right side. It starts with an interaction 112 between the user and the robot. In step 113, facial localization with EVS and facial emotion analysis are performed. The active emotion and system loop 114 operates as outlined above, and the emotion is monitored or regulated 115. In case the user is in a good mood, a higher speed and frequency 116a can be employed. This can cause the more the user is happy, the more the robot interacts with the user 116b. The feedback loop includes steps 115, 116a and 116b for real-time updating. In case the user is in a bad mood, soothing music can be played or the robot’s conversation can be slowed down 117a. If the user feels uplifted (positive feedback) 117b, the feedback loop is closed, returning to step 115. If the user is still in a bad mood (negative feedback) 118a, a relaxing scent is released 118b.
[0081] Figure 12 A flowchart of an example method is shown with additional information Figure 11 For the EVS emotion analysis, since the sensor captures a 3D data volume containing a time component, the captured information can be used to better predict the occurring emotion and to estimate the duration of the emotion. As mentioned for the EVS emotion sensing, it can not only detect the emotion, but also analyze the shift of the emotion, which especially helps the robot to adjust its settings according to the user’s mood.
[0082] The following summarizes some examples:
[0083] (1) A method for adjusting a service consumed by a human, the method comprising capturing optical event data of the human using an event-based sensor; analyzing the optical event data to detect a predefined emotional pattern; determining information about an emotional state of the human based on the emotional pattern detected in the optical event data; and adjusting the service based on the information about the emotional state of the human.
[0084] (2) The method of (1), wherein the service is a game, a movie, a smart home service, or a metaverse.
[0085] (3) The method of (1) or (2), wherein the capturing comprises capturing the optical event data from a plurality of event-based sensors.
[0086] (4) The method of (1) to (3), wherein the capturing comprises capturing the optical event data in a three-dimensional data volume comprising a time component.
[0087] (5) The method of (1) to (4), wherein the optical event data comprises pixel-level polarity data on luminance or contrast changes.
[0088] (6) The method of (1) to (5), wherein the optical event data has a sampling rate of at least 500 frames per second.
[0089] (7) The method of (1) to (6), wherein the analyzing comprises determining micro-movements in the optical event data to detect a predefined emotional pattern.
[0090] (8) The method of (8), wherein the micro-movements refer to one or more elements of the group of a human’s mouth, eyes, eyebrows, eyelids, lips, jaw, skin, cheek, and nose.
[0091] (9) The method of (1) to (8), wherein the analyzing comprises analyzing facial expressions of the human to detect an emotional state.
[0092] (10) The method of (1) to (9), wherein the analyzing comprises analyzing body posture or gestures of the human to detect an emotional state.
[0093] (11) The method of (1) to (10), wherein the adjusting of the service comprises one or more elements of the group of increasing or decreasing agility of a game character, increasing or decreasing volume of audio, increasing or decreasing speed of verbal or audio content, changing style of music, adjusting room temperature, adjusting brightness and / or color of light, providing special effects in the service, adjusting virtual appearance of the human, or introducing other humans to the human.
[0094] (12) A computer program having a program code for performing the method according to any other example herein when the computer program is executed in a computer, processor, or programmable hardware component.
[0095] (13) An apparatus for adjusting a service consumed by a human, comprising circuitry configured to perform one of the methods of (1) to (11).
[0096] (14) The apparatus of (13), further comprising a plurality of event-based optical sensors having different fields of view of the user.
[0097] (15) The apparatus of (13) or (14), wherein the service is a game, a movie, a smart home service, or a metaverse.
[0098] (16) The apparatus of (13) to (15), wherein the circuitry is configured to capture optical event data from the plurality of event-based sensors.
[0099] (17) The device of (13) to (16), wherein the circuitry is configured to capture the optical event data in a three-dimensional data volume comprising a time component.
[0100] (18) The device of (13) to (17), wherein the optical event data comprises pixel-level polarity data regarding luminance or contrast changes.
[0101] (19) The device of (13) to (18), wherein the optical event data has a sampling rate of at least 500 frames per second.
[0102] (20) The device of (13) to (19), wherein the circuitry is configured to determine micro-movements in the optical event data to detect a predefined emotional pattern.
[0103] (21) The device of (20), wherein the micro-movements refer to one or more elements of the group of a human's mouth, eyes, eyebrows, eyelids, lips, jaw, skin, cheek, and nose.
[0104] (22) The device of (13) to (21), wherein the circuitry is configured to analyze facial expressions of the human to detect an emotional state.
[0105] (23) The device of (13) to (22), wherein the circuitry is configured to analyze body posture or gestures of the human to detect an emotional state.
[0106] (24) The device of (13) to (23), wherein the circuitry is configured to adjust a service by one or more elements of the group of increasing or decreasing agility of a game character, increasing or decreasing volume of audio, increasing or decreasing speed of verbal or audio content, changing style of music, adjusting room temperature, adjusting brightness and / or color of light, providing special effects in the service, adjusting a virtual avatar of the human, or introducing the human to another person.
[0107] (25) A robot comprising the device of (13) to (24).
[0108] (26) A vehicle comprising the device of (13) to (24).
[0109] (27) An entertainment system comprising the device of (13) to (24).
[0110] (28) A smart home system comprising the device of (13) to (24).
[0111] Aspects and features described in relation to a particular example of the foregoing examples can also be combined with one or more other examples, to replace a same or similar feature in the other examples, or to additionally introduce the feature into the other examples.
[0112] Examples can also be or relate to a (computer) program including a program code for performing one or more of the above-described methods, when the program is executed in a computer, processor or other programmable hardware component. Recipients of a computer, processor or other programmable hardware component can therefore cause operations to be performed in the computer, processor or other programmable hardware component, to cause the steps, operations or processes described above to be performed. Accordingly, aspects of the above-described different methods can also be performed by a programmed computer, processor or other programmable hardware component. Examples can also encompass a program storage medium (or media) readable by a computer, processor, or other programmable hardware component and encoding a computer program of instructions and / or data structures for performing one or more of the above-described methods. The program storage medium (or media) can include, alone or in combination with one another, a digital data storage medium, a magnetic storage medium, e.g., magnetic disks, magnetic tapes, hard disk drives, or optical storage medium. Other examples can include a computer, processor, control unit, (field) programmable logic array ((F)PLA), (field) programmable gate array ((F)PGA), graphics processor unit (GPU), application-specific integrated circuit (ASIC), integrated circuit (IC), or system on chip (SoC) system programmed to perform the steps of the above-described methods.
[0113] It should also be understood that the disclosure of a number of steps, processes, operations, or functions disclosed in the description or claims that are not specifically dependent on each other should not be construed as implying that these operations must take place in the described order, unless explicitly stated or required by the technical circumstances. The foregoing description therefore does not limit the performance of the plurality of steps or functions to a specific order. Furthermore, in other examples, a single step, function, process or operation can include and / or be broken down into a plurality of sub-steps, sub-functions, sub-processes or sub-operations.
[0114] If some aspects have been described for a device or system, these aspects should also be understood as a description of a corresponding method. For example, a block, device or functional aspect of a device or system can correspond to a feature of a corresponding method, such as a method step. Aspects described in relation to a method should therefore also be understood as a description of corresponding blocks, corresponding elements, properties or functional features of a corresponding device or a corresponding system.
[0115] The following claims are hereby incorporated into the detailed description, wherein each claim can stand on its own as a separate example. Also, note that while the following claims refer to a specific combination of claims, other examples can include a combination of those claims with the subject matter of any other claims not specifically cited in such claim combination. Such combinations are hereby expressly proposed, unless circumstances make such a proposition impossible. In addition, features of one claim can also be included in any other independent claim, even though the claim is not directly dependent on the other independent claim.
Claims
1. A method for adjusting a service consumed by a human, the method comprising: capturing optical event data of the human using event-based sensors; analyzing the optical event data to detect a predefined emotional pattern; determining information about an emotional state of the human based on the emotional pattern detected in the optical event data; adjusting the service based on the information about the emotional state of the human.
2. The method of claim 1, wherein, the service is a game, a movie, a smart home service, or a metaverse.
3. The method of claim 1, wherein, the capturing comprises capturing optical event data from a plurality of event-based sensors.
4. The method of claim 1, wherein, the capturing comprises capturing optical event data in a three-dimensional data volume comprising a time component.
5. The method of claim 1, wherein, the optical event data comprises pixel-level polarity data about luminance or contrast changes.
6. The method of claim 1, wherein, the optical event data has a sampling rate of at least 500 frames per second.
7. The method of claim 1, wherein, the analyzing comprises determining micro-movements in the optical event data to detect the predefined emotional pattern.
8. The method of claim 7, wherein, the micro-movements refer to one or more elements of a group of mouth, eyes, eyebrows, eyelids, lips, jaw, skin, cheeks, and nose of the human.
9. The method of claim 1, wherein, the analyzing comprises analyzing facial expressions of the human to detect the emotional state.
10. The method of claim 1, wherein, the analyzing comprises analyzing body posture or gestures of the human to detect the emotional state.
11. The method of claim 1, wherein, the adjusting of the service comprises one or more elements of a group of: increasing or decreasing agility of a game character, increasing or decreasing volume of audio, increasing or decreasing speed of textual or audio content, changing style of music, adjusting room temperature, adjusting brightness and / or color of light, providing special effects in the service, adjusting a virtual avatar of the human, and introducing other humans to the human.
12. A computer program having a program code for performing the method according to claim 1 when the computer program is executed in a computer, a processor, or a programmable hardware component.
13. An apparatus for adjusting a service consumed by a human, the apparatus comprising circuitry configured to perform the method according to claim 1.
14. The apparatus according to claim 13, further comprising a plurality of event-based optical sensors having different fields of view of a user.
15. A robot comprising the apparatus according to claim 12.
16. A vehicle comprising the apparatus according to claim 12.
17. An entertainment system comprising the apparatus according to claim 12.
18. A smart home system comprising the apparatus according to claim 12.