Methods for interaction between a user and an AI model and information technology system
The method uses pupil size changes to provide emotional feedback to AI models, enabling automatic content adjustment and safer user interaction without cognitive effort, addressing the challenge of distracting user feedback in AI interaction.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- MERCEDES BENZ GROUP AG
- Filing Date
- 2024-08-17
- Publication Date
- 2026-06-25
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
The invention relates to a method for interaction between a user and an AI model and to an information technology system of the type defined in more detail in the preamble of claim 7. With the help of artificial intelligence, particularly in the form of so-called generative AI models, the automatic and artificial generation of content is possible. Such a generative AI model is trained through a process called pre-training, enabling it to generate artificial content such as images, videos, music, speech, text, and the like, based on an input prompt. During pre-training, the generative AI model is trained with extensive datasets, using a wide variety of texts, images, videos, and similar materials as source material. More in-depth information on the training and application of generative AI can be found, for example, in: Devlin et al. (2018). “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”. arXiv:1810.04805v2 [cs.CL]. Vaswani et al. (2017), “Attention Is All You Need”, arXiv:1706.03762 [cs.CL]. Ho et al. (2022), “Imagen Video: High-Definition Video Generation with Diffusion Models”, arXiv:2210.02303 [cs.CV]. Generative AI models are capable of generating many different results from the same input data by incorporating random variables. One such random variable is the so-called "seed." Adjusting the seed alters the artificially generated content. Furthermore, a user can customize the prompt for the AI model, for example, by specifying particular styles, such as "in the style of Pointillism," "in the style of Goya," "in the style of a comic book," or similar. Furthermore, DE 10 2022 121 930 A1 discloses a method and a device for playing music in a motor vehicle, DE 10 2006 015 332 A1 relates to a guest service system for vehicle users, and DE 10 2022 212 854 A1 describes reinforcement learning for automatically adjusting settings in vehicles. Machine-based evaluation of the quality of artificially generated content by the AI model is difficult, which is why user feedback is typically required. The user can provide input to tell the AI model whether they like or dislike the artificially generated content. However, this requires the user to perform specific actions. Providing such feedback can be difficult in situations where the user is distracted, for example, while driving a vehicle, or it can distract the user from their primary task, such as operating the vehicle. This creates the need to specify means that allow a user to give appropriate feedback to an AI model without having to redistribute their attention between different tasks. One approach to reducing the cognitive effort required for users to perform an operation is to eliminate the need for physical controls or virtual controls displayed on a touchscreen. For example, devices and systems can be operated via gesture control or voice control. However, the user still has to perform a corresponding gesture or issue a voice command, which also involves some cognitive effort, albeit reduced. It is also known to control devices via the eyes. For example, the Canon EOS R3 camera has an autofocus system that can be controlled via gaze tracking. The camera is able to determine which area of the viewfinder the photographer is looking at and focus accordingly on that region. It is also known to use gaze tracking as an input device for a computer, allowing paralyzed individuals to move a mouse pointer. For example, letters can be arranged in a geometric pattern, which a paralyzed person would look at sequentially to "type" words. Furthermore, DE 10 2016 104 362 A1 and US 7,744,216 B1 each describe a display device whose brightness can be controlled depending on the pupil size detected by a user. These devices utilize the fact that a viewer's pupil dilates in darkness to allow more light into the eye socket and thus onto the retina, and that the pupil increasingly constricts in high light intensity to prevent retinal burns from excessive light exposure. For this purpose, the eyes of a user viewing the display device are detected using a visual detection device, and the degree of pupil dilation is evaluated. Depending on the degree of pupil dilation, the brightness of the display device is then increased, kept constant, or decreased.The influence of ambient brightness can be taken into account, so that, for example, different limit values for the pupil diameter can be specified for different ambient brightness levels, from which a corresponding reduction or increase in the intensity of the display brightness is carried out. Recent studies also show that pupil size can adjust not only depending on lighting conditions but also on the emotions experienced by a user. For example, the pupil can dilate when a user hears a baby crying or a dog growling. This reaction, in the form of a change in pupil size, occurs immediately after experiencing the corresponding emotions, i.e., after a minimal time interval. Further information can be found in: Andreas Widmann, Erich Schröger, Nicole Wetzel, “Emotion lies in the eye of the listener: Emotional arousal to novel sounds is reflected in the sympathetic contribution to the pupil dilation response and the P3”, Biological Psychology, Volume 133, 2018, Pages 10-17, ISSN 0301-0511, https: / / doi.org / 10.1016 / j.biopsycho.2018.01.010. The present invention is based on the objective of providing a method for interaction between a user and an AI model, which allows the user to give the AI model feedback on content artificially generated by the AI model without the user having to be cognitively active. According to the invention, this problem is solved by a method for interaction between a user and an AI model with the features of claim 1. Advantageous embodiments and further developments, as well as an information technology system for carrying out the method, are described in the dependent claims. The invention describes a method for interaction between a user and an AI model, wherein the AI model artificially generates acoustic and / or visual content and causes the output of the acoustic content via acoustic output devices and / or the output of the visual content via visual output devices, wherein the user consumes the artificially generated content, wherein the pupil size of at least one of the user's eyes is detected during and / or after the output of the artificially generated content, and wherein the acoustic and / or visual content is varied by the AI model depending on the pupil size. The inventive method for user interaction with the AI model utilizes the fact that the user's pupils change in size when experiencing emotions in order to provide the AI model with feedback about the artificially generated content.The applicant recognized that the emotion-induced dilation or reduction of pupil size can be used as a parameter to describe the quality of artificially generated content as perceived by the user. The pupil size changes automatically when experiencing emotions, so the user doesn't have to expend any cognitive effort to provide the AI model with the corresponding feedback. Therefore, the user can remain fully focused on the main task, so there is no risk of the execution quality of the main task suffering, at least not at the moment the user provides the feedback. A wide variety of content can be artificially generated by the AI model and delivered to the user in various ways. For example, the AI model can artificially generate text, such as a weather report, a dictionary entry, a news article, a story, a poem, or similar content. The text can be displayed on a screen and / or read aloud through speakers using computer-generated speech. The AI model can also artificially generate sounds, such as nature sounds like birdsong, rippling water, the sound of the sea, or music, such as the sound of a musical instrument or even complete songs. Furthermore, it can generate image or video content, such as a work of art in a specific style, a landscape photograph, or an animation of a traffic scene.Corresponding visual content can also be output via a display device. The hardware components required for implementing the method according to the invention will be discussed in more detail later. At this point, it should merely be mentioned that the AI model can be executed locally on a computing unit, for example, a PC, or on a server or in a server environment, in particular a high-performance computing cluster. This can also be a distributed computing environment. This also allows the use of complex generative AI models, such as large language models, also known as "Large Language Models" (LLMs). A further advantageous development of the method involves the AI model expanding the acoustic and / or visual content if a change in pupil size during the output of the artificially generated content and / or after a defined period of time following the output of the artificially generated content is smaller than a defined initial size change threshold. This step is based on the idea that the artificially generated content consumed by the user is disliked. If the artificially generated content is disliked by the user, or at least perceived as boring, then no emotions are triggered in the user. Consequently, the change in pupil size either does not occur or is only marginal. This serves as feedback to the AI model that the artificially generated content should be adjusted to better appeal to the user.In the simplest case, "enhanced" in this context means that the artificially generated content is modified by the AI model. This generates variations, increasing the likelihood of producing artificial content that the user might find appealing. More importantly, however, additional stimuli and content are integrated into the artificially generated content. This means that multiple layers of information are gradually incorporated into the artificially generated content, and their intensity is particularly increased. For example, quiet nature sounds could initially be played. If no or only a minimal reaction is detected from the user, the volume of the nature sounds can be gradually increased, and even additional artificially generated sounds, such as birdsong, babbling water, or traffic noise, can be added to the audio content simultaneously or subsequently. It is also conceivable to stop playing nature sounds and switch to music. For instance, classical music could be played initially, which, if there is no user reaction, could then switch to, say, heavy metal or drum and bass.Additional output modalities can be added, such as activating a display device and outputting visual content that matches the acoustic content, such as a picture of a forest, an animation of spinning and smoking tires, or a video of a rock concert. The detected change in pupil size can include both dilation and constriction. For dilation and constriction, an individual initial size change threshold can be defined, such as 3% of the current pupil diameter for both dilation and constriction, or 3% for dilation and 4% for constriction, or similar. It is particularly advantageous to use suitable sensors, such as a brightness sensor or a camera, to detect the amount of light falling on the user's eye area and to take its influence on changes in pupil size into account as a correction factor. This prevents changes in pupil width or pupil size caused by differences in brightness from being mistakenly interpreted as a user reaction. Emotion-related changes in pupil size can occur during or after the output of the artificially generated content. Due to the short reaction time described earlier, typically on the order of a few milliseconds, the corresponding time period is also only a brief moment, for example, less than one second, preferably a few hundred milliseconds. Since the output of the artificially generated content and the evaluation of the pupil size changes are temporally correlated, an initial safety mechanism already exists to prevent a brightness-related change in pupil size from being misinterpreted as a user reaction should such a change be detected without any temporal correlation to the output of the artificially generated content.Furthermore, user feedback is fed into the AI model much faster than with traditional manual user input. This allows for the even faster artificial generation of relevant content. According to a further advantageous embodiment of the method according to the invention, the AI model reduces the acoustic and / or visual content, or at least keeps it constant in scope, if a change in pupil size during the output of the artificially generated content and / or after a defined period of time following the output of the artificially generated content exceeds a defined second size change threshold. A significant change in pupil size indicates that the user is experiencing emotions. This can then be transmitted to the AI model as feedback, informing it that the user enjoys the artificially generated content consumed. In this case, the content artificially generated by the AI model can continue to be output unchanged, or its scope can be reduced. This means that fewer stimuli or pieces of information are incorporated into the artificially generated content. For example, the volume of the audio content output via the acoustic output devices can be reduced, slow music or calming sounds can be played instead of fast-paced music, or the audio output can be stopped and only an image displayed on a screen. Video output could also be stopped and only an image or audio content displayed instead. The defined second size change threshold can be the same as the first size change threshold, for example, 3% of the current pupil size, or it can be larger, for example, 5% of the current pupil size. Therefore, if changes in pupil size are detected that are smaller than the defined first size change threshold, the artificially generated content is preferentially expanded. Advantageously, the artificially generated content can be reduced if a change in pupil size is detected that is larger than the second size change threshold. If changes in pupil size are detected that lie between the first and second size change thresholds, the artificially generated content can preferentially continue to be displayed unchanged. A further advantageous embodiment of the method according to the invention provides that the AI model is formed by a generative model, in particular in the form of a generative pre-trained transformer and / or a diffusion model. Various methods and algorithms for providing artificial intelligence are known. Any model suitable for generating artificial content can be used. A generative model is particularly preferred for the AI model, since generative models are especially suitable for generating artificial content due to their power. In this context, generative pre-trained transformers, also known as "Generative Pre-trained Transformers" (GPTs), have proven effective, particularly in connection with the artificial generation of texts. Large language models are such generative pre-trained transformers. Diffusion models are used particularly for the artificial generation of visual content such as images or videos. The AI model can also be a combination of different algorithms. For example, various artificial neural networks can be interconnected in a cascade. A large language model for formulating input prompts can be used for a diffusion model to generate image or video content. According to the inventive method, at least one AI model parameter is modified to vary the artificially generated content, in particular in the form of a seed and / or a stable diffusion setting, and / or an input prompt is modified, in particular a style setting. Changing the seed and / or the stable diffusion setting allows for influencing random parameters that affect the artificially generated content. The seed and the stable diffusion setting can be changed randomly or according to predefined patterns. It is also possible to modify the input prompt itself, whereby specific style settings can be defined. Various input prompts can be predefined and stored in a database. Different input prompts can then be read from this database, depending on the situation, and passed to the AI model.Especially when multiple AI models are interconnected in a cascade, input prompts can also be generated by an AI located at the beginning of the cascade, particularly in the form of a large language model. This provides even more options for customizing the artificially generated content. According to the inventive method, the extent to which the at least one AI model parameter is changed is designed to depend on the percentage change in pupil size. In particular, if no or only a marginal change in pupil size is detected, the artificially generated content must be adjusted. In this case, the AI model parameter(s) and / or the input prompt can be modified more significantly, thus increasing the variation in the artificially generated content. Particular emphasis is placed on tracking which AI model parameter is changed across multiple iterations, or what adjustments are made to the input prompt. This allows for the identification of changes to AI model parameters and input prompts that generate artificial content particularly reliably evokes emotions in the user. For example, if an AI model parameter is changed without detecting a change in pupil size, the corresponding AI model parameter can be reset or changed in the opposite direction. Conversely, if a change to an AI model parameter has caused a change in pupil size, this AI model parameter, or this aspect of the input prompt, can be left unchanged in the next iteration or further modified in the same direction as in the previous iteration.An information technology system comprising acoustic and / or visual output means, visual detection means for capturing a user's pupil size, and a computing unit configured for controlling the output and detection means and for providing or interacting with an AI model, is further developed according to the invention in that the output means, the detection means, and the computing unit are configured to execute a method described above. The information technology system can include one or more loudspeakers as acoustic output means. The acoustic output means can be directly connected to the computing unit, or additional components, such as an amplifier, a digital-to-analog converter, or the like, can be provided. The information technology system can, for example, include a stereo sound system, a surround sound system, or the like.With the help of such a sound system, it is possible to place virtual sound sources in the room, so that the sound appears to originate from a specific direction for the user. For example, the Dolby Atmos format can be used for this purpose. As a visual output device, the information technology system can include one or more display devices, such as an LCD screen, an OLED screen, and the like. Such a display device can also be touch-sensitive and thus serve to receive manual input. The information technology system could also include one or more lighting devices, such as light strips, as visual output devices. These light strips can be controlled individually to, for example, display animated light patterns. For instance, the color, color temperature, and / or brightness of the individual light sources in such a light strip can be specifically changed. In this way, for example, moving patterns can be created. By switching the lights on and off in a timed sequence, flashing patterns can also be generated. Such light strips can be used to provide ambient lighting in a vehicle.Suitable camera systems can be used to visually capture the user's eyes. These systems can include special lenses so that, by selecting an appropriate focal length, the user's eye area is imaged sufficiently large on a corresponding image sensor. This allows for the reliable detection of even the slightest changes in pupil size. Post-processing of the recorded camera images is also possible. For example, contrast can be adjusted, noise removed, and so on. The visual capture devices can also be capable of detecting light in the non-visible spectrum, such as infrared light. The user can also be illuminated with an active light source, particularly one designed to emit infrared light. This enables reliable eye capture even under adverse lighting conditions.Preferably, a predetermined arrangement of the user relative to the visual detection devices is provided. The visual detection devices are thus positioned in a predetermined installation location opposite the user and specifically aimed at a head resting area. A "head resting area" describes a region of the room where the user's head will typically be located. The computing unit can be, for example, a desktop computer, a laptop, an embedded system (such as a system-on-a-chip, or SoC), a server or server cluster, a tablet computer, a smartphone, or similar device. If the computing unit has relatively powerful components, particularly multi-core processors and multi-core graphics processing units (GPUs), the AI model can run directly on the computing unit. However, if the hardware is relatively weak, the AI model is preferably run in a computing cluster or cloud environment, and interaction between the computing unit and the cloud environment is handled via a communication channel. For example, the computing unit can communicate with the corresponding computing cluster via the internet.This makes it possible to provide the performance of extensive and powerful AI models even on computationally weak hardware. To control the output and input devices and to interact with the AI model, the computing unit includes corresponding machine-interpretable instructions, particularly in the form of a computer program. This computer program is stored in a computer-readable storage medium within the computing unit and is executable by a processor within the computing unit. An advantageous further development of the information technology system according to the invention provides that the information technology system is implemented in a vehicle. The computing unit can then preferably be the control unit of a vehicle subsystem, a central on-board computer, a telematics unit, or the like. Preferably, the AI model is executed in a cloud environment with which the computing unit exchanges data via a mobile network using the internet and an in-vehicle telecommunications unit. Vehicle-integrated loudspeakers can then be used as acoustic output devices. Visual output devices can include, for example, the display of the instrument cluster, the display of the head unit, a dedicated passenger display, the projection surface of a head-up display (HUD), and the like.A preferred visual detection device is an interior camera designed to detect the driver, for example, integrated into the vehicle's dashboard. Such interior cameras can be embedded in the instrument cluster and allow for driver assistance functions, such as gaze direction detection or drowsiness detection. The components required to provide the method according to the invention are therefore already installed in typical production vehicles today. This allows for easy integration into vehicles. According to an advantageous embodiment of the information technology system according to the invention, the computing unit is further configured to receive a vehicle parameter and supply it to the AI model as input data for shaping the artificially generated content. For this purpose, the computing unit can, for example, access a data bus of the vehicle, such as a CAN bus, an Ethernet data line, or the like. This allows it to read sensor data generated by the vehicle using sensors and to receive information from control units. Thus, the computing unit can process vehicle parameters such as the vehicle's speed, longitudinal and / or lateral acceleration acting on the vehicle, steering angle, outside temperature, oil temperature, engine speed, air conditioning setting, an active radio station, an active driving mode, and the like.For example, it's possible to play an artificially generated song in the vehicle, which the user—the driver or other passengers—perceives as appropriate to the driving situation. This means that a suitable song is perceived, for example, as matching the driving style, the route, or the driver's mood. A further advantageous embodiment of the information technology system according to the invention provides that the computing unit is configured to receive an artificially generated control command from the AI model, taking into account the vehicle parameters, and to control a vehicle subsystem with this command. The AI model can thus also generate control commands for vehicle subsystems as artificially generated content. This allows the AI model, for example, to operate the vehicle's power windows, adjust the position of an electrically adjustable seat, change the vehicle's air conditioning settings, and so on. With the aid of appropriate actuators, it is also possible to output artificially generated haptic content to the user.For example, an actuator of a massage seat can be specifically activated to transmit vibrations to the user, the steering wheel of the vehicle can vibrate or perform rotational oscillations, and so on. The AI model is particularly advantageous because it is capable of learning. This allows the AI model to analyze the relationship between variations made to the artificially generated content and observed user feedback, and to deduce which changes to the artificially generated content are likely to trigger an emotion and which are not. This enables the AI model to adapt the artificially generated content even more precisely. To learn how variations in artificially generated content affect the emotions evoked in the user, manual feedback can be requested or voluntarily provided by the user, either as a supplement or alternative. For example, the user can input their manual feedback into a prompt for the AI model. This allows the AI model to generate new content that appeals to the user in a more targeted way. Ideally, this manual feedback should be provided in a calm situation, such as when parking the vehicle. During this process, the last n pieces of content generated by the AI model can be displayed again and evaluated by the user. Furthermore, an automatic termination criterion can be implemented. For example, if no change in pupil size is detected after repeatedly modifying the artificially generated content, the output of artificially generated content in the vehicle can be stopped or at least paused. This prevents cognitive overload for the driver. Further advantageous embodiments of the inventive method for interaction between a user and an AI model, as well as of the inventive information technology system, also become apparent from the exemplary embodiments, which are described in more detail below with reference to the figures. Figure 1 shows a schematic flowchart of an inventive method for user interaction with an AI model; and Figure 2 shows a schematic representation of an application of the method in a vehicle. Using an inventive method for interaction between a user 1, shown in Fig. 2, and an AI model 2, the AI model 2 can receive feedback on whether the artificial content generated by the AI model 2 elicits a reaction from the user 1, without requiring the user 1 to actively provide feedback. This allows the user 1 to devote their full attention to a primary task and thus continue to perform it with particular reliability and safety. The method is based on the finding that changes in pupil size occur when experiencing emotions. This change in pupil size is detected sensorially and fed back to the AI model 2 as feedback. As shown in Fig. 1, in step 101 the AI model 2 first generates corresponding acoustic content 3 and / or visual content 4. The acoustic content 3 can be, for example, noises, tones, harmonies, music, and the like. The visual content 4 can be, in particular, images, videos, animations, and the like. The artificially generated content 3, 4 is then output to the user 1 in step 102 via acoustic output devices 5, for example, loudspeakers, and / or visual output devices 6, such as a display device. During the output of the artificially generated content 3, 4 and / or within a defined time window after the output, the eyes 7 of the user 1 are detected in step 103 using suitable visual detection devices 8, for example, an indoor camera. As shown in Fig. 1, the pupil size can decrease, remain constant, or increase. The change in pupil size is tracked over time, allowing corresponding changes in size to be detected.If this change in pupil size, occurring in relation to the output of the artificially generated content 3,4, is less than a defined first size change threshold, AI model 2 preferentially expands the artificially generated content. Conversely, if the change in pupil size is greater than a defined second size change threshold, AI model 2 keeps the size of the artificially generated content 3,4 constant or reduces it. The influence of brightness-related changes in pupil size can be taken into account. In step 104, the corresponding feedback is fed to the AI model 2, so that the AI model 2 can adjust its behavior to generate the artificially generated content 3,4. Fig. 2 shows an implementation of an information technology system usable for carrying out the method according to the invention in a vehicle 10. In general, however, the information technology system can also be implemented independently of the vehicle 10, for example implemented in a PC, wherein loudspeakers and a monitor are connected to the PC as output means and a camera as a recording means. The vehicle 10, on the other hand, comprises a computing unit 9, which has at least read access to a computer-readable storage medium and contains machine-interpretable instructions which, when executed by a processor of the computing unit 9, cause it to provide the method according to the invention. The vehicle 10 includes corresponding loudspeakers as acoustic output means 5. One or more display devices are provided as visual output means 6. The vehicle 10 may also have so-called ambient lighting, which can likewise be considered a visual output means 6. In addition, the vehicle 10 has an interior camera integrated into the instrument cluster, which serves as a visual detection means 8 for detecting the eye area of the user 1. The output means 5, 6 and the visual detection means 8 are controlled by the computing unit 9. The AI model 2 for generating the artificially generated content 3,4 can be implemented or executed on the computing unit 9. However, Fig. 2 shows an embodiment in which the AI model 2 is provided on a central computing facility 11. The central computing facility 11 is, in particular, a cloud environment, or a server or server cluster. The server cluster has powerful hardware and can therefore also be referred to as a high-performance computing cluster. The computing unit 9 manages the process for generating the artificial content 3,4. The computing unit 9 formulates corresponding input requests for the AI model 2 and transmits them wirelessly, in particular using a mobile internet connection, to the central computing facility 11 and thus to the AI model 2 via a telecommunications unit 12.The AI model 2 now generates the artificial content 3,4 on the central computing unit 11, which then transmits the artificially generated content 3,4 back to the computing unit 9 via the wireless connection. The computing unit 9 then causes the artificially generated content 3,4 to be output in the vehicle 10, so that the user 1 can consume it. During the output of the artificially generated content 3,4, or shortly thereafter, the eyes 7 of the user 1 are monitored. It may be sufficient to track only the change in pupil size of a single eye 7 of the user 1. Depending on the magnitude of the change in pupil size, the computing unit 9 interacts with the AI model 2 in a different way.Preferably, the computing unit 9 varies at least one AI model parameter, in particular in the form of a seed and / or a stable diffusion setting and / or transmits a customized input prompt, in particular taking into account a style specification. For example, user 1 might be driving vehicle 10 along a coastal road. Based on an analysis of the driving situation, processing unit 9 uses AI model 2 to generate an artificially created song that is intended to match the situation. The song is played in vehicle 10, but user 1 perceives it as inappropriate. This results in little to no change in pupil size. Processing unit 9 transmits this information to AI model 2, which then generates a different song, which is now played in vehicle 10. The adapted song, in terms of its harmony and / or the semantics of the vocals, is now more in keeping with the atmosphere of driving along the coastal road at sunset. Corresponding emotions are triggered in user 1, causing their pupils to dilate.This information is also transmitted back to the AI model 2 by the computing unit 9, which then proceeds with the generation of the artificially generated content 3 in the form of the song.
Claims
Method for interaction between a user (1) and an AI model (2), wherein the AI model (2) artificially generates acoustic content (3) and / or visual content (4) and causes the output of the acoustic content (3) via acoustic output means (5) and / or the output of the visual content (4) via visual output means (6), wherein the user (1) consumes the artificially generated content (3, 4), wherein the pupil size of at least one eye (7) of the user (1) is recorded during and / or after the output of the artificially generated content (3, 4), wherein the acoustic content (3) and / or visual content (4) is varied by the AI model (2) depending on the pupil size, and wherein at least one AI model parameter is changed and / or an input prompt is changed to vary the artificially generated content (3, 4), characterized in that the extent,with which at least one AI model parameter and / or the input prompt is changed, depending on the percentage change in pupil size. Method according to claim 1, characterized in that the AI model (2) enhances the acoustic (3) and / or visual content (4) when a change in pupil size during the output of the artificially generated content (3, 4) and / or after a defined period of time following the output of the artificially generated content (3, 4) is smaller than a defined first size change threshold. Method according to claim 1 or 2, characterized in that the AI model (2) reduces the acoustic (3) and / or visual content (4) or at least keeps it constant in scope if a change in pupil size during the output of the artificially generated content (3, 4) and / or after a defined period of time following the output of the artificially generated content (3, 4) is greater than a defined second size change threshold. Method according to one of claims 1 to 3, characterized in that the AI model (2) is formed by a generative model, in particular in the form of a generative pre-trained transformer and / or a diffusion model. Information technology system comprising acoustic (5) and / or visual output means (6), visual detection means (8) for detecting the pupil size of a user (1), and a computing unit (9) configured for controlling the output means (5, 6) and detection means (8) and for providing or interacting with an AI model (2), characterized in that the output means (5, 6), the detection means (8) and the computing unit (9) are configured for performing a method according to one of claims 1 to 4. Information technology system according to claim 5, characterized by an implementation in a vehicle (10). Information technology system according to claim 6, characterized in that the computing unit (9) is configured to receive a vehicle parameter and supply it to the AI model (2) in the form of input data for the design of the artificially generated content (3, 4). Information technology system according to claim 7, characterized in that the computing unit (9) is configured to receive an artificially generated control command from the AI model (2) taking into account the vehicle parameter and to control a vehicle subsystem with the control command.
Citation Information
Patent Citations
CUSTOMIZING AN ELECTRONIC DISPLAY BASED ON EYE TRACKING
DE102016104362A1
Display system intensity adjustment based on pupil dilation
US7744216B1
guest service system for vehicle users
DE102006015332A1
Method and device for playing music in a motor vehicle
DE102022121930A1
Reinforcement learning for automatically adjusting settings in vehicles
DE102022212854A1