System and method for using a psychophysiological state of a user for artificial intelligence

WO2026170137A1PCT designated stage Publication Date: 2026-08-13HARMAN BECKER AUTOMOTIVE SYSTEMS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-02-09
Publication Date
2026-08-13

Smart Images

  • Figure US2026014563_13082026_PF_FP_ABST
    Figure US2026014563_13082026_PF_FP_ABST
Patent Text Reader

Abstract

A method includes receiving a psychophysiological state of a user, executing an artificial-intelligence (AI) model responsive to an input to the AI model to generate an output, and actuating a component according to the output. The AI model uses the psychophysiological state of the user to interpret the input. A controller performs the method.
Need to check novelty before this filing date? Find Prior Art

Description

Atty. Doc. No. HARM1000PCTSYSTEM AND METHOD FOR USING A PSYCHOPHYSIOLOGICAL STATE OF A USER FOR ARTIFICIAL INTELLIGENCE TECHNICAL FIELD

[0001] Aspects disclosed herein generally relate to using a psychophysiological state of a user for enhancing interactions with artificial intelligence (Al). These aspects and others will be discussed in more detail herein.CROSS-REFERENCE TO RELATED APPLICATION

[0002] This application claims priority to provisional U.S. Patent Appl. Nos. 63 / 756,633 and 63 / 756,664, both filed on February 10, 2025, which are hereby incorporated by reference in their entirety.BACKGROUND

[0003] Technologies utilizing artificial intelligence (Al), including large language models (LLMs), are becoming an integral part of everyday life for many people. Although there are various types of Al models (e.g., machine learning (ML) models), LLMs, including unimodal and multimodal LLMs, are currently becoming more widespread and being successfully used to approach diverse and complex problems, thus, advancing capabilities in research and various industries. For a regular user, LLMs present wide-ranging opportunities in text generation and knowledge retrieval. Implementing LLMs for enhancing user experience in interaction with digital systems is a new fast-growing trend in the cutting-edge technology markets.SUMMARY

[0004] User interactions with Al models, such as LLMs, are often implemented as an exchange of information via a text chat or a voice assistant. A user creates a text query / prompt and an LLM provides an answer. To obtain relevant and informative answers from LLMs, users currently need to adjust their queries / prompts by providing meaningful contextual information. This process requires skills and may be time-consuming. When implementing an LLM as part of a solution designed as a voice assistant, it may be desirable to create an experience of a natural conversation requiring low effort from the user. This disclosure seeks to integrate information about the human psychophysiological state that can be derived from biosignals as part of contextual informationAtty. Doc. No. HARM1000PCTavailable to LLMs when processing user queries. An ability to interpret the user state based on biosignals will allow Al to make responses more relevant, informative, and empathetic.

[0005] References to Al responses as disclosed herein may generally relate to both verbal and non-verbal components of communication. Multimodal interactions with Al can include generated audiovisual effects, such as an anthropomorphic avatar which responds to changes in user state via adjustments in its voice and behavior, creating a fuller interpersonal experience.

[0006] Although scaling up LLM architectures is resulting in their growing ability to capture diverse knowledge encoded in text, non-linguistic aspects of communication may lead to successfully solving tasks and deriving the meaning of information from language as set forth in “Large Language Models are Few-Shot Health Learners”, LIU et al., Consumer Health Research Team, Google, 24 May 2023. Objective information about the human state can be obtained from biosignals and used by Al models. Integrating Al into biomedical signal analysis has been at the forefront of innovations within healthcare and wellbeing industries in recent years as noted in, for example, Lee, Y.J., Park, C., Kim, H., et al., “Artificial intelligence on biomedical signals: technologies, applications, and future directions”, Med-X 2, 25 (2024). One of the novel directions in such research is combining biosignals processing and natural language processing in solutions that leverage the rapid growth of LLMs. New systems, such as SignalGPT or BioSignal Copilot have been proposed for reducing medical signals to a freestyle or formatted clinical, technical reports capturing key features and characterization of input signal as noted in, for example, Chunyu Liu, Yongpei Ma, Kavitha Kothur, Armin Nikpour, Omid Kaveei, “BioSignal Copilot: Leveraging the power of LLMs in drafting reports for biomedical signals”, medRxiv 28 JUN.2023.

[0007] Despite these and other developments in the field of combining biosignals processing with LLM solutions, little has been done to implement the use of biosignals to enhance user interactions with Al.

[0008] The disclosed embodiments provide a system and a method for using human biosignals to enhance interactions with Al, including LLMs. The method includes, among other things, collecting biosignals from a user, processing these biosignals to create data representations in formats acceptable for input into Al model, adding formatted biosignals data to input along with the user query, training an Al model (e.g., LLM) to use information about the psychophysiologicalAtty. Doc. No. HARM1000PCTstate derived from biosignals when generating responses to user queries and interacting with the user, which may include verbal and non-verbal communication.

[0009] The system and method of using human biosignals enhances interactions with Al, including LLMs. Although there may be frequent references to LLMs herein as this is a widespread type of Al model used for user interaction at this time, it should be noted that the disclosed embodiments may relate to any type of Al model. One potential implementation of the system and / or method may be in a vehicle. Although the examples disclosed herein may be disclosed in such an implementation as an example, the proposed system and / or method may be utilized for other applications or implementations.

[0010] A method includes receiving a psychophysiological state of a user, executing an artificial-intelligence (Al) model responsive to an input to the Al model to generate an output, and actuating a component according to the output. The Al model uses the psychophysiological state of the user to interpret the input.

[0011] In an example, the input may include a query inputted by the user. In a further example, the output may include a conversational response to the query, and the Al model may determine the conversational response based on the psychophysiological state.

[0012] In an example, the input may include an environmental condition of an environment containing the user. In a further example, the output may adjust a parameter affecting the environmental condition.

[0013] In another further example, the environment may be a passenger compartment of a vehicle. In a still further example, the output may adjust a parameter affecting the environmental condition, the parameter including at least one of volume of a speaker of the vehicle, a setting of a driver assistance system, or a climate-control setting.

[0014] In an example, the Al model may be executed in response to the psychophysiological state exceeding a threshold.

[0015] In an example, the method may further include generating a textual prompt from the input and from the psychophysiological state, the textual prompt being directly inputted into the Al model.

[0016] In an example, the Al model may take multimodal input, and the input and the psychophysiological state may be directly inputted into the Al model as separate modalities. In aAtty. Doc. No. HARM1000PCTfurther example, the Al model may encode the input and the psychophysiological state as an embedding.

[0017] In an example, the output may include changing a behavior of an anthropomorphic avatar being shown to the user by a user interface.

[0018] In an example, the Al model may include a large language model.

[0019] In an example, the Al model may include an artificial neural network with transformers.

[0020] In an example, the psychophysiological state may include at least one of alertness, arousal, cognitive load, stress, fatigue, drowsiness, emotion, eye gaze metrics, heart rate, heart rate variability, or electrodermal activity of the user.

[0021] In an example, the method may further include determining the psychophysiological state based on sensor data depicting the user. In a further example, the method may further include receiving the sensor data from at least one of an optical sensor aimed at the user or a biometric sensor on the user.

[0022] In an example, a controller may include at least one memory device, and the controller may be programmed to perform one of the foregoing methods. In a further example, a system may include the controller.

[0023] In another further example, a vehicle may include the controller.BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The system may be better understood with reference to the following drawings and description. The components in the figures are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the invention. Moreover, in the figures, like-referenced numerals designate corresponding parts throughout the different views.

[0025] Figure l isa block diagram of an example system for executing an artificial-intelligence (Al) model.

[0026] Figure 2 is a flowchart of an example process for executing the Al model.DETAILED DESCRIPTION

[0027] With reference to the Figures, wherein like numerals indicate like parts throughout the several views, a method includes receiving a psychophysiological state of a user, executing an artificial-intelligence (Al) model responsive to an input to the Al model to generate an output, andAtty. Doc. No. HARM1000PCTactuating a component according to the output. The Al model uses the psychophysiological state of the user to interpret the input.

[0028] With reference to Figure 1, the system may be part of a vehicle 100. The vehicle 100 may be any passenger or commercial automobile such as a car, a truck, a sport utility vehicle, a crossover, a van, a minivan, a taxi, a bus, etc. The vehicle 100 may include a controller 105, a communications network 110, sensors 115, a user interface 120, a transceiver 125, one or more advanced driver assistance systems (ADAS) 130, and a climate-control system 135. Although the embodiments noted herein may primarily use a system and method implemented on the vehicle 100 as an example, it is recognized that the disclosed system and method proposed with this disclosure may be potentially used in other environments (e.g., home entertainment systems, operator monitoring systems, etc.).

[0029] The controller 105 is a microprocessor-based computing device such as a generic computing device including a processor and a memory device, an electronic controller or the like, a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a combination of the foregoing, etc. Typically, a hardware description language such as VHDL (VHSIC (Very High Speed Integrated Circuit) Hardware Description Language) is used in electronic design to describe digital and mixed-signal systems such as FPGA and ASIC. For example, an ASIC is manufactured based on VHDL programming provided pre-manufacturing, whereas logical components inside an FPGA may be configured based on VHDL programming (e g., stored in a memory electrically connected to the FPGA circuit). The controller 105 can thus include a processor, a memory device, etc. The memory of the controller 105 can include media for storing instructions executable by the processor as well as for electronically storing data and / or databases, and / or the controller 105 can include structures such as the foregoing by which programming is provided. The controller 105 can include multiple processors and memory devices coupled together.

[0030] The controller 105 may transmit and receive data through the communications network 110. The communications network 110 may be a controller area network (CAN) bus, Ethernet, WiFi, Local Interconnect Network (LIN), onboard diagnostics connector (OBD-II), and / or any other wired or wireless communications network. The controller 105 may be communicatively coupled to the sensors 115, the user interface 120, the transceiver 125, ADAS 130, the climatecontrol system 135, and other components via the communications network 110.Atty. Doc. No. HARM1000PCT

[0031] The sensors 115 may provide data about operation of the vehicle 100, for example, wheel speed, wheel orientation, and engine and transmission data (e.g., temperature, fuel consumption, etc.). The sensors 115 may detect the location and / or orientation of the vehicle 100. For example, the sensors 115 may include global positioning system (GPS) sensors; accelerometers such as piezo-electric or microelectromechanical systems (MEMS); gyroscopes such as rate, ring laser, or fiber-optic gyroscopes; inertial measurements units (IMU); and magnetometers. The sensors 115 may detect the external world, including objects and / or characteristics of surroundings of the vehicle 100, such as other vehicles, road lane markings, traffic lights and / or signs, road users, etc.

[0032] Specifically, for example, the sensors 115 may include optical sensors such as radar sensors, scanning laser range finders, light detection and ranging (lidar) devices, and image processing sensors such as cameras 140, 145. The cameras 140, 145 may include at least one internal camera 140 and at least one external camera 145.

[0033] The internal camera 140 and external camera 145 are configured to detect electromagnetic radiation in some range of wavelengths. For example, the cameras 140, 145 may detect visible light, infrared radiation, ultraviolet light, or some range of wavelengths including visible, infrared, and / or ultraviolet light. For example, the cameras 140, 145 can be charge-coupled devices (CCD), complementary metal oxide semiconductors (CMOS), or any other suitable type. The internal camera 140 may be positioned such that the field of view covers all or a part of a passenger compartment of the vehicle 100. The internal camera 140 may be part of a driver monitoring system. The field of view of the internal camera 140 may encompass a driver of the vehicle 100, as well as other occupants of the vehicle 100. The external cameras 145 may be positioned such that the fields of view cover an external environment around the vehicle 100. The external cameras 145 may be aimed generally horizontally away from the vehicle 100. The fields of view of the external cameras 145 may encompass the external environment in front of, to the sides of, and / or behind the vehicle 100.

[0034] The sensors 115 may include biometric sensors 150. The biometric sensors 150 may be any sensors capable of measuring a biometric characteristic. The biometric sensors 150 may include the internal camera 140, radar positioned in the passenger compartment, wearable devices such as smartwatches, etc. For example, the internal camera 140 may detect facial expression, pupil dilation, eye gaze direction, and / or other movement metrics. The internal camera 140 mayAtty. Doc. No. HARM1000PCTcollect thermovision if equipped for infrared detection. The radar may detect breathing rate or heartrate by tracking chest movement. The wearable device may detect heartrate and / or electrodermal activity. The wearable device may be connected to the communications network 110 via the transceiver 125.

[0035] The user interface 120 presents information to and receives information from a driver or passenger of the vehicle 100. The user interface 120 may be located on a dashboard in the passenger compartment of the vehicle 100, and / or wherever may be readily seen by the driver. The user interface 120 may include output devices such as dials, digital readouts, screens, speakers 155, visual devices 160, and so on for providing information to the driver, such as human-machine interface (HMI) elements such as are known. The user interface 120 may include input devices such as buttons, knobs, keypads, microphone 165, and so on for receiving information from the driver.

[0036] The microphone 165 is a transducer that converts sound to an electrical signal. The microphone 165 can be any suitable type, such as a dynamic microphone, which includes a coil of wire suspended in a magnetic field; a condenser microphone, which uses a vibrating diaphragm as a capacitor plate; a contact microphone, which uses a piezoelectric crystal; etc. The microphone 165 may be positioned to detect speech by the user, for example, in the passenger compartment of the vehicle 100, etc.

[0037] The speakers 155 are electroacoustic transducers that convert an electrical signal into sound. The speakers 155 can be any suitable type for producing sound audible to the occupants (e g., dynamic).

[0038] The visual devices 160 may include screens and / or lamps. The screens can be any suitable type for displaying content legible to the respective occupants, such as light-emitting diode (LED), organic light-emitting diode (OLED), liquid crystal display (LCD), plasma, digital light processing technology (DLPT), etc. The lamps generate light to illuminate the passenger compartment. The lamps may be any suitable type for providing illumination within the passenger compartment, including tungsten, halogen, high-intensity discharge (HID) such as xenon, lightemitting diode (LED), laser, etc. The lamps may be distinct from screens or readouts of the user interface 120 (i.e., are a separate component of the vehicle 100 from the screens or readouts). The lamps may be capable of adjusting their brightness level and / or adjusting their color.Atty. Doc. No. HARM1000PCT

[0039] The transceiver 125 may be adapted to transmit signals wirelessly through any suitable wireless communication protocol, such as cellular, Bluetooth®, Bluetooth® Low Energy (BLE), ultra- wideband (UWB), WiFi, IEEE 802.1 la / b / g / p, cellular-V2X (CV2X), Dedicated Short-Range Communications (DSRC), other RF (radio frequency) communications, etc. The transceiver 125 may be adapted to communicate with a remote server, that is, a server distinct and spaced from the vehicle 100. The remote server may be located outside the vehicle 100. For example, the remote server may be associated with another vehicle (e.g., V2V communications), an infrastructure component (e.g., V2I communications), a first responder, a mobile device or wearable device associated with the operator of the vehicle 100, etc. The transceiver 125 may be one device or may include a separate transmitter and receiver.

[0040] ADAS 130 are electronic technologies that assist drivers in driving and parking functions. Examples of ADAS 130 include forward proximity detection, lane-departure detection, blind-spot detection, braking actuation, adaptive cruise control, and lane-keeping assistance systems.

[0041] The climate-control system 135 provides heating and / or cooling to the passenger compartment of the vehicle 100. The climate-control system 135 may include a compressor, a condenser, a receiver-dryer, a thermal -expansion valve, an evaporator, blowers, fans, ducts, vents, vanes, temperature sensors, and other components that are known for heating or cooling vehicle interiors. The climate-control system 135 may operate to cool the passenger compartment by transporting a refrigerant through a heat cycle to absorb heat from the passenger compartment and expel the heat from the vehicle 100, as is known. The climate-control system 135 may include a heater core that operates as a radiator for an engine of the vehicle 100 by transferring some waste heat from the engine into the passenger compartment, as is known. The climate-control system 135 may include an electrically powered heater such as a resistive heater, positive-temperature-coefficient heater, electrically power heat pump, etc.

[0042] A user of a digital system, such as, for example, one installed in the vehicle 100, may intend to communicate with the digital system verbally in the same way the user may communicate with another human being. Some LLM solutions are predominantly text-based. A user creates a text query (prompt) and an LLM provides an answer in the forms of a text or audible response. Relevant and informative answers from an LLM require meaningful contextual information. One way, for example, is to provide this information in a prompt. Thus, efficient prompting of an LLMAtty. Doc. No. HARM1000PCTbecomes a skill that requires learning and can be time- and energy-consuming and frustrating for users. Another difference between human-human and human-LLM interactions is that a conversation between humans entails a vast amount of non-verbal and non-linguistic information that is perceived and considered by each participant. Thus, automatically supplying contextual information about the human state (e.g., arousal level, emotion expression, etc.) of the user in the embeddings would enhance human-AI interactions, including both verbal and non-verbal aspects of communication, if Al is presented as a virtual agent in the environment that has voice and appearance, for example, such as an anthropomorphic avatar.

[0043] In one possible implementation, a user has an LLM agent on their mobile device or smartphone and ear buds (or headphones). The user communicates with the LLM agent through, for example, a voice assistant which has a virtual appearance on the visual device 160 (e.g., screen) of the mobile device as, for example, an anthropomorphic avatar. Along with providing audio output for the voice assistant, the ear buds also include photoplethysmography (PPG) sensor collecting heart rate (HR) data. The HR data is made available to the LLM agent during periods of querying / prompting and receiving the answer (e.g., the chat between the user and LLM). In the process of generating a response to each prompt, the LLM analyzes information about the dynamics of user HR that may be indicative of the psychophysiological state. Increases in HR are associated with higher arousal levels which can be interpreted by the LLM along with other contextual information when generating responses. Verbal responses, voice characteristics, and visual features of the anthropomorphic avatar are generated based on user queries and HR.

[0044] In another implementation, the user may be a driver of the vehicle 100. A driver voice assistant with an LLM module (or LLM controller) is installed in the vehicle 100 and appears on a digital display as the anthropomorphic avatar. The internal camera 140 is used to collect information about user eye movement, facial expressions, and HR (via remote PPG). This captured video stream can be used by other systems (e.g., controllers) and algorithms to determine various parameters such as gaze directions or HR (via remote PPG). The algorithms can be rulebased or machine learning (ML) / deep learning (DL) algorithms. Usually, such algorithms receive 1 frame or a set of frames as input, and the algorithm processes the received information and prepares the result for transmission via a controller. For example, determining the direction of a user’s gaze can be implemented using only 1 frame, which comes from the internal camera 140 toAtty. Doc. No. HARM1000PCTthe input of a neural network, which produces the coordinates of the gaze direction in a world coordinate system.

[0045] The system may further use environmental data. It is recognized that in everyday life, a person is interacting with various aspects of the environment to achieve their goals. Some of them are not possible to control, such as seasons, time of day, weather, traffic, etc., and individuals need to adapt their behavior, e.g., via calendars, sleeping patterns, clothes, or travel arrangements. Other aspects can be managed to meet individual needs for safety, comfort, and other personal goals. For example, humans adjust lighting and heating in our homes, control speed and direction when we are driving, and choose any number of mechanisms for gathering information or communicating with friends and family. Rapid development of technology, especially recent advancements in Al powered applications, can be used for enhancing interactions with the environment, where it is relevant, to save time and effort for more important tasks.

[0046] In general, when a person is interacting with the environment, various biosignals can be used to determine the human state and multiple sources of information about the environment conditions are also available. Without automation, if a person is aware of any safety concerns or comfort needs, they can manually adjust various parameters of the environment, such as lighting, temperature, and sound volume. However, people are not always aware of their state, and it is not always convenient to manually adjust settings when such needs arise. If the environment is powered by an Al model capable of interpreting the information about the psychophysiological state and making inferences about what is possible to do to bring the person to the most optimal state or make the environment more comfortable, such adjustments could be automated, thus, increasing personal safety and comfort, thus, elevating overall user experience.

[0047] As a general overview, the controller 105 is programmed to execute an artificialintelligence (Al) model responsive to an input to the Al model to generate an output, with the Al model using the psychophysiological state of the user to interpret the input. The controller 105 receives data from the sensors 115, including biosignals from the biometric sensors 150. Sensor data is formatted and stored (e.g., as images, time-series, or numeric data). An additional state detection model, when executed by the controller 105, can be used to determine the user state. In one example, at least one machine learning (ML) unimodal or multimodal model, when executed by the controller 105, can be used to determine the psychophysiological state. For instance, user arousal level can be estimated based on heart rate (HR) data, or cognitive load can be estimatedAtty. Doc. No. HARM1000PCTbased on user HR and eye gaze data collected using the internal camera 140. Alternatively, sensor data can be processed, for example, by the controller 105, without any additional models and added directly to the Al model within the input. The Al model processes the query along with the psychophysiological state and encodes the same. The Al model (via the controller 105) generates a response to an input such as a query taking into account the information about the psychophysiological state of the user derived from the biosignals data. Visual and auditory features of the response can also be adjusted using a model or by the Al model, when executed by the controller 105. The Al model may be trained to use information about the psychophysiological state and the environment to predict what parameters of the environment can be adjusted to increase user safety and comfort. Additional training or finetuning can be used to improve the responses from the Al model.

[0048] For the purposes of this disclosure, “psychophysiological state” is defined as a current psychological, emotional, and / or physiological condition of someone. For example, the psychophysiological state may include at least one of alertness, arousal, cognitive load, stress, fatigue, drowsiness, emotion, eye gaze metrics, heart rate, heart rate variability, or electrodermal activity. Alertness is a state of active attention characterized by high sensory awareness, and it is a psychological and physiological state. Drowsiness is the opposite. Arousal is the physiological and psychological state of being awoken or of sense organs stimulated to a point of perception. It involves activation of the ascending reticular activating system in the brain, which mediates wakefulness, the autonomic nervous system, and the endocrine system, leading to increased heart rate and blood pressure and a condition of sensory alertness, desire, mobility, and reactivity. Cognitive load is the effort being used in the working memory when the user is processing information. Emotions are physical and mental states associated with thoughts, feelings, and a degree of pleasure or displeasure (e.g., happy, excited, scared, sad, angry, etc.). Eye gaze metrics are measures of the gaze direction of the user, such as where the user is looking, how frequently where the user is looking changes, etc. Heart rate is the frequency at which the user’s heart beats, generally measured in beats per minute. Heart rate variability is a measure of how the heart rate changes overtime. Electrodermal activity is the property of the human body that causes continuous variation in the electrical characteristics of the skin, also been known as skin conductance, galvanic skin response, electrodermal response, psychogalvanic reflex, skin conductance response, sympathetic skin response, or skin conductance level. Electrodermal activity can indicate arousal.Atty. Doc. No. HARM1000PCT

[0049] The controller 105 is programmed to receive the psychophysiological state of the user. For example, the controller 105 may receive the psychophysiological state in the form of a vector or listing of numerical values for heart rate, heart rate variability, electrodermal activity, and / or eye gaze metrics. For another example, the controller 105 may receive the psychophysiological state of the user by determining the psychophysiological state based on sensor data depicting the user (e.g., the biometric sensor data) and possibly other data that directly affects the psychophysiological state such as time of day, ambient brightness, ambient noise level, etc. For example, the controller 105 may determine some aspects of the psychophysiological state by direct measurement of the sensor data, such as heart rate, heart rate variability, or electrodermal activity. For another example, the computer may execute a machine-learning algorithm on image data from the internal camera 140, such as an object-recognition algorithm, a facial-detection algorithm, an eye-tracking algorithm, etc. The machine-learning algorithm may output an identification of a facial expression or an emotion or level or alertness or drowsiness associated with the identified facial expression, or a direction of the eye gaze. For another example, the controller 105 may combine multiple direct measurements indicative of a more generalized state such as stress or cognitive load. The direct measurements may include heart rate, electrodermal activity, and / or eye gaze movement. The controller 105 may combine the direct measurements by normalizing one or more of the direct measurements to make the units commensurable and then using a simple or weighted average. The output may be a metric indicating stress, cognitive load, etc. The psychophysiological state may include multiple results from the foregoing determinations.

[0050] The controller 105 is programmed to execute the Al model responsive to the input to the Al model to generate the output. As a general overview, the Al model may be executed in response to a trigger event, such as receiving the input from the user or the psychophysiological state having some value. As one example, the Al model may be a large language model (LLM). The input may include a query from the user and / or environmental data. The controller 105 may combine the input and the psychophysiological state into a textual prompt directly inputted to the Al model, or the input and the psychophysiological state may be directly inputted into the Al model as separate modalities that the Al model encodes as an embedding. The output from the Al model may include a conversational response to the input and / or an adjustment to a component of the vehicle 100.Atty. Doc. No. HARM1000PCT

[0051] The controller 105 may be programmed to execute the Al model in response to a trigger event occurring. For example, the trigger event may include receiving the input (e.g., the query) from the user (e.g., via the user interface 120). For another example, the trigger event may include the psychophysiological state exceeding a threshold. The threshold may be for one of the values constituting the psychophysiological state, e.g., the heartrate. The threshold may be a confidence value for a qualitative assessment of the psychophysiological state from a machine-learning model, e.g., confidence in classifying the user as having high arousal. A continuous functioning of such system can be implemented via periodical probing, for example, the psychophysiological state may be evaluated every, for example, 30 seconds, and adjustments are made when one or more state features are outside the acceptable range defined by thresholds, thus, triggering changes in one or more parameters of the environment as the output, as described below.

[0052] The Al model may be any suitable type of machine learning, such as a large language model, an artificial neural network, a hybrid model incorporating multiple types, etc. As one example, the Al model may be a large language model (LLM). The term “large language model” is used in its machine-learning sense of a computational model for natural language processing tasks having at least thousands of trainable parameters, generally millions or billions. The LLM may be any suitable type, such as a generative pretrained transformer. The LLM may be a single model or may be multiple agent models working together. The agents may each be an LLM.

[0053] LLM solutions may be predominantly text-based or multimodal. LLMs provide responses to supplied prompts. Efficient and successful prompting of an LLM is one way to achieve meaningful and relevant outcomes. LLMs can also be finetuned for specific purposes and perform better in certain tasks they have been exposed to previously.

[0054] LLMs operate by encoding the input text (along with information in other formats, such as images, sounds, etc., in case of multimodal LLMs) into a high-dimensional vector space, where semantic relationships between elements are preserved. The Al model then decodes this representation to generate a response, guided by the learned statistical patterns. The quality of the response can be influenced by various factors, including the prompt provided to the Al model, the Al model’s hyperparameters, and the diversity of the data used to train the Al model.

[0055] Alternatively or additionally, the Al model may include an artificial neural network (ANN) with transformers. An ANN includes a plurality of transformer-based layers configured to process sequential input data through self-attention mechanisms and feed-forward operations, forAtty. Doc. No. HARM1000PCTwhich each transformer layer includes multi-head attention modules for capturing contextual relationships across tokenized representations, positional encoding components for preserving order information, and normalization units for stabilizing gradient flow. The ANN may further incorporate parameter-sharing strategies and scalable architecture enabling parallelized computation across distributed hardware, thereby facilitating efficient training and inference for tasks involving natural language understanding, image recognition, and multimodal data integration.

[0056] The input is data provided to the Al model as part of executing the Al model. The input may include a query inputted by the user, for example, via the user interface 120. The user may type in the query, or the user may speak such that the microphone 165 is able to detect the speech. The audio data detected by the microphone 165 may be converted to a query using a speech-to-text algorithm. The input may further include one or more environmental conditions of an environment containing the user, for example, the passenger compartment of the vehicle 100. The environmental conditions may include time of day, geographic position, weather, in-cabin lighting, sound, temperature, etc. The data indicating the environmental conditions may be provided by the sensors 115 and / or by components of the vehicle 100 responsible for the environmental condition (such as the state of the in-cabin lighting).

[0057] The Al model is programmed to use the psychophysiological state of the user to interpret the input. If the Al model is an LLM, the Al model may determine the conversational response based on the psychophysiological state. The Al model may choose calming phrasing if the user is highly aroused, the Al model may remind the user to stay alert if the user is distracted, etc. Alternatively or additionally, the Al model may control a component of the vehicle 100 based on the psychophysiological state. The Al model may lower the temperature of the climate-control system 135 if the user is highly aroused, the Al model may dim nonessential lighting if the user is distracted, etc. More specifically, the input and the psychophysiological state may be combined into a textual prompt to the Al model, or the input and the psychophysiological state may be provided as separate modes of input, as described in turn below. In either case, the psychophysiological state is processed by the Al model along with the input as part of executing the Al model.

[0058] As one example, the controller 105 may be programmed to generate a textual prompt from the input and from the psychophysiological state. The textual prompt is directly inputted intoAtty. Doc. No. HARM1000PCTthe Al model. As one example, the input (e.g., the query), an introductory statement characterizing the psychophysiological state, and the psychophysiological state may be concatenated together. The introductory statement may described what kind of data is included in the psychophysiological state and include instructions for the Al model on how to use the psychophysiological state (e.g., stating that the psychophysiological state is heartrate data or arousal level and instructing the Al model to generate an output that tends to move the heartrate toward a target heartrate or target arousal level). The introductory statement may be prestored in the memory device of the controller 105. The introductory statement may be crafted using techniques from prompt engineering. Alternatively or additionally, such an introductory statement may be inputted to the Al model earlier in the conversation or session, describing how to use future textual prompts including the input and the psychophysiological state. Prompt engineering, i.e., the process of designing instructions to obtain the optimal output from an Al model, has been a growing field of research accelerated in recent years as noted in Chen et al. , “Unleashing the potential of prompt engineering in Large Language Models: a comprehensive review”, arxiv, Cornell University, 23 Oct. 2023. The environmental conditions may be similarly included in the textual prompt and prefaced by an introductory statement.

[0059] As another example, the Al model may take multimodal input, and the input and the psychophysiological state may be directly inputted into the Al model as separate modalities. A multimodal LLM is able to receive, interpret, and generate output from heterogeneous sources of input. As a multimodal LLM, the Al model may include a plurality of neural network modules including transformer-based encoders and decoders. The Al model may encode the input and the psychophysiological state as an embedding. For example, the Al model may integrate text, image, audio, and structured data inputs through modality-specific embedding layers that convert each input type into a unified latent representation. Embedding models encode input data into vectors, which are then used by LLMs to understand prompts and generate outputs in patterns that are found in their training data. The Al model may further employ cross-attention mechanisms to correlate semantic features across modalities, a multimodal fusion layer to generate a consolidated contextual state, and a generative output engine adapted to produce text or other modality-specific outputs based on the fused representation, thereby enabling context-aware reasoning, content synthesis, and task execution across diverse data formats within a single end-to-end architecture.Atty. Doc. No. HARM1000PCT

[0060] As another example, the previous two examples may be combined. The AT model may receive the psychophysiological state as part of the textual prompt and as a separate mode of input. For example, different aspects of the psychophysiological state may be put into the textual prompt or the separate mode of input (e.g., arousal level as part of the textual prompt and heartrate as a separate mode of input).

[0061] The output generated by the Al model may include a message or other display or sound for the user interface 120, an adjustment of a parameter of a component or system of the vehicle 100, etc. For example, the output includes a conversational response to the query inputted by the user, particularly if the Al model is an LLM. The conversational response may answer a question asked by the user, may describe parameters being adjusted by the output, may notify the user of environmental conditions or the user’s own psychophysiological state, etc.

[0062] For another example, the output may include changing the behavior of an anthropomorphic avatar being shown to the user by the user interface 120. For the purposes of this disclosure, “anthropomorphic” is defined as, for a nonhuman entity, having human form or characteristics. The anthropomorphic avatar may have eyes. Because the avatar has eyes, the avatar is capable of appearing to look at the user, for example, to return the user’s gaze and thereby indicate that the user interface 120 is ready to respond to the user. The change in behavior by the anthropomorphic avatar may include, for example, a specific gesture or motion by the avatar as shown by the visual device 160 of the user interface 120 and / or speaking by the avatar as outputted by the speaker 155 of the user interface 120. As one example, the motion by the avatar may include changing a gaze direction of the anthropomorphic avatar. The gaze direction of the anthropomorphic avatar resulting from the output may be directed toward the eyes of the user. The default gaze direction of the anthropomorphic avatar may be some direction that is not toward the user’s eyes, such as straight ahead. Changing the gaze direction of the avatar toward the user indicates that the avatar is responsive to the user’s initiation in a way that is easy for the user to immediately understand.

[0063] For another example, the output may include adjusting a parameter of one of the components of the vehicle 100, for example, a parameter affecting one or more of the environmental conditions included in the input to the Al model. The parameter may include at least one of volume of the speaker 155, a setting of the ADAS 130 (e.g., following distance, target speed, etc.), or a climate-control setting of the climate-control system 135 (e.g., target temperatureAtty. Doc. No. HARM1000PCTof the passenger compartment, fan speed, etc ). As one example, the adjusted parameters can be safety focused, such as increasing the following distance to compensate for increased reaction time, if it is determined that the driver is fatigued or drowsy. Other parameters can address specific user comfort requirements, for example, customizing lighting and the sound system based on the level of arousal.

[0064] The Al model may be pretrained, as described above, and then finetuned. Pretrained LLMs can be further finetuned to perform better in specific tasks. An LLM with only few-shot tuning may be capable of grounding various physiological and behavioral time-series data and making meaningful inferences on numerous health tasks in clinical and wellness contexts as discussed in, for example, Liu et al., “Large Language Models are Few-Shot Health Learners”, Consumer Health Research Team, Google, 24 May 2023. In another recent study, a research group from Massachusetts Institute of Technology (Kim et al., “Health-LLM: Large Language Models for Health Prediction via Wearable Sensor Data”, Massachusetts Institute of Technology, 12 Jan 2024) tested the capacity of eight state-of-the-art LLMs (Med-Alpaca, PMC-Llama, Asclepius, ClinicalCamel, Flan-T5, Palmyra-Med, GPT-3.5 and GPT-4) to deliver multi-modal health predictions based on contextual information (e.g., user demographics, health knowledge) and physiological data (e.g., resting heart rate, sleep minutes). The framework enabled LLMs to adapt to health predictions by prompting / training via wearable sensor data. Systems, such as SignalGPT or BioSignal Copilot, have been proposed for reducing medical signals to a freestyle or formatted clinical, technical reports capturing key features and characterization of input signal as in, for example, Liu et al, “BioSignal Copilot: Leveraging the power of LLMs in drafting reports for biomedical signals”, medRxiv, 28 June 2023.

[0065] The process of finetuning involves further training of an LLM on a smaller, domainspecific dataset. There are various types of techniques used for finetuning (e.g., see Parthasarathy et al., “The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities”, CeADAR Ireland’s Centre for Al, 23 Aug. 2024). The process involves adapting the pretrained LLM by updating its parameters. Thus, training an LLM to use information about the psychophysiological state when responding to queries requires a carefully prepared dataset. For the described purposes the disclosed embodiments provide the following approaches for dataset creation.Atty. Doc. No. HARM1000PCT

[0066] In one example, data including conversations between two or more people with simultaneous recording of one or more biosignals through wearable devices in at least one person is collected. A wearable device can be, for example, a smart watch with biosignals sensors, such as a photoplethysmography (PPG) sensor. The psychophysiological state can be determined based on heart rate (HR) data extracted from PPG.

[0067] In another example, a dataset can be created from collected video recordings of one or more people during communication. The psychophysiological state can be determined based on information about body movement, pose, facial expressions, voice, eye movement, remote PPG, and other data obtained from video recordings. Sources for video recordings could be a drivermonitoring-system camera installed in a vehicle, interview recordings and videos of real-life interactions accessed through open platforms and servers, such as YouTube®.

[0068] An Al model is subsequently trained to interpret the psychophysiological state from data derived from collected videos and biosignals and to use information about the psychophysiological state when interacting with a user. In general, a combination of finetuning and prompt engineering appears to be an optimal approach to developing LLM solutions that can interpret the human state and use this information as context for generating responses.

[0069] The controller 105 is programmed to actuate a component according to the output. For example, the controller 105 may actuate the visual device 160 of the user interface 120 to display a message (e.g., the conversational response) of the output, the speaker 155 to emit the message (eg., the conversational response) of the output, the user interface 120 to show the behavior of the anthropomorphic avatar, the climate-control system 135 or ADAS 130 according to the adjusted parameter, etc.

[0070] What follows is an extended example of the foregoing techniques. The controller 105 is installed in the vehicle 100. A driver voice assistant with, for example, the anthropomorphic avatar is included. Algorithms may be executed on the on-board computing unit and in the cloud. The controller 105 executes algorithms determining heartrate, stress levels, drowsiness levels, cognitive load levels, and other driver state parameters, which constitute the psychophysiological state. These parameters are responsible for the driver’s performance, their current mental and physical state, etc. For example, a driver who is very drowsy has a chance of falling asleep at the wheel and causing an accident. An additional algorithm may be executed by the controller 105 and may be used to regulate the avatar behavior and the vehicle systems. An LLM may be implementedAtty. Doc. No. HARM1000PCTin the cloud and responsible for processing the data and generating the avatar’s response and actions.

[0071] The biometric sensors 150 collect biosignals used as information about the psychophysiological state. This data is then directed into the Al model, which may be an LLM. The LLM can be finetuned to process biosignals in text mode or at the embedding level. For example, the LLM receives text as input, which the LLM, when executed, transforms into a set of vectors, where each ‘token’ corresponds to its own vector. Such sets of vectors are called “embeddings”. Further, inside the Al model, various operations with these embeddings take place, and the result is generated. Based on this, it is possible to feed biosignals in a text form, for example adding the user’s heartrate to the text: <text>, hr = 75 bpm. Conversely, the Al model may receive a separate input for the biosignal, where this biosignal will be converted into embeddings and used by the model further.

[0072] The LLM is given an introductory statement as part of the prompt, instructing the Al model how to use the driver state data and the task to be solved. An example of an introductory statement is below:

[0073] »You are an Al assistant that knows the psychophysiological characteristics of people and can interact with them using this knowledge. In each request, you will receive a direct query from the user and information about their heart rate (HR). Using the HR information and your knowledge of how HR reflects the psychophysiological state of a person, you should provide your responses accordingly.

[0074] »Your response should not include a mention of the direct HR that you received in the request, nor should the response include an interpretation of this HR. It simply should be such that you take this information into account and form a response so that it is understood as correctly as possible by the user. Obviously, adding HR to the response will not increase the level of understanding, you must decide for yourself how you will use the HR information. You also understand that questions and answers to them cannot always be tied to the current heart rate; do not always try to do this, do it only when it is relevant to the question and its context.

[0075] »When you already have a history of interaction, try to use information about the pattern of heart rate changes from request to request, and use knowledge about the change in dynamics, and not just a specific request, as is known, it is the dynamics that are a greater indicator of the state than the heart rate itselfAtty. Doc. No. HARM1000PCT

[0076] »The request will look like this:

[0077] »Query: “Text of the query”

[0078] »HR: Number representing the heart rate in beats per minute at the time of the query

[0079] An input from the user can be transmitted to the Al model with artificial modification to form the textual prompt. The parameters of the user’s state and the format of the Al model’s response will be first, so that this response corresponds to the agent’s behavior and actions, if necessary. The following example illustrates such a request:

[0080] »You got a message from another system in previous messages. This system measures the driver’s state and optimizes other vehicle systems for that state. This system is a part of you, you can use its information as your information, not provided by user.

[0081] »The driver’s arousal level at the time of request: 0.94.

[0082] »The speed of the car is 50 km / h.

[0083] »Here is the last user request: Recommend me a movie.

[0084] »HR: 104.0 bpm; User state HR while previous assistant answer: HR 101.0 beats per minute.

[0085] The LLM will generate a text response that is tailored to the driver’s current state. To illustrate this, in response to the example message above, the following response could be formulated, taking into account the driver’s high cognitive load:

[0086] »I recommend watching “The Secret Life of Walter Mitty.” It’s an inspiring and uplifting film that might help you relax and feel more at ease.

[0087] Furthermore, the output generated by the Al model may include instructions for modifying the anthropomorphic avatar’s behavior in accordance with the psychophysiological state of the driver. To illustrate this, in instances where the driver is experiencing drowsiness, the avatar can be designed to exhibit heightened levels of energy and high sound level, with the objective of diverting the driver’s attention from the sensation of drowsiness.

[0088] The movements of the avatar can be adjusted based on the psychophysiological state of the driver, reflecting natural changes in the interlocutor’s state and various circumstances. For example, if the driver appears stressed or anxious, the avatar’s movements could be calming and soothing to provide a sense of reassurance and comfort. Conversely, if the driver is alert and engaged, the avatar can mirror this by displaying more dynamic and interactive gestures.Atty. Doc. No. HARM1000PCT

[0089] The generation of avatar movements can be accomplished using various methods. One approach involves using an LLM to produce prompts for a separate Al model dedicated to generating animations as set forth in, for example, Zhou et al., “EMDM: Efficient Motion Diffusion Model for Fast and High-Quality Motion Generation” arxiv, Cornell University, 4 Dec.2024. Another method includes issuing commands to an animation generator, which can create tailored movements based on specific parameters and constraints, as set forth in, for example, Tirinzoni et al., “Zero-Shot Whole-Body Humanoid Control via Behavioral Foundation Models”, Reinforcement Learning, Meta, 12 Dec. 2024. Additionally, an end-to-end solution as set forth in, for example, Zhou et al. might involve directly generating explicit movements for a 3D model by specifying changes in the coordinates of individual parts of the model. These diverse techniques ensure that the anthropomorphic avatar’s behavior is both responsive and realistic, enhancing the overall interaction experience.

[0090] The artificial modifications as noted above generally correspond to the conversion of a user’s message into the expected format for the LLM, as well as the addition of additional parameters. For example, a user’s request could be: “Recommend me a film”. This request will be augmented and converted into this format for example:

[0091] »User request: Recommend me a film.

[0092] »User HR: 65 bpm

[0093] »User HR while previous assistant answer: 68 bpm.

[0094] What follows is another extended example of the foregoing techniques. The user may be the driver of the vehicle 100. The internal camera 140 may be used to collect information about eye movements and heartrate (e.g., via remote PPG algorithms). The user may be wearing a smartwatch that outputs information about their heartrate. The controller 105 may determine the psychophysiological state using a neurosense universal state detector indicating data about user arousal levels and a cognitive load detector indicating data about cognitive load levels. The Al model may be an LLM that receives data about the psychophysiological state of the user derived from collected biosignals and state detection models and processes the data along with the environmental conditions (e.g., time of day, geolocation, etc.) and other contextual information (e.g., navigation details) available within the input. The controller 105 supplies the Al model with an instruction that delineates the data input parameters, ranges of variation in state / biosignalsAtty. Doc. No. HARM1000PCTvalues, and description of the parameters that the AT model is able to control. An example of the input is as follows:

[0095] »You are a system that helps the driver of a car. You can receive information about the speed of the car, as well as the arousal (excitability) indicator, where 0 indicates that the person is asleep, 2 indicates that the person is stressed. In the middle is the optimal state, in which the driver shows the best driving performance.

[0096] »You also have the ability to control the car’s systems. Here is a list of possible actions:

[0097] » <action> waiting < / action> you can skip current action and wait for some changes of driver state.

[0098] » <action> notification, “text of notification” < / action> you can create a notification and say something to the driver to warn the driver about something.

[0099] » <action> music, “name of possible track” < / action> you can control the playlist to play more active or quieter music.

[0100] » <action> turn off music < / action> you can turn off music for increasing driver’s attention in a stressful situation.

[0101] » <action> turn on lane keeping < / action> you can turn on ADAS such as lane keeping for increasing driver safety.

[0102] » <action> turn off lane keeping < / action> you can turn off ADAS such as lane keeping.

[0103] » <action> reduce emergency braking < / action> you can reduce distance to activate emergency braking.

[0104] » <action> increase emergency braking < / action> you can increase distance to activate emergency braking.

[0105] »Input format is the following:

[0106] »Driver arousal: <arousal number>, car speed: <speed number>. Give your recommendation for the next state and apply an action from your list.

[0107] The following is an example of an input to the LLM:

[0108] »Driver arousal: 0.2, car speed: 45. Give your recommendation for the next state and apply an action from your list.Atty. Doc. No. HARM1000PCT

[0109] The LLM processes this information and generates a response for adjusting one or more parameters of the environment and optionally informing the user of changes and the reasons behind them to facilitate safety and comfort. For example, if the LLM infers that the user may be tired, the LLM may generate an output to issue a voice notification: “Please, stay alert if you are feeling tired; consider taking a break”. The LLM may also increase the audio volume and change music to more upbeat and energetic. The LLM may automatically switch the settings of the ADAS 130 to enhance lane keeping and braking. The LLM may also change the visual device 160 of the user interface 120 to reduce the size of items displayed on the screen to bring the driver’s attention to the safety notification.

[0110] Figure 2 is a flowchart illustrating an example process 200 for executing the Al model. The memory device of the controller 105 stores executable instructions for performing the steps of the process 200 and / or programming can be implemented in structures such as mentioned above. As a general overview of the process 200, the controller 105 receives data from the sensors 115 and the user interface 120, determines the psychophysiological state of the user, generates the input for the Al model, executes the Al model to generate an output, and actuates a component according to the output.

[0111] The process 200 begins in a block 205, in which the controller 105 receives data from the sensors 115 and the user interface 120, including biometric data, a query from the user, etc., as described above.

[0112] Next, in a block 210, the controller 105 receives the psychophysiological state of the user based on the data from the block 205, as described above.

[0113] Next, in a block 215, upon the occurrence of a trigger condition, the controller 105 generates the input and / or the rest of the textual prompt to provide to the Al model based on the data from the block 205 and the psychophysiological state from the block 210, as described above.

[0114] Next, in a block 220, the controller 105 executes the Al model to generate the output based on the input from the block 215, as described above.

[0115] Next, in a block 225, the controller 105 actuates a component according to the output from the block 220, as described above. After the block 225, the process 200 ends.

[0116] As required, detailed embodiments of the present invention are disclosed herein; however, it is to be understood that the disclosed embodiments are merely exemplary of the invention that may be embodied in various and alternative forms. The figures are not necessarilyAtty. Doc. No. HARM1000PCTto scale; some features may be exaggerated or minimized to show details of particular components. Therefore, specific structural and functional details disclosed herein are not to be interpreted as limiting, but merely as a representative basis for teaching one skilled in the art to variously employ the present invention.

[0117] The disclosure has been described in an illustrative manner, and it is to be understood that the terminology which has been used is intended to be in the nature of words of description rather than of limitation. Many modifications and variations of the present disclosure are possible in light of the above teachings, and the disclosure may be practiced otherwise than as specifically described. Operations, systems, and methods described herein should always be implemented and / or performed in accordance with an applicable owner’ s / user’s manual and / or safety guidelines.

Claims

Atty. Doc. No. HARM1000PCTCLAIMSWhat is claimed is:

1. A method comprising:receiving a psychophysiological state of a user;executing an artificial-intelligence (Al) model responsive to an input to the Al model to generate an output, the Al model using the psychophysiological state of the user to interpret the input; andactuating a component according to the output.

2. The method of claim 1, wherein the input includes a query inputted by the user.

3. The method of claim 2, wherein the output includes a conversational response to the query, and the Al model determines the conversational response based on the psychophysiological state.

4. The method of claim 1, wherein the input includes an environmental condition of an environment containing the user.

5. The method of claim 4, wherein the output adjusts a parameter affecting the environmental condition.

6. The method of claim 4, wherein the environment is a passenger compartment of a vehicle.

7. The method of claim 6, wherein the output adjusts a parameter affecting the environmental condition, the parameter including at least one of volume of a speaker of the vehicle, a setting of a driver assistance system, or a climate-control setting.

8. The method of claim 1, wherein the Al model is executed in response to the psychophysiological state exceeding a threshold.

9. The method of claim 1, further comprising generating a textual prompt from the input and from the psychophysiological state, the textual prompt being directly inputted into the Al model.Atty. Doc. No. HARM1000PCT10. The method of claim 1 , wherein the Al model takes multimodal input, and the input and the psychophysiological state are directly inputted into the Al model as separate modalities.

11. The method of claim 10, wherein the Al model encodes the input and the psychophysiological state as an embedding.

12. The method of claim 1, wherein the output includes changing a behavior of an anthropomorphic avatar being shown to the user by a user interface.

13. The method of claim 1, wherein the Al model includes a large language model.

14. The method of claim 1, wherein the Al model includes an artificial neural network with transformers.

15. The method of claim 1, wherein the psychophysiological state includes at least one of alertness, arousal, cognitive load, stress, fatigue, drowsiness, emotion, eye gaze metrics, heart rate, heart rate variability, or electrodermal activity of the user.

16. The method of claim 1, further comprising determining the psychophysiological state based on sensor data depicting the user.

17. The method of claim 16, further comprising receiving the sensor data from at least one of an optical sensor aimed at the user or a biometric sensor on the user.

18. A controller comprising:at least one memory device;wherein the controller is programmed to perform the method of one of claims 1-17.

19. A system comprising:the controller of claim 18.

20. A vehicle comprising:the controller of claim 18.