Vehicle interface using generative artificial intelligence
By deploying cloud agents in vehicles in combination with large language models (LLMs), the limitations of existing LLMs in vehicle function control are overcome, enabling efficient interaction and function control between the vehicle and the cloud, and improving the user experience.
Patent Information
- Application Number
- CN202511135299.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-08-07
- Filing Date
- 2025-08-14
- Publication Date
- 2026-03-03
AI Technical Summary
Existing generative large language models (LLMs) have limitations in understanding and generating domain-specific responses, making it difficult to effectively provide functional control for vehicles.
By deploying a cloud agent in the cloud and combining it with the vehicle's voice assistant system using a large language model (LLM), interaction between the vehicle and the cloud can be achieved, providing voice control functions, including vehicle status acquisition and command execution.
It enables efficient interaction between the vehicle and the cloud, improves the accuracy and flexibility of vehicle function control, and enhances the user experience.
Smart Images

Figure CN121590569A_ABST
Abstract
Description
[0001] Related patent applications
[0002] This application claims the benefit of U.S. Provisional Application Serial No. 63 / 683,089, filed on August 14, 2024, entitled “VEHICLE INTERFACE USING GENERATIVE ARTIFICIAL INTELLIGENCE”. Background Technology
[0003] This disclosure relates to the use of cloud LLM to access vehicle functionality. Summary of the Invention
[0004] In one aspect, a computing system is configured to receive text from a vehicle and receive vehicle status from the vehicle. The computing system writes a prompt including the text, at least a portion of the vehicle status, and a set of instructions. The computing system submits the prompt to a machine learning model and receives a response to the prompt from the machine learning model, the response including an instruction selected from the instruction set. The computing system sends the selected instruction to the vehicle.
[0005] In some implementations, the response to the prompt includes one or more arguments to the selected instruction, and the computing system is also configured to send the selected instruction along with one or more arguments to the vehicle.
[0006] In some implementations, at least a portion of the vehicle state is considered as a part of the vehicle state. The computing system is configured to select a portion of the vehicle state based on its relevance to the text.
[0007] In some implementations, the computing system is configured to select an instruction set based on its relevance to the text.
[0008] In some implementations, the machine learning model is a large language model (LLM).
[0009] In some implementations, vehicle status includes the state of charge of the vehicle's battery.
[0010] In some implementations, vehicle status includes the status of the vehicle's climate control system.
[0011] In some implementations, the computing system is a cloud computing platform.
[0012] In another approach, one method includes: receiving text from a vehicle by a computing system, and receiving vehicle status from the vehicle by the computing system. The method includes: composing a prompt by the computing system that includes the text, at least a portion of the vehicle status, and a set of instructions. The method includes: submitting the prompt to a machine learning model by the computing system. The method includes: receiving a response to the prompt from the machine learning model by the computing system, the response including an instruction selected from the instruction set. The method includes: sending the selected instruction to the vehicle by the computing system.
[0013] In some implementations, the response to the prompt includes one or more arguments to the selected instruction, and the computing system is also configured to send the selected instruction along with one or more arguments to the vehicle.
[0014] In some implementations, at least a portion of the vehicle state is considered as a part of the vehicle state. The method may also include: a computing system selecting a portion of the vehicle state based on its relevance to the text.
[0015] In some implementations, the method includes: the computing system selecting a set of instructions based on their relevance to the text.
[0016] In some implementations, the machine learning model is a large language model (LLM).
[0017] In some implementations, vehicle status includes the state of charge of the vehicle's battery.
[0018] In some implementations, vehicle status includes the status of the vehicle's climate control system.
[0019] On the other hand, the vehicle includes a microphone, a computing system, and multiple components. The computing system is configured to detect speech in the microphone's output and generate speech-related context, which includes the state of one or more of the components. The computing system sends a request, including the speech and context, to a remote computing system. The computing system can receive a response to the request from the remote computing system. The computing system executes instructions included in the response regarding one or more components.
[0020] In some implementations, the computing system is configured to convert speech into text.
[0021] In some implementations, one or more components include the vehicle’s climate control system, and the context includes the state of the climate control system.
[0022] In some implementations, one or more components include the vehicle’s climate control system, and the context includes the state of the climate control system.
[0023] In some implementations, the computing system is configured to generate a response to a request from the vehicle, the response including text. The computing system is configured to detect connectivity with the vehicle. When the connectivity with the vehicle has a first quality, the computing system synthesizes a response to generate an audio message and sends the audio message to the vehicle. When the connectivity with the vehicle has a second quality less than the first quality, the computing system sends a response to the vehicle without synthesizing a response to generate an audio message. Attached Figure Description
[0024] Figure 1A Example vehicles that can be operated according to certain implementation schemes are illustrated.
[0025] Figure 1B An example is shown of a chassis of a vehicle with multiple drive units that is operable according to certain embodiments.
[0026] Figure 2 It is a schematic block diagram of components used to operate a vehicle according to certain implementation schemes.
[0027] Figure 3 This is a schematic block diagram illustrating a system for implementing a voice assistant in a vehicle according to certain implementation schemes.
[0028] Figure 4 This is a flowchart of a method for implementing a voice assistant in a vehicle, based on certain implementation schemes.
[0029] Figure 5 This is a schematic diagram illustrating voice quality control based on network connectivity quality according to certain implementation schemes. Detailed Implementation
[0030] Large language models (LLMs) have the ability to provide responses to natural language queries in the form of natural language output. While commercially available LLMs are highly sophisticated in understanding and generating text, their training is often very general, limiting their ability to provide accurate domain-specific responses. Using the method described in this paper, a vehicle, in conjunction with a cloud agent that interacts with the LLM, can leverage the LLM's capabilities to provide voice control functionality within the vehicle.
[0031] Figure 1A An example vehicle 100 in which the methods described herein can be implemented is illustrated. Figure 1A As shown, vehicle 100 has multiple external cameras 102 and one or more front displays 104. Each of these external cameras 102 can capture a specific view or perspective of the exterior of vehicle 100. The images or videos captured by the external cameras 102 can then be displayed on one or more displays in vehicle 100, such as one or more front displays 104, for the driver to view.
[0032] refer to Figure 1B The vehicle 100 may include a chassis 106, which includes a frame 108 that provides the main structural components of the vehicle 100. The frame 108 may be formed by one or more beams or other structural components, or may be integral with the vehicle body (e.g., a monolithic construction).
[0033] In embodiments where vehicle 100 is a battery electric vehicle (BEV) or possibly a hybrid vehicle, a large battery 110 is mounted to the chassis 106 and may occupy a significant portion (e.g., at least 80%) of the area within the frame 108. For example, battery 110 may store 100 to 200 kWh. Battery 110 may be a lithium-ion battery or other types of rechargeable battery. The battery may be substantially planar in shape.
[0034] Power from battery 110 can be supplied to one or more drive units 112. Each drive unit 112 may be formed by an electric motor and, possibly, a gear train providing gear reduction. In some embodiments, a single drive unit 112 is present, which drives the front or rear wheels of vehicle 100. In another embodiment, two drive units 112 are present, each driving the front or rear wheels of vehicle 100. In yet another embodiment, four drive units 112 are present, each driving one of the four wheels of vehicle 100.
[0035] Power from the battery 110 can be supplied to the drive unit 112 via one or more power modules 114 (such as power modules for each drive unit 112 or a pair of drive units 112). The power modules 114 may include inverters configured to convert direct current (DC) from the battery 110 into alternating current (AC) supplied to the motor of the drive unit 112. The power modules 114 also facilitate the operation of the drive unit's motor as a generator to provide regenerative braking. The power modules 114 further facilitate the transfer of regenerative current to the battery 110.
[0036] A drive unit 112 is coupled to two or more wheel hubs 116 to which wheels can be mounted. Each wheel hub 116 includes a corresponding brake 118, such as a disc brake as illustrated. Each wheel hub 116 is further coupled to a frame 108 via a suspension 120. The suspension 120 may include metal or pneumatic springs for absorbing shocks. The suspension 120 may be implemented as a pneumatic or hydraulic suspension capable of adjusting the ground clearance of the chassis 106 relative to a supporting surface. The suspension 120 may include a damper, wherein the characteristics of the damper are fixed or electronically adjustable.
[0037] exist Figure 1BIn the implementation scheme and in the discussion below, vehicle 100 is a battery electric vehicle. However, hybrid electric vehicles can also benefit from the methods described herein. Similarly, non-vehicle applications using inverters or other related power components can also benefit from the methods described herein.
[0038] Figure 2 Examples Figure 1A Example components of vehicle 100. (e.g.) Figure 2 As shown, vehicle 100 includes a camera 102, one or more front displays 104, a user interface 200, one or more sensors 202, a motion sensor 204, and a positioning system 206. The one or more sensors 202 may include ultrasonic sensors, radio detection and ranging (RADAR) sensors, light detection and ranging (LIDAR) sensors, or other types of sensors. The positioning system 206 may be implemented as a Global Positioning System (GPS) receiver. The user interface 200 allows a user (such as a driver or occupant in vehicle 100) to provide input.
[0039] Components of vehicle 100 may include one or more temperature sensors 208. Temperature sensors 208 may include sensors configured to sense ambient air temperature, battery 110 temperature, power module 114 temperature, temperature of each drive unit 112 and / or each motor of each drive unit 112 temperature, temperature of coolant fluid entering or leaving the coolant system, oil temperature within drive unit 112, or the temperature of any other component of vehicle 100. Temperature sensors 208 may include temperature sensors directly mounted to the microprocessor of power module 114, as described in more detail below.
[0040] The control system 214 executes instructions to perform at least some of the actions or functions of the vehicle 100. For example, such as... Figure 2 As shown, the control system 214 may include one or more electronic control units (ECUs) configured to perform at least some of the actions or functions of the vehicle 100, including the functions described below. In some embodiments, each ECU is dedicated to a specific set of functions.
[0041] Some features of the implementation scheme described herein can be controlled by a telematics control module (TCM) ECU. The TCM ECU can provide a wireless vehicle communication gateway to support functionality, by way of example and not limitation, such as over-the-air (OTA) software updates, vehicle-to-Internet communication, vehicle-to-computing device communication, in-vehicle navigation, vehicle-to-vehicle communication, vehicle-to-landscape features (e.g., automatic toll road sensors, automatic toll booths, power distributors at charging stations), or automatic calling functionality.
[0042] Some features of the implementation described herein can be controlled by a Central Gateway Module (CGM) ECU. The CGM ECU serves as the vehicle's communication hub, connecting various ECUs, sensors, cameras, microphones, motors, displays, and other vehicle components, and transmitting data to and from these components. The CGM ECU may include a network switch providing connectivity via a Controller Area Network (CAN) port, a Local Interconnect Network (LIN) port, and an Ethernet port. The CGM ECU can also function as the master controller for different vehicle modes (e.g., road driving mode, parking mode, off-road mode, trailer mode, camping mode), thereby controlling certain vehicle components associated with placing the vehicle in one of these vehicle modes.
[0043] In various implementations, the CGM ECU collects sensor signals from one or more sensors of the vehicle 100. For example, the CGM ECU may collect data from camera 102, sensor 202, motion sensor 204, positioning system 206, and temperature sensor 208. The sensor signals collected by the CGM ECU are then transmitted to the appropriate ECU for processing.
[0044] The control system 214 may also include one or more additional ECUs, as an example and not a limitation, such as a vehicle dynamics module (VDM) ECU, an experience management module (XMM) ECU, a vehicle access system (VAS) ECU, a near field communication (NFC) ECU, a body control module (BCM) ECU, a seat control module (SCM) ECU, a door control module (DCM) ECU, a rear zone control (RZC) ECU, an autonomous control module (ACM) ECU, an autonomous safety module (ASM) ECU, a driver monitoring system (DMS) ECU, and / or a winch control module (WCM) ECU.
[0045] If vehicle 100 is an electric vehicle, one or more ECUs may provide functionality related to the vehicle's battery pack, such as a Battery Management System (BMS) ECU, a Battery Power Isolation (BPI) ECU, a Balanced Voltage and Temperature (BVT) ECU, and / or a Thermal Management Module (TMM) ECU. In various implementations, the XMM ECU sends data to the TCM ECU (e.g., via Ethernet, etc.). Additionally or alternatively, the XMM ECU may send other data (e.g., audio data from microphone 216, etc.) to the TCM ECU.
[0046] refer to Figure 3The control system 214 can execute a voice assistant (VA) service 300 (“Service 300”). The VA service 300 includes various components of the control system 214 and other components 302 of the vehicle (such as…) Figure 2 The service 300 interacts with the ECUs (exemplified in the ECU or any ECU) via an application programming interface (API) 300a. API 300a enables the service 300 to control components and receive and react to events generated by other components.
[0047] Service 300 may include an audio processing module 300b that receives the output of a microphone and performs audio processing to facilitate the interpretation of spoken words in the microphone output. Audio processing module 300b may perform the illustrated functions and produce processed output. Speech model 300c receives the processed output and attempts to detect spoken words in the processed output. Speech model 300c may remain inactive except for attempting to detect a wake word (e.g., “Okay, Rivian”) in the processed output. Upon detection of a wake word, speech model 300c may perform other processing, such as speech-to-text (STT). Speech model 300c may also convert text to be output to a user into synthesized speech, such as text-to-speech (TTS), for output to a speaker.
[0048] Service 300 may include a user interface module 300d that receives input via an interface displayed on the front display 104 and displays output generated by service 300 on the front display 104. Information described herein as being output by a speaker may be supplemented by information displayed on the front display 104, including user interface elements for invoking the functionality of vehicle 100.
[0049] Service 300 may include router 300e. Router 300e routes events to one or more application plugins 300f. For example, an event may be speech detected using speech model 300c and routed to plugin 300f based on keywords included in the speech. Events may be generated by vehicle component 302. Events may include data received from cloud 304, as discussed below. Application plugin 300f may also generate events that are processed by another application plugin 300f.
[0050] Application plugin 300f can receive speech (e.g., text generated from speech) detected using speech model 300c, and receive the status of vehicle 100 from one or more vehicle components 302 and generate a request to be transmitted to cloud 304. The request may include speech and context. The context may include a portion of the vehicle status determined by application plugin 300f to be speech-related. For example, a request including speech relating to the climate in the vehicle's cabin (e.g., text derived from speech or an audio signal including speech) may include context reflecting the current state of the climate control system of vehicle 100.
[0051] In Cloud 304, the voice assistant agent 304a can interact with service 300, receive requests generated by service 300, and produce responses to those requests. Voice assistant agent 304a may include various application plugins 304b that handle events such as requests. Application plugins 304b can receive events via cloud router 304c, which routes events to application plugins 304b that are event-addressed or registered to receive specific types of events. Application plugins 304b can also generate events that are processed by other application plugins 304b.
[0052] The voice assistant agent 304a can interact with software component 306 that executes in or can otherwise be accessed through the cloud 304. Software component 306 may include dedicated logic for component 302 for diagnosing vehicle 100. Software component 306 may also provide responses to different types of queries, such as navigation, providing travel guidance, or other functions.
[0053] The voice assistant agent 304a can interact with the software component 306 via the cloud API 304d. The software component 306 can generate events that are transmitted to the voice assistant agent 304a via the cloud API 304d.
[0054] Application plugin 304b may submit prompts to a large language model (LLM) 304e or other generative artificial intelligence platform. The LLM 304e may include any commercially available LLM (GOOGLE GEMINI, OPENAI's CHATGPT, MICROSOFT COPILOT) or a proprietary LLM. Voice assistant agent 304a may function regardless of the LLM 304e used. Application plugin 304b may receive responses to prompts as events and process events, such as by submitting the responses to software component 306. For example, a prompt may include text along with speech and context from a request received from service 300, instructing the LLM 304e to generate appropriate instructions for software component 306 based on speech, context, and possibly a set of instructions selected from that set.
[0055] Cloud 304 can perform cloud audio processing functions 304f, such as STT and TTS. For example, if the connection has sufficient bandwidth, the request may include an audio segment instead of text, allowing cloud audio processing function 304f to translate the audio segment into text. Cloud audio processing function 304f can also additionally convert a text response generated by application plugin 304b into an audio segment, which can be sent to service 300 for playback on the cabin speakers of vehicle 100.
[0056] Application plugin 304b can also interact with cloud services 304g, such as third-party services including Google Maps, streaming music services (e.g., SPOTIFY), and restaurant reservation services (e.g., OPEN TABLE).
[0057] Figure 4 An example of a method 400, executable to process voice input received from a user of vehicle 100, is illustrated. At step 410, a problem audio is received, for example, from a microphone and audio processing module 300b. At step 412, the problem audio is processed, for example, via a voice model 300c or a cloud audio processing function 304f, to obtain problem text.
[0058] At step 414, the question text can be sent to the voice assistant agent 304a in the cloud 304. Before or simultaneously with step 414, the vehicle status can be sent to the voice assistant agent 304a. For example, when connected to the cloud 304, the control system 214 can maintain the current vehicle status 402a stored in the cloud 304. The vehicle status 402a may include the outputs of any of the vehicle's sensors (see...). Figure 2 (and corresponding descriptions), the status of any ECU in the vehicle's ECU, the vehicle's location, the vehicle's environment (e.g., temperature, rain sensor output, light sensor output, etc.), the vehicle's speed, and the current driving status (stopped, parked, driving, driving mode, etc.).
[0059] The voice assistant agent 304a may request a prompt to be generated by a prompt generator 304b, which may be an application plugin selected by application plugin 304b based on the type of request, keywords included in the question text, application plugin 300f that generates the request, or other criteria. The prompt generator 304b may generate the prompt at step 416. Generating the prompt may include requesting contextual information to generate the prompt. For example, the prompt generator 304b may request the current vehicle state 402a at step 418 and receive at least a portion of the current vehicle state 402a at step 6. Step 418 may include requesting a portion of the vehicle state 402a, such as a portion relevant to the question text, and then receiving that portion at step 420. In the illustrated example, the vehicle state includes state of charge (SOC), temperature, and geographic location. Relevance may include the same information domain corresponding to the question text among multiple domains. Any method used to determine text similarity may be used to select a domain based on the textual similarity between the domain description and the question text.
[0060] At step 422, prompt generator 304b may request driver profile 402b, and at step 424, it receives the driver profile. For example, the driver profile may include the driver of vehicle 100's home address, dietary preferences, or other attributes or preferences. In some embodiments, prompt generator 304b may request information from trip planner 402c.
[0061] The prompt generator 304b combines the question text with some or all of the information obtained at steps 420 and 424 to form a prompt. For example, the prompt generator 304b may extract the portion of the information obtained at steps 420 and 424 that is relevant to the question text. Relevance can be determined using any method used to determine the relevance of one text (e.g., a data object that includes the information obtained at steps 420 and 424) and another text (e.g., the question text). The prompt may additionally include a set of instructions, such as a set of instructions relevant to the question text. The relevance of the instruction set may include performing a textual comparison of the question text with the instruction set, a description of the instruction set, or individual instructions or other data within the instruction set.
[0062] The instruction set may include, for example, function calls to the application programming interface (API) of the control system 214 for controlling vehicle components such as the climate control system, the vehicle's drivetrain (e.g., function calls for selecting driving modes and adjusting suspension 120), and the vehicle's infotainment system (e.g., function calls for interacting with navigation software, controlling media playback, finding points of interest (POI) information, etc.).
[0063] The prompt generator 304b may return a prompt to the voice assistant agent 304a at step 426, which may then send the prompt to the LLM 304e at step 428. A response to the prompt is received at step 430. At step 432, the instruction included in the response to the prompt may be executed. For example, the LLM may select an instruction from a set of instructions. The instruction may include arguments that can also be selected by the LLM 304e. For example, the prompt may list and / or describe the possible arguments of the instructions in the set of instructions, and the prompt may instruct the LLM to select the arguments selected by the LLM 304e for the instruction.
[0064] Then, at step 432, the voice assistant agent 304a may submit the instruction to the vehicle component 302 corresponding to the instruction, the software component 306 in the cloud referenced by the instruction, or other components. For example, the cloud router 304c of the voice assistant agent 304a may route the instruction to the component referenced by the instruction. In the illustrated example, the instruction is provided to the trip planner 402c, which returns one or more points of interest (POIs) and a route to the destination referenced by the instruction at step 434.
[0065] Responses from components can be intents that can be processed in various ways, including invoking the generation of another prompt for LLM 304e. For example, the result of an instruction can be returned to service 300. For instance, a response to service 300 could indicate at step 436 that the user has indicated an intent to proceed along a route, or it could indicate at step 438 that the user intends to find a POI. Service 300 can receive responses from the user (e.g., route selection or POI selection) via verbal responses or input to the user interface module 300d.
[0066] In response, voice assistant agent 304a may send a callback response to LLM 304e at step 440. The callback response may include the intent from steps 436 and / or 438. At step 442, LLM 304e may return a response to voice assistant agent 304a. Voice assistant agent 304a may, for example, use speech model 300c to synthesize the response into a verbal response and return the verbal response to service 300 at step 18. Service 300 may then use the speaker of vehicle 100 to play back the verbal response.
[0067] In some implementations, the result from an instruction (e.g., from vehicle component 302 or software component 306) can be passed to LLM 304e to obtain a summary of the response, and this summary can be returned to the user and played back by the speakers of vehicle 100.
[0068] In some implementations, the response from LLM 304e to a prompt including a verbal question may be a request for additional information. The request for additional information may be sent to service 300 by voice assistant agent 304a before or after it is synthesized into a verbal message. The verbal message may be played back by the vehicle 100's speakers. The verbal response to the verbal message may be forwarded by service 300 to voice assistant agent 304a to be passed to LLM as additional context for the verbal question.
[0069] refer to Figure 5 The way the voice assistant agent 304a interacts with the service 300 can correspond to the quality of the connection between the vehicle 100 and the cloud 304. For example, in the absence of a connection, the service 300 can handle verbal questions through direct interpretation. For example, the verbal question can be converted into text, and the text can be matched against a fixed set of available commands. If a matching command is found, the command is executed. If the result of the command is a verbal response, the response is synthesized locally by the voice model 300c and played back on the speaker of the vehicle 100.
[0070] With Level 1 connectivity available, the voice assistant agent 304a interacts with the service 300 using text and other data. However, any responses from the LLM 304e or the voice assistant agent 304a are synthesized locally as speech by the speech model 300c and played back on the speakers of the vehicle 100. The locally synthesized speech may not be as realistic as the speech synthesized by the cloud audio processing function 304f. Features such as changes in emotion, speed, and volume, breathing, and contextual emphasis may be present in the speech synthesized by the cloud audio processing function 304f but not in the speech synthesized by the speech model 300c.
[0071] When Level 2 connectivity is available (Level 2 is greater than Level 1), cloud audio processing function 304f can be used to synthesize the response sent from voice assistant agent 304a to service 300 to obtain a verbal response, which can then be sent to service 300 for playback on the speakers of vehicle 100. Under Level 2 connectivity, verbal questions can be sent by service 300 as an audio file to voice assistant agent 304a, which is then converted to text using cloud audio processing function 304f, potentially resulting in greater accuracy. Level 2 verbal responses are likely to be more conversational and include more information compared to Level 1 verbal responses or verbal responses without connectivity.
[0072] As Figure 5Examples of methods. In the absence of connectivity, the user might need to explicitly say "set the temperature to X degrees" to increase the set temperature of the climate control system. In the case of second-level connectivity, the user could simply say "I feel cold." This statement can then be used to deduce the user's intention to increase the set temperature of the climate control system.
[0073] Without connectivity, some functions may be unavailable. For example, commands involving software component 306, such as advanced diagnostic functions, may be unavailable. For instance, error codes may be interpreted along with vehicle status to provide a more detailed explanation of the problem, while service 300 may simply output the error code.
[0074] In some implementations, service 300 may implement a smaller, less powerful LLM that handles verbal questions from users, used when connectivity is lost or in the presence of Level 1 connectivity. Service 300 may monitor the signal strength of the connection to cloud 304 and switch between operating modes (direct interpretation, Level 1 functionality, Level 2 functionality, etc.). In some implementations, synchronization may occur when connectivity is restored, where changes in vehicle state and events that occurred during the connectivity loss can be forwarded to cloud 304 to update vehicle state 402a. In some implementations, in the event of anticipated connectivity loss (e.g., the vehicle route passes through an area without connectivity), data such as road conditions, charging station locations, points of interest, restaurant locations, or other data that may be relevant during the period of connectivity loss may be pushed to service 300 by voice assistant agent 304a.
[0075] The system described above can be used to implement various use cases. In the first example, application plugin 300f can use LLM 304e to initiate the creation of a verbal response. For example, in response to an event generated by application plugin 300f, such as an event from vehicle component 302 or software component 306, the methods described above can be used to request the creation of a prompt. For example, weather data may indicate that it is snowing or may snow. Service 300 can invoke the generation of a verbal response, such as, “Would you like to switch to snow driving mode?” In another example, vehicle status received by voice assistant agent 304a may indicate that the first vehicle encountered snow or ice on a first location. In response, voice assistant agent 304a can invoke the replay of a question on a second vehicle 100 with a route passing through the first location, such as, “Snow or ice is expected ahead. Would you like to switch to snow driving mode?”
[0076] In another example use case, service 300 interacts with voice assistant agent 304a to operate in tour guide mode. For example, a user might ask, “Tell me the situation at this location.” The vehicle’s current location can be provided to the LLM along with the question text to obtain a response, which can then be synthesized and played back by the vehicle 100’s speakers.
[0077] In another example use case, vehicle 100 is detected to be within a threshold proximity to the user's house and closer to the user's specific garage door. Service 300 can invoke the generation of the verbal question "Would you like to open [garage door name]?" and process any verbal response using the methods described above to invoke the opening of the garage door.
[0078] In another example use case, the verbal user request is "open the front trunk," and the vehicle status indicates that the attempt to open the front trunk failed. When determining the response to the user request, the vehicle status provided to the voice assistant agent 304a can be used to provide context to the LLM 304e.
[0079] Various embodiments of this disclosure have been described for illustrative purposes. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been chosen to explain the principles of the embodiments, their practical application, or technical improvements to technologies found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0080] In the foregoing, reference has been made to the embodiments presented in this disclosure. However, the scope of this disclosure extends beyond the specifically described embodiments. Rather, any combination of features and elements is contemplated for implementing and practicing the contemplated embodiments, whether or not different embodiments are involved. Furthermore, while the embodiments disclosed herein may achieve advantages over other possible solutions or over the prior art, the embodiments may achieve some advantages or no particular advantages. Therefore, the aspects, features, embodiments, and advantages discussed herein are merely illustrative.
[0081] The various aspects of this disclosure may take the form of a completely hardware implementation, a completely software implementation (including firmware, resident software, microcode, etc.), or a combination of software and hardware implementations, all of which may be collectively referred to herein as “circuit,” “module,” or “system.”
[0082] Various aspects of this disclosure are described by narrative text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in implementations of a computer program product (CPP). Regarding any flowchart, depending on the technology involved, operations can be performed in a different order than that shown in a given flowchart. For example, again depending on the technology involved, two operations shown in consecutive flowchart frames can be performed in reverse order, as a single integrated step, concurrently, or in a manner that at least partially overlaps in time.
[0083] A Computer Program Product Implementation (“CPP Implementation” or “CPP”) is a term used in this disclosure to describe any set of one or more storage media (also referred to as “media”) collectively included in a collection of one or more storage devices, which collectively include machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. A “storage device” is any tangible device capable of holding and storing instructions for use by one or more computer processing devices. Without limitation, a computer-readable storage medium can be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Specific types of storage devices including these media include: magnetic disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory sticks, floppy disks, mechanical encoding devices (such as punched cards or pits / platforms formed in the main surface of the disk), or any suitable combination of the foregoing. As used in this disclosure, computer-readable storage medium refers to a non-transitory storage device rather than the transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides, optical pulses transmitted through fiber optic cables, and electrical signals transmitted through wires and / or other transmitting media. As those skilled in the art will understand, data typically moves at some incidental points in time during the normal operation of the storage device (such as during access, defragmentation, or garbage collection), but the storage device remains non-transitory during these processes because the data remains non-transitory while stored.
[0084] While the foregoing relates to embodiments of this disclosure, other and further embodiments may be devised without departing from the basic scope of this disclosure, the scope of which is defined by the appended claims.
Claims
1. A computing system, the computing system being configured to: Receive text from the vehicle; Receive vehicle status from the vehicle; Compile a prompt that includes the text, at least a portion of the vehicle status, and a set of instructions; Submit the prompt to the machine learning model; Receive a response to the prompt from the machine learning model, the response including an instruction selected from the instruction set; and Send the selected instructions to the vehicle.
2. The computing system of claim 1, wherein the response to the prompt includes one or more arguments of the selected instruction, and the computing system is further configured to send the selected instruction along with the one or more arguments to the vehicle.
3. The computing system of claim 1, wherein the at least portion of the vehicle state is a part of the vehicle state, and the computing system is further configured to select the portion of the vehicle state based on its relevance to the text.
4. The computing system of claim 3, wherein the computing system is configured to select the instruction set based on its relevance to the text.
5. The computing system according to claim 1, wherein the machine learning model is a large language model (LLM).
6. The computing system of claim 1, wherein the vehicle state includes the state of charge of the vehicle's battery.
7. The computing system of claim 1, wherein the vehicle state includes the state of the vehicle's climate control system.
8. The computing system according to claim 1, wherein the computing system is a cloud computing platform.
9. A method, the method comprising: The computing system receives text from the vehicle; The computing system receives the vehicle status from the vehicle; The computing system generates a prompt that includes the text, at least a portion of the vehicle status, and a set of instructions; The computing system submits the prompt to the machine learning model; The computing system receives a response to the prompt from the machine learning model, the response including an instruction selected from the instruction set; and The selected instructions are sent to the vehicle by the computing system.
10. The method of claim 9, wherein the response to the prompt includes one or more arguments to the selected instruction, and the computing system is configured to send the selected instruction along with the one or more arguments to the vehicle.
11. The method of claim 9, wherein the at least portion of the vehicle state is a part of the vehicle state, the method further comprising: The calculation system selects the portion of the vehicle state based on its relevance to the text.
12. The method of claim 11, further comprising the computing system selecting the instruction set based on its relevance to the text.
13. The method of claim 9, wherein the machine learning model is a large language model (LLM).
14. The method of claim 9, wherein the vehicle state includes the state of charge of the vehicle's battery.
15. The method of claim 9, wherein the vehicle state includes the state of the vehicle's climate control system.
16. A vehicle, the vehicle comprising: Multiple components; microphone; and The computing system is configured as follows: Voice is detected via the microphone; Generate a context related to the speech, the context including the state of one or more of the plurality of components; Send a request to a remote computing system, the request including the voice and the context; Receive a response to the request from the remote computing system; as well as The instructions included in the response pertain to the execution of the one or more components.
17. The vehicle of claim 16, wherein the computing system is configured to convert the speech into text.
18. The vehicle of claim 16, wherein one or more components include a climate control system for the vehicle, and the context includes the state of the climate control system.
19. The vehicle of claim 16, wherein one or more components include a climate control system of the vehicle, and the context includes the state of the climate control system.
20. The vehicle of claim 16, wherein the computing system is further configured to synchronize the state of the vehicle with the remote computing system.