Systems and methods for providing context-relevant vehicle instructions

The AI/ML-based vehicle instruction system generates personalized and easily understandable guidance using a 3D vehicle model, addressing the limitations of general user manuals and online videos by providing specific and efficient vehicle component assistance.

DE102025100633A1Pending Publication Date: 2025-07-17FORD GLOBAL TECH LLC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
DE102025100633
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-12
Filing Date
2025-01-09
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing vehicle user manuals and online guidance videos are often general and difficult for users to understand, lacking specific and user-friendly instructions for vehicle component operations and repairs.

Method used

A vehicle instruction generation system using AI/ML to analyze user inputs, generate personalized instruction media content, and display it through a 3D vehicle model at an optimal viewpoint, providing step-by-step guidance for vehicle component operations and repairs.

Benefits of technology

The system provides user-specific, easily understandable, and contextually relevant guidance, enhancing user convenience and efficiency in operating, repairing, or maintaining vehicle components without needing to consult general manuals or search online.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A system for issuing vehicle instructions is disclosed. The system may include a transceiver, a memory, and a processor. The transceiver may be configured to receive user input associated with a vehicle from a user, and the memory may be configured to store a trained machine model. The trained machine model may be trained using training data including a plurality of vehicle component identifiers and a plurality of user command intents. The processor may be configured to acquire the user input and determine a user intent based on the user input.The processor may be further configured to identify a vehicle component associated with the user intent by executing instructions stored in the trained machine model and generate instructional media content based on the vehicle component and the user intent. The processor may further output the instructional media content.
Need to check novelty before this filing date? Find Prior Art

Description

REGIONThe present disclosure relates to systems and methods for providing contextual vehicle instructions to a user.GENERAL STATE OF THE ARTThere are known cases where users seek support when they need to repair their vehicle or require guidance to operate one or more vehicle components. For example, a user may search for assistance when the windshield wipers may not properly clean the windshield or when the user may not know the approach to operating a new vehicle feature or component. In such situations, the users typically look up in the user manuals or search the Internet for guidance video to obtain the necessary support.Searching for guidance videos on the Internet may be awkward and may not be effective because most of these videos are general in nature and do not provide the specific instructions that the users may search for. Further, the user manuals are also general in nature and may not provide support to users in all scenarios. Furthermore, user manuals typically include technical terms that may be difficult for a layman to understand.Therefore, a system is required that provides assistance to vehicle users in an easily comprehensible manner.SUMMARYThe present disclosure describes a vehicle instruction generation system ("system") that may be configured to generate and provide instruction or teaching media content to a vehicle user that may assist the user in efficiently operating, repairing, and / or maintaining one or more vehicle components associated with a vehicle. The user may provide user input to the system when the user desires to search the system for assistance with respect to one or more vehicle components. As an example, the user may provide the user inputs when one or more vehicle components may not function properly or need to be replaced, when the user may search for guidance to operate a new vehicle feature / component, and / or the like. In an exemplary aspect, the user input may be a natural language-based voice command. The system may obtain the user input from the user and may generate instruction or teaching media content (e.g., a "guidance video") based on the user input, which may assist the user in conveniently learning about the procedure for repairing, installing, replacing, and / or operating the associated vehicle component(s).In some aspects, the system may be an artificial intelligence (AI) / machine learning (ML) based system that can analyze user input and accordingly generate optimal instruction media content that can efficiently support the user. The system may include a large language model (LLM) agent, which may include one or more trained machine learning modules and / or algorithms that may analyze user input and generate instruction media content.In an example aspect, in response to obtaining the user input from the user, the system may determine a user intent based on the user input by performing natural language processing on the user input. The system may perform natural language processing on the user input by executing the instructions stored in one or more LLM agent modules. The user intent may indicate an accurate user request (particularly when the user input is vague or implicit) and / or a reason or context that may have caused the user to provide the user input to the system. In response to determining the user intent, the system may execute the instructions stored in the LLM agent to identify one or more vehicle components associated with the user intent and then spontaneously generate the instruction media content based on the identified vehicle component(s) and the user intent. The system may further cause a display screen associated with a human machine interface (MMS) of the vehicle and / or a user device to display the generated instruction media content to enable the user to view the instruction media content comfortably.In some aspects, the system may cause the display screen to display / render the instruction media content using a three-dimensional (3D) digital model of the vehicle exterior and interior (which may be prestored in a system memory). The system may be configured to determine an optimal viewpoint associated with the digital 3D model of the vehicle exterior and interior based on the identified vehicle component(s) and the user's intention, and cause the display screen to display the instruction media content at the optimal viewpoint, such that the user can efficiently and easily capture the media content.The present disclosure discloses a vehicle instruction generation system that generates and provides contextual instruction media content or guidance video to the user based on requests from the user. Through the use of the system, the user does not need to search for general videos on the Internet or look up the instruction manuals / user manuals to learn more about the vehicle components. The system operates by obtaining voice-based commands from the user, thereby greatly improving user convenience. Further, the system generates the instruction media content based on specific needs and requirements of a user, and therefore the generated media content is highly relevant to the user. The system further ensures that the generated media content is displayed to the user at an optimal viewing angle so that the user can efficiently and easily capture the media content.These and other advantages of the present disclosure are provided in detail herein.BRIEF DESCRIPTION OF THE DRAWINGSThe detailed description will be set forth with reference to the accompanying drawings. The use of the same reference numerals may indicate similar or identical elements. For various embodiments, elements and / or components other than those illustrated in the drawings may be used, and some elements and / or components may not be present in various embodiments. The elements and / or components in the figures are not necessarily drawn to scale. Throughout the disclosure, terms in the singular and plural may be used interchangeably depending on the context. FIG. 1 illustrates an example environment in which techniques and structures for providing the systems and methods disclosed herein may be implemented. FIG. 2 depicts a block diagram of an example vehicle instruction generation system according to the present disclosure. FIG. 3 depicts a three-dimensional (3D) digital vehicle model displayed on a human-machine interface (MMS) of the vehicle at a standard angle of view, in accordance with the present disclosure. FIG. 4 depicts a sequence of views displayed on the vehicle MMS illustrating movement of the 3D digital vehicle model from the standard viewpoint to an optimal viewpoint, in accordance with the present disclosure. FIG. 5 depicts a flowchart of an example vehicle instruction generation method according to the present disclosure.DETAILED DESCRIPTIONThe disclosure will be described in more detail hereinafter with reference to the accompanying drawings, in which exemplary embodiments of the disclosure are illustrated, and is not intended to be limiting.FIG. 1 illustrates an example environment 100 in which techniques and structures for providing the systems and methods disclosed herein may be implemented. The environment 100 may include a user 102 seated within a vehicle 104. The vehicle 104 may take the form of any passenger or commercial vehicle, for example, a car, a work vehicle, a crossover vehicle, a truck, a van, a minivan, a taxi, a bus, etc. The vehicle 104 may be a manually driven vehicle and / or may be configured to operate in a semi- or fully autonomous mode, and may include any powertrain, such as a gasoline engine, one or more electrically-actuated engine(s), a hybrid system, etc.In some aspects, the user 102 may need assistance associated with the one or more vehicle components of the vehicle 104. In an example aspect, the user 102 may need assistance when the user 102 may need to install or replace a vehicle component in the vehicle 104, repair a vehicle component, operate a new vehicle component from which the user 102 may not know anything (or for which he may not know the procedure for servicing), perform maintenance of a vehicle component, and / or the like. In such cases, the user 102 may issue a command to the vehicle 104 to view one or more video content that may provide the user 102 with the necessary assistance. The video content may be, for example, instruction or teaching video content or guidance video that may provide the user 102 with stepwise instructions that the user 102 may need to execute / perform to complete the task associated with one or more vehicle components.The vehicle 104 may be communicatively coupled to a vehicle instruction generation system 106 (or system 106), which may be configured to obtain the user command (e.g., a voice-based command or a voice-based user input) and generate instruction media content (e.g., a guidance video) associated with one or more vehicle components and to be provided to the user 102 based on the user input. In some aspects, the system 106 may be part of the vehicle 104. In other aspects, the system 106 may be part of a server (not shown) and communicatively coupled to the vehicle 104 via a wireless network.The system 106 may be an artificial intelligence (AI) / machine learning (ML)-based system configured to spontaneously generate the instruction media content associated with one or more vehicle components based on the user input, regardless of whether the user 102 provides an explicit command associated with the vehicle components or an implicit command. For example, the user 102 may provide the system 106 with an explicit voice command (or user input) that states: "Ceige me how to fill the windshield wiper fluid." In this case, the system 106 may obtain the explicit command from the user 102 and generate instruction media content (e.g., a guidance video) that illustrates to the user 102 the procedure for filling the windshield wiper fluid. In response to generating the instruction media content, the system 106 may cause a display screen 108 associated with a human machine interface (MMS) of the vehicle (or a display screen associated with a user device) to display the generated instruction media content.If the user 102 provides an implicit voice command (or user input), e.g., "why the windshield wipers are not properly cleaning" (as represented by a view 110 in FIG. 1 ), the system 106 may first determine a user intent based on or based on the user input. In some aspects, the system 106 may include a large language model (LLM) agent, which may include one or more algorithms and / or one or more trained machine models (trained using supervised machine learning algorithms) that may assist the system 106 to analyze the implicit user input using natural language processing and determine the user intent. In further aspects, the system 106 may determine the user intent by performing one or more of a sential analysis, an emotional analysis, a user age analysis, and / or the like on the implicit user input by executing instructions associated with one or more natural language processing algorithms of the LLM agent.By determining the user intent, the system 106 may identify an "actual destination" of the user input or one or more possible reasons that caused the user 102 to provide the implicit user input to the system 106. For example, if the user 102 provides the command as shown in view 110, the system 106 may determine that either the vehicle windshield wipers may be defective or need to be repaired or the windshield wiper fluid needs to be replenished.In response to determining the user intent, the system 106 may identify one or more vehicle components associated with the determined user intent. Continuing with the example described above, the system 106 may determine the windshield wipers and / or windshield wiper fluid as the vehicle components associated with the user intent. In other words, the system 106 may identify that the implicit user input may affect windshield wipers and / or windshield wiper fluid, or the user 102 may have provided the implicit user input to the system 106, as the user 102 may have problems with the windshield wipers and / or windshield wiper fluid. In some aspects, the system 106 may identify the vehicle component(s) based on user intent by executing instructions stored in one or more trained machine models, which may be part of the LLM agent. In an example aspect, the LLM agent may include a trained machine model that may be trained using training data including a plurality of user command intents and a plurality of vehicle component identifiers. The trained machine model may be configured to predict or estimate a vehicle component identifier (and thus a corresponding vehicle component, e.g., the windshield wipers, windshield wiper fluids, etc.) based on a user intent that is input to or provided as input to the trained machine model. The system 106 may use the instructions included in the trained machine model to identify the vehicle component(s) to which the user 102 is referring (or may have problems) in the implicit user input based on the determined user intent.In some aspects, the system 106 may further identify the vehicle component(s) (or confirm the vehicle component(s) identified using the LLM agent described above) based on internal vehicle data / information, e.g., measured wiper fluid levels, detected electrical malfunction, diagnostic trouble codes, or DTCs, and / or the like. The system 106 may obtain such internal vehicle data from one or more vehicle sensors and / or individually from the vehicle components.In response to identifying the vehicle component(s), the system 106 may spontaneously generate instruction media content based on the determined vehicle component and the user intent. For example, the system 106 may generate an instruction video illustrating the stepwise procedure for replenishing the windshield wiper fluid when the particular vehicle components may be associated with the windshield wiper fluid.As described above, the system 106 further performs a sential analysis, an emotional analysis, and / or a user age analysis on the user input to determine the user intent. In some aspects, the system 106 performs the sential and / or emotional analysis to determine whether the user 102 may be quiet, applied, confused, eile, antagonised, panic, etc., and performs the user age analysis to determine a possible age profile of the user 102. The system 106 may be configured to generate the instruction media content based on the sential analysis, the emotional analysis, and / or the user age analysis. As an example, the system 106 may generate a more complex and detailed step-by-step instruction video when the system 106 determines that the user 102 may be an adult and may generate a general video when the user 102 may be a vissable child. In additional aspects, the system 106 may determine the user intent based on known user profiles associated with vehicle users, particularly when more than one user may be using the vehicle 104 periodically. The user profiles may include known features, characteristics, etc. of the users, and the system 106 may "learn" such user profiles associated with the vehicle users over time as the users use the vehicle 104. In other aspects, such user profiles may be added / provided to the system 106 by the respective users.In some aspects, system 106 may generate the instruction media content by executing instructions associated with one or more video generation modules / models, which may be part of the LLM agent. The video generation module may generate the instruction media content based on a user manual associated with the vehicle 104 (and may be prestored in a system memory), the particular vehicle component(s) and user intent, one or more prestored video captures associated with the vehicle component(s), a three-dimensional (3D) digital model of the vehicle exterior and interior (which may be prestored in the system memory), and / or the like. As an example, if the vehicle components may be associated with the windshield wiper fluid, the video generation module may use one or more prestored video images associated with a vehicle hood, windshield wiper fluid tank, and / or the like to assemble them and spontaneously generate instruction media content based on the user intent and the prestored video images. In some aspects, the instruction media content may use the vehicle exterior and interior 3D model (or the 3D vehicle model) as a "base," such that the media content is three-dimensional in nature and thus more immersive and helpful in assisting the user 102 (e.g., to conveniently provide the user 102 with the technique for replenishing windshield wiper fluid).In response to generating the instruction media content, the system 106 may cause the display screen 108 to display the instruction media content. In some aspects, the system 106 may additionally determine an optimal viewpoint (or "camera angle") to display the instruction media content on the display screen 108 using the 3D vehicle model so that the user 102 may easily and efficiently understand the instruction media content. The system 106 may determine the optimal viewpoint based on the determined vehicle component(s) and the user intent. As an example, when the vehicle components are associated with the windshield wiper fluid, the system 106 may determine that the optimal viewpoint for displaying the instruction media content using the 3D vehicle model may be a front view of the vehicle from above (as opposed to a side view or a rear view of the 3D vehicle model) such that the user 102 may efficiently and clearly view the steps illustrated in the instruction media content and understand the procedure for replenishing the windshield wiper fluid.In response to determining the optimal viewpoint, the system 106 may cause the display screen 108 to display the 3D vehicle model at the optimal viewpoint and then begin displaying or "rendering" the instruction media content. In some aspects, the system 106 may additionally output teletexts (e.g., text headers, captions, and / or the like) and / or audio prompts associated with the steps included in the instruction media content, while the instruction media content may be displayed / played back on the display screen 108. The image texts and / or audio prompts may improve user convenience in viewing the instruction media content and allow the user 102 to easily grasp the steps displayed on the instruction media content.Further system details are described below in connection with FIG. 2.The vehicle 104 and system 106 implement and / or perform operations as described herein in the present disclosure in accordance with the user manual and security policies. Moreover, any action taken by the user 102 should meet all of the regulations specific to the location and operation of the vehicle 104 (e.g., federal, country, city, etc.). The notifications, recommendations, or media content as provided by the vehicle 104 and / or the system 106 should be treated as suggestions and followed only according to any regulations specific to the location and operation of the vehicle 104.FIG. 2 depicts a block diagram of the vehicle instruction generation system 106 according to the present disclosure. As described above in connection with FIG. 1, in some aspects, the system 106 may be part of the vehicle 104. In other aspects, the system 106 may be part of a server (not shown) that may be communicatively coupled to the vehicle 104 via a wireless network. The wireless network may be and / or include the Internet, a private network, a public network, or other configuration operating using any one or more of any known communication protocols, such as transmission control protocol / Internet protocol (TCP / IP), Bluetooth ®, BLE® Institute of Electrical and Electronics Engineers (IEEE) 802.11 standard, ultra wide band (UWB), and cellular techniques such as time division multiple access (TDMA), code division multiple access (CDMA), high speed packet access (HSPDA), long term evolution (LTE), for example, Global System for Mobile Communications (GSM) and Fifth Generation (5G) to name a few examples. In yet another aspect of the present disclosure, one or more system components may be part of the vehicle 104 and remaining system components may be part of a server. In the description of FIG. 2, reference is made to FIGS. 3 and 4.The vehicle 104 may include a variety of units / components including, but not limited to, a vehicle sensor unit 202, the display screen 108 associated with a vehicle MMS, a vehicle microphone 204, and / or the like. The vehicle sensor unit 202 may be configured to measure / detect inputs associated with a vehicle operating status (e.g., whether the vehicle engine is in the ON or OFF state, whether the vehicle 104 is in motion or stationary, etc.), a vehicle speed, a geographic location of the vehicle (e.g., via signals obtained from a global positioning system), a weather condition (e.g., temperature, presence of snow, rain, sunlight, ambient light intensity, etc.) associated with a vehicle environment, a vehicle occupant status, and / or the like. The vehicle sensor unit 202 may transmit the inputs described above (or "sensor inputs") to the system 106 a predefined number of times.The vehicle microphone 204 may be configured to capture the voice-based command (e.g., an implicit or explicit user input as described above in connection with FIG. 1 ) provided to the vehicle 104 by the user 102. The vehicle 104 may be configured to transmit (via a vehicle transceiver, not shown) the user input detected by the vehicle microphone 204 to the system 106.It will be understood by those of ordinary skill in the art that the vehicle 104 may include a variety of additional units / components not shown in FIG. 2 and not described in the present disclosure. The vehicle units / components depicted in FIG. 2 are illustrative and should not be construed as limiting. The vehicle 104 may include additional units / components without departing from the scope of the present disclosure.The system 106 may include a plurality of units / components including, but not limited to, a transceiver 206, a processor 208, and a memory 210. The transceiver 206 may be configured to receive / transmit data / information / signals / media content from external systems and devices. For example, the transceiver 206 may receive the sensor inputs from the vehicle sensor unit 202, the implicit or explicit user input (provided by the user 102) from the vehicle 104, and / or the like. The transceiver 206 may further transmit the instruction media content and / or the 3D vehicle model to the vehicle 104 to cause the display screen 108 to display the instruction media content and / or the 3D vehicle model. The transceiver 206 may additionally transmit audio requests to the vehicle 104, which may be output by the vehicle 104 along with the instruction media content via one or more vehicle speakers (or via the vehicle MMS).The processor 208 may be an artificial intelligence (AI) / machine learning (ML)-based processor that may be disposed in communication with one or more storage devices disposed in communication with the respective computing systems (e.g., the memory 210 and / or one or more external databases not shown in FIG. 2 ). The processor 208 may use the memory 210 to store programs as code and / or to store data for performing aspects according to the disclosure. The memory 210 may be a transitory computer readable storage medium or a transitory computer readable memory having stored therein program codes that may enable the processor 208 to perform operations in accordance with the present disclosure. The memory 210 may include any one or combination of volatile memory elements (e.g., dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), etc.), and may include any one or more of nonvolatile memory elements (e.g., erasable programmable read only memory (EPROM), flash memory, electronically erasable programmable read only memory (EEPROM), programmable read only memory (PROM), etc.).In some aspects, the memory 210 may include a variety of databases, modules, and agents, including, but not limited to, a large language model (LLM) agent 212, a trained machine module 214 (or a trained machine learning module), a video generation module 216, a natural language processing module 218, and a vehicle information database 220. The vehicle information database 220 may be configured to store information associated with the vehicle 104, e.g., information associated with a user manual of the vehicle 104, a three-dimensional (3D) digital model of the vehicle exterior and interior (or 3D vehicle model), labels / labels of a plurality of vehicle components, pre-generated or pre-loaded video recordings associated with the plurality of vehicle components, and / or the like.In some aspects, the trained engine module 214, the video generation module 216, and the natural language processing module 218 may be part of the LLM agent 212, as shown in FIG. 2. In other aspects, one or more of the trained engine module 214, the video generation module 216, and the natural language processing module 218 may be located external to the LLM agent 212. The trained machine module 214 may be trained (e.g., using a technique / supervised machine learning algorithm) using training data including a plurality of vehicle component identifiers (e.g., labels / labels of a plurality of vehicle components) and a plurality of user command intents. In some aspects, the training data may be stored in the memory 210 or on a remote server, which may be communicatively coupled to the system 106. The trained machine module 214 may be trained by the processor 208 (using the training data) and may be configured to output labels / labels of possible vehicle components when a user intent may be input to the trained machine module 214.The natural language processing module 218 may store instructions associated with one or more natural language processing algorithms that may be configured to determine the user intent based on a natural language voice command-based user input (e.g., the implicit or explicit user input described above in connection with FIG. 1 ) by performing natural language processing, sentiment analysis, emotion analysis, user age analysis, and / or the like on the user input. In additional or alternative aspects, as described above, the system 106 may determine the user intent based on known user profiles associated with vehicle users, particularly when more than one user may be using the vehicle 104 periodically. The details of the user profiles have already been described above in connection with FIG. 1.The video generation module 216 may be configured to generate instruction media content (e.g., guidance video) spontaneously based on the user intent, the vehicle component(s), the information associated with the user manual, and / or the sensor inputs (obtained from the vehicle sensor unit 202). In some aspects, the video generation module 216 may use the video recordings and / or the 3D vehicle model stored in the vehicle information database 220 to generate optimal instruction media content that may provide contextual assistance to the user 102 according to the situation in which the user 102 may be located or based on accurate requirements of the user.In some aspects, LLM agent 212 may include one or more additional modules in addition to the modules described above without departing from the scope of the present disclosure. The LLM agent 212 may use the instructions included in the modules to correctly determine the intention of the user based on the user input (e.g., based on the natural language voice command of the user), generate relevant instruction media content for the user 102, and allow the display screen 108 to optimally display / output the instruction media content such that the user 102 may capture the media content in the most efficient and easily understood manner.In some aspects, the modules included in the LLM agent 212 may be trained by the processor 208 using a supervised machine learning technique. It will be appreciated by those of ordinary skill in the art that machine learning is an artificial intelligence (AI) application, which, when used, may have the ability of the systems or processors (e.g., processor 208) to automatically learn from experiences and improve without being explicitly programmed. Machine learning focuses on using data and algorithms to mimic how people learn. In some aspects, machine learning algorithms may be generated to generate classifications and / or predictions. Machine learning based on systems may be used for a variety of applications including, but not limited to, speech recognition, image or video processing, statistical analysis, natural language processing, content generation, and / or the like.Machine learning may be of various types based on data or signals available to the learning system. For example, the machine learning approach may include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. Supervised learning is an approach that can be supervised by a human. In this approach, the machine learning algorithm may use labeled training data and defined variables. In the case of supervised learning, both the input and the output of the algorithm may be fixed / defined, and the algorithms may be trained to classify data and / or accurately predict results.Generally, supervised learning may be one of two types: "regression" and "classification.". In classification learning, the learning algorithm may assist in dividing the dataset into classes based on different parameters. In this case, a computer program may be trained on the training dataset and the computer program may categorize input data into different classes based on the training. Some known methods used in classification learning include logistic regression, K-nearest neighbor, support vector machines (SVM), kernel SVM, naive bayes, decision tree classification, and random forest classification.In regression learning, the learning algorithm may predict an output value that may be continuous in nature or a real value. Some known methods used in regression learning include simple linear regression, multiple linear regression, polynomial regression, support vector regression, decision tree regression, and random forest regression.Unsupervised learning is an approach that includes algorithms that can be trained with unlabeled data. An unsupervised learning algorithm may analyze the data itself and find patterns in input data. Further, semi-supervised learning is a combination of supervised learning and unsupervised learning. A semi-supervised learning algorithm includes labeled training data; however, the semi-supervised learning algorithm may still find patterns in the input data. Reinforcement learning is a multistage or dynamic process. This model is similar to supervised learning, but cannot be trained using example data. This model can learn "on-the-fly" by using trial and error. In reinforcement learning, a sequence of successful results may be reinforced to develop the best recommendation or policy for a given problem.As described above, the modules of the LLM agent 212 may be trained using a supervised machine learning approach. The modules may be updated (or enhanced) when more training data may be input to the processor 208, e.g., when the processor 208 acquires and "learns" user inputs provided by the user 102 over a period of time, or when the processor 208 acquires updated training data from a remote server.In operation, the user 102 may transmit / provide a natural language voice command or "user input" to the vehicle 104 when the user 102 may need assistance associated with the vehicle 104. As described above in connection with FIG. 1, the user 102 may provide explicit user input or implicit (or vague) user input. The vehicle microphone 204 may detect the user input and the vehicle 104 may transmit the user input to the system 106.The transceiver 206 may receive the user input from the vehicle 104 (e.g., from the vehicle microphone 204). Additionally, the transceiver 206 may receive the sensor inputs from the vehicle sensor unit 202. The transceiver 206 may transmit the received user input and the sensor inputs to the processor 208.The processor 208 may obtain the user input and the sensor inputs from the transceiver 206. In response to obtaining the user input, the processor 208 may determine a user intent based on the user input, as described above in connection with FIG. 1. As described above, processor 208 may additionally determine the user intent based on the user profile / s. In some aspects, the user intent may indicate the precise user request associated with the vehicle 104 and / or a reason or context that may have caused the user 102 to provide the user input to the system 106. As an example, if the user 102 provides a voice command, "Why clean the windshield wipers not right?", the processor 208 may determine that the user intent may be to indicate that the vehicle windshield wipers may be faulty or that windshield wiper fluid may need to be replenished.In some aspects, processor 208 may determine the user intent based on the user input by performing natural language processing. In particular, the processor 208 may determine the user intent based on the user input by executing instructions associated with one or more natural language processing algorithms stored in the natural language processing module 218. In additional aspects, the processor 208 may execute the instructions associated with natural language processing algorithms stored in the natural language processing module 218 to perform a sential analysis, emotional analysis, and / or user age analysis on the user input. As described above in connection with FIG. 1, the processor 208 may perform the analysis described above to determine whether the user 102 may be quiet, applied, confused, eile, antagonised, panic, etc., and / or to determine a possible age profile of the user 102. In some aspects, processor 208 may use additional LLM modules (not shown) along with natural language processing module 218 to determine the user intent based on the user input and / or user profile / s.In response to determining the user intent, the processor 208 may identify one or more vehicle components that may be associated with the user intent. In some aspects, the processor 208 may identify the vehicle components by executing the instructions stored in the trained machine module 214. In particular, the processor 208 may "enter" the user intent into the trained machine module 214, which may output the expected vehicle components associated with the user intent. For example, if the user intent may indicate that the user 102 wishes to replenish the windshield wiper fluid, the processor 208 may identify one or more vehicle components (e.g., a vehicle hood, a windshield wiper fluid tank, and / or the like) that the user 102 may need to access to replenish the windshield wiper fluid. Additionally, as described above in connection with FIG. 1, the processor 208 may determine / confirm the vehicle component(s) based on the internal vehicle data / information.In response to identifying the vehicle components, the processor 208 may generate instruction media content (e.g., a guidance video) based on the vehicle components and the user input, which may efficiently contribute to the user 102 the procedure for performing the task desired by the user 102 (e.g., to replenish the windshield wiper fluid). In some aspects, the generated instruction media content may be video content, which may include instructions to enable or apply to the user 102 to perform one or more of installing the vehicle components, operating the vehicle components, replacing the vehicle components, repairing the vehicle components, and servicing the vehicle components (e.g., replenishing windshield wiper fluid to enable the vehicle windshield wipers to function properly).The processor 208 may generate the instruction media content by executing instructions stored in the video generation module 216. As described above, the video generation module 216 may be configured to generate the instruction media content based on the determined user intent, the identified vehicle components, the information associated with the user manual (which the processor 208 / video generation module 216 may obtain from the vehicle information database 220), and / or the sensor inputs (which the video generation module 216 may obtain from the processor 208). As an example, if the user intent indicates that the user 108 needs to replenish the windshield wiper fluid, the video generation module 216 may generate video content (e.g., by using the video images and / or the 3D vehicle model stored in the vehicle information database 220) that provides the user 102 with the action of actuating one or more vehicle components to conveniently replenish the windshield wiper fluid. In an exemplary aspect, the video content may provide the user 102 with the vehicle hood opening procedure, the wiper fluid tank accessing procedure, the wiper fluid filling procedure, the wiper fluid tank and vehicle hood closing / securing procedure, and / or the like. As another example, if the sensor inputs indicate that the weather conditions in the environment of the vehicle may be dark and cloudy, the video generation module 216 may generate low brightness video content as compared to if the weather conditions indicate a sunny day or much natural or artificial light in the environment of the vehicle. Similarly, the video generation module 216 may generate video content with multiple audible or audio prompts when the sensor inputs indicate that the vehicle 104 may be moving, and may include a lesser number of audio prompts when the vehicle 104 may be stationary. As yet another example, the video generation module 216 may generate detailed step-by-step instruction video content when the user intent indicates that an adult may provide the user input and not a vissual child (in which case the video generation module 216 may generate general video content). In some aspects, the generated video content may be based on the information included in the user manual or instruction manual such that the action included in or taught by the video content matches or authorizes the user manual or instruction manual.In response to generating the instruction media content / video content as described above, the processor 208 may cause the display screen 108 to output / display the instruction media content. In some aspects, to cause the display screen 108 to output / display the instruction media content, the processor 208 may first retrieve the 3D vehicle model from the vehicle information database 220 and cause the display screen 108 to output / display the instruction media content using the 3D vehicle model, as described below.In an example aspect, the processor 208 may first cause a default view of the 3D vehicle model to be displayed on the display screen 108. In some aspects, the default view on the display screen 108 may be a top view angle of the 3D vehicle model, as shown in FIG. 3. In particular, as shown in FIG. 3, a viewing angle 302 from above of the 3D vehicle model may be displayed on the display screen 108, for example, if the generated instruction media content / video content may not be displayed or rendered on the display screen 108. Those of ordinary skill in the art can understand that the default view on the display screen 108 may also be any other viewpoint of the 3D vehicle model that is different from the top viewpoint 302, e.g., a side viewpoint, a rear viewpoint, and / or the like. Hereinafter, the viewing angle 302 from above in the present disclosure is referred to as a standard viewing angle 302.In some aspects, before causing the display screen 108 to output / display or render the generated instruction media content, the processor 208 may determine an optimal gaze angle associated with the 3D vehicle model to display the instruction media content based on the identified vehicle component and the determined user intent. The optimal viewpoint may be the viewpoint at which it is most convenient for the user 102 to understand or grasp the steps represented in the instruction media content. As an example, if the instruction media content may be associated with the wiper fluid refill procedure, the processor 208 may determine that the optimal viewpoint is a top, front, or isometric front view at which the user may conveniently "see" the steps to be performed on the 3D vehicle model. In some aspects, processor 208 may determine the optimal gaze angle associated with the 3D vehicle model by executing the instructions stored in one or more modules of LLM agent 212.In response to determining the optimal viewpoint, the processor 208 may cause the view of the 3D vehicle model that may be displayed on the display screen 108 to rotate / change from the default viewpoint 302 to the optimal viewpoint, as shown in FIG. 4. In particular, as shown in FIG. 4, the processor 208 may cause the display screen 108 to change / rotate the view from the default viewpoint 302 to an optimal viewpoint 402 associated with the 3D vehicle model in a series or sequence of multiple steps (shown as views 404 and 406 in FIG. 4 ).The processor 208 may then cause the display screen 108 to begin outputting / displaying / rendering the instruction media content in response to changing / rotating the default angle of view 302 to the optimal angle of view 402. In some aspects, processor 208 may additionally generate step-by-step image instructions 408 (e.g., by executing instructions stored in LLM agent 212) and cause display screen 108 to display them, which may allow user 102 to conveniently learn the steps required to perform the task user 102 wishes to perform (e.g., fill in windshield wiper fluid). In some aspects, the task may require specific auxiliary materials that may or may not be available in the vehicle 104. If the tools are available in the vehicle 104 (e.g., tire changing kit, portable tire pump, etc.), the instructions 408 may direct the user 102 to the location of those required tools. On the other hand, if the assets are not available in the vehicle 104, the vehicle 104 may query a server to identify a method of obtaining the required assets, e.g., direct the user to a nearby retail location where the required assets are to be purchased.The processor 208 may additionally generate the audio commands / prompts associated with the steps shown on the display screen 108 and cause the vehicle speakers (not shown) to output them to enable the user 102 to easily understand / grasp the steps. An example audio prompt says: "They should replenish the windshield wiper fluid. I point to you how this is." is depicted as prompt 410 in FIG. 4. In an exemplary aspect, the processor 208 may cause the vehicle speakers to output the prompt 410 when the view on the display screen 108 may change / rotate from the default angle of view 302 to the optimal angle of view 402 (or at any time before or after the view changes / rotates). Further, the processor 208 may output the request 410 in response to the user input (which may be, for example, "why cleaning the windshield wipers not right?", as described above in connection with FIG. 1 ). It will be understood by those of ordinary skill in the art that the audio prompts, views, etc. described in the present disclosure and depicted in the figures are for illustration purposes only and should not be construed as limiting.The user 102 may conveniently view the instruction media content on the display screen 108 and accordingly may execute or learn the task to execute the task that the user 102 wishes to perform. In this manner, the system 106 enables the user 102 to more conveniently learn about the various features, components, and / or policies associated with the vehicle 104 without having to manually access the user manual or search for assistance on the Internet. The user 102 may additionally contract and expand the display screen 108 to enlarge or reduce the views of the instructional media content displayed on the display screen 108 (e.g., to focus on particular vehicle components). The user 102 may additionally gather and separate the views based on voice-based commands.In some aspects, once the instruction media content is generated by the processor 208, the processor 208 may additionally transmit the instruction media content to a remote server or cloud via the transceiver 206 such that other vehicles / systems that may require similar instruction media content retrieve (and use) the media content generated by the processor 208 from the remote server or cloud.FIG. 5 depicts a flowchart of an example vehicle instruction generation method 500 in accordance with the present disclosure. FIG. 5 may be described with continued reference to previous figures. The following process is exemplary and is not limited to the steps described below. Moreover, alternative embodiments may include more or fewer steps than shown or described herein, and these steps may include in an order that differs from the order described in the following example embodiments.The method 500 begins at step 502. At step 504, the method 500 may include obtaining, by the processor 208, the user input. At step 506, the method 500 may include determining, by the processor 208, the user intent based on the user input. At step 508, the method 500 may include identifying, by the processor 208, one or more vehicle components associated with the user intent by executing the instructions stored in the trained machine module 214 or the LLM agent 212.At step 510, the method 500 may include generating, by the processor 208, the instruction media content / video content based on the vehicle components and the user intent. At step 512, the method 500 may include outputting, by the processor 208, the instruction media content on the display screen 108.The method 500 may end at step 514.In the foregoing disclosure, reference has been made to the accompanying drawings, which form a part hereof, and illustrate specific implementations in which the present disclosure may be practiced. It should be understood that other implementations may be utilized and structural changes may be made without departing from the scope of the present disclosure. References in the specification to "an embodiment," "an embodiment," etc., indicate that the described embodiment may include one(s) specific feature, structure, or characteristic, but each embodiment does not necessarily have to include that(s) specific feature, structure, or characteristic. Furthermore, such formulations do not necessarily relate to the same embodiment. Further, when a feature(s) structure, or characteristic is described in connection with one embodiment, those skilled in the art will recognize such feature(s) structure, or characteristic in connection with other embodiments whether or not explicitly described.Further, the functions described herein may be performed in one or more of hardware, software, firmware, digital components, or analog components, as appropriate. For example, one or more application specific integrated circuits (ASICs) may be programmed to execute one or more of the systems and procedures described herein. Certain terms used throughout the specification and claims refer to specific system components. It will be appreciated by those skilled in the art that the components may be designated by other terms. In this document, it is not intended to distinguish between components which differ according to the designation, but not in terms of their function.Also, it should be understood that the word "example" as used herein is intended to be non-exclusive and non-limiting. In particular, the word "example" as used herein indicates one of several examples, and it is understood that no undue emphasis or preference is directed to the particular example described.A computer readable medium (also referred to as a processor readable medium) includes any non-transitory (e.g., tangible) medium that participates in providing data (e.g., instructions) that may be read by a computer (e.g., by a processor of a computer). Such a medium may take many forms, including, but not limited to, non-volatile media and volatile media. Computing devices may include computer-executable instructions, where the instructions may be executable by one or more computing devices, such as those listed above, and may be stored on a computer-readable medium.With reference to the processes, systems, methods, heuristics, etc. described herein, it should be understood that although the steps of such processes, etc. have been described as occurring according to a particular ordered sequence, such processes could be practiced with the described steps performed in an order that varies from the order described herein. Further, it should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be omitted. In other words, the descriptions of processes herein are for the purpose of illustrating various embodiments and should not be construed as limiting the claims.Accordingly, it is to be understood that the foregoing description is intended to be illustrative and not restrictive. From reading the foregoing description, many embodiments and applications other than the examples provided will be apparent. The scope should be determined, not with reference to the foregoing description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is to be understood and intended that there will be future developments in the art discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. Overall, it is to be understood that the application may be modified and varied.All terms used in the claims are intended to have their general meaning as known to those skilled in the art of technologies described herein, unless expressly stated to the contrary in this specification. In particular, the use of the singular articles such as "a", "an", "the", "the", etc. is to be understood to mean one or more of the stated elements unless a claim gives an explicit limitation to the contrary. With phrases expressing conditional contexts, such as, but not limited to, "may," "could," "may," or "may possibly," it is generally intended that certain embodiments may include certain features, elements, and / or steps, while other embodiments may not include these unless specifically stated otherwise or otherwise will be apparent from the context used. Thus, such phrases expressing conditional relationships are generally not intended to imply that features, elements, and / or steps are required in any way for one or more embodiments.In one aspect of the invention, determining the user intent further comprises determining the user intent by performing at least one of a sential analysis, an emotional analysis, or an age analysis on the user input by executing the instructions associated with the natural language processing algorithm.In one aspect of the invention, outputting the instruction media content includes displaying the instruction media content on a display screen associated with a human machine interface (MMS) or a user device.In one aspect of the invention, the method includes: retrieving a three-dimensional (3D) digital model of the vehicle interior and exterior; and displaying the instruction media content on the display screen using the digital 3D model of the vehicle interior and exterior.In one aspect of the invention, the method includes: determining an optimal gaze angle associated with the digital 3D model of the vehicle interior and exterior to display the instruction media content based on the vehicle component and the user intent; causing the display screen to rotate a default gaze angle associated with the digital 3D model of the vehicle interior and exterior displayed on the display screen to the optimal gaze angle; and causing the display screen to display the instruction media content responsive to the rotation of the default gaze angle to the optimal gaze angle.According to the present invention, there is provided a non-transitory computer readable storage medium having instructions stored thereon that, when executed by a processor, cause the processor to: obtain, from a user, a user input associated with a vehicle; determine a user intent based on the user input; identify a vehicle component associated with the user intent by executing instructions stored in a trained machine model, wherein the trained machine model is trained using training data comprising a plurality of vehicle component identifiers and a plurality of user command intents; generate instruction media content based on the vehicle component and the user intent; and output the instruction media content.

Claims

A system for outputting vehicle instructions, the system comprising: a transceiver configured to receive user input associated with a vehicle from a user; a memory configured to store a trained machine model, the trained machine model being trained using training data comprising a plurality of vehicle component identifiers and a plurality of user command intents; and a processor communicatively coupled to the transceiver and the memory, the processor configured to: obtain the user input from the transceiver; determine a user intent based on the user input; identify a vehicle component associated with the user intent by executing instructions stored in the trained machine model; generating instruction media content based on the vehicle component and the user intention; and outputting the instruction media content.The system of claim 1, wherein the user input is a natural language voice command.The system of claim 2, wherein the memory is further configured to store instructions associated with a natural language processing algorithm, and wherein the processor determines the user intent based on the user input by executing the instructions associated with the natural language processing algorithm.The system of claim 3, wherein the processor is further configured to determine the user intent by performing at least one of a sential analysis, an emotional analysis, and a user age analysis on the user input by executing the instructions associated with the natural language processing algorithm.The system of claim 1, wherein the instruction media content is video content comprising instructions to perform one or more of installing the vehicle component, operating the vehicle component, replacing the vehicle component, repairing the vehicle component, or vehicle component maintenance.The system of claim 1, wherein the processor outputs the instruction media content by displaying the instruction media content on a display screen associated with a human machine interface (MMS) of the vehicle or a user device.The system of claim 6, wherein the memory is further configured to store a three-dimensional (3D) digital model of the vehicle interior and exterior, and wherein the processor is further configured to: retrieve the digital 3D model of the vehicle interior and exterior from the memory; and display the instruction media content on the display screen using the digital 3D model of the vehicle interior and exterior.The system of claim 7, wherein the processor is further configured to: determine an optimal viewpoint associated with the digital 3D model of the vehicle interior and exterior to display the instruction media content based on the vehicle component and the user's intention; cause the display screen to rotate a default viewpoint associated with the digital 3D model of the vehicle interior and exterior displayed on the display screen to the optimal viewpoint; and cause the display screen to display the instruction media content in response to the rotation of the default viewpoint to the optimal viewpoint.The system of claim 1, wherein the memory is further configured to store information associated with a user manual of the vehicle, and wherein the processor is further configured to generate the instruction media content based on the information associated with the user manual.The system of claim 1, wherein the transceiver is further configured to receive sensor inputs from a vehicle sensor unit, and wherein the sensor inputs comprise inputs associated with at least one of a vehicle operating status, a vehicle speed, a geographic location of the vehicle, or a weather condition associated with an environment of the vehicle.The system of claim 10, wherein the processor is further configured to: obtain the sensor inputs from the transceiver; and generate the instruction media content based on the sensor inputs.The system of claim 1, wherein the system is part of the vehicle.A method for outputting vehicle instructions, the method comprising: obtaining, by a processor, user input associated with a vehicle from a user; determining, by the processor, a user intent based on the user input; identifying, by the processor, a vehicle component associated with the user intent by executing instructions stored in a trained machine model, wherein the trained machine model is trained using training data comprising a plurality of vehicle component identifiers and a plurality of user command intents; and generating, by the processor, instruction media content based on the vehicle component and the user intent; and outputting, by the processor, the instruction media content.The method of claim 13, wherein the user input is a natural language voice command.The method of claim 14, wherein determining the user intent comprises determining the user intent based on the user input by executing instructions associated with a natural language processing algorithm.

Citation Information

Cited By

  • Computer-implemented method for the generative generation and subsequent output of an explanation of a vehicle function, as well as on-board computer and vehicle

    DE102025104465A1