System and method for providing context-related vehicle guidance

Through a vehicle guidance generation system based on artificial intelligence and machine learning, user inputs are analyzed to generate personalized three-dimensional vehicle model operation guidance, solving the problem that vehicle component operation guidance in the prior art is not specific and difficult to understand, and achieving efficient vehicle operation guidance.

CN120348230APending Publication Date: 2025-07-22FORD GLOBAL TECH LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510027889.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-12
Filing Date
2025-01-08
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When users need to repair or operate vehicle components, the existing technical manuals and Internet video guidance are not specific and easy to understand enough, making it difficult for users to obtain effective help.

Method used

A vehicle guidance generation system based on artificial intelligence and machine learning is adopted to analyze the guiding media content related to the generated situation by analyzing user input, and using three-dimensional vehicle models and natural language processing technology to provide personalized operation guidance.

Benefits of technology

It improves the convenience and understanding efficiency of users' operation of vehicle components, and helps users complete repair and operation tasks efficiently through personalized guiding media content and optimal viewing display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120348230A_ABST
    Figure CN120348230A_ABST
Patent Text Reader

Abstract

The invention provides a system and method for providing context-related vehicle guidance. A system for outputting vehicle instructions is disclosed. The system may include a transceiver, a memory, and a processor. The transceiver may be configured to receive a user input associated with the vehicle from a user, and the memory may be configured to store a trained machine model. The trained machine model may be trained by using training data including a plurality of vehicle component identifiers and a plurality of user command intent. The processor may be configured to obtain a user input and determine a user intent based on the user input. The processor may also be configured to identify a vehicle component associated with the user intent by executing instructions stored in the trained machine model, and generate instructive media content based on the vehicle component and the user intent. The processor may also output instructive media content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to systems and methods for providing contextually relevant vehicle guidance to a user. Background Art

[0002] There are known instances where users seek help when they need to repair their vehicle or need guidance to operate one or more vehicle components. For example, when the windshield wipers may not be cleaning the windshield properly or when the user may not know the procedure for operating a new vehicle feature or component, the user may seek help. In such cases, the user typically looks at the owner's manual or searches the Internet for "how-to" videos to obtain the required help.

[0003] Searching the Internet for how-to videos can be a cumbersome operation and may not be effective because most such videos are generic in nature and do not provide the specific guidance that the user may be looking for. Additionally, the owner's manual is also generic in nature and may not be able to help the user in all cases. Further, the owner's manual typically includes technical jargon that may be difficult for a layperson to understand.

[0004] Accordingly, there is a need for a system that can assist vehicle users in an easy-to-understand manner. Summary of the Invention

[0005] The present disclosure describes a vehicle guidance generation system ("system") that may be configured to generate and provide to a vehicle user a guidance or educational media content that can assist the user in effectively operating, repairing, and / or maintaining one or more vehicle components associated with the vehicle. When the user desires to seek help from the system regarding one or more vehicle components, the user may provide user input to the system. As an example, when one or more vehicle components may not be operating properly or need to be replaced, when the user may seek guidance for operating a new vehicle feature / component, etc., the user may provide user input. In an exemplary aspect, the user input may be a natural language-based voice command. The system may obtain the user input from the user and may generate a guidance or educational media content (e.g., a "how-to" video) based on the user input that can assist the user in conveniently understanding the procedure for repairing, installing, replacing, and / or operating the associated vehicle components.

[0006] In some aspects, the system can be an artificial intelligence (AI) / machine learning (ML)-based system that can analyze user input and accordingly generate optimal guiding media content that can efficiently assist the user. The system can include a large language model (LLM) agent, and the LLM agent can include one or more trained machine learning modules and / or algorithms that can analyze user input and generate guiding media content.

[0007] In an exemplary aspect, in response to obtaining user input from the user, the system can determine the user intent based on the user input by performing natural language processing on the user input. The system can perform natural language processing on the user input by executing instructions stored in one or more LLM agent modules. The user intent can indicate the exact user requirements (especially when the user input is vague or implicit) and / or the reason or context that may have caused the user to provide the user input to the system. In response to determining the user intent, the system can execute the instructions stored in the LLM agent to identify one or more vehicle components associated with the user intent, and then dynamically generate guiding media content based on the identified vehicle components and the user intent. The system can also cause a display screen associated with the vehicle human-machine interface (HMI) and / or the user device to display the generated guiding media content so that the user can conveniently view the guiding media content.

[0008] In some aspects, the system can cause the display screen to display / play the guiding media content by using three-dimensional (3D) digital vehicle exterior and interior models, which can be pre-stored in the system memory. The system can be configured to determine the optimal viewing angle associated with the 3D digital vehicle exterior and interior models based on the identified vehicle components and the user intent, and cause the display screen to display the guiding media content at the optimal viewing angle so that the user can efficiently and easily comprehend the media content.

[0009] The present disclosure discloses a vehicle guidance generation system that generates context-related guiding media content or "how-to" videos based on the user's requirements and provides the context-related guiding media content or "how-to" videos to the user. By using the system, the user does not need to search for general videos on the Internet or view the owner / user manual to understand the vehicle components. The system operates by obtaining voice-based commands from the user, thus significantly enhancing the convenience of the user. In addition, the system generates guiding media content based on the specific needs and requirements of the user, and thus the generated media content is highly relevant to the user. The system also ensures that the generated media content is displayed to the user at the optimal viewing angle so that the user can efficiently and easily understand the media content.

[0010] These and other advantages of the present disclosure are provided in detail herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The detailed description is set forth with reference to the accompanying drawings. The use of the same reference numerals may indicate similar or identical items. Various embodiments may utilize elements and / or components in addition to those shown in the drawings, and some elements and / or components may not be present in various embodiments. The elements and / or components in the drawings are not necessarily drawn to scale. Throughout the present disclosure, depending on the context, singular and plural terms may be used interchangeably.

[0012] Figure 1 An example environment is depicted in which the techniques and structures for providing the systems and methods disclosed herein may be implemented.

[0013] Figure 2 A block diagram of an example vehicle guidance generation system in accordance with the present disclosure is depicted.

[0014] Figure 3 A three-dimensional (3D) digital vehicle model is depicted as displayed on a vehicle human-machine interface (HMI) in accordance with the present disclosure from a default perspective.

[0015] Figure 4 A sequence of views is depicted as displayed on a vehicle HMI in accordance with the present disclosure, the sequence of views showing the 3D digital vehicle model moving from a default perspective to an optimal perspective.

[0016] Figure 5 A flowchart of an example vehicle guidance generation method in accordance with the present disclosure is depicted. DETAILED DESCRIPTION

[0017] The present disclosure will be described more fully hereinafter with reference to the accompanying drawings, in which example embodiments of the present disclosure are shown and the example embodiments are not intended to be limiting.

[0018] Figure 1 An example environment 100 is depicted in which the techniques and structures for providing the systems and methods disclosed herein may be implemented. Environment 100 may include a user 102 sitting inside a vehicle 104. Vehicle 104 may take the form of any passenger or commercial vehicle, such as a car, work vehicle, crossover vehicle, truck, van, minivan, taxi, bus, etc. Vehicle 104 may be a manually driven vehicle, and / or may be configured to operate in a partial or fully autonomous mode, and may include any powertrain, such as a gasoline engine, one or more electric actuators, a hybrid system, etc.

[0019] In some aspects, user 102 may need assistance related to one or more vehicle components of vehicle 104. In an exemplary aspect, user 102 may need assistance when user 102 may have to install or replace a vehicle component in vehicle 104, repair a vehicle component, operate a new vehicle component that user 102 may not be familiar with (or may not know the operating procedure), perform maintenance of a vehicle component, etc. In such cases, user 102 can output a command to vehicle 104 to view one or more video contents that can provide the required assistance to user 102. The video content can be, for example, instructional or educational video content or "how-to" videos, which can provide user 102 with step-by-step guidance that user 102 may need to perform / carry out to complete a task associated with one or more vehicle components.

[0020] Vehicle 104 can be communicatively coupled to a vehicle guidance generation system 106 (or system 106), which can be configured to obtain a user command (e.g., a voice-based command or user input), and generate, based on the user input, instructional media content (e.g., "how-to" videos) associated with one or more vehicle components to be provided to user 102. In some aspects, system 106 can be a part of vehicle 104. In other aspects, system 106 can be a part of a server (not shown) and communicatively coupled to vehicle 104 via a wireless network.

[0021] System 106 can be an artificial intelligence (AI) / machine learning (ML)-based system, which is configured to dynamically generate instructional media content associated with one or more vehicle components based on user input, regardless of whether user 102 provides an explicit command associated with a vehicle component or an implicit command. For example, user 102 can provide an explicit voice command (or explicit user input) to system 106 that states "Show me how to add windshield wiper fluid". In this case, system 106 can obtain the explicit command from user 102 and generate instructional media content (e.g., a how-to video) that shows user 102 the procedure for adding windshield wiper fluid. In response to generating the instructional media content, system 106 can cause a display screen 108 associated with the vehicle human machine interface (HMI) (or a display screen associated with the user device) to display the generated instructional media content.

[0022] When user 102 provides an implicit voice command (or implicit user input), such as, "Why isn't the windshield wiper cleaning the windshield properly" (as Figure 1When in the state shown in View 110), the system 106 can first determine the user intention based on or according to the user input. In some aspects, the system 106 can include a large language model (LLM) agent, which can include one or more algorithms and / or trained machine models (trained using supervised machine learning algorithms), and the one or more algorithms and / or trained machine models can help the system 106 analyze implicit user input and determine the user intention by using natural language processing. In additional aspects, the system 106 can determine the user intention by performing one or more of sentiment analysis, emotion analysis, user age analysis, etc. on the implicit user input by executing instructions associated with one or more natural language processing algorithms of the LLM agent.

[0023] By determining the user intention, the system 106 can identify the "true goal" of the user input or one or more possible reasons that caused the user 102 to provide the implicit user input to the system 106. For example, when the user 102 provides a command as shown in View 110, the system 106 can determine that the vehicle wiper may be damaged or in need of repair, or that wiper fluid needs to be added.

[0024] In response to determining the user intention, the system 106 can identify one or more vehicle components associated with the determined user intention. Continuing with the above example, the system 106 can determine the wiper and / or wiper fluid as the vehicle components associated with the user intention. In other words, the system 106 can identify that the implicit user input may be related to the wiper and / or wiper fluid, or that the user 102 may have provided the implicit user input to the system 106 because the user 102 may be facing problems with the wiper and / or wiper fluid. In some aspects, the system 106 can identify vehicle components based on the user intention by executing instructions stored in one or more trained machine models that can be part of the LLM agent. In an exemplary aspect, the LLM agent can include a trained machine model, which can be trained by using training data including multiple user command intentions and multiple vehicle component identifiers. The trained machine model can be configured to predict or estimate vehicle component identifiers (and thus the corresponding vehicle components, such as wipers, wiper fluid, etc.) based on the user intention fed into or provided as input to the trained machine. The system 106 can use the instructions included in the trained machine model to identify the vehicle components that the user 102 may have mentioned (or faced problems with) in the implicit user input based on the determined user intention.

[0025] In some aspects, system 106 can also identify vehicle components (or confirm vehicle components identified by using the above LLM agents) based on internal vehicle data / information (e.g., measured wiper fluid level, detected electrical faults, diagnostic trouble codes or DTCs, etc.). System 106 can obtain such internal vehicle data from one or more vehicle sensors and / or from vehicle components individually.

[0026] In response to identifying a vehicle component, system 106 can dynamically generate instructional media content based on the identified vehicle component and user intent. For example, when the identified vehicle component may be associated with wiper fluid, system 106 can generate an instructional video showing a step-by-step procedure for adding wiper fluid.

[0027] As described above, system 106 also performs sentiment analysis, emotion analysis, and / or user age analysis on user input to determine user intent. In some aspects, system 106 performs emotion and / or sentiment analysis to determine whether user 102 may be calm, irritated, confused, rushed, afraid, panicked, etc., and performs user age analysis to determine the possible age-profile of user 102. System 106 can be configured to generate instructional media content based on sentiment analysis, emotion analysis, and / or user age analysis. As an example, when system 106 determines that user 102 may be an adult, system 106 can generate a more complex and detailed step-by-step instructional video, and when user 102 may be a curious child, a general video can be generated. In additional aspects, system 106 can determine user intent based on known user profiles associated with the vehicle user, especially in cases where more than one user may regularly use vehicle 104. User profiles can include known user characteristics, traits, etc., and when the user uses vehicle 104, system 106 can "learn" such user profiles associated with the vehicle user over time. In other aspects, such user profiles can be added / provided to system 106 by the respective user.

[0028] In some aspects, system 106 can generate instructional media content by executing instructions associated with one or more video generation modules / models that can be part of an LLM agent. The video generation module can generate instructional media content based on a user manual associated with vehicle 104 (which can be pre-stored in the system memory), the identified vehicle components and user intent, one or more pre-stored video clips associated with vehicle components, three-dimensional (3D) digital vehicle exterior and interior models (which can be pre-stored in the system memory), etc. As an example, when a vehicle component may be associated with wiper fluid, the video generation module can "stitch" together one or more pre-stored video clips associated with the vehicle hood, wiper fluid reservoir, etc., and dynamically generate instructional media content based on the user intent and pre-stored video clips. In some aspects, the instructional media content can use the 3D vehicle exterior and interior model (or 3D vehicle model) as a "basis" such that the media content is three-dimensional in nature and thus more immersive and helpful in assisting user 102 (e.g., to conveniently teach user 102 the procedure for adding wiper fluid).

[0029] In response to generating the instructional media content, system 106 can cause display 108 to display the instructional media content. In some aspects, system 106 can additionally determine an optimal viewing angle (or optimal "camera angle") to display the instructional media content on display 108 using the 3D vehicle model such that user 102 can easily and efficiently understand the instructional media content. System 106 can determine the optimal viewing angle based on the identified vehicle components and user intent. As an example, if the vehicle component is associated with wiper fluid, system 106 can determine that the optimal viewing angle to display the instructional media content using the 3D vehicle model may be a front top view of the vehicle (as opposed to a side view or rear view of the 3D vehicle model) such that user 102 can efficiently and clearly view the steps shown in the instructional media content and understand the procedure for adding wiper fluid.

[0030] In response to determining the optimal viewing angle, system 106 can cause display 108 to display the 3D vehicle model at the optimal viewing angle and then start to display or "play" the instructional media content. In some aspects, when the instructional media content may be displayed / played on display 108, system 106 can additionally output instructions (e.g., text instructions, picture instructions, etc.) and / or auditory cues associated with the steps included in the instructional media content. The instructions and / or auditory cues can enhance the convenience for the user to view the instructional media content and can enable user 102 to easily understand the steps being displayed on the instructional media content.

[0031] Additional system details are described below in conjunction with Figure 2 the following.

[0032] Vehicle 104 and system 106 implement and / or perform operations as described herein in this disclosure in accordance with the owner's manual and safety guidelines. Additionally, any actions taken by user 102 should comply with all rules specific to the location and operation of vehicle 104 (e.g., federal, state, country, city, etc.). Notifications, recommendations, or media content provided by vehicle 104 or system 106 should be considered as advice and followed only in accordance with any rules specific to the location and operation of vehicle 104.

[0033] Figure 2 Depicts a block diagram of vehicle guidance generation system 106 in accordance with the present disclosure. As described above in connection with Figure 1 In some examples, system 106 can be part of vehicle 104. In other aspects, system 106 can be part of a server (not shown) that can be communicatively coupled to vehicle 104 via a wireless network. The wireless network can be and / or include the Internet, a private network, a public network, or other configurations operating using any one or more known communication protocols, such as, for example, Transmission Control Protocol / Internet Protocol (TCP / IP), BLE, Wi-Fi based on Institute of Electrical and Electronics Engineers (IEEE) standard 802.11, Ultra-Wideband (UWB), and cellular technologies such as Time Division Multiple Access (TDMA), Code Division Multiple Access (CDMA), High-Speed Packet Access (HSPDA), Long Term Evolution (LTE), Global System for Mobile Communications (GSM), and Fifth Generation (5G), to name just a few examples. In yet another aspect of the present disclosure, one or more system components can be part of vehicle 104 and the remaining system components can be part of a server. In the description Figure 2 When, reference will be made to Figure 3 and Figure 4 .

[0034] Vehicle 104 can include a plurality of units / components, including but not limited to vehicle sensor unit 202, display screen 108 associated with the vehicle HMI, vehicle microphone 204, etc. Vehicle sensor unit 202 can be configured to measure / detect inputs associated with the vehicle operating state (i.e., whether the vehicle engine is on or off, whether vehicle 104 is in motion or stationary, etc.), vehicle speed, vehicle geographical location (e.g., via signals obtained from the Global Positioning System), weather conditions associated with the vehicle's surrounding environment (e.g., temperature, presence of snow, presence of rain, presence of sunlight, ambient light intensity, etc.), vehicle occupant status, etc. Vehicle sensor unit 202 can transmit the above inputs (or "sensor inputs") to system 106 at a predefined frequency.

[0035] The vehicle microphone 204 can be configured to capture voice-based commands provided to the vehicle 104 by the user 102 (e.g., implicit or explicit user input as described above in connection with Figure 1 ). The vehicle 104 can be configured to transmit (via a vehicle transceiver, not shown) the user input captured by the vehicle microphone 204 to the system 106.

[0036] Those of ordinary skill in the art will appreciate that the vehicle 104 can include Figure 2 a plurality of additional units / components not shown in and not described in this disclosure. Figure 2 The vehicle units / components depicted in

[0037] are for illustrative purposes and should not be construed as restrictive. Without departing from the scope of this disclosure, the vehicle 104 can include additional units / components.

[0038] The processor 208 can be a device that can be set to work with one or more memory devices (e.g., the memory 210 and / or Figure 2One or more external databases (not shown in the figure) for communicating with an artificial intelligence (AI) / machine learning (ML)-based processor, and the one or more memory devices are configured to communicate with a corresponding computing system. The processor 208 can utilize the memory 210 to store programs in the form of code and / or store data for performing various aspects in accordance with the present disclosure. The memory 210 can be a non-transitory computer-readable storage medium or memory that stores program code that enables the processor 208 to perform operations in accordance with the present disclosure. The memory 210 can include any one or a combination of volatile memory elements (e.g., dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), etc.), and can include any one or more non-volatile memory elements (e.g., erasable programmable read-only memory (EPROM), flash memory, electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), etc.).

[0039] In some aspects, the memory 210 can include multiple databases, modules, and agents, including but not limited to a large language module (LLM) agent 212, a trained machine module 214 (or a trained machine learning module), a video generation module 216, a natural language processing module 218, and a vehicle information database 220. The vehicle information database 220 can be configured to store information associated with the vehicle 104 (e.g., information associated with the user manual of the vehicle 104), three-dimensional (3D) digital vehicle interior and exterior models (or 3D vehicle models), names / labels of multiple vehicle components, pre-generated or pre-loaded video footage associated with multiple vehicle components, etc.

[0040] In some aspects, the trained machine module 214, the video generation module 216, and the natural language processing module 218 can be part of the LLM agent 212, as Figure 2 shown. In other aspects, one or more of the trained machine module 214, the video generation module 216, and the natural language processing module 218 can be external to the LLM agent 212. The trained machine module 214 can be trained (e.g., by using supervised machine learning techniques / algorithms) by using training data that includes multiple vehicle component identifiers (e.g., names / labels of multiple vehicle components) and multiple user command intents. In some aspects, the training data can be stored in the memory 210 or in a remote server communicatively coupled to the system 106. The trained machine module 214 can be trained by the processor 208 (by using the training data) and can be configured to output names / labels of possible vehicle components when a user intent can be input to the trained machine module 214.

[0041] The natural language processing module 218 may store instructions associated with one or more natural language processing algorithms, which may be configured to determine a user intention based on the user input (e.g., the implicit or explicit user input described above in connection with Figure 1 by performing natural language processing, sentiment analysis, emotion analysis, user age analysis, etc. on a natural language voice command-based user input). In additional or alternative aspects, as described above, the system 106 may determine the user intention based on a known user profile associated with the vehicle user, particularly in cases where more than one user may regularly use the vehicle 104. Details of the user profile have been described above in connection with Figure 1 .

[0042] The video generation module 216 may be configured to dynamically generate instructional media content (e.g., "how-to") based on the user intention, vehicle components, information associated with the user manual, and / or sensor input (obtained from the vehicle sensor unit 202). In some aspects, the video generation module 216 may use video footage and / or 3D vehicle models stored in the vehicle information database 220 to generate optimal instructional media content, which may provide situationally relevant assistance to the user 102 based on the situation the user 102 may be in or based on the user's exact requirements.

[0043] In some aspects, without departing from the scope of the present disclosure, in addition to the above modules, the LLM agent 212 may further include one or more additional modules. The LLM agent 212 may use the instructions included in the modules to correctly determine the user's intention based on the user input (i.e., based on the user's natural language voice command), generate relevant instructional media content for the user 102, and enable the display screen 108 to optimally display / output the instructional media content such that the user 102 can comprehend the media content in the most efficient and understandable manner.

[0044] In some aspects, the modules included in the LLM agent 212 may be trained by the processor 208 using supervised machine learning techniques. Those of ordinary skill in the art will understand that machine learning is an application of artificial intelligence (AI), using which a system or a processor (e.g., the processor 208) can have the ability to automatically learn and improve based on experience without being explicitly programmed. Machine learning focuses on using data and algorithms to imitate the way humans learn. In some aspects, machine learning algorithms may be created for classification and / or prediction. Machine learning-based systems may be used in various applications, including but not limited to speech recognition, image or video processing, statistical analysis, natural language processing, content generation, etc.

[0045] Based on the data or signals available to the learning system, machine learning can be of various types. For example, machine learning methods can include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. Supervised learning is a method that can be supervised by humans. In this method, a machine learning algorithm can use labeled training data and defined variables. In the case of supervised learning, both the input and output of the algorithm can be specified / defined, and the algorithm can be trained to accurately classify the data and / or predict the results.

[0046] Generally, supervised learning can be of two types: "regression" and "classification". In classification learning, the learning algorithm can help divide a dataset into multiple categories based on different parameters. In this case, a computer program can be trained on a training dataset, and based on the training, the computer program can classify the input data into different categories. Some known methods used in classification learning include logistic regression, K-nearest neighbor, support vector machine (SVM), kernel SVM, naive Bayes, decision tree classification, and random forest classification.

[0047] In regression learning, the learning algorithm can predict an output value that may have a continuous nature or real value. Some known methods used in regression learning include simple linear regression, multiple linear regression, polynomial regression, support vector regression, decision tree regression, and random forest regression.

[0048] Unsupervised learning is a method that involves algorithms that can be trained on unlabeled data. Unsupervised learning algorithms can analyze the data by themselves and find patterns in the input data. In addition, semi-supervised learning is a combination of supervised learning and unsupervised learning. Semi-supervised learning algorithms involve labeled training data; however, semi-supervised learning algorithms can still find patterns in the input data. Reinforcement learning is a multi-step or dynamic process. This model is similar to supervised learning but may not be trained using sample data. This model can "learn as it goes" by using a trial-and-error method. A sequence of successful results can be reinforced to develop the best recommendations or strategies for a given problem in reinforcement learning.

[0049] As described above, the modules of the LLM agent 212 can be trained by using supervised machine learning methods. As more training data can be fed into the processor 208, for example, when the processor 208 obtains user input provided by the user 102 over a period of time and "learns" from it, or when the processor 208 obtains updated training data from a remote server, the modules can be updated (or enhanced).

[0050] In operation, when the user 102 may need assistance associated with the vehicle 104, the user 102 can transmit / provide a natural language voice command or "user input" to the vehicle 104. As described above in connection with Figure 1As described, user 102 may provide explicit user input or implicit (or ambiguous) user input. Vehicle microphone 204 may capture the user input, and vehicle 104 may transmit the user input to system 106.

[0051] Transceiver 206 may receive the user input from vehicle 104 (i.e., from vehicle microphone 204). Additionally, transceiver 206 may receive sensor input from vehicle sensor unit 202. Transceiver 206 may transmit the received user input and sensor input to processor 208.

[0052] Processor 208 may obtain the user input and sensor input from transceiver 206. In response to obtaining the user input, processor 208 may determine the user intent based on the user input as described above in connection with Figure 1 As described above, processor 208 may additionally determine the user intent based on the user profile. In some aspects, the user intent may indicate the exact user requirements associated with vehicle 104 and / or the reason or context that may have caused user 102 to provide the user input to system 106. As an example, when user 102 provides the voice command "Why isn't the windshield wiper cleaning the windshield properly?", processor 208 may determine that the user intent may be to state that the vehicle windshield wiper may be malfunctioning or that windshield wiper fluid may need to be added.

[0053] In some aspects, processor 208 may determine the user intent based on the user input by performing natural language processing. Specifically, processor 208 may determine the user intent based on the user input by executing instructions associated with one or more natural language processing algorithms stored in natural language processing module 218. In additional aspects, processor 208 may execute instructions associated with the natural language processing algorithms stored in natural language processing module 218 to perform sentiment analysis, emotion analysis, and / or user age analysis on the user input. As described above in connection with Figure 1 As described, processor 208 may perform the above analyses to determine whether user 102 may be calm, irritated, confused, anxious, afraid, panicked, etc., and / or the possible age profile of user 102. In some aspects, processor 208 may use an additional LLM module (not shown) as well as natural language processing module 218 to determine the user intent based on the user input and / or the user profile.

[0054] In response to determining the user intent, the processor 208 may identify one or more vehicle components that may be associated with the user intent. In some aspects, the processor 208 may identify the vehicle components by executing instructions stored in the trained machine module 214. Specifically, the processor 208 may "input" the user intent into the trained machine module 214, and the trained machine module may output the expected vehicle components associated with the user intent. For example, when the user intent may indicate that the user 102 desires to add windshield wiper fluid, the processor 208 may identify one or more vehicle components (e.g., the vehicle hood, the windshield wiper fluid tank, etc.) that the user 102 may have to access to add the windshield wiper fluid. As described above in connection with Figure 1 the processor 208 may additionally determine / confirm the vehicle components based on internal vehicle data / information.

[0055] In response to identifying the vehicle components, the processor 208 may generate instructional media content (e.g., an operation method video) based on the vehicle components and the user input, and the instructional media content may efficiently teach the user 102 the procedure for performing the task (e.g., adding windshield wiper fluid) that the user 102 desires. In some aspects, the generated instructional media content may be video content that may include instructions for enabling or teaching the user 102 to perform one or more of installing a vehicle component, operating a vehicle component, replacing a vehicle component, repairing a vehicle component, and performing vehicle component maintenance (e.g., adding windshield wiper fluid to enable the vehicle windshield wipers to operate properly).

[0056] The processor 208 can generate the guidance media content by executing the instructions stored in the video generation module 216. As described above, the video generation module 216 can be configured to generate the guidance media content based on the determined user intent, the identified vehicle components, the information associated with the user manual (which the processor 208 / video generation module 216 can obtain from the vehicle information database 220), and / or the sensor input (which the video generation module 216 can obtain from the processor 208). As an example, if the user intent indicates that the user 108 must add windshield wiper fluid, the video generation module 216 can generate video content (e.g., by using the video footage and / or 3D vehicle model stored in the vehicle information database 220), and the video content teaches the user 102 the procedure of operating one or more vehicle components to conveniently add windshield wiper fluid. In an exemplary aspect, the video content can teach the user 102 the procedures of opening the vehicle hood, approaching the windshield wiper fluid tank, filling the windshield wiper fluid, closing / fixing the windshield wiper fluid tank and the vehicle hood, etc. As another example, if the sensor input indicates that the weather condition around the vehicle may be dark and cloudy, the video generation module 216 may generate video content with low brightness compared to when the weather condition may indicate sunny or there is sufficient natural or artificial light in the vehicle's surrounding environment. Similarly, the video generation module 216 can generate video content with multiple auditory or audio cues when the sensor input indicates that the vehicle 104 may be moving, and can include fewer counted audio cues when the vehicle 104 may be stationary. As yet another example, compared to a curious child (in which case the video generation module 216 can generate general video content), when the user intent indicates that an adult may be providing user input, the video generation module 216 can generate detailed step-by-step guidance video content. In some aspects, the generated video content can be based on the information included in the user manual or the owner's manual, such that the procedures included in or taught by the video content can be synchronized with or verified by the user manual or the owner's manual.

[0057] In response to generating the guidance media content / video content as described above, the processor 208 can cause the display screen 108 to output / display the guidance media content. In some aspects, in order to cause the display screen 108 to output / display the guidance media content, the processor 208 can first extract the 3D vehicle model from the vehicle information database 220, and cause the display screen 108 to output / display the guidance media content by using the 3D vehicle model, as described below.

[0058] In an exemplary aspect, the processor 208 can first cause the default view of the 3D vehicle model to be displayed on the display screen 108. In some aspects, the default view on the display screen 108 can be the top-down perspective of the 3D vehicle model, asFigure 3 As shown. Specifically, as Figure 3 shown, when, for example, the generated guidance media content / video content may not be displayed or played on the display screen 108, a top-down view 302 of the 3D vehicle model can be displayed on the display screen 108. Those of ordinary skill in the art can understand that the default view on the display screen 108 can also be any other view of the 3D vehicle model different from the top-down view 302, such as, a side view, a rear view, etc. Hereinafter, the top-down view 302 is referred to as the default view 302 in the present disclosure.

[0059] In some aspects, before causing the display screen 108 to output / display or play the generated instruction media content, the processor 208 can determine the best view for displaying the guidance media content associated with the 3D vehicle model based on the identified vehicle components and the determined user intent. The best view can be the view that the user 102 may find most convenient for understanding or comprehending the steps shown in the guidance media content. As an example, when the guidance media content may be associated with a procedure for adding wiper fluid, the processor 208 can determine the best view as a front top view, a front view, or a front isometric view, in which the user can conveniently "see" the steps to be performed on the 3D vehicle model. In some aspects, the processor 208 can determine the best view associated with the 3D vehicle model by executing instructions stored in one or more modules of the LLM agent 212.

[0060] In response to determining the best view, the processor 208 can cause the view of the 3D vehicle model that may be displayed on the display screen 108 to rotate / change from the default view 302 to the best view, as Figure 4 shown. Specifically, as Figure 4 shown, the processor 208 can cause the display screen 108 to change / rotate the view associated with the 3D vehicle model from the default view 302 to the best view 402 in a series or sequence of multiple steps (shown as views 404 and 406 in Figure 4 ).

[0061] Then, the processor 208 may cause the display screen 108 to start outputting / displaying / playing the directive media content in response to changing / rotating the default perspective 302 to the optimal perspective 402. In some aspects, the processor 208 may additionally generate step-by-step picture guidance 408 (e.g., by executing instructions stored in the LLM agent 212) and cause the display screen 108 to display the step-by-step picture guidance, which may enable the user 102 to conveniently learn the steps required to perform the task that the user 102 desires to perform (e.g., adding windshield wiper fluid). In some aspects, the task may require specific supplies, which may or may not be available in the vehicle 104. If the supplies are available in the vehicle 104 (e.g., a tire change kit, a portable tire inflator, etc.), the guidance 408 may direct the user 102 to the location of those required supplies. On the other hand, if the supplies are not available in the vehicle 104, the vehicle 104 may query the server to identify a way to obtain the required supplies, e.g., directing the user to a nearby retail location where the required supplies can be purchased.

[0062] The processor 208 may additionally generate audio commands / hints associated with the steps shown on the display screen 108 and cause the vehicle speaker (not shown) to output the audio commands / hints, so that the user 102 can easily understand / apprehend the steps. An example audio hint stating "You should add windshield wiper fluid. Let me show you how to add it" is depicted as hint 410 in Figure 4 . In an exemplary aspect, when the view on the display screen 108 may be changing / rotating from the default perspective 302 to the optimal perspective 402 (or at any time before or after the view change / rotation), the processor 208 may cause the vehicle speaker to output hint 410. Additionally, the processor 208 may output hint 410 in response to a user input (which may be, for example, "Why isn't the windshield wiper cleaning the windshield properly?", as described above in conjunction with Figure 1 . Those of ordinary skill in the art will understand that the audio hints, views, etc. described in this disclosure and depicted in the drawings are for illustrative purposes only and should not be construed as limiting.

[0063] User 102 can conveniently view the instructional media content on the display screen 108 and can accordingly perform tasks or learn to perform tasks that user 102 desires to perform. In this way, system 106 enables user 102 to conveniently learn about the different features, components, and / or procedures associated with vehicle 104 without having to manually access a user manual or search the Internet for assistance. User 102 can additionally pinch and zoom the display screen 108 to magnify or reduce the view of the instructional media content that may be displayed on the display screen 108 (e.g., to focus on a specific vehicle component). User 102 can additionally magnify or reduce the view based on voice-based commands.

[0064] In some aspects, once the processor 208 generates the instructional media content, the processor 208 can additionally transmit the instructional media content to a remote server or cloud via the transceiver 206 so that other vehicles / systems that may need similar instructional media content can obtain (and use) the media content generated by the processor 208 from the remote server or cloud.

[0065] Figure 5 A flowchart depicting an example vehicle guidance generation method 500 according to the present disclosure is shown. The following may continue to be described with reference to the previous figures. Figure 5 The following process is exemplary and is not limited to the steps described below. Additionally, alternative embodiments may include more or fewer steps than those shown or described herein and may include those steps in a different order than the order described in the following example embodiments.

[0066] Method 500 begins at step 502. At step 504, method 500 may include obtaining user input by processor 208. At step 506, method 500 may include determining the user intent by processor 208 based on the user input. At step 508, method 500 may include identifying one or more vehicle components associated with the user intent by processor 208 by executing instructions stored in the trained machine module 214 or the LLM agent 212.

[0067] At step 510, method 500 may include generating instructional media content / video content by processor 208 based on the vehicle components and the user intent. At step 512, method 500 may include outputting the instructional media content on the display screen 108 by processor 208.

[0068] Method 500 may end at step 514.

[0069] In the foregoing disclosure, reference has been made to the accompanying drawings, which form a part of the foregoing disclosure, and which illustrate specific embodiments in which the disclosure may be practiced. It is to be understood that other embodiments may be utilized and structural changes may be made without departing from the scope of the disclosure. References in this specification to "one embodiment," "an embodiment," "an example embodiment," etc., indicate that the embodiment described may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a feature, structure, or characteristic is described in connection with an embodiment, whether or not explicitly described, those skilled in the art will recognize such feature, structure, or characteristic in connection with other embodiments.

[0070] In addition, where appropriate, the functions described herein may be performed in one or more of the following: hardware, software, firmware, digital components, or analog components. For example, one or more application specific integrated circuits (ASICs) may be programmed to perform one or more of the systems and programs described herein. Throughout the specification and claims, certain terms are used to refer to particular system components. As those skilled in the art will appreciate, components may be referred to by different names. This document is not intended to distinguish between components that differ in name but not in function.

[0071] It should also be understood that the word "example" as used herein is intended to be non-exclusive and non-restrictive in nature. More specifically, the word "example" as used herein indicates one of a number of examples, and it should be understood that no undue emphasis or preference is given to the particular example described.

[0072] A computer-readable medium (also referred to as a processor-readable medium) includes any non-transitory (e.g., tangible) medium that participates in providing data (e.g., instructions) that can be read by a computer (e.g., by a processor of a computer). Such a medium may take many forms, including but not limited to non-volatile media and volatile media. A computing device may include computer-executable instructions, where the instructions may be executable by one or more computing devices such as those listed above and stored on a computer-readable medium.

[0073] Regarding the processes, systems, methods, heuristics, etc. described herein, it should be understood that although the steps of such processes etc. have been described as occurring according to a certain ordered sequence, such processes may be practiced with the described steps performed in an order different from that described herein. It should also be understood that certain steps may be performed simultaneously, other steps may be added, or certain steps described herein may be omitted. In other words, the description of the processes herein is provided for purposes of illustrating various embodiments and should in no way be construed as limiting the claims.

[0074] Accordingly, it should be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided will be apparent upon reading the above description. The scope should not be determined with reference to the above description, but should be determined with reference to the appended claims and the entire scope of equivalents to which such claims are entitled. It is anticipated and expected that the technology discussed herein will evolve in the future, and the disclosed systems and methods will be incorporated into such future embodiments. In summary, it should be understood that this application is capable of modification and change.

[0075] Unless expressly stated to the contrary herein, all terms used in the claims are intended to be given their ordinary meaning as understood by one of ordinary skill in the art as described herein. Specifically, unless the claims recite a clear limitation to the contrary, the use of the singular articles such as "a," "the," and "said" should be construed to recite one or more of the indicated elements. Conditional language, such as, but not limited to, "can," "could," "might," or "may," is generally intended to convey that certain embodiments can include certain features, elements, and / or steps, while other embodiments may not include certain features, elements, and / or steps. Thus, such conditional language is not generally intended to imply that one or more embodiments necessarily require each feature, element, and / or step.

[0076] In one aspect of the present invention, determining user intent further includes determining user intent by performing one or more of sentiment analysis, emotion analysis, or age analysis on user input by executing instructions associated with natural language processing algorithms.

[0077] In one aspect of the present invention, outputting guiding media content includes displaying guiding media content on a display screen associated with a vehicle human-machine interface (HMI) or a user device.

[0078] In one aspect of the present invention, the method includes: extracting a three-dimensional (3D) digital vehicle interior and exterior model; and displaying guiding media content on a display screen by using the 3D digital vehicle interior and exterior model.

[0079] In one aspect of the present invention, the method includes: determining an optimal viewing angle for displaying guiding media content associated with the 3D digital vehicle interior and exterior model based on vehicle components and user intent; causing the display screen to rotate a default viewing angle associated with the 3D digital vehicle interior and exterior model displayed on the display screen to the optimal viewing angle; and causing the display screen to display guiding media content in response to rotating the default viewing angle to the optimal viewing angle.

[0080] According to the present invention, there is provided a non-transitory computer-readable storage medium having instructions stored thereon, which when executed by a processor cause the processor to: obtain user input associated with a vehicle from a user; determine a user intention based on the user input; identify a vehicle component associated with the user intention by executing instructions stored in a trained machine model, wherein the trained machine model is trained using training data including a plurality of vehicle component identifiers and a plurality of user command intentions; generate guiding media content based on the vehicle component and the user intention; and output the guiding media content.

Claims

1. A system for outputting vehicle instructions, the system comprising: a transceiver configured to receive a user input associated with a vehicle from a user; a memory configured to store a trained machine model, wherein the trained machine model is trained using training data including a plurality of vehicle component identifiers and a plurality of user command intents; and a processor communicatively coupled to the transceiver and the memory, wherein the processor is configured to: obtain the user input from the transceiver; determine a user intent based on the user input; identify a vehicle component associated with the user intent by executing instructions stored in the trained machine model; generate a guiding media content based on the vehicle component and the user intent; and output the guiding media content.

2. The system according to claim 1, wherein the user input is a natural language voice command.

3. The system according to claim 2, wherein the memory is further configured to store instructions associated with a natural language processing algorithm, and wherein the processor determines the user intent based on the user input by executing the instructions associated with the natural language processing algorithm.

4. The system according to claim 3, wherein the processor is further configured to determine the user intent by performing one or more of sentiment analysis, emotion analysis, and age analysis on the user input by executing the instructions associated with the natural language processing algorithm.

5. The system according to claim 1, wherein the guiding media content is video content including instructions for performing one or more of installing the vehicle component, operating the vehicle component, replacing the vehicle component, repairing the vehicle component, or performing vehicle component maintenance.

6. The system according to claim 1, wherein the processor outputs the guiding media content by displaying the guiding media content on a display screen associated with a vehicle human machine interface (HMI) or a user device.

7. The system according to claim 6, wherein the memory is further configured to store three-dimensional (3D) digital vehicle interior and exterior models, and wherein the processor is further configured to: extract the 3D digital vehicle interior and exterior models from the memory; and display the guiding media content on the display screen by using the 3D digital vehicle interior and exterior models.

8. The system according to claim 7, wherein the processor is further configured to: determine an optimal viewing angle for displaying the guiding media content associated with the 3D digital vehicle interior and exterior models based on the vehicle component and the user intent; cause the display screen to rotate a default viewing angle associated with the 3D digital vehicle interior and exterior models displayed on the display screen to the optimal viewing angle; and cause the display screen to display the guiding media content in response to rotating the default viewing angle to the optimal viewing angle.

9. The system according to claim 1, wherein the memory is further configured to store information associated with a user manual of the vehicle, and wherein the processor is further configured to generate the guidance media content based on the information associated with the user manual.

10. The system according to claim 1, wherein the transceiver is further configured to receive sensor inputs from a vehicle sensor unit, and wherein the sensor inputs include inputs associated with at least one of a vehicle operating state, a vehicle speed, a vehicle geographical location, or weather conditions associated with the vehicle surroundings.

11. The system according to claim 10, wherein the processor is further configured to: obtain the sensor inputs from the transceiver; and generate the guidance media content based on the sensor inputs.

12. The system according to claim 1, wherein the system is part of the vehicle.

13. A method for outputting vehicle instructions, the method comprising: obtaining, by a processor, user input associated with a vehicle from a user; determining, by the processor, a user intent based on the user input; identifying, by the processor, a vehicle component associated with the user intent by executing instructions stored in a trained machine model, wherein the trained machine model is trained using training data including a plurality of vehicle component identifiers and a plurality of user command intents; generating, by the processor, guidance media content based on the vehicle component and the user intent; and outputting, by the processor, the guidance media content.

14. The method according to claim 13, wherein the user input is a natural language voice command.

15. The method according to claim 14, wherein determining the user intent includes determining the user intent based on the user input by executing instructions associated with a natural language processing algorithm.