A voice prompting method and vehicle

CN122551467APending Publication Date: 2026-08-11GREAT WALL MOTOR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-05
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]本申请提供了一种语音提示方法及车辆,目的在于当驾驶员注意力分散或视觉条件受限时,通过解析控制器局域网报文生成适配故障的语音提示,并结合用户驾驶行为自适应调整播放策略,以克服现有仪表仅依赖视觉提示难以有效传达报警信息的缺陷

Benefits of technology

[0047]本申请提供的技术方案,监听车辆的控制器局域网报文;当发生指示灯报警事件时,基于控制器局域网报文,确定报警指示灯的诊断故障码;基于语义映射数据库,确定与诊断故障码对应的故障语义描述和故障风险等级;基于故障语义描述和故障风险等级,生成待播报语音;基于实时检测的用户驾驶行为,确定待播报语音的播放策略;根据播放策略,执行待播报语音的播放。本申请通过监听控制器局域网报文直接确定诊断故障码,经语义映射转换为故障描述与风险等级,并结合实时驾驶行为自适应播报,在驾驶员分心时强化提醒、高危时延迟播报,克服了纯视觉报警的局限,降低了故障理解门槛,提升了响应时效与驾驶安全性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551467A_ABST
    Figure CN122551467A_ABST
Patent Text Reader

Abstract

This application discloses a voice prompt method and vehicle, belonging to the field of intelligent cockpit technology. The method involves: monitoring controller local area network (Controller Area Network) messages; when an indicator light alarm event occurs, determining the diagnostic fault code of the alarm indicator light based on the Controller Area Network messages; determining the fault semantic description and fault risk level corresponding to the diagnostic fault code based on a semantic mapping database; generating a voice message to be played based on the fault semantic description and fault risk level; determining the playback strategy of the voice message to be played based on real-time detected user driving behavior; and executing the playback of the voice message to be played according to the playback strategy. This method directly determines the diagnostic fault code by monitoring Controller Area Network messages, converts it into a fault description and risk level through semantic mapping, and adaptively broadcasts it based on real-time driving behavior. It strengthens reminders when the driver is distracted and delays broadcasts during high-risk situations, overcoming the limitations of purely visual alarms, lowering the threshold for fault understanding, and improving response time and driving safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart cockpit technology, and in particular to a voice prompt method and vehicle. Background Technology

[0002] In modern intelligent vehicles, instrument panel indicator lights are the core visual indication system for vehicle malfunctions. However, existing vehicles rely solely on visual cues when these indicator lights illuminate, lacking accompanying voice guidance. When the driver's attention is diverted (e.g., operating the central control system, observing road conditions) or visual conditions are limited (e.g., at night, in rainy or foggy weather), it is difficult to promptly detect or accurately understand the meaning of the warning indicator lights, leading to delays in malfunction handling and posing safety hazards. Summary of the Invention

[0003] This application provides a voice prompt method and vehicle, the purpose of which is to generate a voice prompt adapted to the fault by parsing the controller local area network message when the driver's attention is distracted or the visual conditions are limited, and to adaptively adjust the playback strategy in combination with the user's driving behavior, so as to overcome the shortcomings of existing instruments that rely solely on visual prompts and cannot effectively convey alarm information.

[0004] To achieve the above objectives, this application provides the following technical solution:

[0005] A voice prompt method, comprising:

[0006] Monitor the vehicle's controller area network (LAN) messages;

[0007] When an indicator light alarm event occurs, the diagnostic fault code of the alarm indicator light is determined based on the controller LAN message;

[0008] Based on the semantic mapping database, determine the fault semantic description and fault risk level corresponding to the diagnostic fault code;

[0009] Based on the fault semantic description and the fault risk level, generate the voice to be broadcast;

[0010] Based on real-time detection of user driving behavior, the playback strategy for the voice message to be played is determined.

[0011] According to the playback strategy, the audio to be played is executed.

[0012] Optionally, based on the fault semantic description and the fault risk level, a voice message to be broadcast is generated, including:

[0013] Obtain the current operating status of the vehicle;

[0014] Based on the fault semantic description, fault handling suggestions, and the current operating status, personalized suggestions are generated through a preset large language model; the fault handling suggestions are pre-associated with the diagnostic fault codes and stored in the semantic mapping database;

[0015] Based on the aforementioned fault risk level, determine the corresponding voice playback configuration;

[0016] Based on the personalized suggestions and the voice playback configuration, a voice message to be played is generated.

[0017] Optionally, based on real-time detected user driving behavior, the playback strategy for the voice message to be played is determined, including:

[0018] Obtain multimodal information related to user driving behavior; the multimodal information includes driving behavior data collected in real time by different on-board sensors;

[0019] Based on the multimodal information, the corresponding user driving behavior is determined;

[0020] When the user's driving behavior is distracted, a playback strategy for the voice message to be played is determined based on a first preset strategy; the first preset strategy is to increase the playback volume and repeat the playback twice.

[0021] When the user's driving behavior is high-risk, a playback strategy for the voice message to be played is determined based on a second preset strategy; the second preset strategy is: delay the playback and reduce the playback volume.

[0022] When the user's driving behavior is in a low-risk driving situation, the playback strategy for the voice message to be played is determined based on a third preset strategy; the third preset strategy is to maintain the default volume for playback.

[0023] Optionally, based on the multimodal information, determining the corresponding user driving behavior includes:

[0024] Based on the multimodal information, the corresponding user driving behavior is obtained through a driving behavior analysis model. The construction process of the driving behavior analysis model includes: constructing training samples based on pre-collected sample multimodal information and corresponding sample user driving behavior; and using the training samples to train an initial multimodal large model to obtain the driving behavior analysis model.

[0025] Optionally, after playing the speech to be broadcast according to the playback strategy, the method further includes:

[0026] Collect the user's voice response to the voice to be played;

[0027] Based on the user's voice response, the corresponding reply voice is obtained through the intelligent voice assistant pre-deployed in the vehicle;

[0028] Play the response voice message.

[0029] Optionally, after playing the speech to be broadcast according to the playback strategy, the method further includes:

[0030] Based on the semantic description of the fault, a graphic and textual fault report is generated;

[0031] Based on the vehicle's current location information, repair shop recommendations are obtained through the in-vehicle navigation map;

[0032] The graphic fault report and the recommended repair shop information are sent to a mobile terminal that has been pre-connected to the vehicle.

[0033] Optionally, according to the playback strategy, the playback of the voice to be broadcast is performed, including:

[0034] Determine whether the broadcast priority of the voice message to be broadcast meets the standard; wherein, the broadcast priority is determined based on the fault risk level;

[0035] If the broadcast priority is met, the audio task currently being executed by the vehicle multimedia system is interrupted, and the playback of the voice to be broadcast is executed first according to the playback strategy before the playback of the audio task continues.

[0036] If the broadcast priority is not met, the playback of the audio task will be performed simultaneously with the playback of the voice to be broadcast, according to the playback strategy.

[0037] Optionally, the voice playback configuration includes at least acoustic features and playback priority;

[0038] Based on the personalized suggestions and the voice playback configuration, a voice message to be played is generated, including:

[0039] The personalized suggestion content is used as the text to be broadcast, and the acoustic feature parameters of the text to be broadcast are labeled according to the acoustic features.

[0040] A text-to-speech engine is used to synthesize the speech to be played according to the acoustic feature parameters, and a tag corresponding to the playback priority is set for the speech to be played.

[0041] Optionally, based on the controller local area network (LAN) messages, the diagnostic fault codes of the alarm indicator lights are determined, including:

[0042] Obtain a pre-stored Controller Area Network (CAN) database file; the CAN database file is used to define the physical meaning of each data bit in the CAN message;

[0043] Based on the controller local area network database file, the controller local area network packets are parsed bit by bit to obtain the bit-by-bit parsing results;

[0044] Based on the bit-by-bit parsing results, the diagnostic fault codes corresponding to the alarm indicator lights are extracted.

[0045] A vehicle includes: a processor, a memory, and a bus; the processor and the memory are connected via the bus.

[0046] The memory is used to store a program, and the processor is used to run the program, wherein the program is executed by the processor to perform the voice prompt method.

[0047] The technical solution provided in this application monitors the vehicle's controller area network (CLAN) messages. When an indicator light alarm event occurs, it determines the diagnostic fault code of the alarm indicator light based on the CLAN messages. Based on a semantic mapping database, it determines the corresponding fault semantic description and fault risk level. Based on the fault semantic description and fault risk level, it generates a voice message to be played. Based on real-time detected user driving behavior, it determines the playback strategy for the voice message to be played. According to the playback strategy, it executes the playback of the voice message to be played. This application directly determines the diagnostic fault code by monitoring the CLAN messages, converts it into a fault description and risk level through semantic mapping, and combines it with adaptive playback based on real-time driving behavior. It strengthens reminders when the driver is distracted and delays playback during high-risk situations, overcoming the limitations of purely visual alarms, lowering the threshold for fault understanding, and improving response time and driving safety. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 This is a first flowchart illustrating a voice prompt method provided in an embodiment of this application;

[0050] Figure 2 This is a second flowchart illustrating a voice prompt method provided in an embodiment of this application;

[0051] Figure 3 A third flowchart illustrating a voice prompt method provided in an embodiment of this application;

[0052] Figure 4 A fourth flowchart illustrating a voice prompt method provided in an embodiment of this application;

[0053] Figure 5 A fifth flowchart illustrating a voice prompt method provided in an embodiment of this application;

[0054] Figure 6 A sixth flowchart illustrating a voice prompt method provided in an embodiment of this application;

[0055] Figure 7 A seventh flowchart illustrating a voice prompt method provided in an embodiment of this application;

[0056] Figure 8 The eighth flowchart of a voice prompt method provided in an embodiment of this application;

[0057] Figure 9 This is a schematic diagram of the architecture of a vehicle provided in an embodiment of this application. Detailed Implementation

[0058] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0059] In this application, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0060] Example 1

[0061] like Figure 1The diagram shown is a first flowchart of a voice prompt method provided in an embodiment of this application. It can be applied to the vehicle's onboard computing platform (including but not limited to devices such as intelligent cockpit domain controllers, in-vehicle infotainment systems, and autonomous driving domain controllers), and includes the following steps.

[0062] S101: Monitors the vehicle's controller area network messages.

[0063] Among them, the indicator light drive signals and controller LAN messages generated by the IP (Instrument Panel) can be monitored in real time through the OBD-II (On-Board Diagnostics II, the second generation standard of vehicle diagnostic system) interface or vehicle network gateway.

[0064] In some examples, the OBD-II interface is a standardized diagnostic interface for vehicles, enabling direct access to vehicle internal bus data. When the instrument cluster control unit sends a controller area network (CAN) message related to an indicator light alarm to the bus, the onboard computing platform, as a node on the bus, can directly acquire that message.

[0065] S102: When an indicator light alarm event occurs, determine the diagnostic fault code of the alarm indicator light based on the controller LAN message.

[0066] When a warning light alarm occurs in the vehicle's instrument cluster control unit, the control unit generates a corresponding Controller Area Network (CLAN) message. The onboard computing platform can capture the CLAN message of the indicator light drive signal within a very short time (e.g., within 100ms) of the indicator light alarm illuminating.

[0067] In some examples, the DTC (Diagnostic Trouble Code) of the alarm indicator can be obtained by parsing the captured Controller Area Network (CAN) packets bit by bit according to the definition of different message IDs and data bits in the CAN database file through a CAN protocol parsing module pre-deployed on the vehicle computing platform.

[0068] Optionally, the process by which the Controller Area Network (CAN) protocol parsing module parses CAN packets to obtain the diagnostic fault codes of the alarm indicator lights can be summarized as follows: Figure 6 The steps are shown.

[0069] S103: Based on the semantic mapping database, determine the fault semantic description and fault risk level corresponding to the diagnostic fault code.

[0070] The semantic mapping database includes various diagnostic fault codes for the vehicle, along with their corresponding semantic descriptions and risk levels. Essentially, this database is a lookup table whose core function is to translate machine code into semantic information that the driver can understand.

[0071] In some examples, the types of diagnostic fault codes include, but are not limited to, OBD-II fault codes, the expressions of which can be seen as P0171 (fuel system too lean) and C1234 (wheel speed sensor signal abnormality).

[0072] For example, the semantic description of the diagnostic fault code P0171 could be "engine fuel system too lean", and the fault risk level could be "warning". The fault risk level can be classified as critical (e.g., brake system failure, requiring immediate stop), warning (e.g., engine performance degradation, requiring prompt repair), or reminder (e.g., maintenance reminder, routine maintenance is sufficient).

[0073] In some examples, the fault semantic description includes natural language used to describe the fault indicated by the diagnostic fault code; for example, the fault semantic description could be “transmitter fuel system too lean”.

[0074] In some examples, the fault risk level includes the risk level of the fault indicated by the diagnostic fault code. For example, the fault risk level can be classified as critical (e.g., brake system failure, high-voltage battery thermal runaway, requiring immediate stop), warning (e.g., degraded engine system performance, low tire pressure, requiring prompt repair), and alert (e.g., maintenance reminder, low windshield washer fluid level, requiring routine maintenance).

[0075] In a possible implementation, the semantic mapping database may include physical status codes of various indicator lights and their diagnostic fault codes, as well as semantic tags for each diagnostic fault code (i.e., fault semantic description, fault handling suggestions, and fault risk level), and supports OTA (Over The Air) updates and maintenance to adapt to the continuous emergence of new models and new fault codes.

[0076] It should be noted that the physical status code is the value of a specific bit in the message data field, which directly corresponds to the on / off state of a certain indicator light on the dashboard.

[0077] In some examples, the OTA update and maintenance process can be as follows: The vehicle receives an update package for the semantic mapping database from a cloud server. This update package can be an incremental or full update package. The vehicle establishes a connection with the cloud server via its in-vehicle communication module (T-Box) and receives the update package periodically or upon receiving a push notification. The update package contains newly added or modified indicator light physical status codes, diagnostic fault codes, and their corresponding semantic information. Based on the update package, the semantic mapping database is upgraded online to update the physical status code of at least one indicator light and its corresponding diagnostic fault code, fault semantic description, fault handling suggestions, and fault risk level. After receiving the update package and completing integrity verification, the in-vehicle computing platform uses the update package to perform an online upgrade operation on the locally stored semantic mapping database. The entire process requires no user intervention and does not require the vehicle to be taken to a 4S shop for offline flashing.

[0078] S104: Generate the voice to be broadcast based on the semantic description of the fault and the fault risk level.

[0079] The generation of the speech to be broadcast is the process of transforming structured text information into natural, fluent, and expressive speech for playback through the vehicle's speaker system. The core of this process lies in converting abstract textual semantics into auditoryally perceptible speech signals while preserving necessary tone, emotion, and urgency information.

[0080] In some examples, the semantic description of the fault can be directly substituted into a preset voice template to generate the text to be broadcast, which is then synthesized into speech using a TTS (Text-to-Speech) engine. To further enhance personalization, more accurate suggestions can be generated by combining dynamic information such as the vehicle's current operating status.

[0081] Optionally, the process of generating the voice to be broadcast based on the fault semantic description and fault risk level can be found in [reference needed]. Figure 2 The steps are shown.

[0082] S105: Determine the playback strategy for the voice message to be broadcast based on real-time detected user driving behavior.

[0083] The playback strategy refers to a set of dynamic rules governing the playback parameters of the voice prompts (such as volume, timing, repetition count, and interruption). By monitoring the user's driving behavior in real time, the system can assess the driver's current cognitive load and attention allocation, thereby adaptively adjusting the playback method of the voice prompts to ensure that fault information is effectively delivered to the user without interfering with driving safety. The core objective of this strategy is to achieve context-specific intelligent broadcasting, avoiding additional cognitive burden when the driver is under high load or performing high-risk operations, and preventing missed critical alarms due to insufficient broadcast intensity when the driver is distracted.

[0084] In some examples, user driving behavior can be detected based on real-time collected driving behavior data from different modalities, and a playback strategy for the announced audio can be determined accordingly. This playback strategy is dynamic and designed to ensure that fault information is effectively delivered to the driver while maintaining safety. It can also mitigate the negative impact of announced audio on the user, such as avoiding increased cognitive load during complex lane-changing maneuvers.

[0085] In some examples, user driving behavior can be detected and classified in real time using onboard sensor networks. Common driving behavior categories include, but are not limited to: distracted driving (such as the driver's gaze being off the road for extended periods, operating the central control screen, or talking to passengers), high-risk driving (such as rapid acceleration, sudden deceleration, sharp turns, driving in complex traffic flow or inclement weather, or driver agitation), and low-risk driving (such as driving at a constant speed in a straight line on a clear road, with the driver attentive and showing no obvious signs of stress or fatigue). Different categories of driving behavior correspond to different risk levels and cognitive load levels, thus requiring differentiated playback strategies.

[0086] Optionally, the implementation process of determining the playback strategy for the voice message based on real-time detected user driving behavior can be found in [reference needed]. Figure 3 The steps are shown.

[0087] S106: Play the audio to be played according to the playback strategy.

[0088] The in-vehicle computing platform, based on the playback strategy determined in the previous step, will invoke the audio processing unit and the in-vehicle speaker system to play the voice message to be broadcast. This adaptive playback mechanism, which incorporates user driving behavior perception, enables adaptive playback of the voice message, reducing the negative impact of intelligent voice prompts on user driving behavior and preventing them from becoming a new source of interference in high-risk driving scenarios.

[0089] In some examples, the vehicle's multimedia system is currently performing audio tasks, such as in-vehicle navigation or playing in-vehicle music. Before playing the voice message to be broadcast, it is necessary to consider coordinating the playback order between the audio task and the voice message to be broadcast in order to avoid conflicts that could lead to key information being covered or a decline in user experience.

[0090] Optionally, the implementation process for playing the audio to be broadcast, based on the playback strategy, can be found in [reference needed]. Figure 4 The steps are shown.

[0091] In some examples, after hearing the personalized suggestions displayed in the voice prompt, users may have questions, such as not understanding a certain technical term or wanting to know more detailed handling steps. As a result, users may involuntarily ask how to carry out subsequent fault maintenance operations. Therefore, it is also necessary to provide intelligent Q&A based on the user's voice response to improve the user experience.

[0092] Optionally, the implementation process for intelligent Q&A based on user voice responses can be found in [reference needed]. Figure 5 The steps are shown.

[0093] In some examples, mobile terminals (such as smartphones) establish a remote connection with vehicles, allowing users to manage vehicles through their mobile terminals. To improve vehicle maintenance efficiency, corresponding graphic and textual fault reports can be generated based on the semantic description of the fault, and these reports can be sent to the mobile terminals so that users can apply for after-sales service on their mobile terminals, forming a complete closed loop of "voice prompts → mobile terminal follow-up".

[0094] Optionally, based on the semantic description of the fault, a graphic and text-based fault report can be provided to the mobile terminal to enable rapid appointment of vehicle maintenance. For details on this process, please refer to [link / reference needed]. Figure 7 The steps are shown.

[0095] Through field performance testing, the voice prompt method shown in the embodiments of this application can achieve the following beneficial effects: (1) In terms of fault response time, the average time of traditional manual query is shortened from 47 seconds to within 2-3 seconds of voice prompt; (2) In terms of misjudgment rate, due to the clear guidance of intelligent voice, the misreading rate of the alarm indicator light by users is reduced; (3) In terms of safety, due to the emergency fault voice interruption mechanism, the accident rate caused by ignoring the alarm indicator light during high-speed driving is reduced, and the driver's distraction time is reduced, significantly reducing the driving risk caused by fault handling; (4) In terms of user experience, most users reported that "voice prompts are more intuitive and reassuring than alarm indicator lights", especially in elderly drivers and night driving scenarios; (5) In terms of after-sales efficiency, the personalized suggestions in the voice broadcast include diagnostic fault codes, and maintenance personnel can directly locate the problem quickly based on the voice recording, reducing the diagnosis time; (6) Reduce unnecessary roadside assistance calls and reduce user maintenance costs.

[0096] The processes described in S101-S106 firstly, by monitoring the controller's local area network messages and directly determining the diagnostic fault code when an indicator light alarm event occurs, alarm information can be captured without relying on visual perception. This overcomes the shortcomings of traditional instruments that rely solely on visual cues, making it difficult for drivers to detect alarms in a timely manner when their attention is distracted or their vision is limited. Secondly, based on a semantic mapping database, the diagnostic fault code is converted into a fault semantic description and fault risk level, lowering the threshold for drivers to understand the fault and solving the problem of delayed response due to misreading or misunderstanding of the alarm's meaning. Furthermore, the voice playback strategy is adaptively determined based on real-time detected user driving behavior, enhancing the reminder when the driver is distracted and delaying the broadcast during high-risk driving to avoid interference. This ensures effective transmission of fault information while avoiding negative impacts on driving safety. Finally, by executing the adaptively played voice to be broadcast, the synergy between visual and auditory dual alarms and context-aware broadcasting is achieved, significantly improving alarm response timeliness and driving safety.

[0097] Example 2

[0098] like Figure 2 The diagram shown is a second flowchart of a voice prompt method provided in an embodiment of this application, which includes the following steps.

[0099] S201: Obtain the current operating status of the vehicle.

[0100] The vehicle's current operating status includes, but is not limited to, vehicle speed, gear, ambient temperature, engine speed, and battery charge.

[0101] In some examples, the current operating status of a vehicle can be obtained by accessing its ECU (Electronic Control Unit) or reading relevant status messages via the Controller Area Network (CLAN) bus. For instance, vehicle speed information typically comes from the ABS (Anti-lock Braking System) or ESP (Electronic Stability Program) controller and can be read via CLAN messages.

[0102] S202: Based on the semantic description of the fault, fault handling suggestions, and the current operating status, generate personalized suggestions through a preset large language model.

[0103] Among them, fault handling suggestions are pre-associated with diagnostic fault codes and stored in a semantic mapping database.

[0104] In some examples, the semantic description of the fault, fault handling suggestions, and the current operating status can be used as prompts, which are then substituted into a preset question template to obtain the corresponding question query text. This question query text is then used to guide the large language model to generate personalized suggestions. The large language model has powerful semantic understanding and generation capabilities, enabling it to provide more accurate and contextualized suggestions based on the dynamic vehicle status.

[0105] For example, a complete problem query text could be: "System detected a fault: [Fault semantic description]. Standard handling suggestion: [Fault handling suggestion]. Current vehicle status: speed [80km / h], gear [D], ambient temperature [35℃]. Please generate a short, clear, personalized voice prompt text for the driver that includes the above information." A large language model might generate: "Please note that the engine system has detected an overly lean condition. You are currently driving at high speed. We recommend that you maintain a constant speed and check the fuel tank cap at the next service area as soon as possible."

[0106] S203: Determine the corresponding voice playback configuration based on the fault risk level.

[0107] The voice playback configuration includes at least acoustic features and playback priority. Playback priority determines the playback weight of voice prompts relative to other in-vehicle audio tasks (such as music and navigation), while acoustic features directly affect the user's auditory experience and perception of urgency.

[0108] In some examples, acoustic features include, but are not limited to, intonation, speech rate, and speech packets.

[0109] For example, the tone types include rapid, steady, and gentle; the speech speed types include 1.5x (fast), 1.0x (normal), and 0.8x (slow); and the voice pack types include, but are not limited to, a celebrity's voice, one's own voice, and an anime character's voice. For "critical" level faults, a male voice pack with a speech speed of 1.5x and a rapid tone can be configured; for "alert" level faults, a female voice pack with a speech speed of 0.8x and a gentle tone can be configured.

[0110] It should be noted that after determining the fault risk level, the tone, speech rate, and voice packets corresponding to the fault risk level can be retrieved from the local database, and the retrieved tone, speech rate, and voice packets can be configured as the acoustic features corresponding to the fault risk level.

[0111] In some examples, different risk levels correspond to different broadcast priorities. After determining the fault risk level, the broadcast priority corresponding to that fault risk level can be determined directly based on the preset mapping relationship between risk level and priority. Generally speaking, this mapping relationship can be a preset lookup table.

[0112] For example, the higher the fault risk level, the higher the corresponding broadcast priority; the lower the fault risk level, the lower the corresponding broadcast priority. For instance, "Critical" corresponds to the highest priority level 1, "Warning" corresponds to level 2, and "Tips" corresponds to level 3.

[0113] S204: Generate the voice to be played based on personalized suggested content and voice playback configuration.

[0114] Specifically, based on personalized suggestions, a corresponding natural language speech can be generated using a preset TTS model. Furthermore, based on the speech playback configuration, the acoustic features of this natural language speech can be adjusted using a preset speech intonation modulation engine to obtain a speech to be played carrying a playback priority tag. This tag can be recognized by the audio management module and used for subsequent playback decisions.

[0115] In possible implementations, the raw speech generated by the TTS model typically has neutral or default acoustic features, such as normal speaking speed (approximately 4-5 words per second), moderate volume, and neutral intonation. To reflect the broadcast style corresponding to the fault risk level, the generated raw speech also needs to undergo acoustic feature modulation.

[0116] Optionally, the process of generating the voice to be played based on personalized suggested content and voice playback configuration can be found in [link to documentation]. Figure 8 The steps are shown.

[0117] It should be noted that generating the voice message to be played involves combining personalized suggestions in text form with acoustic parameters in the voice playback configuration, ultimately outputting an audio data stream that can be played through the vehicle's speakers. The core of this step lies in fusing semantic information with expressive style, ensuring that the voice message not only matches the fault scenario and vehicle status in terms of content, but also accurately conveys the urgency of the fault and the intended interaction, thereby improving the user's perception efficiency and willingness to respond to the prompts.

[0118] The processes described in S201-S204 above acquire the vehicle's current operating status, combine it with fault semantic descriptions and handling suggestions, generate personalized suggestion content using a large language model, and determine acoustic features and broadcast priorities based on the fault risk level, thereby generating the voice to be broadcast. This process constructs a voice tone and broadcast priority control mechanism based on fault risk level, achieving auditory level recognition and significantly improving the user's ability to perceive fault risks.

[0119] Example 3

[0120] like Figure 3 The diagram shown is a third flowchart of a voice prompt method provided in an embodiment of this application, including the following steps.

[0121] S301: Obtain multimodal information related to the user's driving behavior.

[0122] Multimodal information includes driving behavior data collected in real time by different onboard sensors. Multimodal fusion aims to comprehensively assess the driver's current state and behavior from multiple dimensions, making it more robust and accurate than judgments from a single sensor.

[0123] In some examples, different vehicle sensors include cameras, lidar, millimeter-wave radar, IMU (Inertial Measurement Unit), human body temperature sensors, facial recognition cameras, recording devices, etc., which are involved in ADAS sensors.

[0124] For example, the various driving behavior data involved in multimodal information include, but are not limited to, user body temperature (reflecting tension or fatigue), video of the user manipulating the steering wheel (acquired through a driver monitoring camera), vehicle acceleration (provided by an IMU), vehicle turning angle, image of the user's eye gaze direction (analyzed through a camera), and user voice (identifying emotions such as anger and anxiety).

[0125] S302: Determine the corresponding user driving behavior based on multimodal information.

[0126] After obtaining multimodal information, the corresponding user driving behavior can be determined by analyzing the multimodal information, such as "distracted driving," "high-risk driving," or "low-risk driving." This is a typical pattern recognition or behavior classification problem.

[0127] Optionally, the process of determining the corresponding user driving behavior based on multimodal information can be as follows: based on multimodal information, the corresponding user driving behavior is obtained through a driving behavior analysis model; the construction process of the driving behavior analysis model includes: constructing training samples based on pre-collected sample multimodal information and corresponding sample user driving behavior; using the training samples, training an initial multimodal large model to obtain the driving behavior analysis model.

[0128] In some examples, multimodal information of the samples can be used as the content of the training samples, and the driving behavior of the sample users can be labeled as the training samples.

[0129] It should be noted that the initial multimodal large model refers to a deep learning system capable of simultaneously processing and understanding information from multiple modalities, such as text, images, audio, video, and sensor data. Its core objective is to achieve more accurate and comprehensive information understanding and decision-making capabilities than single-modal systems through cross-modal fusion and interaction. For example, this model can simultaneously process the driver's eye movement data, heart rate changes, and vehicle lateral deviation to comprehensively determine whether the driver is distracted.

[0130] S303: When the user's driving behavior is distracted, the playback strategy for the voice message to be played is determined based on the first preset strategy.

[0131] The first preset strategy is to increase the broadcast volume and repeat the broadcast twice.

[0132] It should be noted that when a user is driving distractedly, such as when their gaze is off the road for an extended period or they are adjusting the air conditioning, they may not notice the warning indicator light on the instrument panel. To attract the user's attention and ensure that the fault information is effectively received, the announcement volume is increased and the announcement is repeated twice to ensure that the user can clearly hear the voice message.

[0133] S304: When the user's driving behavior is high-risk, the playback strategy for the voice message to be broadcast is determined based on the second preset strategy.

[0134] The second preset strategy is to delay the broadcast and reduce the broadcast volume.

[0135] It should be noted that when a user's driving behavior is at a high risk, such as when the vehicle is swerving at high speed to avoid obstacles, navigating complex intersections, or when the driver is braking suddenly, the user's cognitive load is already extremely high. The user's attention is focused on non-driving activities, such as using a mobile phone while driving at high speed, or other important matters. If suddenly startled by an in-car voice message, it could frighten the user and cause them to lose control of their driving. Therefore, the announcement needs to be delayed and the volume lowered to avoid disturbing the user at that moment. The system will add the announcement to the queue and wait until the vehicle's status stabilizes (e.g., acceleration and steering wheel angle return to normal) before broadcasting it.

[0136] S305: When the user's driving behavior is in a low-risk driving situation, the playback strategy for the voice message to be broadcast is determined based on the third preset strategy.

[0137] The third preset strategy is to maintain the default volume for playback.

[0138] It should be noted that when a user's driving behavior is in a low-risk driving state, such as when the vehicle is traveling at a constant speed in a straight line on a smooth road, and it is determined that the user is currently in a normal driving state with a low cognitive load, the user may be able to hear any voice content broadcast in the car. Therefore, it is sufficient to maintain the default volume for broadcasting.

[0139] The processes shown in S301-S305 above determine the user's driving behavior and its corresponding playback strategy by integrating multimodal information collected from different sensors, thereby achieving intelligent broadcasting tailored to the individual and the context, avoiding information overload or omission, and realizing context-aware voice guidance.

[0140] Example 4

[0141] like Figure 4 The diagram shown is a fourth flowchart of a voice prompt method provided in an embodiment of this application, including the following steps.

[0142] S401: Determine whether the playback priority of the voice to be played meets the standard.

[0143] The broadcast priority is determined based on the fault risk level. If the broadcast priority of the voice to be broadcast meets the standard, S402 is executed; if the broadcast priority of the voice to be broadcast does not meet the standard, S403 is executed.

[0144] In some examples, the speech to be played carries a playback priority label. By parsing this label, the playback priority of the speech can be determined. For instance, if the playback priority is level one or two, the playback priority is considered met; if the playback priority is level three, the playback priority is considered not met. The threshold for "meeting the standard" can be dynamically adjusted based on the vehicle's current safety policy.

[0145] S402: Interrupt the currently executing audio task of the vehicle multimedia system, prioritize the playback of the voice message to be played according to the playback strategy, and then continue the playback of the audio task.

[0146] Among these measures, when the priority of the voice message to be played meets the standard, it is more important to determine the current indicator light alarm event. The fault risk level corresponding to this event is high and poses a direct or potential serious threat to driving safety. Therefore, the audio manager will send a "pause" command to the application currently playing audio (such as a music app or navigation app), interrupting the audio task currently being executed by the vehicle's multimedia system. After prioritizing the playback of the voice message to be played according to the playback strategy, a "resume" command will be sent to the application to continue the audio task playback so that the user can hear the voice message to be played clearly.

[0147] S403: While playing audio tasks, play the voice to be broadcast simultaneously according to the playback strategy.

[0148] In cases where the playback priority of the voice message to be played is not met, and the current indicator light alarm event is deemed relatively minor with a low risk level, it is deemed unnecessary to interrupt the user's immersive experience. Therefore, the audio manager treats the voice message to be played as a new audio stream and, while executing the audio task, simultaneously plays the voice message according to the playback strategy. This is typically achieved through audio mixing techniques, such as setting the volume of the voice message to be played to be higher than the background music to ensure the user can hear it without completely drowning out the background noise.

[0149] The processes described in S401-S403 above determine whether the playback priority of the voice message to be broadcast meets the standard. If it does, the currently executing audio task of the vehicle's multimedia system is interrupted, and the voice message to be broadcast is played first before the original task continues. This process adopts an interrupt-based broadcasting strategy based on the broadcasting priority mechanism to ensure that high-priority alarm information is not missed, thereby improving driving safety.

[0150] Example 5

[0151] like Figure 5 The diagram shown is a fifth flowchart of a voice prompt method provided in an embodiment of this application, including the following steps.

[0152] S501: Collects user voice responses to the audio to be played.

[0153] The system can capture user voice responses using in-vehicle recording equipment or microphone arrays. Microphone arrays, with their beamforming and noise suppression capabilities, can effectively pick up voice commands from drivers or passengers in noisy in-vehicle environments.

[0154] S502: Based on the user's voice response, the corresponding reply voice is obtained through the intelligent voice assistant pre-deployed in the vehicle.

[0155] This process involves inputting the user's voice response into a smart voice assistant to obtain a reply voice output by the assistant. Smart voice assistants typically include modules for speech recognition, natural language understanding, dialogue management, and speech synthesis. They can analyze the user's intent and, based on a knowledge base or predefined dialogue flow, generate appropriate response text, which is then synthesized into a reply voice.

[0156] In some examples, the user's voice response can also be converted into corresponding user voice-text and displayed on the vehicle's multimedia system, enabling collaborative interaction between voice and text, and making it convenient for users to confirm whether their questions have been correctly recognized.

[0157] For example, after the initial voice prompt (i.e., playing the voice message to be broadcast), if the user responds with "What should I do?" or "Continue," the user's voice response is input into the intelligent voice assistant to obtain the response voice output by the intelligent voice assistant, such as "Please turn off the air conditioner, open the hood to check the coolant level. If the level is below the standard line, do not add water and contact roadside assistance immediately."

[0158] S503: Play the response voice.

[0159] In addition to playing the response voice, the system can also convert the response voice into corresponding response voice text and display the response voice text on the vehicle's multimedia screen, providing users with visual reference or making it easier to understand in noisy environments.

[0160] In some examples, multi-turn voice interactions with the user can be recorded and visualized on the vehicle's multimedia system, allowing the user to review the multi-turn voice interaction records on the vehicle's multimedia system and form a complete fault consultation history.

[0161] For example, vehicle multimedia can be an in-vehicle display screen.

[0162] The processes described in S501-S503 above involve playing the voice message to be broadcast, collecting the user's voice response, and then executing the playback after obtaining the reply voice through the vehicle's pre-deployed intelligent voice assistant. This process supports multi-round voice interaction with the user, enabling interactive voice fault querying, meeting the user's need to inquire about handling details, and significantly improving the user experience.

[0163] Example 6

[0164] like Figure 6 The diagram shown is a sixth flowchart of a voice prompt method provided in an embodiment of this application, including the following steps.

[0165] S601: Obtain the pre-stored controller LAN database file.

[0166] The Controller Area Network (CAN) database file is used to define the physical meaning of each data bit in the CAN message, including message identifier, signal start bit, signal length, byte order, data type, scaling factor, and offset.

[0167] In some examples, the Controller Area Network (CAN) Database file is also called a DBC (CAN Database) file, which is a standard format file widely used in vehicle network communication. This DBC file is typically pre-installed on the storage media of the onboard computing platform and can be updated with vehicle software upgrades.

[0168] In a possible implementation, the onboard computing platform obtains the parsing rules for controller area network (CAN) messages by reading a pre-stored DBC file in its local non-volatile memory. The DBC file describes the signal layout corresponding to each message ID in text form. For example, for message ID 0x18FEC000, the 5th bit is defined as the engine fault indicator status signal, with a signal length of 1 bit, a data type of Boolean, a scaling factor of 1, and an offset of 0.

[0169] It should be noted that the Controller Area Network (DAC) database file is the foundation for parsing DAC messages. Different vehicle models or DAC protocol versions may correspond to different DBC files. By pre-storing the DBC file adapted to the current vehicle, accurate message parsing can be ensured, avoiding misreading or parsing failures due to protocol incompatibility.

[0170] S602: Based on the Controller Area Network (CAN) database file, perform bit-by-bit parsing of CAN packets to obtain bit-by-bit parsing results.

[0171] Bit-by-bit parsing refers to extracting the specific values ​​of each signal from the binary data of the original message one by one, according to the definition of each signal bit in the controller LAN database file.

[0172] In some examples, the Controller Area Network (CLAN) packets captured by the in-vehicle computing platform are binary sequences containing a packet ID, data length, and data field. The CLAN protocol parsing module on the in-vehicle computing platform first looks up the corresponding signal layout definition in the DBC file based on the packet ID. Then, according to the start bit, signal length, byte order, and other parameters in the definition, it reads the data bit by bit from the corresponding position in the data field and calculates the physical value by applying the scaling factor and offset.

[0173] In a possible implementation, when a Controller Area Network (CLAN) message with message ID 0x18FEC000 and data field 0x08 is received, the parsing module, according to the definition of the message ID in the DBC file: the engine fault indicator light signal starts at the 5th bit and has a length of 1 bit, so a value of 1 for this bit indicates that the indicator light is active. The parsing module extracts the value 1 from the 5th bit of the data field, thus obtaining the bit-by-bit parsing result.

[0174] In some examples, where multiple signals are mixed in the same message, bit-by-bit parsing can be performed in parallel or serial mode to ensure that all the required signal values ​​are obtained quickly.

[0175] S603: Based on the bit-by-bit parsing results, extract the diagnostic fault codes corresponding to the alarm indicator lights.

[0176] Diagnostic fault codes are a set of standardized codes used to identify the specific type of fault that has occurred in the vehicle. When a certain indicator light status signal is set (e.g., a value of 1) in the bit-by-bit parsing result, the diagnostic fault code corresponding to that indicator light can be extracted according to a preset mapping relationship.

[0177] In some examples, there is a one-to-one correspondence between warning indicator lights and diagnostic fault codes. For instance, if the engine malfunction indicator light is activated, the corresponding diagnostic fault code is P0171 (fuel system too lean); if the brake system malfunction warning indicator light is activated, the corresponding diagnostic fault code is C1234 (wheel speed sensor signal abnormality).

[0178] In one possible implementation, the onboard computing platform maintains a mapping table between indicator light status signals and diagnostic fault codes. When the bit-by-bit parsing result indicates that a certain indicator light status signal is active, the platform quickly extracts the corresponding diagnostic fault code by querying the mapping table for subsequent lookup in the semantic mapping database.

[0179] It should be noted that the extraction of diagnostic fault codes is a crucial step connecting the underlying message parsing with the upper-layer semantic understanding. By converting the raw bit information into standardized diagnostic fault codes, a unified and standardized input can be provided for the subsequent generation of fault semantic descriptions.

[0180] The processes described in S601-S603 above, by acquiring the controller LAN database file and parsing the messages bit by bit, can accurately and efficiently extract the diagnostic fault codes corresponding to the alarm indicator lights from the original controller LAN messages. Compared with traditional methods that rely on visual recognition or polling diagnostic requests, this process achieves low-latency fault code acquisition through "message parsing," laying a data foundation for rapid response to subsequent voice prompts. It also avoids parsing errors caused by protocol incompatibility, improving the system's reliability and compatibility.

[0181] Example 7

[0182] like Figure 7 The diagram shown is a seventh flowchart of a voice prompt method provided in an embodiment of this application, which includes the following steps.

[0183] S701: Generate graphic and textual fault reports based on fault semantic descriptions.

[0184] Among them, graphic fault reports refer to visual information documents that include textual descriptions of faults and related fault icons, used to intuitively show users the type of fault currently occurring in the vehicle and suggested handling measures.

[0185] In some examples, the semantic description of the fault comes from the natural language description corresponding to the diagnostic fault code in the semantic mapping database, such as "engine fuel system too lean". The graphic fault report can be based on the text description, with the corresponding icon (such as engine icon, warning triangle) and color code matching the fault risk level (such as red for critical, yellow for warning, and green for alert).

[0186] In a possible implementation, after the onboard computing platform finishes playing the voice message, it calls the image and text generation module to combine the fault semantic description, fault risk level, and preset handling suggestion templates to render and generate an image or PDF document containing a fault icon, detailed text description, and suggested operating steps. For example, for a fault of "insufficient brake fluid," the image and text fault report may include a brake fluid reservoir icon, the text "Brake fluid level is too low, please add brake fluid as soon as possible and check the braking system," and a shortcut button for "Contact a repair shop."

[0187] It should be noted that the generation of graphic fault reports does not require manual operation by the user; it is automatically triggered by the system. The purpose is to provide users with a visual archive of fault information, making it easy to review carefully after parking or forward to maintenance personnel.

[0188] In some examples, for vehicles that support large language models, more detailed fault analysis content and illustrated instructions can be generated based on fault semantic descriptions and the vehicle's current operating status, further improving the readability and usability of the report.

[0189] S702: Based on the vehicle's current location information, it obtains repair shop recommendations through the in-vehicle navigation map.

[0190] The recommended repair shop information includes the name, distance, address, contact information, and user reviews of nearby repair shops found based on the vehicle's current location.

[0191] In some examples, the in-vehicle navigation map can be the vehicle's built-in navigation system (such as in-vehicle versions of Amap or Baidu Maps), or it can be obtained through real-time communication between the vehicle and a cloud-based map service. The vehicle obtains its current precise latitude and longitude coordinates through the Global Positioning System (GPS) or the BeiDou Navigation Satellite System (BDS) module.

[0192] In a possible implementation, the onboard computing platform uses the vehicle's current location coordinates as the search center point, calls the surrounding search interface of the onboard navigation map, sets the search radius to 10 kilometers, and filters out repair shops with the corresponding fault repair qualifications (e.g., capable of handling engine faults, high-voltage battery faults, etc.), sorting them from nearest to farthest. For critical faults, the system can prioritize recommending the nearest repair shop; for warning-type faults, it can prioritize recommending repair shops with higher ratings.

[0193] In some examples, the recommended repair shops can be further filtered by combining fault handling suggestions from the semantic mapping database. For instance, if the fault semantic description is "high-voltage battery insulation fault," the system will only recommend repair shops with new energy vehicle repair qualifications, excluding ordinary repair shops that do not have such qualifications.

[0194] S703: Sends a text and image fault report along with recommended repair shops to a mobile terminal that has been pre-connected to the vehicle.

[0195] The mobile terminal includes, but is not limited to, devices such as the owner's mobile phone and tablet computer that are linked to the vehicle. Pre-established connections can be achieved through Bluetooth pairing, Wi-Fi Direct, or network communication based on a cloud-based account system.

[0196] In some examples, the vehicle uses its built-in T-Box (Telematics Box, in-vehicle communication module) to send text and image fault reports and repair shop recommendations to the vehicle app installed on the owner's mobile phone via push notifications. The owner can view the text and image fault reports in the app's message center and use their mobile map to navigate to the recommended repair shop with a single click.

[0197] In one possible implementation, the vehicle first uploads a text and image-based fault report and recommended repair shops to a cloud server. The cloud server then uses a push notification service to send the information to the corresponding mobile device based on the user account associated with the vehicle's identifier. Upon receiving the information, the mobile device alerts the user via a pop-up notification or an in-app message.

[0198] It should be noted that sending information to a mobile device can overcome the inconvenience of viewing it on the in-vehicle screen. Users can still check fault reports and repair shop information on their mobile phones at any time after parking or leaving the vehicle, making it easier to arrange subsequent repairs.

[0199] In some examples, after receiving the information, the user can click the navigation button in the repair shop recommendation information to directly call the map app on their phone to plan the route to the repair shop, thus realizing a complete closed loop of "vehicle alarm → mobile phone follow-up → one-click navigation".

[0200] The processes described in S701-S703 above generate a text and image fault report based on the semantic description of the fault, obtain recommended repair shop information by combining the vehicle's current location, and push it to a pre-connected mobile terminal. This process, building upon voice broadcasting, provides users with a visual, storable, and shareable fault information carrier, forming a complete closed-loop handling mechanism of "voice prompts → mobile terminal follow-up." This allows users to easily review the information after parking, forward it to repair personnel, or navigate to a repair shop with one click, effectively improving after-sales efficiency and user experience.

[0201] Example 8

[0202] like Figure 8 The diagram shown is an eighth flowchart of a voice prompt method provided in an embodiment of this application, which includes the following steps.

[0203] S801: The personalized suggestion content is used as the text to be broadcast, and the acoustic feature parameters of the text to be broadcast are labeled according to the acoustic features.

[0204] The personalized suggestions are text generated by a large language model, such as "Please note that the engine fuel system is too lean. Current speed is 80 km / h. We recommend maintaining a constant speed and checking the fuel cap at the next service area as soon as possible." Acoustic features include parameters such as intonation and speech rate, used to control the expressive style of the synthesized speech.

[0205] Acoustic feature parameter annotation refers to adding markers at specific locations (such as characters, words, phrases, or sentences) of the text to be broadcast to guide how the text-to-speech engine pronounces the text. Examples of markers include pitch curves, duration, stress, and pauses.

[0206] In some examples, for critical faults, the acoustic features can be configured as a rapid tone and 1.2x speech rate; for warning faults, a steady tone and normal speech rate; and for prompt faults, a gentle tone and 0.8x speech rate. Acoustic feature parameter annotations can be generated based on preset rules, such as marking final particles as rising intonation, key numerical information as stressed, and handling suggestions as appropriate pauses.

[0207] In a possible implementation, the in-vehicle computing platform first acquires the acoustic features (pitch and speed) from the voice playback configuration, and then performs text analysis on the personalized suggestions. The analysis process includes: determining the overall reading time ratio based on the speed, and inserting corresponding prosodic markers into the text based on the tone type (rapid, steady, gentle).

[0208] It's important to note that acoustic feature parameter annotation is a preprocessing step completed before text-to-speech synthesis, unlike traditional post-synthesis audio processing using equalizers or filters. By incorporating acoustic features during the synthesis stage, more natural and emotionally accurate speech can be generated, avoiding distortion or unnaturalness caused by post-processing.

[0209] In some examples, for systems that support end-to-end neural TTS models, acoustic features can be directly used as conditional inputs to the model (e.g., through speaker coding or style coding), and the end-to-end neural TTS model can automatically learn and generate the required speech without the need for explicit annotation labels.

[0210] S802: Using a text-to-speech engine, synthesize the speech to be played according to the acoustic feature parameters, and set a tag for the speech to be played that corresponds to the playback priority.

[0211] The text-to-speech engine is a software module used to convert annotated text into audio signals. Based on the instructions in the acoustic feature parameter annotations, the engine adjusts the prosodic features of the synthesized speech to generate the final audio data. The tag corresponding to the playback priority is a metadata marker used by the subsequent audio management module to determine the scheduling priority of the voice prompt relative to other audio tasks (such as navigation and music).

[0212] In some examples, the text-to-speech engine can be a parametric engine or a neural network model. For text labeled with rapid intonation, the neural network model generates audio with a faster speech rate, higher pitch, and stronger energy; for text labeled with soft intonation, it generates audio with a slower speech rate and smoother pitch.

[0213] In a possible implementation, after synthesizing the speech to be played, the text-to-speech engine outputs an audio data packet and attaches a tag field to it. The value of this tag field is determined according to the playback priority. For example, the tag is set to "PRIORITY_HIGH" for priority level 1 (critical), "PRIORITY_MEDIUM" for priority level 2 (warning), and "PRIORITY_LOW" for priority level 3 (hint). This tag is sent to the vehicle audio manager along with the audio data packet.

[0214] It's important to note that setting playback priority tags is a crucial prerequisite for implementing interrupted or mixed playback strategies. When multiple audio streams exist simultaneously, the audio manager can determine which audio stream can interrupt the others based on the tag values, thus achieving an interrupted playback mechanism.

[0215] In some examples, for audio frameworks that do not support tag settings, the same effect can be achieved by encoding the playback priority into the channel properties or metadata extension fields of the audio stream.

[0216] The processes described in S801-S802 above use personalized suggestions as the text to be played. Acoustic feature parameters are labeled based on acoustic characteristics, and a text-to-speech engine is used to synthesize speech according to these labels. Simultaneously, tags corresponding to the playback priority are assigned to the synthesized speech. This process achieves deep coupling between content and style, ensuring that the voice prompts not only semantically fit the fault scenario but also accurately convey the urgency of the fault aurally. It also provides a identifiable basis for subsequent playback priority scheduling, thereby enhancing the user's perception of fault risks and improving the system's resource scheduling efficiency.

[0217] Example 9

[0218] like Figure 9The diagram shown is a schematic representation of a vehicle architecture according to an embodiment of this application. The vehicle includes a processor 901, a memory 902, and a bus 903. The processor 901 and the memory 902 are connected via the bus 903. The memory 902 is used to store programs, and the processor 901 is used to run programs. When the program runs, it executes the voice prompt method provided in this application, which can be encapsulated into the unit shown below.

[0219] The message listening unit 100 is used to listen to the vehicle's controller area network messages.

[0220] The message parsing unit 200 is used to determine the diagnostic fault code of the alarm indicator light based on the controller local area network message when an indicator light alarm event occurs.

[0221] Optionally, the message parsing unit 200 is specifically used for: obtaining a pre-stored controller area network (MAN) database file; the MAN database file is used to define the physical meaning of each data bit in the MAN message; parsing the MAN message bit by bit according to the MAN database file to obtain the bit-by-bit parsing result; and extracting the diagnostic fault code corresponding to the alarm indicator based on the bit-by-bit parsing result.

[0222] The semantic query unit 300 is used to determine the fault semantic description and fault risk level corresponding to the diagnostic fault code based on the semantic mapping database.

[0223] The speech generation unit 400 is used to generate speech to be broadcast based on the semantic description of the fault and the fault risk level.

[0224] Optionally, the voice generation unit 400 is specifically used for: obtaining the current operating status of the vehicle; generating personalized suggestion content based on the fault semantic description, fault handling suggestions and the current operating status through a preset large language model; pre-associating the fault handling suggestions with the diagnostic fault codes and storing them in the semantic mapping database; determining the corresponding voice playback configuration based on the fault risk level; and generating the voice to be played based on the personalized suggestion content and the voice playback configuration.

[0225] Optionally, the voice playback configuration includes at least acoustic features and playback priority. The voice generation unit 400 is specifically used to: take the personalized suggestion content as the text to be played, and determine the acoustic feature parameter annotation of the text to be played according to the acoustic features; use a text-to-speech engine to synthesize the voice to be played according to the acoustic feature parameter annotation, and set a tag corresponding to the playback priority for the voice to be played.

[0226] The strategy determination unit 500 is used to determine the playback strategy of the voice to be played based on real-time detected user driving behavior.

[0227] Optionally, the strategy determination unit 500 is specifically used for: obtaining multimodal information related to the user's driving behavior; the multimodal information includes driving behavior data collected in real time by different vehicle sensors; determining the corresponding user driving behavior based on the multimodal information; when the user's driving behavior is distracted driving, determining the playback strategy for the voice message to be played based on a first preset strategy; the first preset strategy is: increasing the playback volume and repeating the playback twice; when the user's driving behavior is high-risk driving, determining the playback strategy for the voice message to be played based on a second preset strategy; the second preset strategy is: delaying the playback and reducing the playback volume; when the user's driving behavior is low-risk driving, determining the playback strategy for the voice message to be played based on a third preset strategy; the third preset strategy is: maintaining the default volume for playback.

[0228] Optionally, the strategy determination unit 500 is specifically used to: obtain the corresponding user driving behavior based on multimodal information through a driving behavior analysis model; the construction process of the driving behavior analysis model includes: constructing training samples based on pre-collected sample multimodal information and corresponding sample user driving behavior; and using the training samples to train an initial multimodal large model to obtain the driving behavior analysis model.

[0229] The voice playback unit 600 is used to play the voice to be played according to the playback strategy.

[0230] Optionally, the voice playback unit 600 is specifically used to: determine whether the playback priority of the voice to be played meets the standard, wherein the playback priority is determined based on the fault risk level; if the playback priority meets the standard, interrupt the audio task currently being executed by the vehicle multimedia, prioritize the playback of the voice to be played according to the playback strategy, and then continue the playback of the audio task; if the playback priority does not meet the standard, while executing the playback of the audio task, simultaneously execute the playback of the voice to be played according to the playback strategy.

[0231] The intelligent interaction unit 700 is used to: collect user voice responses to the voice to be played; obtain corresponding reply voices through a pre-deployed intelligent voice assistant in the vehicle based on the user voice responses; and execute the playback of the reply voices.

[0232] The intelligent recommendation unit 800 is used to: generate a graphic fault report based on the semantic description of the fault; obtain repair shop recommendation information through the vehicle navigation map based on the vehicle's current location information; and send the graphic fault report and repair shop recommendation information to a mobile terminal that has been pre-connected to the vehicle.

[0233] Each unit described above diagnoses fault codes by monitoring and parsing LAN messages, converting them into fault descriptions and risk levels through semantic mapping. Personalized suggestions are generated using a large language model, and playback strategies are adaptively adjusted based on multimodal driving behavior. Acoustic features and priorities are configured based on risk levels; high-priority interruptions trigger emergency multimedia broadcasts, while low-priority audio is mixed to avoid interference. After broadcasting, multi-round voice interaction and Q&A are supported, and a text and image fault report is automatically generated and pushed to the mobile terminal, forming a closed loop of "voice prompts → mobile follow-up." Overall, this achieves end-to-end intelligent voice guidance from fault occurrence to user response, lowering the threshold for fault understanding and reducing misjudgment rates, shortening response time, reducing accident rates and distraction time caused by ignoring visual alarms, improving after-sales efficiency and user experience, and comprehensively overcoming the limitations of purely visual alarms on instrument panels.

[0234] While several specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of this application. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0235] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A voice prompt method, characterized in that, include: Monitor the vehicle's controller area network (LAN) messages; When an indicator light alarm event occurs, the diagnostic fault code of the alarm indicator light is determined based on the controller LAN message; Based on the semantic mapping database, determine the fault semantic description and fault risk level corresponding to the diagnostic fault code; Based on the fault semantic description and the fault risk level, generate the voice to be broadcast; Based on real-time detection of user driving behavior, the playback strategy for the voice message to be played is determined. According to the playback strategy, the audio to be played is executed.

2. The method according to claim 1, characterized in that, Based on the fault semantic description and the fault risk level, a voice message to be broadcast is generated, including: Obtain the current operating status of the vehicle; Based on the fault semantic description, fault handling suggestions, and the current operating status, personalized suggestions are generated through a preset large language model; the fault handling suggestions are pre-associated with the diagnostic fault codes and stored in the semantic mapping database; Based on the aforementioned fault risk level, determine the corresponding voice playback configuration; Based on the personalized suggestions and the voice playback configuration, a voice message to be played is generated.

3. The method of claim 1, wherein, Based on real-time detected user driving behavior, the playback strategy for the voice message to be played is determined, including: Obtain multimodal information related to user driving behavior; the multimodal information includes driving behavior data collected in real time by different on-board sensors; Based on the multimodal information, the corresponding user driving behavior is determined; When the user's driving behavior is distracted, a playback strategy for the voice message to be played is determined based on a first preset strategy; the first preset strategy is to increase the playback volume and repeat the playback twice. When the user's driving behavior is high-risk, a playback strategy for the voice message to be played is determined based on a second preset strategy; the second preset strategy is: delay the playback and reduce the playback volume. When the user's driving behavior is in a low-risk driving situation, the playback strategy for the voice message to be played is determined based on a third preset strategy; the third preset strategy is to maintain the default volume for playback.

4. The method of claim 3, wherein, Based on the multimodal information, the corresponding user driving behavior is determined, including: Based on the multimodal information, the corresponding user driving behavior is obtained through a driving behavior analysis model. The construction process of the driving behavior analysis model includes: constructing training samples based on pre-collected sample multimodal information and corresponding sample user driving behavior; and using the training samples to train an initial multimodal large model to obtain the driving behavior analysis model.

5. The method of claim 1, wherein, After playing the audio to be played according to the playback strategy, the method further includes: Collect the user's voice response to the voice to be played; Based on the user's voice response, the corresponding reply voice is obtained through the intelligent voice assistant pre-deployed in the vehicle; Play the response voice message.

6. The method of claim 1, wherein, After playing the audio to be played according to the playback strategy, the method further includes: Based on the semantic description of the fault, a graphic and textual fault report is generated; Based on the vehicle's current location information, repair shop recommendations are obtained through the in-vehicle navigation map; The graphic fault report and the recommended repair shop information are sent to a mobile terminal that has been pre-connected to the vehicle.

7. The method of claim 1, wherein, According to the playback strategy, the playback of the voice to be broadcast is performed, including: Determine whether the broadcast priority of the voice message to be broadcast meets the standard; wherein, the broadcast priority is determined based on the fault risk level; If the broadcast priority is met, the audio task currently being executed by the vehicle multimedia system is interrupted, and the playback of the voice to be broadcast is executed first according to the playback strategy before the playback of the audio task continues. If the broadcast priority is not met, the playback of the audio task will be performed simultaneously with the playback of the voice to be broadcast, according to the playback strategy.

8. The method of claim 2, wherein, The voice playback configuration includes at least acoustic features and playback priority; Based on the personalized suggestions and the voice playback configuration, a voice message to be played is generated, including: The personalized suggestion content is used as the text to be broadcast, and the acoustic feature parameters of the text to be broadcast are labeled according to the acoustic features. A text-to-speech engine is used to synthesize the speech to be played according to the acoustic feature parameters, and a tag corresponding to the playback priority is set for the speech to be played.

9. The method of claim 1, wherein, Based on the controller LAN messages, determine the diagnostic fault codes of the alarm indicator lights, including: Obtain a pre-stored Controller Area Network (CAN) database file; the CAN database file is used to define the physical meaning of each data bit in the CAN message; Based on the controller local area network database file, the controller local area network packets are parsed bit by bit to obtain the bit-by-bit parsing results; Based on the bit-by-bit parsing results, the diagnostic fault codes corresponding to the alarm indicator lights are extracted.

10. A vehicle characterized by comprising: include: Processor, memory, and bus; The processor and the memory are connected via the bus; The memory is used to store a program, and the processor is used to run the program, wherein the program is executed by the processor to perform the voice prompt method according to any one of claims 1-9.