Vehicle-mounted voice interaction method and device, vehicle, and storage medium

CN117594039BActive Publication Date: 2026-09-04CHERY AUTOMOBILE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311630858.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2026-09-04
Estimated Expiration
2043-11-29

AI Technical Summary

Technical Problem

[0003]而在上述方法中,由于车辆使用场景通常较为公开,若播放的语音信息中涉及用户隐私,公开播放的语音信息会导致用户使用体验降低,因此,当前的语音交互方法的灵活性较低

Benefits of technology

[0058]通过设置语音交互数据库,在车辆的行驶场景满足语音交互数据库中的目标语音触发场景的情况下,基于该目标语音触发场景对应的语音生成规则以及行驶场景的场景信息生成并播放语音,从而实现基于行驶场景的语音自动播放,而无需驾驶员通过语音唤醒,提高语音交互效率。并且,考虑到在车载语音交互过程中,语音播放均为在车内公开播放,可能会导致驾驶员的隐私泄露等问题,影响驾驶员的使用体验,因此,通过设置包括不同语音交互权限对应的语音生成规则的语音生成规则集,在行驶场景满足目标语音触发场景的情况下,能够基于车内人员对应的目标语音交互权限,从语音生成规则集中获取目标语音生成规则,从而在基于语音生成规则生成语音时,能够基于不同的语音交互权限,生成不同的语音并进行播放,进而实现在车内人员权限不同的情况下,生成并播放的语音也存在区别的效果,以提高语音交互的灵活性和私密性,提升驾驶员的使用体验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117594039B_ABST
    Figure CN117594039B_ABST
Patent Text Reader

Abstract

The application discloses a vehicle-mounted voice interaction method and device, equipment and storage medium, and belongs to the technical field of vehicles. The method comprises the following steps: determining a voice interaction database of a driver; in the case that a driving scene of a target vehicle meets a target voice trigger scene, obtaining a voice generation rule set corresponding to the target voice trigger scene from the voice interaction database; determining a target voice interaction permission corresponding to an in-vehicle person, obtaining a target voice generation rule corresponding to the target voice interaction permission from the voice generation rule set corresponding to the target voice trigger scene, and generating and playing voice based on the target voice generation rule and scene information of the driving scene. The application generates and plays different voices based on different voice interaction permissions in the case that the driving scene meets the target voice trigger scene, thereby improving the flexibility and privacy of voice interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle technology, and in particular to an in-vehicle voice interaction method, device, vehicle, and storage medium. Background Technology

[0002] With the rapid development of vehicle technology, in-vehicle functions are gradually becoming more intelligent and comprehensive. Among these, intelligent cockpit technology, represented by in-vehicle voice functionality, is a typical example and has been widely used in vehicles. In these technologies, the in-vehicle voice function is mainly activated by the user sending specific commands or by enabling the vehicle to actively perform voice interaction based on preset scenarios.

[0003] In the above methods, since vehicle usage scenarios are usually quite public, if the voice information played involves user privacy, publicly playing the voice information will lead to a decrease in user experience. Therefore, the current voice interaction methods have low flexibility. Summary of the Invention

[0004] This application provides an in-vehicle voice interaction method, device, equipment, and storage medium, which can improve the flexibility of voice interaction. The technical solution is as follows:

[0005] On the one hand, an in-vehicle voice interaction method is provided, the method comprising:

[0006] A voice interaction database for the driver is determined. The voice interaction database includes at least one voice trigger scenario and a set of voice generation rules corresponding to the voice trigger scenario. The set of voice generation rules includes at least one voice generation rule. Different voice generation rules correspond to different voice interaction permissions, and different voice generation rules generate different voice content.

[0007] If the driving scenario of the target vehicle satisfies the target voice triggering scenario, the voice generation rule set corresponding to the target voice triggering scenario is obtained from the voice interaction database, wherein the target voice triggering scenario is any one of the at least one voice triggering scenario;

[0008] Determine the target voice interaction permission corresponding to the person in the vehicle, obtain the target voice generation rule corresponding to the target voice interaction permission from the voice generation rule set corresponding to the target voice triggering scenario, and generate and play voice based on the target voice generation rule and the scenario information of the driving scenario.

[0009] Optionally, before generating and playing the speech based on the target speech generation rule and the scene information of the driving scenario, the method further includes:

[0010] Obtain the driving status of the target vehicle;

[0011] Based on the environment mapping relationship, the interaction environment corresponding to the driving state is determined. The interaction environment includes a safe environment or a non-safe environment. The environment mapping relationship is used to indicate the interaction environment corresponding to different driving states.

[0012] When the interactive environment is a safe environment, the steps of generating and playing voice based on the target voice generation rules and the scene information of the driving scenario are performed.

[0013] Optionally, after generating and playing the speech based on the target speech generation rule and the scene information of the driving scenario, the method further includes:

[0014] The system receives a function command triggered by the driver and executes the target function corresponding to the function command. The function command includes: a voice command or a button command.

[0015] Optionally, the voice interaction database further includes a voice interaction priority corresponding to each voice triggering scenario, the target voice triggering scenario includes at least two voice triggering scenarios, the target voice generation rule includes at least two voice generation rules, and the at least two voice generation rules are obtained from the voice generation rule set corresponding to the at least two voice triggering scenarios according to the target voice interaction permission; the method further includes:

[0016] Obtain the voice interaction priorities corresponding to the at least two voice triggering scenarios from the voice interaction database;

[0017] The process of generating and playing voice based on the target voice generation rule and the scene information of the driving scenario includes:

[0018] Generate at least two speech entries based on the at least two speech generation rules and the scene information;

[0019] The at least two voice messages are played sequentially in descending order of voice interaction priority corresponding to the at least two voice trigger scenarios.

[0020] Optionally, before determining the driver's voice interaction database, the method further includes:

[0021] The voice interaction settings interface is displayed, which is used to prompt the driver to set the at least one voice triggering scenario and the voice generation rules for different voice interaction permissions corresponding to the voice triggering scenario.

[0022] Obtain the at least one voice triggering scenario and the voice generation rules for different voice interaction permissions corresponding to the voice triggering scenario from the voice interaction settings interface.

[0023] Based on the voice generation rules for different voice interaction permissions corresponding to the at least one voice triggering scenario, a voice generation rule set corresponding to the at least one voice triggering scenario is generated, and the at least one voice triggering scenario and the voice generation rule set corresponding to the voice triggering scenario are stored in the voice interaction database.

[0024] Optionally, determining the target voice interaction permissions corresponding to the occupants in the vehicle includes:

[0025] Retrieve the voice interaction permissions corresponding to each person in the vehicle from the personnel permission database, and obtain at least one voice interaction permission;

[0026] The lowest-level voice interaction permission among the at least one voice interaction permission is determined as the target voice interaction permission.

[0027] Optionally, before determining the target voice interaction permissions corresponding to the occupants in the vehicle, the method further includes:

[0028] The voice interaction settings interface is displayed, which is used to prompt the driver to set voice interaction permissions for different people in the vehicle.

[0029] The voice interaction permissions corresponding to different in-vehicle personnel are obtained from the voice interaction settings interface and stored in the personnel permission database.

[0030] On the other hand, an in-vehicle voice interaction device is provided, the device comprising:

[0031] The database determination module is used to determine the driver's voice interaction database. The voice interaction database includes at least one voice triggering scenario and a voice generation rule set corresponding to the voice triggering scenario. The voice generation rule set includes at least one voice generation rule. Different voice generation rules correspond to different voice interaction permissions, and different voice generation rules generate different voice content.

[0032] The rule set matching module is used to obtain the voice generation rule set corresponding to the target voice triggering scenario from the voice interaction database when the driving scenario of the target vehicle meets the target voice triggering scenario, wherein the target voice triggering scenario is any one of the at least one voice triggering scenario;

[0033] The voice playback module is used to determine the target voice interaction permission corresponding to the person in the vehicle, obtain the target voice generation rule corresponding to the target voice interaction permission from the voice generation rule set corresponding to the target voice triggering scenario, and generate and play voice based on the target voice generation rule and the scenario information of the driving scenario.

[0034] Optionally, the voice playback module is further configured to:

[0035] Obtain the driving status of the target vehicle;

[0036] Based on the environment mapping relationship, the interaction environment corresponding to the driving state is determined. The interaction environment includes a safe environment or a non-safe environment. The environment mapping relationship is used to indicate the interaction environment corresponding to different driving states.

[0037] When the interactive environment is a safe environment, the steps of generating and playing voice based on the target voice generation rules and the scene information of the driving scenario are performed.

[0038] Optionally, the in-vehicle voice interaction device further includes a function execution module, which is used to receive the function command triggered by the driver and execute the target function corresponding to the function command. The function command includes: voice command or button command.

[0039] Optionally, the voice interaction database further includes a voice interaction priority corresponding to each voice triggering scenario, the target voice triggering scenario includes at least two voice triggering scenarios, the target voice generation rule includes at least two voice generation rules, and the at least two voice generation rules are obtained from the voice generation rule set corresponding to the at least two voice triggering scenarios according to the target voice interaction permission; the rule set matching module is further configured to:

[0040] Obtain the voice interaction priorities corresponding to the at least two voice triggering scenarios from the voice interaction database;

[0041] The voice playback module is also used for:

[0042] Generate at least two speech entries based on the at least two speech generation rules and the scene information;

[0043] The at least two voice messages are played sequentially in descending order of voice interaction priority corresponding to the at least two voice trigger scenarios.

[0044] Optionally, the database determination module is further configured to:

[0045] The voice interaction settings interface is displayed, which is used to prompt the driver to set the at least one voice triggering scenario and the voice generation rules for different voice interaction permissions corresponding to the voice triggering scenario.

[0046] Obtain the at least one voice triggering scenario and the voice generation rules for different voice interaction permissions corresponding to the voice triggering scenario from the voice interaction settings interface.

[0047] Based on the voice generation rules for different voice interaction permissions corresponding to the at least one voice triggering scenario, a voice generation rule set corresponding to the at least one voice triggering scenario is generated, and the at least one voice triggering scenario and the voice generation rule set corresponding to the voice triggering scenario are stored in the voice interaction database.

[0048] Optionally, the voice playback module is specifically used for:

[0049] Retrieve the voice interaction permissions corresponding to each person in the vehicle from the personnel permission database, and obtain at least one voice interaction permission;

[0050] The lowest-level voice interaction permission among the at least one voice interaction permission is determined as the target voice interaction permission.

[0051] Optionally, the database determination module is further configured to:

[0052] The voice interaction settings interface is displayed, which is used to prompt the driver to set voice interaction permissions for different people in the vehicle.

[0053] The voice interaction permissions corresponding to different in-vehicle personnel are obtained from the voice interaction settings interface and stored in the personnel permission database.

[0054] On the other hand, a vehicle is provided, the vehicle including a memory and a processor, the memory for storing computer programs, and the processor for executing the computer programs stored in the memory to implement the steps of the in-vehicle voice interaction method described above.

[0055] On the other hand, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, it implements the steps of the above-described in-vehicle voice interaction method.

[0056] On the other hand, a computer program product containing instructions is provided, which, when executed on a computer, cause the computer to perform the steps of the in-vehicle voice interaction method described above.

[0057] The technical solution provided in this application can bring at least the following beneficial effects:

[0058] By establishing a voice interaction database, when the vehicle's driving scenario meets the target voice trigger scenario in the database, voice is generated and played based on the voice generation rules corresponding to the target voice trigger scenario and the scenario information of the driving scenario. This achieves automatic voice playback based on the driving scenario without requiring the driver to wake it up with voice, improving the efficiency of voice interaction. Furthermore, considering that voice playback during in-vehicle voice interaction is always publicly played inside the vehicle, potentially leading to privacy leaks and affecting the driver's experience, a voice generation rule set is established, including voice generation rules corresponding to different voice interaction permissions. When the driving scenario meets the target voice trigger scenario, the target voice generation rule can be obtained from the voice generation rule set based on the target voice interaction permissions of the occupants. Therefore, when generating voice based on these rules, different voices can be generated and played according to different voice interaction permissions. This ensures that the generated and played voices differ depending on the occupants' permissions, improving the flexibility and privacy of voice interaction and enhancing the driver's experience. Attached Figure Description

[0059] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0060] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application;

[0061] Figure 2 This is a flowchart of an in-vehicle voice interaction method provided in an embodiment of this application;

[0062] Figure 3 This is a schematic diagram of the structure of an in-vehicle voice interaction device provided in an embodiment of this application;

[0063] Figure 4 This is a schematic diagram of the structure of a vehicle provided in an embodiment of this application. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0065] Before providing a detailed explanation of the voice interaction provided in the embodiments of this application, the implementation environment involved in the embodiments of this application will be introduced first.

[0066] Please refer to Figure 1, Figure 1 This is a schematic diagram illustrating an implementation environment according to an exemplary embodiment. The implementation environment includes a voice interaction terminal 101, an image acquisition module 102, and a processor 103. The processor 103 can communicate with both the voice interaction terminal 101 and the image acquisition module 102. This communication connection can be wired or wireless; this embodiment does not limit the specific connection.

[0067] The voice interaction terminal 101 is used to enable voice interaction with a user. For example, the voice interaction terminal 101 may include a speaker, through which voice is played when required to trigger a voice response.

[0068] In some embodiments, the voice interaction terminal 101 may further include a microphone to receive voice commands issued by the driver when the driver needs to actively trigger voice interaction.

[0069] The scene acquisition module 102 is used to acquire scene information of the vehicle's driving scene and send the scene information to the processor 103. The scene acquisition module 102 may include a variety of acquisition devices, which can be determined in combination with the acquisition requirements of the driving scene (i.e., the triggering scene in the voice interaction database).

[0070] For example, if the driving scenario can include the environment inside and outside the vehicle, then the scene acquisition module 102 can include an image acquisition device to acquire information about the environment inside and outside the vehicle through image acquisition, and send the acquired environmental information to the processor 103. As another example, if the driving scenario can also include the vehicle's driving status, then the scene acquisition module 102 can include multiple sensors to acquire information such as vehicle speed, remaining fuel, remaining battery power, temperature inside and outside the vehicle, humidity inside and outside the vehicle, and the current location of the vehicle, and send the acquired information to the processor 103. As yet another example, if the driving scenario can also include internet data, then the scene acquisition module 102 can include an internet data acquisition module to acquire user internet data (such as user trip information, user communication information, user mobile terminal software information), current time, weather, etc., and send the acquired data to the processor 103.

[0071] The processor 103 may be equipped with a voice interaction database, which is used to receive scene information of driving scenarios sent by the scene acquisition module 102, and generate voice and send it to the voice interaction terminal 101 when the current driving scenario meets the voice trigger scenario in the voice interaction database, so that the voice can be played through the voice interaction terminal 101.

[0072] The processor 103 can be a general-purpose CPU (Central Processing Unit), NP (Network Processor), microprocessor, or one or more integrated circuits for implementing the scheme of this application, such as ASIC (Application-Specific Integrated Circuit), PLD (Programmable Logic Device), or a combination thereof. The aforementioned PLD can be CPLD (Complex Programmable Logic Device), FPGA (Field-Programmable Gate Array), GAL (Generic Array Logic), or any combination thereof.

[0073] Those skilled in the art should understand that the above-described voice interaction terminal 101, scene acquisition module 102, and processor 103 are merely examples. Other existing or future voice interaction terminals, scene acquisition modules, or processors that are applicable to the embodiments of this application should also be included within the scope of protection of the embodiments of this application, and are hereby incorporated by reference.

[0074] It should be noted that the implementation environment described in the embodiments of this application is for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and does not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, as the implementation environment evolves, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0075] The voice interaction method provided in the embodiments of this application will now be explained in detail.

[0076] Figure 2 This is a flowchart of a voice interaction method provided in an embodiment of this application, which is applied to the processor 103 described above. Please refer to... Figure 2 The method includes the following steps.

[0077] Step 201: Determine the driver's voice interaction database, which includes at least one voice triggering scenario and a voice generation rule set corresponding to the voice triggering scenario. The voice generation rule set includes at least one voice generation rule, with different voice generation rules corresponding to different voice interaction permissions and different voice content generated by different voice generation rules.

[0078] In some embodiments, the driver's voice interaction database can be determined based on the driver's identity information. For example, the driver's identity information can be determined individually or in combination using various forms such as facial features, voiceprint features, fingerprint features, and password sets, and then the driver's voice interaction database can be determined based on the driver's identity information.

[0079] In some embodiments, for any voice-triggered scenario, the voice generation rule set may include multiple voice generation rules. Different voice generation rules are used to indicate different levels of detail in the voice playback for that voice-triggered scenario. For example, if the voice-triggered scenario is receiving an SMS notification, the multiple voice generation rules corresponding to this scenario, in descending order of voice interaction permissions, may include: Voice generation rule A1: Received an SMS from <sender>, the content of which is <SMS content>; Voice generation rule A2: Received an SMS from <sender>; Voice generation rule A3: Received an SMS.

[0080] In some embodiments, the voice generation rule set corresponding to a voice-triggered scenario may include only one voice generation rule, i.e., no distinction is made regarding voice interaction permissions. For example, if the voice-triggered scenario is rain within two hours, the voice generation rule set corresponding to this scenario may only include voice generation rule B: "Weather type is expected to occur within <weather change forecast time>, please take precautions."

[0081] For voice-triggered scenarios and the corresponding voice generation rule sets, it can be the initial voice interaction database configured by the manufacturer when the vehicle leaves the factory, or it can be a voice interaction database that the driver customizes based on actual usage needs. For example, the voice interaction database can be obtained by the driver through the voice interaction settings interface by setting the voice-triggered scenarios and the voice generation rules for different voice interaction permissions corresponding to the voice-triggered scenarios. The specific settings can be combined with actual usage needs.

[0082] In some embodiments, a voice interaction settings interface may be displayed, which prompts the driver to set the at least one voice trigger scenario and the voice generation rules for different voice interaction permissions corresponding to the voice trigger scenario; the at least one voice trigger scenario and the voice generation rules for different voice interaction permissions corresponding to the voice trigger scenario are obtained from the voice interaction settings interface; based on the voice generation rules for different voice interaction permissions corresponding to the at least one voice trigger scenario, a voice generation rule set corresponding to the at least one voice trigger scenario is generated, and the at least one voice trigger scenario and the voice generation rule set corresponding to the voice trigger scenario are stored in the voice interaction database.

[0083] In some embodiments, the processor may include a voice interaction setting module, the voice interaction setting interface being the display interface of the voice interaction setting module, through which the driver can modify (such as add, delete, update, etc.) the voice triggering scenarios in the voice interaction database and the voice generation rule set corresponding to the voice triggering scenarios.

[0084] In some embodiments, before obtaining the at least one voice trigger scenario and the voice generation rules for different voice interaction permissions corresponding to the voice trigger scenario from the voice interaction settings interface, the processor also needs to obtain the driver's identity information to determine the voice interaction database corresponding to the driver based on the identity information. This allows the processor to store the voice trigger scenario and the corresponding voice generation rule set into the driver's voice interaction database after obtaining the voice trigger scenario and the voice generation rule set. The identity information can be set based on actual usage needs. For example, the identity information may include biometric information such as the driver's facial features, voiceprint features, and fingerprint features, or it may be set through password groups, such as gesture passwords or key passwords (which may include physical keys, virtual keys, etc.).

[0085] It should be noted that when prompting the driver to set voice trigger scenarios and voice generation rules for different voice interaction permissions corresponding to voice trigger scenarios through the voice interaction settings interface, the driver can be guided. For example, the driver can be first prompted to select a voice trigger scenario, and then set the voice generation rules corresponding to different voice interaction permissions for the voice trigger scenario. The specific order of prompting the voice trigger scenario, voice interaction permissions, and voice generation rules can be selected according to usage needs, or prompts can be given simultaneously, allowing the driver to flexibly choose the order of settings through the voice interaction settings interface.

[0086] The voice trigger scenario can include various scenario information, such as scenario information of different scenario types, which can be determined based on actual usage needs. For example, the voice trigger scenario can include various scenario types such as vehicle environment, vehicle status, and internet data. The vehicle environment can include the interior environment and the exterior environment. For example, the interior environment can include various environmental information such as the people in the vehicle, the interior temperature, the interior humidity, the interior odor, and the interior sound. The exterior environment can include various environmental information such as road conditions, vehicle status, exterior temperature, exterior humidity, current location, and weather. The vehicle status can include vehicle maintenance information (such as tire status, engine oil, oil filter, spark plugs, battery status, etc.) and the vehicle's current driving status (such as vehicle speed, fuel consumption, power consumption, and remaining driving range, etc.). Internet data can include driver trip information, communication information (such as SMS, phone calls, and emails from mobile terminals), and application push information (such as mobile terminal applications, in-vehicle computer terminal applications, etc.).

[0087] It is understood that the above scenarios are merely illustrative examples. With the development of technology and the enrichment of vehicle voice interaction usage scenarios, other existing or future scenarios that may be applicable to the embodiments of this application should also be included within the protection scope of the embodiments of this application, and are hereby incorporated by reference.

[0088] In some embodiments, the driver can flexibly select the trigger scenario based on one or more of the above specific scenarios. For example, the trigger scenario can be: the vehicle's current location is the target location and the current time is the target time.

[0089] Drivers can edit the voice generation rules themselves based on the voice interaction settings interface and store the voice generation rules in the voice generation rule set corresponding to the voice triggering scenario. Alternatively, they can select from multiple voice generation rules built into the vehicle and store them in the voice generation rule set corresponding to the voice triggering scenario.

[0090] In some embodiments, the voice interaction permission level may include three levels: high, medium and low, or four levels: first, second, third and fourth. The specific number of levels can be determined based on usage requirements, and the embodiments in this application are not limited in this respect.

[0091] Step 202: If the driving scenario of the target vehicle satisfies the target voice trigger scenario, obtain the voice generation rule set corresponding to the target voice trigger scenario from the voice interaction database. The target voice trigger scenario is any one of the at least one voice trigger scenario.

[0092] In some embodiments, the scene information of the current driving scenario of the target vehicle can be matched with the scene information of at least one voice-triggered scenario in the voice interaction database to determine whether the driving scenario of the target vehicle meets the target voice-triggered scenario.

[0093] Specifically, for any voice-triggered scenario, the voice-triggered scenario can be understood as a combination of at least one scenario condition information, and the driving scenario of the target vehicle can be understood as a combination of multiple scenario state information. Logical judgment is performed based on the multiple scenario state information and at least one scenario condition information in each voice-triggered scenario. If a certain scenario state satisfies a certain scenario condition, the judgment result of the scenario condition is determined to be "true"; otherwise, the judgment result of the scenario condition is determined to be "false". Furthermore, if the judgment results of all scenario conditions in the voice-triggered scenario are "true", then the driving scenario of the target vehicle is considered to satisfy the voice-triggered scenario; otherwise, the driving scenario of the target vehicle is considered not to satisfy the voice-triggered scenario.

[0094] For example, one voice trigger scenario in the voice interaction database is that the distance between the vehicle and the park is less than 500 meters, the current date is a non-working day (the criteria for determining a working day can be based on the driver's settings or the default calendar; here, Monday to Friday are taken as working days), and the weather is sunny. Then, based on the target vehicle's driving scenario, it is determined whether all the conditions of this voice trigger scenario are met. If the target vehicle is 300 meters from the nearest park, the current date is Saturday, and the weather is sunny, then the target vehicle's driving scenario is considered to meet the voice trigger scenario. If the target vehicle is 100 meters from the nearest park, the current date is Wednesday, and the current weather is cloudy, since the judgment results for both the date and weather conditions are "false," the target vehicle's driving scenario is considered not to meet the voice trigger scenario.

[0095] It should be noted that the driving scenario in this application embodiment is not the scenario of the vehicle being in motion, but can be understood as the usage scenario of the vehicle, such as the usage scenario after the vehicle is unlocked, after the car door is opened, or after the presence of a driver in the driver's seat is detected. The specific settings can be combined with actual usage needs.

[0096] In some embodiments, in order to improve the driver's user experience and the energy efficiency of the vehicle, the processor can be in a sleep state under normal circumstances. By setting an initial voice trigger scenario and an initial voice, when the driving scenario meets the initial voice trigger scenario, the processor guides the driver to activate the in-vehicle voice interaction function in this embodiment by playing the initial voice. The processor then determines the driver's identity information based on the voice command triggered by the driver, and then determines the driver's voice interaction database.

[0097] For example, the initial voice trigger scenario is when the vehicle is unlocked and a driver is present in the driver's seat. The initial voice is "Do you need to activate the voice assistant?" When the vehicle is unlocked and a driver is present, this initial voice is generated and sent. Upon receiving an affirmative response from the driver (such as "Open," "Okay," or "Sure"), the in-vehicle voice interaction function is activated. If no affirmative response is received from the driver, the processor remains in sleep mode. Upon receiving an affirmative response from the driver, the system can also identify the driver's voiceprint based on the driver's voice, or identify the driver's facial features through in-vehicle image acquisition equipment, guide the driver through fingerprint recognition, or guide the driver to unlock a password group, thereby determining the driver's identity information and establishing the driver's voice interaction database.

[0098] In some embodiments, the processor may also be in a always-on state, and when a driver is present in the driver's seat, it can determine the driver's voice interaction database and identify in real time whether the current vehicle driving scenario meets the voice trigger scenario.

[0099] Step 203: Determine the target voice interaction permission corresponding to the person in the vehicle, obtain the target voice generation rule corresponding to the target voice interaction permission from the voice generation rule set corresponding to the target voice triggering scenario, generate and play the voice based on the target voice generation rule and the scenario information of the driving scenario.

[0100] In some embodiments, the voice interaction permissions corresponding to each person in the vehicle can be obtained from the personnel permission database to obtain at least one voice interaction permission; the voice interaction permission with the lowest level among the at least one voice interaction permission is determined as the target voice interaction permission.

[0101] In some embodiments, the personnel permission database may include at least one identity information and voice interaction permissions corresponding to each identity information. Then, based on the identity information of the in-vehicle personnel, the voice interaction permissions corresponding to the in-vehicle personnel are obtained from the personnel permission database. The identity information of the in-vehicle personnel may include facial features and / or voiceprint features to determine the voice interaction permissions corresponding to each in-vehicle personnel without affecting the perception of the in-vehicle personnel.

[0102] For example, facial features of people inside the vehicle can be acquired using an image acquisition device, and voiceprint features of people inside the vehicle can be acquired using a sound acquisition device such as a microphone. The processor determines the identity information of people inside the vehicle based on the acquired facial features and / or voiceprint features, and then obtains the voice interaction permissions corresponding to the identity information from the personnel permission database.

[0103] Given the high degree of public visibility of vehicle usage scenarios, setting voice interaction permissions for every occupant would result in a large data volume in the occupant permission database, impacting data response speed. Therefore, in some embodiments, the occupant permission database may only include occupant information excluding the lowest level of voice interaction permission. When the occupant's corresponding voice interaction permission cannot be determined based on the occupant permission database (i.e., the database does not contain the occupant's identity information), the occupant's corresponding voice interaction permission is considered to be at the lowest level.

[0104] In some embodiments, a voice interaction settings interface may be displayed, which prompts the driver to set voice interaction permissions for different occupants in the vehicle; the voice interaction permissions for different occupants in the vehicle are obtained from the voice interaction settings interface and stored in the occupant permission database.

[0105] In some embodiments, the identity information of the occupants in the vehicle can be monitored to obtain the identity information of the occupants in the vehicle during each trip, and the identity information of the occupants in the vehicle can be sorted in descending order of the frequency of occurrence. When a setting request is received from the driver through the voice interaction setting interface, the sorting result is displayed on the voice interaction setting interface so that the driver can set the voice interaction permissions of the occupants based on the frequency of occurrence of different occupants.

[0106] In some embodiments, after storing the voice interaction permissions corresponding to different in-vehicle occupants in the occupant permission database, the updated voice interaction permissions corresponding to different in-vehicle occupants can be obtained from the voice interaction settings interface based on the driver's updated settings, and the voice interaction permissions corresponding to different in-vehicle occupants in the occupant permission database can be updated to realize the update of voice interaction permissions.

[0107] In some embodiments, before generating and playing voice based on the target voice generation rule and the scene information of the driving scenario, the driving state of the target vehicle can be obtained; based on the environment mapping relationship, the interaction environment corresponding to the driving state is determined, the interaction environment includes a safe environment or a non-safe environment, and the environment mapping relationship is used to indicate the interaction environment corresponding to different driving states; if the interaction environment is a safe environment, the step of generating and playing voice based on the target voice generation rule and the scene information of the driving scenario is performed.

[0108] If the target vehicle's driving scenario meets the target voice trigger scenario and the interaction scenario corresponding to the driving state is a non-safe environment, then playing the voice at this time would cause the driver's attention to be distracted and pose a threat to driving safety. Therefore, even if the target vehicle's driving scenario meets the target voice trigger scenario at this time, the step of generating and playing the voice based on the target voice generation rule and the scenario information of the driving scenario will not be executed.

[0109] It is understandable that if, after the interaction environment changes from an unsafe environment to a safe environment, the target vehicle's driving environment still meets the target voice trigger command, then the step of generating and playing voice based on the target voice generation rule and the scene information of the driving scenario is executed. Conversely, if, after the interaction environment changes from an unsafe environment to a safe environment, the target vehicle's driving environment no longer meets the target voice trigger scenario, then the step of generating and playing voice based on the target voice generation rule and the scene information of the driving scenario is not executed. In other words, if the target vehicle's driving scenario meets the target voice trigger scenario, but the current interaction environment is unsafe, then the step of generating and playing voice based on the target voice generation rule and the scene information of the driving scenario is skipped, and after the interaction environment becomes safe, the determination of whether the target voice trigger scenario is met is re-based on the target vehicle's driving scenario.

[0110] Taking the voice trigger scenario where the distance between the vehicle and the park is less than 500 meters as an example, if the driving scenario still meets the target voice trigger scenario after the interaction environment changes from an unsafe environment to a safe environment, that is, the distance between the vehicle and the park is still less than 500 meters, then the step of generating and playing the voice based on the target voice generation rule and the scene information of the driving scenario can be executed; if the driving scenario no longer meets the target voice trigger scenario after the interaction environment changes from an unsafe environment to a safe environment, that is, the distance between the vehicle and the park is greater than or equal to 500 meters, then the process ends until the driving scenario of the target vehicle meets the target voice trigger scenario again (such as the distance between the vehicle and the park is less than 500 meters again).

[0111] In some embodiments, parameters that have a significant impact on driving safety can be determined based on expert experience, statistical data, etc., thereby obtaining a driving state used to evaluate the safety of the interactive environment. For example, the driving state may include vehicle speed, weather, road conditions, vehicle condition, etc.

[0112] It is understandable that when a vehicle interacts with a driver via voice, it can cause the driver's attention to be distracted. The more complex the driving conditions, such as higher speeds, worse weather (such as rain, snow, or fog), more complex road conditions (such as one-way streets or muddy roads), and more complex traffic conditions (such as a large number of surrounding vehicles or pedestrians), the more the driver's attention needs to be focused. In such cases, voice interaction with the driver can lead to a distraction of the driver's attention and pose a threat to driving safety.

[0113] Therefore, in some embodiments, based on experimental statistical data, it can be determined whether voice interaction under different driving conditions has an impact on driving safety, that is, the interaction environment corresponding to different driving conditions, and obtain the environment mapping relationship.

[0114] For example, based on experimental statistical data, the interaction environment under different driving conditions (i.e., the interaction environment type under different vehicle speeds, different weather, different road conditions, and different vehicle conditions) can be determined. Then, based on the current driving state, the interaction environment corresponding to the current driving state can be obtained by looking up a table through the environment mapping relationship.

[0115] In some embodiments, in order to improve the accuracy of interactive environment recognition, the correspondence between driving state and interactive safety evaluation value can be determined based on experimental data, and then the interactive safety evaluation value corresponding to the current driving state can be determined based on the correspondence. The interactive safety evaluation value can be understood as the degree of influence of playing voice on driving safety under different driving conditions. The interactive environment corresponding to the current driving state can be determined by the relationship between the interactive safety evaluation value and the safety threshold.

[0116] For example, taking driving status including vehicle speed, weather, road conditions and vehicle condition as an example, the influence coefficients of vehicle speed, weather, road conditions and vehicle condition on interactive safety can be determined based on multiple sets of test data, such as multiple sets of driving status data when accidents occur, so as to obtain the corresponding relationship between driving status and interactive safety evaluation value.

[0117] It should be noted that since weather, road conditions, and vehicle conditions are all non-linear values, calibration values ​​can be used to determine the evaluation values ​​corresponding to different weather, road, and vehicle conditions. For example, the evaluation value of weather on driving safety can be determined based on relevant industry standards; the evaluation value of road conditions on driving safety can be determined based on information such as road surface smoothness and road technical standards, such as the evaluation values ​​corresponding to different road conditions based on information such as urban roads (including expressways, arterial roads, secondary arterial roads, and local roads) and highways (including expressways, first-class highways, second-class highways, third-class highways, and fourth-class highways); the evaluation value corresponding to vehicle conditions is determined based on the number of vehicles and pedestrians around the target vehicle. Then, based on the evaluation values ​​corresponding to the current weather, current road conditions, current vehicle conditions, and current vehicle speed, an interaction safety evaluation value is obtained. When the interaction safety evaluation value is greater than or equal to a safety threshold, the interaction environment corresponding to the current driving state is considered an unsafe environment; when the interaction safety evaluation value is less than the safety threshold, the interaction environment corresponding to the current driving state is considered a safe environment.

[0118] In some embodiments, the voice interaction database further includes a voice interaction priority corresponding to each voice triggering scenario. The target voice triggering scenario includes at least two voice triggering scenarios, and the target voice generation rule includes at least two voice generation rules. The at least two voice generation rules are obtained from the voice generation rule set corresponding to the at least two voice triggering scenarios according to the target voice interaction permission. The voice interaction priorities corresponding to the at least two voice triggering scenarios can also be obtained from the voice interaction database. At least two voices are generated based on the at least two voice generation rules and the scenario information. The at least two voices are played sequentially in descending order of the voice interaction priorities corresponding to the at least two voice triggering scenarios.

[0119] It should be noted that, due to the diverse types of vehicle driving scenarios, such as the vehicle environment, vehicle status, and internet data in the example above, if the voice interaction database includes multiple voice trigger scenarios, especially when there are multiple voice trigger scenarios with different aspects, the target vehicle's driving scenario may simultaneously satisfy multiple voice trigger scenarios. For example, voice trigger scenario A is when the vehicle's battery has a remaining range of 25km, and voice trigger scenario B is when the driver's mobile terminal receives a text message. When the target vehicle's driving scenario has a remaining battery range of 25km and the driver's mobile terminal receives a text message, then the current target vehicle's driving scenario simultaneously satisfies both voice trigger scenario A and voice trigger scenario B.

[0120] In the above situation, in order to avoid voice playback conflicts caused by simultaneously meeting multiple voice trigger scenarios, the voice interaction priority can be set using the above method. This allows multiple voices to be played sequentially based on the voice interaction priority when multiple voice trigger scenarios are met simultaneously.

[0121] In some embodiments, the driver can be prompted to set the voice interaction priority corresponding to the voice triggering scenario through the voice interaction settings interface. For example, when the driver is prompted to set the voice triggering scenario and the voice generation rule corresponding to the voice triggering scenario through the voice interaction settings interface, after the driver sets the voice triggering scenario and the voice generation rule corresponding to the voice triggering scenario, the driver still needs to set the voice interaction priority corresponding to the voice triggering scenario in order to complete the setting and store it in the voice interaction database.

[0122] In some embodiments, a voice playback interval can be set. The next voice message will only be played after the duration of the interval has elapsed following the completion of a single message, thus preventing continuous playback from negatively impacting the passenger experience. This voice playback interval can be set according to actual usage needs; for example, it can be set to a default of 3 seconds, but the driver can also personalize the setting through the voice interaction interface.

[0123] In some embodiments, after generating and playing voice based on the target voice generation rule and the scene information of the driving scenario, the system can also receive a function command triggered by the driver and execute the target function corresponding to the function command. The function command includes: a voice command or a button command.

[0124] It should be noted that the voice generation rules set by the driver can include interrogative voice generation rules and executive voice generation rules. Interrogative voice generation rules allow the execution of functions to be combined with the driver's intention to execute functions in different situations, allowing the driver to determine whether to execute, thereby improving the flexibility of function execution. Executive voice generation rules allow the target function to be executed directly when the conditions for function execution (driving scenario, interaction permissions) are met, thereby improving the efficiency of function execution.

[0125] For example, target voice trigger scenario A is that the driver's mobile terminal receives a text message. In the voice interaction database, the voice interaction permissions corresponding to the voice generation rules for this trigger scenario include three levels: high-level permission, medium-level permission, and low-level permission. If the target voice trigger scenario A is met in a driving scenario and the target voice interaction permission is high-level, then target voice generation rule A1 is: "Received a text message from <sender>, the content of which is <text message content>". If the target voice trigger scenario A is met in a driving scenario and the target voice interaction permission is medium-level, then target voice generation rule A2 is: "Received a text message from <sender>, do you need to play the text message content?". If the target voice trigger scenario A is met in a driving scenario and the target voice interaction permission is low-level, then target voice generation rule A3 is: "Received a text message, do you need to play the sender's name or the text message content?".

[0126] It is understandable that in the above examples, target speech generation rule A1 is an execution-type speech generation rule, which directly plays the sender's name and the content of the text message to notify the driver that the text message has been received; target speech generation rule A2 is an interrogation-type speech generation rule, which plays the sender's name to notify the driver that the text message has been received, and after playing the voice, it needs to receive a function command triggered by the driver (play the text message content) before it can execute the function of playing the text message content; target speech generation rule A3 is also an interrogation-type speech generation rule, which only plays the message that the text message has been received, and after playing the voice, it needs to receive a function command triggered by the driver (play the sender's name and play the text message content) before it can execute the function of playing the sender's name and playing the text message content.

[0127] In this embodiment, by setting up a voice interaction database, when the driving scenario of the target vehicle meets the target voice triggering scenario in the voice interaction database, the target voice triggering scenario and the target voice generation rule corresponding to the target voice interaction permission corresponding to the current occupant are determined. Voice is then generated and played based on the target voice generation rule and the scenario information of the driving scenario, thus achieving active voice playback without driver intervention. Furthermore, by setting voice interaction permissions for each occupant, the target voice interaction permission is determined based on their individual permissions, allowing for different voice content to be played depending on the current occupant, avoiding overly private playback and making the playback more flexible while protecting driver privacy. Moreover, considering that automatically triggered voice playback might affect driver safety, the driving status of the target vehicle is obtained before playback. If the interaction environment is determined to be a safe environment based on the driving status, voice is generated and played to avoid automatically played voice distracting the driver and threatening their safety. In addition, after the voice is generated and played, the driver's function execution intention is determined by monitoring the function commands triggered by the driver. The target function is executed when the driver's function commands are received, so that the executed target function can better match the driver's function execution intention in different situations, improve the driver's user experience, and thus improve the efficiency of voice interaction.

[0128] Figure 3 This is a schematic diagram of the structure of an in-vehicle voice interaction device provided in an embodiment of this application. The in-vehicle voice interaction device can be implemented by software, hardware, or a combination of both, forming part or all of an in-vehicle voice interaction equipment. The in-vehicle voice interaction equipment can be... Figure 1 The processor shown. Please refer to... Figure 3 The device includes: a database determination module 301, a rule set matching module 302, a voice playback module 303, and a function execution module 304.

[0129] The database determination module 301 is used to determine the driver's voice interaction database. The voice interaction database includes at least one voice triggering scenario and a voice generation rule set corresponding to the voice triggering scenario. The voice generation rule set includes at least one voice generation rule. Different voice generation rules correspond to different voice interaction permissions, and different voice generation rules generate different voice content.

[0130] The rule set matching module 302 is used to obtain the voice generation rule set corresponding to the target voice triggering scenario from the voice interaction database when the driving scenario of the target vehicle meets the target voice triggering scenario. The target voice triggering scenario is any one of the at least one voice triggering scenario.

[0131] The voice playback module 303 is used to determine the target voice interaction permission corresponding to the person in the vehicle, obtain the target voice generation rule corresponding to the target voice interaction permission from the voice generation rule set corresponding to the target voice triggering scenario, and generate and play voice based on the target voice generation rule and the scenario information of the driving scenario.

[0132] Optionally, the voice playback module 303 is also used for:

[0133] Obtain the driving status of the target vehicle;

[0134] Based on the environment mapping relationship, the interaction environment corresponding to the driving state is determined. The interaction environment includes a safe environment or a non-safe environment. The environment mapping relationship is used to indicate the interaction environment corresponding to different driving states.

[0135] If the interactive environment is a safe environment, the steps of generating and playing the voice based on the target voice generation rule and the scene information of the driving scenario are executed.

[0136] Optionally, the in-vehicle voice interaction device further includes a function execution module 304, which is used to receive the function command triggered by the driver and execute the target function corresponding to the function command. The function command includes: voice command or button command.

[0137] Optionally, the voice interaction database also includes a voice interaction priority corresponding to each voice triggering scenario. The target voice triggering scenario includes at least two voice triggering scenarios, and the target voice generation rule includes at least two voice generation rules. These at least two voice generation rules are obtained from the voice generation rule set corresponding to the at least two voice triggering scenarios according to the target voice interaction permission. The rule set matching module 302 is further configured to:

[0138] Retrieve the voice interaction priorities corresponding to the at least two voice triggering scenarios from the voice interaction database;

[0139] The voice playback module 303 is also used for:

[0140] Generate at least two speech entries based on the at least two speech generation rules and the scene information;

[0141] The at least two voice prompts are played sequentially in descending order of voice interaction priority corresponding to the at least two voice trigger scenarios.

[0142] Optionally, the database determination module 301 is also used for:

[0143] The voice interaction settings interface is displayed, which prompts the driver to set the voice generation rules for the at least one voice trigger scenario and the different voice interaction permissions corresponding to the voice trigger scenario.

[0144] Obtain the at least one voice trigger scenario and the voice generation rules for different voice interaction permissions corresponding to the voice trigger scenario from the voice interaction settings interface;

[0145] Based on the voice generation rules for different voice interaction permissions corresponding to the at least one voice triggering scenario, a voice generation rule set corresponding to the at least one voice triggering scenario is generated, and the at least one voice triggering scenario and the voice generation rule set corresponding to the voice triggering scenario are stored in the voice interaction database.

[0146] Optionally, the voice playback module 303 is specifically used for:

[0147] Retrieve the voice interaction permissions corresponding to each person in the vehicle from the personnel permission database, and obtain at least one voice interaction permission;

[0148] The lowest-level voice interaction permission among the at least one voice interaction permission is determined as the target voice interaction permission.

[0149] Optionally, the database determination module 301 is also used for:

[0150] The voice interaction settings interface is displayed, which prompts the driver to set voice interaction permissions for different people in the vehicle.

[0151] The voice interaction permissions corresponding to different people in the vehicle are obtained from the voice interaction settings interface and stored in the person's permission database.

[0152] In this embodiment, by setting up a voice interaction database, when the driving scenario of the target vehicle meets the target voice triggering scenario in the database, the target voice triggering scenario and the target voice generation rule corresponding to the target voice interaction permission corresponding to the current occupant are determined. Voice is then generated and played based on the target voice generation rule and the scenario information of the driving scenario, thus achieving active voice playback without driver intervention. Furthermore, by setting voice interaction permissions for each occupant, the target voice interaction permission is determined based on their individual permissions, allowing for different voice content to be played depending on the current occupant. This avoids overly private voice playback, making the playback more flexible and protecting driver privacy. Moreover, considering that automatically triggered voice playback might affect driver safety, the driving status of the target vehicle is obtained before playback. If the interaction environment is determined to be a safe environment based on the driving status, voice is generated and played to avoid automatically played voice distracting the driver and threatening their safety. In addition, after the voice is generated and played, the driver's function execution intention is determined by monitoring the function commands triggered by the driver. The target function is executed when the driver's function commands are received, so that the executed target function can better match the driver's function execution intention in different situations, improve the driver's user experience, and thus improve the efficiency of voice interaction.

[0153] It should be noted that the in-vehicle voice interaction device provided in the above embodiments is only illustrated by the division of the above functional modules when implementing in-vehicle voice interaction. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the in-vehicle voice interaction device and the in-vehicle voice interaction method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0154] Figure 4 This is a structural block diagram of a vehicle 400 provided in an embodiment of this application. Typically, the vehicle 400 includes a processor 401 and a memory 402.

[0155] Processor 401 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 401 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). Processor 401 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 401 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 401 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0156] The memory 402 may include one or more computer-readable storage media, which may be non-transitory. The memory 402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 402 are used to store at least one instruction, which is executed by the processor 401 to implement the in-vehicle voice interaction method provided in the method embodiments of this application.

[0157] In some embodiments, the vehicle 400 may also optionally include a peripheral device interface 403 and at least one peripheral device. The processor 401, memory 402, and peripheral device interface 403 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 403 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 404, a display screen 405, a camera assembly 406, an audio circuit 407, a positioning assembly 408, and a power supply 409.

[0158] Peripheral device interface 403 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 401 and memory 402. In some embodiments, processor 401, memory 402 and peripheral device interface 403 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 401, memory 402 and peripheral device interface 403 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0159] The radio frequency (RF) circuit 404 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 404 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 404 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 404 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 404 can communicate with other vehicles through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 404 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application embodiment.

[0160] Display screen 405 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 405 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 401 for processing. In this case, display screen 405 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, display screen 405 can be a single unit located inside the cabin of vehicle 400; in still other embodiments, display screen 405 can be a flexible display screen. Furthermore, display screen 405 can be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. Display screen 405 can be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0161] The camera assembly 406 is used to capture images or videos.

[0162] The audio circuit 407 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 401 for processing, or input to the radio frequency circuit 404 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, positioned at different locations within the vehicle 400. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 401 or the radio frequency circuit 404 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 407 may also include a headphone jack.

[0163] The positioning component 408 is used to locate the current geographical location of the vehicle 400 in order to enable navigation or LBS (Location Based Service). The positioning component 408 can be a positioning component of GPS (Global Positioning System), BeiDou system or Galileo system.

[0164] Power source 409 is used to supply power to various components in vehicle 400. Power source 409 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power source 409 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0165] Those skilled in the art will understand that Figure 4 The structure shown does not constitute a limitation on vehicle 400 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0166] In some embodiments, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of the in-vehicle voice interaction method described above. For example, the computer-readable storage medium may be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, or optical data storage device.

[0167] It is worth noting that the computer-readable storage medium mentioned in the embodiments of this application can be a non-volatile storage medium, in other words, it can be a non-transient storage medium.

[0168] It should be understood that all or part of the steps of the above embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented wholly or partially in the form of a computer program product. The computer program product includes one or more computer instructions. The computer instructions can be stored in the above-described computer-readable storage medium.

[0169] That is, in some embodiments, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform the steps of the in-vehicle voice interaction method described above.

[0170] It should be understood that "at least one" as mentioned herein refers to one or more, and "multiple" refers to two or more. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. In addition, in order to clearly describe the technical solutions of the embodiments of this application, the terms "first," "second," etc., are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or execution order, and the terms "first," "second," etc., are not necessarily different.

[0171] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in the embodiments of this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0172] The above descriptions are embodiments provided in this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A vehicle-mounted voice interaction method, characterized in that, The method includes: A voice interaction database for the driver is determined. The voice interaction database includes at least one voice triggering scenario, a voice generation rule set corresponding to the voice triggering scenario, and a voice interaction priority corresponding to each voice triggering scenario. The voice generation rule set includes at least one voice generation rule. Different voice generation rules correspond to different voice interaction permissions, and different voice generation rules generate different voice content. When the driving scenario of the target vehicle meets the target voice triggering scenario, the voice generation rule set corresponding to the target voice triggering scenario is obtained from the voice interaction database. The target voice triggering scenario includes at least two voice triggering scenarios. Determine the target voice interaction permission corresponding to the person in the vehicle, and obtain the target voice generation rule corresponding to the target voice interaction permission from the voice generation rule set corresponding to the at least two voice triggering scenarios. The target voice generation rule includes at least two voice generation rules. Obtain the voice interaction priorities corresponding to the at least two voice triggering scenarios from the voice interaction database; generate at least two voices based on the at least two voice generation rules and the scenario information of the driving scenario, and play the at least two voices in descending order of the voice interaction priorities corresponding to the at least two voice triggering scenarios.

2. The method as described in claim 1, characterized in that, The method further includes: Obtain the driving status of the target vehicle; Based on the environment mapping relationship, the interaction environment corresponding to the driving state is determined. The interaction environment includes a safe environment or a non-safe environment. The environment mapping relationship is used to indicate the interaction environment corresponding to different driving states. When the interaction environment is a safe environment, the steps are as follows: generate at least two voices based on the at least two voice generation rules and the scene information of the driving scenario, and play the at least two voices in descending order of voice interaction priority corresponding to the at least two voice trigger scenarios.

3. The method as described in claim 1 or 2, characterized in that, After playing the at least two voice messages sequentially according to their voice interaction priorities from highest to lowest for the at least two voice-triggered scenarios, the method further includes: The system receives a function command triggered by the driver and executes the target function corresponding to the function command. The function command includes: a voice command or a button command.

4. The method as described in claim 1 or 2, characterized in that, Before determining the driver's voice interaction database, the method further includes: The voice interaction settings interface is displayed, which is used to prompt the driver to set the at least one voice triggering scenario and the voice generation rules for different voice interaction permissions corresponding to the voice triggering scenario. Obtain the at least one voice triggering scenario and the voice generation rules for different voice interaction permissions corresponding to the voice triggering scenario from the voice interaction settings interface. Based on the voice generation rules for different voice interaction permissions corresponding to the at least one voice triggering scenario, a voice generation rule set corresponding to the at least one voice triggering scenario is generated, and the at least one voice triggering scenario and the voice generation rule set corresponding to the voice triggering scenario are stored in the voice interaction database.

5. The method as described in claim 1 or 2, characterized in that, The process of determining the target voice interaction permissions corresponding to the occupants of the vehicle includes: Retrieve the voice interaction permissions corresponding to each person in the vehicle from the personnel permission database, and obtain at least one voice interaction permission; The lowest-level voice interaction permission among the at least one voice interaction permission is determined as the target voice interaction permission.

6. The method as described in claim 5, characterized in that, Before determining the target voice interaction permissions corresponding to the occupants in the vehicle, the method further includes: The voice interaction settings interface is displayed, which is used to prompt the driver to set voice interaction permissions for different people in the vehicle. The voice interaction permissions corresponding to different in-vehicle personnel are obtained from the voice interaction settings interface and stored in the personnel permission database.

7. A vehicle-mounted voice interaction device, characterized in that, Applied to the in-vehicle voice interaction method as described in claim 1, the device comprises: The database determination module is used to determine the driver's voice interaction database. The voice interaction database includes at least one voice triggering scenario, a voice generation rule set corresponding to the voice triggering scenario, and a voice interaction priority corresponding to each voice triggering scenario. The voice generation rule set includes at least one voice generation rule. Different voice generation rules correspond to different voice interaction permissions, and different voice generation rules generate different voice content. The rule set matching module is used to obtain the voice generation rule set corresponding to the target voice triggering scenario from the voice interaction database when the driving scenario of the target vehicle meets the target voice triggering scenario. The target voice triggering scenario includes at least two voice triggering scenarios. The voice playback module is used to determine the target voice interaction permissions corresponding to the occupants in the vehicle, obtain the target voice generation rule corresponding to the target voice interaction permission from the voice generation rule set corresponding to the at least two voice triggering scenarios, the target voice generation rule including at least two voice generation rules; obtain the voice interaction priority corresponding to the at least two voice triggering scenarios from the voice interaction database; generate at least two voices based on the at least two voice generation rules and the scenario information of the driving scenario, and play the at least two voices in descending order of voice interaction priority corresponding to the at least two voice triggering scenarios.

8. A vehicle, characterized in that, The vehicle includes a memory and a processor, the memory being used to store a computer program, and the processor being used to execute the computer program stored in the memory to implement the steps of the method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the steps of the method described in any one of claims 1-6.

Citation Information

Patent Citations

  • Vehicle voice control method, device and equipment, vehicle and storage medium

    CN111653277A

  • Error correction interaction method and system based on vehicle-mounted voice scene and vehicle

    CN115294976A