Vehicle voice interaction method and apparatus, electronic device, storage medium, and vehicle

By identifying the emotional needs of users and basic functional needs and generating stylized natural language templates, the problem of low anthropomorphism in vehicle-mounted voice interaction is solved, the degree of anthropomorphism of voice interaction and user emotional comfort is improved, and resource consumption is reduced.

WO2025148303A1PCT designated stage expired Publication Date: 2025-07-17CHINA FAW CO LTD +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/111142
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-08
Filing Date
2024-08-09
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

The existing vehicle-mounted voice interaction methods lack the "contextual emotion" modification in language structure, resulting in low degree of anthropomorphism of voice interactions, which can easily trigger the "uncanny valley" effect and affect user experience.

Method used

By obtaining user voice data, identifying emotional needs and basic functional needs, generating emotional natural language prompt projects, and conducting AIGC model training, generating stylized natural language templates, and finally generating voice interactive copy that adapts to user emotional needs and basic functional needs.

Benefits of technology

It improves the degree of anthropomorphism of voice interaction, enhances the user's emotional comfort, reduces the consumption of large model resources, and improves compatibility and adaptability in different vehicle application contexts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024111142_17072025_PF_FP_ABST
    Figure CN2024111142_17072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses a vehicle voice interaction method, a vehicle voice interaction apparatus, an electronic device, a storage medium, and a vehicle. The method comprises: acquiring voice data of a user; analyzing the voice data, and identifying a user emotional demand and a basic function demand; on the basis of the user emotional demand and the basic function demand, generating an emotional natural language prompt project; on the basis of the emotional natural language prompt project, performing AIGC model training to generate a stylized natural language template; and on the basis of the stylized natural language template, generating a voice interaction copywriting corresponding to the user emotional demand and the basic function demand. By means of the solution, the stylized natural language template is generated on the basis of identifying the user emotional demand and the basic function demand, and the voice interaction copywriting is efficiently generated on the basis of the stylized natural language template, thereby improving the personification of voice interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Vehicle voice interaction method, device, electronic device, storage medium and vehicle Technical Field

[0001] The present application relates to the field of voice interaction, and in particular to a vehicle voice interaction method, a vehicle voice interaction device, an electronic device, a storage medium, and a vehicle. Background Art

[0002] Existing in-vehicle voice interaction methods mostly extract slot information. For example, a user asks, "Turn on the air conditioner?" and the car computer responds, "The air conditioner is on, and the current temperature is 'temperature' degrees" (where "temperature" refers to the slot information). AIGC, or large models, offer significant advantages in semantic understanding and dialogue generation, enabling relatively accurate extraction of slot information and enabling corresponding control tasks.

[0003] The above method lacks "contextual emotion". Although its "contextual emotion" can be adjusted through intonation, timbre, tone, and even pre-recording, it lacks the "language style" modification of the language structure, which makes the stereotyped feeling of voice interaction abrupt. With the advancement of technical routes such as intonation, timbre, tone, and even pre-recording, it gradually approaches the "uncanny valley", which in turn suppresses the user's motivation for the anthropomorphic demand for voice interaction.

[0004] Therefore, a vehicle voice interaction solution is needed that can achieve highly real-time "contextual emotion" voice interaction without consuming too many large model resources, overcome the "uncanny valley" effect caused by the low degree of anthropomorphic realism, and improve the emotional comfort of voice interaction control of vehicle-mounted equipment.

[0005] Summary of the Invention

[0006] The object of the present invention is to provide a vehicle voice interaction method, a vehicle voice interaction device, an electronic device, a storage medium and a vehicle, which at least solve one of the above-mentioned technical problems.

[0007] The present invention provides the following solutions:

[0008] According to one aspect of the present invention, a vehicle voice interaction method is provided, the vehicle voice interaction method comprising:

[0009] Get user voice data;

[0010] Analyze the voice data to identify the user's emotional needs and basic functional needs;

[0011] Generate an emotional natural language prompt project based on the user's emotional needs and basic functional requirements;

[0012] According to the emotional natural language prompting project, AIGC model training is performed to generate stylized natural language templates;

[0013] Based on the stylized natural language template, a voice interaction copy corresponding to the user's emotional needs and basic functional needs is generated.

[0014] Furthermore, generating a voice interaction text corresponding to the user's emotional needs and basic functional needs based on the stylized natural language template includes:

[0015] The voice data corresponds to an in-vehicle application context;

[0016] Analyze the voice data according to the in-vehicle application context to identify the user's emotional needs and basic functional needs;

[0017] Generate an emotional natural language prompt project based on the user's emotional needs, basic functional requirements and the in-vehicle application context;

[0018] According to the emotional natural language prompting project, AIGC model training is performed to generate stylized natural language templates;

[0019] According to the stylized natural language template corresponding to the in-vehicle application context, a voice interaction copy corresponding to the user's emotional needs and basic functional needs is generated.

[0020] Furthermore, generating a voice interaction text corresponding to the user's emotional needs and basic functional needs based on the stylized natural language template further includes:

[0021] Identify user emotional needs and basic functional requirements based on local user voice data;

[0022] Analyze the voice data based on the in-vehicle application context of the local user to identify the user's emotional needs and basic functional needs;

[0023] Generate an emotional natural language prompt project based on the user's emotional needs, basic functional needs, and the in-vehicle application contexts of multiple local end users;

[0024] According to the emotional natural language prompting project, AIGC model training is performed to generate cloud-based stylized natural language templates;

[0025] Synchronizing a local stylized natural language template according to the cloud-based stylized natural language template;

[0026] Based on the local stylized natural language template, generate voice interaction copy corresponding to the user's emotional needs and basic functional needs.

[0027] Furthermore, generating a voice interaction text corresponding to the user's emotional needs and basic functional needs based on the stylized natural language template further includes:

[0028] Obtain user feature information;

[0029] Based on the user feature information, obtaining user voice data;

[0030] Analyzing the user's voice data based on the user's characteristic information to identify the user's emotional needs and basic functional needs;

[0031] Generate an emotional natural language prompt project based on the user's emotional needs and basic functional needs corresponding to the user characteristics;

[0032] According to the emotional natural language prompting project, AIGC model training is performed to generate stylized natural language templates;

[0033] Based on the stylized natural language template, a voice interaction copy corresponding to the user's emotional needs and basic functional needs is generated.

[0034] Furthermore, analyzing the voice data and identifying the user's emotional needs and basic functional needs includes:

[0035] Generate TTS document information based on the user voice data;

[0036] According to the TTS document information, the slot information of the user's basic functional requirements and the slot information of the user's emotional requirements are obtained;

[0037] According to the slot information of the basic functional requirements, the slot information of the user's emotional requirements and the stylized natural language template, a voice interaction copy corresponding to the user's emotional requirements and basic functional requirements is generated.

[0038] Furthermore, the performing of AIGC model training according to the emotional natural language prompting project to generate a stylized natural language template includes:

[0039] Generating a stylized natural language template library according to the stylized natural language template;

[0040] The stylized natural language template library includes TTS text corresponding to the stylized natural language template;

[0041] Outputting a TTS text corresponding to the stylized natural language template according to the emotional natural language prompting project;

[0042] Determining, based on the slot information and the TTS text, whether the output TTS text corresponds to the user's emotional needs;

[0043] If the output TTS text corresponds to the user's emotional needs, the stylized natural language template is activated according to the TTS text.

[0044] According to two aspects of the present invention, a vehicle voice interaction device is provided, the vehicle voice interaction device comprising:

[0045] User data module, used to obtain user voice data;

[0046] A data analysis module is used to analyze the voice data and identify the user's emotional needs and basic functional needs;

[0047] A prompt engineering module is used to generate an emotional natural language prompt engineering according to the user's emotional needs and basic functional requirements;

[0048] A template generation module is used to perform AIGC model training based on the emotional natural language prompting project to generate a stylized natural language template;

[0049] The interactive copy module is used to generate voice interactive copy corresponding to the user's emotional needs and basic functional needs based on the stylized natural language template.

[0050] According to three aspects of the present invention, there is provided an electronic device, comprising: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;

[0051] A computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the vehicle voice interaction method.

[0052] According to four aspects of the present invention, a computer-readable storage medium is provided, comprising: a computer program that can be executed by an electronic device is stored therein, and when the computer program runs on the electronic device, the electronic device executes the steps of the vehicle voice interaction method.

[0053] According to five aspects of the present invention, there is provided a vehicle comprising:

[0054] An electronic device for implementing the steps of the vehicle voice interaction method;

[0055] a processor, the processor running a program, and executing the steps of the vehicle voice interaction method based on data output by the electronic device when the program is running;

[0056] The storage medium is used to store a program, and when the program is running, it executes the steps of the vehicle voice interaction method for data output from the electronic device.

[0057] Through the above solution, the following beneficial technical effects are achieved:

[0058] This application generates stylized natural language templates based on identifying users' emotional needs and basic functional needs, and efficiently generates voice interaction copy based on the stylized natural language templates, thereby improving the degree of humanization of voice interaction.

[0059] This application generates a stylized natural language template based on the user's characteristics, so that the stylized natural language template can adapt to users with different characteristics, thereby enriching the adaptability and uniqueness of the stylized natural language template.

[0060] This application uses the AIGC model to process sample data of speech in the context of in-vehicle applications, train stylized natural language templates for voice interaction, and iterate the stylized natural language templates based on the sample data of speech in the context of in-vehicle applications. Compared with methods relying on timbre, tone, or even pre-recording, the degree of humanization of voice interaction is rapidly improved.

[0061] This application reduces the frequency of cloud data upload and download and reduces the consumption of channel resources for AIGC model training of stylized natural language templates by training a stylized natural language template library and then updating the local stylized natural language template library through the stylized natural language template library.

[0062] This application distinguishes different in-vehicle application contexts, so that the trained stylized natural language templates are automatically adjusted according to the different in-vehicle application contexts, thereby improving the compatibility of stylized natural language templates in different in-vehicle application contexts and improving the emotional comfort of users when switching between different in-vehicle application contexts. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] FIG1 is a flow chart of a vehicle voice interaction method provided by one or more embodiments of the present invention.

[0064] FIG2 is a structural diagram of a vehicle voice interaction device provided by one or more embodiments of the present invention.

[0065] FIG3 is a flow chart of a vehicle voice interaction method according to a specific embodiment of the present invention.

[0066] FIG4 is a structural diagram of a vehicle voice interaction system according to a specific embodiment of the present invention.

[0067] FIG5 is a structural diagram of a vehicle voice interaction device according to a specific embodiment of the present invention.

[0068] FIG6 is a schematic diagram showing a comparison of voice interaction improvements according to a specific embodiment of the present invention.

[0069] FIG7 is a schematic diagram of a TTS text library according to a specific embodiment of the present invention.

[0070] FIG8 is a schematic diagram of updating a local stylized natural language template library according to a specific embodiment of the present invention.

[0071] FIG9 is a schematic diagram of a large model return example according to a specific embodiment of the present invention.

[0072] FIG10 is a schematic diagram of a speech training architecture according to a specific embodiment of the present invention.

[0073] FIG11 is a block diagram of an electronic device structure of a vehicle voice interaction method provided by one or more embodiments of the present invention. DETAILED DESCRIPTION

[0074] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0075] FIG1 is a flow chart of a vehicle voice interaction method provided by one or more embodiments of the present invention.

[0076] The vehicle voice interaction method shown in FIG1 includes:

[0077] Step S1, obtaining user voice data;

[0078] Step S2: parsing voice data to identify user emotional needs and basic functional needs;

[0079] Step S3: Generate an emotional natural language prompt project based on the user's emotional needs and basic functional requirements;

[0080] Step S4: Perform AIGC model training based on the emotional natural language prompting project to generate a stylized natural language template;

[0081] Step S5: Generate voice interaction text corresponding to the user's emotional needs and basic functional needs based on the stylized natural language template.

[0082] Specifically, voice recognition not only identifies users' basic functional needs but also their need for emotional convenience. For example, when controlling in-vehicle equipment through voice interaction, the control instructions for the in-vehicle equipment also include the emotional needs of the voice interaction. An emotional natural language prompting project is generated based on users' emotional needs and basic functional needs. The emotional natural language prompting project is an automated language processing project used to extract users' emotional needs and basic functional needs. The AIGC model inputs user emotional needs and basic functional requirements information and generates stylized natural language templates based on model training. The stylized natural language module performs screening and processing based on a pre-established template library. The stylized natural language module reserves semantic slots, extracts semantic slot information from the voice text, and combines it with the stylized natural language template to generate voice interaction copy corresponding to the user's emotional needs and basic functional needs.

[0083] In this embodiment, generating voice interaction text corresponding to the user's emotional needs and basic functional needs based on the stylized natural language template includes:

[0084] Voice data corresponds to the in-vehicle application context;

[0085] Analyze voice data based on the in-vehicle application context to identify users' emotional needs and basic functional requirements;

[0086] Generate emotional natural language prompts based on user emotional needs, basic functional requirements, and in-vehicle application context;

[0087] Based on the emotional natural language prompt project, AIGC model training is performed to generate stylized natural language templates;

[0088] Based on the stylized natural language template corresponding to the in-vehicle application context, voice interaction copy corresponding to the user's emotional needs and basic functional needs is generated.

[0089] Specifically, based on different in-vehicle application contexts, the semantics expressed by the same voice are different. According to the in-vehicle application context, voice data is analyzed to identify the user's emotional needs and basic functional needs. This is the voice data analysis corresponding to the current in-vehicle application context. Based on the analysis results, an emotional natural language prompt project is generated, and the AIGC model is trained to generate a stylized natural language template; the reserved semantic slots on the stylized natural language template are filled with the semantic slot information extracted from the voice text in the same in-vehicle application context and combined on the stylized natural language template to generate voice interaction copy corresponding to the user's emotional needs and basic functional needs, which also corresponds to the in-vehicle application context.

[0090] In this embodiment, generating voice interaction text corresponding to the user's emotional needs and basic functional needs based on the stylized natural language template also includes:

[0091] Identify user emotional needs and basic functional requirements based on local user voice data;

[0092] Analyze voice data based on the local user's in-vehicle application context to identify the user's emotional needs and basic functional requirements;

[0093] Generate emotional natural language prompts based on user emotional needs, basic functional requirements, and the in-vehicle application context of multiple local users;

[0094] Based on the emotional natural language prompt project, AIGC model training is performed to generate cloud-based stylized natural language templates;

[0095] Synchronize local stylized natural language templates based on cloud-based stylized natural language templates;

[0096] Based on local stylized natural language templates, generate voice interaction copy corresponding to users' emotional needs and basic functional requirements.

[0097] Specifically, there are multiple local terminals, and users at different local terminals have different emotional needs and basic functional needs. Moreover, each local terminal is in a different in-vehicle application context. Based on the above premise, the voice data is parsed to identify the user's emotional needs and basic functional needs. The collected data is used to generate an emotional natural language prompt project. According to the emotional natural language prompt project, the AIGC model is trained to generate a cloud-based stylized natural language template. The data input from multiple local terminals is aggregated to generate a cloud-based stylized natural language template. As the cloud-based stylized natural language template is updated, the local stylized natural language template is synchronized, so that the localized stylized natural language module obtains a rich set of stylized natural language templates. Based on the richer local stylized natural language templates, voice interaction copywriting corresponding to the user's emotional needs and basic functional needs is generated. The differences between the templates can be compared more finely, and a more suitable voice interaction copywriting can be selected.

[0098] In this embodiment, generating voice interaction text corresponding to the user's emotional needs and basic functional needs based on the stylized natural language template also includes:

[0099] Obtain user feature information;

[0100] Based on user feature information, obtain user voice data;

[0101] Analyze user voice data based on user feature information to identify user emotional needs and basic functional requirements;

[0102] Generate emotional natural language prompts based on user characteristics corresponding to user emotional needs and basic functional needs;

[0103] Based on the emotional natural language prompt project, AIGC model training is performed to generate stylized natural language templates;

[0104] Based on stylized natural language templates, generate voice interaction copy corresponding to users' emotional needs and basic functional requirements.

[0105] Specifically, due to different user characteristics, the resulting emotional needs and basic functional needs differ slightly. For example, age, gender, and attire lead to different basic functional needs for air conditioning, and the resulting emotional needs also differ. For example, older men typically wear thicker clothing in cold weather, which can lead to slower reactions. The basic functional needs for car air conditioning are relatively sensitive to very low temperatures, leading to emotional needs that favor clear, accurate, and patient "gentle" voice interactions, rather than "middle school" voice interactions with a lot of "slang."

[0106] In this embodiment, parsing voice data and identifying user emotional needs and basic functional needs includes:

[0107] Generate TTS document information based on the user voice data;

[0108] Based on the TTS document information, obtain the slot information of the user's basic functional needs and the slot information of the user's emotional needs;

[0109] Based on the slot information of basic functional requirements, the slot information of user emotional needs and the stylized natural language template, voice interaction copy corresponding to user emotional needs and basic functional needs is generated.

[0110] Specifically, voice data is converted into TTS document information, which can be cut, compared, and spliced. Based on the slot information for basic functional requirements and the slot information for user emotional needs, the selected stylized natural language template is used to find the slots where empty slots are needed for splicing. Voice interaction copy corresponding to the user's emotional needs and basic functional requirements is generated. Voice interaction is then performed, giving the interactive voice a style such as "middle school" or "gentle", meeting both basic functional requirements and emotional needs.

[0111] In this embodiment, AIGC model training is performed based on the emotional natural language prompting project to generate stylized natural language templates, including:

[0112] generating a stylized natural language template library according to the stylized natural language template;

[0113] The stylized natural language template library includes TTS texts corresponding to stylized natural language templates;

[0114] Based on the emotional natural language prompt project, output TTS copy corresponding to the stylized natural language template;

[0115] Based on the slot information and TTS copy, determine whether the output TTS copy meets the user's emotional needs;

[0116] If the output TTS text corresponds to the user's emotional needs, the stylized natural language template is activated according to the TTS text.

[0117] Specifically, stylized natural language templates are used to generate a stylized natural language template library, wherein the stylized natural language template library generates corresponding TTS texts, such as "whether it is a second-order interaction mode" in the TTS text, which is output to the human-computer interaction terminal. By obtaining user feedback (determining whether the output TTS text corresponds to the user's emotional needs), the stylized natural language templates in the stylized natural language template library are activated to generate voice interaction texts.

[0118] FIG2 is a structural diagram of a vehicle voice interaction device provided by one or more embodiments of the present invention.

[0119] The vehicle voice interaction device shown in FIG2 includes: a user data module, a data analysis module, a prompt engineering module, a template generation module, and an interactive copywriting module;

[0120] User data module, used to obtain user voice data;

[0121] Data analysis module, used to analyze voice data and identify users' emotional needs and basic functional requirements;

[0122] The prompt engineering module is used to generate emotional natural language prompt engineering based on user emotional needs and basic functional requirements;

[0123] The template generation module is used to train the AIGC model based on the emotional natural language prompt project and generate stylized natural language templates;

[0124] The interactive copywriting module is used to generate voice interaction copywriting corresponding to users' emotional needs and basic functional requirements based on stylized natural language templates.

[0125] It is worth noting that although this system only discloses the user data module, data analysis module, prompt engineering module, template generation module, and interactive copy module, relatively speaking, what the present invention wants to express is that, based on the above-mentioned basic functional modules, those skilled in the art can arbitrarily add one or more functional modules in combination with the existing technology to form an infinite number of embodiments or technical solutions. In other words, this system is open rather than closed. Just because this embodiment only discloses individual basic functional modules, it cannot be considered that the scope of protection of the claims of the present invention is limited to the above-mentioned basic functional modules.

[0126] Through the above solution, the following beneficial technical effects are achieved:

[0127] This application generates stylized natural language templates based on identifying users' emotional needs and basic functional needs, and efficiently generates voice interaction copy based on the stylized natural language templates, thereby improving the degree of humanization of voice interaction.

[0128] This application generates a stylized natural language template based on the user's characteristics, so that the stylized natural language template can adapt to users with different characteristics, thereby enriching the adaptability and uniqueness of the stylized natural language template.

[0129] This application uses the AIGC model to process sample data of speech in the context of in-vehicle applications, train stylized natural language templates for voice interaction, and iterate the stylized natural language templates based on the sample data of speech in the context of in-vehicle applications. Compared with methods relying on timbre, tone, or even pre-recording, the degree of humanization of voice interaction is rapidly improved.

[0130] This application reduces the frequency of cloud data upload and download and reduces the consumption of channel resources for AIGC model training of stylized natural language templates by training a stylized natural language template library and then updating the local stylized natural language template library through the stylized natural language template library.

[0131] This application distinguishes different in-vehicle application contexts, so that the trained stylized natural language templates are automatically adjusted according to the different in-vehicle application contexts, thereby improving the compatibility of stylized natural language templates in different in-vehicle application contexts and improving the emotional comfort of users when switching between different in-vehicle application contexts.

[0132] FIG3 is a flow chart of a vehicle voice interaction method according to a specific embodiment of the present invention.

[0133] FIG4 is a structural diagram of a vehicle voice interaction system according to a specific embodiment of the present invention.

[0134] FIG5 is a structural diagram of a vehicle voice interaction device according to a specific embodiment of the present invention.

[0135] In a specific embodiment, the vehicle voice interaction method shown in FIG3 includes:

[0136] Step S11, obtaining voice data corresponding to the in-vehicle application context;

[0137] Step S12, identifying the semantics of the in-vehicle application context speech;

[0138] Step S13: training a stylized natural language template for voice interaction based on the semantics of the voice in the in-vehicle application context;

[0139] Step S14: modifying the language style of the feedback speech in the voice interaction according to the stylized natural language template.

[0140] Specifically, the car computer is equipped with multiple in-car applications (e.g., software applications including APPs), targeting different application environments, such as in-car applications for controlling music, in-car applications for controlling air conditioning, and in-car applications for controlling navigation. When different in-car applications are activated, the semantics of the speech in the context of the in-car applications are identified. For example, in the air conditioning application, the user issues a voice command of "a little louder". Through the context of the air conditioning application, the voice command is interpreted as the semantics of increasing the cooling air of the air conditioner. Based on the semantics of the speech in the context of the in-car application, stylized natural language templates for voice interaction are trained, such as friendly, gentle, and middle school language styles. In addition to the intonation, timbre, and tone as features of the template, the stylized natural language templates are also modified or defined in terms of grammatical structure and word formation. During the voice interaction with the in-car application, the stylized natural language templates are used to modify the semantics of the speech when responding to the user's speech output. For example, the execution of feedback or commands is converted into the semantics of TTS text, which is then modified with a stylized natural language template. The modified TTS text is then converted into speech, and the speech is further modified in terms of intonation, timbre, and tone to match the wording and grammatical structure of the TTS text language style.

[0141] In this embodiment, it also includes:

[0142] Based on the voice data of the in-vehicle application context, we train stylized natural language templates for voice interaction corresponding to the in-vehicle application context.

[0143] Acquire voice data of an in-vehicle application context of one or more vehicles based on a stylized natural language template for voice interaction;

[0144] Iterate the training of stylized natural language templates for voice interaction based on voice data in the context of in-vehicle applications.

[0145] Specifically, the vehicle has one or more in-vehicle applications, and the user switches between different in-vehicle applications according to the needs of use. In addition to whether the software APP (in-vehicle application) is working in the foreground, it also includes the situation where multiple in-vehicle applications are in the current interactive state at the same time. At this time, the voice data of the in-vehicle application context of one or more vehicles is obtained, and the stylized natural language template of voice interaction is trained to adapt to the in-vehicle application context where multiple in-vehicle applications are online at the same time. Based on the voice data of the in-vehicle application context, the training of the stylized natural language template of voice interaction is iterated to make the language style have the characteristics of coherence and continuity, or give people the feeling of multiple interactive objects, or give people the feeling of one interactive object. By adjusting the language style to the in-vehicle application context, the user's voice interaction process is flexibly and accurately promoted.

[0146] For example, if the air conditioning app and music app are currently open and the user issues a voice command "Turn up the volume," the car computer can respond with a voice response, using a pre-selected "middle school" voice style template: "Old man, turn up the volume. Are you trying to freeze yourself to death or shock yourself to death?" The user then issues a voice command: "What do you think?" The car computer continues with the pre-selected "middle school" voice style template: "I think you're annoying me. I'll turn up the volume so you don't have to hear my nonsense." The music app's control interface is then activated, turning up the volume.

[0147] During the training of stylized natural language templates for iterative voice interaction, the user's input voice and historical interaction data are parsed to obtain the user's semantic intent. For example, sample data is collected regarding the control object preferences generated by different tones of voice under the same user voice command. For example, sample data is collected regarding the in-vehicle application's subsequent responses to user acceptance and rejection of a voice command issued by the user in the same vehicle environment.

[0148] In another specific embodiment, the vehicle voice interaction system shown in FIG2 includes:

[0149] Cloud systems and local systems;

[0150] A cloud-based system for recognizing the semantics of speech data and training a library of stylized natural language templates based on the semantic data;

[0151] A local system for voice interaction based on a library of local stylized natural language templates and the context of the in-vehicle application;

[0152] The local system performs voice interaction based on the context of the in-vehicle application, obtains voice data corresponding to the context of the in-vehicle application, and sends it;

[0153] The cloud system receives voice data corresponding to the in-vehicle application context, recognizes the semantics of the voice data, and trains a stylized natural language template library based on the semantic data, and sends it;

[0154] The local system receives data from the cloud system and updates the data of the local stylized natural language template library based on the data of the stylized natural language template library.

[0155] Specifically, the vehicle voice interaction system mainly includes a cloud system and a local system. The local system is located at one end of the vehicle, and the cloud system is located at one end of the cloud server, and is connected through wireless platforms such as V2X.

[0156] The local stylized natural language template library is a subset of the stylized natural language template library and updates data from the stylized natural language template library.

[0157] The cloud-based system recognizes the semantics of speech data and uses it to train a stylized natural language template library. The speech data comes from real-time human-computer interaction speech data from the local system. The cloud-based system first parses the speech into semantics and then translates the semantics into text-to-speech (TTS) text. Based on this TTS text, the stylized natural language template library is trained to obtain speech interaction templates with linguistic style.

[0158] In this embodiment, the local system includes: a dialogue management module, a language generation module, a local stylized natural language template library, and an in-vehicle application module;

[0159] In-vehicle application module, used to open the port for voice interaction control of in-vehicle applications;

[0160] The dialogue management module is used to convert voice data, generate control instructions or obtain feedback data, and control voice interaction on the in-vehicle application port;

[0161] A local stylized natural language template library, used to provide voice interaction templates with language styles;

[0162] The language generation module is used to process the voice feedback in voice interaction to conform to the preset language style;

[0163] The vehicle application module opens the port for voice interaction control of vehicle applications;

[0164] The dialogue management module obtains the voice interaction data of the vehicle application module as the voice interaction object;

[0165] The dialogue management module converts voice data, generates control instructions or obtains feedback data, and controls the voice interaction of the in-vehicle application port;

[0166] Among them, the dialogue management module sends the voice data fed back in the voice interaction;

[0167] The language generation module receives data from the dialogue management module and processes the speech feedback during the voice interaction;

[0168] Processing the voice feedback in voice interaction includes calling the voice interaction template from the local stylized natural language template library;

[0169] A voice interaction template that conforms to a preset language style is provided based on a local stylized natural language template library, and the language generation module sends voice feedback in the voice interaction that conforms to the preset language style.

[0170] Specifically, the local system on the vehicle side constitutes the actual in-vehicle application scenario, and the in-vehicle application module opens the port for voice interaction to control the in-vehicle application, that is, the control port for the in-vehicle application through voice interaction.

[0171] The dialogue management module converts voice data, generates control instructions or obtains feedback data, and controls the voice interaction in the vehicle application port, including the use of voice interaction templates for voice feedback during the interaction process.

[0172] In this embodiment, the cloud system includes: a cloud-based speech semantic recognition engine and a stylized natural language template library;

[0173] Cloud-based speech semantic recognition engine, used to identify the semantics of speech data;

[0174] Stylized natural language template library for generating voice interaction templates;

[0175] The cloud-based speech semantic recognition engine receives data from the dialogue management module and identifies the semantics of the speech data;

[0176] Stylized natural language template library, generating voice interaction templates and sending them;

[0177] The local stylized natural language template library receives the stylized natural language template library data, generates a voice interaction template, and updates the voice interaction template of the local stylized natural language template library.

[0178] Specifically, the cloud-based system includes a cloud-based speech semantic recognition engine and a stylized natural language template library. The cloud-based speech semantic recognition engine receives data from the dialogue management module, namely, user voice data from the vehicle side, and generates voice feedback during the voice interaction process based on the semantics of the recognized voice data.

[0179] In this embodiment, the cloud system further includes: a large model engine;

[0180] Large model engine for semantic language style training;

[0181] The large model engine obtains semantic data from the cloud-based speech and semantic recognition engine;

[0182] Based on semantic data, conduct semantic language style training and iteratively stylize the voice interaction templates of the natural language template library;

[0183] Among them, the semantic data of the cloud-based speech semantic recognition engine comes from the voice interaction of voice feedback under the voice interaction template.

[0184] Specifically, when the user's voice interaction process is affected by the voice interaction template generated using the TTS copy library, the language style training is adjusted based on the preset degree of impact on the user's voice interaction. For example, if a "zhongerbi" style is preset to adjust the humor of the interaction, but an excessive "zhongerbi" style is "offensive" to the user, the input samples of the language style training can be adjusted based on the user's feedback during the interaction process, such as adjusting the weight of one or more of the intonation, timbre, and tone, and adjusting the scope of word formation and grammatical structure.

[0185] In another specific embodiment, the vehicle voice interaction device shown in FIG3 includes: a voice data module, a voice recognition module, a template training module, and a modified voice module;

[0186] Voice data module, used to obtain voice data corresponding to the in-vehicle application context;

[0187] Speech recognition module, used to identify the semantics of speech in the context of in-vehicle applications;

[0188] The template training module is used to train stylized natural language templates for voice interaction based on the semantics of speech in the in-vehicle application context;

[0189] The speech modification module is used to modify the language style of the feedback speech in the speech interaction according to the stylized natural language template.

[0190] FIG6 is a schematic diagram showing a comparison of voice interaction improvements according to a specific embodiment of the present invention.

[0191] FIG7 is a schematic diagram of a TTS text library according to a specific embodiment of the present invention.

[0192] FIG8 is a schematic diagram of updating a local stylized natural language template library according to a specific embodiment of the present invention.

[0193] FIG9 is a schematic diagram of a large model return example according to a specific embodiment of the present invention.

[0194] FIG10 is a schematic diagram of a speech training architecture according to a specific embodiment of the present invention.

[0195] In a specific embodiment, as shown in Figure 6, the voice interaction improvements compare that in-vehicle voice interaction currently uses pre-set responses. For example, if a user asks, "Turn on the air conditioner," the vehicle computer responds, "The air conditioner is on, and the current temperature is 'temperature' degrees" (where temperature is the temperature of the air conditioner).

[0196] AIGC, or large models, have significant advantages in semantic understanding and dialogue generation. However, large models take a long time to output dialogues and are relatively expensive, making them less suitable for cost-sensitive scenarios with high real-time requirements.

[0197] In another specific embodiment, under the voice training architecture shown in FIG10 , the TTS copy library shown in FIG7 is developed to add voice style to the feedback of vehicle-side voice interaction.

[0198] In another specific embodiment, under the voice training architecture shown in Figure 10, the local stylized natural language template library is updated as shown in Figure 8, the large model engine trains the stylized natural language template library, and the local stylized natural language template library is used as a subset of the stylized natural language template library. The required stylized natural language templates are selected and downloaded locally, thereby reducing the interaction frequency between the local system and the cloud system and reducing costs.

[0199] In another specific embodiment, in another specific embodiment, under the voice training framework shown in Figure 8, a large model return example as shown in Figure 9 is carried out.

[0200] 1. The user has a need to adjust the voice copywriting style, such as "Can you answer in a more childish way?"

[0201] 2. The speech semantic recognition engine identifies user needs, and the copywriting style is "zhong er"; in this logic, the semantic engine needs to normalize the semantics; similarly, needs such as enthusiasm and fiery are normalized as "enthusiasm"; similarly, needs such as "zhong er" and "funny" are normalized as "zhong er".

[0202] 3. The dialogue management module sends the copywriting style to the language generation module;

[0203] 4. The language generation module queries the local stylized natural language template library for the target style. If the target style is found, it directly switches to the target style copy library. In this scenario, the voice directly responds: "The conversation style has been adjusted to 'Conversational Style'."

[0204] 5. The language generation module queries the local stylized natural language template library for the target style. If not, it queries the cloud for the target style. If so, it synchronizes the cloud style with the local copy library and switches to the target style. In this scenario, the voice response is: "The conversation style has been adjusted to 'Conversational Style'."

[0205] 6. The language generation module queries the local stylized natural language template library for the target style. If not, it queries the cloud for the target style. If not, it stylizes the natural language template library, enters the style into the AIGC prompt copy, and generates a new style copy by accessing the large model. The copy is then sent to the local stylized natural language template library and switched to the target style. In this scenario, the voice first responds: "The 'conversational style' style is being generated for you."

[0206] After waiting for AIGC to be generated, proactively initiate a conversation to inform the user that "a 'conversational style' has been generated for you."

[0207] FIG11 is a block diagram of an electronic device structure of a vehicle voice interaction method provided by one or more embodiments of the present invention.

[0208] As shown in FIG11 , the present application provides an electronic device, comprising: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;

[0209] A computer program is stored in the memory. When the computer program is executed by the processor, the processor executes the steps of a vehicle voice interaction method.

[0210] The present application also provides a computer-readable storage medium, which stores a computer program that can be executed by an electronic device. When the computer program runs on the electronic device, the electronic device executes the steps of a vehicle voice interaction method.

[0211] The present application also provides a vehicle, comprising:

[0212] An electronic device for implementing the steps of the vehicle voice interaction method;

[0213] a processor that runs a program and, when the program is running, executes the steps of the vehicle voice interaction method based on data output by the electronic device;

[0214] The storage medium is used to store a program, which, when running, executes the steps of the vehicle voice interaction method for data output from the electronic device.

[0215] The communication bus mentioned in the electronic device mentioned above may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.

[0216] The electronic device includes a hardware layer, an operating system layer running on the hardware layer, and an application layer running on the operating system. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory. The operating system can be any one or more computer operating systems that control electronic devices through processes, such as the Linux operating system, the Unix operating system, the Android operating system, the iOS operating system, or the Windows operating system. In the embodiments of the present invention, the electronic device can be a handheld device such as a smartphone or a tablet computer, or an electronic device such as a desktop computer or a portable computer, which is not particularly limited in the embodiments of the present invention.

[0217] The execution subject of the electronic device control in the embodiment of the present invention can be an electronic device, or a functional module in the electronic device that can call a program and execute the program. The electronic device can obtain the firmware corresponding to the storage medium. The firmware corresponding to the storage medium is provided by the supplier. The firmware corresponding to different storage media can be the same or different, and is not limited here. After the electronic device obtains the firmware corresponding to the storage medium, it can write the firmware corresponding to the storage medium into the storage medium, specifically, burn the firmware corresponding to the storage medium into the storage medium. The process of burning the firmware into the storage medium can be implemented using existing technology and will not be described in detail in the embodiment of the present invention.

[0218] The electronic device can also obtain a reset command corresponding to the storage medium. The reset command corresponding to the storage medium is provided by the supplier. The reset commands corresponding to different storage media can be the same or different, and are not limited here.

[0219] In this case, the storage medium of the electronic device is a storage medium in which the corresponding firmware is written. The electronic device can respond to the reset command corresponding to the storage medium in which the corresponding firmware is written, thereby resetting the storage medium in which the corresponding firmware is written according to the reset command corresponding to the storage medium. The process of resetting the storage medium according to the reset command can be implemented in the existing technology and will not be described in detail in the embodiments of the present invention.

[0220] For the convenience of description, the above devices are described as various units and modules according to their functions. Of course, when implementing this application, the functions of each unit and module can be implemented in the same or multiple software and / or hardware.

[0221] Those skilled in the art will understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art in the art to which the present invention pertains. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with those in the context of the prior art and, unless specifically defined, will not be interpreted in an idealized or overly formal sense.

[0222] For simplicity of description, the method embodiments are described as a series of actions. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because certain steps can be performed in other orders or simultaneously according to the embodiments of the present invention. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0223] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general-purpose hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.

[0224] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A vehicle voice interaction method, characterized in that, The vehicle voice interaction method includes: Obtain user voice data; Parse the voice data to identify the user's emotional needs and basic functional needs; Generate an emotional natural language prompt project according to the user's emotional needs and basic functional needs; Perform AIGC model training according to the emotional natural language prompt project to generate a stylized natural language template; Generate a voice interaction copy corresponding to the user's emotional needs and basic functional needs according to the stylized natural language template.

2. The vehicle voice interaction method according to claim 1, wherein, The generating a voice interaction copy corresponding to the user's emotional needs and basic functional needs according to the stylized natural language template includes: The voice data corresponds to the in-vehicle application context; Parse the voice data according to the in-vehicle application context to identify the user's emotional needs and basic functional needs; Generate an emotional natural language prompt project according to the user's emotional needs, basic functional needs, and the in-vehicle application context; Perform AIGC model training according to the emotional natural language prompt project to generate a stylized natural language template; Generate a voice interaction copy corresponding to the user's emotional needs and basic functional needs according to the stylized natural language template corresponding to the in-vehicle application context.

3. The vehicle voice interaction method according to claim 1 or 2, characterized in that The generating a voice interaction copy corresponding to the user's emotional needs and basic functional needs according to the stylized natural language template further includes: Identify the user's emotional needs and basic functional needs based on the local user voice data; Parse the voice data according to the in-vehicle application context of the local user to identify the user's emotional needs and basic functional needs; Generate an emotional natural language prompt project according to the user's emotional needs, basic functional needs, and the in-vehicle application contexts of multiple local users; Perform AIGC model training according to the emotional natural language prompt project to generate a cloud stylized natural language template; Synchronize the local stylized natural language template according to the cloud stylized natural language template; Generate a voice interaction copy corresponding to the user's emotional needs and basic functional needs according to the local stylized natural language template.

4. The vehicle voice interaction method according to claim 3, wherein The generating a voice interaction copy corresponding to the user's emotional needs and basic functional needs according to the stylized natural language template further includes: Obtain user characteristic information; Obtain user voice data based on the user characteristic information; Parse the user's voice data according to the user characteristic information to identify the user's emotional needs and basic functional needs; Generate an emotional natural language prompt project corresponding to the user's emotional needs, basic functional needs, and the user characteristics; Perform AIGC model training according to the emotional natural language prompt project to generate a stylized natural language template; Generate a voice interaction copy corresponding to the user's emotional needs and basic functional needs according to the stylized natural language template.

5. The vehicle voice interaction method according to claim 4, wherein The parsing the voice data to identify the user's emotional needs and basic functional needs includes: Generate TTS document information according to the obtained user voice data; Obtain the slot information of the user's basic functional requirements and the slot information of the user's emotional requirements according to the TTS document information; Generate a voice interaction copy corresponding to the user's emotional requirements and basic functional requirements according to the slot information of the basic functional requirements, the slot information of the user's emotional requirements, and the stylized natural language template.

6. The vehicle voice interaction method according to claim 5, characterized in that The generation of the stylized natural language template by performing AIGC model training according to the emotional natural language prompting engineering includes: Generate a stylized natural language template library according to the stylized natural language template; The stylized natural language template library includes the TTS copy corresponding to the stylized natural language template; Output the TTS copy corresponding to the stylized natural language template according to the emotional natural language prompting engineering; Judge whether the output TTS copy corresponds to the user's emotional requirements according to the slot information and the TTS copy; If the output TTS copy corresponds to the user's emotional requirements, activate the stylized natural language template according to the TTS copy.

7. A vehicle voice interaction device, characterized in that, The vehicle voice interaction device includes: A user data module for obtaining user voice data; A data parsing module for parsing the voice data to identify the user's emotional requirements and basic functional requirements; A prompting engineering module for generating an emotional natural language prompting engineering according to the user's emotional requirements and basic functional requirements; A template generation module for performing AIGC model training according to the emotional natural language prompting engineering to generate a stylized natural language template; An interaction copy module for generating a voice interaction copy corresponding to the user's emotional requirements and basic functional requirements according to the stylized natural language template.

8. An electronic device, characterized in that, Includes: A processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; A computer program is stored in the memory. When the computer program is executed by the processor, the processor executes the steps of the vehicle voice interaction method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, Includes: It stores a computer program executable by an electronic device. When the computer program runs on the electronic device, the electronic device executes the steps of the vehicle voice interaction method according to any one of claims 1 to 6.

10. A vehicle, characterized in that, Includes: An electronic device for implementing the steps of the vehicle voice interaction method according to any one of claims 1 to 6; A processor, the processor runs a program. When the program runs, it executes the steps of the vehicle voice interaction method according to any one of claims 1 to 6 on the data output from the electronic device; A storage medium for storing a program. When the program runs, it executes the steps of the vehicle voice interaction method according to any one of claims 1 to 6 on the data output from the electronic device.

Citation Information

Patent Citations

  • Voice interaction method, vehicle, server, system and storage medium

    CN111767021A

  • Voice interaction method and device, electronic equipment and storage medium

    CN112382287A

  • Conversation processing method and device thereof and electronic equipment

    CN113205811A

  • Speech generation method and device based on pre-training language model, equipment and medium

    CN116364055A

  • Multi-mode reply generation method and device, electronic equipment and storage medium

    CN116401349A