Vehicle-mounted voice interaction method and device, computer readable medium and electronic equipment
By recognizing the age, speech rate, and conversation habits of in-vehicle terminal users, the system determines the intent of the conversation and generates personalized interactive responses and speech rates. This solves the problem that in-vehicle voice interaction cannot adapt to different user habits and improves the user experience.
Patent Information
- Application Number
- CN202411427833.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-14
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2044-10-14
AI Technical Summary
Existing in-vehicle voice interaction functions cannot adapt to the different conversational habits of users, affecting the user experience.
By receiving voice input from target users, the system identifies user age, speaking speed, and conversation habits, determines the conversational intent, and generates adaptive interactive responses and their playback speed based on these characteristics.
It enables personalized voice interaction experiences based on the different dialogue characteristics of different users, thereby improving user satisfaction.
Smart Images

Figure CN119517017B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of voice interaction, and particularly relates to a vehicle-mounted voice interaction method and device, a computer readable medium and an electronic device. BACKGROUND
[0002] With the development of artificial intelligence and vehicle networking, vehicle-mounted voice interaction functions have gradually been popularized. In the current technical solution, the voice interaction function configured by the vehicle-mounted terminal is relatively monotonous and unified when performing voice interaction. Although the user can replace the voice interaction role to improve the interaction experience, the voice interaction can only be performed in a specific mode, and cannot adapt to the conversation habits of different drivers or passengers, affecting the user experience. Therefore, how to adapt to the conversation habits of different users and ensure the user experience has become a technical problem to be solved. SUMMARY
[0003] Embodiments of the present application provide a vehicle-mounted voice interaction method, device, computer readable medium and electronic device, which can at least adapt to the conversation habits of different users and ensure the user experience to some extent.
[0004] Other characteristics and advantages of the present application will become apparent from the following detailed description, or will be learned by practice of the present application.
[0005] According to an aspect of an embodiment of the present application, a vehicle-mounted voice interaction method is provided, comprising:
[0006] receiving a voice input of a target user, the target user being any one of all persons in the current vehicle;
[0007] identifying according to the voice input to determine a user age, a conversation speed and a conversation habit of the target user, the conversation habit being used to represent a content detail degree of the target user when conversing;
[0008] performing semantic recognition on text information corresponding to the voice input to determine a conversation intention of the target user;
[0009] determining corresponding interaction reply content and a playing speed thereof according to the conversation intention, the user age, the conversation speed and the conversation habit;
[0010] performing voice playing on the interaction reply content according to the playing speed.
[0011] In some embodiments of the present application, based on the foregoing scheme, determining corresponding interaction reply content and a playing speed thereof according to the conversation intention, the user age, the conversation speed and the conversation habit comprises:
[0012] query according to the dialogue intention, determine a plurality of candidate interactive reply templates corresponding to the dialogue intention, different candidate interactive reply templates correspond to different content details;
[0013] determine a target interactive reply template from the plurality of candidate interactive reply templates according to the user age, the dialogue speed and / or the dialogue habit;
[0014] generate corresponding interactive reply content based on the target interactive reply template;
[0015] determine a playback speed corresponding to the interactive reply content according to the user age and / or the dialogue speed.
[0016] In some embodiments of the present application, based on the foregoing scheme, the target interactive reply template is determined from the plurality of candidate interactive reply templates according to the user age, the dialogue speed and / or the dialogue habit, comprising:
[0017] calculate the weighted sum of the user age, the dialogue speed and the dialogue habit, and take the weighted sum as an evaluation score;
[0018] select a candidate interactive reply template corresponding to a numerical range of the evaluation score from the plurality of candidate interactive reply templates as the target interactive reply template according to the numerical range.
[0019] In some embodiments of the present application, based on the foregoing scheme, the playback speed corresponding to the interactive reply content is determined according to the user age and / or the dialogue speed, comprising:
[0020] if the user age is less than or equal to a first age threshold, obtain a first preset playback speed as the playback speed corresponding to the interactive reply content;
[0021] if the user age is greater than or equal to a second age threshold, obtain a second preset playback speed as the playback speed corresponding to the interactive reply content, the second age threshold being greater than the first age threshold
[0022] if the user age is greater than the first age threshold and less than the second age threshold, determine the dialogue speed as the playback speed corresponding to the interactive reply content.
[0023] In some embodiments of the present application, based on the foregoing scheme, the target interactive reply template is determined from the plurality of candidate interactive reply templates according to the user age, the dialogue speed and / or the dialogue habit, comprising:
[0024] if the user age is less than or equal to the first age threshold or the user age is greater than or equal to the second age threshold, selecting a candidate interactive reply template with detailed content from the plurality of candidate interactive reply templates as the target interactive reply template;
[0025] if the user age is greater than the first age threshold and less than the second age threshold, selecting a candidate interactive reply template corresponding to the conversation habit from the plurality of candidate interactive reply templates as the target interactive reply template.
[0026] In some embodiments of the present application, based on the foregoing scheme, the conversation habit of the target user is determined according to the voice input, including:
[0027] The voice input is converted into corresponding text information;
[0028] The text information is processed by word segmentation to obtain at least one keyword;
[0029] According to the number of keywords corresponding to the target part of speech in at least one keyword, the conversation habit of the target user is determined.
[0030] According to an aspect of an embodiment of the present application, a vehicle-mounted voice interaction device is provided, including:
[0031] The receiving module is configured to receive voice input of a target user, the target user being any one of all people in a current vehicle;
[0032] The first identification module is configured to determine, according to the voice input, a user age, a conversation speed and a conversation habit of the target user, the conversation habit being used to represent a content detail level of the target user in conversation;
[0033] The second identification module is configured to determine, according to text information corresponding to the voice input, a conversation intention of the target user;
[0034] The determination module is configured to determine, according to the conversation intention, the user age, the conversation speed and the conversation habit, corresponding interactive reply content and a playing speed thereof;
[0035] The processing module is configured to perform voice playing on the interactive reply content according to the playing speed.
[0036] According to an aspect of an embodiment of the present application, a computer readable medium having a computer program stored thereon is provided, the computer program being executed by a processor to implement the vehicle-mounted voice interaction method as described in the foregoing embodiments.
[0037] According to an aspect of an embodiment of the present application, an electronic device is provided, comprising: one or more processors; a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the vehicle-mounted voice interaction method as described in the above embodiments.
[0038] According to an aspect of an embodiment of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes the vehicle-mounted voice interaction method provided in the above embodiments.
[0039] Based on the technical solutions provided in the present application, by receiving a voice input of a target user, the target user being any one of all people in the current vehicle, the voice input is recognized to determine a user age, a conversation speed and a conversation habit of the target user, the conversation habit being used to represent a content detail degree of the target user in conversation, and semantic recognition is performed on text information corresponding to the voice input to determine a conversation intention of the target user, then, the conversation habit, the user age, the conversation speed and the conversation habit are used to determine corresponding interaction reply content and a playing speed thereof, so as to perform voice playing on the interaction reply content according to the playing speed. In this way, the corresponding interaction reply content and playing speed can be determined according to the conversation characteristics of different users, so as to meet the conversation needs of different users and ensure user experience.
[0040] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and are not limiting to the present application. BRIEF DESCRIPTION OF DRAWINGS
[0041] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application. It is clear that the accompanying drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings. In the drawings:
[0042] Figure 1 A flowchart of a vehicle-mounted voice interaction method according to an embodiment of the present application is shown;
[0043] Figure 2 A block diagram of a vehicle-mounted voice interaction device according to an embodiment of the present application is shown;
[0044] Figure 3A structural schematic of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown. DETAILED DESCRIPTION
[0045] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of protection of the present application.
[0046] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the present application. One skilled in the relevant art will recognize, however, that the technical solutions of the present application can be practiced without one or more of the specific details, or with other methods, components, devices, steps, etc. In other instances, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0047] The block diagrams shown in the drawings are merely functional entities, and do not necessarily correspond to physically independent entities. That is, the functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0048] The flowcharts shown in the drawings are merely exemplary illustrations, and do not necessarily include all contents and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be further decomposed, and some operations / steps can be combined or partially combined, so the actual execution order can be changed according to actual conditions.
[0049] In the description of the present application, it should be understood that the terms "first", "second" are used only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, unless otherwise specified, the meaning of "multiple" is two or more.
[0050] Figure 1A flowchart of a vehicle-mounted voice interaction method according to an embodiment of the present application is shown. The method can be applied to a terminal device or a server, wherein the terminal device can include one or more of a vehicle-mounted terminal, a smartphone, a tablet computer, a portable computer, and a desktop computer; and the server can be a physical server or a cloud server.
[0051] As shown in Figure 1 The vehicle-mounted interaction method includes steps S110 to S150, which are described in detail as follows (the method is applied to a vehicle-mounted terminal as an example, hereinafter referred to as "terminal"):
[0052] In step S110, a voice input of a target user is received, wherein the target user is any one of all people in a current vehicle.
[0053] In this embodiment, the vehicle-mounted terminal can receive the voice input of the target user in real time through its own voice receiving device (such as a microphone). It is worth noting that the target person can be any one of all people in the current vehicle (such as the driver or a passenger), that is, any person in the current vehicle can interact with the vehicle-mounted terminal by voice, so that the vehicle-mounted terminal can interact with different people according to their conversation characteristics in the subsequent, and meet the conversation needs of different people.
[0054] In step S120, the voice input is recognized to determine a user age, a conversation speed, and a conversation habit of the target user, wherein the conversation habit is used to represent a content detail degree of the target user when conversing.
[0055] In this embodiment, after receiving the voice input of the target user, the vehicle-mounted terminal can recognize the received voice input to extract the conversation characteristics of the target user, which can include the user age, the conversation speed, and the conversation habit of the target user, wherein the conversation habit can be used to represent the content detail degree of the target user when conversing. In an example, the content detail degree can be divided into three types: concise, medium, and detailed. The vehicle-mounted terminal can recognize and classify based on the voice input of the target user, so as to match the target user with the corresponding type of content detail degree.
[0056] Thus, by extracting the conversation characteristics of different users, and providing a voice interaction experience that is more in line with their personal habits based on the conversation characteristics in the subsequent, the pertinence of voice interaction can be ensured, and the user experience can be improved.
[0057] In step S130, semantic recognition is performed on text information corresponding to the voice input to determine a conversation intention of the target user.
[0058] In this embodiment, the vehicle-mounted terminal can first convert the received voice input into corresponding text information, and perform semantic recognition based on the text information, so as to determine the dialog intention of the target user, i.e., the purpose or requirement expressed by the target user in the dialog, such as playing music or map navigation, etc.
[0059] In an example, the vehicle-mounted terminal can first determine the keywords existing in the text information, and match the keywords with a predefined list of intention keywords, and determine the dialog intention of the target user according to the matching result. In other examples, the vehicle-mounted terminal can also use machine learning algorithm or deep learning algorithm to recognize the text information, so as to determine the dialog intention of the target user.
[0060] It should be noted that the skilled in the art can also determine the dialog intention in other ways, which is not specially limited in the present application.
[0061] In step S140, the corresponding interactive reply content and its playing speed are determined according to the dialog intention, the user age, the dialog speed and the dialog habit.
[0062] In this embodiment, it should be understood that different users may have different dialog requirements when interacting with voice, for example, young users who speak concisely and at a fast speed will expect the vehicle-mounted terminal to also interact at a relatively fast speed and with concise content, while users who speak relatively slowly or children and the elderly will expect the vehicle-mounted terminal to interact at a slower speed and with more detailed content. Therefore, after determining the dialog intention of the target user, the vehicle-mounted terminal can comprehensively consider the dialog intention of the target user, the user age, the dialog speed and the dialog habit, so as to determine the corresponding interactive reply content and the playing speed, so as to achieve the pertinence in voice interaction.
[0063] In step S150, the interactive reply content is played according to the playing speed.
[0064] Thus, based on the dialog intention of the target user, the user age, the dialog speed and the dialog habit, the vehicle-mounted terminal can determine the corresponding interactive reply content and the playing speed, so as to achieve the pertinence in voice interaction. Figure 1In the illustrated embodiment, by receiving a voice input of a target user, which is any one of all people in the current vehicle, identifying according to the voice input, determining a user age, a conversation speed and a conversation habit of the target user, the conversation habit being used to represent a content detail degree of the target user when conversing, and performing semantic recognition on text information corresponding to the voice input to determine a conversation intention of the target user, then, according to the conversation habit, the user age, the conversation speed and the conversation habit, determining corresponding interactive reply content and a playing speed thereof, and performing voice playing on the interactive reply content according to the playing speed. In this way, the corresponding interactive reply content and the playing speed can be determined according to the conversation characteristics of different users, so as to meet the conversation needs of different users and ensure user experience.
[0065] In some embodiments of the present application, according to the conversation intention, the user age, the conversation speed and the conversation habit, determining corresponding interactive reply content and a playing speed thereof, comprises:
[0066] According to the conversation intention, querying to determine a plurality of candidate interactive reply templates, different candidate interactive reply templates corresponding to different content detail degrees;
[0067] According to the user age, the conversation speed and / or the conversation habit, determining a target interactive reply template from the plurality of candidate interactive reply templates;
[0068] Based on the target interactive reply template, generating corresponding interactive reply content;
[0069] According to the user age and / or the conversation speed, determining a playing speed corresponding to the interactive reply content.
[0070] In this embodiment, a person skilled in the art can preset a plurality of candidate interactive reply templates for each conversation intention, and different candidate interactive templates corresponding to the same conversation intention correspond to different content detail degrees. In this way, when voice interaction is needed, the vehicle-mounted terminal can first query according to the conversation intention of the target user, so as to determine a plurality of candidate interactive reply templates corresponding to the conversation intention, and then the vehicle-mounted terminal can select one of the plurality of candidate interactive reply templates as a target interactive reply template according to the user age, the conversation speed and / or the conversation habit.
[0071] Notably, the in-vehicle terminal can select the target interaction reply template from the plurality of candidate interaction reply templates according to one or more of the user age, the conversation speed, and the conversation habit. In an example, different age groups correspond to interaction reply templates of different content detail levels, and the in-vehicle terminal can select the target interaction reply template according to only the magnitude of the user age, and determine the candidate interaction reply template corresponding to the age group in which the user age falls as the target interaction reply template. In another example, the in-vehicle terminal can also select the candidate interaction reply template corresponding to the content detail level according to the conversation habit of the target user, and determine the candidate interaction reply template as the target interaction reply template. In yet another example, the in-vehicle terminal can comprehensively consider any two or three of the user age, the conversation speed, and the conversation habit, and then determine the target interaction reply template from the plurality of candidate interaction reply templates.
[0072] After determining the target interaction reply template, the in-vehicle terminal can generate the corresponding interaction reply content based on the target interaction reply template. In an example, when the content detail level corresponding to the target interaction reply template is medium or detailed, it can include at least one slot, and the in-vehicle terminal can extract a specific type of keyword (such as a place name or a song name, etc.) from the text information corresponding to the voice input to fill in the corresponding slot, thereby generating the corresponding interaction reply content. For example, when the target user needs to perform map navigation, the in-vehicle terminal can extract the keyword for indicating the destination from the text information to fill in the target interaction reply template, thereby generating the interaction reply content. When the content detail level corresponding to the target interaction reply template is concise, it can also not include a slot, i.e., the target interaction reply template is the interaction reply content. For example, when the target user needs to perform map navigation, the target interaction reply template of the content detail level of concise can be "start navigation" or "has started", etc.
[0073] Next, the in-vehicle terminal can also determine the playback speed of the interaction reply content according to the user age and / or the conversation speed of the target user. It should be understood that the playback speed is the TTS (Text To Speech) speed, i.e., the in-vehicle terminal can convert the interaction reply content in the form of text into corresponding speech for playing according to the playback speed.
[0074] In an example, the in-vehicle terminal can determine the playback speed of the interaction reply content according to only the user age. Those skilled in the art can pre-set the corresponding playback speed for different age groups, so that in actual use, the corresponding playback speed can be queried and determined according to the age group in which the user age falls.
[0075] In another example, the in-vehicle terminal can also directly determine the conversation speed of the target user as the playback speed of the interaction reply content, thereby adapting to the conversation needs of the target user.
[0076] In some embodiments of the present application, the target interaction reply template is determined from the plurality of candidate interaction reply templates according to the user age, the conversation speed and / or the conversation habit, comprising:
[0077] The weighted sum of the user age, the conversation speed and the conversation habit is calculated, and the weighted sum is taken as an evaluation score;
[0078] According to the numerical range in which the evaluation score falls, a candidate interaction reply template corresponding to the numerical range is selected from the plurality of candidate interaction reply templates as the target interaction reply template.
[0079] In this embodiment, the in-vehicle terminal can comprehensively consider the user age, the conversation speed and the conversation habit of the target user to determine the corresponding target interaction reply template. Specifically, the person skilled in the art can pre-set the respective weights of the user age, the conversation speed and the conversation habit, and after obtaining the specific values of the three, the weighted sum can be calculated according to the respective weights, and the calculation result is taken as the evaluation score of the target user.
[0080] The person skilled in the art can pre-determine the numerical ranges corresponding to different candidate interaction reply templates, and after obtaining the evaluation score of the target user, the corresponding candidate interaction reply template can be found according to the numerical range in which the evaluation score falls to be the target interaction reply template, so as to ensure the accuracy of the determination of the target interaction reply template.
[0081] In some embodiments of the present application, the playback speed corresponding to the interaction reply content is determined according to the user age and / or the conversation speed, comprising:
[0082] If the user age is less than or equal to a first age threshold, a first preset playback speed is obtained as the playback speed corresponding to the interaction reply content;
[0083] If the user age is greater than or equal to a second age threshold, a second preset playback speed is obtained as the playback speed corresponding to the interaction reply content, the second age threshold being greater than the first age threshold
[0084] If the user age is greater than the first age threshold and less than the second age threshold, the conversation speed is determined as the playback speed corresponding to the interaction reply content.
[0085] In this embodiment, the vehicle-mounted terminal can determine the playing speed according to the user age of the target user. Specifically, a person skilled in the art can pre-set a first age threshold and a second age threshold according to prior experience, and the second age threshold is greater than the first age threshold. For example, the first age threshold can be 13 years old, and the second age threshold can be 60 years old.
[0086] The vehicle-mounted terminal can compare the user age with the first age threshold and the second age threshold. If the user age is less than or equal to the first age threshold, it indicates that the target user is relatively young. The vehicle-mounted terminal can obtain a first preset playing speed as the playing speed of the interactive reply content. The first preset playing speed can be a playing speed suitable for a user who is relatively young, which is pre-set by a person skilled in the art according to prior experience.
[0087] If the user age is greater than or equal to the second age threshold, it indicates that the target user is relatively old. The speed is too fast and may not be clear. Therefore, the vehicle-mounted terminal can obtain a second preset playing speed as the playing speed of the interactive reply content. Similarly, the second preset playing speed can also be a playing speed suitable for a user who is relatively old, which is pre-set by a person skilled in the art according to prior experience.
[0088] Furthermore, if the user age is between the first age threshold and the second age threshold, that is, the user age is greater than the first age threshold and less than the second age threshold, the vehicle-mounted terminal can directly determine the conversation speed as the playing speed of the interactive reply content, so as to adapt to the conversation needs of the target user and ensure the user experience.
[0089] Based on the foregoing embodiments, in some embodiments of the present application, the target interactive reply template is determined from the plurality of candidate interactive reply templates according to the user age, the conversation speed, and / or the conversation habit, comprising:
[0090] If the user age is less than or equal to the first age threshold, or the user age is greater than or equal to the second age threshold, a candidate interactive reply template with detailed content is selected from the plurality of candidate interactive reply templates as the target interactive reply template;
[0091] If the user age is greater than the first age threshold and less than the second age threshold, a candidate interactive reply template corresponding to the conversation habit in content detail is selected from the plurality of candidate interactive reply templates as the target interactive reply template.
[0092] In this embodiment, the candidate interactive reply template with the detailed content detail level can be used as the target interactive reply template for a target user who is younger (i.e., less than or equal to the first age threshold) or older (i.e., greater than or equal to the second age threshold), so that more detailed feedback can be provided for the target user to make the target user clear about his / her operation behavior.
[0093] For a target user with an intermediate age (i.e., greater than the first age threshold and less than the second age threshold), the in-vehicle terminal can give a reply with a content detail level corresponding to the conversation habit of the target user, i.e., select a candidate interactive reply template with a content detail level corresponding to the conversation habit of the target user as the target interactive reply template, so as to meet the conversation habit of the target user and ensure user experience.
[0094] In some embodiments of the present application, the conversation habit of the target user is determined according to the voice input, including:
[0095] The voice input is converted into corresponding text information;
[0096] The text information is processed by word segmentation to obtain at least one keyword;
[0097] The conversation habit of the target user is determined according to the number of keywords corresponding to a target part of speech in the at least one keyword.
[0098] In this embodiment, when determining the conversation habit of the target user, the in-vehicle terminal can first convert the voice input of the target user into corresponding text information, and process the text information by word segmentation to obtain at least one keyword. Then, the in-vehicle terminal can determine the part of speech of each keyword and determine the number of keywords corresponding to a target part of speech, so as to determine the conversation habit of the target user based on the number.
[0099] In an example, a person skilled in the art can pre-set a first number threshold and a second number threshold. If the number of keywords of the target part of speech is less than the first number threshold, it indicates that the target user's conversation is concise, and it can be determined that the conversation habit of the target user is concise. If the number of keywords of the target part of speech is greater than or equal to the first number threshold and less than or equal to the second number threshold, it can be determined that the conversation habit of the target user is intermediate. If the number of keywords of the target part of speech is greater than the second number threshold, it can be determined that the conversation habit of the target user is detailed.
[0100] Thus, the in-vehicle terminal can accurately identify the conversation habit of the target user, and then can ensure the adaptability of the subsequently determined interactive reply content to the target user and ensure user experience.
[0101] The device embodiments of the present application are introduced below, which can be used to execute the vehicle-mounted voice interaction method in the above-mentioned embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the above-mentioned embodiments of the vehicle-mounted voice interaction method.
[0102] Figure 2 A block diagram of a vehicle-mounted voice interaction device according to one embodiment of the present application is shown.
[0103] Referring to Figure 2 According to one embodiment of the present application, the vehicle-mounted voice interaction device includes:
[0104] A receiving module is configured to receive a voice input of a target user, the target user being any one of all people in a current vehicle;
[0105] A first identifying module is configured to identify the voice input to determine a user age, a conversation speed and a conversation habit of the target user, the conversation habit being used to represent a content detail degree of the target user when conversing;
[0106] A second identifying module is configured to perform semantic identification on text information corresponding to the voice input to determine a conversation intention of the target user;
[0107] A determining module is configured to determine corresponding interaction reply content and a playing speed thereof according to the conversation intention, the user age, the conversation speed and the conversation habit;
[0108] A processing module is configured to perform voice playing on the interaction reply content according to the playing speed.
[0109] In some embodiments of the present application, determining corresponding interaction reply content and a playing speed thereof according to the conversation intention, the user age, the conversation speed and the conversation habit includes:
[0110] Querying according to the conversation intention to determine a plurality of candidate interaction reply templates, different candidate interaction reply templates corresponding to different content detail degrees;
[0111] Determining a target interaction reply template from the plurality of candidate interaction reply templates according to the user age, the conversation speed and / or the conversation habit;
[0112] Generating corresponding interaction reply content based on the target interaction reply template;
[0113] Determining a playing speed of the interaction reply content according to the user age and / or the conversation speed.
[0114] In some embodiments of the present application, the target interaction reply template is determined from the plurality of candidate interaction reply templates according to the user age, the conversation speed, and / or the conversation habit, including:
[0115] A weighted sum of the user age, the conversation speed, and the conversation habit is calculated, and the weighted sum is taken as an evaluation score;
[0116] According to a numerical range in which the evaluation score falls, a candidate interaction reply template corresponding to the numerical range is selected from the plurality of candidate interaction reply templates as the target interaction reply template.
[0117] In some embodiments of the present application, the playback speed corresponding to the interaction reply content is determined according to the user age and / or the conversation speed, including:
[0118] If the user age is less than or equal to a first age threshold, a first preset playback speed is obtained as the playback speed corresponding to the interaction reply content;
[0119] If the user age is greater than or equal to a second age threshold, a second preset playback speed is obtained as the playback speed corresponding to the interaction reply content, the second age threshold being greater than the first age threshold
[0120] If the user age is greater than the first age threshold and less than the second age threshold, the conversation speed is determined as the playback speed corresponding to the interaction reply content.
[0121] In some embodiments of the present application, the target interaction reply template is determined from the plurality of candidate interaction reply templates according to the user age, the conversation speed, and / or the conversation habit, including:
[0122] If the user age is less than or equal to the first age threshold, or the user age is greater than or equal to the second age threshold, a candidate interaction reply template with detailed content is selected from the plurality of candidate interaction reply templates as the target interaction reply template;
[0123] If the user age is greater than the first age threshold and less than the second age threshold, a candidate interaction reply template corresponding to the conversation habit in terms of content detail is selected from the plurality of candidate interaction reply templates as the target interaction reply template.
[0124] In some embodiments of the present application, the conversation habit of the target user is determined according to the voice input, including:
[0125] The voice input is converted into corresponding text information;
[0126] performing word segmentation on the text information to obtain at least one keyword;
[0127] determining the conversation habit of the target user according to a number of keywords corresponding to a target part-of-speech in the at least one keyword.
[0128] Figure 3 A structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown.
[0129] It should be noted that, Figure 3 The computer system of the electronic device shown is only an example and should not impose any limitation on the functions and use range of the embodiments of the present application.
[0130] As Figure 3 shown, the computer system includes a central processing unit (CPU) 301 which can perform various appropriate actions and processes, such as executing the methods described in the above embodiments, according to programs stored in a read-only memory (ROM) 302 or loaded from a storage section 308 into a random access memory (RAM) 303. Various programs and data required for system operation are also stored in the RAM 303. The CPU 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0131] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, etc.; an output section 307 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as necessary. A removable recording medium 311 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 310 as necessary, so that a computer program read out therefrom is installed in the storage section 308 as necessary.
[0132] In particular, the processes described above with reference to the flow charts can be implemented as a computer software program in accordance with embodiments of the present application. For example, embodiments of the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program comprising computer instructions for performing the methods illustrated by the flow charts. In such embodiments, the computer program can be downloaded and installed from a network via the communication section 309 and / or installed from the removable media 311. When the computer program is executed by the central processing unit (CPU) 301, various functions defined in the system of the present application are performed.
[0133] It should be noted that the computer readable medium shown in the embodiments of the present application can be a computer readable signal medium or a computer readable storage medium or any combination thereof. The computer readable storage medium may, for example, be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this application, the computer readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In this application, the computer readable signal medium can include a data signal carried in a baseband or as part of a carrier wave, in which the computer readable program is carried. Such a propagated data signal can take any of a variety of forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device. The computer program contained in the computer readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0134] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or combinations of hardware and software.
[0135] The units described in the embodiments of the present application can be implemented by software, or by hardware, or by a combination of software and hardware. The units described may
[0136] As another aspect, the present application also provides a computer readable medium, which can be included in the electronic device described in the above embodiments, or can exist separately without being assembled into the electronic device. The computer readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to implement the method described in the above embodiments.
[0137] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, the division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into a plurality of modules or units.
[0138] From the above description of the embodiments, those skilled in the art will readily appreciate that the example embodiments described herein can be implemented by software and / or by hardware coupled with software. Accordingly, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, or the like) or on a network, and includes a number of instructions for causing a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to perform the methods according to the embodiments of the present application.
[0139] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. It is intended that the application embrace any and all variations of the present application that fall within the scope of the general inventive concept as defined by the appended claims and their equivalents. It is intended that the application not be limited to the examples described herein and that upon reading this disclosure, persons of ordinary skill in the art will appreciate still other implementations that are within the scope of the present application. Those skilled in the art will readily recognize from this disclosure that alternative embodiments of the present application can be employed without departing from the characteristics of the application. Accordingly, the application is not to be limited as illustrated and described herein, but is only limited as by the claims which follow.
[0140] It is to be understood that the application is not limited to the precise details of construction and the arrangement of components described above and illustrated in the drawings. Various modifications and changes can be made thereunto without departing from the scope of the application. The scope of the application is limited only by the claims that follow.
Claims
1. A vehicle-mounted voice interaction method, characterized in that, include: Receive voice input from a target user, where the target user is any one of all people currently in the vehicle; Based on the voice input, the user's age, speaking speed, and dialogue habits are determined, and the dialogue habits are used to characterize the level of detail in the content of the dialogue. Semantic recognition is performed based on the text information corresponding to the voice input to determine the target user's dialogue intent; Based on the dialogue intent, a query is performed to determine multiple candidate interactive response templates, with different candidate interactive response templates corresponding to different levels of content detail. Based on the user's age, the conversation speed, and / or the conversation habits, a target interactive response template is determined from a plurality of candidate interactive response templates; When the content of the target interactive response template is of medium or detailed level, it contains at least one slot. Specific keywords of a certain type are extracted from the text information corresponding to the voice input and filled into the corresponding slot to generate the corresponding interactive response content. When the content of the target interactive response template is of concise level, it does not contain any slots, and the target interactive response template is the interactive response content. The playback speed corresponding to the interactive response content is determined based on the user's age and / or the conversation speed. The interactive response content is played back by voice according to the playback speed.
2. The method according to claim 1, characterized in that, Based on the user's age, the conversation speed, and / or the conversation habits, a target interactive response template is determined from a plurality of candidate interactive response templates, including: Calculate a weighted sum of the user's age, the conversation speed, and the conversation habits, and use the weighted sum as the evaluation score; Based on the numerical range of the evaluation score, a candidate interactive response template corresponding to the numerical range is selected from multiple candidate interactive response templates as the target interactive response template.
3. The method according to claim 1, characterized in that, Determining the playback speed corresponding to the interactive response content based on the user's age and / or the conversation speed includes: If the user's age is less than or equal to a first age threshold, a first preset playback speed is obtained as the playback speed corresponding to the interactive response content; If the user's age is greater than or equal to a second age threshold, a second preset playback speed is obtained as the playback speed corresponding to the interactive response content. The second age threshold is greater than the first age threshold. If the user's age is greater than the first age threshold and less than the second age threshold, the dialogue speed is determined to be the playback speed corresponding to the interactive response content.
4. The method according to claim 3, characterized in that, Based on the user's age, the conversation speed, and / or the conversation habits, a target interactive response template is determined from a plurality of candidate interactive response templates, including: If the user's age is less than or equal to the first age threshold, or if the user's age is greater than or equal to the second age threshold, select the interactive response template with the highest level of detail from the multiple candidate interactive response templates as the target interactive response template; If the user's age is greater than the first age threshold and less than the second age threshold, select the candidate interaction response template whose content detail corresponds to the dialogue habit from among the multiple candidate interaction response templates as the target interaction response template.
5. The method according to any one of claims 1-4, characterized in that, Based on the voice input, the target user's dialogue habits are determined, including: Convert the voice input into corresponding text information; The text information is segmented to obtain at least one keyword; The target user's conversation habits are determined based on the number of keywords corresponding to the target part of speech in at least one of the keywords.
6. A vehicle-mounted voice interaction device, characterized in that, include: The receiving module is used to receive voice input from a target user, who is any one of the people currently in the vehicle. The first recognition module is used to recognize the voice input and determine the target user's age, speech rate and dialogue habits. The dialogue habits are used to characterize the level of detail in the content of the dialogue. The second recognition module is used to perform semantic recognition based on the text information corresponding to the voice input to determine the dialogue intent of the target user. The determination module is used to query based on the dialogue intent and determine multiple corresponding candidate interactive response templates, with different candidate interactive response templates having different levels of detail in their content; Based on the user's age, the conversation speed, and / or the conversation habits, a target interactive response template is determined from a plurality of candidate interactive response templates. When the content of the target interactive response template is of medium or detailed level, it contains at least one slot. Specific types of keywords are extracted from the text information corresponding to the voice input and filled into the corresponding slot to generate the corresponding interactive response content. When the content of the target interactive response template is of concise level, it does not contain any slots, and the target interactive response template is the interactive response content. Based on the user's age and / or the conversation speed, the playback speed corresponding to the interactive response content is determined. The processing module is used to play the interactive response content by voice according to the playback speed.
7. The apparatus according to claim 6, characterized in that, Based on the stated dialogue intent, the user's age, the dialogue speed, and the dialogue habits, determine the corresponding interactive response content and its playback speed, including: Based on the dialogue intent, a query is performed to determine multiple candidate interactive response templates, with different candidate interactive response templates corresponding to different levels of content detail. Based on the user's age, the conversation speed, and / or the conversation habits, a target interactive response template is determined from a plurality of candidate interactive response templates; Based on the target interactive response template, generate the corresponding interactive response content; The playback speed corresponding to the interactive response content is determined based on the user's age and the conversation speed.
8. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the in-vehicle voice interaction method as described in any one of claims 1 to 5.
9. An electronic device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the in-vehicle voice interaction method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Intelligent voice interaction implementation methods and devices, computer equipment, and storage medium
CN108711423A
Voice interaction method and device, electronic equipment and storage medium
CN115329057A