Voice interaction method, device, electronic device and storage medium

By conducting semantic analysis of user interaction voice, determining user intention information and outputting personalized and emotional voice responses, the problem of insufficient voice interaction in the prior art is solved and the user experience is improved.

CN114093356BActive Publication Date: 2025-08-01APOLLO INTELLIGENT CONNECTIVITY (BEIJING) TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111297097.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-03
Publication Date
2025-08-01
Estimated Expiration
2041-11-03

AI Technical Summary

Technical Problem

Existing voice interaction technologies are difficult to provide personalized and emotional voice responses based on user intention information, resulting in insufficient user experience.

Method used

By conducting semantic analysis of user interaction speech, user intention information is determined, and target voice response mode is determined based on the intention information, including voice response style and pronunciation characteristics of the pronunciation person, and personalized and emotional voice response information is output.

Benefits of technology

It improves the personalization and fun of voice interaction, meets users' personalized needs, and improves users' user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114093356B_ABST
    Figure CN114093356B_ABST
Patent Text Reader

Abstract

The present disclosure provides a voice interaction method, apparatus, electronic device, and storage medium, relating to the field of computer technology, and particularly to the fields of autonomous driving, intelligent cockpit, and vehicle networking technologies. The specific implementation solution is as follows: in response to receiving an interaction voice from a user, perform semantic analysis on the interaction voice to obtain user intention information; determine a target voice response manner based on the user intention information; and output voice response information according to the target voice response manner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of computer technology and Internet technology, and particularly to the field of intelligent voice technology. Specifically, it relates to a voice interaction method, a voice interaction device, an electronic device, a storage medium, and a program product. Background Art

[0002] With the development of technology, voice interaction technology is widely used in various intelligent voice devices such as intelligent robots, intelligent speakers, intelligent vehicles, and intelligent appliances. The intelligent voice device can perform corresponding operations according to the interaction voice issued by the user, such as answering questions in the user's interaction voice, starting or stopping the device, etc. Summary of the Invention

[0003] The present disclosure provides a method for voice interaction, a device for voice interaction, an electronic device, a storage medium, and a program product.

[0004] According to one aspect of the present disclosure, there is provided a voice interaction method, including: in response to receiving an interaction voice from a user, performing semantic analysis on the interaction voice to obtain user intention information; based on the user intention information, determining a target voice response mode; and outputting voice response information according to the target voice response mode.

[0005] According to another aspect of the present disclosure, there is provided a voice interaction device, including: an intention determination module, configured to perform semantic analysis on the interaction voice in response to receiving an interaction voice from a user to obtain user intention information; a response mode determination module, configured to determine a target voice response mode based on the user intention information; and a response output module, configured to output voice response information according to the target voice response mode.

[0006] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.

[0007] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method as described above.

[0008] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements the method as described above.

[0009] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. Description of the Drawings

[0010] The drawings are used to better understand the present solution and do not constitute a limitation to the present disclosure. Among them:

[0011] Figure 1 Schematically shows an exemplary system architecture to which the voice interaction method and apparatus according to an embodiment of the present disclosure can be applied;

[0012] Figure 2 Schematically shows a flowchart of the voice interaction method according to an embodiment of the present disclosure;

[0013] Figure 3 Schematically shows an application scenario diagram of the voice interaction method according to an embodiment of the present disclosure;

[0014] Figure 4 Schematically shows an application scenario diagram of a user performing voice interaction according to an embodiment of the present disclosure;

[0015] Figure 5 Schematically shows a block diagram of the voice interaction apparatus according to an embodiment of the present disclosure;

[0016] Figure 6 Shows a schematic block diagram of an example electronic device that can be used to implement the embodiments of the present disclosure. Detailed Embodiments

[0017] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0018] The present disclosure provides a voice interaction method, a voice interaction apparatus, an electronic device, a storage medium, and a program product.

[0019] According to an embodiment of the present disclosure, the voice interaction method includes: in response to receiving an interaction voice from a user, performing semantic analysis on the interaction voice to obtain user intention information; determining a target voice response manner based on the user intention information; and outputting voice response information according to the target voice response manner.

[0020] According to an embodiment of the present disclosure, by performing semantic analysis on the user's interactive voice to determine the user intention information, and based on the user intention information, determining the target voice response manner, the target voice response manner can be determined on the basis of fully considering the user intention information; outputting the voice response information according to the target voice response manner can meet the personalized needs of the user according to the user intention, and improve the interest and the user experience.

[0021] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information are all in compliance with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0022] Figure 1 An exemplary system architecture to which the voice interaction method and apparatus according to an embodiment of the present disclosure can be applied is schematically shown.

[0023] It should be noted that Figure 1 The illustration is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios. For example, in another embodiment, the exemplary system architecture to which the voice interaction method and apparatus can be applied may include a terminal device, but the terminal device can implement the voice interaction method and apparatus provided by the embodiments of the present disclosure without interacting with the server.

[0024] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0025] The user can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as examples).

[0026] The terminal devices 101, 102, 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0027] The server 105 may be a server that provides various services, such as a background management server (merely an example) that supports the content browsed by users using the terminal devices 101, 102, and 103. The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0028] It should be noted that the voice interaction method provided by the embodiments of the present disclosure can generally be executed by the terminal devices 101, 102, or 103. Correspondingly, the voice interaction device provided by the embodiments of the present disclosure can also be disposed in the terminal devices 101, 102, or 103.

[0029] Alternatively, the voice interaction method provided by the embodiments of the present disclosure can generally also be executed by the server 105. Correspondingly, the voice interaction device provided by the embodiments of the present disclosure can generally be disposed in the server 105. The voice interaction method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Correspondingly, the voice interaction device provided by the embodiments of the present disclosure can also be disposed in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.

[0030] For example, when the user issues an interaction voice, the terminal devices 101, 102, 103 can acquire the interaction voice from the user, and then send the acquired interaction voice from the user to the server 105. The server 105 performs semantic analysis on the interaction voice to obtain user intention information; based on the user intention information, determines the target voice response method; and outputs the voice response information according to the target voice response method. Or a server or a server cluster capable of communicating with the terminal devices 101, 102, 103 and / or the server 105 performs semantic analysis on the interaction voice, and finally realizes outputting the voice response information according to the target voice response method.

[0031] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the servers in

[0032] Figure 2 are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.

[0033] As Figure 2 shown, the method includes operations S210 to S230.

[0034] In operation S210, in response to receiving an interactive voice from a user, semantic analysis is performed on the interactive voice to obtain user intent information.

[0035] In operation S220, based on the user intent information, a target voice response mode is determined.

[0036] In operation S230, voice response information is output according to the target voice response mode.

[0037] According to an embodiment of the present disclosure, the user intent information may include information reflecting the user's needs. For example, if the interactive voice from the user is "Please play the songs of singer Axx.", semantic analysis of this interactive voice can determine that the user intent information includes information reflecting the user's need to enjoy the songs of singer Axx. However, it is not limited to this. The user intent information may also be information reflecting other needs of the user, such as information on the need to obtain weather information, information on the need to listen to audiobooks, etc.

[0038] It should be noted that semantic analysis can be performed on the interactive voice through a neural network module in related technologies. The neural network module can be, for example, constructed based on a long short-term memory network, or can also be a model constructed based on a statistical algorithm. The specific technical means for performing semantic analysis on the interactive voice in the embodiments of the present disclosure are not limited.

[0039] According to an embodiment of the present disclosure, the voice response mode may refer to the voice response style, the speaker of the voice response, etc. For example, the voice response style may include styles in aspects such as intonation, speech rate, and volume, and the speaker of the voice response may refer to the timbre characteristics of the speaker, etc.

[0040] According to an embodiment of the present disclosure, the target voice response mode may refer to a voice response mode that matches the user intent information. For a variety of different user intent information, target voice response modes that respectively match the various different user intent information can be matched.

[0041] According to an embodiment of the present disclosure, the voice response information may include voice information for answering the user intent information. For example, when the user intent information is information reflecting the user's need to enjoy the songs of singer Axx, the voice response information may be "Okay, here is the song 'ABC' of singer Axx for you."

[0042] According to an embodiment of the present disclosure, the target voice response manner can be determined by determining the obtained user intention information, and the target voice response manner matching the user intention information can be adjusted according to different user intention information. Thus, by outputting the voice response information according to the target voice response manner, the voice response information can have different voice information characteristics, such as different voice response styles, different emotional colors, etc. Furthermore, the personalized needs of users can be met, and the interest and user experience can be improved.

[0043] The following will further illustrate the data evaluation method of the embodiments of the present disclosure in conjunction with specific embodiments and with reference to Figures 3 to 4 ,.

[0044] According to an embodiment of the present disclosure, for operation S220, determining the target voice response manner based on the user intention information may include:

[0045] Determining the application scenario information based on the user intention information; and determining the target voice response manner based on the application scenario information.

[0046] According to an embodiment of the present disclosure, the application scenario information may refer to the category information of the terminal application program or the scenario of performing operations using the terminal application program. For example, a map application program is used to play navigation information, a music playback application program is used to play music, an audiobook playback application program is used to play novels, or a query application program is used to query weather information. Based on different user intention information, different application scenario information can be determined, and the application scenario information can correspond to the user intention information to meet the user's needs. For example, when the user's need reflected by the user intention information is to appreciate the songs of singer Axx, the application scenario information may be to use a music playback application program to play songs.

[0047] According to an embodiment of the present disclosure, when the application scenario information is to use a map application program to play navigation information, the target voice response manner can be determined as an urgent response style. For example, by increasing the speaking speed to characterize the urgency of the response style, and outputting the voice response information according to the urgent response style can improve the user's attention to the voice response information and serve as a warning to the user.

[0048] According to an embodiment of the present disclosure, the corresponding target voice response manner can be determined based on different application scenario information, thereby enhancing the applicability of the target voice response manner to different application scenario information and meeting the personalized needs of users.

[0049] According to an embodiment of the present disclosure, the voice interaction method may further include, after operation S230, that is, after outputting the voice response information according to the target voice response manner:

[0050] Based on the user intention information, determine the subject content information of the operation to be performed.

[0051] According to an embodiment of the present disclosure, the subject content information of the operation to be performed may include the played voice information. For example, it may include the song content information of the song "ABC" sung by singer Axx, but is not limited thereto. The subject content information may also include navigation content information, weather forecast information, cross-talk content information, or audiobook content information, etc.

[0052] According to an embodiment of the present disclosure, different user intention information, which can reflect different user demand information, based on the user intention information, determine the corresponding subject content information of the operation to be performed, can meet the personalized needs of users and improve the user experience.

[0053] Figure 3 Schematically shows an application scenario diagram of the voice interaction method according to an embodiment of the present disclosure.

[0054] As Figure 3 shown, the interactive voice 310 from the user may be "Navigate from AA Mall to BB Cinema." In response to receiving the interactive voice 310 from the user, perform semantic analysis on the interactive voice 310, and it can be determined that the user intention information includes the demand for reflecting the user to obtain navigation information. According to this user intention information, it can be determined that the application scenario information is to play navigation information, and then the target voice response method can be determined based on the application scenario information. For example, it may include the accent information of the local dialect in the navigation area. Output the voice response information 320 according to the target voice response method of the local dialect in the navigation area: "Okay, start navigating for you." Then, according to the user intention information, determine the subject content information 330 of the operation to be performed as: "Please go straight ahead,...".

[0055] According to an embodiment of the present disclosure, by determining the user intention information, for example, reflecting the user's demand for obtaining navigation information, it can be determined that the application scenario information is to play navigation information using a map application. According to the application scenario information, determine the target voice response method as, for example, using the local dialect voice response method in the navigation area, and output the voice response information according to this target voice response method. Thus, while meeting the user's demand for obtaining navigation information, provide the user with the voice response information output according to the target voice response method of the local dialect in the navigation area, and further meet the personalized needs of users, improving the fun and user experience. According to an embodiment of the present disclosure, the target voice response method includes the response style.

[0056] Determining a target voice response manner based on application scenario information may include: in response to the application scenario information being first predetermined scenario information, identifying the emotional information in the theme content information; and based on the emotional information, determining the response style of the target voice response manner.

[0057] According to an embodiment of the present disclosure, it may be determined whether the application scenario information matches the first predetermined scenario information, and in the case where it is determined that the application scenario information is the first predetermined scenario information, the target voice response manner is determined according to the rule matching the first predetermined scenario information.

[0058] For example, the first predetermined scenario information may include playing music using a song playing application, playing a novel using an audiobook playing application, etc. The theme content information may include songs, cross talks corresponding to the first predetermined scenario information. In the case where the application scenario information is the first predetermined scenario information, the emotional information in the theme content information may be identified. For example, the emotional information in the theme content information of the song or novel to be played may be cheerful, sad, etc.

[0059] According to an embodiment of the present disclosure, the response style may include a response style representing voice emotion. For example, it may be a response style representing cheerfulness or sadness. Based on the emotional information in the theme content information, determining the response style of the target voice response manner can unify the response style with the emotional information in the theme content information, so that the voice response information is consistent with the emotion of the theme content information of the subsequent execution operation, can render the emotional color in advance, and further improve the intelligence and interest of voice interaction.

[0060] It should be noted that the emotional information in the theme content information may be identified by parsing the semantic information of the theme content information. For example, the semantic information of the theme content information is parsed through a neural network model to identify the emotional information in the theme content information, but not limited thereto. The emotional information in the theme content information may also be identified by parsing the frequency and amplitude of the sound in the theme content information. The specific technical means for identifying the emotional information in the theme content information in the embodiments of the present disclosure are not limited.

[0061] According to an embodiment of the present disclosure, the target voice response manner includes response voice characteristics.

[0062] Determining a target voice response manner based on application scenario information includes:

[0063] In response to the application scenario information being second predetermined scenario information, identifying the voice characteristics of the speaker in the theme content information; and based on the voice characteristics of the speaker, determining the response voice characteristics of the target voice response manner.

[0064] According to an embodiment of the present disclosure, the second predetermined scenario information may include playing music using a song-playing application, playing a novel using an audiobook-playing application, and the like. The theme content information may include voice information of a song or cross-talk corresponding to the second predetermined scenario information. The voice characteristics of the speaker may include the timbre of the speaker.

[0065] For example, when the second predetermined scenario information is playing a novel using an audiobook-playing application, the theme content information may be the cross-talk work "EDF" by the cross-talk actor Byy, and the speaker of the cross-talk work "EDF" may be Byy. By identifying the voice characteristics of the speaker Byy in the cross-talk work "EDF", the voice characteristics of the speaker Byy can be determined as the response voice characteristics of the target voice response manner.

[0066] It should be noted that the voice characteristics of the speaker in the theme content information can be identified by analyzing the voiceprint characteristics of the voice information in the theme content information.

[0067] According to an embodiment of the present disclosure, based on the voice characteristics of the speaker, determining the response voice characteristics of the target voice response manner may be using the voice characteristics of the speaker as the response voice characteristics. By outputting voice response information according to the voice characteristics of the speaker in the theme content information, the voice response information has the same voice characteristics as the theme content information, thereby improving the intelligence of voice interaction.

[0068] According to an exemplary embodiment of the present disclosure, the target voice response manner includes response voice characteristics and response styles. In the case where the application scenario information is both the first predetermined scenario information and the second predetermined scenario information, the emotional information in the theme content information can be identified. Based on the emotional information, the response style of the target voice response manner is determined. Identify the voice characteristics of the speaker in the theme content information; and based on the voice characteristics of the speaker, determine the response voice characteristics of the target voice response manner. For example, the target voice response manner may include a happy response style and female response voice characteristics. Thus, outputting voice response information according to the response voice characteristics and response styles may be outputting voice response information according to the target voice response manner of "happy girl".

[0069] It should be understood that the first predetermined scenario information and the second predetermined scenario information may be the same or different, and those skilled in the art can set the first predetermined scenario information and the second predetermined scenario information according to actual needs.

[0070] Figure 4 An application scenario diagram of a user performing voice interaction according to an embodiment of the present disclosure is schematically shown.

[0071] As Figure 4As shown, the interactive voice 411 sent by the user 410 can be "Please play the song 'ABC' by singer Axx". The voice interaction device can be set inside the vehicle 420. For example, the voice interaction device is an in-vehicle terminal device. It should be understood that the user 410 can be inside the vehicle 420, or the user 410 can also be outside the vehicle 420, as long as the interactive voice 411 sent by the user 410 can be received by the voice interaction device.

[0072] The voice interaction device of the vehicle 420 can perform semantic analysis on the interactive voice 411, and can determine that the user intention information includes the need to reflect the user's appreciation of the song by singer Axx. According to this user intention information, it can be determined that the first predetermined application scenario can be to play music using a music playback application program, and according to this user intention information, it can be determined that the theme content information is the song 'ABC' 422.

[0073] The emotional information in the song 'ABC' 422 is identified as sad. Based on the emotional information in the song 'ABC' 422, it can be determined that the response style can be sad. According to this response style, it can be determined that the target voice response method includes a sad response style, and then a voice response message 421 with a sad response style is output: "Okay, playing the song 'ABC' by Axx for you."

[0074] Further, when the second predetermined application scenario is the same as the first predetermined application scenario, the voice characteristics of the speaker in the voice information 422 of the song 'ABC' can also be identified, that is, the timbre of the singer Axx of the song 'ABC' is identified. Based on the voice characteristics of the speaker, it can be determined that the response voice characteristics are the timbre of the singer Axx. According to this response voice characteristics, it can be determined that the target voice response method includes the timbre of the singer Axx. Then, a voice response message 421 can be output according to the sad response style and according to the timbre of the singer Axx: "Okay, playing the song 'ABC' by Axx for you." After the voice response message 421 is output, the song 'ABC' 422 can be played.

[0075] By unifying the response style with the emotional information of the song 'ABC' and outputting the voice response message according to the timbre of the speaker of the song 'ABC', it can help the user enter the psychological state of appreciating the song 'ABC' in advance before playing the voice information of the song 'ABC', thereby further meeting the personalized needs of the user and improving the user experience.

[0076] According to an embodiment of the present disclosure, outputting the voice response message according to the target voice response method includes:

[0077] In response to searching for a target voice response mode from multiple template voice response modes, output voice response information according to the target voice response mode; and in response to not searching for a target voice response mode from multiple template voice response modes, determine a template voice response mode from multiple template voice response modes as the target voice response mode according to the attention recommendation rule.

[0078] According to an embodiment of the present disclosure, the attention can include popularity or heat, and can be the attention generated based on user browsing operations, retrieval operations, or like operations. The attention recommendation rule can be to determine a template voice response mode as the template voice response mode from the N template voice response modes with the highest attention. For example, if the target voice response mode includes the voice of singer Axx and the target voice response mode including the voice of singer Axx is not searched for among multiple template voice response modes, the template voice response mode with the highest attention can be determined from multiple template voice response modes in the cloud server as the target voice response mode, and voice response information can be output according to this target voice response mode. However, it is not limited to this. A template voice response mode can also be determined from multiple template voice response modes according to the user's interest level as the target voice response mode.

[0079] Using the voice interaction method provided by the embodiment of the present disclosure, multiple determination channels for the target voice response mode are provided, enriching the diversity of the target voice response mode, expanding the applicable range of the voice interaction method, and thus enhancing the user experience.

[0080] Figure 5 A block diagram of a voice interaction device according to an embodiment of the present disclosure is schematically shown.

[0081] As Figure 5 shown, the voice interaction device 500 may include: an intention determination module 510, a response mode determination module 520, and a response output module 530.

[0082] The intention determination module 510 is configured to, in response to receiving an interaction voice from a user, perform semantic analysis on the interaction voice to obtain user intention information.

[0083] The response mode determination module 520 is configured to determine a target voice response mode based on the user intention information.

[0084] The response output module 530 is configured to output voice response information according to the target voice response mode.

[0085] According to an embodiment of the present disclosure, the response determination module includes: an application scenario determination sub-module and a response mode determination sub-module.

[0086] An application scenario determination sub-module, configured to determine application scenario information based on user intention information.

[0087] A response mode determination sub-module, configured to determine a target voice response mode based on the application scenario information.

[0088] According to an embodiment of the present disclosure, the voice interaction device further includes, after outputting voice response information according to the target voice response mode:

[0089] An operation determination module, configured to determine the subject content information of the operation to be performed based on the user intention information.

[0090] According to an embodiment of the present disclosure, the target voice response mode includes a response style.

[0091] The response mode determination sub-module includes: a first recognition unit and a response style determination unit.

[0092] The first recognition unit is configured to recognize the emotional information in the subject content information in response to the application scenario information being the first predetermined scenario information.

[0093] The response style determination unit is configured to determine the response style of the target voice response mode based on the emotional information.

[0094] According to an embodiment of the present disclosure, the target voice response mode includes response voice features;

[0095] The response mode determination sub-module includes: a second recognition unit and a voice feature determination unit.

[0096] The second recognition unit is configured to recognize the voice features of the speaker in the subject content information in response to the application scenario information being the second predetermined scenario information.

[0097] The voice feature determination unit is configured to determine the response voice features of the target voice response mode based on the voice features of the speaker.

[0098] According to an embodiment of the present disclosure, the response output module includes: a first voice response sub-module and a second voice response sub-module.

[0099] The first voice response sub-module is configured to output voice response information according to the target voice response mode in response to searching for the target voice response mode from multiple template voice response modes.

[0100] The second voice response sub-module is configured to determine a template voice response mode as the target voice response mode from multiple template voice response modes according to the attention recommendation rule in response to not searching for the target voice response mode from multiple template voice response modes.

[0101] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0102] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.

[0103] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.

[0104] According to an embodiment of the present disclosure, a computer program product includes a computer program that, when executed by a processor, implements the method as described above.

[0105] Figure 6 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smartphone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0106] As Figure 6 shown, the device 600 includes a computing unit 601 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0107] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as a keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as a disk, optical disc, etc.; and communication unit 609, such as a network card, modem, wireless communication transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0108] Computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 601 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 601 executes the various methods and processes described above, such as the voice interaction method. For example, in some embodiments, the voice interaction method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by computing unit 601, one or more steps of the voice interaction method described above can be executed. Alternatively, in other embodiments, computing unit 601 can be configured to execute the voice interaction method by any other suitable means (e.g., by means of firmware).

[0109] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0110] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0111] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0112] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0113] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0114] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0115] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0116] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A voice interaction method, comprising: responding to an interaction voice received from a user, performing semantic analysis on the interaction voice to obtain user intent information; determining application scenario information based on the user intent information, wherein the application scenario information characterizes the category of a terminal application program for performing an operation; responding to the application scenario information being second predetermined scenario information, identifying the voice characteristics of the speaker in the subject content information of the operation to be performed, wherein the subject content information is determined based on the user intent information; determining the response voice characteristics of a target voice response manner based on the voice characteristics of the speaker; and outputting voice response information according to the target voice response manner.

2. The method according to claim 1, further comprising, after outputting the voice response information according to the target voice response manner: determining the subject content information of the operation to be performed based on the user intent information.

3. The method according to claim 2, wherein The target voice response manner includes a response style; The method further comprises: responding to the application scenario information being first predetermined scenario information, identifying the emotion information in the subject content information; and determining the response style of the target voice response manner based on the emotion information.

4. The method according to any one of claims 1 to 3, wherein, The outputting the voice response information according to the target voice response manner includes: responding to the target voice response manner being searched from a plurality of template voice response manners, outputting the voice response information according to the target voice response manner; and responding to the target voice response manner not being searched from the plurality of template voice response manners, determining a template voice response manner from the plurality of template voice response manners as the target voice response manner according to the attention recommendation rule.

5. A voice interaction device, comprising: an intent determination module, configured to respond to an interaction voice received from a user, perform semantic analysis on the interaction voice to obtain user intent information; a response manner determination module, configured to determine a target voice response manner based on the user intent information; and a response output module, configured to output voice response information according to the target voice response manner; wherein, the response manner determination module includes: an application scenario determination sub-module, configured to determine application scenario information based on the user intent information, wherein the application scenario information characterizes the category of a terminal application program for performing an operation; and a response manner determination sub-module, configured to determine a target voice response manner based on the application scenario information; wherein, the target voice response manner includes response voice characteristics; the response manner determination sub-module includes: a second identification unit, configured to respond to the application scenario information being second predetermined scenario information, identify the voice characteristics of the speaker in the subject content information of the operation to be performed, wherein the subject content information is determined based on the user intent information; and a voice characteristic determination unit, configured to determine the response voice characteristics of the target voice response manner based on the voice characteristics of the speaker.

6. The apparatus according to claim 5, further comprising, after outputting the voice response information according to the target voice response manner: An operation determination module, configured to determine the subject content information of the operation to be performed based on the user intention information.

7. The device according to claim 6, wherein, The target voice response manner includes a response style; The response manner determination sub-module includes: A first recognition unit, configured to recognize the emotion information in the subject content information in response to the application scenario information being the first predetermined scenario information; And A response style determination unit, configured to determine the response style of the target voice response manner based on the emotion information.

8. The device according to any one of claims 5 to 7, wherein, The response output module includes: A first voice response sub-module, configured to output voice response information according to the target voice response manner in response to searching for the target voice response manner from a plurality of template voice response manners; and A second voice response sub-module, configured to determine a template voice response manner from the plurality of template voice response manners as the target voice response manner according to the attention recommendation rule in response to not searching for the target voice response manner from the plurality of template voice response manners.

9. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 4.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 4.

11. A computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Intelligent voice interaction realization method and device, computer equipment and storage medium

    CN108597509A

  • Speech synthesis method and related equipment

    CN108962217A

  • Voice interaction method, device and system, and voice processing method and device

    CN109346076A