A method, device, electronic device and storage medium for voice interaction

By identifying voice information, determining target applications and generating enable commands, the problem of insufficient expressiveness in existing intelligent voice interaction devices when handling complex voice tasks is solved, and higher voice interaction accuracy and user experience are achieved.

CN113555014BActive Publication Date: 2025-05-20BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010329240.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-04-23
Publication Date
2025-05-20
Estimated Expiration
2040-04-23

AI Technical Summary

Technical Problem

Existing smart voice interaction devices with screens are poor in handling complex voice interaction tasks, making it difficult to meet users' complex needs such as shopping or search.

Method used

Through methods of voice information recognition, target application determination and command generation, the target application that matches the user's intentions are automatically selected and opened to improve the accuracy of voice interaction.

Benefits of technology

It realizes automatic selection and opening of target applications based on voice information, so that the enabled applications meet users' psychological expectations, and improves the accuracy and user experience of voice interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113555014B_ABST
    Figure CN113555014B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, electronic device and storage medium for voice interaction, and relates to the field of intelligent voice interaction. The received voice information is recognized to obtain a recognition result; according to the recognition result, a target application is determined from candidate applications that match the recognition result; and an instruction to open the target application is generated according to the recognition result. Through the above scheme, the target application can be automatically selected and opened according to the voice information. The opened target application meets the user's psychological expectations and improves the accuracy of voice interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and particularly to the field of intelligent voice interaction. Background Art

[0002] Smart voice interaction devices with screens have appeared in more and more households. Existing technologies only support relatively simple voice interactions, such as querying the weather, querying the time, etc. When users have more complex needs such as shopping or searching, existing smart voice interaction devices with screens often have poor performance. Summary of the Invention

[0003] Embodiments of this application provide a method, device, electronic device, and storage medium for voice interaction to solve one or more technical problems in the prior art.

[0004] In a first aspect, this application provides a method for voice interaction, including the following steps:

[0005] Recognize the received voice information to obtain a recognition result;

[0006] Determine a target application program from the candidate application programs that match the recognition result according to the recognition result;

[0007] Generate an instruction to start the target application program according to the recognition result.

[0008] Through the above solution, the target application program can be automatically selected and started according to the voice information. The started target application program meets the user's psychological expectations and improves the accuracy of voice interaction.

[0009] In a second aspect, this application provides a device for voice interaction, including the following components:

[0010] A voice information recognition module for recognizing the received voice information to obtain a recognition result;

[0011] A target application program determination module for determining a target application program from the candidate application programs that match the recognition result according to the recognition result;

[0012] An instruction generation module for generating an instruction to start the target application program according to the recognition result.

[0013] In a third aspect, embodiments of this application provide an electronic device, including:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores instructions executable by at least one processor. The instructions are executed by at least one processor to enable the at least one processor to execute the method provided in any embodiment of the present application.

[0017] In a fourth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method provided in any embodiment of the present application.

[0018] Other effects of the above optional manners will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings are used to better understand the solution and do not limit the present application. Among them:

[0020] Figure 1 is a flowchart of a method for voice interaction according to an embodiment of the present application;

[0021] Figure 2 is a schematic diagram of an interaction interface presented after a target application is opened according to an embodiment of the present application;

[0022] Figure 3 is a flowchart of a method for determining a user's usage habit according to an embodiment of the present application;

[0023] Figure 4 is a schematic diagram of a device for voice interaction according to an embodiment of the present application;

[0024] Figure 5 is a schematic diagram of a target application determination module according to an embodiment of the present application;

[0025] Figure 6 is a block diagram of an electronic device for implementing the method for voice interaction in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The following describes exemplary embodiments of the present application with reference to the accompanying drawings. Various details of the embodiments of the present application are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0027] As Figure 1 shown, in one implementation manner, a method for voice interaction is provided, including the following steps:

[0028] S101: Identify the received voice information to obtain an identification result.

[0029] S102: Determine a target application from the candidate applications that match the recognition result according to the recognition result.

[0030] S103: Generate an instruction to open the target application according to the recognition result.

[0031] The above method in the embodiments of this application can be executed by a screen-equipped intelligent voice interaction device or by a cloud server. Taking the execution by the cloud server as an example, when the screen-equipped intelligent voice interaction device receives voice information, it sends the voice information to the cloud server. The cloud server performs recognition on it to obtain a recognition result. The recognition result may include relevant information of the user and / or the user's intention, etc.

[0032] The relevant information of the user may include the user's identity, usage habits or preferences for candidate applications, etc.

[0033] The user's intention may include a wake-up intention, a query intention, etc.

[0034] For example, the voice information may be "Xiaodu, Xiaodu". In this case, it can be determined that it is a wake-up intention, and then the voice interaction mode can be entered.

[0035] For another example, the voice information may be "I want to buy something". In this case, although the wake-up word is not included, through intention recognition, it can be determined that the voice information is a voice interaction instruction issued by the user. Then, the shopping application that matches the instruction can be opened to query items to respond to the instruction.

[0036] In addition, the recognition result of the voice information may also be noise, such as ambient background noise or the user's phone voice, etc. In this case, no corresponding operation is performed.

[0037] For the cases where the recognition result is noise or wake-up intention, they are not specifically described in this embodiment.

[0038] According to the recognition result of the voice information, candidate applications that match the recognition result can be determined, and a target application can be determined from the candidate applications. For example, the voice information may be "Play a song for me", then the corresponding recognition result is to play a song. Based on this, music-playing applications can be determined as candidate applications. For another example, the voice information may be "I want to buy something", then the corresponding recognition result is shopping. Based on this, shopping applications can be determined as candidate applications.

[0039] In the case where multiple candidate applications that match the recognition result are pre-installed in the screen-equipped intelligent voice interaction device, one of them can be randomly selected as the determined target application.

[0040] In addition, the target application can be determined according to the usage habits of the current user. The usage habits can be obtained through the operation records in the screen-equipped intelligent voice interaction device, or through the background server of the candidate application.

[0041] Taking a shopping application as an example of the candidate application, for example, in the case where multiple shopping applications are pre-installed in the screen-equipped intelligent voice interaction device. The candidate application that the current user often uses can be used as the target application.

[0042] In addition, the target application can also be determined according to the characteristics of the candidate application.

[0043] The characteristics of the candidate application can be activities such as limited-time free experience launched by the candidate application, shopping discounts, etc.

[0044] Taking a music-playing candidate application as an example, by using technologies such as web crawlers, it can be obtained that the first candidate application conducts a limited-time free audition activity, then the first candidate application can be used as the target application. Or, taking a shopping candidate application as an example, by using technologies such as web crawlers, the promotion activities of each shopping candidate application can be obtained. For example, the promotion activity of the first candidate application is "200 off 40 for the whole store", and the promotion activity of the second candidate application is "10% off for the whole store". By comparison, the candidate application with a greater promotion intensity can be selected as the target application.

[0045] In addition, the target application can also be determined according to the usage habits of the current user and the characteristics of the candidate application.

[0046] For example, according to the usage habits of the current user in the past month, the first score values of two candidate applications are calculated. According to the characteristics of the candidate application, the second score values of the two candidate applications are calculated. Further, weights can also be set for the usage habits of the user and the characteristics of the candidate application. Finally, the cumulative result of the score values calculated using the weights is used as the final score value of the two candidate applications. The candidate application with the higher final score value is used as the target application.

[0047] Convert the recognition result of the voice information into an instruction to open the target application. As Figure 2 Shown is a schematic diagram of the interaction interface presented after the target application is opened. After the target application is opened, a voice message can also be broadcast, such as "Welcome to XX Shopping, select XX high-quality goods, come and shop quickly".

[0048] Through the above solution, the target application can be automatically selected and opened according to the voice information. The opened target application meets the user's psychological expectations and improves the accuracy of voice interaction.

[0049] In one implementation, step S102 includes:

[0050] When the recognition result includes the query intention and the user's usage habit, determine a target application that matches the query intention from the candidate applications according to the user's usage habit.

[0051] Based on the voice information, the query intention and the user's usage habit of the current user can be obtained. For example, when the identity of the current user is determined according to the voice information, the operation record of the current user can be obtained. The usage habit of the user can be determined according to the operation record, so as to determine the candidate application that conforms to the user's usage habit as the target application.

[0052] Take shopping as an example of the query intention. When multiple shopping applications are pre-installed in the screen-equipped intelligent voice interaction device, then multiple shopping applications can all be used as candidate applications. For example, when it is determined through the operation record that the user prefers to place an order for shopping from the first candidate application, the first candidate application can be determined as the target application.

[0053] Through the above solution, the target application can be determined in combination with the user's usage habit, so that the determined target application better meets the user's psychological expectations.

[0054] As Figure 3 shown, in one implementation, the method for determining the user's usage habit includes:

[0055] S301: Identify the current user according to the voice information.

[0056] S302: Determine the user's usage habit according to the time for the current user to browse the candidate applications and / or according to the frequency of the current user to open the candidate applications.

[0057] A matching relationship pair between the user and the voiceprint can be established in advance. By performing voiceprint recognition on the voice information, the current user can be determined.

[0058] The usage habit of the current user can be determined according to the operation record of the current user. For example, the usage habit of the current user is determined according to records such as the time for the current user to browse the candidate applications and the frequency of the current user to open the candidate applications. That is, the candidate application that the user prefers more is determined.

[0059] Through the above solution, the target application can be determined more accurately in combination with the user's usage habit, so that the determined target application better meets the user's psychological expectations.

[0060] In one implementation, it further includes:

[0061] Continuously recognize the received voice information, and when the recognition result is to adjust the display state of the target application, send the instruction corresponding to the recognition result to the target application.

[0062] Adjusting the display state of the target application may include entering the next page, returning to the previous page, returning to the home page, exiting the target application, etc. Convert the above recognition result into a control instruction for the target application and send it to the target application. Thus, direct voice control of the target application can be achieved.

[0063] Through the above solution, the target application can be directly controlled by voice information. Especially for target applications that do not support voice information control, the above technical solution can also upgrade them to applications that can respond to voice information control.

[0064] In one embodiment, the target application is a shopping application.

[0065] Through the above solution, when the recognition result includes a shopping intention, the target application that meets the shopping intention can be directly recommended to the user without the user performing other operations. It can more intelligently meet the user's shopping needs on the screen-enabled intelligent voice interaction device.

[0066] As Figure 4 shown, in one embodiment, a voice interaction device is provided, including the following components:

[0067] A voice information recognition module 401, configured to recognize the received voice information to obtain a recognition result.

[0068] A target application determination module 402, configured to determine a target application from candidate applications that match the recognition result according to the recognition result.

[0069] An instruction generation module 403, configured to generate an instruction to start the target application according to the recognition result.

[0070] In one embodiment, the target application determination module 402 is specifically configured to: when the recognition result includes a query intention and user usage habits, determine a target application that matches the query intention from the candidate applications according to the user usage habits.

[0071] As Figure 5 shown, in one embodiment, the target application determination module 402 includes:

[0072] A current user recognition sub-module 4021, configured to recognize the current user according to the voice information;

[0073] A user usage habit determination sub-module 4022 is configured to determine the user usage habit according to the time when the current user browses the candidate application and / or according to the frequency of the current user starting the candidate application.

[0074] In one implementation, it further includes:

[0075] A voice information recognition module is configured to continuously recognize the received voice information, and when the recognition result is to adjust the display state of the target application, send the instruction corresponding to the recognition result to the target application.

[0076] In one implementation, the target application is a shopping application.

[0077] For the functions of the modules in each device of the embodiments of the present application, reference may be made to the corresponding descriptions in the above methods, which will not be elaborated herein.

[0078] According to the embodiments of the present application, the present application also provides an electronic device and a readable storage medium.

[0079] As Figure 6 shown, it is a block diagram of an electronic device for a voice interaction method according to an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processing, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0080] As Figure 6 shown, the electronic device includes: one or more processors 610, a memory 620, and an interface for connecting each component, including a high-speed interface and a low-speed interface. Each component is interconnected using different buses and can be installed on a common motherboard or otherwise installed as needed. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In other embodiments, if necessary, multiple processors and / or multiple buses can be used with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (such as, as a server array, a group of blade servers, or a multi-processor system). Figure 6 One processor 610 is taken as an example herein.

[0081] The memory 620 is the non-transitory computer-readable storage medium provided by this application. Among them, the memory stores instructions executable by at least one processor, so that the at least one processor executes the voice interaction method provided by this application. The non-transitory computer-readable storage medium of this application stores computer instructions, and these computer instructions are used to make a computer execute the voice interaction method provided by this application.

[0082] As a non-transitory computer-readable storage medium, the memory 620 can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the voice interaction method in the embodiments of this application (for example, Figure 4 the voice information recognition module 401, the target application determination module 402, and the instruction generation module 403 shown in the appendix). By running the non-transitory software programs, instructions, and modules stored in the memory 620, the processor 610 executes various functional applications and data processing of the server, that is, implements the voice interaction method in the above method embodiments.

[0083] The memory 620 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the electronic device for voice interaction, etc. In addition, the memory 620 may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 620 may optionally include a memory remotely set relative to the processor 610, and these remote memories can be connected to the electronic device for voice interaction through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0084] The electronic device for the voice interaction method may further include: an input device 630 and an output device 640. The processor 610, the memory 620, the input device 630, and the output device 640 can be connected through a bus or other means, Figure 6 taking the connection through the bus as an example.

[0085] The input device 630 can receive input digital or character information and generate key signal inputs related to user settings and function controls of the electronic device for voice interaction, such as input devices like touchscreens, keypads, mice, trackpads, touchpads, pointing sticks, one or more mouse buttons, trackballs, joysticks, etc. The output device 640 can include display devices, auxiliary lighting devices (e.g., LEDs), and tactile feedback devices (e.g., vibration motors), etc. The display device can include, but is not limited to, liquid crystal displays (LCDs), light emitting diode (LED) displays, and plasma displays. In some embodiments, the display device can be a touchscreen.

[0086] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuitry, integrated circuit systems, dedicated ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0087] These computational programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, device, and / or apparatus (e.g., disks, optical disks, memories, programmable logic devices (PLDs)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal for providing machine instructions and / or data to a programmable processor.

[0088] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0089] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0090] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other.

[0091] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved, and no limitation is imposed herein.

[0092] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the protection scope of this application.

Claims

1. A voice interaction method, characterized in that: include: Recognize the received voice information and obtain a recognition result; According to the recognition result, determining a target application from candidate applications matching the recognition result; Generate an instruction to start the target application according to the recognition result; The step of determining a target application from candidate applications matching the recognition result includes: Determine the target application according to the current user's usage habits and the characteristics of the candidate application; wherein the characteristics of the candidate application are limited-time free trials and shopping discounts launched by the candidate application; The target application is determined according to the current user's usage habits and the characteristics of the candidate application, including: Weights are set for the current user's usage habits and the characteristics of the candidate applications; the cumulative result of the score values ​​calculated using the weights is used as the final score value of the candidate applications, and the candidate application with the highest final score value is used as the target application; Wherein, the method further comprises: Continuously recognizing the received voice information, and when the recognition result is to adjust the display state of the target application, sending an instruction corresponding to the recognition result to the target application; Among them, the methods for determining user usage habits include: Identifying the current user according to the voice information; Determining the user usage habits according to the time when the current user browses the candidate application and / or according to the frequency with which the current user opens the candidate application; Wherein, identifying the current user according to the voice information includes: A pre-established matching relationship pair between a user and a voiceprint is obtained; based on the matching relationship pair, voiceprint recognition is performed on the voice information to determine the current user.

2. The method according to claim 1, characterized in that Determining a target application from candidate applications matching the recognition result according to the recognition result also includes: In the case where the recognition result includes the query intention and the user's usage habits, a target application matching the query intention is determined from the candidate applications according to the user's usage habits.

3. The method according to any one of claims 1 to 2, characterized in that: The target application is a shopping application.

4. A voice interaction device, characterized in that: include: The voice information recognition module is used to recognize the received voice information and obtain the recognition result; a target application determining module, configured to determine a target application from candidate applications matching the recognition result according to the recognition result; An instruction generation module, used to generate an instruction to start the target application according to the recognition result; The target application determination module is specifically used to determine the target application according to the usage habits of the current user and the characteristics of the candidate application; wherein the characteristics of the candidate application are limited-time free trials and shopping discount activities launched by the candidate application; The target application is determined according to the current user's usage habits and the characteristics of the candidate application, including: Weights are set for the current user's usage habits and the characteristics of the candidate applications; the cumulative result of the score values ​​calculated using the weights is used as the final score value of the candidate applications, and the candidate application with the highest final score value is used as the target application; Wherein, the device further comprises: A voice information recognition module, used to continuously recognize the received voice information, and when the recognition result is to adjust the display state of the target application, send an instruction corresponding to the recognition result to the target application; Wherein, the target application determination module includes: A current user identification submodule, used to identify the current user according to the voice information; A user usage habit determination submodule, used to determine the user usage habit according to the time when the current user browses the candidate application and / or according to the frequency of the current user opening the candidate application; The current user identification submodule is further configured to obtain a pre-established matching relationship pair between a user and a voiceprint; and based on the matching relationship pair, perform voiceprint recognition on the voice information to determine the current user.

5. The device according to claim 4, characterized in that The target application determination module is further specifically used for: when the recognition result includes the query intent and the user's usage habits, determining a target application matching the query intent from candidate applications according to the user's usage habits.

6. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 3.

7. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Method for intelligently recommending application program and mobile terminal

    CN103942270A

  • Method and apparatus for performing tracking processing on use of application

    CN105653434A

  • Voice control method, intelligent equipment and storage medium

    CN107492374A

  • Method, apparatus, terminal, and storage medium for displaying application program icons

    CN109343926A