A voice interaction method and an electronic device
By identifying and simulating user voices and images, electronic devices realize personalized voice interaction, solving the problem that personalized interaction cannot be provided in the prior art, and improving user experience.
Patent Information
- Application Number
- CN202010232268.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-03-27
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2040-03-27
AI Technical Summary
Existing smart devices cannot provide a personalized voice interaction experience in voice interaction, cannot recognize the user identity of voice information and simulate user voice for personalized conversations.
The electronic device recognizes the user's voice information, simulates the voice of the first user, and conducts voice conversations according to the dialogue between the first user and the second user, displays the image or expression of the first user, and provides a personalized voice interaction experience.
It improves the voice interaction performance of electronic devices, provides a real experience of face-to-face communication with the first user, and enhances user's interactive interest and satisfaction.
Smart Images

Figure CN113449068B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical fields of artificial intelligence and speech processing, and particularly to a speech interaction method and an electronic device. Background Art
[0002] Most existing intelligent devices can receive voice information (such as voice commands) issued by a user and perform operations corresponding to the voice information. Exemplarily, the above intelligent devices can be devices such as mobile phones, intelligent robots, smart watches, or smart home devices (such as smart TVs). For example, a mobile phone can receive a voice command "lower the volume" issued by a user and then automatically lower the volume of the mobile phone.
[0003] Some intelligent devices can also provide a voice interaction function. For example, an intelligent robot can receive voice information from a user and conduct a voice conversation with the user based on the voice information, thereby implementing the voice interaction function. However, when existing intelligent devices conduct a voice conversation with a user, they can only give some patterned voice responses according to a set voice mode. The interaction performance between the intelligent device and the user is poor, and a personalized voice interaction experience cannot be provided for the user. Summary of the Invention
[0004] The present application provides a speech interaction method and an electronic device, which improve the interaction performance between the electronic device and the user, thereby providing a personalized voice interaction experience for the user.
[0005] To achieve the above technical objectives, the present application adopts the following technical solutions:
[0006] In a first aspect, the present application provides a speech interaction method, which may include: an electronic device can receive a first voice message sent by a second user; and in response to the first voice message, the electronic device can recognize the first voice message. Among them, the first voice message is used to request a voice conversation with a first user. Based on the electronic device recognizing that the first voice message is the voice message of the second user, the electronic device can simulate the voice of the first user and conduct a voice conversation with the second user in the way that the first user conducts a voice conversation with the second user.
[0007] In the above solution, the electronic device can receive the first voice message and recognize that the first voice message is sent by the second user. Since the first voice message is a request for a voice conversation with the first user, the electronic device can recognize that the first voice message is that the second user wants to have a voice conversation with the first user. In this way, the electronic device can simulate the voice of the first user and have a voice conversation with the second user intelligently in the conversation mode of the first user and the second user having a voice conversation. Thus, the electronic device can simulate the first user and provide the second user with an exchange experience of having a real voice conversation with the first user. This voice interaction method improves the interaction performance of the electronic device and can provide a personalized voice interaction experience for the user.
[0008] In a possible implementation manner, the above conversation mode is used to indicate the tone and diction of the first user and the second user having a voice conversation.
[0009] The electronic device has a voice conversation with the second user in the conversation mode of the first user and the second user having a voice conversation. That is to say, the electronic device has a voice conversation with the first user in the tone and diction when the first user and the second user have a conversation. It provides the second user with a more real exchange experience of having a voice conversation with the first user and improves the interaction performance of the electronic device.
[0010] In another possible implementation manner, the electronic device may store the image information of the first user. Then, when the electronic device simulates the voice of the first user and has a voice conversation with the second user in the conversation mode of the first user and the second user, the electronic device can also display the image information of the first user.
[0011] If the electronic device can display images and the electronic device stores the image information of the first user. Then when the electronic device simulates the first user and the second user having a voice conversation, it displays the image information of the first user. Thus, when the electronic device simulates the first user and the second user having a voice conversation, the second user can not only hear the voice of the first user but also see the image of the first user. Through this solution, it can provide the user with an exchange experience similar to having a face-to-face voice conversation with the first user.
[0012] In another possible implementation manner, the electronic device may store the face model of the first user. Then, when the electronic device simulates the voice of the first user and has a voice conversation with the second user in the conversation mode of the first user and the second user, the electronic device can simulate the expressions of the first user and the second user having a voice conversation and display the face model of the first user. Among them, the expression of the first user in the face model displayed by the electronic device can change dynamically.
[0013] If a face model of a first user is stored in an electronic device, when the electronic device simulates a voice interaction between the first user and a second user, the electronic device displays the face model of the first user. Moreover, the displayed face model of the electronic device can change dynamically, making the user think that they are having a voice conversation with the first user. In this way, when the electronic device simulates a voice conversation between the first user and the second user, the second user can not only hear the voice of the first user but also see the facial expressions of the first user during the voice conversation with the first user. Through this solution, a more realistic experience of having a face-to-face voice conversation with the first user can be provided to the user.
[0014] In another possible implementation manner, before the electronic device receives the first voice message, the above method may further include: the electronic device may further obtain a second voice message, where the second voice message is the voice message during a voice conversation between the first user and the second user. The electronic device analyzes the obtained second voice message, so as to obtain the voice features during the voice conversation between the first user and the second user, and saves the voice features.
[0015] It can be understood that the voice features may include voiceprint features, tone features, and word-using features. The tone features are used to indicate the tone during the voice conversation between the first user and the second user; the word-using features are used to indicate the habitual vocabulary during the voice conversation between the first user and the second user. A more realistic communication experience of having a voice conversation with the first user is provided for the second user, further improving the interaction performance of the electronic device.
[0016] Among them, before the electronic device simulates a voice interaction between the first user and the second user, the electronic device obtains a second voice message, where the second voice message is the voice message of the voice conversation between the first user and the second user. The electronic device can analyze the voice features during the voice conversation between the first user and the second user according to the second voice message. In this way, when the electronic device simulates the conversation mode of the voice conversation between the first user and the second user, the electronic device can send out a voice conversation similar to that of the first user, thereby providing a personalized voice interaction experience for the user.
[0017] In another possible implementation manner, in the second voice message, the electronic device may also save the voice conversation record of the electronic device simulating the voice conversation between the first user and the second user.
[0018] In another possible implementation, when the electronic device identifies that the first voice message is the voice message of the second user, the electronic device can simulate the voice of the first user and have a voice conversation with the second user in the conversation mode of the voice conversation between the first user and the second user. Specifically, the electronic device identifies that the first voice is the voice message of the second user, the electronic device simulates the voice of the first user, and sends a voice response message to the first voice in the conversation mode of the voice conversation between the first user and the second user. If, after sending the voice response message to the first voice, the electronic device receives a third voice message, and the electronic device identifies that the third voice is the voice message of the second user. Then the electronic device identifies that the third voice is the voice message of the second user, and the electronic device can simulate the voice of the first user and send a voice response message to the third voice message in the conversation mode of the voice conversation between the first user and the second user.
[0019] It can be understood that when the electronic device responds to the first voice message by simulating the conversation mode between the first user and the second user, after receiving the third voice message, the electronic device needs to identify that the third voice is sent by the second user. Then, after the electronic device identifies that the third voice message is the voice message of the second user, it sends a response message to the third voice message. Suppose there are other users sending voice messages in the environment where the electronic device has a voice conversation with the second user. After receiving the third voice message, if the electronic device identifies that the third voice message is sent by the second user, it can have a better voice conversation with the second user. Thereby improving the voice interaction function and enhancing the user experience.
[0020] In another possible implementation, the electronic device can obtain the schedule information of the first user, and this schedule information is used to refer to the schedule arrangement of the first user. The voice response message sent by the electronic device to the third voice can be that the electronic device refers to this schedule information and sends a voice response message to the third voice message.
[0021] If the third voice is sent by the second user to inquire about the schedule arrangement of the first user, since the electronic device has obtained the schedule information of the first user, the electronic device can directly respond to the third voice message according to the schedule information. Thereby providing a personalized interaction experience for the first user.
[0022] In another possible implementation, the electronic device can save the voice conversation record of the electronic device simulating the voice of the first user and having a voice conversation with the second user, and the electronic device can also send this voice conversation record to the electronic device of the first user.
[0023] The electronic device sends the voice conversation record to the electronic device of the first user, so that the first user can understand the conversation content. The electronic device provides a more personalized voice interaction for the second user.
[0024] In another possible implementation, the electronic device stores the voice conversation record in which the electronic device simulates the first user and the second user. The electronic device can also extract keywords in the voice conversation from the above voice conversation record. The electronic device can send the keyword to the electronic device of the first user.
[0025] In another possible implementation, the electronic device simulates the voice of the first user and interacts with the second user in voice according to the conversation mode of the voice conversation between the first user and the second user. The electronic device can also obtain the image information and action information of the second user, and store the image information and action information of the second user.
[0026] Wherein, when the electronic device simulates the voice conversation between the first user and the second user and obtains the image information and action information of the second user, it can learn the expressions and actions of the second user during the voice conversation with the first user. So that the electronic device simulates the conversation mode of the second user and the first user in voice.
[0027] In a second aspect, the present application further provides an electronic device, which may include a memory, a voice module, and one or more processors. The memory, the voice module, and one or more processors are coupled.
[0028] The microphone can be used to receive the first voice information. The memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the processor is used to, in response to the first voice information, identify the first voice information, and the first voice information is used to request a voice conversation with the first user. Based on the first voice being recognized as the voice information of the second user, simulate the voice of the first user and have a voice conversation with the second user according to the conversation mode of the voice conversation between the first user and the second user.
[0029] In a possible implementation, the electronic device may further include a display screen, and the display screen is coupled to the processor. The display screen is used to display the image information of the first user.
[0030] In another possible implementation, a face model of the first user is stored in the electronic device. The display screen in the electronic device is further used to simulate the expression of the first user and the second user during the voice conversation and display the face model; wherein, the expression of the first user in the face model changes dynamically.
[0031] In another possible implementation, the microphone is further used to obtain the second voice information, and the second voice information is the voice information during the voice conversation between the first user and the second user.
[0032] The processor is further used to analyze the second voice information, obtain the voice characteristics during the voice conversation between the first user and the second user, and store the voice characteristics.
[0033] Among them, the voice features include voiceprint features, tone features, and word - using features. The tone features are used to indicate the tone when the first user has a voice conversation with the second user, and the word - using features are used to indicate the habitual vocabulary when the first user has a voice conversation with the second user.
[0034] In another possible implementation, the processor is further configured to save the voice conversation record in which the electronic device simulates the voice conversation between the first user and the second user in the second voice information.
[0035] In another possible implementation, the microphone is further configured to receive the third voice information. The processor is further configured to recognize the third voice information in response to the third voice information. Based on the fact that the third voice information is recognized as the voice information of the second user, the electronic device simulates the voice of the first user, and in the conversation mode of the voice conversation between the first user and the second user, the speaker is further configured to emit the voice response information of the third voice information.
[0036] In another possible implementation, the processor is further configured to obtain the schedule information of the first user, and the schedule information is used to indicate the schedule arrangement of the first user. Among them, emitting the voice response information of the third voice information includes: the electronic device refers to the schedule information and emits the voice response information of the third voice information.
[0037] In another possible implementation, the processor is further configured to save the voice conversation record in which the electronic device simulates the voice of the first user and the second user; and send the voice conversation record to the electronic device of the first user.
[0038] In another possible implementation, the processor is further configured to save the voice conversation record in which the electronic device simulates the voice conversation between the first user and the second user. Extract the keywords of the voice conversation in which the electronic device simulates the voice conversation between the first user and the second user from the voice conversation record; and send the keywords to the electronic device of the first user.
[0039] In another possible implementation, the electronic device further includes a camera, and the camera is coupled to the processor; the camera is used to obtain the image information and action information of the second user, and the processor is further configured to save the image information and action information of the second user.
[0040] In a third aspect, the present application further provides a server, and the server may include: a memory and one or more processors. The memory is coupled to the one or more processors. Among them, the memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the server is made to execute the methods in the first aspect and any of its possible implementations above.
[0041] Fourthly, the present application also provides a computer-readable storage medium, including computer instructions, which, when running on an electronic device, enable the electronic device to execute the methods in the first aspect and any possible implementation manners thereof as described above.
[0042] Fifthly, the present application also provides a computer program product, which, when running on a computer, enables the computer to execute the methods in the first aspect and any possible implementation manners thereof as described above.
[0043] It can be understood that for the beneficial effects that can be achieved by the electronic device in the second aspect, the server in the third aspect, the computer-readable storage medium in the fourth aspect, and the computer program product provided by the present application as described above, reference can be made to the beneficial effects in the first aspect and any possible design manners thereof, which will not be elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1A FIG. [X] is a system architecture diagram provided by an embodiment of the present application;
[0045] Figure 1B FIG. [X] is another system architecture diagram provided by an embodiment of the present application;
[0046] Figure 2A FIG. [X] is a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0047] Figure 2B FIG. [X] is a schematic software structure diagram of an electronic device provided by an embodiment of the present application;
[0048] Figure 3A FIG. [X] is a flowchart of a voice interaction method provided by an embodiment of the present application;
[0049] Figure 3B FIG. [X] is a schematic diagram of a display interface of a smart speaker provided by an embodiment of the present application;
[0050] Figure 4 FIG. [X] is a schematic structural diagram of a smart speaker provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this embodiment, unless otherwise specified, "a plurality" means two or more.
[0052] Note: The "FIG. [X]" in the translation is a placeholder for the actual figure number which should be filled in according to the specific content of the original patent text.Generally, an electronic device with a voice interaction function can issue a corresponding voice response according to the recognized voice information. However, the electronic device cannot identify which user issued the voice information. That is to say, when the electronic device is in the voice interaction function, once the voice information is recognized, a corresponding voice response will be issued. In addition, the corresponding voice response issued by the electronic device is also fixed. The voice interaction function of the electronic device enables the electronic device to have a voice conversation with the user. If the electronic device can identify the user who issued the voice information, it can issue a corresponding voice response according to the user who issued the voice information and specifically. Then it can provide a personalized voice interaction experience for the user, thereby increasing the user's interest in using the electronic device for voice interaction.
[0053] In addition, generally, an electronic device cannot "act as" another user. Herein, "acting as" means that when the electronic device has a voice interaction with user 2, it simulates the voice of user 1 and uses the conversation mode between user 1 and user 2 to have a voice interaction with user 2. In some actual situations, such as when parents need to go out to work and cannot communicate with their children at any time. If the electronic device can "act as" a father or mother to have a voice conversation with the child to meet the child's idea of wanting to communicate with the parent. It enables the electronic device to provide a more personalized and user-friendly voice interaction for the child.
[0054] The embodiment of the present application provides a voice interaction method applied to an electronic device. It enables the electronic device to "act as" user 1 to have a voice interaction with user 2. It improves the voice interaction performance of the electronic device and can also provide a personalized interaction experience for user 2.
[0055] Exemplarily, the electronic device in the embodiment of the present application can be a mobile phone, a television, a smart speaker, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, a vehicle-mounted device, an ultra-mobile personal computer (UMPC), a netbook, as well as a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) / virtual reality (VR) device, etc. The embodiment of the present application does not impose any special restrictions on the specific form of the electronic device.
[0056] Hereinafter, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings.
[0057] Please refer to Figure 1A , which is a system architecture diagram provided by the embodiment of the present application. Assume that the electronic device "acts as" user 1 to have a voice interaction with user 2. As Figure 1AThe electronic device can collect the voice information sent by User 2. The electronic device can interact with a remote server via the Internet, send the voice information of User 2 to the server, generate a response message corresponding to the voice information by the server, and send the response message corresponding to the generated voice information to the electronic device. The electronic device is used to play the response message corresponding to the voice information to achieve the purpose of "acting as" User 1 to have a voice interaction with User 2. That is to say, the electronic device can collect and recognize the voice information sent by User 2, and can play the response message corresponding to the voice information. In this implementation, the server connected to the electronic device recognizes the voice information of User 2 and generates a response message corresponding to the voice information. The electronic device plays the response message corresponding to the voice information, which can reduce the computing requirements of the electronic device and the production cost of the electronic device.
[0058] Please refer to Figure 1B , which is another system architecture diagram provided by the embodiments of the present application. Assume that the electronic device "acts as" User 1 and has a voice interaction with User 2. As Figure 1B The electronic device can collect the voice information sent by User 2. The electronic device recognizes that the voice information is the voice information of User 2 according to the voice information, and this voice information is a request to have a voice conversation with User 1. The electronic device generates a corresponding response message according to the voice information and plays the response message. In this implementation, the electronic device can achieve voice interaction and reduce the dependence of the electronic device on the Internet.
[0059] Please refer to Figure 2A , which is a schematic structural diagram of an electronic device 200 provided by the embodiments of the present application. As Figure 2A shown, the electronic device 200 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a sensor module 280, a camera 293, a display screen 294, etc.
[0060] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may include more or fewer components than shown, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0061] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0062] Among them, the controller may be the nerve center and command center of the electronic device 200. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0063] A memory may also be provided in the processor 210 for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory may save the instructions or data that the processor 210 has just used or recycled. If the processor 210 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.
[0064] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0065] It can be understood that the interface connection relationships among the modules illustrated in the embodiments of the present invention are only illustrative descriptions and do not constitute a structural limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may also adopt different interface connection manners in the above embodiments, or a combination of multiple interface connection manners.
[0066] The external memory interface 220 can be used to connect an external memory card, such as a Micro SD card, to implement the storage capacity expansion of the electronic device 200. The external memory card communicates with the processor 210 through the external memory interface 220 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0067] The internal memory 221 can be used to store computer-executable program codes, and the executable program codes include instructions. The processor 210 executes various functional applications and data processing of the electronic device 200 by running the instructions stored in the internal memory 221. The internal memory 221 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the electronic device 200 (such as audio data, a phone book, etc.). In addition, the internal memory 221 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0068] The charging management module 240 is used to receive a charging input from a charger. Among them, the charger can be a wireless charger or a wired charger. The power management module 241 is used to connect the battery 242, the charging management module 240, and the processor 210. The power management module 241 receives the inputs from the battery 242 and / or the charging management module 240 to supply power to the processor 210, the internal memory 221, the external memory, the display screen 294, the wireless communication module 260, the audio module 270, etc.
[0069] The wireless communication function of the electronic device 200 can be implemented through antenna 1, antenna 2, the mobile communication module 250, the wireless communication module 260, the modulation and demodulation processor, and the baseband processor, etc.
[0070] The mobile communication module 250 may provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the electronic device 200. The mobile communication module 250 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 250 may receive electromagnetic waves through the antenna 1, filter and amplify the received electromagnetic waves, and then transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 250 may also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through the antenna 1 for radiation.
[0071] The wireless communication module 260 may provide solutions for wireless communications such as wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. applied to the electronic device 200. Among them, the wireless communication module 260 may be one or more devices integrating at least one communication processing module. The wireless communication module 260 receives electromagnetic waves through the antenna 2, performs frequency modulation and filtering on the electromagnetic wave signals, and sends the processed signals to the processor 210. The wireless communication module 260 may also receive the signal to be transmitted from the processor 210, perform frequency modulation and amplification on it, and convert it into electromagnetic waves through the antenna 2 for radiation.
[0072] The display screen 294 is used to display images, videos, etc. The display screen 294 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 200 may include one or N display screens 294, where N is a positive integer greater than 1.
[0073] The camera 293 is used to capture still images or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB, YUV, etc. In some embodiments, the electronic device 200 may include one or N cameras 293, where N is a positive integer greater than 1.
[0074] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 200 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0075] The video codec is used to compress or decompress digital videos. The electronic device 200 can support one or more video codecs. In this way, the electronic device 200 can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0076] The electronic device 200 can implement audio functions through the audio module 270, the speaker 270A, the microphone 270B, and the application processor, etc. Such as music playback, recording, etc.
[0077] The audio module 270 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 270 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 270 can be disposed in the processor 210, or some functional modules of the audio module 270 can be disposed in the processor 210.
[0078] The speaker 270A, also known as a "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 200 can listen to music or a hands-free call through the speaker 270A. In some embodiments, the speaker 270A can play a response message to the voice message.
[0079] The microphone 270B, also known as a "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak close to the microphone 270B with their mouth to input the sound signal into the microphone 270B. For example, the microphone 270B can collect the voice message sent by the user. The electronic device 200 can be provided with at least one microphone 270B. In some embodiments, the electronic device 200 can be provided with two microphones 270B, which can not only collect sound signals but also implement a noise reduction function. In other embodiments, the electronic device 200 can also be provided with three, four or more microphones 270B to collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording.
[0080] The software system of the electronic device 200 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of the present invention, the Android system with a layered architecture is taken as an example to exemplarily illustrate the software structure of the electronic device 200.
[0081] Figure 2B It is a software structure block diagram of the electronic device 200 in the embodiments of the present invention.
[0082] The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, namely the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0083] The application layer can include a series of application packages.
[0084] As Figure 2B shown, the application packages can include applications such as a camera, a gallery, a calendar, WLAN, voice conversation, Bluetooth, music, and video.
[0085] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.
[0086] As Figure 2B shown, the application framework layer may include a window manager, a content provider, a view system, a telephone manager, a resource manager, a notification manager, etc.
[0087] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.
[0088] The content provider is used to store and obtain data, and make this data accessible to applications. The data may include videos, images, audio, dialed and answered calls, browsing history and bookmarks, phone books, etc. For example, the data may be the voiceprint feature of user 2, the relationship between user 2 and user 1, etc.
[0089] The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views.
[0090] The telephone manager is used to provide the communication function of the electronic device 200. For example, the management of call status (including connection, hanging up, etc.).
[0091] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, and so on.
[0092] The notification manager enables applications to display notification information in the status bar. It can be used to convey notification-type messages, which can automatically disappear after a short stay without user interaction. For example, the notification manager is used to notify download completion, message reminder, Bluetooth pairing success reminder, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as the notification of a background-running application, and can also be a notification that appears on the screen in the form of a dialogue window. For example, prompt text information in the status bar, emit a prompt sound, the electronic device vibrates, the indicator light flashes, etc.
[0093] Android Runtime includes a core library and a virtual machine. Android runtime is responsible for the scheduling and management of the Android system.
[0094] The core library contains two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core library of Android.
[0095] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files in the application layer and the application framework layer as binary files. The virtual machine is used to manage the object life cycle, stack management, thread management, security and exception management, and garbage collection and other functions.
[0096] The system library can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing library (e.g., OpenGL ES), 2D graphics engine (e.g., SGL), etc.
[0097] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications.
[0098] The media library supports the playback and recording of various common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0099] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.
[0100] The 2D graphics engine is a drawing engine for 2D drawing.
[0101] The kernel layer is the layer between hardware and software. The kernel layer at least includes a display driver, a camera driver, an audio driver, and a sensor driver.
[0102] The methods in the following embodiments can all be implemented in an electronic device with the above hardware structure.
[0103] Please refer to Figure 3A , which is a flowchart of the voice interaction method provided by the embodiments of the present application. Among them, in the embodiments of the present application, taking the electronic device as a smart speaker, and the smart speaker "acting as" user 1 to have a voice conversation with user 2 as an example, the voice interaction method is specifically described. As Figure 3A shown, the method includes steps 301-step 304.
[0104] Step 301: User 2 sends a first voice message to the smart speaker.
[0105] Among them, the first voice message is used to request the smart speaker to "act as" user 1 to have a voice conversation with itself (user 2).
[0106] In a possible scenario, when the parents go out to work and the child at home needs the company of the parents and wants to have a voice conversation with them, the child can send a voice message to the smart speaker at home and request the smart speaker to "act as" the father or mother to accompany himself / herself. For example, the first voice message can be "I want to talk to Dad" or "Smart speaker, I want to talk to Dad".
[0107] It can be understood that the smart speaker needs to be woken up to work, and the wake-up voice of the smart speaker can be fixed. In some embodiments, before step 301, user 2 can first send a wake-up word to the smart speaker to make the smart speaker in a wake-up state.
[0108] Implementation method 1, the wake-up word can be "Smart speaker", "Intelligent speaker" or "Voice speaker", etc. The wake-up word can be pre-configured in the smart speaker or set by the user in the smart speaker. In this embodiment, the above first voice message may not include the wake-up word. For example, the first voice message can be "I want to talk to Dad".
[0109] Implementation method 2, the first voice message can include the above wake-up word and also include the voice command sent by user 2 to the smart speaker. For example, the first voice message can be "Smart speaker, I want to talk to Dad".
[0110] Step 302: The smart speaker receives the first voice message sent by user 2.
[0111] Among them, when the smart speaker is not woken up, it is in a sleep state. When user 2 wants to use the smart speaker, the voice assistant can be woken up by voice. The voice wake-up process can include: the smart speaker monitors voice data through a low-power digital signal processor (DSP). When the DSP monitors that the similarity between the voice data and the above wake-up word meets certain conditions, the DSP hands over the monitored voice data to the application processor (AP). The AP performs text verification on the above voice data to determine whether the voice data can wake up the smart speaker.
[0112] It can be understood that when the smart speaker is in a sleep state, it can listen to the voice messages sent by the user at any time. If the voice message is not the wake-up voice to wake itself (the smart speaker) up, the smart speaker will not respond to the voice message nor record the voice message.
[0113] In the above implementation method 1, when the smart speaker is in a wake-up state, the first voice message may not include the wake-up word of the smart speaker, and the smart speaker receives the first voice message and responds to the first voice message.
[0114] In the above implementation 2, when the smart speaker is in the sleep state, the first voice message includes the wake-up word of the smart speaker. The smart speaker is woken up upon receiving the first voice message and responds to the first voice message.
[0115] Step 303: The smart speaker responds to the first voice message, recognizes the first voice message, and determines that the first voice message is used to request a voice conversation with User 1.
[0116] Among them, the smart speaker can perform text recognition on the first voice message. According to the result of the text recognition, it is determined that the first voice message is used to request a voice conversation with User 1. That is to say, the first voice message includes the name or title of the role that the smart speaker is to "play", enabling the smart device to recognize the role to be "played" based on the name or title. When the smart speaker recognizes the name in the first voice message, it can determine the role to be "played". For example, if the first voice message is "I want to talk to Li Ming", the smart speaker can determine that the role to be "played" is Li Ming. When the smart speaker recognizes the title in the first voice message, it can determine that the sender of the first voice message is User 2, and the smart speaker determines the role to be "played" based on the relationship between User 2 and the title in the first voice message.
[0117] Taking the usage scenario of the smart speaker being a home environment as an example, family member relationships can be pre-stored in the smart speaker. Then, after receiving the first voice message, the smart speaker can determine the role to be "played" based on the family member relationships.
[0118] Example 1: The first voice message sent by User 2 received by the smart speaker is "I want to talk to Dad". After the smart speaker recognizes the first voice message and recognizes the title "Dad", it can determine that the relationship between User 2 and the role to be played is father-son. The smart speaker can recognize that the first voice message is sent by the child "Li Xiaoming", and based on the father-son relationship between Li Xiaoming and Li Ming in the pre-stored family member relationships, it determines that the role to be "played" is "Li Ming" (Dad).
[0119] Example 2: The first voice message sent by User 2 received by the smart speaker is "I want to talk to Li Ming". The smart speaker recognizes the name "Li Ming" included in the first voice message. The smart speaker can also recognize that the first voice message is sent by "Li Xiaoming". The smart speaker determines the father-son relationship between Li Ming and Li Xiaoming based on the pre-stored family member relationships, and the smart speaker determines that the role to be "played" is "Dad" (Li Ming) of Li Xiaoming.
[0120] Taking the application of a smart speaker in a home scenario as an example, when the smart speaker is initially set up, the relationships between family members in the home need to be entered into the smart speaker. In one possible implementation, it is sufficient for the smart speaker to obtain the relationship between one family member and another known family member, and the smart speaker can infer the relationships between this family member and other family members. For example, among the family members, there are: grandfather, grandmother, father, mother, and child. If grandfather, grandmother, and mother have already been entered, when father is entered, it can simply be stated that father and mother are husband and wife. The smart speaker can infer that father and grandfather are father-son relationship, and father and grandmother are mother-son relationship based on the relationship between mother and father. Among them, the above reasoning can be achieved through technologies such as knowledge graphs.
[0121] In some embodiments, the pre-stored family member information may include: name, age, gender, communication method, voice information, image information, preferences, and personality, etc. At the same time, record the relationship information between this family member and the existing family members. Among them, when recording the information of each family member, the title of this member can also be recorded, such as "father", "grandfather", "Mr. Li", and "Mr. Li", etc. Among them, "father" and "Mr. Li" both refer to Li Ming; "grandfather" and "Mr. Li" both refer to Li Ming's father. For example, Li Xiaoming is User 2, and the first voice message is "I want to talk to Mr. Li" or the first voice message is "I want to talk to my father", the smart speaker can determine that it is "playing" the role of father Li Ming.
[0122] Step 304: Based on the recognition that the first voice message is the voice message of User 2, the smart speaker can simulate the voice of User 1 and send a response message to the first voice message in the dialogue manner of the voice conversation between User 1 and User 2.
[0123] It can be understood that the smart speaker recognizes that the first voice is the voice message sent by User 2, and the smart speaker can generate a response message to the first voice message in the dialogue manner of the conversation between User 1 and User 2. That is to say, the smart speaker can "play" the role of User 1 and infer the possible response message after User 1 hears the first voice message sent by User 2.
[0124] Among them, the voice of User 1 and the conversation manner between User 1 and User 2 are pre-stored in the smart speaker. The smart speaker sends a voice message in the conversation manner between User 1 and User 2 and simulates the voice of User 1, making User 2 think that they are really having a voice conversation with User 1. Thus, a personalized voice interaction experience is provided for User 2.
[0125] On the one hand, the smart speaker can analyze the voice of User 1, including analyzing the voiceprint characteristics of User 1. Among them, the voiceprint characteristics of each person are unique, so the speaker can be identified according to the voiceprint characteristics in the voice. The smart speaker analyzes the voice of User 1 and saves the voiceprint characteristics of User 1 so that the smart speaker can simulate the voice of User 1 when it needs to "act as" User 1.
[0126] Specifically, when the smart speaker receives the voice information of User 1, it can analyze the voiceprint characteristics of User 1. The smart speaker saves the voiceprint characteristics of User 1, so that when the smart speaker determines to "act as" User 1, it can simulate the voice of User 1 according to the stored voiceprint characteristics of User 1. It can be understood that when the smart speaker has a voice conversation with User 1, it can update the voiceprint characteristics of User 1 according to the voice changes of User 1. Or, as time goes by, the smart speaker can update the voiceprint characteristics of User 1 when having a voice conversation with User 1 after a preset time interval.
[0127] On the other hand, the conversation style between User 1 and User 2 can reflect the language expression characteristics of User 1. The conversation style between User 1 and User 2 includes the tone and word usage of the voice conversation between User 1 and User 2. Among them, when a person has a voice conversation with different people, the tone may be different because the people in the conversation are different. For example, a person's tone is gentle when communicating with their lover and respectful when communicating with the elders at home. Therefore, the smart speaker can infer the tone of User 1 to be played according to the relationship between the role to be "played" and User 2. The word usage in the voice conversation between User 1 and User 2 can also reflect the language expression characteristics of User 1, so that the response information of the first voice message generated by the smart speaker according to the word usage in the voice conversation between User 1 and User 2 is closer to the language expression of User 1. The smart speaker can simulate the voice of User 1 and send the response information of the first voice message in the conversation style between User 1 and User 2, so that User 2 thinks they are having a voice conversation with User 1.
[0128] Specifically, the conversation style between User 1 and User 2 in a voice conversation can include: the tone, word usage habits (such as catchphrases), and voice expression habits when User 1 and User 2 are talking. The tone when User 1 and User 2 are talking includes being serious, gentle, strict, slow, and aggressive, etc. Word usage habit is the language expression characteristic when a person is speaking. For example, when speaking, they are used to using words such as "then", "that is", "yes", "do you understand?" etc. Voice expression habit can reflect a person's language expression characteristic. For example, some people like to use inverted sentences when speaking, such as "Have you eaten your meal?" and "I'll leave first then".
[0129] Exemplarily, the smart speaker can pre-store the voice conversations between User 1 and User 2, and the smart speaker can learn this voice conversation. Understand information such as the tone, word usage habits, and language expression characteristics when User 1 and User 2 have a voice conversation, and save the information learned to the conversation information of this person. Among them, the conversation information can save the conversation information of this person's conversation with other tasks. If the smart speaker receives a request from User 2 to ask the smart speaker to "act as" User 1 in a conversation, the smart speaker can send a voice conversation according to the stored conversation method between User 1 and User 2.
[0130] It can be understood that the more voice conversations there are between User 1 and User 2 obtained by the smart speaker, the more accurate the conversation method information learned and summarized by the smart speaker about User 1 and User 2. When the smart speaker "acts as" User 1, the response information of the first voice message sent by the smart speaker is closer to the voice reply that User 1 would give. Similarly, the smart speaker can also learn the conversation method when User 2 and User 1 have a voice conversation from the voice conversation between User 1 and User 2, and store the conversation method of User 2 as the information of User 2 in the conversation information of User 2.
[0131] Also exemplarily, if there is no pre-stored voice conversation between User 1 and User 2 in the smart speaker, the smart speaker can infer the possible tone that User 1 may use based on the relationship between User 1 and User 2. For example, if the smart speaker identifies that the relationship between User 1 and User 2 is father and son, and the smart speaker is to "act as" the father, the smart speaker can default the tone of User 1 to be severe.
[0132] Among them, the smart speaker infers the tone when User 1 sends a voice response based on the relationship between User 1 and User 2. The tone of User 1 inferred by the smart speaker can be at least one. For example, if the smart speaker determines that the relationship between User 1 and User 2 is grandparent and grandchild, the tone of User 1 inferred by the smart speaker is doting, slow, and happy.
[0133] In some embodiments, if the smart speaker has a display screen, when the smart speaker "acts as" User 1 and has a voice conversation with User 2, a photo of User 1 can be displayed on the display screen. As Figure 3B shown, Figure 3B a photo of User 1 is displayed on the display screen of the smart speaker. Or, if the smart speaker stores a face model of User 1, when the smart speaker "acts as" User 1 and has a voice conversation with User 2, the dynamic change of the expression of User 1 can be displayed on the display screen.
[0134] In addition, when the smart speaker "acts as" User 1 and User 2 during a voice interaction, the smart speaker can also turn on the camera to obtain the image information of User 2. The smart speaker recognizes the obtained image information of User 2, that is, obtains information such as the appearance and actions of User 2. In this way, the smart speaker can establish a user model of User 2 based on the voice information and image information of User 2. Establishing a user model of User 2 by the smart speaker can facilitate a more vivid and lifelike "acting as" User 2 in the future.
[0135] Exemplarily, when the smart speaker "acts as" User 1 and User 2 during a voice interaction, the smart speaker can also turn on the camera to obtain User 2's expressions, actions, etc. So that the smart speaker can establish a user model of User 2 based on the voice information and image information of User 2, and determine the action information and expression information when User 2 is having a conversation with User 1.
[0136] Suppose during the voice interaction between User 2 and the smart speaker, the smart speaker receives the voice information of User 2 asking about User 1's schedule. The smart speaker can obtain User 1's schedule information, which is used to indicate User 1's schedule. In this way, the smart speaker can respond to the voice information asking about the schedule according to User 1's schedule information. For example, the smart speaker "acts as" Li Ming and has a voice conversation with his son Li Xiaoming. Li Xiaoming sends a voice message asking about his father's schedule. Suppose the voice message is "Will you come to my graduation ceremony on Friday?", and the smart speaker determines through querying User 1 (i.e., Dad)'s schedule information that Dad has to go on a business trip on Friday. The smart speaker can reply, "Son, Dad just received a notice from the company and has to go on a business trip to Beijing for an important meeting. Maybe I can't attend your graduation ceremony on Friday."
[0137] It is worth mentioning that the smart speaker can also save the conversation information of each "acting as" a role. And during the next "acting as" a role, if relevant schedule information is involved, the smart speaker can feedback the updated schedule to User 2. Another example is that after the above-mentioned voice conversation where the smart speaker "acts as" Dad and Li Xiaoming ends. The smart speaker "acts as" Xiaoming and has a voice conversation with Xiaoming's mother (User 2). The voice message sent by Xiaoming's mother is "Son, your father and I will accompany you to attend the graduation ceremony on Friday.", and the smart speaker can reply according to the previous voice conversation where it "acted as" Dad, "My dad said he has to go on a business trip to Beijing for a meeting and can't attend my graduation ceremony."
[0138] It should be noted that the above steps 301 - 304 are a conversation between User 2 and the smart speaker. After step 304, the smart speaker can continue the voice conversation with User 2. For example, if User 2 sends another voice message to the smart speaker, after receiving this voice message, based on the fact that this voice message is sent by User 2, the smart speaker continues to simulate the voice of User 1 and has a voice conversation with User 2 in the way of the conversation between User 1 and User 2. That is to say, the smart speaker will continue to simulate the voice of User 1 and send voice messages in the way of the conversation between User 1 and User 2 only when it continues to receive the voice message of User 2. If this voice message is not sent by User 2, the smart speaker may not simulate the voice of User 1.
[0139] In some embodiments, after the smart speaker responds to the voice message of User 2 each time, it can wait for a preset time. Herein, the preset waiting time is the reaction time of User 2, so that the smart speaker can maintain the voice conversation with User 2. If the smart speaker does not receive the voice message of User 2 within the preset time, the smart speaker can end this voice conversation.
[0140] Exemplarily, if the smart speaker determines that the voice conversation with User 2 has ended, it can send the content of this voice conversation to the electronic device of User 1 for User 1 to understand the details of the conversation in which the smart speaker "acts as" him (User 1) and User 2. Or, if the smart speaker determines that the voice conversation with User 2 has ended, the smart speaker can summarize the abstract of this voice conversation and send the abstract of the voice conversation to the electronic device of User 1. So that User 1 can simply understand the situation of the conversation in which the smart speaker "acts as" him (User 1) and User 2.
[0141] In one embodiment, after receiving the end of the voice conversation, the smart speaker can send the abstract of the voice conversation to the electronic device of User 1 after a preset time. For example, if User 2 is Xiaoming's mother and the User 1 played by the smart speaker is Xiaoming. If Xiaoming's mother is about to go out to buy groceries and says to the smart speaker, "Mom is going out to buy groceries. You can't watch TV until you finish your homework first." Later, if User 2 is Xiaoming's grandmother and the User 1 played by the smart speaker is Xiaoming. If Xiaoming's grandmother is about to go out for a walk and says to the smart speaker, "Grandma is going for a walk. I left a cake in the fridge for you. Remember to get it and eat it." After the preset time, the smart speaker makes a text summary and aggregation of the conversations between different roles and Xiaoming, and then generates a conversation summary, which can be "Mom reminded to finish homework in time, and Grandma left a cake in the fridge for you." The smart speaker can send this conversation summary to Xiaoming's mobile phone by means of communication (such as text message).
[0142] In the above manner, the smart speaker can identify that the first voice message is sent by User 2 and can recognize that the first voice message instructs the smart speaker to "act as" User 1. In response to the first voice message, the smart speaker can simulate the voice of User 1 and send a response message to the first voice message in the conversation mode between User 1 and User 2. In this way, the purpose of the smart speaker "acting as" User 1 to have a voice conversation with User 2 is achieved. This voice interaction method improves the interaction performance of the smart speaker and can provide a personalized voice interaction experience for User 2.
[0143] It can be understood that, in order to implement the above functions, the above smart speaker includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed in this article, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the manner of hardware or computer software driving hardware depends on the specific application and design constraint conditions of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0144] The embodiments of the present application can divide the functional modules of the smart speaker according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation.
[0145] As Figure 4 shown, it is a possible structural schematic diagram of the smart speaker involved in the above embodiments. The smart speaker may include: a voice recognition module 401, a relationship reasoning module 402, a role-playing module 403, a knowledge pre-storage module 404, a role information knowledge base 405, and an audio module 406. Optionally, the smart speaker may further include a camera module, a communication module, a sensor module, etc.
[0146] Among them, the speech recognition module 401 is used to recognize the first speech information received by the smart speaker. The relationship reasoning module 402 is used to infer the relationship between the newly entered person and the existing family members based on the existing family relationship. The role-playing module 403 enables the smart speaker to simulate the voice of user 1 and send a response message corresponding to the first speech information. The knowledge pre-storage module 404 is used to store the information of each user, so that the role-playing module 403 can obtain the user information and generate a response message corresponding to the speech information according to the user information. The role information knowledge base 405 is used to store the conversation information of the user and can generate a response message for the speech information according to the first speech information.
[0147] In some embodiments, the smart speaker may further include a summary module. The summary module is used to extract keywords in the conversation information and use the keywords as the summary of the conversation information; or, it is used to summarize the information of the conversation. Among them, the summary module can send the summary of the conversation information to the smart device of user 1 "played" by the smart speaker. Or, the communication module in the smart speaker sends the keywords extracted from the conversation information by the summary module to the smart device of user 1 "played" by the smart speaker.
[0148] Of course, the unit modules in the above smart speaker include but are not limited to the above speech recognition module 401, relationship reasoning module 402, role-playing module 403, knowledge pre-storage module 404, role information knowledge base 405, and audio module 406, etc. For example, the smart speaker may further include a storage module. The storage module is used to save the program code and data of the electronic device.
[0149] The embodiment of the present application also provides a computer-readable storage medium, in which computer program code is stored. When the above processor executes the computer program code, the smart speaker can execute Figure 3A the relevant method steps to implement the method in the above embodiment.
[0150] The embodiment of the present application also provides a computer program product. When the computer program product runs on a computer, it enables the computer to execute Figure 3A the relevant method steps to implement the method in the above embodiment.
[0151] Among them, the smart speaker, computer storage medium, or computer program product provided by the embodiment of the present application are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.
[0152] From the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0153] In several embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0154] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0155] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROMs, magnetic disks, or optical discs that can store program codes.
[0156] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A voice interaction method, characterized in that, The method includes: The electronic device receives first voice information. In response to the first voice information, the electronic device recognizes the first voice information, which is used to request a voice conversation with a first user. Based on the first voice information being recognized as the voice information of a second user, the electronic device simulates the voice of the first user, and in accordance with the conversation manner of the first user and the second user having a voice conversation, sends a response message to the first voice information and has a voice conversation with the second user; the conversation manner is used to indicate the tone and diction of the first user and the second user having a voice conversation; the tone is determined according to the relationship between the first user and the second user; the response message includes information determined based on the relationship between the first user and the second user in the current voice conversation and the voice conversation record simulated by the electronic device in the previous voice conversation; the voice conversation record is the response message sent by the electronic device when simulating the first user in the previous voice conversation; the relationship between the first user and the second user in the previous voice conversation is different from the relationship in the current voice conversation.
2. The method according to claim 1, characterized in that, Image information of the first user is stored in the electronic device; the method further includes: The electronic device displays the image information of the first user.
3. The method according to claim 1, characterized in that, A face model of the first user is stored in the electronic device; the method further includes: The electronic device simulates the expressions of the first user and the second user having a voice conversation and displays the face model; wherein, the expressions of the first user in the face model change dynamically.
4. The method according to any one of claims 1 to 3, characterized in that, Before the electronic device receives the first voice information, the method further includes: The electronic device obtains second voice information, which is the voice information when the first user and the second user have a voice conversation. The electronic device analyzes the second voice information to obtain the voice characteristics when the first user and the second user have a voice conversation and stores the voice characteristics. Wherein, the voice characteristics include voiceprint characteristics, tone characteristics, and diction characteristics. The tone characteristics are used to indicate the tone when the first user and the second user have a voice conversation, and the diction characteristics are used to indicate the habitual vocabulary when the first user and the second user have a voice conversation.
5. The method according to claim 4, wherein The method further includes: The electronic device stores the voice conversation record of the electronic device simulating the first user and the second user in the second voice information.
6. The method according to claim 1, characterized in that, Based on the first voice information being recognized as the voice information of the second user, the electronic device simulates the voice of the first user and has a voice conversation with the second user in accordance with the conversation manner of the first user and the second user having a voice conversation, including: Based on the first voice information being recognized as the voice information of the second user, the electronic device simulates the voice of the first user and in accordance with the conversation manner of the first user and the second user having a voice conversation, sends a voice response message to the first voice information. The electronic device receives third voice information; In response to the third voice information, the electronic device recognizes the third voice information; Based on the recognition that the third voice information is the voice information of the second user, the electronic device simulates the voice of the first user and sends out a voice response message to the third voice information in the dialogue manner of the first user having a voice conversation with the second user.
7. The method according to claim 6, wherein The method further includes: The electronic device obtains the schedule information of the first user, and the schedule information is used to indicate the schedule arrangement of the first user; Wherein, sending out the voice response message to the third voice information includes: The electronic device refers to the schedule information and sends out the voice response message to the third voice information.
8. The method according to claim 1, wherein The method further includes: The electronic device saves the voice conversation record of the electronic device simulating the voice of the first user and the second user; The electronic device sends the voice conversation record to the electronic device of the first user.
9. The method according to claim 1, wherein The method further includes: The electronic device saves the voice conversation record of the electronic device simulating the first user and the second user; The electronic device extracts keywords of the voice conversation of the electronic device simulating the first user and the second user from the voice conversation record; The electronic device sends the keywords to the electronic device of the first user.
10. The method according to claim 1, characterized in that, The method further includes: The electronic device obtains the image information and action information of the second user, and saves the image information and action information of the second user.
11. An electronic device, characterized in that, The electronic device includes: a memory, a microphone, a speaker, and a processor; the memory, the microphone, and the speaker are coupled to the processor; the microphone is used to receive first voice information; the memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the above computer instructions, The processor is configured to, in response to the first voice information, recognize the first voice information, where the first voice information is used to request a voice conversation with the first user; Based on the recognition that the first voice information is the voice information of the second user, simulate the voice of the first user and send out a response message to the first voice information in the dialogue manner of the first user having a voice conversation with the second user, and have a voice conversation with the second user; the dialogue manner is used to indicate the tone and diction of the first user having a voice conversation with the second user; the tone is determined according to the relationship between the first user and the second user; the response message includes information determined according to the relationship between the first user and the second user in the current voice conversation and the voice conversation record simulated by the electronic device in the previous voice conversation; the voice conversation record is the response message sent by the electronic device when simulating the first user in the previous voice conversation; the relationship between the first user and the second user in the previous voice conversation is different from the relationship in the current voice conversation; The speaker is used to send out the response message corresponding to the first voice information.
12. The electronic device according to claim 11, characterized in that, The electronic device further includes a display screen, which is coupled to the processor; the display screen is used to display the image information of the first user.
13. The electronic device according to claim 12, wherein The face model of the first user is stored in the electronic device; The display screen is further used to simulate the expressions of the first user during a voice conversation with the second user and display the face model; wherein, the expressions of the first user in the face model change dynamically.
14. The electronic device according to any one of claims 11-13, characterized in that The microphone is further used to obtain second voice information, which is the voice information during the voice conversation between the first user and the second user; The processor is further used to analyze the second voice information to obtain the voice characteristics during the voice conversation between the first user and the second user, and store the voice characteristics; Wherein, the voice characteristics include voiceprint characteristics, tone characteristics and word-using characteristics. The tone characteristics are used to indicate the tone during the voice conversation between the first user and the second user, and the word-using characteristics are used to indicate the habitual vocabulary during the voice conversation between the first user and the second user.
15. The electronic device according to claim 14, wherein The processor is further used to store the voice conversation record of the electronic device simulating the voice conversation between the first user and the second user in the second voice information.
16. The electronic device according to claim 11, characterized in that The microphone is further used to receive third voice information; The processor is further used to recognize the third voice information in response to the third voice information; Based on the recognition that the third voice information is the voice information of the second user, the electronic device simulates the voice of the first user and, in the dialogue manner of the voice conversation between the first user and the second user, the speaker is further used to emit the voice response information of the third voice information.
17. The electronic device according to claim 16, wherein The processor is further used to obtain the schedule information of the first user, and the schedule information is used to indicate the schedule arrangement of the first user; Wherein, the emission of the voice response information of the third voice information includes: The electronic device refers to the schedule information and emits the voice response information of the third voice information.
18. The electronic device according to claim 11, wherein The processor is further used to store the voice conversation record of the electronic device simulating the voice of the first user and the second user; wherein, the voice conversation record is used to indicate the content of the voice conversation; Send the voice conversation record to the electronic device of the first user.
19. The electronic device according to claim 11, wherein The processor is further used to Store the voice conversation record of the electronic device simulating the voice conversation between the first user and the second user; Extract the keywords of the voice conversation of the electronic device simulating the voice conversation between the first user and the second user from the voice conversation record; Send the keywords to the electronic device of the first user.
20. The electronic device according to claim 11, characterized in that, The electronic device further includes a camera, which is coupled to the processor; The camera is used to obtain the image information and action information of the second user, and the processor is further used to store the image information and action information of the second user.
21. A server, characterized in that, including a memory and one or more processors; the memory and the one or more processors are coupled; wherein, the memory is used for storing computer program code, the computer program code includes computer instructions, and when the processor executes the computer instructions, the server is caused to execute the method according to any one of claims 1-10.
22. A computer-readable storage medium, characterized in that, including computer instructions, and when the computer instructions are run on an electronic device, the electronic device is caused to execute the method according to any one of claims 1-10.
Citation Information
Patent Citations
Voice interaction method and device, equipment and medium
CN110633357A
Electronic personal interactive device
US20110283190A1
Cited By
Image forming apparatus
US12725613B2
Image forming apparatus
US20250037713A1