An interaction method, an electronic device, and a medium

By transmitting only the voice wake-up status animation data and displaying the interactive interface content when necessary in the scenario of mobile phone and vehicle interconnection, the problem of interface interruption when the voice assistant is woken up is solved, improving user experience and device resource utilization efficiency.

CN118280355BActive Publication Date: 2026-01-16HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211727325.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2026-01-16
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

In scenarios where mobile phones and in-vehicle systems are interconnected, how can we improve the speed and convenience of interaction while ensuring driving safety, especially by avoiding interrupting the user's interaction with the in-vehicle interface when waking up the voice assistant?

Method used

Instead of directly projecting the interface, the first electronic device detects the voice assistant's wake-up command and sends the voice wake-up status animation data to the second electronic device. The second electronic device then draws the voice wake-up status animation and displays the interactive interface content when necessary, avoiding the transmission of unnecessary interface data.

Benefits of technology

It improves the user experience, avoids interruptions in interface interaction, saves device resources, and can accurately execute user intent commands, thus improving the speed and security of interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118280355B_ABST
    Figure CN118280355B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of communication, and discloses an interaction method, an electronic device and a medium, wherein the interaction method comprises the following steps: a first electronic device and a second electronic device establish a connection, and the second electronic device displays a first interface, and the first interface comprises first display content; the first electronic device detects a wake-up instruction of a voice assistant; the first electronic device sends first data corresponding to the wake-up instruction to the second electronic device, wherein the first data comprises data corresponding to a voice wake-up state animation effect; and the second electronic device displays a second interface based on the first data and the first display content, and the second interface comprises the first display content and second display content corresponding to the first data. Based on the above scheme, when a user wakes up the voice assistant of the first electronic device, the first electronic device only sends data corresponding to the voice wake-up state animation effect to the second electronic device, so that the interaction of the user with the interface of the second electronic device is not interrupted, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of communication, in particular to an interaction method, an electronic device and a medium. BACKGROUND

[0002] With the development and popularization of electronic devices such as mobile phones, electronic devices such as mobile phones have been carried by users everywhere. At the same time, with the development of automobile intelligence, more and more cars are equipped with vehicle entertainment systems (or called central control screen systems, car machines). Among them, the hardware computing capability of the car machine of the car is relatively long in the iteration cycle, so the application and service of the car machine are not rich enough. And the electronic device has the latest computing hardware platform, the latest software platform, the latest high-speed mobile data network connection capability, and various user habit applications and services, etc., so the car machine and the electronic device are generally connected to realize hardware assistance or application ecological sharing of the car machine and the electronic device. For example, after the electronic device and the car machine are connected through wired or wireless mode, the user can control and use the application on the electronic device through the input and output devices of the car (such as car machine large screen, button knob, car microphone, loudspeaker, camera, etc.).

[0003] In the mobile phone and car machine interconnection scene, how to ensure driving safety while improving the convenience of interaction has become the research direction of the industry. SUMMARY

[0004] In order to realize the driving safety while improving the convenience of interaction in the mobile phone and car machine interconnection scene, the present application provides an interaction method, an electronic device and a medium.

[0005] In the first aspect, the present application provides an interaction method, comprising: a first electronic device and a second electronic device establish a connection, and the second electronic device displays a first interface, the first interface comprising first display content; the first electronic device detects a wake-up instruction of a voice assistant; the first electronic device sends first data corresponding to the wake-up instruction to the second electronic device, wherein the first data comprises data corresponding to a voice wake-up state animation; the second electronic device displays a second interface based on the first data and the first display content, the second interface comprising the first display content and second display content corresponding to the first data.

[0006] Based on the above scheme, when the user wakes up the voice assistant of the first electronic device, the first electronic device will not directly project the current display interface to the second electronic device for display, but only send data corresponding to the voice wake-up state animation to the second electronic device, and the second electronic device draws data corresponding to the voice wake-up state animation. In this way, the interaction of the user with the interface of the second electronic device is not interrupted, and the user experience is improved.

[0007] In some embodiments, the first electronic device can be a mobile phone, and the second electronic device can be a car machine, wherein the first display content can be display content of a current display interface of the car machine, and the second interface can refer to displaying a voice wake-up state animation on the current display interface of the car machine.

[0008] In a possible implementation, the interaction method further includes, in response to the user interacting with the first electronic device, the second electronic device displaying a third interface, the third interface including the first display content and third display content corresponding to the interaction operation.

[0009] It can be understood that the above-mentioned user interaction operation with the first electronic device can refer to the user interacting with the voice assistant of the first electronic device through a voice instruction or the like. In some embodiments, the second electronic device can receive the voice instruction, send data corresponding to the voice instruction to the first electronic device, and the first electronic device can perform ASR recognition, intent understanding, and the like on the voice instruction. In addition, the first electronic device can send data corresponding to the third display content, i.e., text data obtained by ASR recognition, result data of intent understanding (multi-round user selection GUI), voice interaction content data such as intelligent tips, and voice interaction state animation data (e.g., listening animation, broadcast animation) to the car machine for display.

[0010] During the user interaction with the first electronic device, the first electronic device can only send voice interaction content and voice interaction state animation, without sending interface data corresponding to the application interface of the mobile phone. In this way, the user's interaction with the interface of the second electronic device will not be interrupted, and the user experience is improved.

[0011] In a possible implementation, the third display content includes voice interaction content and voice interaction state animation corresponding to the interaction operation.

[0012] In a possible implementation, in response to the user interacting with the first electronic device, the second electronic device displaying the third interface includes: in response to the user interacting with the first electronic device, the first electronic device sending data corresponding to the voice interaction content and data corresponding to the voice interaction state animation to the second electronic device; and the second electronic device displaying the third interface based on the data corresponding to the voice interaction content, the data corresponding to the voice interaction state animation, and the first display content.

[0013] In a possible implementation, the voice interaction content includes text corresponding to a user voice instruction, intelligent tips, and intent option content; and the voice interaction state animation includes a listening animation or a broadcast animation.

[0014] In a possible implementation, the first electronic device detects a user intention instruction, determines an execution device of the intention instruction based on a type of the intention instruction and current running states of the first electronic device and the second electronic device; and when the execution device of the intention instruction is determined to be the second electronic device, the first electronic device sends the intention instruction to the second electronic device.

[0015] It can be understood that, when detecting the user intention instruction, the first electronic device determines the execution device of the intention instruction based on the type of the intention instruction and the current running states of the first electronic device and the second electronic device, so that when the intention instruction is an intention that the first electronic device cannot execute, the second electronic device can execute the intention instruction, ensuring execution of the intention instruction, and in addition, the execution device of the intention instruction can be determined based on the current running states of the first electronic device and the second electronic device to be the device that is most suitable for executing the user intention instruction at present, improving user experience. The current running states of the first electronic device and the second electronic device can refer to applications that the first electronic device and the second electronic device are currently running, to determine whether an application that can execute the intention exists in a running state in the first electronic device and the second electronic device.

[0016] In a possible implementation, the second electronic device acquires an intention parameter and a slot parameter based on the intention instruction, and executes the intention instruction based on the intention parameter and the slot parameter.

[0017] It can be understood that, in the embodiments of the present application, the second electronic device executes the intention instruction based on the slot parameter, which can enable accurate execution of the user intention instruction. For example, the function in the application can be accurately controlled.

[0018] In a possible implementation, determining the execution device of the intention instruction based on the type of the intention instruction and the current running states of the first electronic device and the second electronic device includes: when the type of the intention instruction is an exclusive type, determining the execution device of the intention instruction to be an execution device corresponding to the exclusive type; and when the intention instruction is not of the exclusive type, determining the execution device of the intention instruction based on the current running states of the first electronic device and the second electronic device.

[0019] In a possible implementation, determining the execution device of the intention instruction based on the current running states of the first electronic device and the second electronic device includes: taking a device that is running a first application or a device that is displaying an interface of the first application as the execution device of the intention instruction, the first application being an application that can execute the intention instruction.

[0020] It can be understood that, taking the device that is running the application that can execute the intention instruction or the device that is displaying the interface of the application that can execute the intention instruction as the execution device of the intention instruction can save device resources without restarting the corresponding application of another device.

[0021] In a possible implementation, when the execution device of the intent instruction is the first electronic device, and the type of the intent instruction is a preset type of execution process without screen projection display, the first electronic device does not send data corresponding to a display interface in the execution process of the intent instruction to the second electronic device.

[0022] In some embodiments, the preset type of execution process without screen projection display can be a type such as music playing type without the need to watch the screen. In this way, the transmission of unnecessary interface data and the drawing of the related interface can be effectively saved, and the device resources can be saved.

[0023] In a possible implementation, when the execution device of the intent instruction is the first electronic device, and the type of the intent instruction is a preset type of execution process with screen projection display, the first electronic device sends first data corresponding to a display interface in the execution process of the intent instruction to the second electronic device; and the first electronic device displays a fourth interface based on the first data corresponding to the display interface.

[0024] In a possible implementation, the fourth interface includes first display content and fourth display content corresponding to the first data.

[0025] In a possible implementation, the first display content and the fourth display content are displayed in different regions of the screen of the first electronic device.

[0026] It can be understood that the second electronic device can display the first display content and the fourth display content in split screen.

[0027] In a possible implementation, the fourth interface includes fourth display content corresponding to the first data.

[0028] It can be understood that the second electronic device can display the first display content in full screen.

[0029] In a possible implementation, the first data does not include interface data corresponding to a current application interface of the first electronic device.

[0030] In a possible implementation, the first electronic device is a mobile phone, and the second electronic device is a car machine.

[0031] In some embodiments, the car machine can include a car machine host CPU, that is, the car machine is a control system capable of executing the interactive method of the present application. In some other embodiments, the car machine can also be a combined structure including a car machine host CPU and a microphone, a loudspeaker, a steering wheel voice key, a central control large screen, a USB, a Bluetooth physical component, a Wi-Fi physical component, a network antenna physical component, and the like.

[0032] In some embodiments, the mobile phone can comprise a memory for storing a computer program comprising program instructions; and a processor for executing the program instructions to cause the mobile phone to perform the interaction method mentioned in the present application. The car machine can comprise a memory for storing a computer program comprising program instructions; and a processor for executing the program instructions to cause the mobile phone to perform the interaction method mentioned in the present application.

[0033] In a second aspect, the present application provides an electronic device, the electronic device being a first electronic device, the first electronic device establishing a connection with a second electronic device; the first electronic device being configured to detect a wake-up instruction of a voice assistant; the first electronic device being configured to send first data corresponding to the wake-up instruction to the second electronic device, wherein the first data comprises data corresponding to a voice wake-up state effect, and does not comprise interface data corresponding to a current application interface of the first electronic device.

[0034] In a possible implementation, the first electronic device is configured to, in response to a user interacting with the voice assistant, send data corresponding to voice interaction content and data corresponding to a voice interaction state effect to the second electronic device.

[0035] In a possible implementation, the first electronic device is configured to detect a user intent instruction, determine an execution device of the intent instruction based on a type of the intent instruction and a current running state of the first electronic device and the second electronic device, and send the intent instruction to the second electronic device when it is determined that the execution device of the intent instruction is the second electronic device.

[0036] In a third aspect, the present application provides an electronic device, the electronic device being a second electronic device, the second electronic device establishing a connection with a first electronic device; the second electronic device being configured to display a first interface, the first interface comprising first display content; the second electronic device being configured to display a second interface based on first data sent by the first electronic device and the first display content, the second interface comprising the first display content and second display content corresponding to the first data, wherein the first data comprises data corresponding to a voice wake-up state effect, and does not comprise data corresponding to a current application interface of the first electronic device.

[0037] In a possible implementation, the second electronic device is configured to, in response to a user interacting with a voice assistant of the first electronic device, display a third interface, the third interface comprising the first display content and third display content corresponding to the interaction operation.

[0038] In a fourth aspect, the present application provides an electronic device, the electronic device being a first electronic device, the first electronic device being connected with a second electronic device; the first electronic device comprising: a voice assistant application, configured to detect a wake-up instruction of a voice assistant, and send first data corresponding to the wake-up instruction to a first voice assistant atomic service module of the first electronic device, wherein the first data comprises data corresponding to a voice wake-up state animation, and does not comprise data corresponding to a current application interface of the first electronic device; and the first voice assistant atomic service module, configured to send the data corresponding to the voice wake-up state animation to a second voice assistant atomic service module of the second electronic device.

[0039] In a possible implementation, the voice assistant application is configured to, in response to an interactive operation of a user with the voice assistant, send data corresponding to voice interaction content and data corresponding to a voice interaction state animation to the first voice assistant atomic service module; and the first voice assistant atomic service module is configured to send the data corresponding to the voice interaction content and the data corresponding to the voice interaction state animation to the second voice assistant atomic service module of the second electronic device.

[0040] In a possible implementation, the second electronic device further comprises a first voice intent collaborative distribution module; the voice assistant application is configured to detect a user intent instruction, determine an execution device of the intent instruction based on a type of the intent instruction and a current running state of the first electronic device and the second electronic device, and send the intent instruction to the first voice intent collaborative distribution module when it is determined that the execution device of the intent instruction is the second electronic device; and the first voice intent collaborative distribution module is configured to send the intent instruction to a voice intent collaborative distribution module of the second electronic device.

[0041] In a fifth aspect, the present application provides an electronic device, the electronic device being a second electronic device, the second electronic device being connected with the second electronic device; the second electronic device comprising: a voice assistant distributed collaborative UI module, configured to control a screen of the second electronic device to display a first interface, the first interface comprising first display content; and the voice assistant distributed collaborative UI module, configured to display a second interface based on first data sent by the first electronic device and the first display content, the second interface comprising the first display content and second display content corresponding to the first data, wherein the first data comprises data corresponding to a voice wake-up state animation, and does not comprise data corresponding to a current application interface of the first electronic device.

[0042] In a possible implementation, the voice assistant distributed collaborative UI module is configured to, in response to an interactive operation of a user with a voice assistant of the first electronic device, control the screen of the second electronic device to display a third interface, the third interface comprising the first display content and third display content corresponding to the interactive operation.

[0043] In a sixth aspect, the present application provides an electronic device, comprising: a memory, configured to store a computer program, the computer program comprising program instructions; and a processor, configured to execute the program instructions, so that the electronic device performs the interaction method mentioned in the present application.

[0044] In a seventh aspect, the present application provides a computer-readable storage medium, which stores a computer program, the computer program comprising program instructions, and the program instructions are run by an electronic device to make the electronic device perform the interaction method mentioned in the present application. BRIEF DESCRIPTION OF DRAWINGS

[0045] Figure 1a According to some embodiments of the present application, a hardware structure schematic diagram of an electronic device is shown;

[0046] Figure 1b According to some embodiments of the present application, a software structure schematic diagram of an electronic device is shown;

[0047] Figure 1c According to some embodiments of the present application, a structure schematic diagram of a car and a mobile phone is shown;

[0048] Figure 1d According to some embodiments of the present application, a structure schematic diagram of a car and a mobile phone is shown;

[0049] Figure 2 According to some embodiments of the present application, a process schematic diagram of interaction between a car machine and a mobile phone is shown;

[0050] Figure 3a According to some embodiments of the present application, a scenario schematic diagram of mobile phone driving mode desktop screen projection is shown;

[0051] Figure 3b According to some embodiments of the present application, a scenario schematic diagram of mobile phone driving mode desktop screen projection is shown;

[0052] Figure 4 According to some embodiments of the present application, an interaction mode between a mobile phone and a car machine is shown;

[0053] Figure 5 According to some embodiments of the present application, a process schematic diagram of interaction between a car machine and a mobile phone is shown;

[0054] Figure 6 According to some embodiments of the present application, a flow schematic diagram of an interaction method is shown;

[0055] Figure 7 According to some embodiments of the present application, a scenario schematic diagram of an interaction process is shown;

[0056] Figure 8According to some embodiments of the present application, a scenario diagram of an interaction process is shown.

[0057] Figure 9a According to some embodiments of the present application, a delivery of voice interaction process data and a car machine reverse control process diagram is shown.

[0058] Figure 9b According to some embodiments of the present application, a process diagram of intent coordination distribution is shown.

[0059] Figure 9c According to some embodiments of the present application, a flow diagram of an interaction method is shown.

[0060] Figure 10 According to some embodiments of the present application, a flow diagram of an interaction method is shown.

[0061] Figure 11a According to some embodiments of the present application, a deployment scheme diagram of voice mutual assistance capability implementation is shown.

[0062] Figure 11b According to some embodiments of the present application, a flow diagram of an interaction method is shown. DETAILED DESCRIPTION

[0063] The illustrative embodiments of the present application include but are not limited to an interaction method, an electronic device, and a medium.

[0064] Before the interaction method of the present application is described in detail, first, the electronic device mentioned in the present application is introduced. The electronic device can be the first electronic device mentioned in the present application, or the second electronic device mentioned in the present application. Among them, the first electronic device and the second electronic device can include but are not limited to a communication module capable of executing the interaction method of the present application, or a mobile phone, a personal computer, a tablet computer, a wearable device (such as a smart watch, a smart bracelet, etc.), a car machine, etc. including the above-mentioned communication module.

[0065] As Figure 1aAs shown, the electronic device can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0066] It can be understood that the processor 110 can be configured to run an operating system of the electronic device to perform the steps of the electronic device side in the interaction method of the embodiments of the present application.

[0067] In some embodiments, the processor 110 can include one or more interfaces. The interface can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0068] The USB interface 130 is an interface conforming to the USB standard specification, and can be a Mini USB interface, a Micro USB interface, a USB Type-C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device, and can also be used to transmit data between the electronic device and a peripheral device. The interface can also be used to connect a headset to play audio through the headset. The interface can also be used to connect other electronic devices, such as an AR device, etc.

[0069] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a structural limitation on the electronic device. In other embodiments of the present application, the electronic device can also use different interface connection methods or combinations of multiple interface connection methods.

[0070] The charging management module 140 is used to receive charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from a wired charger through the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input through a wireless charging coil of the electronic device. The charging management module 140 can charge the battery 142 while also supplying power to the electronic device through the power management module 141.

[0071] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to supply power to the processor 110, the internal memory 121, the display screen 194, the camera 193, and the wireless communication module 160, etc. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, battery health status (leakage, impedance), etc. In other embodiments, the power management module 141 can also be disposed in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be disposed in the same device.

[0072] The wireless communication function of the electronic device can be realized through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, etc.

[0073] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in combination with a tuning switch.

[0074] The mobile communication module 150 can provide a solution for wireless communication including 2G / 3G / 4G / 5G, etc. The mobile communication module 150 can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via the antenna 1, and perform filtering, amplification, etc. on the received electromagnetic waves, and transfer the processed signals to the modem processor for demodulation. The mobile communication module 150 can also amplify signals modulated by the modem processor, and radiate the amplified signals as electromagnetic waves via the antenna 1. In some embodiments, at least part of the functions of the mobile communication module 150 can be provided in the processor 110. In some embodiments, at least part of the functions of the mobile communication module 150 can be provided in the same device as at least part of the processor 110.

[0075] The modem processor can include a modulator and a demodulator. The modulator can modulate a low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator can demodulate a received electromagnetic wave signal into a low-frequency baseband signal. The demodulator can then transfer the demodulated low-frequency baseband signal to the baseband processor for processing. The low-frequency baseband signal processed by the baseband processor can be transferred to the application processor. The application processor can output a sound signal through an audio device (not limited to the speaker 170A, the microphone 170B, etc.), or display an image or a video through the display screen 194. In some embodiments, the modem processor can be a separate device. In other embodiments, the modem processor can be provided in the same device as the mobile communication module 150 or other functional modules, independently of the processor 110.

[0076] The wireless communication module 160 can provide a solution for wireless communication including wireless local area networks (WLAN) (e.g., a wireless fidelity (Wi-Fi) network), Bluetooth (BT), a global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc. The wireless communication module 160 can be one or more devices that integrate at least one communication processing module. The wireless communication module 160 can receive electromagnetic waves via the antenna 2, perform frequency modulation and filtering on the electromagnetic wave signals, and transmit the processed signals to the processor 110. The wireless communication module 160 can also receive signals to be transmitted from the processor 110, perform frequency modulation and amplification, and radiate the processed signals as electromagnetic waves via the antenna 2.

[0077] In some embodiments, the antenna 1 and the mobile communication module 150 are coupled, and the antenna 2 and the wireless communication module 160 are coupled, so that the electronic device can communicate with a network and other devices through wireless communication technology.

[0078] The electronic device implements a display function through a GPU, a display screen 194, and an application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs that execute program instructions to generate or change display information.

[0079] The display screen 194 can also be referred to as a screen, and is used to display images, videos, etc. In some embodiments, the electronic device can include 1 or N display screens 194, N being a positive integer greater than 1.

[0080] The electronic device can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor, etc.

[0081] The ISP is used to process data fed back by the camera 193. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing to convert it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be provided in the camera 193.

[0082] The camera 193 is used to capture still images or videos.

[0083] The digital signal processor is used to process digital signals, in addition to being able to process digital image signals, it can also process other digital signals. For example, when the electronic device selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0084] The video codec is used to compress or decompress digital videos. The electronic device can support one or more video codecs. In this way, the electronic device can play or record videos in multiple encoding formats, such as: moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.

[0085] The NPU is a neural-network (NN) computing processor. By drawing on the structure of a biological neural network, for example, by drawing on the transmission mode between human brain neurons, the NPU can quickly process input information and can also constantly self-learn. Through the NPU, intelligent cognition and other applications of an electronic device can be realized, for example, image recognition, face recognition, voice recognition, text understanding, and the like.

[0086] The external memory interface 120 can be used to connect an external memory card, such as a MicroSD card, to realize the expansion of the storage capacity of the electronic device. The external memory card communicates with the processor 110 through the external memory interface 120 to realize a data storage function. For example, music, video, and the like are saved in the external memory card.

[0087] The internal memory 121 can be used to store computer executable program code, which includes instructions. The internal memory 121 can include a program storage area and a data storage area.

[0088] The electronic device can realize an audio function through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the earphone interface 170D, and the application processor, and the like. For example, music playing, recording, and the like.

[0089] The audio module 170 is used to convert digital audio information into an analog audio signal output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used to encode and decode an audio signal. In some embodiments, the audio module 170 can be disposed in the processor 110, or part of the functions of the audio module 170 can be disposed in the processor 110.

[0090] The speaker 170A, also known as a "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device can listen to music or listen to a hands-free call through the speaker 170A.

[0091] The receiver 170B, also known as a "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device answers a call or a voice message, the receiver 170B can be held close to the ear to listen to the voice.

[0092] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Electronic devices can have at least one microphone 170C. In some embodiments, electronic devices can have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic devices can have three, four, or more microphones 170C, enabling sound signal collection, noise reduction, sound source identification, and directional recording, among other functions.

[0093] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. The electronic device can receive button input and generate key signal inputs related to user settings and function control of the electronic device.

[0094] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can correspond to touch operations performed on different applications (such as taking photos, playing audio, etc.). Motor 191 can also correspond to different vibration feedback effects for touch operations performed on different areas of the display screen 194. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.

[0095] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.

[0096] In some embodiments, the software system of an electronic device may adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application uses the layered architecture of the Android system as an example to exemplify the software structure of a mobile phone. Figure 1b As shown, the layered architecture divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0097] The application layer can include a series of application packages. These application packages can include applications such as voice assistants, gallery, calendar, calling, maps, navigation, WLAN, Bluetooth, music, video, and SMS.

[0098] The application framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The application framework layer includes some pre-defined functions.

[0099] The application framework layer can include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like. In some embodiments, the voice wake-up module, the voice assistant atomization service module, and the voice intent collaborative distribution module can be arranged in the operating system according to actual needs, for example, can be arranged in the framework layer of the application.

[0100] The window manager is used to manage window programs. The window manager can obtain the size of a display screen, determine whether there is a status bar, lock a screen, and the like.

[0101] The content provider is used to store and obtain data, and make the data accessible to applications. The data can include videos, images, audios, dialed and received calls, browsing history and bookmarks, phone books, and the like.

[0102] The view system includes visual controls, such as a control for displaying text and a control for displaying pictures. The view system can be used to build an application. A display interface can be composed of one or more views. For example, a display interface including a short message notification icon can include a view for displaying text and a view for displaying pictures.

[0103] The phone manager is used to provide communication functions of a mobile phone. For example, management of a call state (including call connection and call hang-up).

[0104] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, and the like.

[0105] The notification manager enables an application to display notification information in a status bar. The notification manager can be used to convey a message of the notification type, which can automatically disappear after a short stay without user interaction. For example, the notification manager is used to notify a download completion and a message reminder. The notification manager can also be a notification in the form of a chart or a scroll bar text appearing in the top status bar of the system, for example, a notification of an application running in the background, and can also be a notification in the form of a dialogue window appearing on the screen. For example, a text information is prompted in the status bar, a prompt sound is emitted, the electronic device vibrates, a light flashes, and the like.

[0106] The Android runtime includes a core library and a virtual machine. The Android runtime is responsible for scheduling and management of the Android system.

[0107] The core library includes two parts: one part is the function function called by the java language, and the other part is the core library of Android.

[0108] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform the functions of object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0109] The system library can include a plurality of function modules. For example: surface manager, media library, three-dimensional graphics processing library (for example: OpenGLES), 2D graphics engine (for example: SGL) and the like.

[0110] The surface manager is used to manage the display subsystem, and provides a plurality of applications with the fusion of 2D and 3D layers.

[0111] The media library supports a plurality of commonly used audio, video format playback and recording, and static image files and the like. The media library can support a plurality of audio and video coding formats, for example: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG and the like.

[0112] The three-dimensional graphics processing library is used to realize three-dimensional graphics drawing, image rendering, synthesis, and layer processing and the like. The 2D graphics engine is a drawing engine for 2D drawing.

[0113] The kernel layer is the layer between hardware and software. The kernel layer at least includes display driver, camera driver, audio driver, sensor driver.

[0114] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the first electronic device and the second electronic device can include more or less components than the illustration, or combine certain components, or split certain components, or different component arrangements, or have similar function component arrangements, etc. The illustrated components can be implemented in hardware, software or a combination of software and hardware.

[0115] For example, in some embodiments, the first electronic device is a mobile phone. As shown in Figure 1c The mobile phone can include CPU, microphone, loudspeaker, USB interface, Bluetooth physical component, Wi-Fi physical component, network antenna physical component, screen and the like.

[0116] The CPU is configured to run an operating system of the mobile phone to implement the interactive method of the present application. For example, the voice assistant application, the interconnection protocol, and the related service application (e.g., music application, navigation application, etc.) in the operating system can be run. The architecture of the operating system of the mobile phone is described in detail below, and thus is not described herein.

[0117] The USB interface can be used to realize wired connection between the mobile phone and the electronic device such as the car machine, for example, connection by a connection line. The USB interface is an interface conforming to the USB standard specification, and can be a MiniUSB interface, a MicroUSB interface, a USB Type C interface, etc. The USB interface can be used to connect a charger to charge the mobile phone, and can also be used to transmit data between the mobile phone and the peripheral device. It can also be used to connect a headset to play audio through the headset.

[0118] The Bluetooth physical component is used to realize Bluetooth short-distance communication (generally within 10 m) between the mobile phone and other electronic devices such as the car machine.

[0119] The Wi-Fi physical component is used to realize Wi-Fi communication between the mobile phone and other electronic devices such as the car machine.

[0120] The network antenna physical component can include a first antenna and a second antenna for transmitting and receiving electromagnetic wave signals. The first antenna and the second antenna can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, the first antenna can be multiplexed as a diversity antenna of a wireless local area network. In some other embodiments, the first antenna and the second antenna can be used in combination with a tuning switch.

[0121] The screen is used to display the human-computer interaction interface, images, videos, etc. The screen includes a display panel.

[0122] The microphone and the loudspeaker can be used for voice interaction. Specifically, the loudspeaker, also known as the speaker, is used to convert an audio electrical signal into an acoustic signal. The mobile phone can play voice through the loudspeaker, and can receive user voice through the microphone.

[0123] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the mobile phone. In some other embodiments of the present application, the mobile phone can include more or fewer components than the illustration, or combine certain components, or split certain components, or different component arrangements, or component arrangements with similar functions, etc. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0124] In some embodiments, the second electronic device can be a car machine. The partial structure of the car containing the car machine is introduced as follows. Figure 1cAs shown, the car can include: a microphone, a loudspeaker, a steering wheel voice button, a car machine, a central control large screen, a USB, a Bluetooth physical component, a Wi-Fi physical component, a network antenna physical component.

[0125] The car machine can include a car machine host CPU, which can include one or more processing units. Different processing units can be independent devices or integrated in one or more processors. Storage units can be provided in the processors for storing instructions and data. The car machine host CPU can be used to run the operating system of the car machine to execute the interactive method of the present application, such as the voice assistant application in the operating system, the interconnection protocol, and the related business application (such as music application, navigation application, etc.), and also to control the central control large screen to display the HMI interface. The architecture of the operating system of the car machine is described in detail later, and will not be described here.

[0126] The USB interface can be used for physical connection with electronic devices such as car machines, such as connection line connection. The USB interface is an interface that meets the USB standard specification, and can be a MiniUSB interface, a MicroUSB interface, a USB Type C interface, etc. The USB interface can be used to connect a charger to charge a mobile phone, and can also be used to transmit data between a mobile phone and a peripheral device. It can also be used to connect a headset to play audio through the headset.

[0127] The Bluetooth physical component is used to realize Bluetooth short-distance communication (generally within 10m) between the mobile phone and other electronic devices such as the car machine.

[0128] The Wi-Fi physical component is used to realize Wi-Fi communication between the mobile phone and other electronic devices such as the car machine.

[0129] The network antenna physical component can include a first antenna and a second antenna for transmitting and receiving electromagnetic wave signals. The first antenna and the second antenna can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example: the first antenna can be multiplexed as a diversity antenna of a wireless local area network. In some other embodiments, the first antenna and the second antenna can be used in combination with a tuning switch.

[0130] The central control large screen is used to display the HMI, and the central control large screen includes a display panel.

[0131] The microphone and the loudspeaker can be used for voice interaction. Specifically, the loudspeaker, also known as the "speaker", is used to convert audio electrical signals into sound signals. The mobile phone can play voice through the loudspeaker, and can receive user voice through the microphone.

[0132] The steering wheel voice button is used to wake up the voice assistant.

[0133] In some embodiments, the car machine can include a car machine host CPU, that is, the car machine is a control system capable of executing the interactive method of the present application, in other embodiments, the car machine can also be a combination structure including the car machine host CPU and the microphone, the loudspeaker, the steering wheel voice key, the central control large screen, the USB, the Bluetooth physical component, the Wi-Fi physical component, the network antenna physical component and the like. Or it can include more or less components than those shown in Figure 1 and mentioned above, or combine some components, or split some components, or different component arrangement, or component arrangement with similar functions, etc. The components shown can be realized in hardware, software or a combination of software and hardware.

[0134] In some embodiments, as shown in Figure 1d The operating system of the mobile phone can include:

[0135] The voice assistant application is used to recognize the user voice instruction, understand the user voice instruction, determine the execution device of the user's intention instruction, and respond to the voice.

[0136] The voice wake-up module is used to listen to the user voice wake-up instruction and draw the corresponding wake-up state animation of the mobile phone voice assistant.

[0137] The voice assistant atomization service module is used to send the voice interaction state data and interface data to the car machine end.

[0138] The voice intention collaborative distribution module is used to determine the execution device of the user's intention instruction, and send the intention instruction to the car machine when the execution device of the user's intention instruction is the car machine. And used to receive the intention instruction sent by the car machine that needs to be executed by the mobile phone.

[0139] The business application is used to execute the corresponding intention instruction, and send the video stream data obtained by encoding the application interface in the execution process to the voice assistant atom service module of the mobile phone.

[0140] The operating system of the car machine can include:

[0141] The voice assistant application is used to recognize the user voice instruction, understand the user voice instruction, determine the execution device of the user's intention instruction, and respond to the voice.

[0142] The voice wake-up module is used to listen to the user voice wake-up instruction and draw the corresponding wake-up state animation of the car machine voice assistant.

[0143] The voice assistant distributed collaborative UI module is used to draw the corresponding interface on the car machine display interface according to the received voice interaction state data and interface data.

[0144] The voice assistant atomic service data module is configured to receive voice interaction state data and interface data sent by the mobile phone connected to the car machine, and send the received voice interaction state data and interface data to the distributed collaborative UI module.

[0145] The voice intent collaborative distribution module is configured to determine an execution device of the user's intent instruction, and send the intent instruction to the mobile phone when it is determined that the execution device of the user's intent instruction is the mobile phone. The voice intent collaborative distribution module is also configured to receive the intent instruction sent by the mobile phone and required to be executed by the car machine.

[0146] The service application is configured to execute the corresponding intent instruction.

[0147] The following describes the interaction process between distributed devices in some embodiments by taking the interaction process between the mobile phone and the car machine as an example. Figure 2 As shown in FIG. 1, the interaction process between the car machine and the mobile phone can include the following steps.

[0148] 101: The mobile phone and the car machine establish a connection.

[0149] It can be understood that the mobile phone and the car machine can establish a connection in any implementable manner, such as through Bluetooth, a wireless fidelity (Wi-Fi) network, or a connection line.

[0150] 102: The user wakes up the mobile phone voice assistant through a side control button.

[0151] It can be understood that the mobile phone voice assistant can be woken up in any manner, such as through a voice wake-up word or through a car side control button.

[0152] 103: The car machine plays a response sound.

[0153] It can be understood that in the embodiments of the present application, the mobile phone sends the response audio data to the car machine, and the car machine plays the response sound based on the audio data.

[0154] 104: The mobile phone projects a driving mode interface to the car machine, and the car machine displays a voice wake-up animation effect.

[0155] It can be understood that the mobile phone driving mode interface (desktop) refers to the current display interface of the mobile phone in the connected state with the car machine. In some embodiments, the mobile phone driving mode desktop can also be a display interface obtained by removing some non-core data from the current display interface of the mobile phone. For example, a display interface obtained by removing some controls or icons from the current display interface.

[0156] 105: The user performs voice interaction.

[0157] It can be understood that the user can issue an intent instruction in a voice manner, for example, the user issues a voice instruction of "navigate to address A", and the car machine can receive the intent instruction of the user through a microphone or the like.

[0158] 106: The mobile phone performs ASR recognition and intent understanding.

[0159] In some embodiments, the car machine can send the intent instruction to the mobile phone. The mobile phone performs ASR recognition, intent understanding, and the like on the intent instruction, obtains voice interaction information, and screens the current interface of the mobile phone including the voice interaction information to the car machine.

[0160] 107: The mobile phone screens and displays the voice interaction information.

[0161] It can be understood that the car machine can display the current interface of the mobile phone, and the voice interaction information can be displayed on the current interface of the mobile phone.

[0162] For example, the mobile phone displays the text corresponding to the intent instruction and confirms the intent instruction. For example, when the user issues a voice instruction of "navigate to address A", the mobile phone can display the text corresponding to the voice instruction of "navigate to address A" on the current interface, understand the voice instruction, and display multiple options after semantic understanding, such as "navigate to address A1", "navigate to address A2", "navigate to address A", and the like, to facilitate the user to select and confirm. The display interface in the above-mentioned mobile phone and user interaction process can be screened and displayed to the car machine.

[0163] 108: The mobile phone executes the voice intent instruction of the user.

[0164] It can be understood that after confirming the intent instruction of the user, the mobile phone can execute the intent instruction of the user, for example, when the user selects "navigate to address A", the navigation application can be opened to navigate.

[0165] 109: The car machine performs voice broadcast.

[0166] In some embodiments, the mobile phone can send voice broadcast data to the car machine, and the car machine performs voice broadcast.

[0167] It can be understood that in the above-mentioned embodiments, when the mobile phone is connected with the car machine, the mobile phone will automatically enter the driving mode. When the mobile phone voice assistant is woken up, if the mobile phone driving mode desktop is not displayed on the car central screen, the mobile phone driving mode desktop will be directly applied to the car machine screen, wherein, as described above, the mobile phone driving mode desktop refers to the current display interface of the mobile phone in the state that the mobile phone is connected with the car machine, or a display interface obtained by removing some non-core data from the current display interface of the mobile phone. The mobile phone driving mode desktop can include an application shortcut navigation bar (such as a voice wake-up ball), state information, service card, application window interface, and the like.

[0168] When the phone driving mode desktop is projected to the car machine, as shown in Figure 3a , the car machine screen will directly display the phone driving mode desktop and the voice state animation. As shown in Figure 3b , during the interaction between the user and the voice assistant through the wake-up icon of the projection interface, the phone projection driving mode desktop interface is projected to the car machine screen, and the car machine can display the voice state animation, the ASR text or tips during the interaction. In this way, the current human-machine interface (HMI) of the car machine will be covered by the phone driving mode desktop, interrupting the interaction between the user and the HMI of the car machine. At this time, the user may only want to wake up the phone assistant and need the phone assistant to perform some functions, such as playing music, and does not have the need to watch the current interface of the phone, but wants to continue the interaction with the HMI of the car machine, for example, wants to continue to watch the navigation interface on the car machine screen, but at this time, the interface projected by the phone covers the application interface of the car machine, which will result in a poor user experience and in some cases, will also affect the driving safety.

[0169] In addition, in some car machine and phone interconnection solutions, the projection mode adopted is to project the application interface of the phone, that is, the display interface of the car machine only has the application interface window of the phone that is projected, and does not include other function interfaces of the driving mode desktop. In this way, since the voice interface is not projected, the voice assistant on the car machine end cannot be used for collaborative functions.

[0170] In addition, in some car machine and phone interconnection solutions, after the phone voice assistant is woken up, as shown in Figure 4 , the graphical user interface (GUI) of the voice interaction process, such as the voice interaction GUI (including automatic speech recognition (ASR) text display, intelligent tips, multi-round user selection GUI, etc.) and the voice image state animation GUI (including voice wake-up (listening) state animation, voice broadcast state animation, etc.), is displayed after the GUI interface is drawn on the phone and is projected to the car machine screen for display depending on the current display interface of the phone. After detecting the touch screen operation of the user, the car machine can send the corresponding touch screen event to the corresponding application of the phone to realize the reverse control of the corresponding application of the phone. In the above solution, the interface projected by the phone also covers the HMI of the car machine, interrupting the interaction between the user and the HMI of the car machine, resulting in a poor user experience.

[0171] To solve the above problems, the embodiment of the present application provides an interaction method, comprising: a first electronic device and a second electronic device establish a connection, when the voice assistant of the first electronic device is woken up, the first electronic device sends voice assistant wake-up state animation data to the second electronic device, without sending the current display interface data of the first electronic device, and the second electronic device displays the voice assistant wake-up state animation based on the voice assistant wake-up state animation data. In the subsequent process of voice interaction between the user and the first electronic device, the first electronic device only sends voice interaction data (i.e. the data corresponding to the aforementioned voice interaction GUI and voice image state animation GUI, etc.) of the voice interaction process to the second electronic device, and also does not send the current display interface of the first electronic device. The second electronic device can draw the voice interaction GUI and the voice image state animation GUI on the current display interface of the second electronic device based on the voice interaction data. After the voice assistant of the first electronic device receives the user intent instruction, if it is determined that the user has the demand to watch the current application interface of the first electronic device during the execution of the user intent instruction, the corresponding application interface is projected to the second electronic device for display.

[0172] Based on the above scheme, when the user wakes up the voice assistant of the first electronic device, the first electronic device will not directly project the current display interface to the second electronic device for display, but only send the voice wake-up state animation to the second electronic device, and the second electronic device draws the voice wake-up state animation. In this way, the user's interaction with the interface of the second electronic device is not interrupted, and the user experience is improved. For example, as shown in Figure 5 When the first electronic device is a mobile phone and the second electronic device is a car machine, when the user wakes up the voice assistant of the mobile phone by the wake-up word "Xiao Yi", the mobile phone can send data corresponding to the wake-up state animation to the car machine, and the car machine displays the wake-up state animation and intelligent tips "What day is today" on the original car machine HMI based on the data corresponding to the wake-up state animation.

[0173] When the user interacts with the first electronic device by voice, the first electronic device can send data corresponding to the voice interaction content and data corresponding to the voice interaction state animation to the second electronic device, and the second electronic device displays the voice interaction content and the voice interaction animation on the current display interface of the second electronic device based on the data corresponding to the voice interaction content and the data corresponding to the voice interaction state animation. In some embodiments, the voice interaction content can include text corresponding to the user voice instruction, intelligent tips, and intent understanding content, and the voice interaction state animation can include broadcast animation, listening animation, etc.

[0174] In some embodiments, the interaction operation of the user with the first electronic device can include an operation of the user issuing a voice instruction to a voice assistant of the first electronic device. For example, when the first electronic device is a mobile phone and the second electronic device is a car machine, during the interaction of the user with the voice assistant of the mobile phone, for example, when the user issues a voice instruction of "navigate to the train station", the car machine can receive the voice instruction of the user based on a microphone or the like, and can send the received voice instruction of the user to the mobile phone. When the mobile phone receives the voice instruction of the user, the mobile phone obtains the recognized text corresponding to the voice instruction, and the mobile phone can send the text "navigate to the train station" corresponding to the voice instruction to the car machine, and the car machine can display the text "navigate to the train station".

[0175] In some embodiments, the second electronic device can also perform a recognition process of the voice instruction when detecting the voice instruction of the user, for example, directly recognizing the voice instruction, understanding the intention, and the like, to obtain the corresponding voice instruction text for display. When the intention is executed, the first electronic device can project the corresponding application interface to the second electronic device according to the user's demand, so as to facilitate the user to watch and improve the user experience. Moreover, the voice interaction data includes wake-up state animation data, so that the second electronic device obtains the entrance of the wake-up voice assistant, and the user can control the application of the first electronic device through the second electronic device, that is, the second electronic device and the first electronic device can use the voice assistant cooperative function.

[0176] In addition, as mentioned above, the GUI of the voice interaction process and the voice image state animation GUI, in the scheme of completing the GUI interface drawing on the first electronic device (for example, a mobile phone) and then displaying it on the screen of the second electronic device (for example, a car machine) by relying on the current display interface of the first electronic device, when the user wakes up the voice assistant of the second electronic device, the corresponding application of the first electronic device can only be opened through the voice assistant of the second electronic device, but the specific function of the first electronic device application cannot be controlled through the voice assistant of the second electronic device, for example, only the navigation application can be opened through the voice assistant of the second electronic device, but the first electronic device map application cannot be navigated to a certain destination address through the voice assistant of the second electronic device, and again, only the music playing application can be opened through the voice assistant of the second electronic device, but the first electronic device music application cannot be controlled to play a song of a certain singer. The application of the first electronic device interconnected and shared to the second electronic device cannot be consistent with the original application of the second electronic device in the voice interaction experience of the second electronic device.

[0177] To solve the above problems, in the embodiments of the present application, when the woken-up voice assistant receives a user intent instruction, the device executing the user intent instruction can be determined according to the type of the user intent, whether there is an application capable of executing the intent currently running, and the like. For example, after the voice assistant of the first electronic device is woken up, when the user intent instruction is received, if it is determined that the first electronic device is the intent execution device, the first electronic device can execute the user intent instruction, and send the corresponding application interface data in the execution process to the second electronic device, so that the second electronic device draws the corresponding application interface. When it is determined that the second electronic device is the intent execution device, the intent instruction can be sent to the second electronic device for execution.

[0178] The manner in which the woken-up voice assistant determines the intent execution device can be: first determining whether the intent instruction type is a dedicated type, the dedicated type being an intent that can only be executed by the second electronic device or the first electronic device. When it is determined that the type of the intent instruction is a dedicated type, the device corresponding to the dedicated type is determined to be the execution device. When it is determined that the type of the intent instruction is not a dedicated type, the execution device can be determined according to the current running state of the second electronic device and the first electronic device, for example, it can be determined whether there is a device running an application capable of executing the current intent instruction. If there is, the device running the application capable of executing the current intent instruction is determined to be the execution device.

[0179] For example, when the first electronic device is a mobile phone and the second electronic device is a car machine, if the mobile phone receives a user intent instruction to turn on the vehicle air conditioner, it is determined that the type of the user intent is a vehicle control type, i.e., a dedicated type of intent executed by the car machine, and the device executing the user intent is determined to be the car machine. When the user instruction received is "navigate to address A", it is determined that the user intent is a navigation type, which is not a dedicated type of intent. At this time, the mobile phone can determine whether there is a device running a navigation application, for example, if the mobile phone is running a navigation application and the car machine is not running a navigation application, the mobile phone is used as the user intent execution device.

[0180] In addition, in some embodiments, the execution device can parse the user intention instruction to obtain an intention parameter and a slot parameter corresponding to the intention instruction, and call the intention parameter and the slot parameter corresponding to the user intention instruction to execute the user intention instruction. The intention parameter can be a characteristic parameter representing a user intention category, such as map navigation, air conditioner control, music control, etc. The slot parameter can be a specific detail characteristic parameter corresponding to the characteristic parameter of the user intention category. For example, the specific detail characteristic parameter corresponding to the map navigation can be a navigation destination, the specific detail characteristic parameter corresponding to the air conditioner control can be a specific temperature, and the specific detail characteristic parameter corresponding to the music control can be a song name, an album name, a singer, a music tag (such as language, instrument, style, emotion, age, singer gender, ranking list, etc.), a playing application name, etc.

[0181] For example, the user intention instruction is "navigate to address A", the intention parameter corresponding to the user intention instruction can be "map navigation", and the slot parameter can be the destination address "address A" that needs to be navigated in the application.

[0182] For example, the user intention instruction is "set the air conditioner to 26 degrees", the intention parameter corresponding to the user intention instruction can be "air conditioner control", and the slot parameter can be the temperature "26 degrees".

[0183] For example, the user intention instruction is "use Huawei Music to play Mr. Liu's music A", the intention parameter corresponding to the user intention instruction can be "music control", and the slot parameter can be music A, Mr. Liu, and Huawei Music.

[0184] It can be understood that, in the embodiments of the present application, the execution device executes the intention instruction based on the intention parameter and the slot parameter, which can accurately execute the user intention instruction. Moreover, the accurate control of the application in another device through the voice assistant of one of the distributed collaborative devices can be realized.

[0185] Based on the above method, the accurate control of the application in another device through the voice assistant of one of the distributed collaborative devices can be realized, so that the voice interaction experience of the application in another device shared by one of the devices and the native application of another device on another device remains consistent, and the user experience is improved.

[0186] The following takes the first electronic device as a mobile phone, the second electronic device as a car machine, and the user waking up the voice assistant of the mobile phone as an example to describe the interaction method in the embodiments of the present application in detail. Figure 6 An interaction method in the embodiments of the present application is shown in the schematic diagram. As shown in Figure 6 The interaction method can include:

[0187] 601: The mobile phone and the car machine establish a connection.

[0188] It can be understood that the manner in which the mobile phone and the vehicle machine establish a connection can be any implementable connection manner such as Bluetooth, WIFI, or USB.

[0189] 602: The mobile phone detects an instruction of a user to wake up the voice assistant.

[0190] The manner in which the voice assistant of the mobile phone is woken up can be any implementable manner such as a manner in which the user presses a button of the automobile or a voice wake-up word corresponding to the mobile phone.

[0191] 603: The mobile phone sends data corresponding to a voice assistant wake-up state animation to the vehicle machine.

[0192] In the embodiments of the present application, when the user wakes up the voice assistant of the mobile phone, the mobile phone does not directly cast the current display interface to the vehicle machine for display, but only sends voice wake-up state animation data or voice wake-up state animation data and smart tips data to the vehicle machine, and the vehicle machine draws the voice wake-up state animation or the voice wake-up state animation and the smart tips on the original display interface. In this way, the interaction of the user with the interface of the vehicle machine is not interrupted, and the user experience is improved.

[0193] For example, when the user wakes up the voice assistant of the mobile phone by using the wake-up word "Xiaoyi", the mobile phone can send data corresponding to a wake-up state animation to the vehicle machine, as shown in FIG. 6B. Figure 5 The vehicle machine displays the wake-up state animation in the wake-up listening state and the smart tips "What day is today" and the like on the original vehicle machine HMI interface based on the data corresponding to the wake-up state animation.

[0194] 604: The vehicle machine draws the voice assistant wake-up state animation on the current display interface.

[0195] In some embodiments, the vehicle machine can display the voice assistant wake-up state animation on the current original interface of the vehicle machine based on the data corresponding to the voice assistant wake-up state animation. In some embodiments, the vehicle machine can display the smart tips in addition to the voice assistant wake-up state animation on the current original interface.

[0196] 605: The mobile phone sends voice interaction data of a voice interaction process to the vehicle machine.

[0197] In some embodiments, the voice interaction data of the voice interaction process can refer to voice interaction content and corresponding voice interaction state animation generated by the mobile phone in response to an interactive operation performed by the user with the voice assistant of the mobile phone. The interactive operation performed by the user with the voice assistant of the mobile phone can be an operation in which the user issues a voice intent instruction to the voice assistant of the mobile phone.

[0198] It can be understood that the car machine can receive the voice intent instruction of the user in the process of voice interaction with the voice assistant of the mobile phone through a microphone or the like, and send the voice intent instruction corresponding data to the mobile phone for ASR recognition, intent understanding and the like, and the mobile phone can send the voice interaction content data, i.e. the text data obtained by ASR recognition, the result data of intent understanding (multi-round user selection GUI), intelligent tips and the like, and voice interaction state animation data (such as listening animation, broadcast animation) to the car machine for display.

[0199] For example, after the mobile phone voice assistant is woken up, the car machine displays the original interface and the wake-up state animation (also known as the listening state animation), and when the user issues the voice intent instruction "navigate to the train station", the car machine can send the voice intent instruction corresponding data to the mobile phone. When the mobile phone receives the voice intent instruction data of the user, the voice intent instruction is subjected to ASR recognition to obtain the text corresponding to the voice intent instruction, and the text corresponding to the voice intent instruction is sent to the car machine. As shown in Figure 7 , the car machine can display the original interface, the listening state animation and the text "navigate to the train station".

[0200] In addition, the mobile phone can perform intent understanding on the voice intent instruction to obtain the intent understanding result, such as the multiple options "1. Shenzhen train station", "2. Shenzhen North Station" and the like after intent understanding, and send the intent understanding result and the corresponding broadcast state animation to the car machine for display. As shown in Figure 7 , the car machine can display the original interface of the car machine and the multiple options "1. Shenzhen train station", "2. Shenzhen North Station" after semantic understanding, and the broadcast state animation.

[0201] In the embodiments of the present application, the mobile phone can only send the voice interaction data to the car machine, and not send the current application interface data of the mobile phone to the car machine, so that the car machine can directly display the interaction content (i.e. the voice interaction interface) based on the voice interaction data on the original interface (i.e. the current display interface) of the car machine, without completely blocking the original interface of the car machine, thereby improving the user experience.

[0202] In some embodiments, the car machine can receive the voice intent instruction of the user in the process of voice interaction with the voice assistant of the mobile phone through a microphone or the like, and directly perform ASR recognition, intent understanding and the like, and the car machine can directly display the text data obtained by ASR recognition, the result data of intent understanding (multi-round user selection GUI), intelligent tips and the like voice interaction content, and voice interaction state animation data (such as listening animation, broadcast animation) on the original interface of the car machine for display.

[0203] 606: The car machine draws a voice interaction interface on the current display interface based on the voice interaction data.

[0204] 607: The mobile phone determines the user's intended command.

[0205] In some embodiments, such as Figure 8 As shown, users can select Shenzhen North Station by sending a voice command to their phone's voice assistant between options "1. Shenzhen Railway Station" and "2. Shenzhen North Station". The vehicle's infotainment system can receive the user's voice command via a microphone or other device and send the corresponding data to the phone. The phone determines the user's command to navigate to Shenzhen North Station based on this data. The system then determines the execution device using the method described in step 608. If the phone is determined to be the execution device, it sends navigation interface data during the execution of the command to the vehicle's infotainment system. The system can then draw and display the corresponding navigation interface based on this data.

[0206] In some embodiments, the vehicle's infotainment system can directly perform ASR (Automatic Speech Recognition) and intent understanding processes on the user's voice intent commands to determine the user's intent command, and directly display the ASR-recognized text, intent understanding results, and other voice interaction content. For example, the vehicle's infotainment system can determine the user's intent command as navigation to Shenzhen North Station based on the user's voice intent command of selecting Shenzhen North Station.

[0207] In some embodiments, the vehicle infotainment system can display the navigation interface in full-screen mode, for example, by switching the current interface of the vehicle infotainment system to the navigation interface.

[0208] In some embodiments, the navigation interface can also be displayed in a floating manner on the vehicle's infotainment system. For example, a floating window for the navigation application can be displayed on the current screen of the vehicle's infotainment system to show the corresponding navigation interface.

[0209] In some embodiments, the navigation interface can also be displayed in a split-screen manner, such as displaying the navigation interface and other application interfaces in different display areas of the vehicle's central control screen.

[0210] It should be noted that, in this embodiment of the application, the display method of the application interface drawn by the vehicle system based on the application interface data during the execution of the mobile phone can be any feasible display method such as full-screen display, floating window display, split-screen display, etc.

[0211] 608: When the mobile phone determines that it is an execution device, it executes the user's intended command.

[0212] It can be understood that in the embodiments of the present application, after the voice assistant of the mobile phone determines the user's intention instruction, the execution device of the intention instruction can be determined. In some embodiments, the device of the woken-up voice assistant can determine the device executing the user's intention according to the type of the user's intention, the running state of the mobile phone and the car machine, for example, whether the current mobile phone and car machine have an application capable of executing the intention in a running state, etc.

[0213] In some embodiments, the woken-up device can first determine whether the type of the intention instruction is a dedicated type, wherein the dedicated type is an intention type that can only be executed by the car machine or the mobile phone. The non-dedicated type is an intention type that can be executed by both the mobile phone and the car machine. For example, vehicle control type intentions, such as vehicle air conditioning, light control, etc., can only be executed by the car machine, and the vehicle control type intention is the dedicated type of intention corresponding to the car machine. For example, the intention instruction is to open the B service of the A application, and the car machine does not install the A application, but the mobile phone installs the A application, and the type of the intention instruction is the dedicated type of intention corresponding to the mobile phone. For example, the intention instruction is to perform navigation, which requires opening the navigation application, and both the car machine and the mobile phone install the navigation application and can execute the intention instruction, and the type of the intention instruction is a non-dedicated type or not a dedicated type.

[0214] When it is determined that the type of the intention instruction is a dedicated type, the device corresponding to the dedicated type is determined as the execution device. When it is determined that the type of the intention instruction is not a dedicated type, the execution device can be determined based on the running state of the device, such as whether there is an application running on the device that can execute the current intention instruction or whether there is an application displayed on the interface of the device that can execute the current intention instruction. If there is an application running on the device that can execute the current intention instruction or there is an application displayed on the interface of the device that can execute the current intention instruction, the device running the application that can execute the current intention instruction or the device displaying the interface of the application that can execute the current intention instruction is determined as the execution device.

[0215] For example, after the voice assistant of the mobile phone is woken up, the mobile phone receives a user instruction to turn on the vehicle air conditioner, which can only be executed by the car machine, which is the dedicated type corresponding to the car machine. Therefore, the device executing the user's intention is determined as the car machine, and the intention instruction is sent to the car machine.

[0216] When the phone receives a user instruction "navigate to address A", it is determined that the user's intention is navigation, which is not a dedicated intention. At this time, the phone can determine whether there is a device running a navigation application or whether the interface displayed by the device is an application that can execute the current intention instruction, for example, if the phone is running a navigation application and the car machine is not running a navigation application, the phone is used as the user intention execution device. Alternatively, the phone is running a navigation application in the background, but the navigation application interface is not displayed in the foreground, and the car machine foreground interface displays the navigation application interface, so the car machine can be preferentially selected as the intention instruction execution device.

[0217] When receiving a user instruction to play music B, it is determined that the user's intention is a music playing intention, which is not a dedicated type. At this time, the phone can determine whether there is a device running a music application or whether there is an application currently playing music, or whether the current audio application of the device is the top application or the focus application, etc. to determine the execution device. For example, if the phone is running a music application and the car machine is not running a music application, the phone is used as the user intention execution device.

[0218] It can be understood that the above examples of the present application are only illustrative, and the present application can include but is not limited to the above-mentioned determination of the execution device. In the embodiments of the present application, the execution device can be determined based on the type of intention instruction and the current running state of the device to determine the execution device that is more in line with the user's habits and improve the user experience.

[0219] In addition, in some embodiments, the execution device can parse the user intention instruction to obtain the intention parameters and slot parameters corresponding to the user intention instruction, and call the intention parameters and slot parameters corresponding to the user intention instruction to execute the user intention instruction. The intention parameter can be a characteristic parameter representing the type of user intention, such as map navigation, air conditioning control, music control, etc. The slot parameter can be a specific detail characteristic parameter corresponding to the characteristic parameter of the type of user intention, for example, the specific detail characteristic parameter corresponding to map navigation can be the navigation destination, the specific detail characteristic parameter corresponding to air conditioning control can be the specific temperature, and the specific detail characteristic parameter corresponding to music control can be the song name, album name, singer, music tag (such as language, instrument, style, emotion, age, singer gender, ranking list, etc.), playing application name, etc.

[0220] For example, the user intention instruction is "navigate to address A", the intention parameter corresponding to the user intention instruction is "map navigation", and the slot parameter is the destination address "address A" that needs to be navigated in the application.

[0221] For example, the user intention instruction is "set the air conditioner to 26 degrees". The intention parameter corresponding to the user intention instruction is "air conditioning control", and the slot parameter is the temperature "26 degrees".

[0222] For example, the user intention instruction is "play music A of Mr. Liu with Huawei Music", the corresponding intention parameter can be "music control", and the slot parameter can be music A, Mr. Liu and Huawei Music.

[0223] It can be understood that, in the embodiments of the present application, the execution device executes the intention instruction based on the intention parameter and the slot parameter, which can accurately execute the user intention instruction. Moreover, accurate control of the application function of another device through the voice assistant of one of the devices can be realized.

[0224] In some embodiments, the determination of the execution device can also be performed by the car machine.

[0225] 609: The mobile phone sends the video stream data obtained by encoding the application interface during the execution process to the car machine.

[0226] The mobile phone can project the interface during the execution of the intention instruction to the car machine for display. For example, the mobile phone can project the interface of the navigation application to the car machine for display during the execution of the navigation application.

[0227] In some embodiments, the mobile phone can encode the application interface into video stream data and send the video stream data to the car machine. After receiving the video stream data, the car machine can decode the video stream data, obtain decoded data, and display the corresponding application interface based on the decoded data.

[0228] In some embodiments, the mobile phone can also select whether to project the interface during the execution of the intention instruction to the car machine for display according to the user demand. For example, the mobile phone can store the intention types that need to project the interface during the execution of the intention instruction to the car machine for display and the intention types that do not need to project the interface during the execution of the intention instruction to the car machine for display. For example, the intention types that need to send the interface data during the execution can include navigation and the like. The intention types that do not need to send the interface data during the execution can include music playing and the like.

[0229] In some embodiments, the mobile phone can also issue an inquiry instruction through voice broadcast or other any implementable manner, so as to enable the user to select whether to need to project the interface during the execution of the intention instruction to the car machine for display. For example, the mobile phone can broadcast "whether to need to project to the car machine for display". When the user selects to need, the mobile phone can project the interface during the execution of the intention instruction to the car machine for display.

[0230] 610: The car machine decodes the video stream data and displays the corresponding interface based on the decoded data.

[0231] It can be understood that, in the embodiments of the present application, the mode of the car machine display interface can be set according to actual needs, for example, the full-screen display of the mobile phone screen interface can be set, or the split-screen display can be set, for example, the first screen area displays the original interface of the car machine, and the second screen area displays the screen interface of the mobile phone, and so on, so that the user can watch the interface of the car machine and the mobile phone, and the user experience is improved.

[0232] Based on the above scheme, when the user wakes up the mobile phone assistant, the mobile phone will not directly project the current display interface data to the car machine, but only send the voice interaction data of the voice interaction process to the car machine, and the car machine draws the interaction interface such as interaction animation and interaction text. In this way, the user's interaction with the car machine HMI is not interrupted, and the user experience is improved.

[0233] When the intent is executed, the mobile phone can project the corresponding application interface to the car machine display according to the user's needs, so that the user can watch and the user experience is improved. Moreover, the voice interaction data includes wake-up animation data, so that the car machine can obtain the entrance to wake up the voice assistant, and the user can control the application of the mobile phone through the car machine, that is, the voice assistant collaborative function can be used.

[0234] In addition, based on the above scheme, precise control of the application of another device can be realized by waking up the voice assistant of one of the devices, for example, the application of the mobile phone shared to the car machine can be kept consistent with the native application of the car machine in the car machine voice interaction experience, and the user experience is improved.

[0235] It can be understood that, in the embodiments of the present application Figure 6 The steps of the interaction method shown in the embodiments of the present application can include more or fewer steps than those described above, and although each step in the flowchart of the embodiments of the present application is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the order of the arrow. The execution of these steps has no strict order limitation, and they can be executed in other arbitrary order. Moreover, at least part of the steps in the figure can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or sub-steps or stages of other steps.

[0236] Based on the above software architecture of the mobile phone and the car machine, the transmission of the voice interaction process data (including voice assistant state data and voice interaction interface data) in the mobile phone of the present application and the reverse control process of the car machine are briefly introduced. As shown in Figure 9a The process can include:

[0237] (1) The voice assistant of the mobile phone sends voice assistant state data (or data corresponding to the voice image state dynamic effect GUI) and voice interaction interface data (or voice interaction GUI) to the voice assistant atomic service module of the mobile phone.

[0238] (2) The voice assistant atomic service module of the mobile phone sends the voice assistant state data and the voice interaction interface data to the voice assistant atomic service module of the vehicle machine.

[0239] (3) The voice assistant atomic service module of the vehicle machine sends the voice assistant state data and the voice interaction interface data to the voice assistant distributed collaborative UI module of the vehicle machine, and the voice assistant distributed collaborative UI module of the vehicle machine draws the voice assistant state dynamic effect and the voice interaction interface based on the voice assistant state data and the voice interaction interface data.

[0240] (4) The voice assistant distributed collaborative UI module of the vehicle machine detects the user's control of the voice interaction interface and sends the corresponding control information to the voice assistant atomic service module of the vehicle machine.

[0241] (5) The voice assistant atomic service module of the vehicle machine sends the corresponding control information to the voice assistant atomic service module of the mobile phone.

[0242] (6) The voice assistant atomic service module of the mobile phone sends the corresponding control information to the voice assistant of the mobile phone, and the voice assistant of the mobile phone performs the corresponding response.

[0243] The process of voice assistant intent collaborative distribution is briefly introduced as follows. Figure 9b As shown in the figure, the voice assistant of the mobile phone is woken up, and after the user voice instruction is understood and the user intent instruction is determined, the execution device of the intent instruction can be judged based on the context scene perception, that is, through the type of the intent instruction (whether it belongs to the exclusive type that can only be executed by the vehicle machine or the mobile phone), the current running state of the vehicle and the mobile phone (for example, whether there is an application currently running on the device that can execute the user intent instruction, etc.), etc. When the mobile phone voice assistant determines that the execution device is the vehicle machine, the intent instruction is sent to the voice intent collaborative distribution module of the mobile phone, and the voice intent collaborative distribution module of the mobile phone sends the intent instruction to the voice intent collaborative distribution module of the vehicle machine. The voice intent collaborative distribution module of the vehicle machine can call the corresponding application to execute the intent instruction. When the mobile phone voice assistant determines that the execution device is the mobile phone, the corresponding mobile phone application can be directly called to execute the intent instruction. Among them, the execution application can call the intent parameter and slot parameter corresponding to the user intent instruction to execute the user intent.

[0244] Similarly, when the voice assistant of the car machine is determined to execute the user's intention after being woken up, the execution device can still be determined based on the same manner as the mobile phone, and the intention instruction can be distributed to the corresponding execution device through the voice intention collaborative distribution module of the car machine and the voice intention collaborative distribution module of the mobile phone.

[0245] The following takes the first electronic device as a mobile phone and the second electronic device as a car machine, and takes the mobile phone's voice assistant being woken up as an example to illustrate the interaction method in the embodiments of the present application in combination with the software architecture of the mobile phone and the car machine. Figure 9c A schematic diagram of an interaction method in the embodiments of the present application is shown. As shown in Figure 9c The interaction method can include the following steps.

[0246] 901: The voice assistant of the mobile phone and the voice assistant of the car machine determine that the mobile phone and the car machine establish a connection.

[0247] It can be understood that the mobile phone and the car machine can establish a connection in any implementable manner, such as through Bluetooth, WIFI, or USB.

[0248] 902: The voice assistant of the mobile phone detects an instruction for waking up the voice assistant.

[0249] The mobile phone's voice assistant can be woken up in any implementable manner, such as through a car control button or using a corresponding voice wake-up word of the mobile phone.

[0250] 903: The voice assistant of the mobile phone sends the wake-up state animation data to the voice assistant atomic service module of the mobile phone.

[0251] 904: The voice assistant atomic service module of the mobile phone sends the wake-up state animation data to the voice assistant atomic service module of the car machine.

[0252] 905: The voice assistant atomic service module of the car machine sends the wake-up state animation data to the voice assistant distributed collaborative UI module of the car machine.

[0253] 906: The voice assistant distributed collaborative UI module of the car machine draws the wake-up state animation based on the wake-up state animation data.

[0254] It can be understood that if the car machine currently displays a first interface including first display content, the voice assistant distributed collaborative UI module of the car machine can draw the wake-up state animation (i.e., second display content) on the first display content of the car machine based on the wake-up state animation data when receiving the wake-up state animation data. At this time, the car machine displays a second interface, and the second interface includes the first display content and the second display content.

[0255] 907: The voice assistant of the mobile phone sends the data corresponding to the voice interaction content to the voice assistant atomic service module of the mobile phone.

[0256] It can be understood that in the embodiments of the present application, the voice interaction content can include a voice interaction GUI, for example, a voice prompt GUI and a voice image state dynamic effect GUI (for example, a voice listening state dynamic effect, a voice broadcast state dynamic effect), etc. Among them, the voice prompt GUI can include user ASR text display, intelligent tips, multi-round user selection GUI, etc., and the voice image state dynamic effect GUI can include a voice listening state dynamic effect, a voice broadcast state dynamic effect, etc.

[0257] 908: The voice assistant atomic service module of the mobile phone sends the data corresponding to the voice interaction content to the voice assistant atomic service module of the vehicle machine.

[0258] 909: The voice assistant atomic service module of the vehicle machine sends the data corresponding to the voice interaction content to the voice assistant distributed collaborative UI module of the vehicle machine.

[0259] 910: The distributed collaborative interface module of the vehicle machine draws a voice interaction interface based on the data corresponding to the voice interaction content.

[0260] It can be understood that the distributed collaborative UI module of the vehicle machine can draw a voice interaction interface on the current display interface of the vehicle machine based on the data corresponding to the voice interaction content.

[0261] 911: The voice assistant of the mobile phone determines the user's intention instruction.

[0262] 912: The voice assistant of the mobile phone judges whether the execution device of the intention instruction is the mobile phone. If yes, go to 917, if not, go to 913.

[0263] It can be understood that in the embodiments of the present application, the mobile phone voice assistant can determine the execution device of the intention instruction after determining the user's intention instruction. In some embodiments, the mobile phone voice assistant can determine the device for executing the user's intention according to the type of the user's intention, the running state of the mobile phone and the vehicle machine, for example, whether the current mobile phone and vehicle machine have an application capable of executing the intention in a running state, etc.

[0264] In some embodiments, the mobile phone voice assistant can first determine whether the intent instruction type is a proprietary type, which is an intent that can only be executed by the car machine or the mobile phone. When it is determined that the type of the intent instruction is a proprietary type, it is determined that the device corresponding to the proprietary type is the execution device. When it is determined that the type of the intent instruction is not a proprietary type, the execution device can be determined based on the device running state, such as whether there is an application running on the device that can execute the current intent instruction or whether the application displayed on the interface of the device is an application that can execute the current intent instruction. If there is an application running on the device that can execute the current intent instruction or the application displayed on the interface of the device is an application that can execute the current intent instruction, the device running the application that can execute the current intent instruction or the device displaying the interface of the application that can execute the current intent instruction is determined as the execution device.

[0265] 913: The voice assistant of the mobile phone sends the intent instruction to the voice intent collaborative distribution module of the mobile phone.

[0266] 914: The voice intent collaborative distribution module of the mobile phone sends the intent instruction to the voice intent collaborative distribution module of the car machine.

[0267] 915: The voice intent collaborative distribution module of the car machine sends the intent instruction to the corresponding application of the car machine.

[0268] 916: The corresponding application of the car machine executes the intent instruction.

[0269] 917: The voice assistant of the mobile phone sends the intent instruction to the corresponding application of the mobile phone.

[0270] 918: The corresponding application of the mobile phone executes the intent instruction.

[0271] 919: The corresponding application of the mobile phone sends the video stream data obtained by encoding the application interface during execution to the voice assistant atomic service module of the mobile phone.

[0272] In some embodiments, the mobile phone application can also select whether to send the interface data during execution to the mobile phone voice assistant atomic service module according to user needs. For example, the intent types that require sending interface data during execution and the intent types that do not require sending interface data during execution can be stored in the mobile phone, for example, the intent types that require sending interface data during execution can include navigation and the like. The intent types that do not require sending interface data during execution can include music playing and the like. When the current intent instruction of the user is an intent type that requires sending interface data during execution, the interface data during execution of the intent instruction is sent to the mobile phone voice assistant atomic service module of the mobile phone, so that the mobile phone voice assistant atomic service module of the mobile phone sends the interface data during execution to the voice assistant atomic service module of the car machine.

[0273] 920: The voice assistant atomic service module of the mobile phone sends the video stream data to the voice assistant atomic service module of the car machine.

[0274] 921: The voice assistant atomic service module of the car machine sends the video stream data to the distributed collaborative UI module of the car machine.

[0275] 922: The distributed collaborative UI module of the car machine decodes the video stream data and draws an interface based on the decoded data.

[0276] It can be understood that the distributed collaborative UI module of the car machine can draw a corresponding interface based on the interface data drawing, and the car machine can display the mobile phone screen projection interface.

[0277] It can be understood that in the embodiments of the present application, the way in which the car machine displays the interface can be set according to actual needs, for example, the mobile phone screen projection interface can be displayed in full screen, or it can also be displayed in split screen, for example, the first screen area displays the original interface of the car machine, and the second screen area displays the mobile phone screen projection interface, and so on. In this way, the user can watch the interfaces of the car machine and the mobile phone, and the user experience is improved.

[0278] It can be understood that in the embodiments of the present application Figure 9c The steps of the interaction method shown in the embodiments of the present application can include more or fewer steps than those described above, and although each step in the flowchart in the embodiments of the present application is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the order of the arrow. The execution of these steps has no strict order limitation, and they can be executed in other arbitrary order. Moreover, at least part of the steps in the figure can include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or other steps or stages.

[0279] In the following, taking the first electronic device as a mobile phone and the second electronic device as a car machine, and combining the software architecture of the mobile phone and the car machine, the voice assistant of the car machine is taken as an example to illustrate the interaction method in the embodiments of the present application. Figure 10 A schematic diagram of an interaction method in the embodiments of the present application is shown. As Figure 10 shown, the interaction method can include:

[0280] 1001: The voice assistant of the mobile phone and the voice assistant atomic service module of the car machine determine that the mobile phone and the car machine establish a connection.

[0281] 1002: The voice assistant of the car machine detects a user's instruction to wake up the voice assistant.

[0282] The way in which the vehicle machine voice assistant is awakened can be any implementable way, such as a way in which a user controls a button of the vehicle or a corresponding voice wake-up word of a mobile phone.

[0283] 1003: The vehicle machine voice assistant draws a voice interaction interface in a voice interaction process and a wake-up state effect.

[0284] 1004: The vehicle machine voice assistant determines a user intention instruction.

[0285] 1005: The vehicle machine voice assistant determines whether an execution device of the intention instruction is a mobile phone. If yes, go to 1006, and if no, go to 1014.

[0286] 1006: The vehicle machine voice assistant sends the intention instruction to a voice intention collaborative distribution module of the vehicle machine.

[0287] 1007: The voice intention collaborative distribution module of the vehicle machine sends the intention instruction to a voice intention collaborative distribution module of the mobile phone.

[0288] 1008: The voice intention collaborative distribution module of the mobile phone sends an execution instruction to a corresponding application of the mobile phone.

[0289] 1009: The corresponding application of the mobile phone executes the intention instruction.

[0290] 1010: The corresponding application of the mobile phone sends video stream data obtained by encoding an application interface in an execution process to a voice assistant atomization service module of the mobile phone.

[0291] 1011: The voice assistant atomization service module of the mobile phone sends the video stream data to a voice assistant atom service module of the vehicle machine.

[0292] 1012: The voice assistant atom service module of the vehicle machine sends the video stream data to a distributed collaborative UI module of the vehicle machine.

[0293] 1013: The distributed collaborative UI module of the vehicle machine decodes the video stream data and draws an interface based on the decoded data.

[0294] 1014: The vehicle machine voice assistant sends the intention instruction to a corresponding application of the vehicle machine.

[0295] 1015: The corresponding application of the vehicle machine executes the intention instruction.

[0296] It can be understood that the embodiments of the present application Figure 10The interactive method shown may include more or fewer steps than described above. Although the steps in the flowcharts of this application embodiment are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. There is no strict order restriction on the execution of these steps; they can be executed in any other order. Moreover, at least some of the steps in the figure may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. Their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.

[0297] The following is combined Figure 11a and Figure 11b The following provides further supplementary explanations of the voice assistant collaborative interaction method in the embodiments of this application.

[0298] Figure 11a This illustrates a specific deployment scheme for implementing voice-assisted communication capabilities, such as... Figure 11a As shown, the specific deployment scheme in this application embodiment involves deploying a voice assistant atomic service module and a voice assistant distributed collaborative UI (or distributed Voice-HMI) module on the vehicle-mounted system, and opening voice atomication capabilities on the corresponding mobile phone (e.g., deploying the voice assistant atomic service module on the mobile phone). Specifically, for the voice assistant atomic service module deployed on the vehicle-mounted system, the mobile phone opens the voice assistant state capability (i.e., it can send voice assistant state data to the vehicle-mounted system) and the mobile phone's voice assistant interaction GUI interface data capability (i.e., it can send the data corresponding to the voice interaction GUI to the vehicle-mounted system) to realize the voice assistant mutual assistance capability between the vehicle-mounted system and the mobile phone. For example, the voice assistant state capability and the voice assistant interaction GUI capability enable the vehicle-mounted system to receive voice assistant state data (e.g., idle state data, wake-up listening state data, and broadcast state data) and draw corresponding state animations (e.g., idle state animations, wake-up listening state animations, and broadcast state animations), and also enable the vehicle-mounted system to receive voice interaction interface data (e.g., ASR text data / smart tips data / multi-turn data and selection, etc.) and draw the interaction interface.

[0299] Figure 11b The diagram illustrates an interaction method based on the in-vehicle infotainment system's interface display, further demonstrating the collaborative interaction process between the mobile phone and the in-vehicle system. For example... Figure 11b As shown, the interaction methods may include:

[0300] 1101: The distributed collaborative UI module of the mobile phone and the distributed collaborative UI module of the vehicle system determine the connection between the mobile phone and the vehicle system.

[0301] 1102: The distributed collaborative UI module of the car machine detects that the user wakes up the phone voice assistant through the key.

[0302] In some embodiments, the user can wake up the phone voice assistant through the car side control key.

[0303] 1103: The distributed collaborative UI module of the car machine sends wake-up information to the phone voice assistant.

[0304] 1104: The voice assistant of the phone enters a voice wake-up state and starts to receive audio.

[0305] 1105: The voice assistant of the phone sends response audio stream data, voice wake-up state animation data, and intelligent voice tips data to the distributed UI module of the car machine.

[0306] 1106: The distributed UI module of the car machine controls the car machine to play a response, and displays voice wake-up state animation and intelligent tips.

[0307] 1107: The distributed collaborative UI module of the car machine receives audio and sends audio stream data to the phone voice assistant.

[0308] 1108: The voice assistant of the phone performs voice recognition based on the audio stream data.

[0309] 1109: The voice assistant of the phone sends the recognized ASR text data to the distributed collaborative UI module of the car machine.

[0310] 1110: The distributed collaborative UI module of the car machine controls the car machine to display the ASR text.

[0311] 1111: The voice assistant of the phone performs intent understanding and obtains audio stream data for voice broadcast.

[0312] 1112: The voice assistant of the phone sends the audio stream data for voice broadcast and voice broadcast state animation to the distributed collaborative UI module of the car machine.

[0313] 1113: The distributed collaborative UI module displays voice broadcast state animation and controls the car machine to perform voice broadcast.

[0314] 1114: The voice assistant of the phone obtains GUI data of a multi-round dialogue.

[0315] 1115: The voice assistant of the phone sends the GUI data of the multi-round dialogue to the distributed collaborative UI module of the car machine.

[0316] 1116: The distributed collaborative UI module of the car machine controls the car machine to display the GUI card of the multi-round dialogue.

[0317] 1117: The distributed collaborative UI module of the car machine acquires the intent instruction selected by the user.

[0318] 1118: The distributed collaborative UI module of the car machine sends the intent instruction selected by the user to the mobile phone.

[0319] 1119: The voice assistant of the mobile phone determines the execution device based on the intent instruction.

[0320] 1120: When the voice assistant of the mobile phone determines that the execution device is the mobile phone, the voice assistant sends the intent instruction to the corresponding application of the mobile phone.

[0321] 1121: When the voice assistant of the mobile phone determines that the execution device is the car machine, the voice assistant sends the intent instruction to the corresponding application of the car machine.

[0322] It can be understood that the voice assistant of the mobile phone sends the intent instruction to the voice intent collaborative distribution module of the mobile phone. The voice intent collaborative distribution module of the mobile phone sends the intent instruction to the voice intent collaborative distribution module of the car machine. The voice intent collaborative distribution module of the car machine sends the intent instruction to the application of the car machine.

[0323] 1122: The corresponding application of the mobile phone casts the interface in the process of executing the intent instruction to the car machine according to the actual needs.

[0324] In some embodiments, the mobile phone can also select whether to cast the interface in the process of executing the intent instruction to the car machine display according to the user needs. For example, the mobile phone can store the intent types that need to cast the interface in the process of executing the intent instruction to the car machine display and the intent types that do not need to cast the interface in the process of executing the intent instruction to the car machine display. For example, the intent types that need to send the interface data in the execution process can include navigation and the like. The intent types that do not need to send the interface data in the execution process can include music playing and the like.

[0325] 1123: The voice assistant of the mobile phone detects that the Voice Activity Detection (VAD) timeout time is greater than a set value, and exits the voice state.

[0326] 1124: The voice assistant of the mobile phone sends the exit voice state animation data to the car machine.

[0327] 1125: The distributed collaborative UI module of the car machine controls the car machine to display the exit voice state animation.

[0328] Based on the above scheme, when the user wakes up the mobile phone assistant, the mobile phone will not directly cast the current display interface data to the car machine, but only send the voice interaction data of the voice interaction process to the car machine, and the car machine draws the interaction interface such as the interaction animation and the interaction text. In this way, the user's interaction with the car machine HMI is not interrupted, and the user experience is improved.

[0329] When the intent is executed, the mobile phone can project the corresponding application interface to the vehicle display according to the user's demand, so as to facilitate the user to watch and improve the user experience. Moreover, the voice interaction data includes wake-up animation data, so that the vehicle can obtain the entrance of the wake-up voice assistant, and the user can control the application of the mobile phone through the vehicle, that is, the voice assistant collaborative function can be used.

[0330] In addition, based on the above scheme, precise control of the application of another device can be realized by waking up the voice assistant of one of the devices, for example, the application of the mobile phone shared to the vehicle can be consistent with the native application of the vehicle in the vehicle voice interaction experience, and the user experience can be improved.

[0331] It can be understood that the steps of the interaction method shown in the embodiments of the present application Figure 6 The steps of the interaction method shown in the embodiments of the present application

[0332] The embodiments disclosed in the present application can be implemented in hardware, software, firmware or a combination of these implementation methods. The embodiments of the present application can be implemented as computer programs or program codes executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memories and / or storage elements), at least one input device and at least one output device.

[0333] The program code can be applied to input instructions to execute the functions described in the present application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purpose of the present application, the processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC) or a microprocessor.

[0334] The program code can be implemented in a high-level programming language or an object-oriented programming language to communicate with the processing system. When necessary, the program code can also be implemented in assembly language or machine language. In fact, the mechanisms described in the present application are not limited to the scope of any specific programming language. In any case, the language can be a compiled language or an interpreted language.

[0335] In some cases, the disclosed embodiments can be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments can also be implemented as instructions carried by or stored on a transitory or non-transitory machine-readable (e.g., computer-readable) medium, which can be read and executed by one or more processors. For example, the instructions can be distributed over the network or by other computer readable media. Thus, a machine-readable medium can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including without limitation, floppy diskettes, optical disks, optical fiber, ROMs, RAMs, EPROMs, EEPROMs, magnetic or optical cards, flash memory, or tangible or other machine-readable media. Accordingly, a machine-readable medium includes any type of media mechanism that is suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).

[0336] In the drawings, some of the structural or methodological features can be shown in particular arrangements and / or orders. However, it should be understood that such particular arrangements and / or orders can not be required. Instead, in some embodiments, the features can be arranged differently than shown in the illustrative figures. Also, inclusion of a structural or methodological feature in a particular figure does not imply that the feature is required in all embodiments, and in some embodiments, the feature can not be included or can be combined with other features.

[0337] It should be noted that each unit / module mentioned in the embodiments of the devices in the present application is a logical unit / module, and in the physical world, one logical unit / module can be a physical unit / module, or a part of a physical unit / module, or a combination of multiple physical unit / modules, and the physical implementation of the logical unit / module itself is not the most important, and the combination of the functions implemented by the logical unit / module is the key to solving the technical problems proposed in the present application. In addition, in order to highlight the innovative part of the present application, the above-mentioned device embodiments of the present application do not introduce the units / modules that are not closely related to solving the technical problems proposed in the present application, which does not mean that the above-mentioned device embodiments do not have other units / modules.

[0338] It has to be noted that, in the description of the application, the terms "first", "second", etc. are used only for distinguishing between similar elements, and do not connote any order, sequence or priority. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0339] While the application has been illustrated and described in detail in the drawings and foregoing description, the same is to be considered as illustrative and not restrictive in character, it being understood that only the preferred embodiments have been shown and described and that all changes and modifications that come within the spirit of the application are desired to be protected.

Claims

1. An interaction method, characterized in that, The method comprises the following steps: A first electronic device and a second electronic device establish a connection, and the second electronic device displays a first interface, wherein the first interface comprises first display content corresponding to a human-computer interface of the second electronic device; The first electronic device detects a wake-up instruction of a voice assistant; The first electronic device sends first data corresponding to the wake-up instruction to the second electronic device, wherein the first data comprises data corresponding to a voice wake-up state effect; The second electronic device displays a second interface based on the first data and the first display content, wherein the second interface comprises the first display content and second display content corresponding to the first data.

2. The method of claim 1, wherein, In response to an interaction operation of a user with the first electronic device, the second electronic device displays a third interface, wherein the third interface comprises the first display content and third display content corresponding to the interaction operation.

3. The method of claim 2, wherein, The third display content comprises voice interaction content and a voice interaction state effect corresponding to the interaction operation.

4. The method of claim 3, wherein, The voice interaction content comprises text, intelligent prompts, and intent understanding content corresponding to a user voice instruction; The voice interaction state effect comprises a listening effect or a broadcast effect.

5. The method according to claim 3 or 4, characterized in that, The second electronic device displays the third interface in response to the interaction operation of the user with the first electronic device comprises: The first electronic device sends data corresponding to the voice interaction content and data corresponding to the voice interaction state effect to the second electronic device in response to the interaction operation of the user with the first electronic device; The second electronic device displays the third interface based on the data corresponding to the voice interaction content, the data corresponding to the voice interaction state effect, and the first display content.

6. The method according to any one of claims 1 to 4, characterized in that, Further comprising: The first electronic device detects a user intent instruction, determines an execution device of the intent instruction based on a type of the intent instruction and a current running state of the first electronic device and the second electronic device; When it is determined that the execution device of the intent instruction is the second electronic device, the first electronic device sends the intent instruction to the second electronic device.

7. The method of claim 6, wherein, Further comprising: The second electronic device acquires intent parameters and slot parameters based on the intent instruction, and executes the intent instruction based on the intent parameters and the slot parameters.

8. The method of claim 6, wherein, The determination of the execution device of the intent instruction based on the type of the intent instruction and the current running state of the first electronic device and the second electronic device comprises: When the type of the intent instruction is a specific type, the execution device of the intent instruction is determined to be an execution device corresponding to the specific type; When the type of the intent instruction is not a specific type, the execution device of the intent instruction is determined based on the current running state of the first electronic device and the second electronic device.

9. The method of claim 8, wherein, The determination of the execution device of the intent instruction based on the current running state of the first electronic device and the second electronic device comprises: The device running the first application or the device displaying the interface of the first application among the first electronic device and the second electronic device is the execution device of the intent instruction, and the first application is an application capable of executing the intent instruction.

10. The method according to claim 8 or 9, characterized in that, When the execution device of the intent instruction is the first electronic device, and the type of the intent instruction is a preset type not requiring a screen projection display execution process, the first electronic device does not send data corresponding to a display interface in the execution process of the intent instruction to the second electronic device.

11. The method according to claim 8 or 9, characterized in that, When the execution device of the intent instruction is the first electronic device, and the type of the intent instruction is a preset type requiring a screen projection display execution process, the first electronic device sends first data corresponding to a display interface in the execution process of the intent instruction to the second electronic device. The first electronic device displays a fourth interface based on the first data corresponding to the display interface.

12. The method of claim 11, wherein, The fourth interface includes the first display content and fourth display content corresponding to the first data.

13. The method of claim 12, wherein, The first display content and the fourth display content are displayed in different regions of a screen of the first electronic device.

14. The method of claim 11, wherein, The fourth interface includes fourth display content corresponding to the first data.

15. The method according to any one of claims 1-4, 7-9 and 12-14, characterized in that, The first data does not include interface data corresponding to a current application interface of the first electronic device.

16. The method according to any one of claims 1-4, 7-9 and 12-14, characterized in that, The first electronic device is a mobile phone, and the second electronic device is a car machine.

17. An electronic device, comprising: Comprise: a memory for storing a computer program, the computer program comprising program instructions; a processor for executing the program instructions to cause the electronic device to perform the interaction method of any one of claims 1-16.

18. A computer readable storage medium, characterized in that, The computer readable storage medium stores a computer program, the computer program comprising program instructions, the program instructions being run by an electronic device to cause the electronic device to perform the interaction method of any one of claims 1-16.

Citation Information

Patent Citations

  • Human-computer interaction method, electronic equipment and system

    CN114255745A