Function execution method and device

By implementing the function execution method in smart vehicles, the problems of poor privacy and high resource utilization are solved. By locally processing audio streams and uploading audio streams with only certain roads, the effects of privacy protection and resource conservation are achieved.

CN120183393APending Publication Date: 2025-06-20HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311759635.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-18
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The awake-free ability of smart vehicles has problems such as poor privacy and high resource utilization.

Method used

By implementing a function execution method in an intelligent vehicle, local processing is performed from multiple receiving audio streams during the first round of dialogue to determine the first way; in the second round and later conversations, only the first audio stream is uploaded to the server to reduce the network bandwidth and cloud-side computing resources of the multiple audio streams to the cloud simultaneously.

Benefits of technology

It improves the privacy and resource usage of smart vehicles without wake-up capabilities, protects user privacy, reduces network bandwidth and cloud-side computing resources, and has good expansion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183393A_ABST
    Figure CN120183393A_ABST
Patent Text Reader

Abstract

The invention discloses a function execution method and device. The method comprises the following steps: receiving a first audio stream from multiple paths; determining a first path according to the first audio stream, and executing a first function corresponding to the first path of the first audio stream; receiving a second audio stream from the multiple paths, the second audio stream including a third audio stream from the first path; sending the third audio stream to a server; and receiving a first message sent by the server, and executing a second function corresponding to the third audio stream according to the first message, so that the problems that the privacy of the wakeup-free capability of the intelligent vehicle is poor and the resource occupation is relatively high can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a function execution method and device. Background Art

[0002] With the development of technology, manufacturers have replicated the hands-free wake-up ability of the main driver in intelligent vehicles to the whole vehicle, laying the foundation for free interaction and further forming a core competitiveness. Hands-free wake-up does not mean that the user actively wakes up the voice assistant application of the vehicle with a wake-up word, but extracts high-frequency corpus as a trigger word, which requires real-time audio stream end-cloud simultaneous transmission. There is privacy anxiety for users; and the resource occupancy of multi-channel audio stream end-cloud simultaneous transmission is relatively high. Therefore, the hands-free wake-up ability of intelligent vehicles has poor privacy and high resource occupancy. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a function execution method and device, which can improve the problems of poor privacy and high resource occupancy of the hands-free wake-up ability of intelligent vehicles.

[0004] In a first aspect, an embodiment of the present invention provides a function execution method. The method includes: during the first-round conversation with the user, the first device receives multiple first audio streams, determines a first path from the multiple first audio streams, and executes a first function corresponding to the first audio stream of the first path; during the second-round and subsequent conversations with the user, the first device receives multiple second audio streams, where the second audio streams include third audio streams from the first path, only uploads the third audio streams to the server and receives a first message returned by the server, and executes a second function corresponding to the third audio streams according to the first message. Compared with uploading all audio streams to the server (also known as "going to the cloud"), the present application protects the privacy of users, reduces the network bandwidth and cloud-side computing resources occupied by simultaneous uploading of multiple audio streams; with the increase in the number of screens and seats of the first device, there is good scalability.

[0005] In combination with the first aspect, in some implementation manners of the first aspect, the first device determines the first path according to multiple first audio streams, including: merging the multiple first audio streams to obtain a first merged audio stream; obtaining a first merged text corresponding to the first merged audio stream according to the first merged audio stream; and determining the first path according to the first merged text. Before processing the multiple first audio streams, the first device of the present application merges the multiple first audio streams into one first merged audio stream. Compared with simultaneously identifying multiple first audio streams, the CPU, NPU, and memory resources occupied by the identification engine of the first device are reduced. With the increase in the number of screens and seats of the first device, reducing resource occupancy can bring good scalability.

[0006] In combination with the first aspect, in certain implementations of the first aspect, the first device determines a first path based on a first merged text, including: determining a first text corresponding to a first audio stream according to the first merged text; if the first text includes a target field, determining the path where the first audio stream corresponding to the first text is located as the first path. In the first device of the present application, it is determined whether the path where the first audio stream is located is the first path based on whether the first text corresponding to the first audio stream includes a target field. If the first text includes a target field, the path where the first audio stream corresponding to the first text is located is the first path.

[0007] In combination with the first aspect, in certain implementations of the first aspect, the first device executes a first function corresponding to the first audio stream of the first path, including: determining a first intent according to the first text corresponding to the first audio stream; obtaining a first instruction according to the first intent; and executing the first function according to the first instruction. In the first device of the present application, the audio stream can be converted into text, the text can be converted into intent, and the intent can be converted into an instruction through an identification engine. Then, the first device executes the first function according to the first instruction, so that the audio stream can be locally processed to execute the function corresponding to the audio stream.

[0008] In combination with the first aspect, in certain implementations of the first aspect, the first function includes playing a first dialogue corresponding to the first audio stream and / or enabling a first corresponding function corresponding to the first audio stream. In the present application, the first device can communicate with the user and / or enable a corresponding function according to what the user says.

[0009] In combination with the first aspect, in certain implementations of the first aspect, the first message includes information about a second function. In the present application, the first message returned by the server to the first device includes a second function, so that the first device can execute the second function according to the first message.

[0010] In combination with the first aspect, in certain implementations of the first aspect, the second function includes playing a second dialogue corresponding to the third audio stream and / or enabling a second corresponding function corresponding to the third audio stream. In the present application, the first device uploads the audio stream to the cloud and can execute the function corresponding to the audio stream by receiving the message returned by the server, that is, communicate with the user and / or enable a corresponding function according to what the user says.

[0011] In combination with the first aspect, in certain implementations of the first aspect, during the first-round dialogue between the first device and the user, after receiving the first audio stream from multiple paths, the first device also determines a second path according to the first audio stream and executes a first function corresponding to the first audio stream of the second path. Therefore, when multiple people speak simultaneously, the first device can execute the functions corresponding to the audio streams of multiple users simultaneously.

[0012] In combination with the first aspect, in some implementations of the first aspect, the first device determines a second path based on a first audio stream, including: if the first text corresponding to the first audio stream does not include a target field, determining the path where the first audio stream corresponding to the first text is located as the second path. In this application, only the path where the first audio stream corresponding to the first text including the target field is located is determined as the first path, and the path where the first audio stream corresponding to the first text not including the target field is located is determined as the second path, so that there is only one first path in the first device at the same time, ensuring that only one path goes to the cloud during the second and subsequent conversations with the user, protecting the user's privacy, reducing the network bandwidth and cloud-side computing resources occupied by multiple audio streams going to the cloud at the same time; with the increase in the number of screens and seats of the first device, there is good scalability.

[0013] In combination with the first aspect, in some implementations of the first aspect, during the second and subsequent conversations between the first device and the user, after receiving multiple second audio streams, the first device also executes a third function corresponding to the second audio stream according to the second audio stream. Therefore, when multiple people speak at the same time, the first device can execute the functions corresponding to the audio streams of multiple users simultaneously. At the same time, in this application, the second audio stream of the second path does not go to the cloud and is only processed locally to reduce false triggering. In combination with the first aspect, in some implementations of the first aspect, the first device executes a third function corresponding to the second audio stream according to the second audio stream, including: merging multiple second audio streams to obtain a second merged audio stream; obtaining a second merged text corresponding to the second merged audio stream according to the second merged audio stream; determining a second text corresponding to the second audio stream according to the second merged text; determining a second intention according to the second text; obtaining a third instruction according to the second intention; and executing a third function according to the third instruction. Before processing multiple second audio streams, the first device in this application merges multiple second audio streams into one second merged audio stream, which reduces the CPU, NPU, and memory resources occupied by the recognition engine of the first device during recognition compared to recognizing multiple second audio streams simultaneously. With the increase in the number of screens and seats of the first device, reducing resource occupancy can bring good scalability. The first device in this application can convert an audio stream into text, convert the text into an intention, convert the intention into an instruction through the recognition engine, and then the first device executes a first function according to the first instruction, so that the function corresponding to the audio stream can be executed by local processing of the audio stream.

[0014] In combination with the first aspect, in some implementations of the first aspect, the third function includes playing a third conversation corresponding to the second audio stream and / or enabling a third corresponding function corresponding to the second audio stream. In this application, the first device can converse with the user and / or enable a corresponding function according to what the user says.

[0015] In combination with the first aspect, in certain implementations of the first aspect, the first device includes a first application; the first device receives a first audio stream and a second audio stream from multiple sources through the first application. In this application, the first device is installed with the first application and receives audio streams from multiple sources through the first application.

[0016] In combination with the first aspect, in certain implementations of the first aspect, the first device includes a screen, namely the first screen, and a first application, and the first application is installed on the first screen; or

[0017] The first device includes multiple screens and a first application, namely the first screen and at least one second screen. The first application is installed on the first screen, and the first device moves the display content of the first screen to the second screen. In this application, when the first device has multiple screens and only one first application is installed, the display interfaces of the first application can be made to appear on all screens by means of screen moving, enabling users to view the first application through the nearby screens and enhancing the user experience.

[0018] In combination with the first aspect, in certain implementations of the first aspect, the first device includes multiple first applications and multiple screens, and the first applications are installed on the screens, with one first application corresponding to one screen; the first electronic device receives a first audio stream and a second audio stream from multiple sources through the multiple first applications, and the first application is used to receive the first audio stream and the second audio stream from at least one source. In this application, when the first electronic device has multiple screens and each screen is installed with a first application, each first application receives audio streams from at least one source, enabling each screen to receive the audio streams in its vicinity.

[0019] In combination with the first aspect, in certain implementations of the first aspect, the multiple screens communicate with each other. In this application, when the first device includes multiple screens, the multiple screens communicate with each other, thereby ensuring that only one path is the first path at the same time.

[0020] In a second aspect, an embodiment of the present invention provides a device, including a processor and a memory. Among them, the memory is used to store a program, and when the processor runs the program, the device is caused to execute the steps of the method as described above.

[0021] In a third aspect, an embodiment of the present invention provides a readable storage medium, and the readable storage medium stores a program, and when the program is run by a device, the device is caused to execute the method as described above.

[0022] In a fourth aspect, an embodiment of the present invention provides a computer program product, and the computer program product contains instructions. When the computer program product runs on a computer or any at least one processor, the computer is caused to execute the functions / steps in the method as described above.

[0023] In the technical solution of the function execution method and device provided by the embodiments of the present invention, the method includes: receiving multiple paths of first audio streams; determining a first path according to the multiple paths of first audio streams, and executing a first function corresponding to at least one path of the first audio streams; receiving the multiple paths of second audio streams; sending the second audio stream of the first path to a server, receiving a first message sent by the server, and executing a second function corresponding to the second audio stream of the first path according to the first message; and executing a third function corresponding to at least one path of the second audio streams according to the multiple paths of second audio streams, which can improve the problems of poor privacy and high resource occupancy of the wake-up-free ability of intelligent vehicles. Description of the Drawings

[0024] Figure 1 It is a schematic structural diagram of a device provided by an embodiment of the present invention;

[0025] Figure 2 It is a software structure block diagram of device 100 according to an embodiment of the present invention;

[0026] Figure 3 It is a schematic diagram of a function execution system;

[0027] Figure 4 It is Figure 3 A schematic diagram of an application scenario of the function execution system in

[0028] Figure 5 It is an architecture diagram of a function execution system provided by an embodiment of the present invention;

[0029] Figure 6 It is another architecture diagram of a function execution system provided by an embodiment of the present invention;

[0030] Figure 7 It is a signaling interaction diagram of a function execution method provided by an embodiment of the present invention;

[0031] Figure 8 It is a schematic diagram of an application scenario of an embodiment of the present invention;

[0032] Figure 9 It is a flowchart of a function execution method provided by an embodiment of the present invention;

[0033] Figure 10 It is a schematic structural diagram of a first device provided by an embodiment of the present invention. Detailed Embodiments

[0034] To better understand the technical solution of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0035] It should be clear that the described embodiments are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative work fall within the scope of protection of the present invention.

[0036] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0037] It should be understood that the term " / and" used herein is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A / and B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally indicates that the associated objects before and after are in an "or" relationship.

[0038] Figure 1 A schematic structural diagram of device 100 is shown.

[0039] Device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a sensor module 180, a key 190, a display screen 194, etc. Among them, the sensor module 180 may include a touch sensor 180K, etc.

[0040] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on device 100. In other embodiments of the present application, device 100 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0041] Processor 110 may include one or more processing units. For example, processor 110 may include an application processor, a modem processor, a graphics processor, an image signal processor, a controller, a video codec, a digital signal processor, a baseband processor, and / or a neural network processor, etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0042] The controller can generate operation control signals according to the instruction operation code and timing signals to complete the control of instruction fetching and execution.

[0043] A memory can also be set in the processor 110 to store instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the said memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0044] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0045] It can be understood that the interface connection relationship between the modules illustrated in the embodiments of the present invention is only for illustrative purposes and does not constitute a structural limitation on the device 100. In other embodiments of the present application, the device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0046] The charging management module 140 is used to receive charging input from a charger. Among them, the charger can be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 140 can receive the charging input of the wired charger through the USB interface 130. In some embodiments of wireless charging, the charging management module 140 can receive the wireless charging input through the wireless charging coil of the device 100. While charging the battery 142, the charging management module 140 can also supply power to the device through the power management module 141.

[0047] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the display screen 194, the wireless communication module 160, etc. The power management module 141 can also be used to monitor parameters such as the battery capacity, the number of battery cycles, and the battery health status (leakage, impedance). In some other embodiments, the power management module 141 can also be provided in the processor 110. In some other embodiments, the power management module 141 and the charging management module 140 can also be provided in the same device.

[0048] The wireless communication function of the device 100 can be implemented by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.

[0049] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 1 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0050] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the device 100. The mobile communication module 150 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves by the antenna 1, filter and amplify the received electromagnetic waves, and then transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signals modulated by the modulation and demodulation processor and convert them into electromagnetic waves through the antenna 1 for radiation. In some embodiments, at least some functional modules of the mobile communication module 150 can be provided in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 can be provided in the same device.

[0051] The wireless communication module 160 may provide solutions for wireless communications applied to the device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The wireless communication module 160 may be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 may also receive signals to be sent from the processor 110, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through the antenna 2 for radiation.

[0052] In some embodiments, the antenna 1 of the device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, enabling the device 100 to communicate with the network and other devices through wireless communication technologies.

[0053] The device 100 implements the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.

[0054] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some embodiments, the device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1.

[0055] The device 100 implements the shooting function through the ISP, the camera, the video codec, the GPU, the display screen 194, and the application processor, etc.

[0056] The external memory interface 120 may be used to connect to an external memory card, such as a MicroSD card, to expand the storage capacity of the device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, music, video, and other files are saved in the external memory card.

[0057] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc. The data storage area can store data created during the use of the device 100, etc. In addition, the internal memory 121 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, universal flash memory, etc. The processor 110 executes various functional applications and data processing of the device 100 by running the instructions stored in the internal memory 121 and / or the instructions stored in the memory provided in the processor.

[0058] The device 100 can implement audio functions through the audio module 170, speaker 170A, receiver 170B, microphone 170C, and application processor, etc. For example, music playback, recording, etc.

[0059] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert analog audio input into a digital audio signal. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.

[0060] The speaker 170A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The device 100 can listen to music or hands-free calls through the speaker 170A.

[0061] The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the device 100 answers a call or a voice message, the voice can be listened to by placing the receiver 170B close to the human ear.

[0062] The microphone 170C, also known as the "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak by bringing the mouth close to the microphone 170C to input the sound signal into the microphone 170C. The device 100 can be provided with at least one microphone 170C. In some other embodiments, the device 100 can be provided with two microphones 170C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the device 100 can also be provided with three, four or more microphones 170C to implement functions such as collecting sound signals, noise reduction, identifying the sound source, and implementing a directional recording function.

[0063] The button 190 includes a power-on button, volume buttons, etc. The button 190 can be a mechanical button or a touch button. The device 100 can receive button inputs and generate key signal inputs related to the user settings and function controls of the device 100.

[0064] The software system of the device 100 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of the present invention, taking the Android system with a layered architecture as an example, the software structure of the device 100 is exemplarily described.

[0065] Figure 2 It is a software structure block diagram of the device 100 in the embodiments of the present invention.

[0066] The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, namely the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0067] The application layer may include a series of application packages.

[0068] As Figure 2 shown, the application packages may include applications such as a calendar, a call, a map, a navigation, a WLAN, a Bluetooth, music, a video, a short message, etc.

[0069] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.

[0070] As Figure 2 shown, the application framework layer may include a window manager, a content provider, a view system, a telephone manager, a resource manager, a notification manager, etc.

[0071] The window manager is used to manage window programs. The window manager can obtain the size of the display screen, determine whether there is a status bar, lock the screen, capture the screen, etc.

[0072] The content provider is used to store and obtain data and enable these data to be accessed by applications. The data may include videos, images, audio, dialed and received calls, browsing history and bookmarks, a phone book, etc.

[0073] The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a text message notification icon can include a view for displaying text and a view for displaying pictures.

[0074] The phone manager is used to provide the communication functions of device 100. For example, the management of call status (including answering, hanging up, etc.).

[0075] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, and so on.

[0076] The notification manager enables applications to display notification information in the status bar. It can be used to convey notification-type messages, which can automatically disappear after a short stay without user interaction. For example, the notification manager is used to inform that the download is complete, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as the notification of a background-running application, or a notification that appears in the form of a dialogue window on the screen. For example, it can prompt text information in the status bar, emit a prompt sound, vibrate the device, blink the indicator light, etc.

[0077] Android Runtime includes core libraries and a virtual machine. Android runtime is responsible for the scheduling and management of the Android system.

[0078] The core libraries contain two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core libraries of Android.

[0079] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.

[0080] The system libraries can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing libraries (such as: OpenGL ES), 2D graphics engines (such as: SGL), etc.

[0081] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications.

[0082] The media library supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0083] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.

[0084] The 2D graphics engine is a drawing engine for 2D drawing.

[0085] The kernel layer is the layer between hardware and software. The kernel layer at least includes a display driver, an audio driver, and a sensor driver.

[0086] With the gradual development of intelligent cockpits, each manufacturer is deploying more screens in the intelligent cockpit to support richer audio-visual and entertainment capabilities, aiming to build the intelligent cockpit into a "mobile home".

[0087] As a crucial application in the intelligent cockpit, the voice assistant application can greatly simplify various operations in the intelligent cockpit and enhance the user experience. Among them, the ability of the voice assistant application to wake up without a wake-up word for the driver simplifies the method for the driver to use the voice assistant application and greatly improves the driver's voice service experience.

[0088] With the development of technology, manufacturers have replicated the driver's hands-free wake-up ability to the entire vehicle, laying the foundation for free interaction and further forming a core competitiveness. Some manufacturers already offer similar capabilities, with an experimental switch for users to control whether to enable this function.

[0089] Figure 3 It is a schematic diagram of a function execution system. The function execution system includes a first device and a server. The first device and the server are wirelessly connected. Exemplarily, the first device includes an intelligent vehicle; the server includes a voice cloud server. The first device includes multiple microphones and multiple screens. The multiple screens include a center control screen and at least one other screen. The center control screen is installed with a first application and a side recognition engine, and at least one other screen is not installed with the first application and the side recognition engine. Cross-screen interaction occurs between the center control screen and at least one other screen. Exemplarily, the first application is a voice assistant application. In the first device, the multiple screens share the same operating system (OS), the same user space, and a single first application.

[0090] Figure 3 Just an example, such as Figure 3As shown, the first device includes two screens, namely the central control screen and the co-pilot screen. The central control screen is installed with a first application and an identification engine. The co-pilot screen is installed with a desktop application. The first application and the desktop application interact across screens. The first device also includes 4 microphones, namely the driver microphone, the co-pilot microphone, the left rear microphone, and the right rear microphone. As Figure 4 shown, the first application receives 4 audio streams sent by the 4 microphones, and then sends the 4 audio streams to the identification engine for identification, and at the same time sends the 4 audio streams to the server, that is, uploads the 4 audio streams to the cloud; the first application receives the identification results of the 4 audio streams by the identification engine and executes corresponding functions for the identification results, and also receives the processing results of the 4 audio streams by the server and executes corresponding functions for the processing results. Therefore, Figure 3 in the function execution system shown, the multi-channel concurrent audio streams are realized by the first application on the central control screen for simultaneous end-cloud transmission.

[0091] Figure 4 is Figure 3 a schematic diagram of an application scenario of the function execution system in Figure 4 shown. As shown, the vehicle includes 4 seats. When the driver user, the co-pilot user, and the left rear user all speak to the voice assistant application using the wake-up-free statement, the voice assistant application can respond to the driver user, the co-pilot user, and the left rear user. As Figure 4 shown, the driver user says "Close the window", and the voice assistant application says "Okay" to the driver user and at the same time controls the vehicle to close the window; the co-pilot user says "Turn on the air conditioner", and the voice assistant application says "Okay" to the co-pilot user and at the same time controls the vehicle to turn on the air conditioner; the left rear user says "I want to watch Peppa Pig", and the voice assistant application says "No problem" to the left rear user, and at the same time controls the vehicle to play Peppa Pig. However, the voice assistant application sends the audio streams of the driver user, the co-pilot user, and the left rear user to the identification engine for identification separately, and the identification engine needs to identify 3 times; the voice assistant application uploads the audio streams of the driver user, the co-pilot user, and the left rear user to the server separately, and the voice assistant needs to upload 3 times; and, whether it is the first round of conversation between the user and the voice assistant application or the subsequent multiple rounds of conversation, the voice assistant application sends the audio stream of each user to the identification engine and the server separately.

[0092] There are the following 4 disadvantages in realizing the simultaneous end-cloud transmission of multi-channel concurrent audio streams by the first application on the central control screen:

[0093] 1. For users, there is privacy anxiety in the real-time simultaneous end-cloud transmission of audio streams;

[0094] 2. When there are four audio streams for simultaneous end-cloud transmission, compared with the simultaneous end-cloud transmission of one audio stream, the resource occupancy of the first device and the server quadruples;

[0095] 3. When the number of vehicle seats increases from 4 to 6 or even 7, it does not have good scalability.

[0096] 4. As long as one user wakes up the first application with a wake-up word, other users can interact with the first application at will, increasing the probability of accidental intrusion.

[0097] Based on the above technical problems, an embodiment of the present invention provides a function execution system. Figure 5 It is an architecture diagram of a function execution system provided by an embodiment of the present invention. The function execution system includes a first device 200 and a server 300. The first device 200 and the server 300 are wirelessly connected. Among them, for the hardware structure and software structure of the device provided by the embodiment of the present invention, reference can be made to Figure 1 and Figure 2 the relevant descriptions of the device 100 therein.

[0098] Exemplarily, the first device 200 includes an intelligent vehicle.

[0099] Exemplarily, the server 300 includes a voice cloud server.

[0100] Exemplarily, the first device 200 includes a first operating system 210, a central control user space 220, and a headless user space 240.

[0101] The first operating system 210 includes multiple microphones. Exemplarily, as Figure 5 shown, the first operating system 210 includes a driver's seat microphone 211, a passenger seat microphone 212, a left rear microphone 213, and a right rear microphone 214.

[0102] Exemplarily, the position of the driver's seat microphone 211 is relatively close to the driver's seat, mainly used for collecting the voice of the driver's seat user and converting the voice into an audio stream.

[0103] Exemplarily, the position of the passenger seat microphone 212 is relatively close to the passenger seat, mainly used for collecting the voice of the passenger seat user and converting the voice into an audio stream.

[0104] Exemplarily, the position of the left rear microphone 213 is relatively close to the seat behind the driver's seat, mainly used for collecting the voice of the left rear user and converting the voice into an audio stream. The left rear user is the user sitting in the seat behind the driver's seat.

[0105] Exemplarily, the position of the right rear microphone 214 is relatively close to the seat behind the passenger seat, mainly used for collecting the voice of the right rear user and converting the voice into an audio stream. The right rear user is the user sitting in the seat behind the passenger seat.

[0106] As Figure 5As shown, the central control user space 220 includes a first screen 221, and the first screen 221 is the central control screen of the first device 200. A first application 222 is installed on the first screen 221. For example, the first application 222 is a voice assistant application.

[0107] The first application 222 is used to receive audio streams sent by at least one microphone. It should be understood that the first application 222 can receive audio streams from at least one path, and one path can be understood as one microphone. Exemplarily, as Figure 5 shown, the first application 222 can receive three audio streams sent by the driver's seat microphone 211, the left rear microphone 213, and the right rear microphone 214.

[0108] Optionally, the first device 200 further includes at least one other user space. The other user space includes a second screen, and the second screen installs the first application.

[0109] Exemplarily, the other user space can be the co-pilot user space, the rear row user space, the left rear user space, or the right rear user space, etc.

[0110] Exemplarily, as Figure 5 shown, the other user space is the co-pilot user space 230. The co-pilot user space 230 includes a second screen 231. The second screen 231 is the co-pilot screen of the first device 200. A first application 232 is installed on the second screen 231. For example, the first application 232 is a voice assistant application. The first application 232 and the first application 222 are the same application, but are installed on different screens.

[0111] The first application 232 is used to receive audio streams sent by at least one microphone. It should be understood that the first application 232 can receive audio streams from at least one path, and one path can be understood as one microphone. Exemplarily, as Figure 5 shown, the first application 232 can receive one audio stream sent by the co-pilot microphone 212.

[0112] It should be understood that the second screen 231 where the first application 232 is located belongs to the co-pilot user space 230. Therefore, the second screen 231 is closest to the co-pilot user. Therefore, the first application 232 can receive one audio stream sent by the co-pilot microphone 212. If the first device further includes a left rear user space, the first application in the second screen of the left rear user space can receive one audio stream sent by the left rear microphone 213.

[0113] The first device 200 includes multiple first applications and multiple screens. The first applications are installed on the screens, and the first applications and the screens are in one-to-one correspondence. The multiple screens communicate with each other. Figure 5 This is just an example. The multiple screens in the first device 200 are under the same OS and different user spaces.

[0114] As shown Figure 5 in the figure, the headless user space 240 includes a merging module 241 and an identification engine 242.

[0115] The merging module 241 is used to receive the audio streams sent by the first application 222 and the first application 232, merge the above multiple audio streams, and then send them to the identification engine 242.

[0116] The identification engine 242 is used to identify the merged multiple audio streams to obtain the text corresponding to the audio streams, and return the text corresponding to the audio streams to the first application. Specifically, if the audio stream is sent from the first application 222 to the identification engine 242 of the first application 222, the identification engine 242 returns the text corresponding to the audio stream to the first application 222; if the audio stream is sent from the first application 232 to the identification engine 242, the identification engine 242 returns the text corresponding to the audio stream to the first application 232.

[0117] The identification engine 242 is also used to receive the text corresponding to the audio stream sent by the first application, obtain the intent according to the text, and return the intent to the first application. Specifically, if the audio stream is sent from the first application 222 to the identification engine 242, the identification engine 242 returns the intent corresponding to the audio stream to the first application 222; if the audio stream is sent from the first application 232 to the identification engine 242, the identification engine 242 returns the intent corresponding to the audio stream to the first application 232.

[0118] The identification engine 242 is also used to receive the intent corresponding to the audio stream sent by the first application, obtain the instruction according to the intent, and return the instruction to the first application. Specifically, if the audio stream is sent from the first application 222 to the identification engine 242, the identification engine 242 returns the instruction corresponding to the audio stream to the first application 222; if the audio stream is sent from the first application 232 to the identification engine 242, the identification engine 242 returns the instruction corresponding to the audio stream to the first application 232.

[0119] The first application is also used to execute corresponding operations according to the instruction corresponding to the audio stream returned by the identification engine 242, so that the first device 200 executes corresponding functions. For example, the first application sends a dialogue corresponding to the audio stream to the speaker (not shown in the figure) of the first device 200 according to the instruction corresponding to the audio stream, so that the speaker of the first device 200 plays the dialogue corresponding to the audio stream, such as the dialogue being "Okay". For example, the first application sends a control instruction to the controller (not shown in the figure) of the first device 200 according to the instruction corresponding to the audio stream, so that the first device 200 executes the basic device function according to the control instruction, for example, the basic device function is to close the window.

[0120] The first application is further configured to determine a first path according to the text corresponding to the audio stream returned by the recognition engine 242, and upload the subsequent received audio stream of the first path to the server 300, and the subsequent received audio streams of other paths are not uploaded to the server 300.

[0121] The first application is further configured to receive a message sent by the server 300 and perform a corresponding function according to the message.

[0122] Figure 6 This is an architecture diagram of another function execution system provided by an embodiment of the present invention. Compared with Figure 5 the function execution system shown, Figure 6 in [the related content], the first device 200 only includes a first application 222, and the first application 222 is installed in the first screen 231. The first application 222 is configured to receive the audio streams of all microphones of the first device 200.

[0123] Optionally, the first device 200 further includes at least one other user space. The other user space includes a second screen, and the first application is not installed on the second screen.

[0124] Taking the first device 200 including one other user space as an example, as Figure 6 shown, the other user space is the co-pilot user space 230. The co-pilot user space 230 includes a second screen 231. The second screen 231 is the co-pilot screen of the first device 200. The first device 200 is configured to move the display content of the first screen 221 to the second screen 231.

[0125] Based on Figure 5 - Figure 6 the function execution system shown, an embodiment of the present invention provides a function execution method, which can improve the problems of poor privacy and high resource occupation of the wake-up-free ability of intelligent vehicles.

[0126] Figure 7 This is a signaling interaction diagram of a function execution method provided by an embodiment of the present invention. As Figure 7 shown, the method includes:

[0127] Step 402, the first application receives the first audio stream from multiple paths.

[0128] In this step, the first application of the first device receives the first audio stream from multiple paths. The operating system of the first device includes multiple microphones, and the multiple-path first audio streams are from the first audio streams sent by the multiple microphones. One audio stream can be understood as the audio stream sent by one microphone.

[0129] Exemplarily, the first audio stream may be the audio stream collected by the microphone during the first-round conversation between the user and the first application. It should be understood that when the microphone of the first device collects the first audio stream, the first device automatically starts the first application, and the first application receives the first audio stream sent by the microphone.

[0130] Exemplarily, when the first device includes multiple first applications, each first application receives the first audio stream from at least one path. For example, as Figure 5 shown, the first application on the first screen receives the first audio stream from three paths, namely the first audio streams sent by the driver's seat microphone, the left rear microphone, and the right rear microphone; the first application on the second screen receives the first audio stream from one path, namely the first audio stream sent by the co-driver's seat microphone. In the present application, when the first electronic device has multiple screens and each screen is installed with a first application, each first application receives the audio stream from at least one path, so that each screen can receive the audio stream near it.

[0131] Exemplarily, when the first device includes one first application, the first application receives the first audio stream from multiple paths. For example, as Figure 6 shown, the first application on the first screen receives the first audio stream from four paths, namely the first audio streams sent by the driver's seat microphone, the co-driver's seat microphone, the left rear microphone, and the right rear microphone.

[0132] Step 404: The first application sends the first audio streams from multiple paths to the merging module.

[0133] In this step, the first application of the first device sends the first audio streams from multiple paths to the merging module. The first device further includes a headless user space, and the headless user space includes a merging module and an identification engine.

[0134] Exemplarily, when the first device includes multiple first applications, each first application sends the first audio stream received from at least one path to the merging module. For example, as Figure 5 shown, the first application on the first screen sends the first audio stream received from three paths to the merging module; the first application on the second screen sends the first audio stream received from one path to the merging module.

[0135] Exemplarily, when the first device includes one first application, the first application sends the first audio stream received from multiple paths to the merging module. For example, as Figure 6 shown, the first application on the first screen sends the first audio stream received from four paths to the merging module.

[0136] Step 406: The merging module merges the first audio streams from multiple paths to obtain a first merged audio stream.

[0137] In this step, the merging module of the first device merges the first audio streams from multiple paths to obtain a first merged audio stream.

[0138] Step 408: The merging module sends the first merged audio stream to the recognition engine.

[0139] In this step, the merging module of the first device sends the first merged audio stream to the recognition engine.

[0140] Step 410: The recognition engine obtains the first merged text corresponding to the first merged audio stream according to the first merged audio stream, and determines the first text corresponding to the first audio stream according to the first merged text.

[0141] In this step, the recognition engine of the first device can convert the audio stream into text.

[0142] Exemplarily, the first audio stream includes first information, and the first information includes an identifier for indicating the microphone corresponding to the first audio stream. The first merged audio stream and the first merged text also include the first information of multiple first audio streams, and the first text corresponding to the first audio stream in the first merged text can be determined according to the first information.

[0143] Before processing multiple first audio streams, the first device of the present application merges the multiple first audio streams into one first merged audio stream. Compared with recognizing multiple first audio streams simultaneously, the central processing unit (CPU), neural-network processing unit (NPU), and memory resources occupied by the recognition engine of the first device are reduced. With the increase in the number of screens and seats of the first device, reducing resource occupancy can bring good scalability.

[0144] Step 412: The recognition engine sends the first text corresponding to the first audio stream to the first application.

[0145] In this step, the recognition engine of the first device sends the first text corresponding to the first audio stream to the first application.

[0146] Exemplarily, when the first device includes multiple first applications, the recognition engine sends the first text corresponding to the first audio stream to the first application corresponding to the first audio stream. For example, as Figure 5 shown, if the first audio stream is sent from the first application on the first screen to the recognition engine, the recognition engine returns the first text corresponding to the first audio stream to the first application on the first screen; if the first audio stream is sent from the first application on the second screen to the recognition engine, the recognition engine returns the first text corresponding to the first audio stream to the first application on the second screen.

[0147] Exemplarily, when the first device includes a first application, the recognition engine sends the first text corresponding to the first audio stream to the first application. For example, as Figure 6 shown, all the first audio streams are sent from the first application on the first screen to the recognition engine, and the recognition engine returns the first text corresponding to the first audio stream to the first application on the first screen.

[0148] Step 414: The first application determines the first path according to the first text corresponding to the first audio stream.

[0149] In this step, the first application of the first device determines the first path according to the first text corresponding to the first audio stream.

[0150] In some possible embodiments, step 414 specifically includes: If the first text includes a target field, the first application determines the path where the first audio stream corresponding to the first text is located as the first path.

[0151] Exemplarily, the target field includes the wake-up word of the first application. Thus, in this application, only the path in the first audio stream that contains the wake-up word is determined as the first path.

[0152] Exemplarily, the first application of the first device only determines one path among multiple paths as the first path. In the first device, there can only be one first path, and there cannot be multiple first paths at the same time.

[0153] Exemplarily, when the first device includes multiple first applications, if a first application determines the first path, it notifies other first applications based on the communication between the screens. For example, as Figure 5 shown, if the first application on the first screen determines the path of the driver's microphone as the first path, it notifies the first application on the second screen based on the communication between the first screen and the second screen. Thus, in this application, only the path in the first audio stream that contains the wake-up word is determined as the first path. In this application, when the first device includes multiple screens, the multiple screens communicate with each other, so that there is only one first path at the same time.

[0154] Exemplarily, when the first device includes a first application, if the first application on the first screen determines the first path, there is no need to notify the second screen. The first device moves the display content of the first screen to other screens. For example, as Figure 6 shown, the first application on the first screen receives the first audio streams of all microphones. If the first application on the first screen determines the path of the driver's microphone as the first path, there is no need to notify the second screen. Since there is no first application on the second screen, the first device moves the display content of the first screen to other screens. In this application, when the first device has multiple screens and only one first application is installed, by moving the screen, the interface of the first application can be displayed on all screens, enabling the user to see the first application through the nearby screen and improving the user experience.

[0155] In this application, if the first text does not include the target field, the first application determines the path where the first audio stream corresponding to the first text is located as the second path. In this application, only the path where the first audio stream corresponding to the first text including the target field is located is determined as the first path, and the path where the first audio stream corresponding to the first text not including the target field is located is determined as the second path, so that there is only one first path on the first device at the same time, ensuring that only one path goes to the cloud during the second and subsequent conversations with the user, protecting the user's privacy, reducing the network bandwidth and cloud-side computing resources occupied by multiple audio streams going to the cloud at the same time; with the increase in the number of screens and seats of the first device, it has good scalability.

[0156] Step 416: The first application sends the first text corresponding to the first audio stream to the recognition engine.

[0157] In this step, the first application of the first device sends the first text corresponding to the first audio stream returned by the recognition engine to the recognition engine again.

[0158] It should be understood that step 414 and step 416 can be executed simultaneously, or step 414 is executed after step 416, or step 414 is executed before step 416.

[0159] Step 418: The recognition engine determines the first intention according to the first text corresponding to the first audio stream.

[0160] In this step, the recognition engine of the first device can convert the text into an intention.

[0161] Step 420: The recognition engine sends the first intention to the first application.

[0162] Exemplarily, when the first device includes multiple first applications, the recognition engine sends the first intention corresponding to the first audio stream to the first application corresponding to the first audio stream. For example, as Figure 5 shown, if the first audio stream is sent from the first application on the first screen to the recognition engine, the recognition engine returns the first intention corresponding to the first audio stream to the first application on the first screen; if the first audio stream is sent from the first application on the second screen to the recognition engine, the recognition engine returns the first intention corresponding to the first audio stream to the first application on the second screen.

[0163] Exemplarily, when the first device includes one first application, the recognition engine sends all the first texts corresponding to the first audio stream to the first application. For example, as Figure 6 shown, if all the first audio streams are sent from the first application on the first screen to the recognition engine, the recognition engine returns all the first intentions corresponding to the first audio stream to the first application on the first screen.

[0164] Step 422, the first application sends the first intent corresponding to the first audio stream to the recognition engine.

[0165] In this step, the first application on the first device sends the first intent corresponding to the first audio stream to the recognition engine.

[0166] Step 424, the recognition engine obtains a first instruction according to the first intent corresponding to the first audio stream.

[0167] In this step, the recognition engine on the first device obtains a first instruction according to the first intent corresponding to the first audio stream.

[0168] Exemplarily, the first instruction includes information for representing a first function.

[0169] Exemplarily, the first function includes playing a first conversation corresponding to the first audio stream and / or enabling a first corresponding function corresponding to the first audio stream. In this application, the first device can communicate with the user and / or enable a corresponding function according to what the user says.

[0170] For example, the first text corresponding to the first audio stream is "turn on the air conditioner", the information of the first conversation corresponding to the first audio stream is "ok", and the information of the first corresponding function to be enabled is "turn on the air conditioner".

[0171] Step 426, the recognition engine sends the first instruction corresponding to the first audio stream to the first application.

[0172] In this step, the recognition engine on the first device sends the first instruction corresponding to the first audio stream to the first application.

[0173] Exemplarily, when the first device includes multiple first applications, the recognition engine sends the first instruction corresponding to the first audio stream to the first application corresponding to the first audio stream. For example, as Figure 5 shown, if the first audio stream is sent from the first application on the first screen to the recognition engine, the recognition engine returns the first instruction corresponding to the first audio stream to the first application on the first screen; if the first audio stream is sent from the first application on the second screen to the recognition engine, the recognition engine returns the first instruction corresponding to the first audio stream to the first application on the second screen.

[0174] Exemplarily, when the first device includes one first application, the recognition engine sends all the first instructions corresponding to the first audio stream to the first application. For example, as Figure 6 shown, if all the first audio streams are sent from the first application on the first screen to the recognition engine, the recognition engine returns all the first instructions corresponding to the first audio stream to the first application on the first screen.

[0175] Step 428, the first application performs a first operation according to the first instruction so that the first device performs a first function.

[0176] In this step, the first application of the first device executes a first operation according to a first instruction, so that the first device executes a first function. Through the recognition engine, the first device of the present application can convert an audio stream into text, convert the text into an intention, convert the intention into an instruction, and then the first device executes the first function according to the first instruction, so that the audio stream can be locally processed to execute the function corresponding to the audio stream. When multiple people speak simultaneously, the first device can execute the functions corresponding to the audio streams of multiple users at the same time.

[0177] Exemplarily, the first operation is to send the information of the first conversation corresponding to the first audio stream to the speaker of the first device and / or send a control instruction to the controller of the first device to turn on the first corresponding function according to the information of the first corresponding function to be turned on.

[0178] For example, the first operation is that the first application sends the information of the first conversation to the speaker of the first device. The speaker of the first device plays the first conversation according to the information of the first conversation, so that the first device plays the first conversation.

[0179] For example, the first operation is to send a control instruction to the controller of the first device to turn on the first corresponding function according to the information of the first corresponding function to be turned on. The controller of the first device turns on the first corresponding function according to the control instruction, so that the first device turns on the first corresponding function.

[0180] In summary, when the first device has a first-round conversation with at least one user and the first application, it does not upload the audio stream to the server and only performs local processing, protecting the privacy of users, reducing the network bandwidth and cloud-side computing resources occupied by multiple audio streams going to the cloud at the same time.

[0181] Step 430: The first application receives a second audio stream from multiple channels, and the second audio stream includes a third audio stream, and the third audio stream comes from the first channel.

[0182] In this step, the first application of the first device receives a second audio stream from multiple channels. The operating system of the first device includes multiple microphones, and the second audio streams from multiple channels come from the second audio streams sent by multiple microphones. An audio stream can be understood as an audio stream sent by one microphone.

[0183] Exemplarily, the second audio stream can be the audio stream collected by the microphone in the second-round conversation and subsequent conversations between the user and the first application.

[0184] Exemplarily, when the first device includes multiple first applications, each first application receives a second audio stream from at least one channel. For example, as Figure 5As shown, the first application on the first screen receives the second audio stream from three sources, namely, the second audio streams sent by the driver's microphone, the left-rear microphone, and the right-rear microphone; the first application on the second screen receives the second audio stream from one source, namely, the second audio stream sent by the co-driver's microphone.

[0185] Exemplarily, when the first device includes a first application, the first application receives the second audio stream from multiple sources. For example, as Figure 6 shown, the first application on the first screen receives the second audio stream from four sources, namely, the second audio streams sent by the driver's microphone, the co-driver's microphone, the left-rear microphone, and the right-rear microphone.

[0186] Step 432: The first application sends the third audio stream to the server.

[0187] In this step, the first application of the first device sends the third audio stream to the server. In this application, during the second-round conversation between the user and the first application and subsequent conversations, only the audio stream of the first path is uploaded to the server, that is, to the cloud, which protects the user's privacy, reduces the network bandwidth and cloud-side computing resources occupied by multiple audio streams being uploaded to the cloud simultaneously; with the increase in the number of screens and seats of the first device, it has good scalability.

[0188] Step 434: The server sends a first message to the first application.

[0189] In this step, after receiving the third audio stream, the server sends a first message to the first application of the first device.

[0190] Exemplarily, the first message contains information about the second function corresponding to the third audio stream. In this application, the first message returned by the server to the first device contains the second function, enabling the first device to execute the second function according to the first message.

[0191] Step 436: The first application executes a second operation according to the first message to enable the first device to execute the second function.

[0192] In this step, the first application of the first device executes a second operation according to the first message to enable the first device to execute the second function.

[0193] Exemplarily, the second function includes playing the second conversation corresponding to the third audio stream and / or enabling the second corresponding function corresponding to the third audio stream. In this application, the first device uploads the audio stream to the cloud and can execute the function corresponding to the audio stream by receiving the message returned by the server, that is, to have a conversation with the user and / or enable the corresponding function according to what the user says.

[0194] Exemplarily, the first operation is to send the information of the second conversation to the speaker of the first device and / or send a control instruction to the controller of the first device to turn on the second corresponding function according to the information of the second corresponding function to be turned on.

[0195] For example, the text corresponding to the first audio stream of the first path is "Navigate to Joy City", and the first conversation played by the first device is "Do you want to go to Joy City in Nankai District or Peace District?"; the text corresponding to the second audio stream of the first path is "Joy City in Nankai District", and the first application of the first device sends "Navigate to Joy City in Nankai District" to the first server, and the first server returns "Okay" and the navigation route to Joy City in Nankai District to the first application of the first device. Among them, "Okay" is the information of the second conversation, and the navigation route is the information of the second corresponding function to be turned on. The second operation performed by the first application includes sending "Okay" to the speaker of the first device and starting the navigation application of the first device according to the navigation route, so that the first device plays "Okay" and navigates according to the navigation route through the navigation application.

[0196] Step 438: The first application sends the second audio streams of multiple paths to the merging module.

[0197] In this step, the first application of the first device sends the second audio streams of multiple paths to the merging module.

[0198] Exemplarily, when the first device includes multiple first applications, each first application sends the second audio stream received from at least one path to the merging module. For example, as Figure 5 shown, the first application on the first screen sends the second audio stream received from three paths to the merging module; the first application on the second screen sends the second audio stream received from one path to the merging module.

[0199] Exemplarily, when the first device includes one first application, the first application sends the second audio streams received from multiple paths to the merging module. For example, as Figure 6 shown, the first application on the first screen sends the second audio stream received from four paths to the merging module.

[0200] It should be understood that step 438 and step 432 can be executed simultaneously, or step 438 is executed before step 432, or step 438 is executed after step 432.

[0201] Among them, in step 438, the first application sends the second audio stream of the first path and the second audio stream of the second path to the merging module. It can be seen from step 414 that the first audio stream of the second path does not include the target field, and the target field includes the wake-up word of the first application. Therefore, this application only performs local processing on the second audio stream of the second path, only supports interacting with the first application with wake-up-free statements, and reduces accidental activation.

[0202] Step 440: The merging module merges multiple second audio streams to obtain a second merged audio stream.

[0203] In this step, the merging module of the first device merges multiple second audio streams to obtain a second merged audio stream.

[0204] Step 442: The merging module sends the second merged audio stream to the recognition engine.

[0205] In this step, the merging module of the first device sends the second merged audio stream to the recognition engine.

[0206] Step 444: The recognition engine obtains a second merged text corresponding to the second merged audio stream based on the second merged audio stream, and determines a second text corresponding to the second audio stream based on the second merged text.

[0207] In this step, the recognition engine of the first device can convert the audio stream into text.

[0208] Exemplarily, the second audio stream includes second information, and the second information includes an identifier for indicating a microphone corresponding to the second audio stream. The second information of multiple second audio streams is also included in the second merged audio stream and the second merged text, and the second text corresponding to the second audio stream in the second merged text can be determined according to the second information.

[0209] Before processing multiple second audio streams, the first device in this application merges the multiple second audio streams into one second merged audio stream. Compared with recognizing multiple second audio streams simultaneously, it reduces the CPU, NPU, and memory resources occupied by the recognition engine of the first device. With the increase in the number of screens and seats of the first device, reducing resource occupancy can bring good scalability.

[0210] Step 446: The recognition engine sends the second text corresponding to the second audio stream to the first application.

[0211] In this step, the recognition engine of the first device sends the second text corresponding to the second audio stream to the first application.

[0212] Exemplarily, when the first device includes multiple first applications, the recognition engine sends the second text corresponding to the second audio stream to the first application corresponding to the second audio stream. For example, as Figure 5 shown, if the second audio stream is sent from the first application on the first screen to the recognition engine, the recognition engine returns the second text corresponding to the second audio stream to the first application on the first screen; if the second audio stream is sent from the first application on the second screen to the recognition engine, the recognition engine returns the second text corresponding to the second audio stream to the first application on the second screen.

[0213] Exemplarily, when the first device includes a first application, the recognition engine sends the second text corresponding to the second audio stream to the first application. For example, as Figure 6 shown, if the second audio stream is sent from the first application on the first screen to the recognition engine, the recognition engine returns the second text corresponding to the second audio stream to the first application on the first screen.

[0214] Step 448: The first application sends the second text corresponding to the second audio stream to the recognition engine.

[0215] In this step, the first application on the first device sends the second text corresponding to the second audio stream returned by the recognition engine to the recognition engine again.

[0216] Step 450: The recognition engine determines a second intent according to the second text corresponding to the second audio stream.

[0217] In this step, the recognition engine on the first device can convert the text into an intent.

[0218] Step 452: The recognition engine sends the second intent to the first application.

[0219] Exemplarily, when the first device includes multiple first applications, the recognition engine sends the second intent corresponding to the second audio stream to the first application corresponding to the first audio stream. For example, as Figure 5 shown, if the second audio stream is sent from the first application on the first screen to the recognition engine, the recognition engine returns the second intent corresponding to the second audio stream to the first application on the first screen; if the second audio stream is sent from the first application on the second screen to the recognition engine, the recognition engine returns the second intent corresponding to the second audio stream to the first application on the second screen.

[0220] Exemplarily, when the first device includes a first application, the recognition engine sends the second text corresponding to the second audio stream to the first application. For example, as Figure 6 shown, if the second audio stream is sent from the first application on the first screen to the recognition engine, the recognition engine returns the second intent corresponding to the second audio stream to the first application on the first screen.

[0221] Step 454: The first application sends the second intent corresponding to the second audio stream to the recognition engine.

[0222] In this step, the first application on the first device sends the second intent corresponding to the second audio stream to the recognition engine.

[0223] Step 456: The recognition engine obtains a third instruction according to the second intent corresponding to the second audio stream.

[0224] In this step, the recognition engine on the first device obtains a third instruction according to the second intent corresponding to the second audio stream.

[0225] Exemplarily, the third instruction includes information for representing a third function.

[0226] Exemplarily, the third function includes playing a third conversation corresponding to the second audio stream and / or activating a third corresponding function corresponding to the second audio stream. In this application, the first device can converse with the user and / or activate a corresponding function according to what the user says.

[0227] For example, the third instruction includes information about the third conversation and / or information about the third corresponding function.

[0228] For example, the second text corresponding to the second audio stream is "turn off the air conditioner", the information about the third conversation corresponding to the second audio stream is "okay", and the information about the third corresponding function to be activated is "turn off the air conditioner".

[0229] Step 458: The recognition engine sends the third instruction corresponding to the second audio stream to the first application.

[0230] In this step, the recognition engine of the first device sends the third instruction corresponding to the second audio stream to the first application.

[0231] Exemplarily, when the first device includes multiple first applications, the recognition engine sends the third instruction corresponding to the second audio stream to the first application corresponding to the second audio stream. For example, as Figure 5 shown, if the second audio stream is sent from the first application on the first screen to the recognition engine, the recognition engine returns the third instruction corresponding to the second audio stream to the first application on the first screen; if the second audio stream is sent from the first application on the second screen to the recognition engine, the recognition engine returns the third instruction corresponding to the second audio stream to the first application on the second screen.

[0232] Exemplarily, when the first device includes one first application, the recognition engine sends all the third instructions corresponding to the second audio stream to the first application. For example, as Figure 6 shown, if the second audio stream is all sent from the first application on the first screen to the recognition engine, the recognition engine returns all the third instructions corresponding to the second audio stream to the first application on the first screen.

[0233] Step 460: The first application performs a third operation according to the third instruction so that the first device performs the third function.

[0234] In this step, the first application of the first device performs a third operation according to a third instruction, so that the first device performs a third function. The first device of the present application can convert an audio stream into text, convert the text into an intention, and convert the intention into an instruction through an identification engine. Then, the first device performs a first function according to the first instruction, so that the audio stream can be locally processed to execute the function corresponding to the audio stream. When multiple people speak simultaneously, the first device can simultaneously execute the functions corresponding to the audio streams of multiple users. At the same time, the second audio stream of the second path in the present application is not sent to the cloud, but only locally processed, reducing misoperation.

[0235] Exemplarily, the third operation is to send the information of the third conversation corresponding to the second audio stream to the speaker of the first device and / or send a control instruction to turn on the third corresponding function to the controller of the first device according to the information of the third corresponding function to be turned on. For example, the third operation is that the first application sends the information of the third conversation to the speaker of the first device. The speaker of the first device plays the third conversation according to the information of the third conversation, so that the first device plays the third conversation.

[0236] For example, the third operation is to send a control instruction to turn on the third corresponding function to the controller of the first device according to the information of the third corresponding function to be turned on. The controller of the first device turns on the third corresponding function according to the control instruction, so that the first device turns on the third corresponding function.

[0237] In summary, the first device determines the first path during the first-round conversation. In the conversations after the first-round conversation, only the audio stream of the first path is uploaded to the server, ensuring that only one audio stream is sent to the cloud, protecting the privacy of users, reducing the network bandwidth and cloud-side computing resources occupied by multiple audio streams being sent to the cloud simultaneously; at the same time, the first device locally processes the audio streams of all paths, so both the first path and the other paths except the first path support wake-free speech to interact with the first application; and before locally processing the audio streams of all paths, the first device merges multiple audio streams into one, reducing the CPU, NPU, and memory resources occupied by the identification engine of the first device; with the increase in the number of screens and seats of the first device, it has good scalability.

[0238] Figure 8 It is a schematic diagram of an application scenario of an embodiment of the present invention. As Figure 8As shown, the vehicle includes 4 seats. Assuming the wake-up word of the voice assistant is "Xiaodu", the driver and the front passenger speak to the voice assistant application for the first time. The driver says "Xiaodu, navigate to Joy City", and the front passenger says "Turn on the air conditioner"; the vehicle asks the driver "Do you want to go to Joy City in Nankai District or Joy City in Heping District?", and says "Okay" to the front passenger and turns on the air conditioner. In the first conversation, the voice assistant application does not upload the audio stream to the server; and before the recognition engine recognizes the audio stream, the two audio streams have been merged, and the recognition engine only performs the process of recognizing the audio stream as text once.

[0239] After that, the driver and the front passenger speak to the voice assistant application for the second time. As Figure 8 shown, the driver says "Joy City in Nankai District", and the front passenger says "Turn on the aromatherapy"; the vehicle says "Okay" to the driver and turns on the navigation to navigate to Joy City in Nankai District, and says "Okay" to the front passenger and turns on the aromatherapy. In the second conversation, the voice assistant application only uploads the driver's audio stream to the server; and before the recognition engine recognizes the audio stream, the two audio streams have been merged, and the recognition engine only performs the process of recognizing the audio stream as text once; for the first path, it supports uploading to the cloud, and for other paths, it only supports wake-free speech.

[0240] In the technical solution of the function execution method provided by the embodiment of the present invention, the method includes: receiving a first audio stream from multiple paths; determining a first path according to the first audio stream, and executing a first function corresponding to the first audio stream of the first path; receiving a second audio stream from multiple paths, the second audio stream includes a third audio stream, and the third audio stream comes from the first path; sending the third audio stream to the server; receiving a first message sent by the server, and executing a second function corresponding to the third audio stream according to the first message, which can improve the problems of poor privacy and high resource occupancy of the wake-free ability of intelligent vehicles.

[0241] Figure 9 It is a flowchart of a function execution method provided by an embodiment of the present invention. As Figure 9 shown, the method includes:

[0242] Step 502, the first device receives a first audio stream from multiple paths.

[0243] In some possible embodiments, the first device includes a first application; step 502 specifically includes: receiving a first audio stream from multiple paths through the first application.

[0244] Exemplarily, the first device includes a first screen, and the first application is installed on the first screen; or the first device includes a first screen and at least one second screen, the first application is installed on the first screen, and the first device moves the display content of the first screen to the second screen.

[0245] In some possible embodiments, the first device includes a plurality of first applications and a plurality of screens, the first applications are installed on the screens, and the first applications and the screens are in one-to-one correspondence; step 502 specifically includes: receiving the first audio stream from multiple paths through the plurality of first applications, and the first applications are used to receive the first audio stream from at least one path.

[0246] Exemplarily, the plurality of screens communicate with each other.

[0247] Step 504: Determine the first path according to the first audio stream, and execute the first function corresponding to the first audio stream of the first path.

[0248] Exemplarily, the first function includes playing the first conversation corresponding to the first audio stream and / or enabling the first corresponding function corresponding to the first audio stream.

[0249] In some possible embodiments, determining the first path according to the first audio stream specifically includes: merging the first audio streams of multiple paths to obtain a first merged audio stream; obtaining a first merged text corresponding to the first merged audio stream according to the first merged audio stream; and determining the first path according to the first merged text.

[0250] In some possible embodiments, determining the first path according to the first merged text specifically includes: determining a first text corresponding to the first audio stream according to the first merged text; and if the first text includes a target field, determining the path where the first audio stream corresponding to the first text is located as the first path.

[0251] In some possible embodiments, executing the first function corresponding to the first audio stream of the first path specifically includes: determining a first intention according to the first text corresponding to the first audio stream of the first path; obtaining a first instruction according to the first intention; and executing the first function according to the first instruction.

[0252] Step 506: Receive a second audio stream from multiple paths, the second audio stream includes a third audio stream, and the third audio stream comes from the first path.

[0253] In some possible embodiments, the first device includes a first application; step 506 specifically includes: receiving the second audio stream from multiple paths through the first application.

[0254] In some possible embodiments, the first device includes a plurality of first applications and a plurality of screens, the first applications are installed on the screens, and the first applications and the screens are in one-to-one correspondence; step 506 specifically includes: receiving the second audio stream from multiple paths through the plurality of first applications, and the first applications are used to receive the second audio stream from at least one path.

[0255] Step 508: Send the third audio stream to the server, receive a first message sent by the server, and execute a second function corresponding to the third audio stream according to the first message.

[0256] Exemplarily, the first message includes information about the second function.

[0257] Exemplarily, the second function includes playing a second dialogue corresponding to a third audio stream and / or activating a second corresponding function corresponding to the third audio stream.

[0258] Optionally, after step 502, the method further includes: step 503.

[0259] Step 503: Determine a second path according to the first audio stream, and execute a first function corresponding to the first audio stream on the second path.

[0260] It should be understood that step 503 may be executed simultaneously with step 504.

[0261] Wherein, determining the second path according to the first audio stream includes: if the first text corresponding to the first audio stream does not include a target field, determining the path where the first audio stream corresponding to the first text is located as the second path.

[0262] Wherein, the method for executing the first function corresponding to the first audio stream on the second path is similar to the method for executing the first function corresponding to the first audio stream on the first path in step 504, and will not be elaborated herein.

[0263] Optionally, after step 506, the method further includes: step 507.

[0264] Step 507: Execute a third function corresponding to the second audio stream according to the second audio stream.

[0265] It should be understood that step 507 may be executed simultaneously with step 508.

[0266] In some possible embodiments, step 510 specifically includes: merging multiple paths of the second audio streams to obtain a second merged audio stream; obtaining a second merged text corresponding to the second merged audio stream according to the second merged audio stream; determining a second text corresponding to the second audio stream according to the second merged text; determining a second intention according to the second text; obtaining a third instruction according to the second intention; and executing a third function according to the third instruction.

[0267] Exemplarily, the third function includes playing a third dialogue corresponding to the second audio stream and / or activating a third corresponding function corresponding to the second audio stream.

[0268] In the technical solution of the function execution method provided by the embodiments of the present invention, the method includes: receiving a first audio stream from multiple channels; determining a first path according to the first audio stream, and executing a first function corresponding to the first audio stream of the first path; receiving a second audio stream from multiple channels, where the second audio stream includes a third audio stream, and the third audio stream comes from the first path; sending the third audio stream to a server; receiving a first message sent by the server, and executing a second function corresponding to the third audio stream according to the first message, which can improve the problems of poor privacy and high resource occupation in the wake-up-free ability of intelligent vehicles.

[0269] Figure 10 FIG. 600 is a schematic structural diagram of a first device provided by an embodiment of the present invention. It should be understood that the first device 600 can execute each step of the first device in the above function execution method. To avoid repetition, details are not described herein again. The first device 600 includes: a transceiver unit 601 and a processing unit 602.

[0270] The transceiver unit 601 is configured to receive a first audio stream from multiple channels;

[0271] The processing unit 602 is configured to determine a first path according to the first audio stream, and execute a first function corresponding to the first audio stream of the first path;

[0272] The transceiver unit 601 is further configured to receive a second audio stream from the multiple channels, where the second audio stream includes a third audio stream, and the third audio stream comes from the first path; send the third audio stream to a server, and receive a first message sent by the server,

[0273] The processing unit 602 is further configured to execute a second function corresponding to the third audio stream according to the first message.

[0274] Optionally, the processing unit 602 is specifically configured to merge the first audio streams of the multiple channels to obtain a first merged audio stream; obtain a first merged text corresponding to the first merged audio stream according to the first merged audio stream; and determine the first path according to the first merged text.

[0275] Optionally, the processing unit 602 is specifically configured to determine a first text corresponding to the first audio stream according to the first merged text; if the first text includes a target field, determine the path where the first audio stream corresponding to the first text is located as the first path.

[0276] Optionally, the processing unit 602 is specifically configured to determine a first intention according to the first text corresponding to the first audio stream of the first path; obtain a first instruction according to the first intention; and execute the first function according to the first instruction.

[0277] Optionally, the first function includes playing a first dialogue corresponding to the first audio stream and / or activating a first corresponding function corresponding to the first audio stream.

[0278] Optionally, the first message contains information about the second function.

[0279] Optionally, the second function includes playing the second dialogue corresponding to the third audio stream and / or activating the second corresponding function corresponding to the third audio stream.

[0280] Optionally, after the transceiver unit 601 receives the first audio stream from multiple channels, the processing unit 602 is further configured to determine a second channel according to the first audio stream and execute a first function corresponding to the first audio stream on the second channel.

[0281] Optionally, the processing unit 602 is specifically configured to, if the first text corresponding to the first audio stream does not include a target field, determine the channel where the first audio stream corresponding to the first text is located as the second channel.

[0282] Optionally, after the transceiver unit 601 receives the second audio stream from the multiple channels, the processing unit 602 is further configured to execute a third function corresponding to the second audio stream according to the second audio stream.

[0283] Optionally, the processing unit 602 is specifically configured to merge the second audio streams of the multiple channels to obtain a second merged audio stream; obtain a second merged text corresponding to the second merged audio stream according to the second merged audio stream; determine a second text corresponding to the second audio stream according to the second merged text; determine a second intention according to the second text; obtain a third instruction according to the second intention; and execute the third function according to the third instruction.

[0284] Optionally, the third function includes playing a third dialogue corresponding to the second audio stream and / or activating a third corresponding function corresponding to the second audio stream.

[0285] Optionally, the first device includes a first application; the transceiver unit 601 is specifically configured to receive the first audio stream from the multiple channels through the first application; and receive the second audio stream from the multiple channels through the first application.

[0286] Optionally, the first device includes a first screen, and the first application is installed on the first screen; or

[0287] The first device includes a first screen and at least one second screen, the first application is installed on the first screen, and the processing unit 602 is further configured to move the display content of the first screen to the second screen.

[0288] Optionally, the first device includes a plurality of first applications and a plurality of screens, the first applications are installed on the screens, and the first applications and the screens are in one-to-one correspondence; the transceiver unit 601 is specifically configured to receive the first audio stream from the multiple channels through the plurality of first applications, and the first application is configured to receive the first audio stream from at least one channel; receive the second audio stream from the multiple channels through the plurality of first applications, and the first application is configured to receive the second audio stream from at least one channel.

[0289] Optionally, the plurality of screens communicate with each other.

[0290] It should be understood that the first device 600 here is embodied in the form of a functional unit. The term "unit" here can be implemented in the form of software and / or hardware, and no specific limitation is made thereto. For example, the "unit" can be a software program, a hardware circuit, or a combination of the two to implement the above functions. The hardware circuit may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor, or a group of processors, etc.) for executing one or more software or firmware programs, a memory, a combined logic circuit, and / or other suitable components supporting the described functions.

[0291] Therefore, the units of the examples described in the embodiments of the present invention can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0292] An embodiment of the present application provides a device, which can be a terminal device or a circuit device built into the terminal device. The device can be used to execute the functions / steps in the above method embodiments.

[0293] An embodiment of the present application provides a readable storage medium, in which instructions are stored, and when the instructions run on a device, the device is caused to execute the functions / steps in the above method embodiments.

[0294] An embodiment of the present application further provides a program product including instructions, and when the program product runs on a device or any at least one processor, the device is caused to execute the functions / steps in the above method embodiments.

[0295] In the embodiments of the present application, "at least one" means one or more, and "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent the situations where A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0296] Those of ordinary skill in the art can realize that the various units and algorithm steps described in the embodiments disclosed herein can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0297] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0298] In several embodiments provided by the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0299] The above is only the specific implementation manner of the present application. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application and should be covered by the protection scope of the present application. The protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A method for function execution, characterized in that, Applied to a first device, the method includes: Receiving a first audio stream from multiple channels; Determining a first channel according to the first audio stream and performing a first function corresponding to the first audio stream of the first channel; Receiving a second audio stream from the multiple channels, the second audio stream including a third audio stream from the first channel; Sending the third audio stream to a server; Receiving a first message sent by the server and performing a second function corresponding to the third audio stream according to the first message.

2. The method according to claim 1, characterized in that, The determining a first channel according to the first audio stream includes: Merging the first audio streams of the multiple channels to obtain a first merged audio stream; Obtaining a first merged text corresponding to the first merged audio stream according to the first merged audio stream; Determining the first channel according to the first merged text.

3. The method according to claim 2, characterized in that, The determining the first channel according to the first merged text includes: Determining a first text corresponding to the first audio stream according to the first merged text; If the first text includes a target field, determining the channel where the first audio stream corresponding to the first text is located as the first channel.

4. The method according to any one of claims 1 - 3, characterized in that, The performing a first function corresponding to the first audio stream of the first channel includes: Determining a first intent according to the first text corresponding to the first audio stream of the first channel; Obtaining a first instruction according to the first intent; Performing the first function according to the first instruction.

5. The method according to any one of claims 1 - 4, characterized in that, The first function includes playing a first dialogue corresponding to the first audio stream and / or enabling a first corresponding function corresponding to the first audio stream.

6. The method according to any one of claims 1 - 5, characterized in that, The first message contains information about the second function.

7. The method according to claim 6, characterized in that, The second function includes playing a second dialogue corresponding to the third audio stream and / or enabling a second corresponding function corresponding to the third audio stream.

8. The method according to any one of claims 1 - 7, characterized in that, After receiving the first audio stream from multiple channels, it further includes: Determining a second channel according to the first audio stream and performing a first function corresponding to the first audio stream of the second channel.

9. The method according to claim 8, characterized in that, The determining a second channel according to the first audio stream includes: If the first text corresponding to the first audio stream does not include a target field, determining the channel where the first audio stream corresponding to the first text is located as the second channel.

10. The method according to any one of claims 1 - 9, characterized in that, After receiving the second audio stream from the multiple channels, it further includes: Performing a third function corresponding to the second audio stream according to the second audio stream.

11. The method according to claim 10, characterized in that, The performing a third function corresponding to the second audio stream according to the second audio stream includes: Merging the second audio streams of the multiple channels to obtain a second merged audio stream; Obtaining a second merged text corresponding to the second merged audio stream according to the second merged audio stream; Determining a second text corresponding to the second audio stream according to the second merged text; Determining a second intent according to the second text; Obtaining a third instruction according to the second intent; Performing the third function according to the third instruction.

12. The method according to claim 10 or 11, wherein, The third function includes playing a third dialogue corresponding to the second audio stream and / or enabling a third corresponding function corresponding to the second audio stream.

13. The method according to any one of claims 1-12, wherein, The first device includes a first application; The receiving a first audio stream from multiple channels includes: Receiving the first audio stream from the multiplex through the first application; Receiving the second audio stream of the multiplex from the multiplex, including: Receiving the second audio stream from the multiplex through the first application.

14. The method according to claim 13, wherein, The first device includes a first screen, and the first application is installed on the first screen; or The first device includes a first screen and at least one second screen, the first application is installed on the first screen, and the method further includes: shifting the display content of the first screen to the second screen.

15. The method according to any one of claims 1-12, wherein, The first device includes a plurality of first applications and a plurality of screens, the first applications are installed on the screens, and the first applications and the screens correspond one by one; Receiving the first audio stream from the multiplex, including: Receiving the first audio stream from the multiplex through the plurality of first applications, where the first application is used to receive the first audio stream from at least one path; Receiving the second audio stream from the multiplex, including: Receiving the second audio stream from the multiplex through the plurality of first applications, where the first application is used to receive the second audio stream from at least one path.

16. The method according to claim 15, wherein, The plurality of screens communicate with each other.

17. A first device, wherein, Including a processor and a memory, where the memory is used to store a program, the program includes program instructions, and when the processor runs the program instructions, the first device executes the steps of the method according to any one of claims 1-16.

18. A readable storage medium, wherein, The readable storage medium stores a program, the program includes program instructions, and when the program instructions are run by a device, the device executes the method according to any one of claims 1-16.