A multi-device voice control system and method
Patent Information
- Application Number
- CN202210272315.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-18
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2042-03-18
AI Technical Summary
由此,相关技术在多设备协同进行业务处理下,存在操作繁琐等问题
[0053]上述第二方面至第九方面中任一方面的有益效果请具体参阅上述第一方面中各种可能的设计的有益效果,在此不再赘述。
Smart Images

Figure CN116805488B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio technology, and more particularly to a multi-device voice control system and method. Background Technology
[0002] With the development of semiconductor and software technologies, terminal devices have taken on various forms, such as mobile phones, tablets, televisions, in-vehicle devices, and various home appliances. Currently, many businesses may involve multiple devices, requiring them to work together to process data; for example, during video conferencing, a portable device's screen can be projected onto a large screen, multiple devices can log into the same account, and a video playing on a mobile phone can be cast to a television.
[0003] In related technologies, scenarios requiring multiple devices to collaborate on business processes typically necessitate manual user intervention. For example, if a user is already logged into an account on their mobile phone and needs to log into the same account on a computer, they generally need to manually authorize the login via their mobile phone. Similarly, if a user wants to cast a video playing on their mobile phone to a TV, they need to manually perform the casting operation on their mobile phone. Therefore, these technologies suffer from cumbersome operations when handling multi-device collaborative business processes.
[0004] Therefore, it is of research significance to explore how to reduce the operational complexity in scenarios where multiple devices collaborate on business processing. Summary of the Invention
[0005] This application provides a multi-device voice control system and method, which provides a technical solution for simultaneously controlling multiple adjacent devices via voice, thereby reducing the operational complexity in scenarios where multiple devices collaborate on business processing.
[0006] In a first aspect, embodiments of this application provide a multi-device voice control system, wherein a first terminal device receives and responds to a user's first voice command, and uploads first voice request information to a first server device, the first voice request information including the first voice command; and a second terminal device receives and responds to a user's second voice command, and uploads second voice request information to the first server device, the second voice request information including the second voice command; the first server device, based on the first voice command and the second voice command, if it determines that the first voice request information and the second voice request information are related, performs at least one of the following processes: generating and sending a first control command to the first terminal device, the first control command being used to execute a related operation of the first voice command, the related operation of the first voice command being used for collaborative processing of services with the second terminal device; generating and sending a second control command to the second terminal device, the second control command being used to execute a related operation of the second voice command; the related operation of the second voice command being used for collaborative processing of services with the first terminal device.
[0007] In this method, multiple terminal devices receive voice commands from the user. The server device can analyze the voice commands to identify the relevant terminal devices, enabling users to control multiple devices via voice commands. The server device can then instruct each terminal device to execute the received voice commands, achieving collaborative business processing across multiple devices. Compared to related technologies that require manual or even multiple manual operations from the user in multi-device collaborative business processing scenarios, the method provided in this application reduces the complexity of user operations, thereby improving the user experience.
[0008] In one possible design, the first server device determines that the first voice request information and the second voice request information are related, including but not limited to at least one of the following methods: (1) determining that the first voice instruction and the second voice instruction are the same; (2) determining that the similarity between the first voice instruction and the second voice instruction is greater than a first specified threshold; (3) determining that the first voice instruction and the second voice instruction correspond to each other.
[0009] In this design, the server device can identify whether the terminal devices need to coordinate business processing by receiving voice commands uploaded from multiple terminal devices. This allows for the control of multiple devices based on the user's voice commands, thereby reducing the complexity of user operations.
[0010] In one possible design, the first voice request information and the second voice request information may further include, but are not limited to, at least one of the following: timestamp information, voiceprint information, and terminal device status information.
[0011] In this design, when a terminal device uploads voice service request information to a server device, in addition to voice command information, it can also include other information that can help determine whether there is a relationship between terminal devices, thereby enabling more accurate voice control of multiple devices based on the user's voice commands.
[0012] In one possible design, the first server device determines that the first voice request information and the second voice request information are related, including but not limited to one or more of the following methods: determining that the timestamp information uploaded by the first terminal device and the timestamp information uploaded by the second terminal device are the same or have a similarity greater than a second specified threshold; determining that the voiceprint information uploaded by the first terminal device and the voiceprint information uploaded by the second terminal device are the same or have a similarity greater than a third specified threshold.
[0013] In this design, to improve the accuracy of identifying whether terminal devices need to collaborate on business processing, the server device can also combine other information uploaded by the terminal devices to determine whether the terminal devices are related, thereby improving the accuracy of voice control of multiple devices; and by judging voiceprint information, the security of voice control of multiple devices can also be improved.
[0014] In one possible design, the first server device performs at least one of the following processes based on the first voice command and the second voice command: the first server device performs semantic analysis on the first voice command and the second voice command; and the first server device determines the state of the first terminal device and the second terminal device based on the terminal device state information corresponding to the first terminal device and the terminal device state information corresponding to the second terminal device; the first server device performs at least one of the following processes based on the result of the semantic analysis and the state of the first terminal device and the second terminal device.
[0015] In this design, when the server device identifies the correlation between terminal devices, it can further determine the relevant operations that each terminal device needs to perform based on the user's intent corresponding to the user's voice commands and instruct the corresponding terminal devices. This allows for voice control of multiple terminal devices based on the user's voice commands. Furthermore, by combining the status of each terminal device, the server device can generate more accurate control commands, thus ensuring the accuracy of voice control across multiple devices.
[0016] In one possible design, if the process performed by the first server device is to generate and send a first control command to the first terminal device, then the first control command includes the device identifier of the second terminal device. The device identifier of the second terminal device is used by the first terminal device to perform the relevant operation of the first voice command based on the device identifier of the second terminal device. Alternatively, if the process performed by the second server device is to generate and send a second control command to the second terminal device, then the second control command includes the device identifier of the first terminal device. The device identifier of the first terminal device is used by the second terminal device to perform the relevant operation of the second voice command based on the device identifier of the first terminal device.
[0017] In this design, the server-side device, based on the user intent corresponding to the voice command, determines in some possible scenarios that certain terminal devices need to perform corresponding operations. These operations typically require collaborative processing with another group of terminal devices. In this case, the server-side device can carry the device identifiers of the other group of terminal devices in the control commands sent to the first group, enabling them to determine which terminal devices they need to collaborate with. Thus, this design eliminates the need for the first and second terminal devices to be connected or on the same local area network; either the first or second terminal device can achieve collaborative processing with the other terminal device based on the device identifier carried in the control command.
[0018] In one possible design, the first server device generates a first identification code, which is used to identify the relationship between the first terminal device and the second terminal device.
[0019] In this design, to facilitate the control of related first and second terminal devices, a first identification code ensures efficient voice control of multiple devices. For example, if the first and second terminal devices need to interact with the first or second server device or other devices after receiving control commands, the first identification code can be used to quickly identify the relationship between the first and second terminal devices.
[0020] In one possible design, the system further includes a second server device, wherein: the first terminal device sends a first request instruction to the second server device according to the first control instruction, the first control instruction and the first request instruction carrying the first identification code; and the second terminal device sends a second request instruction to the second server device according to the second control instruction, the second control instruction and the second request instruction carrying the first identification code; the second server device performs at least one of the following processes according to the first identification code: sending a first response instruction to the first terminal device and sending a second response instruction to the second terminal device.
[0021] In this design, voice control of the first and second terminal devices can also be achieved through other server-side devices in some possible scenarios. In this case, by generating a first identifier for the first and second terminal devices, the relationship between the first and second terminal devices can be quickly determined, thereby ensuring the processing efficiency of voice control across multiple devices.
[0022] In one possible design, the first request instruction is used to request login to a specified platform, and the second request instruction is used to request authorization to log in to the specified platform; the processing performed by the second terminal device is to send a second response instruction to the second terminal device, and the second response instruction is used to instruct the first terminal device to authorize the second terminal device to log in to the specified platform.
[0023] This design presents a scenario where voice control of a first terminal device and a second terminal device is used to log in to a designated platform. Compared to related technologies, which typically require users to scan a QR code generated on an unlogged-in device while already logged in, the method provided in this application can achieve voice control of multiple devices based on user voice commands, thereby reducing the complexity of user operations.
[0024] In one possible design, the first voice command and the second voice command are used to indicate, but are not limited to, any of the following scenarios: logging into a designated platform on the first terminal device or the second terminal device, connecting the first terminal device to the second terminal device, or connecting the second terminal device to the first terminal device.
[0025] This design presents a scenario where voice control of multiple devices can be achieved based on user voice commands. The method provided in this application enables voice control of multiple devices based on user voice commands, thereby reducing the complexity of user operations.
[0026] In one possible design, the first voice command and the second voice command are based on the same voice command from the user, and are received by the first terminal device and the second terminal device, respectively.
[0027] In this design, by receiving user voice commands from multiple terminal devices simultaneously, it is possible to more accurately achieve voice control of multiple devices based on user voice commands. This not only reduces the complexity of user operations but also improves the accuracy of voice control of multiple devices.
[0028] Secondly, embodiments of this application also provide a voice control method for multiple devices, comprising: a first terminal device receiving a first voice command from a user; the first terminal device responding to the first voice command by uploading first voice request information to a first server device, the first voice request information including the first voice command; the first terminal device receiving a first control command sent by the first server device, the first control command being used to execute related operations of the first voice command; the related operations of the first voice command being used to perform collaborative processing of services with a second terminal device; wherein, the first control command is generated by the first server device when it determines that the first voice request information is related to second voice request information uploaded by the second terminal device.
[0029] In one possible design, the first voice request information and the second voice request information may further include, but are not limited to, at least one of the following: timestamp information, voiceprint information, and terminal device status information.
[0030] In one possible design, the first control command includes the device identifier of the second terminal device, which is used by the first terminal device to perform the relevant operation of the first voice command based on the device identifier of the second terminal device.
[0031] In one possible design, the first control command includes a first identification code; the first identification code is used to identify the association between the first terminal device and the second terminal device.
[0032] In one possible design, the method further includes: the first terminal device sending a first request instruction to the second server device according to the first control instruction, wherein the first control instruction and the first request instruction carry the first identification code, so that the second server device performs at least one of the following processes according to the first identification code: sending a first response instruction to the first terminal device and sending a second response instruction to the second terminal device.
[0033] In one possible design, the first request instruction is used to request login to a specified platform; or, the first request instruction is used to request authorization to log in to a specified platform; the method further includes: if the first request instruction is used to request authorization to log in to a specified platform, the first terminal device receives a first response instruction, the first response instruction being used to instruct the second terminal device to authorize the first terminal device to log in to the specified platform.
[0034] In one possible design, the second voice instruction contained in the first voice instruction and the second voice request information is used to indicate, but is not limited to, any of the following scenarios: logging into a designated platform on the first terminal device or the second terminal device, connecting the first terminal device to the second terminal device, or connecting the second terminal device to the first terminal device.
[0035] In one possible design, the second voice command contained in the first voice command and the second voice request information is based on the same voice command from the user, and is received by the first terminal device and the second terminal device, respectively.
[0036] Thirdly, embodiments of this application also provide a voice control method for multiple devices, comprising: a first server device receiving first voice request information uploaded by a first terminal device, the first voice request information including a first voice instruction; and the first server device receiving second voice request information uploaded by a second terminal device, the second voice request information including a second voice instruction; the first server device, based on the first voice instruction and the second voice instruction, if it is determined that the first voice request information and the second voice request information are related, performing at least one of the following processes: generating and sending a first control instruction to the first terminal device, the first control instruction being used to execute a related operation of the first voice instruction; the related operation of the first voice instruction being used to perform collaborative processing of services with the second terminal device; generating and sending a second control instruction to the second terminal device, the second control instruction being used to execute a related operation of the second voice instruction; the related operation of the second voice instruction being used to perform collaborative processing of services with the first terminal device.
[0037] In one possible design, the first server device determines that the first voice request information and the second voice request information are related, including but not limited to at least one of the following methods: determining that the first voice instruction and the second voice instruction are the same; determining that the similarity between the first voice instruction and the second voice instruction is greater than a first specified threshold; determining that the first voice instruction and the second voice instruction correspond to each other.
[0038] In one possible design, the first voice request information and the second voice request information may further include, but are not limited to, at least one of the following: timestamp information, voiceprint information, and terminal device status information.
[0039] In one possible design, the first server device determines that the first voice request information and the second voice request information are related, including but not limited to one or more of the following methods: determining that the timestamp information uploaded by the first terminal device and the timestamp information uploaded by the second terminal device are the same or have a similarity greater than a second specified threshold; determining that the voiceprint information uploaded by the first terminal device and the voiceprint information uploaded by the second terminal device are the same or have a similarity greater than a third specified threshold.
[0040] In one possible design, the first server device performs at least one of the following processes based on the first voice command and the second voice command: the first server device performs semantic analysis on the first voice command and the second voice command; and the first server device determines the state of the first terminal device and the second terminal device based on the terminal device state information corresponding to the first terminal device and the terminal device state information corresponding to the second terminal device; the first server device performs at least one of the following processes based on the result of the semantic analysis and the state of the first terminal device and the second terminal device.
[0041] In one possible design, if the process performed by the first server device is to generate and send a first control command to the first terminal device, then the first control command includes the device identifier of the second terminal device. The device identifier of the second terminal device is used by the first terminal device to perform the relevant operation of the first voice command based on the device identifier of the second terminal device. Alternatively, if the process performed by the second server device is to generate and send a second control command to the second terminal device, then the second control command includes the device identifier of the first terminal device. The device identifier of the first terminal device is used by the second terminal device to perform the relevant operation of the second voice command based on the device identifier of the first terminal device.
[0042] In one possible design, the method further includes: the first server device generating a first identification code, the first identification code being used to identify the association between the first terminal device and the second terminal device.
[0043] In one possible design, the first terminal device sends a first request instruction to the second server device according to the first control instruction, wherein the first control instruction and the first request instruction carry the first identification code; and the second terminal device sends a second request instruction to the second server device according to the second control instruction, wherein the second control instruction and the second request instruction carry the first identification code; the second server device performs at least one of the following processes according to the first identification code: sending a first response instruction to the first terminal device and sending a second response instruction to the second terminal device.
[0044] In one possible design, the first request instruction is used to request login to a specified platform, and the second request instruction is used to request authorization to log in to the specified platform; the processing performed by the second terminal device is to send a second response instruction to the second terminal device, and the second response instruction is used to instruct the first terminal device to authorize the second terminal device to log in to the specified platform.
[0045] In one possible design, the first voice command and the second voice command are used to indicate, but are not limited to, any of the following scenarios: logging into a designated platform on the first terminal device or the second terminal device, connecting the first terminal device to the second terminal device, or connecting the second terminal device to the first terminal device.
[0046] In one possible design, the first voice command and the second voice command are based on the same voice command from the user, and are received by the first terminal device and the second terminal device, respectively.
[0047] Fourthly, embodiments of this application also provide a terminal device, including: one or more processors; one or more memories; the one or more memories being used to store one or more computer programs and data information; wherein the one or more computer programs include instructions; when the instructions are executed by the one or more processors, the terminal device causes the terminal device to perform the method described in any of the possible designs in the second aspect above.
[0048] Fifthly, embodiments of this application also provide a server device, including: one or more processors; one or more memories; the one or more memories being used to store one or more computer programs and data information; wherein the one or more computer programs include instructions; when the instructions are executed by the one or more processors, the server device causes the server device to perform the method described in any of the possible designs in the third aspect above.
[0049] Sixthly, embodiments of this application also provide a multi-device voice control system, including at least two terminal devices as described in the fourth aspect above, and a server device as described in the fifth aspect above.
[0050] In a seventh aspect, embodiments of this application provide a computer-readable storage medium storing a computer program (also referred to as code or instructions) that, when run on a computer, causes the computer to perform the method in any of the possible designs in the second or third aspects described above.
[0051] Eighthly, embodiments of this application provide a computer program product, which includes a computer program (also referred to as code or instructions) that, when run, causes a computer to perform the method in any of the possible designs in the second or third aspects described above.
[0052] In a ninth aspect, embodiments of this application also provide a graphical user interface on a terminal device, the terminal device having a display screen, one or more memories, and one or more processors, the one or more processors being configured to execute one or more computer programs stored in the one or more memories, the graphical user interface including a graphical user interface displayed when the terminal device executes any possible design of the second aspect of embodiments of this application.
[0053] For details on the beneficial effects of any of the second to ninth aspects mentioned above, please refer to the beneficial effects of the various possible designs in the first aspect mentioned above; they will not be repeated here. Attached Figure Description
[0054] Figure 1a This is a schematic diagram illustrating an application scenario for multi-device management provided in an embodiment of this application.
[0055] Figure 1b The corresponding embodiment provided in this application Figure 1a The flowchart illustrating the application scenario is shown.
[0056] Figure 2 This is a schematic diagram of the hardware structure of a possible terminal device provided in an embodiment of this application;
[0057] Figure 3 A software structure block diagram of a terminal device provided in an embodiment of this application;
[0058] Figure 4 This is one of the application scenario diagrams of a multi-device voice control method provided in the embodiments of this application;
[0059] Figure 5 This is the second application scenario diagram of a multi-device voice control method provided in the embodiments of this application;
[0060] Figure 6 This is one of the interactive flow diagrams of a multi-device voice control method provided in an embodiment of this application;
[0061] Figure 7 A flowchart illustrating a multi-device voice control method provided in an embodiment of this application;
[0062] Figure 8a A second schematic diagram of the interaction flow of a multi-device voice control method provided in an embodiment of this application;
[0063] Figure 8b A second schematic diagram of the interaction flow of a multi-device voice control method provided in an embodiment of this application;
[0064] Figure 9 This is the third schematic diagram of the interaction flow of a multi-device voice control method provided in an embodiment of this application. Detailed Implementation
[0065] With the rapid development of society, terminal devices are taking on increasingly diverse forms, such as mobile phones, tablets, and televisions; and they are becoming increasingly ubiquitous. Terminal devices not only have communication functions but also powerful processing capabilities, storage capacity, and camera functions. Through an operating system, terminal devices execute corresponding applications, allowing users to make calls, send text messages, browse the web, watch videos, and more. Furthermore, to facilitate collaborative business processing across different terminal devices, various methods exist for managing multiple devices, such as multiple devices logging into the same account and screen mirroring between devices.
[0066] Based on the background information, in related technologies, in scenarios where multiple devices need to collaborate on business processing, manual operation by the user is usually required.
[0067] For example, see Figure 1a This is a schematic diagram illustrating an application scenario for multi-device management provided in an embodiment of this application. In this scenario, it is assumed that a user is... Figure 1a The mobile phone shown in (a) has an application (APP) installed and a user account is logged in. If the same user account needs to be logged into on another terminal device using the same APP, it can typically be done by using the device already logged in to scan a QR code on the non-logged-in device, thus enabling the logged-in device to authorize login to the non-logged-in device. For example... Figure 1a As shown, users can manually create a QR code on the tablet in (b) to request authorization for login; then, using the mobile phone shown in (a), they can scan the code to authorize login to devices that are not currently logged in.
[0068] based on Figure 1a The scenario shown below is explained through... Figure 1b The flowchart shown illustrates the specific implementation process.
[0069] S101. The backend server corresponding to the unlogged-in device requests authorization login from the APP development platform.
[0070] S102, QR code returned by the APP development platform.
[0071] S103. The backend server corresponding to the unlogged-in device controls the display of the QR code on the unlogged-in device.
[0072] S104. The user scans the QR code using a logged-in device. This process is understood to be a manual operation by the user.
[0073] S105. The user authorizes login through the already logged-in device and instructs the APP development platform.
[0074] S106. The APP development platform informs that the backend server corresponding to the unlogged-in device has been authorized.
[0075] S107. The backend server corresponding to the unlogged-in device requests user account data from the APP development platform.
[0076] S108. The APP development platform returns user account data to the backend server corresponding to the unlogged-in device.
[0077] Through the above implementation process, it can be seen that in this scenario, the user needs to manually operate the device to complete the authorization login for the user to log in to the non-logged-in device through the logged-in device.
[0078] In view of this, embodiments of this application provide a multi-device voice control system and method, which can associate multiple nearby devices through voice and simultaneously realize the collaborative management of multiple nearby devices to complete business tasks that require collaborative processing by multiple devices. The main design concept is that multiple nearby devices simultaneously acquire the user's voice commands and send voice request information, including but not limited to the voice commands, to a server-side device. The server-side device can associate multiple nearby devices that report the same voice command based on the received voice request information from multiple devices and generate corresponding control commands for each device. Therefore, the system or method provided by this application has the characteristics of simple operation and more convenient interaction. Here, "nearby devices" refers to multiple terminal devices that can simultaneously receive the same voice command from the user, such as a mobile phone and a television in the same room, or a computer and a mobile phone on the same desktop.
[0079] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0080] It is understood that the terminal device in this application embodiment can be a device with voice command input capability, such as a smart home device (e.g., a smart TV, a smart screen, a smart speaker, etc.), a mobile phone, a tablet computer, a wearable device (e.g., a watch, a helmet, headphones, etc.), an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. It is understood that this application embodiment does not impose any limitations on the specific type of terminal device.
[0081] The terminal devices to which this application's embodiments can be applied include, but are not limited to, those equipped with... Alternatively, it could be a portable terminal device with another operating system. The aforementioned portable terminal device could also be other portable terminal devices, such as a laptop computer with a touch-sensitive surface (e.g., a touch panel).
[0082] Figure 2 A schematic diagram of a possible hardware structure for a terminal device is shown. The terminal device 200 includes components such as a radio frequency (RF) circuit 210, a power supply 220, a processor 230, a memory 240, an input unit 250, a display unit 260, an audio circuit 270, a communication interface 280, and a wireless-fidelity (Wi-Fi) module 290. Those skilled in the art will understand that... Figure 2 The hardware structure of the terminal device 200 shown in the figure does not constitute a limitation on the terminal device 200. The terminal device 200 provided in the embodiments of this application may include more or fewer components than shown, may combine two or more components, or may have different component configurations. Figure 2 The various components shown can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.
[0083] The following is combined Figure 2 The various components of the terminal device 200 are described in detail below:
[0084] The RF circuit 210 can be used for receiving and transmitting data during communication or a call. Specifically, after receiving downlink data from the base station, the RF circuit 210 sends it to the processor 230 for processing; additionally, it sends uplink data to be transmitted to the base station. Typically, the RF circuit 210 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc.
[0085] Furthermore, the RF circuit 210 can also communicate with other devices via a wireless communication network. The wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, and Short Messaging Service (SMS).
[0086] Wi-Fi technology is a short-range wireless transmission technology. The terminal device 200 can connect to an access point (AP) via the Wi-Fi module 290, thereby enabling access to the data network. The Wi-Fi module 290 can be used for receiving and sending data during communication.
[0087] The terminal device 200 can physically connect to other devices through the communication interface 280. Optionally, the communication interface 280 can be connected to the communication interfaces of other devices via a cable to enable data transmission between the terminal device 200 and other devices.
[0088] Since the terminal device 200 in this embodiment can implement communication services and interact with server-side devices (e.g., including but not limited to voice service servers, account servers, etc.), the terminal device 200 needs to have data transmission capabilities, that is, the terminal device 200 needs to include a communication module. Although Figure 2The RF circuit 210, the Wi-Fi module 290, and the communication interface 280 are shown, but it is understood that the terminal device 200 contains at least one of the above components or other communication modules (such as a Bluetooth module) for data transmission.
[0089] For example, when the terminal device 200 is a mobile phone, the terminal device 200 may include the RF circuit 210, the Wi-Fi module 290, or a Bluetooth module. Figure 2 (Not shown in the image); when the terminal device 200 is a computer, the terminal device 200 may include the communication interface 280, and may also include the Wi-Fi module 290, or may include a Bluetooth module (not shown in the image); Figure 2 (Not shown in the image); when the terminal device 200 is a tablet computer, the terminal device 200 may include the Wi-Fi module, or may include a Bluetooth module (not shown in the image); Figure 2 (Not shown in the image).
[0090] The memory 240 can be used to store software programs and modules. The processor 230 executes various functional applications and data processing of the terminal device 200 by running the software programs and modules stored in the memory 240. Optionally, the memory 240 may mainly include a program storage area and a data storage area. The program storage area may store the operating system (mainly including the software programs or modules corresponding to the kernel layer, system layer, application framework layer, and application layer).
[0091] In addition, the memory 240 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0092] The input unit 250 can be used to receive editing operations on various types of data objects, such as numbers or characters, input by the user, and to generate key signal inputs related to user settings and function control of the terminal device 200. Optionally, the input unit 250 may include a touch panel 251 and other input devices 252.
[0093] The touch panel 251, also known as a touch screen, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel 251), and drive the corresponding connection device according to a pre-set program.
[0094] Optionally, the other input device 252 may include, but is not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.
[0095] The display unit 260 can be used to display information input by the user or information provided to the user, as well as various menus of the terminal device 200. The display unit 260 is the display system of the terminal device 200, used to present the interface and realize human-computer interaction. The display unit 260 may include a display panel 261. Optionally, the display panel 261 can be configured using a liquid crystal display (LCD), organic light-emitting diode (OLED), or other similar forms. In this embodiment, the display unit 260 may not be provided on the terminal device; for example, a smart speaker device does not require a display screen. Alternatively, a display unit 260 may be provided on the terminal device, and the display unit 260 displays the content corresponding to the voice command received by the terminal device 200 through the microphone 271. For example, if the voice command received by the microphone 271 is "Open and log in to instant messaging application A", the display interface of the corresponding instant messaging application A can be displayed on the display panel 261.
[0096] The processor 230 is the control center of the terminal device 200. It connects various components via various interfaces and lines, and executes software programs and / or modules stored in the memory 240, as well as calling data stored in the memory 240, to perform various functions and process data of the terminal device 200, thereby realizing multiple services based on the terminal device 200. In this embodiment, the processor 230 is used to implement the method provided in this embodiment, thus providing a technical solution for simultaneously controlling multiple nearby devices via voice, thereby reducing the operational complexity in scenarios where multiple devices are processing services.
[0097] The terminal device 200 also includes a power supply 220 (such as a battery) for supplying power to various components. Optionally, the power supply 220 can be logically connected to the processor 230 through a power management system, thereby enabling the power management system to manage functions such as charging, discharging, and power consumption.
[0098] like Figure 2As shown, the terminal device 200 also includes an audio circuit 270, a microphone 271, and a speaker 272, providing an audio interface between the user and the terminal device 200. The audio circuit 270 converts audio data into signals recognizable by the speaker 272 and transmits the signals to the speaker 272, where the speaker 272 converts them into sound signals for output. The microphone 271 collects external sound signals (such as human speech or other sounds) and converts the collected external sound signals into signals recognizable by the audio circuit 270, sending them to the audio circuit 270. The audio circuit 270 can also convert the signals sent by the microphone 271 into audio data, then output the audio data to the RF circuit 210 for transmission to, for example, another terminal device, or output the audio data to the memory 240 for further processing. In this embodiment, the trigger scenario for the microphone 271 to collect external sound signals can be triggered by the user clicking a voice input control (such as a smart assistant or voice assistant) on the display interface of the terminal device 200, or by the user waking up the device using a preset wake-up word; this application does not limit this.
[0099] Although not shown, the terminal device 200 may also include at least one sensor, a camera, etc., which will not be described in detail here. The at least one sensor may include, but is not limited to, a pressure sensor, a barometric pressure sensor, an accelerometer, a distance sensor, a fingerprint sensor, a touch sensor, a temperature sensor, etc.
[0100] The operating system (OS) involved in this application embodiment is the most basic system software running on the terminal device 200. Taking a mobile phone as an example, the operating system can be HarmonyOS, Android, or iOS. The software system of the terminal device 200 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses an operating system adopting a layered architecture as an example to illustrate the software structure of the terminal device 200.
[0101] Figure 3 This is a software structure block diagram of a terminal device provided in an embodiment of this application. For example... Figure 3 As shown, the software architecture of a terminal device can be a layered architecture. For example, the software can be divided into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the operating system is divided into five layers, from top to bottom: the application layer, the application framework layer (framework, FWK), the runtime and system libraries, the kernel layer, and the hardware layer.
[0102] The application layer can include a series of application packages. For example... Figure 3As shown, the application layer can include camera, settings, skin modules, user interface (UI), third-party applications, etc. Third-party applications can include WLAN, music, call, Bluetooth, video, etc.
[0103] In some embodiments of this application, the application layer can be used to implement the presentation of the editing interface, which can be used by the user to view or perform operations. For example, if the mobile phone includes a display panel 261, the user can display the relevant interface of the instant messaging application on the main interface displayed on the display panel 261.
[0104] In one possible implementation, the application can be developed using Java, by calling the application programming interface (API) provided by the application framework layer. Developers can then interact with the underlying operating system layers (such as the hardware layer and kernel layer) to develop their own applications. This application framework layer primarily consists of a series of services and management systems within the operating system.
[0105] The application framework layer provides application programming interfaces and a programming framework for applications within the application layer. The application framework layer includes some predefined functions. For example... Figure 3 As shown, the application framework layer may include a shortcut icon management module, a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, etc.
[0106] The shortcut icon management module is used to manage the shortcut icons displayed on the terminal device, such as creating shortcut icons, removing shortcut icons, and monitoring whether shortcut icons meet the display conditions.
[0107] The window manager is used to manage windowed applications. It can obtain the screen size, determine if a status bar is present, lock the screen, and capture screenshots. The content provider stores and retrieves data, making this data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.
[0108] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0109] A phone manager is used to provide communication functions for terminal devices. For example, it manages call status (including connection and disconnection).
[0110] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0111] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of download completion or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating the device, and flashing indicator lights.
[0112] In some embodiments of this application, the application framework layer is mainly responsible for calling the service interface for communication with the hardware layer to pass the user's operation request to the hardware layer. The operation request may include the user's operation request to open or log in to a certain APP through voice command.
[0113] The runtime includes the core libraries and the virtual machine. The runtime is responsible for the scheduling and management of the operating system.
[0114] The core library consists of two parts: one part contains the functionalities that the Java language needs to call, and the other part contains the core libraries of the operating system. The application layer and application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0115] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0116] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.
[0117] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0118] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0119] A 2D graphics engine is a graphics engine for 2D drawing.
[0120] In some embodiments, a 3D graphics processing library can be used to draw 3D motion trajectory images, and a 2D graphics engine can be used to draw 2D motion trajectory images.
[0121] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.
[0122] The hardware layer can include various types of sensors, such as accelerometers, gyroscopes, and touch sensors.
[0123] Typically, a terminal device 200 can run multiple applications simultaneously. In a simpler scenario, one application corresponds to one process; in a more complex scenario, one application can correspond to multiple processes. Each process has a unique process ID.
[0124] In combination with the above Figure 2 The text describes the hardware structure of the terminal device, and... Figure 3 The software framework of the terminal device is described below. With reference to multiple embodiments and accompanying drawings, the working principle of the software and hardware of the terminal device executing the multi-device voice control method proposed in the embodiments of this application is illustrated.
[0125] It should be understood that in the embodiments of this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, a and b, a and c, b and c, or a, b, and c, where a, b, and c can be single or multiple.
[0126] The multiple instances mentioned in the embodiments of this application refer to two or more.
[0127] In addition, it should be understood that in the description of this application, the words "first" and "second" are used only for the purpose of distinguishing descriptions and should not be construed as indicating or implying relative importance or order.
[0128] Furthermore, in the embodiments of this application, "terminal device" and "device" can be used interchangeably, that is, to refer to various devices that can be used to implement the embodiments of this application; "application" and "application program" in the embodiments of this application can also be used interchangeably, both referring to programs or clients with certain service provision capabilities, that is, applications and clients can also be used interchangeably, for example, video clients and instant messaging clients can also be called video applications or instant messaging applications, etc.
[0129] It should be understood that the hardware structure of a terminal device can be as follows: Figure 2 As shown, the software architecture can be as follows: Figure 3 As shown, the software programs and / or modules corresponding to the software architecture in the terminal device can be stored in the memory 240, and the processor 230 can run the software programs and applications stored in the memory 240 to execute the flow of a multi-device voice control method provided in the embodiments of this application.
[0130] To facilitate understanding of the multi-device voice control method provided in this application, the following is combined with... Figures 4 to 9 The content shown describes the implementation process of the method provided in this application.
[0131] This application's embodiments are applicable to application scenarios requiring multi-device collaborative business processing. First, the application scenarios to which this application's embodiments are applicable are illustrated through the following examples. It should be understood that this application is not limited to the following application scenarios.
[0132] In one possible application scenario, such as Figure 1a As shown, in scenarios where a user needs to log in to the same user account on a tablet using a mobile phone already logged in, related technologies typically require the user to use their mobile phone to scan a QR code to log in to the tablet. In this application, the user can control both the mobile phone and tablet simultaneously using voice commands, thus enabling different controls on the two devices using the same voice commands, and ultimately allowing the user to log in to the same user account on the tablet as on the mobile phone.
[0133] See Figure 4 This is an application scenario diagram of a multi-device voice control method provided in an embodiment of this application. Figure 4 The mobile phone shown in (a) has already been logged in. The WeChat account (i.e., the mobile phone is a logged-in device); while in such Figure 4 In the example shown in (b), the tablet computer (also referred to as "tablet") is not logged into a WeChat account (i.e., the tablet is an unlogged-in device). In this scenario, the phone and tablet can simultaneously or almost simultaneously receive the user's voice commands, such as... Figure 4The command is "Log in to your WeChat account on your tablet." Then, both the phone and tablet can upload the voice request information containing this voice command to a first server device (such as a voice service server) for processing. This voice request information may include, but is not limited to, timestamp information, voiceprint information, and terminal device status information.
[0134] After receiving voice request information from multiple terminal devices, the voice service server can analyze the voice request information uploaded by each terminal device. For example, based on the voice commands uploaded by mobile phones and tablets, it can determine whether the voice request information uploaded by mobile phones and tablets is related. For instance, if the analysis shows that the voice commands uploaded by mobile phones and tablets are the same, or the similarity is greater than a specified threshold, or they correspond, then the voice request information uploaded by mobile phones and tablets is determined to be related. In addition, it can also combine timestamp information and voiceprint information for further precise judgment.
[0135] Furthermore, the voice service server can control the mobile phone and tablet based on semantic analysis of the voice commands uploaded by the mobile phone and tablet. In one optional example, the voice service server can generate a first control command corresponding to the mobile phone and a second control command corresponding to the tablet; the first control command may instruct the mobile phone to initiate an authorization request to the account server, and the second control command may instruct the tablet to initiate a login request to a second server device (such as the account server). Thus, after receiving the login request from the tablet and the authorization request from the mobile phone, the account server can respond to the tablet's login request based on the mobile phone's authorization request, thereby enabling WeChat login on the tablet. In another optional example, the voice service server can also generate a first control command corresponding to the mobile phone. In this case, the first control command may include the tablet's device identifier and instruct the mobile phone to send authorization information containing the user account and key to the tablet based on the tablet's device identifier. Thus, after receiving the authorization information from the mobile phone, the tablet can directly log in using the user account and key.
[0136] Another possible application scenario, besides requiring account login via an account server, is screen projection, such as projecting the phone's display onto a television. Related technologies typically require the phone and television to be connected to the same local area network (LAN), and the user must manually operate a screen mirroring control on the phone to project the phone's interface onto the television. In this application, the phone and television do not need to be connected to the same LAN or be in a connected state. The user can control both the phone and television simultaneously via voice commands, allowing for different controls on the phone and television using the same voice commands, thus displaying the phone's interface on the television.
[0137] See Figure 5This diagram illustrates an application scenario of a multi-device voice control method provided in this application. In this scenario, the mobile phone and the television can simultaneously acquire the user's voice commands, such as... Figure 5 The user's voice command is "cast your phone screen to the TV," such as... Figure 5 (a) shows the mobile phone receiving the user's voice command "cast the phone screen to the TV", and as shown in Figure (a). Figure 5 The television shown in (b) also receives the user's voice command "cast your phone screen onto the TV". Then, the phone and the television can each upload the voice request information containing the voice command to a server device (such as a voice service server) for processing; the voice request information may also include, but is not limited to, timestamp information and voiceprint information.
[0138] After receiving voice request information from multiple terminal devices, the voice service server can analyze the voice request information uploaded by each terminal device. For example, based on the voice commands uploaded by mobile phones and televisions, it can determine whether the voice request information uploaded by mobile phones and televisions is related. For instance, if the analysis shows that the voice commands uploaded by mobile phones and televisions are the same, or the similarity is greater than a specified threshold, or they correspond, then the voice request information uploaded by mobile phones and televisions is determined to be related. In addition, it can also combine timestamp information and voiceprint information for further precise judgment.
[0139] Furthermore, the voice service server can control the mobile phone and television based on semantic analysis of the voice commands uploaded by the mobile phone and television. In one optional example, the voice service server can generate control commands corresponding to the mobile phone, such as sending a screen mirroring command to the mobile phone; wherein, the screen mirroring command may include, but is not limited to, the television's device identifier (e.g., access address). Thus, after receiving the screen mirroring command sent by the voice service server, the mobile phone can connect to the television based on the television's access address, without requiring the mobile phone and television to be connected to the same local area network. In this implementation, authentication of the mobile phone and television can be achieved based on the user's voice commands, thereby further enabling the projection of the content corresponding to the mobile phone's display interface onto the television for display. In another optional example, the voice service server can also generate control commands corresponding to the television, such as sending a screen mirroring acceptance command to the television; wherein, the screen mirroring acceptance command may include, but is not limited to, the mobile phone's device identifier, and an instruction to the television to instruct the mobile phone to send the data content of the display page to the television after connecting with the mobile phone.
[0140] It is understood that, in the implementation of this application, the control method of the voice service server on the two terminal devices is not limited after determining that the voice request information of the two terminal devices is related. In actual implementation, control instructions corresponding to the two terminal devices or any one of the terminal devices can be generated according to the semantic analysis results of the voice instructions in order to realize the user intent corresponding to the user's voice instructions.
[0141] Furthermore, this application does not limit the number of times or the method of sending the first control command from the server device to the first terminal device or the second control command from the second terminal device. The control command can be sent once or multiple times. If multiple control commands are required to realize the user's intent corresponding to the user's voice command, the voice service server can send multiple control commands. For example, if the terminal device cannot correctly receive the control command when the voice service server sends it for the first time, it can wait for a preset time and then send it again. As another example, if the user's intent requires periodic control from the voice service server, the voice service server can send a control command to the terminal device each time the period arrives.
[0142] Based on combination Figure 4 and Figure 5 The illustrated content describes the possible application scenarios to which the embodiments of this application may be applicable. The design concept of the method provided in this application is that multiple neighboring devices can receive the same (or nearly the same or corresponding) voice commands, which are then uploaded by each neighboring device to the server device. The server device can then generate different control commands for different neighboring devices. The interaction process of the method provided in this application is described in detail below.
[0143] See Figure 6 This is a schematic diagram of the interaction flow of a multi-device voice control method provided in an embodiment of this application. It should be noted that this embodiment uses a first terminal device and a second terminal device as examples. In implementation of this application, the type and number of terminal devices are not limited; in specific implementations, more terminal devices may be included in the voice control. If more terminal devices are included, the interaction flow of each terminal device can be referred to the implementation process of the first terminal device or the second terminal device. The interaction flow includes:
[0144] Step 601a: The first terminal device receives the input of the first voice command.
[0145] Step 601b: The second terminal device receives the input of the second voice command.
[0146] The first and second voice commands can be derived from the same user voice command and are received by the first and second terminal devices respectively, thus enabling the server device to recognize that the voice request information uploaded by the first and second terminal devices is related. Combined with... Figure 4 The application scenario shown illustrates that the user's voice command corresponding to the first and second voice commands can be "Log in to your WeChat account on the tablet," where the first terminal device can be a mobile phone and the second terminal device can be a tablet computer. Combined with... Figure 5 In the application scenario shown, the user voice command corresponding to the first voice command and the second voice command can be "cast your phone screen onto the TV". The first terminal device can be a mobile phone, and the second terminal device can be a TV.
[0147] For example, before receiving the user's voice command input, the first terminal device or the second terminal device may have been woken up by a wake word, or by a specified control or gesture on the terminal device. This application embodiment does not limit the wake-up process of the terminal device.
[0148] In specific implementations of this application, the execution order of steps 601a and 601b is not limited. Optionally, steps 601a and 601b can be executed simultaneously; for example, the user can first wake up the first terminal device and the second terminal device, so that the first terminal device and the second terminal device can simultaneously receive the user's voice command input. It can also be understood that the first voice command and the second voice command are the input of the same user voice command on two different terminal devices, that is, the first voice command and the second voice command originate from the same user voice command. Alternatively, step 601a can be executed before step 601b, or step 601b can be executed before step 602a. It should be noted that, in implementation, the time difference between steps 601a and 601b can be less than a specified time threshold, and the first voice command received by the first terminal device and the second terminal device must be the same (or nearly the same). For example, the user can first wake up the first terminal device and output the first voice command "cast the phone screen onto the TV," so that the first terminal device receives the user's voice command. Then, the user can wake up the second terminal device and output the second voice command "cast the phone screen onto the TV," so that the second terminal device also receives the user's voice command. In this way, after the first and second terminal devices upload the voice request information containing the voice command to the first server device, the first server device can determine that the first and second terminal devices have an association in receiving the same user voice command. It is understandable that executing steps 601a and 601b simultaneously can achieve a more accurate voice control effect across multiple devices. If executed step by step, although the commands originate from the same user, the differences are significant because they are implemented through the same content but at different times. Therefore, the first server device can set a lower threshold when judging similarity.
[0149] Step 602a: The first terminal device uploads the first voice request information to the first server device, wherein the first voice request information includes the first voice command.
[0150] Step 602b: The second terminal device uploads the second voice request information to the first server device, wherein the second voice request information includes the second voice command.
[0151] In implementation of this application, the first terminal device and the second terminal device may also carry other information in the voice request information (first voice request information or second voice request information) according to the actual application scenario. For example, the voice request information may also include, but is not limited to: timestamp information, voiceprint information, and terminal device status information.
[0152] 1) The timestamp information can be used to identify the time when the terminal device receives the first voice command, so that the first server device can combine the timestamp information to determine whether the first terminal device and the second terminal device receive the same voice command in the same application scenario.
[0153] 2) Voiceprint information can be used to identify user identity information, which in turn makes it easier for the first server device to combine the voiceprint information to determine whether the voice commands received by the first terminal device and the second terminal device come from the same person.
[0154] 3) Terminal device status information can be used, but is not limited to, to identify the account login status on a terminal device. For example, it can identify whether a WeChat account is logged in on the first terminal device and not logged in on the second terminal device. This allows the first server device to combine the terminal device status information to determine the role of each terminal device (for example, it can be considered that...). Figure 4 The mobile phone shown is the "source device" and the tablet is the "target device"; for example, it can be considered that... Figure 5 The mobile phone shown is the "initiator" and the TV is the "receiver" (etc.). Then, different control commands can be generated according to the roles of each terminal device.
[0155] For example, referring to Tables 1a and 1b below, the first server device (voice service server) receives voice request information from the first terminal device (mobile phone), the second terminal device (tablet computer), and the television, respectively. Table 1a compares the voice request information from the mobile phone and the tablet computer, and Table 1b compares the voice request information from the mobile phone and the television, as follows:
[0156] Table 1a
[0157]
[0158] Based on the information shown in Table 1a above, the voice service server compares the voice request information from mobile phones and tablets. If the voice commands, timestamp information, and voiceprint information are the same, or if the similarity is greater than a specified threshold, then the voice request information from the mobile phone and tablet is determined to be related. A higher similarity indicates a greater probability that the two voice request information are related; similarity indicates that the two voice request information are related. Furthermore, the voice service server can further analyze the semantics of the voice commands and generate a first control command for the mobile phone and / or a second control command for the tablet based on the analysis results.
[0159] It should be noted that determining whether voice commands, timestamp information, and voiceprint information are identical or have a similarity greater than a specified threshold can be implemented by separately determining whether the first voice command and the second voice command are identical or have a similarity greater than a first specified threshold, whether the timestamp information uploaded by the first terminal device and the timestamp information uploaded by the second terminal device are identical or have a similarity greater than a second specified threshold, and whether the similarity between the voiceprint information uploaded by the first terminal device and the voiceprint information uploaded by the second terminal device is greater than a third specified threshold. Specifically, determining the similarity of voice commands can be implemented, for example, by judging the time-domain parameters and frequency-domain parameters of the audio corresponding to the voice command; a smaller distance indicates a greater similarity between the voice commands uploaded by the first and second terminal devices. Determining the similarity of timestamp information can be implemented, for example, by judging the time difference; a smaller time difference indicates a greater similarity between the voice commands uploaded by the first and second terminal devices. Determining the similarity of voiceprint information can be implemented, for example, by extracting voiceprint features based on artificial intelligence technology and then comparing the similarity of the voiceprint features.
[0160] Furthermore, if a user uses voice commands to control both the first and second terminal devices sequentially to achieve the same user intent, the user's voice commands do not need to be exactly the same, but rather correspond to each other. Combined with... Figure 5 For example, a user's voice command to a mobile phone might be "cast the phone screen onto the TV," while a voice command to the TV might be "display the phone screen on the TV." Although the two voice commands are not identical and not very similar, they correspond to the same user. Therefore, the first voice command received by the mobile phone and the second voice command received by the TV correspond. In other words, the voice service server can determine that the mobile phone is the "initiator," the TV is the "receiver," and the intent is screen casting based on the mobile phone's command "cast the phone screen onto the TV." Similarly, it can determine that the mobile phone is the "initiator," the TV is the "receiver," and the intent is screen casting based on the TV's command "display the phone screen on the TV." Therefore, the voice commands from the mobile phone and the tablet are corresponding.
[0161] Table 1b
[0162]
[0163] According to Table 1b above, the voice service server compares the voice request information from mobile phones and tablets. If the similarity of the voice command is not greater than a specified threshold, or the similarity of the timestamp information is not greater than a specified threshold, or the similarity of the voiceprint information is not greater than a specified threshold, then the voice request information from the mobile phone and the TV is determined to be unrelated.
[0164] Furthermore, the first server device can typically receive multiple voice request messages. In determining which voice request messages belong to the same application scenario, the first server device can make judgments one by one based on the information contained in the voice request messages. For example, the first server device can first determine relevance based on the similarity of timestamp information between terminal devices. If the similarity is not greater than a specified threshold, it can continue to judge other information such as voiceprint information; otherwise, it can be determined that they are irrelevant. And, if it is finally determined that the similarity of voice commands between terminal devices is not greater than a specified threshold, then the voice request messages uploaded by the terminal devices can be determined to be relevant. It should be noted that, in the implementation of this application, the order in which the first server device makes judgments is not limited. In this way, by filtering or excluding some information in the voice request messages, the processing efficiency of determining whether the voice request messages uploaded by terminal devices are relevant can be improved.
[0165] Step 603: If the first voice request information and the second voice request information are determined to be related based on the first voice command and the second voice command, the first server device performs at least one of the following steps 604a and 604b.
[0166] For example, after receiving voice commands from various terminal devices, the first server device performs a similarity comparison as shown in Tables 1a and 1b to determine the relevant voice request information, i.e., multiple terminal devices receiving the same voice command in the same application scenario. Simultaneously, it performs semantic analysis and slot processing. This can be implemented by first recognizing the voice commands from each terminal device as text content, and then performing user intent understanding or slot parsing based on the obtained text content to determine the role and intent of each terminal device in the application scenario. For instance, based on the first voice request information from the mobile phone and the second voice request information from the tablet, the first server device can determine that multiple terminal devices receiving the same voice command in the same application scenario are receiving the same voice command. Based on this, the first server device can further determine, based on the semantic analysis of a voice command such as "Log in to WeChat account on the tablet," that the mobile phone is the "source device," the tablet is the "target device," and the intent is "Log in to WeChat account" in this application scenario. Here, the first server device can be, for example, a voice service server, mainly used to process the voice request information uploaded by each terminal device.
[0167] Furthermore, the first server device can generate a first control command for the first terminal device and / or a second control command for the corresponding second terminal device based on the determined roles and intentions of each terminal device in the relevant situation.
[0168] Step 604a: The first server device generates and sends the first control command to the first terminal device. The first control command is used to execute the related operations of the first voice command; the related operations of the first voice command are used for collaborative processing of services with the second terminal device. For example, the related operations of the first voice command can be determined by the first server device based on the voice analysis results of the first voice command. For instance, in an application scenario requiring WeChat account login, the first server device can also obtain the terminal device status information (i.e., the login status of the WeChat account) from the voice request information of each terminal device. Figure 4 If the mobile phone shown has a WeChat account logged in, but the tablet does not, then the relevant operation of the first voice command can be to generate a first control command that instructs the mobile phone to initiate an authorization request to the account server and a second control command that instructs the tablet to initiate a login request to the account server.
[0169] Step 604b: The first server device generates and sends the second control command to the second terminal device. The second control command is used to execute the related operations of the second voice command; the related operations of the second voice command are used for collaborative processing of services with the first terminal device. Similarly, the related operations of the second voice command can be determined by the voice analysis results of the second voice command by the first server device. For example, in an application scenario that requires screen mirroring from a mobile phone to a TV, the first server device can send a first control command to the mobile phone based on the semantic analysis results of the voice command; wherein, the first control command may contain the device identifier of the tablet, such as the access address (for example, the access address can be a MAC address or IP address, etc.). In this case, the related operations of the second voice command can be to instruct the mobile phone to access the tablet according to the tablet's access address.
[0170] In this way, by using the voice request information from each terminal device, the first server device can control each terminal device to work together to complete business processing, thereby reducing the complexity of operation for users.
[0171] It is understandable that, based on step 603, it can be determined whether to execute step 604a, step 604b, or both steps 604a and 604b.
[0172] Furthermore, in application scenarios where a designated platform is logged into on either the first terminal device or the second terminal device, optionally, if the first server device instructs the first and second terminal devices to interact with the second server device, a first identifier (e.g., a voice fingerprint) uniquely identifying the first and second terminal devices can be generated. This can be achieved, for example, through a universally unique identifier (UUID) algorithm. See also... Figure 7 This is a flowchart illustrating a multi-device voice control method provided in an embodiment of this application.
[0173] Step 701a: The first server device obtains the first voice command from the first terminal device.
[0174] Step 701b: The first server device obtains the second voice command from the second terminal device.
[0175] Step 701a can be obtained based on the first voice request information received in step 602a above, and step 701b can be obtained based on the second voice request information received in step 602b above.
[0176] Step 702: The first server device determines whether the first terminal device and the second terminal device are related. This can be implemented by determining that the first voice command of the first terminal device and the first voice command of the second terminal device belong to the same voice input (e.g., the similarity of the voice command, timestamp information, and voiceprint information is greater than a specified threshold), and then proceeding to step 703.
[0177] Step 703: The first server device generates a first identifier code, which is used to identify the first terminal device and the second terminal device.
[0178] When this application is implemented, the first server device can carry the first identification code through the first control command and the second control command. Also, when the first terminal device and the second terminal device interact with another server, they can carry the first identification code.
[0179] In this configuration, the first terminal device and the second terminal device each interact with another server. This can be implemented such that the first terminal device, based on the first control instruction, sends a first request instruction to the second server device (e.g., an account server), whereby the first request instruction is used to request login to a specified platform (e.g., an authorization request); and the second terminal device, based on the second control instruction, sends a second request instruction to the second server device, whereby the second request instruction is used to request authorization to log in to the specified platform (e.g., a login request). For example, as... Figure 4In the scenario shown, the mobile phone can send an authorization request to the account server with the first identification code according to the first control instruction; the tablet can also send a login request to the account server with the first identification code according to the second control instruction.
[0180] In this way, the account server can perform at least one of the following processes based on the first identification code received from the mobile phone and tablet: sending a first response instruction to the first terminal device and sending a second response instruction to the second terminal device; wherein the second response instruction is used to instruct the first terminal device to authorize the second terminal device to log in to the designated platform; for example, sending a second response instruction to the tablet to determine that the authorization request from the mobile phone is used to respond to the login request from the tablet.
[0181] In another optional embodiment, in an application scenario where a user logs into a designated platform on either the first terminal device or the second terminal device, the first server device can also instruct the first terminal device to send the user account information of the designated platform to the second terminal device without requiring interaction from the second server device. Specifically, the processing performed by the first server device after step 603 can be sending a first control command to the first terminal device. Optionally, the first server device can carry the device identifier of the second terminal device in the first control command, so that the first terminal device can send the user account information of the designated platform to the second terminal device based on the device identifier of the second terminal device. The device identifier can be the access address of the second terminal device (e.g., MAC address or IP address).
[0182] In application scenarios where the first terminal device is connected to the second terminal device or the second terminal device is connected to the first terminal device, after the first server device receives voice request information from the mobile phone and the TV respectively, since it is determined that the intention of the application scenario is to project the display screen of the mobile phone onto the TV, the server device can send only the first control command to the mobile phone without sending the second control command to the TV; wherein, the first control command sent to the mobile phone in this application scenario can be an instruction for the access address of the TV.
[0183] To facilitate understanding of the methods provided in the embodiments of this application, the following descriptions are in conjunction with... Figure 4 and Figure 5 The application scenarios shown provide a detailed description of the methods provided in the embodiments of this application.
[0184] See Figure 8a This is another interactive flow diagram of a multi-device voice control method provided in an embodiment of this application. Figure 4 The application scenario shown illustrates the interaction process between the first terminal device (mobile phone), the second terminal device (tablet), and the server-side devices (voice service server and account server), including:
[0185] Step 801a: The user inputs a first voice command into the mobile phone. The first voice command may be as follows: Figure 8a The text shows "Log in to your WeChat account on your tablet".
[0186] Step 801b: The user inputs a second voice command to the tablet.
[0187] Step 802a: The mobile phone sends a first voice request message to the voice service server. The first voice request message may include, but is not limited to: the first voice command, timestamp information, voiceprint information, and terminal device status information (such as the login status of the WeChat account on the mobile phone; if logging into other APP accounts is required, it can be the login status of other APP accounts on the mobile phone).
[0188] Step 802b: The tablet sends a second voice request message to the voice service server. The second voice request message may include, but is not limited to: the second voice command, timestamp information, voiceprint information, and terminal device status information (such as the login status of the WeChat account on the tablet).
[0189] For example, after receiving multiple voice request messages (including the first voice request message and the second voice request message) from multiple terminal devices, the voice service server can determine the relevant terminal devices based on each voice request message, that is, determine whether the voice request messages from the mobile phone and tablet are related. Then, for each group of relevant terminal devices, the voice service server can generate control instructions corresponding to each terminal device, such as a first control instruction and / or a second control instruction corresponding to the mobile phone. In addition, for each group of relevant terminal devices, the voice service server can also generate unique identifiers for multiple terminal devices in that application scenario.
[0190] Step 803a: The voice service server sends the first control command to the mobile phone. For example... Figure 4 The application scenario shown is based on the fact that the WeChat account on the mobile phone is already logged in. The first control instruction can be used to instruct the mobile phone to send an authorization request (first request instruction) to the (third-party) account server. That is, it instructs the mobile phone to notify the account server that manages the WeChat account-related data that it can authorize the tablet to log in.
[0191] Step 803b: The voice service server sends a second control command to the tablet. For example... Figure 4 The application scenario shown is based on the fact that the WeChat account on the tablet is not logged in. The second control command can be used to instruct the tablet to send a login request (second request command) to the (third-party) account server.
[0192] Step 804a: The mobile phone sends an authorization request to the (third-party) account server according to the first control instruction.
[0193] Step 804b: The tablet sends a login request to the (third-party) account server according to the second control instruction.
[0194] Step 805: The (third-party) account server authorizes the tablet to log in based on the authorization request and the login request (second response instruction). Specifically, when the mobile phone sends the authorization request, it may carry a first identifier code indicated by the voice service server; when the tablet sends the login request, it may also carry a first identifier code indicated by the voice service server. Since the first identifier codes are the same for both, the account server can match the current authorization request from the mobile phone with the current login request from the tablet based on the first identifier code.
[0195] See Figure 8b This is another interactive flow diagram of a multi-device voice control method provided in an embodiment of this application. Still as... Figure 4 The illustrated application scenario describes the interaction process between a first terminal device (mobile phone), a second terminal device (tablet), and server-side devices (voice service server and account server). Steps 801a to 802b are related to... Figure 8a The same features shown in the diagram will not be repeated here. Different interaction flows include at least the following:
[0196] Step 803: The voice service server sends a first control command to the mobile phone. The first control command includes a tablet identifier (device identifier of the second terminal device).
[0197] Step 804: The mobile phone sends user account information to the tablet based on the tablet's identifier. This allows the tablet to log in using the user account information sent by the mobile phone. For example, if the mobile phone determines that it is connected to the tablet based on the tablet identifier, the mobile phone can directly send the user account information through the corresponding connection channel; wherein, the connection can be achieved through, but is not limited to, one of the following methods: Bluetooth connection, Wi-Fi Direct connection. Another example: if the mobile phone and tablet are not connected, the mobile phone can send the user account information to the tablet based on the tablet access address indicated by the tablet identifier.
[0198] In the above implementation process, this application can identify scenarios requiring collaborative business processing across multiple devices based on user voice commands, associate the multiple terminal devices involved in the scenario, and generate corresponding control commands for each terminal device, thereby enabling convenient operation based on user voice. Compared to related technologies that require users to use a first terminal device and scan a QR code to authorize login to a second terminal device, this reduces the complexity of user operations.
[0199] See Figure 9 This is another interactive flow diagram of a multi-device voice control method provided in an embodiment of this application. Figure 5 The application scenario shown illustrates the interaction process between the first terminal device (mobile phone), the second terminal device (television), and the server device (voice service server), including:
[0200] Step 901a: The user inputs a first voice command to the mobile phone via voice.
[0201] Step 901b: The user inputs a second voice command to the TV. The first and second voice commands can be as follows: Figure 9 The text shows "casting the screen to the TV".
[0202] Step 902a: The mobile phone sends a first voice request message to the voice service server. The first voice request message may include, but is not limited to, the first voice command, timestamp information, and voiceprint information. It is understood that terminal device status information is not required in this scenario, and the voice request message can be set according to the specific scenario.
[0203] Step 902b: The television sends a second voice request message to the voice service server. The second voice request message may include, but is not limited to: the second voice command, timestamp information, and voiceprint information.
[0204] For example, the voice service server can determine that the voice request information from the mobile phone and the television are related based on their respective information. Therefore, the voice service server can generate a first control command that instructs the mobile phone to specify the television's access address by analyzing the first and second voice request information.
[0205] Step 903: The voice service server sends a first control command to the mobile phone. This first control command may include, but is not limited to, the television's device identifier (such as its access address).
[0206] Step 904: The mobile phone connects to the TV according to the first control command.
[0207] The method provided in this application embodiment enables convenient operation in screen projection scenarios based on user voice. Compared to related technologies, which require the first and second terminal devices to be connected to the same local area network and the user to manually perform screen projection on the first terminal device, this method reduces the user's operational complexity. Furthermore, in this application, it is not required that the first and second terminal devices be connected to the same local area network. The first and second terminal devices can be connected or disconnected. The first server device can authenticate whether the first and second terminal devices are related based on the first voice command uploaded by the first terminal device and the second voice command uploaded by the second terminal device, thereby enabling multi-device collaborative processing.
[0208] Based on the above embodiments, this application also provides a terminal device, which includes multiple functional modules; the multiple functional modules interact to implement the functions performed by the first terminal device or the second terminal device in the methods described in the embodiments of this application. For example, [the following is an example of implementation]. Figure 6 In the illustrated embodiment, the first terminal device performs step 601a, or performs... Figure 6 Step 601b is performed by the second terminal device in the illustrated embodiment. The plurality of functional modules can be implemented based on software, hardware, or a combination of software and hardware, and the plurality of functional modules can be arbitrarily combined or divided based on specific implementations.
[0209] Based on the above embodiments, this application also provides a terminal device, which includes at least one processor and at least one memory. The at least one memory stores computer program instructions. When the terminal device is running, the at least one processor executes the functions performed by the terminal device in the various methods described in the embodiments of this application. For example, when executing... Figure 6 In the illustrated embodiment, the first terminal device performs step 601a, or performs... Figure 6 Step 601b is performed by the second terminal device in the illustrated embodiment.
[0210] Based on the above embodiments, this application also provides a server-side device, which includes multiple functional modules; the multiple functional modules interact to implement the functions performed by the first server-side device or the second server-side device in the methods described in the embodiments of this application. For example, ... Figure 6 Steps 602a to 604b are executed by the first server device in the illustrated embodiment. The plurality of functional modules can be implemented based on software, hardware, or a combination of both, and these modules can be arbitrarily combined or divided based on specific implementations.
[0211] Based on the above embodiments, this application also provides a server-side device, which includes at least one processor and at least one memory. The at least one memory stores computer program instructions. When the server-side device is running, the at least one processor executes the functions performed by the server-side device in the various methods described in the embodiments of this application. For example, when executing... Figure 6 Steps 602a to 604b are performed by the first server device in the illustrated embodiment.
[0212] Based on the above embodiments, this application also provides a multi-device voice control system, which includes at least two terminal devices and a server device; wherein the at least two terminal devices perform collaborative processing of services. For example, the at least two terminal devices can be the first terminal device and the second terminal device in the above embodiments, and the server device can be the first server device and the second server device in the above embodiments.
[0213] Based on the above embodiments, this application also provides a computer program product, which includes a computer program (also referred to as code or instructions) that, when run, causes a computer to perform the methods described in the embodiments of this application.
[0214] Based on the above embodiments, this application also provides a computer-readable storage medium storing a computer program, which, when executed by a computer, causes the computer to perform the methods described in the embodiments of this application.
[0215] Based on the above embodiments, this application also provides a chip for reading computer programs stored in a memory to implement the methods described in the embodiments of this application.
[0216] Based on the above embodiments, this application provides a chip system including a processor for supporting a computer device in implementing the methods described in the embodiments of this application. In one possible design, the chip system further includes a memory for storing necessary programs and data of the computer device. This chip system may be composed of chips or may include chips and other discrete devices.
[0217] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0218] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.
[0219] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0220] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.
[0221] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of protection of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A multi-device voice control system, characterized in that, include: The first terminal device receives and responds to the user's first voice command, and uploads first voice request information to the first server device, wherein the first voice request information includes the first voice command; as well as, The second terminal device receives and responds to the user's second voice command, and uploads second voice request information to the first server device, wherein the second voice request information includes the second voice command; If the first server device determines that the first voice request information and the second voice request information are related based on the first voice command and the second voice command, it performs at least one of the following processes: A first control command is generated and sent to the first terminal device. The first control command is used to execute the related operations of the first voice command. The related operations of the first voice command are used to perform collaborative processing of services with the second terminal device. A second control command is generated and sent to the second terminal device. The second control command is used to execute the related operations of the second voice command. The related operations of the second voice command are used to perform collaborative processing of services with the first terminal device.
2. The system according to claim 1, characterized in that, The first server device determines the correlation between the first voice request information and the second voice request information by at least one of the following methods: Determine that the first voice command and the second voice command are the same; Determine that the similarity between the first voice command and the second voice command is greater than a first specified threshold; Determine that the first voice command and the second voice command correspond.
3. The system according to claim 1 or 2, characterized in that, The first voice request information and the second voice request information each further include at least one of the following: timestamp information, voiceprint information, and terminal device status information.
4. The system according to claim 3, characterized in that, The first server device determines the correlation between the first voice request information and the second voice request information by one or more of the following methods: Determine that the timestamp information uploaded by the first terminal device and the timestamp information uploaded by the second terminal device are the same or have a similarity greater than a second specified threshold. It is determined that the voiceprint information uploaded by the first terminal device is the same as or has a similarity greater than a third specified threshold with respect to the voiceprint information uploaded by the second terminal device.
5. The system according to claim 3, characterized in that, The first server device performs at least one of the following processes based on the first voice command and the second voice command: The first server device performs semantic analysis on the first voice command and the second voice command; and the first server device determines the status of the first terminal device and the second terminal device based on the terminal device status information corresponding to the first terminal device and the terminal device status information corresponding to the second terminal device. The first server device performs at least one of the following processes based on the results of the semantic analysis and the states of the first terminal device and the second terminal device.
6. The system according to claim 1 or 2, characterized in that, If the process performed by the first server device is to generate and send the first control command to the first terminal device, then the first control command contains the device identifier of the second terminal device. The device identifier of the second terminal device is used by the first terminal device to perform the relevant operation of the first voice command based on the device identifier of the second terminal device. or, If the process performed by the first server device is to generate and send a second control command to the second terminal device, then the second control command contains the device identifier of the first terminal device. The device identifier of the first terminal device is used by the second terminal device to perform the relevant operation of the second voice command based on the device identifier of the first terminal device.
7. The system according to claim 1 or 2, characterized in that, Also includes: The first server device generates a first identifier code, which is used to identify the first terminal device and the second terminal device.
8. The system according to claim 7, characterized in that, The system also includes a second server device, wherein: The first terminal device sends a first request instruction to the second server device according to the first control instruction, wherein the first control instruction and the first request instruction carry the first identification code; and, The second terminal device sends a second request instruction to the second server device according to the second control instruction, wherein the second control instruction and the second request instruction carry the first identification code; The second server device performs at least one of the following processes based on the first identifier: sending a first response instruction to the first terminal device and sending a second response instruction to the second terminal device.
9. The system according to claim 8, characterized in that, The first request instruction is used to request login to the specified platform, and the second request instruction is used to request authorization to log in to the specified platform; The process performed by the second terminal device is to send a second response instruction to the first terminal device, the second response instruction being used to instruct the first terminal device to authorize the second terminal device to log in to the designated platform.
10. The system according to claim 1 or 2, characterized in that, The first voice command and the second voice command are used to instruct any of the following scenarios: logging into a designated platform on the first terminal device or the second terminal device, connecting the first terminal device to the second terminal device, or connecting the second terminal device to the first terminal device.
11. The system according to claim 1 or 2, characterized in that, The first voice command and the second voice command are based on the same voice command from the user, and are received by the first terminal device and the second terminal device, respectively.
12. A voice control method for multiple devices, characterized in that, include: The first terminal device receives the user's first voice command; In response to the first voice command, the first terminal device uploads first voice request information to the first server device, the first voice request information including the first voice command; The first terminal device receives a first control command sent by the first server device, the first control command being used to execute related operations of the first voice command; the related operations of the first voice command being used to perform collaborative processing of services with the second terminal device; The first control instruction is generated by the first server device when it determines that the first voice request information is related to the second voice request information uploaded by the second terminal device.
13. The method according to claim 12, characterized in that, The first voice request information and the second voice request information each further include at least one of the following: timestamp information, voiceprint information, and terminal device status information.
14. The method according to claim 12 or 13, characterized in that, The first control command includes the device identifier of the second terminal device. The device identifier of the second terminal device is used by the first terminal device to perform the relevant operations of the first voice command based on the device identifier of the second terminal device.
15. The method according to claim 12 or 13, characterized in that, The first control command includes a first identification code; the first identification code is used to identify the relationship between the first terminal device and the second terminal device.
16. The method according to claim 15, characterized in that, The method further includes: The first terminal device sends a first request instruction to the second server device according to the first control instruction. The first control instruction and the first request instruction carry the first identification code, so that the second server device performs at least one of the following processes according to the first identification code: sending a first response instruction to the first terminal device and sending a second response instruction to the second terminal device.
17. The method according to claim 16, characterized in that, The first request instruction is used to request login to a specified platform; or, the first request instruction is used to request authorization to log in to a specified platform; the method further includes: If the first request instruction is used to request authorization to log in to the specified platform, the first terminal device receives a first response instruction, which instructs the second terminal device to authorize the first terminal device to log in to the specified platform.
18. The method according to claim 12 or 13, characterized in that, The second voice instruction contained in the first voice instruction and the second voice request information is used to indicate any of the following scenarios: logging into a designated platform on the first terminal device or the second terminal device, connecting the first terminal device to the second terminal device, or connecting the second terminal device to the first terminal device.
19. The method according to claim 12 or 13, characterized in that, The second voice command contained in the first voice command and the second voice request information is based on the same voice command from the user, and is received by the first terminal device and the second terminal device respectively.
20. A voice control method for multiple devices, characterized in that, include: The first server device receives a first voice request information uploaded by the first terminal device, the first voice request information containing a first voice command; as well as, The first server device receives a second voice request information uploaded by the second terminal device, the second voice request information containing a second voice command; If the first server device determines that the first voice request information and the second voice request information are related based on the first voice command and the second voice command, it performs at least one of the following processes: A first control command is generated and sent to the first terminal device. The first control command is used to execute the related operations of the first voice command. The related operations of the first voice command are used to perform collaborative processing of services with the second terminal device. A second control command is generated and sent to the second terminal device. The second control command is used to execute the related operations of the second voice command. The related operations of the second voice command are used to perform collaborative processing of services with the first terminal device.
21. The method according to claim 20, characterized in that, The first server device determines the correlation between the first voice request information and the second voice request information by at least one of the following methods: Determine that the first voice command and the second voice command are the same; Determine that the similarity between the first voice command and the second voice command is greater than a first specified threshold; Determine that the first voice command and the second voice command correspond.
22. The method according to claim 20 or 21, characterized in that, The first voice request information and the second voice request information each further include at least one of the following: timestamp information, voiceprint information, and terminal device status information.
23. The method according to claim 22, characterized in that, The first server device determines the correlation between the first voice request information and the second voice request information by one or more of the following methods: Determine that the timestamp information uploaded by the first terminal device and the timestamp information uploaded by the second terminal device are the same or have a similarity greater than a second specified threshold. It is determined that the voiceprint information uploaded by the first terminal device is the same as or has a similarity greater than a third specified threshold with respect to the voiceprint information uploaded by the second terminal device.
24. The method according to claim 22, characterized in that, The first server device performs at least one of the following processes based on the first voice command and the second voice command: The first server device performs semantic analysis on the first voice command and the second voice command; and the first server device determines the status of the first terminal device and the second terminal device based on the terminal device status information corresponding to the first terminal device and the terminal device status information corresponding to the second terminal device. The first server device performs at least one of the following processes based on the results of the semantic analysis and the states of the first terminal device and the second terminal device.
25. A terminal device, characterized in that, include: One or more processors; one or more memories; The one or more memories are used to store one or more computer programs and data information; wherein the one or more computer programs include instructions; When the instructions are executed by the one or more processors, the terminal device performs the method as described in any one of claims 12 to 19.
26. A server-side device, characterized in that, Includes one or more processors; one or more memories; The one or more memories are used to store one or more computer programs and data information; wherein the one or more computer programs include instructions; When the instruction is executed by the one or more processors, the server device performs the method as described in any one of claims 20 to 24.
27. A multi-device voice control system, characterized in that, It includes at least two terminal devices as described in claim 25 and a server device as described in claim 26.
28. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program that, when run on a computer, causes the computer to perform the method as described in any one of claims 12 to 24.
Citation Information
Patent Citations
Voice control method and device, server, terminal equipment and storage medium
CN113127609A
Voice control method of terminal equipment, terminal equipment and server
CN113450792A