Method for voice interaction containing reference word, related apparatus, and communication system
By synchronizing information across devices through a multi-device collaborative system, the problem of electronic devices being unable to respond to voice commands containing pronouns is solved, cross-device information synchronization and service provision are realized, and the convenience and naturalness of voice interaction are improved.
Patent Information
- Application Number
- PCT/CN2025/082432
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-15
- Filing Date
- 2025-03-13
- Publication Date
- 2025-09-18
AI Technical Summary
In the prior art, when an electronic device receives a voice command containing a pronoun, it cannot find the corresponding referent object locally, resulting in an inability to respond to the user's intention and provide corresponding services.
Through the multi-device collaborative system, information is synchronized across devices, the referent of the pronoun is obtained, and the appropriate device is determined to perform the operation, thereby realizing cross-device information synchronization and service provision.
Even if the electronic device does not have the object corresponding to the pronoun stored locally, it can still respond to the user's voice commands, improving the convenience and naturalness of voice interaction and reducing user interaction restrictions.
Smart Images

Figure CN2025082432_18092025_PF_FP_ABST
Abstract
Description
Voice interaction method, related device and communication system including pronouns
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on March 15, 2024, with application number 202410306465.3 and application name “Voice interaction method, related device and communication system containing pronouns”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of terminals, and in particular to a voice interaction method, related devices and communication systems including pronouns. Background Art
[0003] Voice assistants have become popular on more and more smart devices, such as mobile phones, cockpits, and smart screens. Users can describe the services they need through voice and let the device complete the operation quickly. Summary of the Invention
[0004] This application provides a voice interaction method, related device, and communication system containing pronouns. This method can be applied to a communication system containing multiple electronic devices, where the devices in the communication system can understand the user's intentions and respond to voice commands to provide services that meet the user's intentions.
[0005] In a first aspect, a voice interaction method including a pronoun is provided, which is applied to a first device, and includes: obtaining a first voice, the first voice indicating a first operation, the first voice also including a first pronoun that undergoes the first operation; determining a referent object of the first pronoun from a second device; and instructing a third device to perform the first operation on the referent object, the third device being different from the second device.
[0006] In implementing the method of the first aspect, in a communication system, the electronic device where the referent is located and the electronic device that performs the operation in response to the voice command can be different devices. In this way, even if the referent corresponding to the referent does not exist locally on the electronic device, it can still respond to the user's voice command and perform the corresponding operation. This enables cross-device information synchronization between various devices in the communication system, providing services to users across devices.
[0007] Exemplarily, the first device can be determined by the computing power of the device, the operating status of the device, etc. It can also be determined by multiple strategies, including: 1. Determine the first device from multiple electronic devices based on the stability of resources, the information interaction method provided or supported by the device, and one or more factors of the devices commonly used by the user. 2. Determine the preset device among multiple devices as the first device. 3. Determine the device selected by the user as the first device. 4. Determine the first device from multiple devices based on the historical interaction information of each device. In different scenarios, the first device includes different electronic devices. For example, in a navigation scenario, the first device includes a mobile phone. For example, in a smart office scenario, the first device includes a mobile phone and a smart screen. For example, in a smart office scenario, the first device includes a mobile phone and a laptop.
[0008] Exemplarily, after receiving the message querying the reference object, the second device returns a message including the reference object to the first device. In some embodiments, the second device and the first device can be the same device. For example, in a navigation scenario, a smart home scenario, and a smart office scenario, the second device includes a mobile phone.
[0009] Exemplarily, the third device can be determined by the first device through the first operation, and can also be determined by the device that performs the first operation included in the voice instruction. In an embodiment of the present application, the third device and the second device are different devices. In different scenarios, the third device includes different electronic devices. For example, in a navigation scenario, the third device includes a car computer. For example, in a smart home scenario, the third device includes a smart screen. For example, in a smart office scenario, the third device includes a projector.
[0010] In conjunction with the first aspect, the first device may acquire the first voice in the following ways: 1. The first device may acquire the first voice itself, i.e., the first voice is acquired by the first device. 2. The first device may also acquire the first voice acquired by another device, such as when the first device receives the first voice sent by a fourth device, i.e., the first voice is acquired by the fourth device.
[0011] In this way, the first device can collect the first voice, and the fourth device can also collect the first voice, ensuring that the voice command can be collected by the devices in the communication system, thereby ensuring that the first device obtains the first voice.
[0012] For example, the fourth device can determine the first device based on its computing power, operating status, and other factors, and send voice commands to the first device. In different scenarios, the fourth device includes different electronic devices. For example, in a navigation scenario, the fourth device includes a mobile phone and a car computer. For example, in a smart home scenario, the fourth device includes a mobile phone and a smart speaker. For example, in a smart office scenario, the fourth device includes a mobile phone and a laptop computer.
[0013] In combination with the previous embodiment, before the first device receives the first voice sent by the fourth device, the fourth device collects the wake-up word.
[0014] In this way, when the first device and the fourth device are the same device, the first device collects the user's voice command only after receiving the wake-up word, avoiding the first device receiving useless voice and wasting resources. When the first device and the fourth device are different devices, even if the first device does not directly collect the voice command, the first device can still receive voice commands obtained by other devices in the communication system, improving the stability of the solution.
[0015] In combination with the first aspect, the obtaining of the first voice specifically includes: the first device collecting the first voice.
[0016] In combination with the previous embodiment, before the first device collects the first voice, the method further includes: the first device collects a wake-up word.
[0017] In this way, the first device can directly collect the voice command, reducing the occupation of communication resources in the communication system.
[0018] In combination with the first aspect, in all the aforementioned implementation methods of the first aspect, the second device and the first device are the same device.
[0019] In this way, the first device can also query the referenced object locally on the first device, which shortens the time for searching the referenced object and reduces the occupation of communication resources in the communication system.
[0020] In combination with the first aspect, in some implementation methods of the first aspect, the first device does not include the referenced object, and the second device is a different device from the first device.
[0021] In this way, the first device can query the referent object in other devices in the communication system. When the referent object does not exist in the first device, a clear voice instruction can be determined.
[0022] In combination with the first aspect, in all the aforementioned implementation methods of the first aspect, determining the referent object of the first pronoun from the second device specifically includes: the first device sends a first message to the second device; the first device receives a second message sent by the second device, the second message includes information about the referent object, and the referent object is determined by the second device in response to the first message from the storage content and / or displayed interface content of the second device.
[0023] In this way, the reference object that meets the user's intention can be data displayed in the second device, or data pre-stored by the user in the second device. The first device can ensure that the reference object is obtained from the second device.
[0024] In combination with the previous embodiment, the first message includes the type of the object that undergoes the first operation, and the object referred to by the first pronoun belongs to the type of the object that undergoes the first operation.
[0025] In combination with the previous one or two implementation embodiments, the first device sends a first message to the second device, specifically including: the first device sends a first message to multiple devices, the multiple devices including the second device; the first device receives a second message sent by the second device, specifically including: the first device receives a third message sent by some or all of the multiple devices, the third message including information of one or more candidate objects, the one or more candidate objects are determined by the corresponding device from the stored content and / or displayed interface content, and the third message sent by the second device is the second message.
[0026] In this way, the first device can obtain the referenced object from multiple devices in the communication system, thereby avoiding omission when the first device obtains the referenced object and ensuring that the first device can obtain the referenced object.
[0027] In combination with the previous embodiment, after the first device receives the third message sent by some or all of the multiple devices, the method also includes: the first device determines multiple candidate objects from the received one or more third messages; the first device sends the information of the multiple candidate objects to the fifth device, so that the fifth device outputs the multiple candidate objects; the first device receives the information of the reference object sent by the fifth device, and the reference object is selected by the user from the multiple candidate objects.
[0028] In this way, through the user's selection operation, interference from multiple candidate objects is eliminated, ensuring that the first device obtains the referent object that meets the user's intention.
[0029] For example, the fifth device can facilitate receiving user operations. In different scenarios, the fifth device includes different electronic devices. For example, in a navigation scenario, the fifth device includes a car computer. For example, in a smart home scenario, the fifth device includes a mobile phone. For example, in a smart office scenario, the fifth device includes a laptop.
[0030] In combination with the first aspect, in all the aforementioned implementation methods of the first aspect, the first voice indicates a device that performs the first operation, and the device that performs the first operation is the third device.
[0031] In combination with the first aspect, in some implementation methods of the first aspect, the first voice does not indicate the device that performs the first operation, and the third device is determined by the first device based on the first operation.
[0032] In combination with the first aspect, in some implementation methods of the first aspect, the first voice indicates the type of device performing the first operation, the multiple devices detected by the first device all belong to this device type, the third device is determined by the first device among the multiple devices based on the first operation, or the third device is a device selected by the user from the multiple devices.
[0033] In this way, if the third device is determined by the first device based on the first operation among the multiple devices, it can be guaranteed that the third device is the best device among the multiple devices to perform the first operation. If the third device is determined based on the user's selection among the multiple devices, it can be guaranteed that the device selected by the user can perform the first operation.
[0034] In combination with the first aspect, in all the aforementioned implementation methods of the first aspect, the method also includes: the first device uses the reference object to replace the first reference word in the first voice to obtain a first instruction; instructs the third device to perform the first operation on the reference object, specifically including: the first device sends the first instruction to the third device, and the first instruction is used to instruct the third device to perform the first operation on the reference object.
[0035] In combination with the first aspect, in all the aforementioned implementation methods of the first aspect, the first voice also includes the execution condition of the first operation, instructing the third device to perform the first operation on the referred object, specifically including: instructing the third device to perform the first operation on the referred object after detecting that the execution condition is met; or, after detecting that the execution condition is met, instructing the third device to perform the first operation on the referred object.
[0036] In this way, the first device can detect the execution conditions of the first operation, and the third device can also detect the execution conditions of the first operation, ensuring that the execution conditions in the voice instruction can be detected by the devices in the communication system. This avoids the problem of the first device being unable to determine the execution conditions and being unable to instruct the third device to perform the first operation under the condition that the user's intention is met, and ensures that the execution conditions are met when the third device performs the first operation.
[0037] In combination with the previous embodiment, the execution condition includes: arriving at the first time and / or arriving at the first location.
[0038] In combination with the first aspect, in all the aforementioned implementation methods of the first aspect, the first operation is any one of the following: a navigation operation, a play operation, or a phone call operation.
[0039] In combination with the first aspect, in all the aforementioned implementation methods of the first aspect, the first device and the third device are the same device.
[0040] In this way, the first device can directly execute the first operation, reducing the occupation of communication resources in the communication system.
[0041] In combination with the first aspect, in all the aforementioned implementation methods of the first aspect, the third device executing the first operation can also be regarded as the third device providing the user with a service corresponding to the first operation.
[0042] In a second aspect, an electronic device is provided, which includes one or more memories and one or more processors; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the electronic device to execute a method as described in the first aspect or any one of the embodiments of the first aspect.
[0043] In a third aspect, a communication system is provided, which includes: a first device, a second device, and a third device, wherein the first device is used to execute the method as described in the first aspect or any one of the embodiments of the first aspect.
[0044] In a fourth aspect, a chip is provided, which is applied to an electronic device. The chip includes one or more processors, which are used to call computer instructions to enable the electronic device to execute the method as described in the first aspect or any one of the embodiments of the first aspect.
[0045] In a fifth aspect, a computer-readable storage medium is provided, comprising instructions, which, when executed on an electronic device, enable the electronic device to execute a method as described in the first aspect or any one of the embodiments of the first aspect.
[0046] In a sixth aspect, a computer program product is provided, which includes computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the method as described in the first aspect or any one of the embodiments of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] FIG1 is a schematic diagram of a scenario in which a vehicle computer provides navigation services to a user through voice commands according to an embodiment of the present application;
[0048] FIG2 is a schematic diagram of a scenario in which a mobile phone provides navigation services to a user through voice commands containing pronouns, provided in an embodiment of the present application;
[0049] FIG3 is a schematic diagram of a scenario in which multiple devices are unable to provide navigation services to users through voice commands containing pronouns, provided by an embodiment of the present application;
[0050] FIG4 is a schematic diagram of a communication system provided in an embodiment of the present application;
[0051] FIG5A is a flow chart of a voice interaction method including pronouns provided in an embodiment of the present application;
[0052] FIG5B is a flowchart of another voice interaction method including pronouns provided in an embodiment of the present application;
[0053] FIG6 is a schematic diagram of a scenario in which multiple devices provide navigation services to users through voice commands containing pronouns, provided in an embodiment of the present application;
[0054] FIG7 is a schematic diagram of a scenario in which multiple devices provide a video playback service to a user through a voice instruction containing a pronoun, provided in an embodiment of the present application;
[0055] FIG8 is a schematic diagram of a scenario in which multiple devices provide a dialing service to a user through a voice instruction containing a pronoun, provided in an embodiment of the present application;
[0056] FIG9 is a schematic diagram of another scenario in which multiple devices provide a video playback service to a user through a voice instruction containing a pronoun, provided by an embodiment of the present application;
[0057] FIG10 is a schematic diagram of a scenario in which multiple devices provide file operation services to users through voice commands containing pronouns, provided in an embodiment of the present application;
[0058] FIG11 is a schematic diagram of the software structure of an electronic device provided in an embodiment of the present application;
[0059] FIG12 is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0060] The technical solutions in the embodiments of the present application will be described clearly and in detail below with reference to the accompanying drawings.
[0061] Voice assistant is a function provided by electronic devices such as mobile phones, tablet computers, smart watches, etc., which supports electronic devices to detect voice commands input by users, use voice recognition technology, natural language processing (such as semantic understanding) technology, etc. to identify the user's intentions, and respond in a targeted manner, such as voice answering, launching applications, performing navigation operations, performing video playback operations, performing dialing operations, etc. The electronic device can start the voice assistant after detecting the wake-up word. The wake-up word can be user-defined or the default of the electronic device, for example, it can be "Xiaoyi Xiaoyi". Alternatively, the electronic device can also start the voice assistant in response to the user operation acting on the electronic device, such as long pressing the power button of the mobile phone or pressing the voice button on the car computer. In some other embodiments, the electronic device can also always turn on the voice assistant, which is not limited here.
[0062] The voice commands input by users usually include three important parts: the operation that the user wants the electronic device to perform, the object of the operation, and the electronic device that performs the operation. For example, the voice command "navigate the phone to ** route" indicates the navigation operation, the object that performs the navigation operation is ** route, and the electronic device that performs the navigation operation is the phone. For another example, the voice command "play ** video on tablet" indicates the play operation, the object that performs the play operation is ** video, and the electronic device that performs the play operation is the tablet. For another example, the voice command "call contact ** on watch" indicates the dialing operation, the object that performs the dialing operation is contact **, and the electronic device that performs the dialing operation is the watch.
[0063] In some cases, the voice command can omit the electronic device that performs the operation, and the electronic device can be defaulted to the device that currently receives the voice command. For example, if a user enters the voice command "Navigate to ** Road" into a mobile phone, the mobile phone will act as the executor of the navigation operation and execute the navigation to ** Road.
[0064] The methods by which users interact with electronic devices through voice assistants may include the following:
[0065] 1. The user inputs a precise voice command, and the electronic device can respond to the voice command.
[0066] Precise voice commands clearly define the intended action. For example, "Navigate to route **" or "Play video **" are examples of precise voice commands. After receiving a precise voice command, the electronic device can understand the user's intent and perform the corresponding action in response to the command.
[0067] Figure 1 illustrates the first interaction method described above. As shown in Figure 1 (a), a user can enter a voice command to the vehicle computer, "Navigate to Times Community." As shown in Figure 1 (b), the vehicle computer responds to the voice command and navigates to Times Community, providing navigation services to the user.
[0068] 2. The user inputs a voice command containing a pronoun. If the electronic device contains the referent object corresponding to the pronoun, the electronic device can respond to the voice command.
[0069] Pronouns are words with ambiguous meanings and unclear referents. Examples include the Chinese words "this," "that," "these," "those," "here," and "there," and the English words "this," "that," "these," and "those." They also include combinations of these words and object categories, such as "this person," "this video," "this file," and "this address." They also include general terms that don't specify specific objects, such as "his," "home," and "company."
[0070] Voice commands containing pronouns may include, for example, "navigate here," "navigate home," "play this video," "call this contact," and the like.
[0071] After receiving a voice command containing a pronoun, the electronic device can search locally to see whether there is a referent object corresponding to the pronoun. If the corresponding referent object is found, the electronic device can understand the user's intention and can perform the corresponding operation in response to the voice command.
[0072] Figure 2 illustrates the second interaction method described above. As shown in Figure 2 (a), a user can launch a mobile phone's map application and search for a location through the application. The user can then enter the voice command "Navigate here" into the phone. Upon receiving this voice command, the phone can search locally for the pronoun "here," which corresponds to the location currently displayed on the phone. As shown in Figure 2 (b), the phone can then navigate to that location, providing navigation services to the user.
[0073] 3. If the user inputs a voice command containing a pronoun, and the electronic device does not contain the referent object corresponding to the pronoun, the electronic device cannot respond to the voice command.
[0074] After the electronic device receives a voice command containing a pronoun, if the referent object corresponding to the pronoun is not found locally, the electronic device cannot understand the user's intention and cannot respond to the voice command to perform the corresponding operation.
[0075] Figure 3 exemplifies the third interaction method mentioned above. In the scenario shown in Figure 3, the user can bring a mobile phone into the car. As shown in Figure 3 (a), the user can start the map application of the mobile phone and search for a certain place through the map application. After that, the user can input the voice command "navigate here" to the car computer. From the user's point of view, the "here" in the voice command input is the place he searched for on the mobile phone. However, after the car computer receives the voice command, since the referent object exists in the mobile phone, the car computer cannot search for the referent object corresponding to the referent "here" locally. As shown in Figure 3 (b), the car computer cannot perform the operation of navigating to the place and cannot provide navigation services to the user. As shown in Figure 3 (b), the car computer can further ask the user for the specific location of navigation.
[0076] It often happens that users input voice commands containing pronouns, and the electronic device that receives the voice command may not have the referent object corresponding to the pronoun locally. For a user, there may be multiple electronic devices around him, and the referent object corresponding to the pronoun input by the user may exist in one of the electronic devices. When the user inputs a voice command containing a pronoun, it is believed that the pronoun refers to the referent object in the electronic device. However, the device that receives the voice command may be another electronic device. The other electronic device cannot know the above-mentioned referent object and therefore cannot respond to the voice command, that is, it cannot provide the corresponding service to the user.
[0077] To address the above-mentioned issues, the present application provides a voice interaction method, related apparatus, and communication system including pronouns. This method can be applied to a communication system including multiple electronic devices. An electronic device in the communication system can receive a voice command containing a pronoun input by a user, and the electronic device can search for the referent corresponding to the pronoun from multiple electronic devices in the communication system, and then understand the user's intention. Afterwards, the electronic device can find a suitable electronic device from the multiple electronic devices in the communication system to respond to the voice command and perform the corresponding operation to provide services to the user.
[0078] In the solution provided by this application, the electronic device where the referent object is located and the electronic device that performs the operation in response to the voice command can be different devices. In this way, even if the referent object corresponding to the referent word does not exist locally on the electronic device, it can still respond to the voice command input by the user and perform the corresponding operation. This achieves cross-device information synchronization between various devices in the communication system and provides services to users across devices.
[0079] This solution supports users to interact with electronic devices through voice commands containing pronouns. Even if the user inputs a voice command containing a pronoun, the electronic device can respond to the voice command to provide services to the user, which can enhance the user's voice interaction experience with the electronic device. In addition, using this solution, the user does not need to pre-set the referential object in an electronic device and input voice commands to the electronic device. This solution uses multiple devices to complete information synchronization. As long as there is an electronic device in the communication system that contains the referential object corresponding to the pronoun, for example, the referential object corresponding to the pronoun is displayed on the display screen of the electronic device, the device in the communication system can understand the user's intention and respond to the voice command to provide services that meet the user's intention. It can be seen that this voice interaction method reduces the limitations of the interaction between users and electronic devices, making voice interaction simpler, natural and convenient.
[0080] Below, the communication system provided by the embodiment of the present application is first introduced.
[0081] As shown in Figure 4, the communication system provided in the embodiment of the present application is also called a multi-device collaborative system 10. The multi-device collaborative system 10 may include multiple electronic devices.
[0082] The multiple electronic devices in the multi-device collaborative system 10 can be of various types, and the embodiments of the present application do not limit this. The multiple electronic devices may include mobile phones, tablet computers, desktop computers, desktop computers, laptop computers, handheld computers, notebook computers, smart screens, wearable devices, augmented reality (AR) devices, virtual reality (VR) devices, artificial intelligence (AI) devices, car computers, smart headphones, game consoles, digital cameras and other smart devices, and may also include smart speakers, smart lamps, smart air conditioners, ovens, coffee machines, cameras, doorbells, millimeter wave sensors and other Internet of Things (IoT) devices or smart home devices, and may also include printers, scanners, fax machines, copiers, projectors and other office equipment.
[0083] The multi-device collaborative system 10 may include movable electronic devices such as mobile phones, smart watches, car computers, tablets, smart bracelets, etc., and may also include immovable smart screens, smart lamps, smart air conditioners, smart speakers, projectors and other devices. In the embodiment of the present application, the multi-device collaborative system 10 may include different devices in different scenarios. The scenarios may include smart travel scenarios, smart home scenarios, smart office scenarios, sports and health scenarios, audio and video entertainment scenarios, etc. For example, in the driving scenario shown in Figure 4, the multi-device collaborative system 10 may include a mobile phone 100-1, a smart watch 100-2, and a car computer 100-3, etc. in sequence. In the smart home scenario shown in Figure 4, the multi-device collaborative system 10 may include a mobile phone 100-1, a smart watch 100-2, a smart screen 100-5, a smart speaker 100-6, etc. in sequence. In the smart office scenario shown in Figure 4, the multi-device collaborative system 10 may include a mobile phone 100-1, a smart watch 100-2, a laptop computer 100-4, a projector 100-7, etc. The multi-device collaboration system 10 of the present application may include any number of electronic devices, not limited to the number in the above example.
[0084] The multi-device collaboration system 10 may include electronic devices produced by the same manufacturer, or may include electronic devices produced by different manufacturers, which is not limited in this embodiment of the present application.
[0085] The multiple electronic devices in the multi-device collaborative system 10 may include a software operating system (OS). The OS configured for each electronic device may be different, including but not limited to Harmony OS. Each electronic device can also be configured with the same software operating system, for example, Harmony
[0086] The steps of adding an electronic device to the multi-device collaboration system 10 may include the following two steps:
[0087] 1. The electronic device and the electronic devices in the multi-device collaboration system 10 are logged in to the same user account, family account, or associated account.
[0088] 2. The electronic device is detected as being online by the electronic devices in the multi-device collaboration system 10.
[0089] Multiple electronic devices can log in to the same account, family account, or linked account. For example, they can log in to the same system account (such as a Huawei account). Each electronic device can then communicate with the server that maintains the system account (such as a server provided by Huawei) via cellular network technologies such as 3G, 4G, and 5G, or wide area network technologies, and then communicate through this server. A family account is an account used by family members together. Linked accounts refer to multiple accounts that are linked together.
[0090] Going online can mean being detected or sensed by other electronic devices. Once an electronic device in the multi-device collaborative system 10 detects that it is online, it can be considered to have joined the multi-device collaborative system 10. The electronic device's online detection can be done through short-range communication technology or by a server. Please refer to the following two paragraphs for details.
[0091] Electronic devices can use near-field communication (NFC) technology to detect online activity. For example, an electronic device can use NFC to send a broadcast message containing account information. Other nearby devices, upon receiving the broadcast message, can compare it with their own logged-in account information. If they match, the other device can send a confirmation message back to the electronic device. This allows the electronic device to detect that the other device has come online, and the other device and the other device have joined the same multi-device collaboration system 10.
[0092] Electronic devices can also detect online access through the server. For example, if an electronic device and other devices use the same account to log in to the server, the server can send information about each device logged in to the same account to these devices. In this way, each device logged in to the same account can sense each other and thus join the same multi-device collaborative system 10.
[0093] The electronic devices in the multi-device collaborative system 10 are connected to some or all other electronic devices and can communicate based on the connection. In other words, any two electronic devices in the multi-device collaborative system 10 can be directly connected and communicate with each other, or can communicate indirectly through another electronic device, or can have no connection or communication relationship at all.
[0094] For example, in the multi-device collaborative system 10, the mobile phone 100-1 can communicate directly with the car computer 100-3, the smart watch 100-2 and the car computer 100-3 can communicate indirectly through the mobile phone 100-1, and the smart screen 100-5 and the laptop computer 100-4 may not have a connection relationship and cannot communicate directly.
[0095] Connections between electronic devices can be established in a variety of ways, for example, a connection can be established under user triggering, or a connection can be established actively by the device, which is not limited in this application.
[0096] The electronic devices in the multi-device collaborative system 10 can establish connections and communicate with each other through any one or more of the following technologies: wireless local area network (WLAN), Wi-Fi direct / Wi-Fi peer-to-peer (Wi-Fi P2P), Bluetooth (BT), near field communication (NFC), infrared (IR), ZigBee, ultra wideband (UWB), hotspot, Wi-Fi softAP, cellular network, wired technology or remote connection technology, etc. Among them, Bluetooth can be classic Bluetooth or Bluetooth low energy (BLE).
[0097] For example, an electronic device can communicate with other devices in the same wireless local area network (WLAN) through a wireless local area network (WLAN). For another example, an electronic device can discover other nearby devices through short-range communication technologies such as BT and NFC, and communicate with other devices after establishing a communication connection. For another example, an electronic device can operate in wireless access point (AP) mode and create a wireless local area network. After other electronic devices connect to the wireless local area network created by the electronic device, the electronic device and other devices can communicate through Wi-Fi softAP.
[0098] In the embodiment of the present application, a communication system may have multiple different connection modes. For example, as shown in FIG4 , the mobile phone 100 - 1 and the vehicle computer 100 - 3 may communicate via Bluetooth, and the mobile phone 100 - 1 and the laptop computer 100 - 4 may communicate via Wi-Fi.
[0099] Each electronic device in the multi-device collaboration system 10 can synchronize or share device information with each other based on inter-device communication connections. This device information may include, but is not limited to, any one or more of the following: device capability information, device type information, user status information collected by the device (such as the distance between the device and the user), device operating status information, and environmental status information.
[0100] In an embodiment of the present application, the multi-device collaborative system 10 can determine the electronic device that collects voice commands, the electronic device that analyzes voice commands, the electronic device that provides services to users, etc. based on the device information of each device.
[0101] Describing the electronic devices included in the multi-device collaboration system 10, the present application solution includes but is not limited to the following three aspects:
[0102] 1. The multi-device collaborative system 10 provides a communication system for the voice interaction method containing pronouns in this application, and the electronic devices in the multi-device collaborative system 10 can obtain voice instructions. The voice instruction is a first voice, and the first voice includes but is not limited to the operation to be performed (i.e., the first operation) and the pronoun (i.e., the first pronoun). The electronic devices in the multi-device collaborative system 10 can directly obtain the voice instruction through a microphone, or indirectly obtain the voice instruction through a communication connection between different devices. For example, the user says the voice instruction "navigate here", and the voice instruction contains the navigation operation and the pronoun "here". The car computer 100-3 can directly obtain the voice instruction spoken by the user through the microphone, and the mobile phone 100-1 can indirectly obtain the voice instruction through the communication connection (such as WLAN) with the car computer 100-3.
[0103] 2. The electronic device in the multi-device collaborative system 10 can also determine the referent of the pronoun in the voice command. The referent is the recipient of the first operation and can exist in the electronic device of the multi-device collaborative system 10. The electronic device in the multi-device collaborative system 10 obtains the interface content and / or storage content of the local device or other devices, and finds the referent of the pronoun. For example, the mobile phone 100-1 can recognize the address of "Times Community" in its own display interface. This address can be the specific address of the navigation operation, that is, it can be the referent of the pronoun "here".
[0104] 3. The electronic device in the multi-device collaborative system 10 can also determine the electronic device that performs the first operation. The electronic device in the multi-device collaborative system 10 can determine the electronic device that performs the first operation in any one or more of the following ways: determining whether the device has the ability to perform the first operation, obtaining the strength of the device's ability to perform the first operation, obtaining the distance between the user and the device, etc. For example, if the mobile phone 100-1 obtains that the car machine 100-3 has the ability to provide navigation services to the user, and the display interface of the car machine 100-3 is larger than the mobile phone 100-1, it is determined that the car machine 100-3 performs the operation of navigating to the "Times Community".
[0105] Each electronic device in the multi-device collaborative system 10 can install and run a voice assistant, which runs after the microphone of the electronic device receives the wake-up word voice. The voice assistant allows the electronic device to perform any one or more of the following tasks: obtaining a voice instruction, determining the referent of the pronoun in the voice instruction, and determining the electronic device that performs the first operation. For the specific description of the various tasks, please refer to the aforementioned description of the present application scheme from the electronic devices included in the multi-device collaborative system 10, which will not be repeated here.
[0106] Based on the multi-device collaborative system 10 described in FIG4 , the voice interaction method including pronouns provided by the present application is introduced below.
[0107] FIG5A exemplarily shows a flow chart of a voice interaction method including pronouns provided in an embodiment of the present application.
[0108] As shown in FIG5A , the method may include the following steps:
[0109] Phase 1 ( S101 ): Each device in the multi-device coordination system 10 synchronizes device information.
[0110] S101, each device in the multi-device collaboration system 10 obtains device information of other devices.
[0111] The multi-device collaboration system 10 involved in the embodiment of the present application can be established through the following two steps: 1. Multiple electronic devices log in to the same user account, family account, or associated account. 2. Multiple electronic devices detect that each other is online.
[0112] For the specific process of establishing the multi-device collaborative system 10, you can refer to the steps of adding electronic devices to the multi-device collaborative system 10 in Figure 4 above, which will not be repeated here.
[0113] In an embodiment of the present application, multiple electronic devices can be organized into a multi-device collaborative system 10 in the manner described above, and then each device in the multi-device collaborative system 10 obtains device information of other devices in the multi-device collaborative system 10, that is, device information synchronized by the electronic devices in the multi-device collaborative system 10.
[0114] In some embodiments, the electronic devices in the multi-device collaboration system 10 may also synchronize device information with each other periodically or aperiodically according to preset rules, for example, once every 30 seconds or once every minute.
[0115] In an embodiment of the present application, each electronic device in the multi-device collaborative system 10 can synchronize device information with each other based on the connection between the devices. For example, if each electronic device in the multi-device collaborative system 10 is connected to the same WLAN, the device information can be synchronized with each other through the WLAN (for example, through a router). For another example, if the electronic devices in the multi-device collaborative system 10 are connected via Bluetooth, the device information can be synchronized with each other based on the Bluetooth connection. For another example, if each electronic device in the multi-device collaborative system 10 is remotely connected by logging into the same account, the device information can be transferred through the server that manages the account. If the multi-device collaborative system 10 contains two electronic devices that are not directly connected, the two electronic devices can synchronize device information with each other through an intermediate device in the multi-device collaborative system 10.
[0116] In an embodiment of the present application, the device information synchronized between the electronic devices in the multi-device collaborative system 10 may indicate any one or more of the following: device capabilities, device type, distance between the device and the user, device operating status, and environmental status.
[0117] The information on device capabilities is used to characterize or describe the capabilities supported by the electronic device. The capabilities of the electronic device may specifically include any one or more of the following: computing power, sound pickup capability, display capability, playback capability or storage capability, etc. Different electronic devices may have different capabilities. For example, a mobile phone has computing power, sound pickup capability, display capability, playback capability and storage capability, and a smart speaker has computing power, sound pickup capability, playback capability and storage capability. This application does not limit the electronic device capabilities possessed by different devices. In some embodiments, the capability information of the device may also include component performance parameters of the supported capabilities. For example, for computing power, it may include the computing power parameters of the central processing unit (CPU) of the corresponding device, such as the number of CPU cores, the frequency of a single CPU core, etc. For display capability, it may include the size and frame rate of the display screen of the corresponding device, etc. For sound pickup capability, it may include the sensitivity and signal-to-noise ratio of the microphone of the corresponding device, etc. In some embodiments, information on whether each capability is currently occupied and whether the current electronic device supports each capability may also be added.
[0118] Device type information is used to characterize or describe the type of electronic device. Device types may include any one or more of the following: mobile devices such as mobile phones, smart watches, and laptops; driving devices such as car computers; and fixed devices such as smart screens, smart speakers, and projectors. In the embodiments of this application, this application does not limit the specific information on device types.
[0119] The distance between the user and the device is measured in straight lines. This distance can be obtained using a distance sensor. For example, an electronic device can measure distance using infrared or laser sensors.
[0120] The device's operating status indicates how the device's resources (such as computing power, display, microphone, etc.) are occupied.
[0121] Environmental status information is used to describe the device's surroundings. For example, if a device's humidity sensor detects a high humidity value on a rainy day, this high humidity value can indicate that the device's surroundings are humid. Another example is a device's air pressure sensor, which can detect a gradual decrease in atmospheric pressure as the device's altitude gradually increases. This decrease in atmospheric pressure can indicate that the device is in an environment with rising altitude. Another example is a device's ambient light sensor, which detects low ambient light on a rainy day, which can indicate that the device is in a dark environment.
[0122] Phase 2 (S102-S104): Wake up the voice assistants of the devices in the multi-device collaborative system 10.
[0123] S102, the devices in the multi-device collaborative system 10 collect the wake-up words input by voice.
[0124] In an embodiment of the present application, a device with sound pickup capability in the multi-device collaborative system 10, before receiving a wake-up word, only collects voice instructions through a microphone to determine whether the voice instruction contains a wake-up word. If the device determines that the voice instruction does not contain a wake-up word, it does not respond to the voice instruction. If the device determines that the voice instruction contains a wake-up word, it responds to the voice instruction and performs subsequent operations. Before the device recognizes the wake-up word, the device runs the relevant hardware (such as a chip for detecting the wake-up word) and software (such as a software program running in the chip that detects the wake-up word) for recognizing the wake-up word, and the device power consumption is low. After the device detects the wake-up word, subsequent operations are run, such as the device determines the device suitable for answering, starts the voice assistant of the device, etc., for details, refer to steps S103-S104 below. For example, mobile phones, smart watches, car computers, laptops, smart speakers and other devices in the multi-device collaborative system 10 all have sound pickup capabilities and can obtain the wake-up word input by the voice instruction. Among them, any one or more devices including mobile phones, smart watches, car computers, laptops and smart speakers receive the wake-up word. This application does not limit the type and number of devices that receive the wake-up word input by the voice instruction.
[0125] In an embodiment of the present application, the wake-up word may include, for example, a default wake-up word (such as "Xiaoyi Xiaoyi"), and may also include a user-defined wake-up word.
[0126] In some other embodiments, S102 can also be replaced by the device in the multi-device collaborative system 10 receiving a gesture wake-up word, for example, it may include receiving a long press operation of a mobile phone power button, volume button, car computer power button, volume button, etc., and may also include a gesture wake-up word (such as an "OK" gesture). This application does not limit the specific content of the wake-up word.
[0127] S103, each device in the multi-device collaborative system 10 determines a device suitable for responding to the wake-up word.
[0128] In the embodiment of the present application, the device suitable for responding to the wake-up word can be determined based on any one or more of the following: whether the sound pickup capability is available, the strength of the sound pickup capability, the distance between the device and the user, the operating status of the device, etc. The above information can be obtained by each device in S101 from the device information synchronized by the multi-device collaborative system 10.
[0129] In some embodiments, each device can obtain whether its own sound pickup capability is idle, and whether the sound pickup capabilities of other devices in the multi-device collaborative system 10 are idle, and then compare them with each other to determine the device with idle sound pickup capability as the device suitable for responding to the wake-up word.
[0130] In some embodiments, each device can obtain information about the strength of its own sound pickup capability, as well as the strength of the sound pickup capability of other devices in the multi-device collaborative system 10, and then compare them with each other. If the device's own sound pickup capability is the strongest in the multi-device collaborative system 10, the device itself is considered suitable for responding to the wake-up word, otherwise it is not suitable for responding to the wake-up word.
[0131] In some embodiments, each device can obtain the distance between itself and the user, as well as the distance between other devices and the user in the multi-device collaborative system 10, and then compare them with each other. If the distance between the device itself and the user is the distance of the device closest to the user in the multi-device collaborative system 10, the device itself is considered suitable for responding to the wake-up word, otherwise it is not suitable for responding to the wake-up word.
[0132] In some embodiments, each device can obtain the operating status of its own device and the operating status of other devices in the multi-device collaborative system 10. If the device itself is idle, that is, the resources in the device are less occupied, the device itself can respond to the wake-up word; otherwise, a device that is idle in the multi-device collaborative system 10 is selected to respond to the wake-up word.
[0133] In some embodiments, each device can also combine multiple of the above factors to determine whether the device itself is suitable for responding to the wake-up word. For example, in a driving environment, if the car computer's sound pickup capability is idle and it is the device closest to the user, the wake-up word will be answered through the car computer. For another example, in a smart home scenario, if the distance between the mobile phone and the smart speaker and the user is less than a first threshold (the first threshold is a value preset by the device before determining the answer to the wake-up word), and the smart speaker's sound pickup capability is stronger than that of the mobile phone, the wake-up word will be answered through the smart speaker.
[0134] For example, referring to the multi-device collaborative system 10 of Figure 4, in a driving scenario, both the mobile phone 100-1 and the car computer 100-3 collect the wake-up word input by the user. If the user is closer to the car computer 100-3 and the sound pickup ability of the car computer 100-3 is stronger than that of the mobile phone 100-1, the car computer 100-3 determines that it is suitable for responding to the wake-up word, and the mobile phone 100-1 determines that it is not suitable for responding to the wake-up word.
[0135] In some embodiments, each device can have a preset priority relationship for devices that respond to the wake-up word. For example, the priority of the devices that respond to the wake-up word is prioritized in descending order, such as the car computer, mobile phone, smart speaker, and laptop computer. In this way, after each device receives the wake-up word, the device with the highest priority in the multi-device collaborative system 10 can be determined as the device that responds to the wake-up word based on the priority relationship.
[0136] In some embodiments, after a device in the multi-device collaborative system 10 determines a device suitable for responding to a wake-up word, it can send a response signal to the determined device. This response signal is used to instruct the device to activate the voice assistant and respond to the user. In this way, even if a device in the multi-device collaborative system 10 suitable for responding to a wake-up word does not receive the wake-up word, it can still activate the voice assistant and respond to the user.
[0137] In the embodiment of the present application, a device suitable for responding to the wake-up word determined in S103 can be referred to as the fourth device. In different scenarios, the fourth device can be different devices. For example, in a smart home scenario, the sound pickup ability of a smart speaker is stronger than that of a mobile phone, and the distance between the smart speaker and the user is closer, then the fourth device is the smart speaker. For another example, in a smart office scenario, the distance between the mobile phone and the user is closer, and the distance between the laptop and the user is farther, then the fourth device is the mobile phone.
[0138] S104: The fourth device starts the voice assistant.
[0139] In an embodiment of the present application, the voice assistant is started, that is, an application (APP) named voice assistant is run, and the voice instructions received by the microphone are saved and sent to the voice assistant.
[0140] In some embodiments, the fourth device may respond to the user first. The methods for responding to the user may include any one or more of the following: 1. The fourth device broadcasts a prompt voice message, such as outputting the voice message "I'm here" or "What do you need?" 2. The fourth device displays a prompt message on the display screen, such as a pop-up window containing prompt text, such as "I'm here." 3. The fourth device may output a vibration prompt via a motor. This application does not limit the specific method by which the fourth device responds to the user.
[0141] Stage 3 (S105-S106): The user inputs a voice instruction containing a pronoun.
[0142] S105: The fourth device collects a voice instruction, where the voice instruction indicates a first operation and further includes a pronoun for the first operation.
[0143] In an embodiment of the present application, a voice instruction indicates an operation that the user wants the electronic device to perform, and the operation that the user wants the electronic device to perform is the first operation. The first operation may be, for example, a navigation operation, a playback operation, a dialing operation, and the like. In a specific voice instruction, for example, in the voice instruction "mobile phone navigates here", the user intends the mobile phone to perform a navigation operation, and the first operation is a navigation operation. For another example, in the voice instruction "laptop play this video", the user intends the laptop to perform a playback operation, and the first operation is a playback operation. For another example, in the voice instruction "car computer call him", the user intends the car computer to perform a phone call operation, and the first operation is a dialing operation. This application does not limit the specific operation of the first operation.
[0144] The voice command includes a pronoun that receives the first operation, such as "here," "this," or "he." For example, in the voice command "Mobile navigation here," the first operation is navigation, and the pronoun that receives the first operation is "here." For another example, in the voice command "Laptop play this video," the first operation is playback, and the pronoun that receives the first operation is "this video." For another example, in the voice command "Car call him," the first operation is dialing, and the pronoun that receives the first operation is "he."
[0145] In the embodiment of the present application, the voice instruction collected by the fourth device may also be referred to as the first voice. The pronoun in the voice instruction that undergoes the first operation may also be referred to as the first pronoun.
[0146] In some embodiments, the voice instruction can also specify the device that performs the first operation. For example, "mobile phone navigates here", this voice instruction specifies that the device that performs the first operation is a mobile phone.
[0147] In some embodiments, the voice instruction does not specify the device for the first operation. One case is that the voice instruction omits the device for performing the first operation, for example, "navigate here", this voice instruction omits the device for performing the navigation operation. Another case is that the voice instruction only specifies the type of device for performing the first operation, but does not specify the specific device. For example, "play video on mobile device", the mobile devices in this voice instruction include mobile phones, smart watches, tablet computers and other mobile devices. If the multi-device collaborative system 10 includes the above-mentioned multiple mobile devices, it is necessary to determine the device that executes the voice instruction among the multiple mobile devices. In this embodiment, the multi-device collaborative system 10 needs to determine the executing device before executing the first operation.
[0148] In some embodiments, the fourth device receives a voice command via a microphone. The voice command input by the user may include, but is not limited to, a wake-up word, a first operation, and a reference to the first operation. For example, the user may input "Xiaoyi Xiaoyi, please navigate here." In this case, the devices in the multi-device collaboration system 10 may receive the wake-up word "Xiaoyi Xiaoyi," and the fourth device may receive the voice command "Please navigate here."
[0149] S106: The fourth device sends a voice command to the first device.
[0150] In some embodiments, the fourth device determines the device for analyzing the voice command, which may be referred to as the first device. The fourth device may determine the first device by including but not limited to any one or more of the following: the computing power of the device, the operating status of the device. The fourth device determines the device for analyzing the voice command by the computing power of all devices in the multi-device collaborative system 10 and the operating status of the corresponding devices. In the driving scenario, the computing power of the mobile phone in the multi-device collaborative system 10 is the strongest. If the computing unit of the mobile phone (such as the CPU) is not in a busy state, the mobile phone is used as the first device for analyzing the voice command. If the computing unit of the mobile phone is in a busy state, the device with the second strongest computing power (such as the car computer) is used as the first device for analyzing the voice command. The busy state indicates that the resources in the device (such as CPU computing power, etc.) are largely occupied by other processes that require computing power.
[0151] In some embodiments, the first device and the fourth device are the same device. That is, the first device determined by the fourth device is the fourth device itself, and step S106 can be understood as interaction between internal modules of the fourth device to send a voice command.
[0152] In some implementations, the first device and the fourth device may be different devices, and this application does not limit whether the first device and the fourth device are the same device.
[0153] Phase 4 (S107-S112): The first device determines an object to be subjected to the first operation.
[0154] S107 , the first device sends a first message to multiple devices in the multi-device collaboration system 10 , where the first message is used to query candidate objects that can undergo the first operation.
[0155] In an embodiment of the present application, after the first device obtains the voice message sent by the fourth device, semantic analysis is performed on the voice message to obtain the first operation and the pronoun that bears the first operation. The first device performs semantic analysis on the voice message, including but not limited to the following methods: after obtaining the voice message, the first device starts the voice assistant of the first device and analyzes the voice message through the voice assistant. Based on the obtained first operation and the pronoun that bears the first operation, the first device determines the first message sent to multiple devices in the multi-device collaboration system 10.
[0156] In an embodiment of the present application, the first device obtains the first operation and the pronoun that bears the first operation, determines the object type that can bear the first operation, and sends the first message to multiple devices in the multi-device collaboration system 10. The first message includes any one or more of the following: the object type of the first operation, the indication information of the first operation. The first operation is used to obtain the candidate object that bears the first operation. For example, when the first device obtains the voice instruction "Navigate here", through the voice analysis of the voice assistant of the first device, the first operation is obtained as the navigation operation, and "here" is the pronoun that bears the navigation operation. At the same time, the first device determines that the object that bears the navigation operation is of the address type (such as ** Road, ** Community, ** Park, etc.), and the address type referred to by "here" is not clear. Therefore, the first device sends the first message for obtaining the candidate object of the address type to multiple devices in the multi-device collaboration system 10.
[0157] In an embodiment of the present application, the first device sending the first message to itself can be understood as the interaction between internal modules of the first device.
[0158] In an embodiment of the present application, the first device sends the first message to itself and queries the candidate object locally. After it is obtained that the candidate object that bears the first operation is not queried in the first device, the first message is sent to other devices in the multi-device collaboration system 10. For example, if the mobile phone is the first device, and the referring object of the navigation operation "** Road" is displayed on the display interface of the car machine, then the mobile phone does not query the referring object of the navigation operation locally. The mobile phone will send the first message to multiple devices in the multi-device collaboration system 10 including the car machine and the smart watch. The present application does not limit the specific manner in which the first device sends the first message.
[0159] With this embodiment, the first device can not only obtain candidate objects locally on the device, but also initiate a request to query candidate objects across devices, expanding the query scope of candidate objects and enabling the acquisition of candidate objects that meet the user's intention. At the same time, the first device preferentially queries whether there are candidate objects locally. When there are candidate objects locally on the first device, it can directly obtain the candidate objects without sending a message to other devices to query candidate objects, improving the efficiency of the first device in querying reference objects in the multi-device collaboration system 10.
[0160] S108. Multiple devices that receive the first message query locally for candidate objects that can withstand the first operation.
[0161] In the embodiments of the present application, after a device in the multi-device collaboration system 10 receives the first message, the device queries locally for candidate objects that can withstand the first operation. The device's query for candidate objects that can withstand the first operation includes but is not limited to the following two methods: 1. The device identifies candidate objects in the display interface. Among them, the display interface of the device displays candidate objects. For example, in a driving scenario, the display interface of the mobile phone displays the specific address of "** Community". When the mobile phone receives the first message and identifies the specific address of "** Community" in the display interface, it uses this specific address as the reference object, and this reference object can be called a candidate object. 2. The device identifies candidate objects in the stored content. Among them, the stored content of the device includes but is not limited to pre-set pronouns. For example, in a map application, it may include the pre-saved address of home as the reference object of the pronoun "home". When the pronoun received by the device is the same as the pre-set pronoun, the reference object of the pre-set pronoun is used as the candidate object.
[0162] In the embodiments of the present application, when a device in the multi-device collaboration system 10 queries for candidate objects locally, it can query for candidate objects using methods including but not limited to the following two: 1. The first message received by the device includes the reference type that can withstand the first operation, and the device obtains the candidate object through this reference type. For example, the mobile phone simultaneously displays "** Road" and "** Video". When the first message received by the mobile phone indicates obtaining a reference object of the address type, the mobile phone uses "** Road" as the reference object, that is, the candidate object. 2. The first message received by the device includes the first operation, and the device obtains candidate objects that match the first operation. For example, the first message received by the mobile phone includes a navigation operation, and the mobile phone displays "** Road", and this "** Road" can be used as the object that can withstand the navigation operation, then the mobile phone uses "** Road" as the candidate object. The present application does not limit the method of querying candidate objects.
[0163] In some embodiments, the candidate object subject to the first operation includes one or more. The multiple candidate objects can come from multiple devices querying candidate objects, or can also come from a single device querying multiple candidate objects. If there are multiple candidate objects, it is necessary to determine the referent object subject to the first operation in subsequent operations. If there is only one candidate object, then the candidate object is the referent object subject to the first operation. This application does not limit the number of candidate objects.
[0164] S109: The device that finds the candidate object sends a third message to the first device, where the third message includes information about the candidate object.
[0165] In some embodiments, some or all of the devices that receive the first message may locally search for one or more candidate objects. The devices that locally search for candidate objects include the sixth device and the second device. The candidate objects include a referent object, and the device where the referent object resides may be referred to as the second device. The number of the sixth device and the second device may be one or more.
[0166] The sixth device and the second device may send a third message to the first device, where the third message includes information about candidate objects found locally by the device. The candidate object information includes information about one or more candidate objects found locally by the device.
[0167] In some embodiments, the number of candidate objects received by the first device may include one or more. If the first device receives only one candidate object, the subsequent steps S110-S112 do not need to be performed, and the first device determines that candidate object as the referent object. For example, if the first device is a mobile phone, and the mobile phone only obtains the candidate object of "** Road" on the mobile phone display interface, then "** Road" is the referent object for the navigation operation. If the first device receives multiple candidate objects, the subsequent steps S110-S112 are performed, and one of the multiple candidate objects is selected through user operation as the referent object that receives the first operation. For example, if the first device is a mobile phone, and the mobile phone obtains the candidate object of "** Road" on the mobile phone display interface and the candidate object of "** Street" on the vehicle display interface, then the mobile phone needs to determine the referent object of "** Road" and "** Street" through the subsequent steps S110-S112 to receive the navigation operation. This application does not limit the number of candidate objects received.
[0168] In some embodiments, if the first device fails to receive a candidate object, the first device may notify the user that the first operation cannot be performed through voice and / or text display. For example, when the first device is a mobile phone, if the mobile phone fails to receive a candidate object, the mobile phone's voice assistant may announce "No location 'here' detected" and display a text prompt at the bottom of the mobile phone's display interface.
[0169] In some embodiments, after the device that receives the first message fails to query for candidate objects, it may send information indicating that no candidate objects have been queried to the first device. The information received by the first device in this application is not restricted.
[0170] S110. The first device sends information of multiple candidate objects to the fifth device.
[0171] In some embodiments, the first device determines a device that is suitable for receiving user operations and determines a referent object among multiple candidate objects. This device may be referred to as the fifth device. In different scenarios, the fifth device may be different devices. For example, in a driving scenario, the in-vehicle computer has stronger display capabilities than a mobile phone, and the display screen of the in-vehicle computer is larger than that of the mobile phone, which is convenient for receiving user operations. Then the fifth device is the in-vehicle computer. Another example is that in a smart home scenario, the mobile phone is closer to the user, which is convenient for receiving user operations. Then the fifth device is the mobile phone. The information of the candidate objects is used to indicate the candidate objects. For example, the first device sends information with address types such as "** Road" and "** Street" to indicate candidate objects of the address type. Another example is that the first device sends information with video types such as "** Video" and "** Program" to indicate candidate objects of the video type.
[0172] In some embodiments, the first device also sends information of the device corresponding to the candidate object to the fifth device. The information of the device corresponding to the candidate object indicates the device that sent the candidate object received by the first device. For example, the first device receives the candidate object "** Road" from the mobile phone, and the mobile phone is the information of the device corresponding to the candidate object "** Road".
[0173] In some embodiments, the first device also sends indication information of the first operation to the fifth device. The indication information of the first operation is used to indicate the operation that the user intends to perform. The indication information includes the first operation. For example, the indication information sent by the first device includes the voice command "Navigate here", and the navigation operation in this voice command is the first operation.
[0174] In some embodiments, the fifth device and the first device may be different devices or the same device. When the fifth device and the first device are the same device, sending information of multiple candidate objects can be understood as: a module with multiple candidate objects inside this device sends information of multiple candidate objects to the display module. This application does not restrict whether the first device and the fifth device are the same device.
[0175] S111. The fifth device outputs multiple candidate objects, receives user operations, and determines a referent object that bears the first operation from among the multiple candidate objects.
[0176] In some embodiments, the fifth device may output multiple candidate objects by including any one or more of the following: 1. The fifth device displays an icon for indicating a candidate object through a display interface. 2. The fifth device announces the candidate object through a speaker. This application does not limit the specific manner in which the fifth device outputs candidate objects.
[0177] In some embodiments, the user operation includes any one or more of the following: the user's touch operation on the device display interface, the user's voice input to the device, the user's gesture operation on the device. The user's operation is an operation to determine the referential object that bears the first operation from multiple candidate objects.
[0178] In some embodiments, the fifth device may receive the user's operation through any one or more of the following ways: 1. The fifth device determines the referential object that bears the first operation by receiving the user's touch operation. 2. The fifth device determines the referential object that bears the first operation by receiving the user's voice instruction. 3. The fifth device determines the referential object that bears the first operation by receiving the user's gesture operation. This application does not limit the specific manner in which the fifth device receives the user's operation.
[0179] S112. The fifth device sends information containing the referential object to the first device.
[0180] In some embodiments, the information containing the referential object may include, but is not limited to, the following forms: The referential object is the complete referential object that bears the first operation. For example, the information of the referential object is the information of "** Road". Or, the referential object is the information referring to a certain candidate object among multiple candidate objects. For example, the information of the referential object is the candidate object from the mobile phone among multiple candidate objects, and this candidate object from the mobile phone is "** Road", that is, the referential object is "** Road".
[0181] In some embodiments, the fifth device and the first device may be different devices or the same device. When the fifth device and the first device are the same device, the fifth device sending information containing the referential object to the first device can be understood as: The module for sending information inside the fifth device sends the information containing the referential object to the module for receiving information. This application does not limit whether the fifth device and the first device are the same device.
[0182] Through the various steps of stage 4, the multi-device collaboration system 10 can determine the referential object that bears the first operation. In the embodiments of this application, among the devices that receive the first message in S107, the device where the referential object is located can be called the second device. Among the devices that send the third message to the first device in S109, the third message sent by the second device can also be called the second message.
[0183] Phase 5 (S113): The first device determines a device that performs the first operation.
[0184] S113: The first device determines a device for performing the first operation.
[0185] The device determined by the first device to perform the first operation may also be referred to as a third device.
[0186] In some embodiments, the first device determines the device that performs the first operation through the capabilities of the electronic devices in the multi-device collaborative system 10. The first device determines the device that performs the first operation, including but not limited to determining the device that performs the first operation through any one or more of the following: the capabilities of the device, the operating status of the device, and the distance between the device and the user. For example, in a smart home scenario, the multi-device collaborative system 10 includes a mobile phone, a smart speaker, and a smart screen. When the mobile phone obtains the user's intention to perform a playback operation, the mobile phone obtains that the display capability of the smart screen is the strongest, and then determines that the smart screen performs the playback operation. This application does not limit the specific manner in which the first device determines the device that performs the first operation.
[0187] In some implementations, the first device may also determine the executing device through the first operation. In a driving scenario, if the user intends to obtain navigation services and the mobile phone is pre-set with a rule that prioritizes providing navigation services on the vehicle computer, the vehicle computer is identified as the third device providing navigation services.
[0188] In some embodiments, if the voice command received by the first device includes a device that performs the first operation, the first device identifies the device as the third device. For example, the voice command is "Mobile phone, navigate here," where "mobile phone" indicates that the mobile phone in the multi-device collaboration system 10 provides navigation services. This application does not limit the specific content of the voice command.
[0189] Phase 6 (S114-S115): The third device performs the first operation.
[0190] S114: The first device instructs the third device to perform a first operation on the referred object.
[0191] In some embodiments, the first device instructs the third device to perform the first operation on the reference object, including any one or more of the following methods: 1. The first device first replaces the reference word in the voice instruction with the reference object determined in stage 4 to obtain a precise voice instruction, and then sends the precise voice instruction to the third device. For example, the first device sends the precise voice instruction "Navigate to ** Road" to the third device, and the third device receives the precise voice instruction and executes the precise voice instruction through the voice assistant. 2. The first device generates a new message based on the aforementioned first operation obtained and the reference object that undergoes the first operation, and the message is used to instruct the third device to perform the first operation on the reference object. The first device sends the message to the third device, and the third device receives the message and runs the application corresponding to the first operation to perform the first operation. This application does not limit the specific method in which the first device instructs the third device to perform the first operation.
[0192] In some embodiments, the voice instruction may also include execution conditions for the first operation. After determining that the execution conditions for the first operation are met, the first device sends the instruction information to the third device. The execution conditions for the first operation may include any one or more of the following: time conditions, location conditions, and environmental conditions. For a detailed description of the execution conditions for the first operation, please refer to the description in the following three paragraphs. This application does not limit the specific content of the execution conditions for the first operation.
[0193] The specific process of the first device determining whether the time condition is met is as follows: the first device obtains the current time and compares it with the time in the voice instruction. If the current time meets the time condition in the voice instruction, the third device is instructed to perform the first operation on the referred object. Otherwise, the first device waits for the current time to meet the time condition in the voice instruction before performing the above operation. For example, the voice instruction is "Call him at 12 noon", and the mobile phone can obtain the execution condition in the voice instruction as the time condition, and the time in the voice instruction is 12 noon. The mobile phone obtains the current time. If the current time is 12 noon, the third device is instructed to perform the dialing operation. If the current time is not 12 o'clock (such as 11 o'clock), the third device is instructed to perform the dialing operation after the current time reaches 12 noon.
[0194] The specific process of the first device determining whether the location condition is met is as follows: the first device obtains the current location and compares it with the location in the voice instruction. If the current location meets the location condition in the voice instruction, the first device instructs the third device to perform the first operation on the referred object. Otherwise, the first device waits for the current location to meet the location condition in the voice instruction before performing the above operation. For example, the voice instruction is "Call him after arriving at ** Road". The mobile phone can obtain the execution condition in the voice instruction as the location condition, and the location in the voice instruction is ** Road. The mobile phone obtains the current location. If the current location obtained by the mobile phone is ** Road, the mobile phone instructs the third device to perform the dialing operation. If the current location obtained by the mobile phone is not ** Road, the mobile phone waits until it arrives at ** Road before instructing the third device to perform the dialing operation.
[0195] The specific process of the first device determining whether the environmental conditions are met is as follows: the first device obtains information about the current environment and compares it with the information about the environment in the voice instruction. If the information about the current environment meets the environmental conditions in the voice instruction, the third device is instructed to perform the first operation on the referred object. Otherwise, the first device waits for the information about the current environment to meet the environmental conditions in the voice instruction before performing the above operation. For example, the voice instruction is "Turn on the wipers when it rains", and the execution condition obtained by the mobile phone in the voice instruction is the environmental condition, and the environmental condition in the voice instruction is rainy. The mobile phone obtains information about the current environment, including but not limited to weather information and humidity information. If the weather information obtained by the mobile phone shows that it is raining, and the humidity value obtained by the mobile phone is relatively large, the mobile phone instructs the third device to turn on the wipers. Otherwise, the mobile phone waits to obtain weather information showing that it is raining, and obtains a relatively large humidity value before performing the above operation.
[0196] In some embodiments, the first device may also send the execution conditions of the first operation to the third device, including any one or more of the following methods: 1. The first device carries the execution conditions of the first operation in a precise voice instruction and sends it to the third device. The precise voice instruction is obtained by the first device replacing the referent in the voice instruction with the referent object determined in stage 4. The precise voice instruction contains the execution conditions of the first operation. For example, the first device carries the time condition in the precise voice instruction "Navigate to ** Road at 12 noon" and sends it to the third device. 2. The first device generates a new message based on the aforementioned execution conditions and sends the message to the third device. The message contains the execution conditions. This application does not limit whether the first device sends the execution conditions of the first operation to the third device.
[0197] S115: The third device performs a first operation on the referred object.
[0198] In some embodiments, the third device may execute the first operation in any one or more of the following ways: 1. The third device launches a voice assistant and executes a voice command containing the first operation through the voice assistant. For example, the vehicle computer may receive a voice command such as "Navigate to Route **" through the voice assistant and initiate a navigation operation to a map application. 2. The third device launches an application that supports the first operation and executes the first operation within that application. For example, the vehicle computer may send a navigation operation directly to the map application via a received message.
[0199] In some embodiments, after executing the first operation, the third device may further output a prompt message to the user, which may include a voice prompt, text, or an interface display. This application does not limit the type of prompt message output by the third device after executing the first operation.
[0200] In some implementations, the third device may also receive the execution condition sent by the first device.
[0201] In some embodiments, after receiving the execution condition sent by the first device, the third device determines whether the execution condition is satisfied. If the third device determines that the execution condition is satisfied, it executes the first operation. Otherwise, the third device waits to execute the first operation until it determines that the execution condition is satisfied. For a description of the third device determining whether the execution condition is satisfied, refer to the detailed description of the first device determining whether the execution condition is satisfied in step S114, and will not be repeated here.
[0202] FIG. 5B shows another voice interaction method including pronouns.
[0203] In another embodiment, the multi-device collaborative system 10 includes a central control device that uniformly schedules multiple devices in the multi-device collaborative system 10. The central control device can select a device in the multi-device collaborative system 10 to collect a user's input voice command containing a pronoun, and the central control device can also obtain the referent corresponding to the pronoun from certain devices, and the central control device can also select a device to respond to the voice command. The multiple devices in the multi-device collaborative system 10 are used to determine the central control device from the multiple devices in any of the following situations:
[0204] (1) The multi-device collaborative system 10 can determine the central control device under user triggering. For example, a user can input an operation on a central device (e.g., a mobile phone) in the multi-device collaborative system 10, triggering the central device to notify other devices in the multi-device collaborative system 10 via broadcast or other means to determine the central control device.
[0205] (2) The multi-device collaborative system 10 may determine the central control device periodically or aperiodically according to a preset rule. For example, each device in the multi-device collaborative system 10 may determine the central control device once a week or once a month.
[0206] (3) When a device joins or leaves the multi-device collaborative system 10. When a device goes online or offline, the old central control device can be used without re-determining the central control device. When the old central control device goes offline, each electronic device in the multi-device collaborative system 10 re-elects a central control device.
[0207] (4) After the preset duration of multiple devices forming the multi-device collaborative system 10. That is, after forming the multi-device collaborative system 10, the multi-device collaborative system 10 can delay determining the central control device, so that the multi-device collaborative system 10 can collect more comprehensive device information to elect a central control device, so as to elect a more suitable central control device. The preset duration can be set according to actual needs, and the embodiment of the present application is not limited to this. For example, the preset duration can be set to 10 seconds, 1 minute, 1 hour, 12 hours, 1 day, 2 days, 3 days, etc.
[0208] Strategies for multiple devices in the multi-device collaborative system 10 to determine the central control device may include the following:
[0209] Strategy 1: Select one or more devices as central control devices from multiple electronic devices in the multi-device collaborative system 10 based on one or more factors such as the stability of resources, the information interaction methods provided or supported by the devices, and the devices commonly used by users. Resources may include but are not limited to camera resources, microphone resources, sensor resources, display resources, or computing resources. Information interaction methods may include voice interaction, display interaction, light interaction, vibration interaction, and the like. For example, devices with relatively stable computing resources, devices with relatively stable memory resources, devices with relatively stable power supplies, devices with more information interaction methods provided or supported, or devices commonly used by users can be determined as central control devices.
[0210] Strategy 2: Determine a preset device among multiple devices as the central control device. For example, if the preset device is a smart screen, then the smart screen is determined as the central control device.
[0211] Strategy 3: Determine the device selected by the user as the central control device. This allows you to determine the central control device based on the user's actual needs.
[0212] Strategy 4: Determine the central control device from multiple devices based on the historical interaction information of each device.
[0213] In some embodiments, the historical interaction information of the device may include, but is not limited to, any one or more of the following: device identification, device type, power consumption, available resources, information interaction mode, usage status, online information, offline information, device location (such as room, living room, etc.), orientation, and the type of environment in which the device is located (such as office, home range, etc.). For example, the electronic device in the multi-device collaborative system 10 can count the number of devices online in a unit time (such as one day, one week, etc.), and determine the electronic device with the largest average value of the number of devices as the central control device. If the same device goes online multiple times in a unit time, it can be counted only once, and the number of times it goes online is not accumulated. For example, assuming that the statistical time period is from February 1 to February 7, the number of online devices counted by the electronic device during the statistical time period is shown in Table 1:
[0214] Table 1
[0215] At this time, the average number of online devices per day is (3+4+5+2+6+8+7) / 7=5.
[0216] The usage status may include, for example, the currently enabled application or hardware of the device, etc. The online information may include the number, time, and duration of the electronic device's online access, etc. Similarly, the offline information may include the number, time, and duration of the electronic device's offline access, etc.
[0217] The various electronic devices in the multi-device collaborative system 10 can negotiate, elect, make decisions or determine the central control device based on the connection between the devices through broadcasting, multicasting, querying, etc.
[0218] The number of central control devices determined by the multiple devices in the multi-device collaborative system 10 includes multiple central control devices, and multiple central control devices are connected to all devices in the multi-device collaborative system 10 at the same time or in the same space. In this way, the central control device can directly interact with other devices in the multi-device collaborative system 10, thereby fully utilizing the information of each device to provide services to users.
[0219] In the above-mentioned voice interaction method including pronouns, the method may include the following steps:
[0220] Phase 1 ( S201 ): The central control device synchronizes device information with devices in the multi-device collaboration system 20 .
[0221] In step S201 , after the multi-device collaborative system 10 is constructed, the central control device obtains device information of the devices in the multi-device collaborative system 10 .
[0222] Stage 2 (S202-S205) and Stage 3 (S206-S207): The central control device wakes up the voice assistant, and the user inputs a voice command containing a pronoun.
[0223] In step S202, the devices in the multi-device collaborative system 10 collect the wake-up word input by voice and send a message carrying the wake-up word to the central control device.
[0224] In steps S203-S204, the central control device obtains the voice input wake-up word collected by the devices in the multi-device collaborative system 10, determines the device that responds to the user, and sends a message to the device that responds to the user. The device that responds to the user is also called the fourth device. The central control device also instructs the device to turn on the voice assistant and respond to the user. Among them, how the central control device determines the device that responds to the user can refer to the description of each device in the multi-device collaborative system 10 in the aforementioned step S103, which will not be repeated here.
[0225] In steps S205-S207, the device responding to the user activates the voice assistant and responds to the user's voice input. When the device receives the user's voice command, it sends the voice command (i.e., the first voice) to the central control device. The voice command includes the first operation and the reference word that receives the first operation.
[0226] Phase 4 (S208-S214): The central control device determines the object to be subjected to the first operation.
[0227] In step S208, the central control device receives the voice command. Subsequently, the central control device sends a message to multiple devices in the multi-device collaborative system 10, querying for candidate objects that can undergo the first operation. This message is also referred to as the first message. For details on how the central control device sends the first message, please refer to the description of the first device in step S107 above and will not be repeated here.
[0228] In steps S209-S210, multiple devices in the multi-device collaborative system 10 that receive the first message query candidate objects locally. For the device that queries the candidate object, the information of the candidate object is returned to the central control device. The candidate object is an object that can withstand the first operation, and the returned message is also called the third message. For the device that queries the candidate object, it can include a sixth device and a second device. Furthermore, there is a reference object in the candidate object, and the device where the reference object is located can be called the second device. The message returned by the second device is also called the second message. Among them, how the multiple devices that receive the first message query the candidate object locally can refer to the relevant description in the aforementioned step S108. At the same time, for how the device that queries the candidate object returns the third message, refer to the relevant description in the aforementioned step S109, which will not be repeated here.
[0229] In steps S211-S212, the central control device obtains the candidate object returned in step S210. If there is one candidate object, the candidate object is the reference object that undergoes the first operation. If there are multiple candidate objects, the central control device determines the device (i.e., the fifth device) that selects the reference object from the multiple candidate objects, and sends information about the multiple candidate objects to the fifth device, waiting to receive information about the reference object returned by the device. The multiple candidate objects can come from one or more devices, or one or more objects from one device. Among them, how the central control device determines the fifth device and how the central control device sends information about the multiple candidate objects to the fifth device can refer to the relevant description of the first device in the aforementioned step S110, which will not be repeated here.
[0230] In step S213, the fifth device outputs multiple candidate objects and receives a user operation (such as a touch operation on the icons of the multiple candidate objects displayed). Based on the user operation, the fifth device determines the referent object that is subject to the first operation from the multiple candidate objects. The device then sends information about the determined referent object to the central control device. For details on how the fifth device determines the referent object, please refer to the relevant description in step S111 above and will not be repeated here.
[0231] In step S214, the fifth device sends information about the reference object to the central control device. After receiving the information about the reference object, the central control device determines the reference object in the user voice command that is subject to the first operation.
[0232] Stage 5 (S215) and Stage 6 (S216-S217): The central control device determines the device that performs the first operation, which is the third device, and the third device performs the first operation.
[0233] In steps S215-S216, the central control device determines a third device for performing the first operation and instructs the third device to perform the first operation on the referenced object. For details on how the central control device determines the third device for performing the first operation, refer to the description of the first device in step S113 above. For details on how the central control device instructs the third device to perform the first operation on the referenced object, refer to the description of the first device in step S114 above, which will not be repeated here.
[0234] In step S217, the third device performs a first operation on the referred object.
[0235] In the above implementation scheme, the central control device and the second device, third device, fourth device, fifth device, and sixth device can be different devices or the same device. When the central control device and the above devices are the same device, the central control device sending messages to the corresponding above devices and receiving messages returned by the corresponding above devices can be regarded as interactions between the central control device and the modules that can execute tasks corresponding to the above devices. This application does not limit whether the central control device and the above devices are the same device.
[0236] Based on the voice interaction method including pronouns described in FIG5A , the specific implementation in different scenarios is shown below.
[0237] Driving scenarios
[0238] Example 1: A scenario in which multiple devices provide navigation services to users through voice commands containing pronouns.
[0239] Refer to FIG6 , which shows the scenario in Example 1.
[0240] As shown in Figure 6(a), the user is currently in the driver's seat of a vehicle, representing a driving scenario. Mobile phone 100-1 and vehicle-mounted computer 100-3 constitute a multi-device collaborative system 10. Mobile phone 100-1 displays the address of the user's driving destination, "Shidai Community." The user intends for vehicle-mounted computer 100-3 to perform a navigation operation, with the destination address displayed on mobile phone 100-1. The user speaks a wake-up word and a voice command containing a pronoun, "Xiaoyi Xiaoyi, navigate to this address." Both mobile phone 100-1 and vehicle-mounted computer 100-3 receive the wake-up word "Xiaoyi Xiaoyi." Mobile phone 100-1 and vehicle-mounted computer 100-3 determine that vehicle-mounted computer 100-3 should respond to the user. Vehicle-mounted computer 100-3 then receives the voice command containing the pronoun, "Navigate to this address." Vehicle-mounted computer 100-3 then sends the collected voice command containing the pronoun to mobile phone 100-1, which performs semantic analysis on the voice command containing the pronoun.
[0241] As shown in Figure 6(b), mobile phone 100-1 obtains the destination address displayed by its own map application and uses the displayed address of "Times Community" as the reference for "this address." Mobile phone 100-1 determines that the device performing the navigation operation is vehicle-mounted device 100-3. Mobile phone 100-1 instructs vehicle-mounted device 100-3 to navigate to Time Community. Vehicle-mounted device 100-3 displays the navigation route and voice assistant prompts on its main display, satisfying the user's request to use vehicle-mounted device 100-3's main display for navigation.
[0242] Based on the scenario provided in Example 1 in which multiple devices provide navigation services to users through voice instructions containing pronouns, the voice assistant in the vehicle computer 100-3 can obtain the navigation address information displayed by the mobile phone 100-1, and determine the voice instructions containing pronouns received by the vehicle computer 100-3 as clear voice instructions through information across devices (such as the mobile phone 100-1), and execute them on the device that meets the user's needs, thereby improving the efficiency of human-computer voice interaction.
[0243] Example 2: A scenario in which multiple devices provide video playback services to users through voice commands containing pronouns.
[0244] Refer to FIG. 7 , which shows a scenario in Example 2. FIG.
[0245] As shown in Figure 7(a), the user is currently in the passenger seat of a vehicle, representing a driving scenario. Mobile phone 100-1 and vehicle-mounted computer 100-3 constitute a multi-device collaborative system 10. The user is in the passenger seat, and a video titled "Brave Step" is displayed on mobile phone 100-1. The user wishes to view the video displayed on mobile phone 100-1 on vehicle-mounted computer 100-3. The user speaks a wake-up word and a voice command containing a pronoun, "Xiaoyi Xiaoyi, play this video." Both mobile phone 100-1 and vehicle-mounted computer 100-3 receive the wake-up word "Xiaoyi Xiaoyi." Mobile phone 100-1 and vehicle-mounted computer 100-3 determine that vehicle-mounted computer 100-3 should respond to the user, and then vehicle-mounted computer 100-3 receives the voice command containing the pronoun, "Play this video." Vehicle-mounted computer 100-3 then sends the collected voice command containing the pronoun to mobile phone 100-1, which performs semantic analysis on the voice command containing the pronoun. Among them, for the description of the user waking up the voice assistant and the user inputting a voice command containing a pronoun, please refer to the aforementioned steps S102-S106.
[0246] As shown in (b) of FIG7 , the mobile phone 100-1 obtains the video displayed by the device's own video application, and uses the displayed "Step Bravely" video as the reference object of "this video" in the voice instruction containing the pronoun. The mobile phone 100-1 determines that the device that performs the video playback operation is the vehicle computer 100-3. For the relevant description of the mobile phone 100-1 determining that the vehicle computer 100-3 performs the video playback service, please refer to the aforementioned step S113. The mobile phone 100-1 instructs the vehicle computer 100-3 to perform the video playback operation, and plays the video in the mobile phone 100-1 on the secondary display screen of the vehicle computer 100-3, which meets the user's requirement to watch the video on the secondary display screen of the vehicle computer 100-3.
[0247] Based on the scenario provided in Example 2 where multiple devices provide video playback services to users through voice commands containing pronouns, the voice assistant in the car computer 100-3 can obtain information about the video played by the mobile phone 100-1, thereby converting the voice commands containing pronouns received by the car computer 100-3 into clear voice commands, satisfying the user's need to watch videos on the secondary display screen of the car computer 100-3 and improving the efficiency of human-computer voice interaction.
[0248] Example 3: A scenario in which multiple devices provide dialing services to users through voice commands containing pronouns.
[0249] Refer to FIG8 , which shows the scenario in Example 3.
[0250] As shown in (a) of Figure 8 , the user is currently in the driver's seat of a vehicle, representing a driving scenario. Mobile phone 100-1 and vehicle-mounted device 100-3 constitute a multi-device collaborative system 10. Mobile phone 100-1 displays the user's contact "Andy," and vehicle-mounted device 100-3 is performing navigation. At this point, the user wishes for vehicle-mounted device 100-3 to call the contact displayed on mobile phone 100-1 upon arrival at the destination. The user speaks a wake-up word and a voice command containing a pronoun, "Xiaoyi Xiaoyi, wait until I get to my destination and call this person." Both mobile phone 100-1 and vehicle-mounted device 100-3 receive the wake-up word "Xiaoyi Xiaoyi." Mobile phone 100-1 and vehicle-mounted device 100-3 determine that vehicle-mounted device 100-3 should respond to the user, and vehicle-mounted device 100-3 receives the voice command containing the pronoun, "Wait until I get to my destination and call this person." The vehicle computer 100 - 3 sends the collected voice instructions containing the pronouns to the mobile phone 100 - 1 , and the mobile phone 100 - 1 performs semantic analysis on the voice instructions containing the pronouns.
[0251] As shown in Figure 8(b), mobile phone 100-1 retrieves the contact information displayed in its own contacts app and uses the displayed "Andy" as the referent for "this person" in the voice command containing a pronoun. Mobile phone 100-1 then learns from the navigation operation being executed by vehicle-mounted computer 100-3 that the navigation destination is "Shidai Community" and uses this address as the referent for "destination." Mobile phone 100-1 detects that the current scene is driving and determines that vehicle-mounted computer 100-3 is the device performing the dialing operation. Before instructing vehicle-mounted computer 100-3 to dial "Andy," mobile phone 100-1 needs to determine whether the user's vehicle has reached "Shidai Community." If the user is still driving and has not reached their destination, mobile phone 100-1 can instruct vehicle-mounted computer 100-3 to reply to the user, "OK, I'll call Andy when I reach Shidai Community." If the user reaches their destination, vehicle-mounted computer 100-3 executes the call to Andy, satisfying the user's desire to dial using vehicle-mounted computer 100-3.
[0252] Based on the scenario provided in Example 3 where multiple devices provide dialing services to users through voice instructions containing pronouns, the voice assistant in the vehicle computer 100-3 can obtain the contact information displayed by the mobile phone 100-1, and determine the voice instructions containing pronouns received by the vehicle computer 100-3 into clear voice instructions through the information of the mobile phone 100-1, and execute them on the device that meets the user's needs, thereby improving the efficiency of human-computer voice interaction.
[0253] Smart home scenarios
[0254] Example 4: A scenario in which multiple devices provide video playback services to users through voice commands containing pronouns.
[0255] Refer to FIG9 , which shows a scenario in Example 4.
[0256] As shown in (a) of Figure 9, the current user is in an indoor scene at home, which belongs to the smart home scene. Mobile phone 100-1, smart speaker 100-6 and smart screen 100-5 constitute a multi-device collaborative system 10. The video "Brave Step" played by the user is displayed on mobile phone 100-1. At this time, the user wants to watch the video displayed by mobile phone 100-1 on the car smart screen 100-5. The user says the wake-up word and the voice command containing the pronoun "Xiaoyi Xiaoyi, play this video". Mobile phone 100-1, smart speaker 100-6 and smart screen 100-5 all receive the wake-up word "Xiaoyi Xiaoyi". Mobile phone 100-1, smart speaker 100-6 and smart screen 100-5 determine that mobile phone 100-1 responds to the user, and smart speaker 100-6 receives the voice command containing the pronoun "Play this video". The smart speaker 100 - 6 sends the acquired ambiguous voice to the mobile phone 100 - 1 , and the mobile phone 100 - 1 performs semantic analysis on the voice command containing the pronoun.
[0257] As shown in (b) of Figure 9, mobile phone 100-1 obtains the video displayed by the device's own video application and uses the displayed "Bravely Step Forward" video as the reference object of "this video" in the voice instruction containing the pronoun. Mobile phone 100-1 obtains that the current scene is a smart home scene and determines that the device that performs the video playback operation is smart screen 100-5. Mobile phone 100-1 instructs smart screen 100-5 to perform the video playback operation. Smart screen 100-5 plays the video "Bravely Step Forward" in mobile phone 100-1 on the display screen, satisfying the user's desire to use smart screen 100-5 to watch videos.
[0258] Based on the scenario provided in Example 4, multiple devices provide users with video playback services through voice commands containing pronouns. The voice assistant in mobile phone 100-1 can obtain the video information played by mobile phone 100-1, thereby sending clear voice commands to smart screen 100-5, satisfying the user's needs for watching videos on smart screen 100-5 and improving the efficiency of human-computer voice interaction.
[0259] Smart office scene
[0260] Example 5: A scenario in which multiple devices provide file operation services to users through voice commands containing pronouns.
[0261] Refer to FIG10 , which shows the scenario in Example 5.
[0262] As shown in Figure 10(a), the user is currently in an indoor office environment, representing a smart office scenario. Mobile phone 100-1, laptop computer 100-4, and projector 100-7 constitute a multi-device collaborative system 10. Mobile phone 100-1 displays the user's open file "Notepad," while laptop computer 100-4 displays the desktop through a projector. The user wishes to view the file displayed on mobile phone 100-1 on the projector. The user speaks a wake-up word and a voice command containing a pronoun, "Xiaoyi Xiaoyi, open this file." Both mobile phone 100-1 and laptop computer 100-4 receive the wake-up word "Xiaoyi Xiaoyi." Mobile phone 100-1, laptop computer 100-4, and projector 100-7 determine that mobile phone 100-1 is the user's recipient. Mobile phone 100-1 then receives the voice command containing the pronoun, "Open this file." Mobile phone 100-1 then sends the voice command containing the pronoun to laptop computer 100-4, which performs semantic analysis on the voice command containing the pronoun.
[0263] As shown in Figure 10(b), laptop computer 100-4 retrieves the file displayed by mobile phone 100-1 and uses the displayed "Notebook" file as the reference for "this file" in the voice command containing the pronoun. Laptop computer 100-4 determines that the device performing the file operation is projector 100-7. Laptop computer 100-4 instructs projector 100-7 to perform the file display operation. Projector 100-7 displays the file "Notepad" on mobile phone 100-1, satisfying the user's request to display the file using the projector. For a detailed description of laptop computer 100-4 providing the video playback service via the projector, please refer to steps S114-S115 described above.
[0264] Based on the scenario provided in Example 5 where multiple devices provide file operation services to users through voice commands containing pronouns, the voice assistant in the laptop computer 100-4 can obtain the file information opened by the mobile phone 100-1 and display the file on the projector 100-7, thereby meeting the user's need to display the file on the projector 100-7 and improving the efficiency of human-computer voice interaction.
[0265] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a microservice architecture, or a cloud architecture. The embodiment of the present application takes the Android system with a layered architecture as an example to illustrate the software structure of the electronic device 100. The software architecture of the electronic device 100 can also be adopted for the electronic devices involved in this application. Only the electronic device 100 is used as an example for description here, and the software architecture of other electronic devices is not limited.
[0266] FIG. 11 exemplarily shows the software structure of the electronic device 100 .
[0267] As shown in Figure 11, the operating system installed on the electronic device 100 adopts a layered architecture. The layered architecture divides the operating system into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the operating system is divided into four layers, namely, the application layer, the application framework layer, the system service layer, and the kernel layer from top to bottom. In other embodiments, the operating system can also be divided into other number of layers, which is not limited in this application.
[0268] The application layer can include a series of applications such as system applications and extended applications (or third-party applications). The distributed implementation method of applications provided in this application is applicable to the capability call of each application (including system applications and extended applications) in the application layer. System applications include desktop, settings, camera, wireless LAN, Bluetooth, navigation, etc.; extended applications include voice assistants, navigation applications, video applications and other software applications developed by third parties.
[0269] The application framework layer provides a multi-language framework for the application layer, including a user interface (UI) framework, a user program framework, and a capability framework, as well as a multi-language framework application programming interface (API) and framework APIs for multiple programming languages. The application framework layer includes some pre-defined functions.
[0270] Among them, the UI framework includes the window manager, content provider, view system, phone manager, resource manager, notification manager, etc., which will not be described in detail here.
[0271] The user program framework and capability framework are multi-language frameworks provided by the application framework layer for applications. For example, they provide the capabilities of each device in the communication system required for the application, that is, these devices are used to implement various application functions of the application layer. The capability framework may include but is not limited to computing power (including CPU computing power, graphics processing unit (GPU) computing power, image signal processor (ISP) computing power, etc.), sound pickup capability (including microphone sound pickup capability, voice recognition capability, etc.), security capabilities in terms of device security protection (including trusted operating environment security level, etc.), display capability (including screen resolution, screen size, etc.), playback capability (including sound amplification capability, stereo sound capability, etc.), and storage capability (including device memory capability, random access memory (RAM) capability, etc.), etc., which are not limited in this application.
[0272] The system service layer is the core of the operating system. It provides services to applications in the application layer through the application framework layer. The system service layer includes a distributed device management module, a capability classification module, a distributed task decision module, a virtual device management module, and a distributed soft bus.
[0273] The distributed device management module is used to centrally manage static device information, such as configuration parameters, of electronic devices 100 interconnected via the underlying network, as well as dynamic device information, such as the current operational status data of each electronic device 100. As described above, the static and dynamic device information centrally managed by the distributed device management module is collected based on the underlying network (e.g., a distributed soft bus).
[0274] The capability grading module is used to compare the capabilities of the electronic devices 100 using a unified capability grading standard, obtain the capability strength relationship of different devices, and provide a decision basis for the distributed task decision module to make quick decisions.
[0275] The distributed task decision module is used to determine the electronic device that performs the operation, such as an electronic device that responds to the wake-up word, an electronic device that queries the candidate object, an electronic device that determines the reference object, an electronic device that performs the first operation, etc. The distributed task decision module can determine the electronic device that performs the operation by including any one or more of the following: electronic device capabilities, information on the type of electronic device, the distance between the electronic device and the user, and information on the operating status and environmental status of the electronic device. For example, the distributed task decision module compares and analyzes the capabilities of different electronic devices, and obtains the operating status of the electronic device, and finally determines an electronic device with stronger capabilities and less occupied resources in the device. Afterwards, the distributed task decision module sends a message (such as a first message) to the determined electronic device, which is used to instruct the electronic device to perform the operation contained in the information (such as the operation of querying the candidate object contained in the first message).
[0276] The virtual device management module is used to respond to the message sent by the distributed task decision module and instruct the electronic device to execute the operation contained in the message.
[0277] The distributed soft bus, as an example structure of the underlying network, is used to collect static device information such as configuration parameters of the electronic device 100, as well as dynamic device information of each electronic device 100, including current operating status data of the electronic device 100.
[0278] Among them, the understanding of the distributed soft bus function can refer to the computer hardware bus. For example, the distributed soft bus builds an "invisible" bus between 1+8+N devices (1 is a mobile phone; 8 represents car computers, speakers, headphones, watches, bracelets, tablets, large screens, personal computers (PCs), augmented reality devices, virtual reality devices; N refers to other IoT devices). It has the characteristics of automatic discovery, instant connection, self-organizing networking (heterogeneous network networking), high bandwidth, low latency, and high reliability. In other words, through distributed soft bus technology, different electronic devices can not only share all data, but also achieve instant interconnection with any device on the same local area network or connected to it via Bluetooth. In addition, the distributed soft bus can also share files between heterogeneous networks such as Bluetooth and wireless fidelity (Wi-Fi) networks (for example, receiving files via Bluetooth on the one hand and transmitting files via Wi-Fi on the other).
[0279] In the operating system, the UI framework, user program framework and capability framework in the above-mentioned application framework layer and the distributed device management module, capability classification module, distributed task decision module, virtual device management module and distributed soft bus in the system service layer can together constitute the system basic capability subsystem set, which is not restricted in this application.
[0280] The kernel layer is the layer between hardware and software. The kernel layer of the operating system includes: kernel subsystem and driver subsystem.
[0281] The kernel subsystem, which allows operating systems to adopt multi-core designs, supports the selection of appropriate OS kernels for different resource-constrained devices. The Kernel Abstraction Layer (KAL) above the kernel subsystem shields multi-core differences and provides upper layers with basic kernel capabilities, including process / thread management, memory management, file system management, network management, and peripheral management.
[0282] The driver subsystem operating system's driver framework is the foundation for the open distributed system hardware ecosystem, providing unified peripheral access capabilities and a driver development and management framework. The driver subsystem includes at least display drivers, camera drivers, audio drivers, and sensor drivers.
[0283] When the first device provided in this application is implemented as the software structure shown in Figure 11, the distributed task decision module of the first device is used to determine the third device. The distributed device management module of the first device provides a decision basis for the distributed task decision module of the first device. For example, if the first device obtains that the display capability of the third device is the strongest, it determines that the third device performs the playback operation. The distributed soft bus of the first device is used to support the first device to communicate with other devices and send a first message, and is also used to instruct the third device to perform the first operation, and is also used to receive the first voice sent by the fourth device to the first device, and is also used to receive the third message sent by the second device to the first device. As in the method on the first device side in Figure 5A.
[0284] When the other devices provided in this application (including but not limited to the second device, the third device, the fourth device, the fifth device, and the sixth device) have the software structure shown in FIG11 , the specific interaction process of the software structure of the other devices in executing the aforementioned steps S101-S115 can refer to the interaction process of the software structure of the aforementioned first device, which will not be repeated here.
[0285] Electronic equipment can be equipped or portable terminal devices with other operating systems, such as mobile phones, tablet computers, desktop computers, laptop computers, handheld computers, notebook computers, ultra-mobile personal computers (UMPCs), netbooks, as well as cellular phones, personal digital assistants (PDAs), AR devices, VR devices, AI devices, wearable devices, in-vehicle devices, smart home devices and / or smart city devices, etc.
[0286] Figure 12 shows the hardware structure of an electronic device provided in an embodiment of the present application. The electronic device is used to execute the methods described in the above embodiments 1 to 5.
[0287] For the electronic devices involved in this application, the hardware structure of the electronic device 100 may also be adopted. Here, only the electronic device 100 is used as an example for introduction, and no limitation is imposed on the hardware structures of other electronic devices.
[0288] The electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0289] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0290] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0291] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0292] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.
[0293] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0294] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0295] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.
[0296] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.) or displays an image or video through the display screen 194. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.
[0297] The wireless communication module 160 can provide wireless communication solutions including WLAN (such as Wi-Fi network), Bluetooth, global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. applied to the electronic device 100. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, demodulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.
[0298] In some embodiments, the antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with a network and other devices through wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS).
[0299] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0300] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD). The display screen panel can also be made of an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniLED, a microLED, a micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 can include one or N display screens 194, where N is a positive integer greater than one.
[0301] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.
[0302] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise and brightness. It can also optimize parameters such as exposure of the captured scene. In some embodiments, the ISP can be located within camera 193.
[0303] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 100 may include 1 or N cameras 193, where N is a positive integer greater than 1.
[0304] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.
[0305] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. This allows electronic device 100 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.
[0306] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU can enable intelligent cognitive applications in electronic device 100, such as image recognition, face recognition, speech recognition, and text comprehension.
[0307] The internal memory 121 may include one or more random access memories (RAM) and one or more non-volatile memories (NVM).
[0308] Random access memory may include static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM, for example, the fifth generation of DDR SDRAM is generally referred to as DDR5 SDRAM), etc.
[0309] Non-volatile memory may include disk storage devices and flash memory.
[0310] Flash memory can be divided into NOR FLASH, NAND FLASH, 3D NAND FLASH, etc. according to the operating principle; can be divided into single-level cell (SLC), multi-level cell (MLC), triple-level cell (TLC), quad-level cell (QLC), etc. according to the storage cell potential level; can be divided into universal flash storage (UFS), embedded multi media card (eMMC), etc. according to the storage specification.
[0311] The random access memory can be directly read and written by the processor 110, and can be used to store executable programs (such as machine instructions) of the operating system or other running programs, and can also be used to store user and application data.
[0312] The non-volatile memory may also store executable programs and user and application data, etc., and may be loaded into the random access memory in advance for direct reading and writing by the processor 110 .
[0313] The external memory interface 120 can be used to connect to an external non-volatile memory to expand the storage capacity of the electronic device 100. The external non-volatile memory communicates with the processor 110 via the external memory interface 120 to implement data storage. For example, files such as music and videos can be stored in the external non-volatile memory.
[0314] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.
[0315] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.
[0316] The speaker 170A, also called a "speaker", is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or listen to hands-free calls through the speaker 170A.
[0317] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the electronic device 100 receives a call or a voice message, the user can place the receiver 170B close to the ear to hear the voice.
[0318] Microphone 170C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the source of sound, realize directional recording function, etc.
[0319] The headphone jack 170D is used to connect a wired headphone and can be the USB interface 130 or a 3.5mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0320] When the first device provided in the present application is implemented as the electronic device structure shown in Figure 12, the processor 110 is used to call the corresponding software and hardware modules to execute the various steps implemented on the first device side of the method of Figure 5A. The display screen 194 can display the interface of the electronic device as described in Figures 6 to 10. Specifically, the display screen 194 can be used to display the reference object that undergoes the first operation. The wireless communication module 160, the mobile communication module 150, the antenna 1, the antenna 2, etc. can be used to receive voice instructions sent by the fourth device, and can also be used to send a first message to other devices in the multi-device collaborative system 10, and can also be used to receive a third message sent by other devices in the multi-device collaborative system 10, and can also be used to instruct a third device to perform the first operation. For a detailed introduction to the interface of the first device, please refer to the previous introduction to Figures 6 to 10, which will not be repeated here.
[0321] When the other devices provided in this application (including but not limited to the second device, the third device, the fourth device, the fifth device, and the sixth device) have the hardware structure shown in FIG12 , the specific interaction process of the hardware structure of the other devices in executing the aforementioned steps S101-S115 can refer to the interaction process of the hardware structure of the aforementioned first device, which will not be repeated here.
[0322] It should be understood that each step in the above method embodiments provided herein can be implemented by hardware integrated logic circuits in a processor or by software instructions. The method steps disclosed in the embodiments of this application can be directly implemented as being executed by a hardware processor, or by a combination of hardware and software modules in a processor.
[0323] The present application also provides an electronic device, which may include: a memory and a processor. The memory may be used to store a computer program; the processor may be used to call the computer program in the memory so that the electronic device executes the method in any one of the above embodiments.
[0324] The present application also provides a chip system, which includes at least one processor for implementing the functions involved in the method executed by the electronic device in any of the above embodiments.
[0325] In one possible design, the chip system further includes a memory, which is used to store program instructions and data, and the memory is located inside or outside the processor.
[0326] The chip system can be composed of chips, or can include chips and other discrete devices.
[0327] Optionally, there may be one or more processors in the chip system. The processor may be implemented in hardware or software. When implemented in hardware, the processor may be a logic circuit, an integrated circuit, etc. When implemented in software, the processor may be a general-purpose processor implemented by reading software code stored in a memory.
[0328] Optionally, the memory in the chip system may be one or more. The memory may be integrated with the processor or may be provided separately from the processor, which is not limited in the embodiments of the present application. For example, the memory may be a non-transient processor, such as a read-only memory (ROM), which may be integrated with the processor on the same chip or provided on different chips. The embodiments of the present application do not specifically limit the type of memory or the configuration of the memory and the processor.
[0329] Exemplarily, the chip system may be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processor unit (CPU), a network processor (NP), a digital signal processor (DSP), a microcontroller unit (MCU), a programmable logic device (PLD) or other integrated chips.
[0330] The present application also provides a computer program product, which includes: a computer program (also referred to as code, or instruction), which, when executed, enables a computer to execute the method executed by the electronic device in any of the above embodiments.
[0331] The present application also provides a computer-readable storage medium storing a computer program (also referred to as code or instruction). When the computer program is executed, the computer executes the method executed by the electronic device in any of the aforementioned embodiments.
[0332] In the description of the embodiments of the present application, unless otherwise specified, “ / ” means or, for example, A / B can mean A or B; “and / or” in the text is merely a way to describe the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0333] In the description of the embodiments of the present application, the terms "first" and "second" are used for descriptive purposes only and should not be understood as implying or suggesting relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more.
[0334] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.
[0335] The term "user interface (UI)" in the above embodiments of the present application refers to a medium interface for interaction and information exchange between an application or operating system and a user, which realizes the conversion between the internal form of information and the form acceptable to the user. The user interface is a source code written in a specific computer language such as Java and extensible markup language (XML). The interface source code is parsed and rendered on an electronic device and finally presented as content that the user can recognize. The commonly used form of user interface is graphical user interface (GUI), which refers to a user interface related to computer operations that is displayed in a graphical manner. It can be a visual interface element such as text, icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, widgets, etc. displayed on the display screen of an electronic device.
[0336] The various implementation modes of this application can be combined arbitrarily to achieve different technical effects.
[0337] In the aforementioned embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in this application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive).
[0338] Those skilled in the art will appreciate that all or part of the processes in the aforementioned method embodiments can be implemented by a computer program instructing the relevant hardware. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the aforementioned method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
[0339] In short, the above description is only an embodiment of the technical solution of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made based on the disclosure of this application should be included in the scope of protection of this application.
Claims
1. A voice interaction method including pronouns, characterized in that: Applied to a first device, the method includes: Acquire a first voice, where the first voice indicates a first operation and further includes a first pronoun that is subject to the first operation; determining, from a second device, a referent of the first pronoun; Instruct a third device to perform the first operation on the referenced object, the third device being different from the second device.
2. The method according to claim 1, characterized in that The obtaining of the first voice specifically includes: The first device receives the first voice sent by the fourth device, and the first voice is collected by the fourth device.
3. The method according to claim 2, characterized in that Before the first device receives the first voice sent by the fourth device, the fourth device collects the wake-up word.
4. The method according to claim 1, wherein The obtaining of the first voice specifically includes: The first device collects the first voice.
5. The method according to claim 4, characterized in that Before the first device collects the first voice, the method further includes: The first device collects a wake-up word.
6. The method according to any one of claims 1 to 5, characterized in that The second device and the first device are the same device.
7. The method according to any one of claims 1 to 5, characterized in that The first device does not include the referenced object, and the second device is a different device from the first device.
8. The method according to any one of claims 1 to 7, characterized in that Determining the referent of the first pronoun from the second device specifically includes: The first device sends a first message to the second device; The first device receives a second message sent by the second device, where the second message includes information of the referenced object. The referenced object is determined by the second device in response to the first message from the storage content and / or displayed interface content of the second device.
9. The method according to claim 8, characterized in that The first message includes the type of the object subject to the first operation, and the object referred to by the first pronoun belongs to the type of the object subject to the first operation.
10. The method according to claim 8 or 9, characterized in that The step of sending, by the first device, a first message to the second device specifically includes: The first device sends a first message to a plurality of devices, where the plurality of devices include the second device; The step of receiving, by the first device, a second message sent by the second device specifically includes: The first device receives a third message sent by some or all of the multiple devices, where the third message includes information of one or more candidate objects, and the one or more candidate objects are determined by the corresponding device from the stored content and / or displayed interface content. The third message sent by the second device is the second message.
11. The method according to claim 10, characterized in that After the first device receives the third message sent by some or all of the multiple devices, the method further includes: The first device determines a plurality of candidate objects from the received one or more third messages; The first device sends information about the multiple candidate objects to a fifth device, so that the fifth device outputs the multiple candidate objects; The first device receives the information of the reference object sent by the fifth device, and the reference object is selected by a user from the multiple candidate objects.
12. The method according to any one of claims 1 to 11, characterized in that The first voice indicates a device that performs the first operation, and the device that performs the first operation is the third device.
13. The method according to any one of claims 1 to 11, characterized in that The first voice does not indicate a device that performs the first operation, and the third device is determined by the first device according to the first operation.
14. The method according to any one of claims 1 to 11, characterized in that The first voice indicates the type of device performing the first operation, and the multiple devices detected by the first device all belong to the device type. The third device is determined by the first device from the multiple devices based on the first operation, or the third device is a device selected by the user from the multiple devices.
15. The method according to any one of claims 1 to 14, characterized in that The method further comprises: The first device uses the referent to replace the first referent in the first speech to obtain a first instruction; Instructing the third device to perform the first operation on the referred object specifically includes: The first device sends the first instruction to the third device, where the first instruction is used to instruct the third device to perform the first operation on the referred object.
16. The method according to any one of claims 1 to 15, characterized in that The first voice also includes an execution condition for the first operation, instructing the third device to perform the first operation on the referenced object, specifically including: instructing the third device to perform the first operation on the referred object after detecting that the execution condition is satisfied; Alternatively, after detecting that the execution condition is met, instruct the third device to perform the first operation on the referred object.
17. The method according to claim 16, characterized in that The execution conditions include: arriving at a first time and / or arriving at a first location.
18. The method according to any one of claims 1 to 17, characterized in that The first operation is any one of the following: a navigation operation, a play operation, or a phone call operation.
19. The method according to any one of claims 1 to 18, characterized in that The first device and the third device are the same device.
20. An electronic device, characterized in that: The electronic device includes one or more memories and one or more processors; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the electronic device to execute the method described in any one of claims 1-19.
21. A communication system, comprising: The first device, the second device, and the third device are characterized in that the first device is used to execute the method according to any one of claims 1 to 19.
22. A chip, applied to electronic equipment, characterized in that: The chip includes one or more processors, and the processors are configured to call computer instructions so that the electronic device executes the method according to any one of claims 1 to 19.
23. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 19.
24. A computer program product, characterized in that The computer program product comprises computer instructions, and when the computer instructions are run on an electronic device, the electronic device is caused to perform the method according to any one of claims 1 to 19.
Citation Information
Patent Citations
Semantic parsing method and server
CN110111787A
Command execution method, device and equipment
CN110798506A
Digital assistant hardware abstraction
CN112732622A
Voice interaction method, server and computer readable storage medium
CN115457959A
Voice instruction processing method, device and system and storage medium
CN117373445A