Cross-device user interaction method, device and cross-device interaction system

By using cross-device user interaction methods, the system can detect user location in real time and enable the flow of response content, solving the problem of users not being able to listen to response content after leaving the device and providing a better content listening experience.

CN120935247APending Publication Date: 2025-11-11QINDAO HAIER REFRIGERATOR CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410584858.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-11
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing smart devices continue to output content even after the user leaves the vicinity of the device, but the user cannot hear it, resulting in a less than ideal content listening experience.

Method used

By using cross-device user interaction methods, it can detect in real time whether a user enters or leaves the effective range of a device. When a user leaves the effective range of one device and enters the effective range of another device, it can realize the flow of response content, ensuring that the user can hear the response content without interruption.

Benefits of technology

This allows users to continue listening to content even after leaving the vicinity of the device, providing a better content listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120935247A_ABST
    Figure CN120935247A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent device interaction, and discloses a cross-device user interaction method and device, and a cross-device interaction system, for a first user who has been in an interaction binding state with a first device, under the condition that the first device outputs response content in response to first voice interaction information sent by the first user, the response content is sent to the first device. When the first user exceeds the effective range of the first device and enters the effective range of the second device, if the first user meets the interactive binding condition with the second device, the second device interactively binds with the first user, and meanwhile, the first device stops outputting the response content and flows the first response content to the second device. Therefore, the second device can output the remaining content, which is not output by the first device, in the first response content. Therefore, after the first user leaves the nearby area of the first device, the first user can continuously hear the response content, and better content listening experience is provided for the first user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the technical field of smart device interaction, such as a cross-device user interaction method, device, and cross-device interaction system. Background Technology

[0002] Currently, devices equipped with intelligent voice modules (such as smart speakers and smart refrigerators) are increasingly appearing in users' homes. These devices can respond to users' voice interaction information and output response content (such as music and news), providing users with a good service experience.

[0003] However, in related technologies, during the process of a device responding to a user's voice interaction, if the user leaves the vicinity of the device, although the device continues to output response content, the user cannot hear it. In these technologies, the device's output response content cannot follow the user's real-time location, thus failing to provide a better listening experience.

[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or describe the scope of protection of these embodiments, but rather as a prelude to the detailed description that follows.

[0006] This disclosure provides a cross-device user interaction method, device, and cross-device interaction system that allows users to continuously listen to response content even after leaving the vicinity of the device, providing users with a better content listening experience.

[0007] According to a first aspect of this disclosure, a cross-device user interaction method is provided, applied to a first device, wherein the first device and a first user are in an interaction-bound state, comprising:

[0008] In response to the first voice interaction message from the first user, output the first response content;

[0009] If, during the duration of the response, it is determined that the first user has exceeded its own valid scope, it is determined whether a content transfer notification has been received;

[0010] Upon receiving a content transfer notification, the system unbinds the interaction with the first user and stops outputting the first response content, and transfers the first response content to the second device so as to output the remaining content in the first response content that was not output by the first device. The first user is currently in an interactive binding state with the second device.

[0011] If no content flow notification is received, continue to output the first response content.

[0012] In some embodiments, even without receiving a content streaming notification, the first response content continues to be output, including:

[0013] In the absence of a content transfer notification, the system detects whether the first user sends a second voice interaction message and whether other users have entered its effective range.

[0014] If it is determined that the first user has not sent a second voice interaction message and other users have not entered its effective range, the first response content will continue to be output.

[0015] In some embodiments, the cross-device user interaction method further includes:

[0016] If it is determined that the first user has not issued a second voice interaction message and other users have been detected entering its effective range, the first response content will continue to be output.

[0017] Upon completion of the first response, the system continues to detect whether the first user has issued a second voice interaction message.

[0018] If no second voice interaction information is detected from the first user within the first preset time period, the system will interact and bind with the other user.

[0019] In some embodiments, after interacting and binding with the other user, the method further includes: receiving response content from other devices that have previously interacted and bound with the other user, and outputting the remaining content in the response content that was not output by the other device.

[0020] In some embodiments, the cross-device user interaction method further includes:

[0021] If it is determined that the first user has issued the second voice interaction information, the type of resource requested by the second voice interaction information shall be determined.

[0022] When the resource type is a streaming resource, do not respond to the second voice interaction information;

[0023] If the resource type is a non-streaming resource, determine whether it is the device with the smallest distance to the first user among all devices;

[0024] When it determines that it is the device with the smallest distance from the first user among all devices, it responds to the second voice interaction information and outputs the second response content.

[0025] According to a second aspect of this disclosure, a cross-device user interaction method is provided, applied to a second device, characterized in that it includes:

[0026] When the first user is detected to have entered its effective range, the feature information of the first user is obtained, and the first user is determined to meet the interaction binding conditions based on the feature information of the first user, wherein the first user is already in the interaction binding state with the first device.

[0027] When the first user meets the interaction binding conditions, the device interacts with the first user and sends a content transfer notification so that after receiving the content transfer notification, the first device unbinds from the first user and stops outputting the first response content, and transfers the first response content to the second device.

[0028] Receive the first response content from the first device and output the remaining content in the first response content that was not output by the first device.

[0029] In some embodiments, obtaining the characteristic information of a first user and determining whether the first user meets the interaction binding conditions based on the characteristic information of the first user includes:

[0030] Determine whether the first user has entered its own face recognition range, where the face recognition range is smaller than the effective range;

[0031] Once it is determined that the first user has entered the facial recognition range, the facial information of the first user is obtained and compared with the facial information in the database.

[0032] If the obtained facial information of the first user is consistent with the facial information of the first user in the database, the first user is determined to meet the interaction binding conditions.

[0033] In some embodiments, the cross-device user interaction method further includes:

[0034] If the first user's facial information does not exist in the database, or if it is determined that the obtained first user's facial information is inconsistent with the first user's facial information in the database, then obtain the first user's voiceprint information.

[0035] Once the voiceprint information of the first user is obtained, it is compared with the voiceprint information in the database.

[0036] If the obtained voiceprint information of the first user is consistent with the voiceprint information of the first user in the database, the first user is determined to meet the interaction binding conditions.

[0037] In some embodiments, the cross-device user interaction method further includes:

[0038] If it is determined that the first user has not entered the facial recognition range, obtain the voiceprint information of the first user;

[0039] Once the voiceprint information of the first user is obtained, it is compared with the voiceprint information in the database.

[0040] If the obtained voiceprint information of the first user is consistent with the voiceprint information of the first user in the database, the first user is determined to meet the interaction binding conditions.

[0041] According to a third aspect of this disclosure, a first device is provided, including a processor and a memory storing program instructions, the processor being configured to execute the cross-device user interaction method provided in the first aspect of this disclosure when running the program instructions.

[0042] According to a fourth aspect of this disclosure, a second device is provided, including a processor and a memory storing program instructions, the processor being configured to execute the cross-device user interaction method provided in the second aspect of this disclosure when running the program instructions.

[0043] According to a fifth aspect of this disclosure, a cross-device interaction system is provided, the cross-device interaction system including a first device provided in a third aspect of this disclosure and a second device provided in a fourth aspect of this disclosure.

[0044] The cross-device user interaction method, device, and cross-device interaction system provided in this disclosure can achieve the following technical effects:

[0045] The cross-device user interaction method provided in this disclosure allows each device in the cross-device interaction system to detect in real time whether a user enters or exceeds its effective range. For a first user already in an interactive binding state with a first device, when the first device outputs response content in response to the first user's first voice interaction information, if the first user exceeds the effective range of the first device and enters the effective range of a second device, and if the first user meets the interactive binding conditions with the second device, then the second device will interactively bind with the first user. Simultaneously, the first device will stop outputting response content and transfer the first response content to the second device, thereby enabling the second device to output the remaining content in the first response content that was not output by the first device. In this way, the first user can continue to listen to the response content even after leaving the vicinity of the first device, providing a better content listening experience.

[0046] The above general description and the description below are exemplary and illustrative only and are not intended to limit this disclosure. Attached Figure Description

[0047] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations and drawings do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are shown as similar elements. The drawings are not to be scaled. And wherein:

[0048] Figure 1 This is a schematic diagram of a cross-device interaction system provided in an embodiment of this disclosure;

[0049] Figure 2 This is a schematic diagram of another cross-device interaction system provided in an embodiment of this disclosure;

[0050] Figure 3 This is a flowchart illustrating a cross-device user interaction method provided in an embodiment of this disclosure;

[0051] Figure 4 This is a flowchart illustrating another cross-device user interaction method provided in an embodiment of this disclosure;

[0052] Figure 5 This is a flowchart illustrating another cross-device user interaction method provided in an embodiment of this disclosure;

[0053] Figure 6 This is a flowchart illustrating another cross-device user interaction method provided in an embodiment of this disclosure;

[0054] Figure 7 This is a schematic diagram of the structure of a first device provided in an embodiment of this disclosure;

[0055] Figure 8This is a schematic diagram of the structure of a second device provided in an embodiment of this disclosure. Detailed Implementation

[0056] To provide a more detailed understanding of the features and technical content of the embodiments of this disclosure, the implementation of the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for illustrative purposes only and are not intended to limit the embodiments of this disclosure. In the following technical description, for ease of explanation, several details are used to provide a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be simplified in their depiction to simplify the drawings.

[0057] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.

[0058] Unless otherwise stated, the term "multiple" means two or more.

[0059] In this embodiment of the disclosure, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.

[0060] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.

[0061] The term "correspondence" can refer to an association or interactive binding relationship. The correspondence between A and B means that there is an association or interactive binding relationship between A and B.

[0062] Currently, devices equipped with intelligent voice modules (such as smart speakers and smart refrigerators) are increasingly appearing in users' homes. These devices can respond to users' voice interaction information and output response content (such as music and news), providing users with a good service experience.

[0063] However, in related technologies, during the process of a device responding to a user's voice interaction, if the user leaves the vicinity of the device, although the device continues to output response content, the user cannot hear it. In these technologies, the device's output response content cannot follow the user's real-time location, thus failing to provide a better listening experience.

[0064] The cross-device user interaction method provided in this disclosure allows each device in the cross-device interaction system to detect in real time whether a user enters or exceeds its effective range. For a first user already in an interactive binding state with a first device, when the first device outputs response content in response to the first user's first voice interaction information, if the first user exceeds the effective range of the first device and enters the effective range of a second device, and if the first user meets the interactive binding conditions with the second device, then the second device will interactively bind with the first user. Simultaneously, the first device will stop outputting response content and transfer the first response content to the second device, thereby enabling the second device to output the remaining content in the first response content that was not output by the first device. In this way, the first user can continue to listen to the response content even after leaving the vicinity of the first device, providing a better content listening experience.

[0065] Combination Figure 1 As shown, this disclosure provides a cross-device interaction system, which includes multiple devices, such as device A, device B, device C, and device D. The devices in the cross-device interaction system can be smart terminals, smart home appliances, etc. For example, the devices in the cross-device interaction system can be televisions, smart screens, refrigerators, washing machines, air conditioners, and smart speakers.

[0066] In a cross-device interaction system, devices can interact with users (such as User 1 and User 2) via voice. Specifically, users can issue voice interaction messages, and devices can respond by outputting corresponding response content. For example, if a user issues the voice interaction message "Play song A," the device can respond by playing song A.

[0067] In this embodiment, the response content output by the device can be either a streaming resource or a non-streaming resource. Streaming resources typically refer to multimedia data resources, such as audio and video files, that are transmitted and played in a streaming form over a network. These resources are characterized by the ability to begin playback or use during data transmission, without waiting for the entire file to be downloaded. Non-streaming resources are a data acquisition method relative to streaming resources. In the internet or computer fields, non-streaming resources typically refer to digital resources that need to be completely downloaded to the local device before they can be used. For example, when downloading a document, music, movie, or application, the user must wait for the entire file to finish downloading before opening or using it.

[0068] In this embodiment, the device can acquire a user's facial image, collect the user's voice, and detect the distance between the user and the device. Using sensors, it can achieve distance detection, facial recognition, and voiceprint recognition functions. The sensor integrates at least a sound source localization sensor, a voiceprint recognition sensor, and a facial recognition sensor.

[0069] This disclosure defines various characteristic parameters of the device based on the device's reference points. The characteristic parameters involved in this disclosure are explained below.

[0070] The facial recognition range is the maximum area within which the device can acquire a user's facial information using sensors. The facial recognition range is defined with the device's reference point as the center and a distance r as the radius. f A spherical region with radius r, at a distance r f It is the maximum distance at which the sensor can acquire the user's facial information.

[0071] Voiceprint recognition range refers to the maximum range within which the device can acquire a user's voiceprint information using sensors. The voiceprint recognition range is defined by a reference point on the device, extending outwards at a distance r. v A spherical region with radius r, at a distance r v It is the maximum distance at which the sensor can acquire the user's voiceprint information.

[0072] The moment of user interaction binding is the moment when the user successfully completes the interaction binding with the device.

[0073] The moment a user interaction is created corresponds to the moment of the first character in the user's valid interaction corpus. The meaning of the user's valid interaction corpus is: voice content in the user's voice interaction information to the device that has a normal word order, clear intent, and practical meaning.

[0074] The moment when a user interaction ends, corresponding to the last word in the user's valid interaction corpus.

[0075] User interaction time difference, the time difference between the end of a user interaction and the creation of a user interaction.

[0076] The device response start time is the time corresponding to the first character of the device's response content.

[0077] The device response completion time, corresponding to the last character in the device's response content.

[0078] Device response time difference is the time difference between the completion time of the device response and the start time of the device response.

[0079] It should be noted that when the device outputs interactive content as streaming resources, the device's response completion time is the moment the termination command is triggered when the user actively triggers a termination command. When the device outputs interactive content as non-streaming resources, the device's response completion time is the moment the new user's interaction is bound to the device.

[0080] The user interaction triple consists of a unique interaction identifier, the user's facial information, and the user's voiceprint information. The unique interaction identifier is a string of numbers of a specified length generated based on preset rules; the facial information can be a facial feature vector; and the voiceprint information can be a voiceprint feature vector.

[0081] In this embodiment of the disclosure, the user interaction triplet can be stored in a database.

[0082] Combination Figure 2 As shown, this disclosure provides another cross-device interaction system, in Figure 1 Based on the cross-device interaction system shown, the cross-device interaction system also includes a cloud server, which can communicate with various devices, and the database is deployed on the cloud server.

[0083] Optionally, the unique interaction identifier consists of the time the user interaction was created and the sequential number of the voice interaction information in the cloud server logs. For example, the unique interaction identifier for a user bound to a device could be 20230926103735000120527.

[0084] In this embodiment of the disclosure, the voiceprint recognition range is greater than or equal to the face recognition range, that is, the maximum distance at which the sensor can acquire the user's voiceprint information is greater than or equal to the distance at which the sensor can (rv≥rf).

[0085] In this embodiment, the device is equipped with sensors that can acquire user facial images, collect user voice, and detect the distance between the user and the device. These sensors enable distance detection, facial recognition, and voiceprint recognition functions. This embodiment uses the sensor coordinates as the device's reference point.

[0086] In this embodiment of the disclosure, a user and a device can be interactively bound together. Under normal circumstances, the device can respond to the user's voice interaction information and output corresponding response content. The interaction binding process between the user and the device is described below.

[0087] (1) If the user is within the device's facial recognition range at the time the user interaction is created, the device will continuously acquire the user's facial information until it successfully acquires the facial information or the user leaves the device's facial recognition range (at which point acquiring facial information fails). In addition, the device will acquire the user's voiceprint recognition features at the end of the user interaction and form a user interaction triplet. The user and device can complete the interaction binding based on the user interaction triplet. Specifically, if the device fails to acquire facial information, the facial information in the user interaction triplet will be empty, meaning the database will not contain the user's facial information.

[0088] (2) If the user is not within the device's face recognition range at the time the user interaction is created, but is within the device's voiceprint recognition range, the device will directly wait until the user interaction ends to obtain the user's voiceprint recognition features and form a user interaction triplet. The user and the device can complete the interaction binding based on the user interaction triplet. In this case, the face information in the user interaction triplet is empty, meaning that the database will not contain the user's face information.

[0089] (3) If the user is not within the device's voiceprint recognition range at the time the interaction is created, the device will not acquire the user's voiceprint features and there is no need to form a user interaction triplet. In this case, the user and the device will not bind to each other.

[0090] In this embodiment, if multiple users (e.g., user 1 and user 2) are simultaneously within the face (voiceprint) recognition range of the same device (e.g., device A) within a unit of time, the priority for binding with the device is first determined based on the order in which each user's interaction creation time occurs. Generally, the user whose interaction creation time is earliest will have priority in binding with the device. If multiple users' interaction creation times are the same, the priority for binding with the device is determined based on the distance between each user and the device. Generally, the user closest to the device will have priority in binding with the device. If the distance between the user and the device is the same as the distance to the device, a user is randomly selected to have priority in binding with the device.

[0091] In this embodiment, if multiple users (e.g., user 1 and user 2) are simultaneously within the face (voiceprint) recognition range of multiple devices (e.g., device A and device B) within a unit of time, the distance between each user and each device is first calculated, and interaction binding is performed based on the principle of closest distance. Taking user 1, user 2, device A, and device B as an example, the distances between user 1 and device A, user 1 and device B, user 2 and device A, and user 2 and device B are calculated. The user and device corresponding to the smallest distance are bound (e.g., user 1 is bound to device A), and the remaining users and devices are bound (e.g., user 2 is bound to device B). If all calculated distances are equal, the user and device are randomly bound together.

[0092] In this embodiment of the disclosure, various interactive scenarios can be defined, such as music scenarios, recipe scenarios, refrigerator control scenarios, radio scenarios, and casual chat scenarios.

[0093] In conjunction with the cross-device interaction system provided in the embodiments of this disclosure, this disclosure provides a cross-device user interaction method. In implementing this user interaction method, a device in the cross-device interaction system that has completed interaction binding with a user can be defined as a first device, and a user in an interaction binding state with the first device can be defined as a first user. It is understood that any device in the cross-device interaction system that has completed interaction binding with a user can be considered as a first device.

[0094] In this embodiment of the disclosure, an effective range can be set for each device in the cross-device interaction system. When a user who is interacting with a device goes beyond the effective range of the device, it indicates that there is a need for the user to interact with other devices.

[0095] Optionally, the voiceprint recognition range of the device can be used as the effective range of the device, or a preset range larger than the voiceprint recognition range can be set as the effective range of the device.

[0096] Combination Figure 3 As shown, this disclosure provides a cross-device user interaction method, wherein the execution subject of the user interaction method is a first device, and includes the following steps:

[0097] S301, the first device responds to the first voice interaction information issued by the first user and outputs the first response content.

[0098] In this embodiment of the disclosure, the most recent voice interaction information issued by the first user is defined as the first voice interaction information, and the output response content of the first device in response to the first voice interaction information is defined as the first response content.

[0099] As mentioned above, the first device and the first user are in an interactive binding state. When the first user is within the voiceprint recognition range of the first device, if the user issues a first voice interaction message, the first device responds to the first voice interaction message issued by the first user and outputs a first response content. For example, if the user issues a voice interaction message of "play song A", the device can respond to the voice interaction message of "play song A" and play song A.

[0100] S302, when the first device determines during the response period that the first user is outside its own valid range, it determines whether it has received a content transfer notification.

[0101] Optionally, the voiceprint recognition range of the first device can be used as the effective range of the first device, or a preset range larger than the voiceprint recognition range can be set as the effective range of the first device.

[0102] In this embodiment of the disclosure, the first device can detect the distance between itself and the first user in real time, thereby determining whether the first user has exceeded its effective range. During the response period, if the first device determines that the first user has exceeded its effective range, it can detect in real time whether it has received a content streaming notification.

[0103] In this embodiment of the disclosure, the content transfer notification is sent by the second device. Wherein, if a user, after leaving the effective range of the first device, enters the effective range of another device (other than the first device), then that other device is defined as the second device.

[0104] In this embodiment, when the second device detects that the first user has entered its effective range, it acquires the first user's feature information and determines whether the first user meets the interaction binding conditions based on the feature information. If the first user meets the interaction binding conditions, the second device performs an interaction binding with the first user and sends a content flow notification.

[0105] Optionally, the second device can directly send the content streaming notification to the first device.

[0106] Optionally, the second device can send the content streaming notification to the cloud server, which then sends the content streaming notification to the first device.

[0107] S303, upon receiving a content transfer notification, the first device unbinds itself from the first user and stops outputting the first response content, and transfers the first response content to the second device.

[0108] In this embodiment of the disclosure, the content transfer notification is used to instruct the first device to unbind its interaction with the first user and to transfer the first response content to the second device. Therefore, upon receiving the content transfer notification, the first device will immediately unbind its interaction with the first user, stop outputting the first response content, and transfer the first response content to the second device.

[0109] After receiving the first response content, the second device will continue to output the remaining content in the first response content that was not output by the first device. In this way, the user can continue to listen to the first response content from the second device, so that the first user can listen to the first response content without interruption and provide the first user with a better content listening experience.

[0110] In this embodiment of the disclosure, after the first device transfers the first response content to the second device, if the first response content is a streaming resource, the first device enters an interaction pause state. If the first response content is a non-streaming resource, the first device enters an interaction termination state.

[0111] S304, the first device continues to output the first response content even if it does not receive a content transfer notification.

[0112] In this embodiment of the disclosure, if the first device determines that the first user has exceeded its valid range during the response duration, and if it has not received a content transfer notification, it will continue to output the first response content until a content transfer notification is received or the first response content output is completed.

[0113] In some embodiments, if no content transfer notification is received, the first response content is continued to be output, including: if no content transfer notification is received, detecting whether the first user has issued a second voice interaction message and detecting whether other users have entered its effective range; if it is determined that the first user has not issued a second voice interaction message and other users have not entered its effective range, the first response content is continued to be output.

[0114] Combination Figure 4 As shown, this disclosure provides another cross-device user interaction method, wherein the execution subject of this user interaction method is a first device, and includes the following steps:

[0115] S401, the first device responds to the first voice interaction information issued by the first user and outputs the first response content.

[0116] S402, when the first device determines during the response period that the first user is outside its own valid range, it determines whether to receive a content transfer notification.

[0117] S403, upon receiving a content transfer notification, the first device unbinds itself from the first user and stops outputting the first response content, and transfers the first response content to the second device.

[0118] S404, the first device detects whether the first user has issued a second voice interaction message and whether other users have entered its effective range when it has not received a content transfer notification.

[0119] In this embodiment of the disclosure, the voice interaction information issued by the first user during the response period of the first device to the first voice interaction information, when the first user is outside the effective range of the first device, is defined as the second voice interaction information.

[0120] Even without receiving a content transfer notification, the first device can continuously detect whether the first user sends out a second voice interaction message, and continuously detect whether other users enter its effective range.

[0121] S405, if the first device determines that the first user has not issued the second voice interaction information and other users have not entered its effective range, the first device continues to output the first response content.

[0122] In this embodiment of the disclosure, when the first device determines that the first user has exceeded its effective range during the response period, if it does not receive a content transfer notification, it continuously detects whether the first user has issued a second voice interaction message and whether other users have entered its effective range. If the first user has not issued a second voice interaction message and other users have not entered its effective range, it will continue to output the first response content until a content transfer notification is received or the first response content output is completed. If it detects that the first user has issued a second voice interaction message or that other users have not entered its effective range, it may continue to output the first response content until a content transfer notification is received or the first response content output is completed, or it may output new response content in response to the second voice interaction message.

[0123] In some embodiments, if the first device determines that the first user has not issued a second voice interaction message and detects that another user has entered its effective range, the first device continues to output the first response content. Upon ending the first response content, the first device begins to continue detecting whether the first user has issued a second voice interaction message. If the first device does not detect the first user issuing a second voice interaction message within a first preset time period, it establishes an interactive binding with that other user.

[0124] In this embodiment, the first device prioritizes maintaining its interaction binding with the first user. Even if the first user moves beyond its effective range and another user is detected entering its effective range, the device will not immediately unbind from the first user, even after the first response to the first voice interaction message has ended. Instead, it will continue to detect whether the first user sends a second voice interaction message within a first preset time period. If no second voice interaction message is detected from the first user within the first preset time period, then the device will re-bind with that other user. This ensures that when the first user sends a second voice interaction message within the first preset time period, the first device can respond promptly, prioritizing the content listening experience of the first user who is in an interaction binding state with the first device. This better aligns with user usage logic and priority rules.

[0125] In this embodiment of the disclosure, the first preset duration is the largest device response time difference among the first N device response time differences. Here, N is a positive integer, and the value of N can be determined according to actual needs; for example, N can be 3, 4, 5, etc.

[0126] Taking N as a value of 3 as an example, the time differences between the first three device responses are 20 seconds, 80 seconds, and 60 seconds, respectively. Therefore, the first preset duration should be 80 seconds. In this case, when the first device finishes its first response, it continues to detect whether the first user has issued a second voice interaction message. If the first device does not detect the first user issuing a second voice interaction message within 80 seconds, it establishes an interaction binding with that other user.

[0127] In some embodiments, when the first device does not receive a content streaming notification, it detects whether the first user has issued a second voice interaction message and whether other users have entered its effective range. If the first device determines that the first user has issued the second voice interaction message, it determines the type of resource requested by the second voice interaction message. If the resource type is a streaming resource, it does not respond to the second voice interaction message. If the resource type is a non-streaming resource, the first device determines whether it is the device with the smallest distance from the first user among all devices. If the first device determines that it is the device with the smallest distance from the first user among all devices, it responds to the second voice interaction message and outputs a second response content.

[0128] Furthermore, if the first device determines that the first user has not issued a second voice interaction message and detects that other users have entered its effective range, it continues to output the first response content. Upon ending the first response content, the first device continues to detect whether the first user has issued a second voice interaction message. When the first device detects that the first user has issued a second voice interaction message within a first preset time period, it determines the type of resource requested by the second voice interaction message. If the resource type is a streaming resource, it does not respond to the second voice interaction message. If the resource type is a non-streaming resource, the first device determines whether it is the device with the smallest distance to the first user among all devices. If the first device determines that it is the device with the smallest distance to the first user among all devices, it responds to the second voice interaction message and outputs the second response content.

[0129] Taking N as a value of 3 as an example, the time differences between the first three device responses are 20 seconds, 80 seconds, and 60 seconds, respectively. Therefore, the first preset duration should be 80 seconds. In this case, when the first response ends, the first device determines the resource type requested by the second voice interaction information. If the resource type is a streaming resource, it does not respond to the second voice interaction information. If the resource type is a non-streaming resource, the first device determines whether it is the device with the smallest distance from the first user among all devices. When the first device determines that it is the device with the smallest distance from the first user among all devices, it responds to the second voice interaction information and outputs the second response content.

[0130] In this embodiment of the disclosure, when the first user sends the second voice interaction information, if the resource type is a non-streaming resource and the first device determines that it is not the device with the smallest distance from the first user among all devices, the device with the smallest current distance from the first user (e.g., the second device) responds to the second voice interaction information and outputs the second response content.

[0131] The cross-device user interaction method provided in this disclosure can ensure that the unique and optimal device is selected to interact with the user by comparing the distance between the user and multiple devices. This not only improves the real-time performance of user-device interaction, but also avoids multiple devices being woken up by the user at the same time, reducing the probability of false device wake-up.

[0132] In this embodiment of the disclosure, after the first device has interacted and bound itself with the other user, it can output corresponding response content in response to the voice interaction information issued by the other user.

[0133] In this embodiment, for a new user who has recently interacted and bound to the first device, if that user has previously interacted and bound to other devices, then after the first device interacts and binds with that user, the first device can receive response content from the other devices that previously interacted and bound to that user, and output the remaining content in the response content that was not output by the other devices. Further, after the first device interacts and binds with that user, the first device can send a content transfer notification to the other devices that previously interacted and bound to that user. Upon receiving the content transfer notification, the other device unbinds from the other user, stops outputting response content, and transfers the response content to the first device, whereby the first device outputs the remaining content in the response content that was not output by the other devices.

[0134] In implementing the cross-device user interaction method provided in the embodiments of this disclosure, after defining a device in the cross-device interaction system that has completed interaction with the user as the first device, other devices in the cross-device interaction system can be regarded as the second devices.

[0135] Combination Figure 5 As shown, this disclosure provides another cross-device user interaction method, in which the execution subject is a second device, and includes the following steps:

[0136] S501, when the second device detects that the first user has entered its effective range, it acquires the first user's feature information and determines whether the first user meets the interaction binding conditions based on the first user's feature information.

[0137] Optionally, the voiceprint recognition range of the second device can be used as the effective range of the second device, or a preset range larger than the voiceprint recognition range can be set as the effective range of the second device.

[0138] In this embodiment, the second device can detect in real time whether a user enters its effective range. When a first user is detected entering its effective range, the device acquires the first user's feature information and determines whether the first user meets the interaction binding conditions based on the first user's feature information.

[0139] S502, when the first user meets the interaction binding conditions, the second device interacts and binds with the first user and sends a content flow notification.

[0140] Optionally, the second device can directly send the content streaming notification to the first device.

[0141] Optionally, the second device can send the content streaming notification to the cloud server, which then sends the content streaming notification to the first device.

[0142] In this embodiment of the disclosure, after receiving the content transfer notification, the first device unbinds itself from the first user and stops outputting the first response content, and then transfers the first response content to the second device.

[0143] S503, the second device receives the first response content transferred from the first device and outputs the remaining content in the first response content that was not output by the first device.

[0144] In some embodiments, obtaining feature information of a first user and determining whether the first user meets the interaction binding conditions based on the feature information of the first user includes: determining whether the first user has entered its own face recognition range, wherein the face recognition range is smaller than the effective range; if it is determined that the first user has entered its own face recognition range, obtaining the face information of the first user and comparing the obtained face information of the first user with the face information in the database; if it is determined that the obtained face information of the first user is consistent with the face information of the first user in the database, determining that the first user meets the interaction binding conditions.

[0145] Combination Figure 6 As shown, this disclosure provides another cross-device user interaction method, in which the execution subject is a second device, and includes the following steps:

[0146] S601, when the second device detects that the first user has entered its effective range, it determines whether the first user has entered its face recognition range.

[0147] In this embodiment, the face recognition range is smaller than the effective range. The second device can detect in real time whether a user has entered its effective range. When a first user is detected entering its effective range, it continues to detect whether the first user has entered its face recognition range.

[0148] S602, when the second device determines that the first user has entered its face recognition range, it acquires the face information of the first user and compares the acquired face information of the first user with the face information in the database.

[0149] As mentioned earlier, after a user interacts and binds with a device, the user is configured with a user interaction triplet, which can be stored in a database. The user interaction triplet consists of a unique interaction identifier, the user's facial information, and the user's voiceprint information.

[0150] In this embodiment, the second device acquires the first user's facial information when it determines that the first user has entered its facial recognition range. Then, it searches the database for the first user's user interaction triplet. If the first user's user interaction triplet is found, the acquired facial information of the first user is compared with the facial information in the first user's user interaction triplet.

[0151] S603, if the second device determines that the obtained facial information of the first user is consistent with the facial information of the first user in the database, it determines that the first user meets the interaction binding conditions.

[0152] S604, when the first user meets the interaction binding conditions, the second device interacts and binds with the first user and sends a content flow notification.

[0153] S605, the second device receives the first response content transferred from the first device and outputs the remaining content in the first response content that was not output by the first device.

[0154] In some embodiments, if the device fails to acquire facial information during the user's interaction and binding process with the device, the facial information in the user interaction triplet will be empty, meaning the database will not contain the user's facial information. Therefore, for the first user, their user interaction triplet may not contain their facial information, meaning the database may not contain the first user's facial information.

[0155] When the second device determines that the first user has entered its facial recognition range, it acquires the first user's facial information and compares it with facial information in the database. If the first user's facial information is not found in the database, or if the acquired facial information does not match the database, the second device acquires the first user's voiceprint information. If the first user's voiceprint information is acquired, it compares it with the database. If the acquired voiceprint information matches the database, the second device determines that the first user meets the interaction binding conditions.

[0156] In this embodiment of the disclosure, if the database does not contain the face information of the first user, or if it is determined that the face information of the first user obtained is inconsistent with the face information of the first user in the database, the second device will guide the first user to record a voice so as to obtain the voiceprint information of the first user obtained by the second device.

[0157] Specifically, the second device guides the first user to record their voice. If the first user responds within a specified time, the second device can record the first user's voice and obtain the first user's voiceprint information. Then, it searches for the first user's user interaction triplet in the database. If the first user's user interaction triplet is found, the obtained voiceprint information of the first user is compared with the voiceprint information in the first user's user interaction triplet.

[0158] If the obtained voiceprint information of the first user matches the voiceprint information of the first user in the user interaction triplet, then the first user is determined to meet the interaction binding condition. At this point, a new user interaction triplet for the first user can be formed and saved in the database.

[0159] If the obtained voiceprint information of the first user is consistent with the voiceprint information of the first user in the user interaction triplet, or if the voiceprint information of the first user does not exist in the database, it is determined that the first user does not meet the interaction binding conditions. At this time, the first device continues to output the first response content.

[0160] After the second device guides the first user to record sound, if the first user does not respond within a specified time, the second device determines that the first user does not meet the interaction binding conditions, and at this time the first device continues to output the first response content.

[0161] In some embodiments, when the second device is not in an interactive binding state with any user, if the first user remains within the effective range of the second device and sends voice interaction information after the first user does not respond within a specified time, the second device can obtain the voiceprint information of the first user. If the obtained voiceprint information of the first user matches the voiceprint information of the first user in the user interaction triplet, the second device determines that the first user meets the interactive binding condition. The second device can also output a response interaction core in response to the voice interaction information sent by the first user. In addition, a new user interaction triplet for the first user can be formed and stored in a database.

[0162] In some embodiments, when the second device detects that the first user has entered its effective range, it determines whether the first user has entered its face recognition range. If the second device determines that the first user has not entered its face recognition range, it acquires the first user's voiceprint information. The second device then compares the acquired voiceprint information with voiceprint information in a database. If the acquired voiceprint information matches the voiceprint information in the database, it determines that the first user meets the interaction binding conditions.

[0163] Similarly, the second device will guide the first user to record voice. If the first user responds within a specified time, the second device can record the first user's voice and obtain the first user's voiceprint information. Then, it will search for the first user's user interaction triplet in the database. If the first user's user interaction triplet is found, the obtained first user's voiceprint information will be compared with the voiceprint information in the first user's user interaction triplet.

[0164] If the obtained voiceprint information of the first user matches the voiceprint information of the first user in the user interaction triplet, then the first user is determined to meet the interaction binding condition. At this point, a new user interaction triplet for the first user can be formed and saved in the database.

[0165] If the obtained voiceprint information of the first user is consistent with the voiceprint information of the first user in the user interaction triplet, or if the voiceprint information of the first user does not exist in the database, it is determined that the first user does not meet the interaction binding conditions. At this time, the first device continues to output the first response content.

[0166] After the second device guides the first user to record sound, if the first user does not respond within a specified time, the second device determines that the first user does not meet the interaction binding conditions, and at this time the first device continues to output the first response content.

[0167] In some embodiments, when the second device is not in an interactive binding state with any user, if the first user remains within the effective range of the second device and sends voice interaction information after the first user does not respond within a specified time, the second device can obtain the voiceprint information of the first user. If the obtained voiceprint information of the first user matches the voiceprint information of the first user in the user interaction triplet, the second device determines that the first user meets the interactive binding condition. The second device can also output a response interaction core in response to the voice interaction information sent by the first user. In addition, a new user interaction triplet for the first user can be formed and stored in a database.

[0168] In some embodiments, when the second device detects that the first user has entered its effective range, the second device and the second user are in an interactive binding state. In this case, when the first user meets the interactive binding conditions, the second device needs to determine whether the second user meets the interactive binding release conditions. If it is determined that the second user meets the interactive binding release conditions, the second device releases the interactive binding with the second user, and then interactively binds with the first user and sends a content transfer notification. This allows the first device to receive the content transfer notification, release the interactive binding with the first user, stop outputting the first response content, and transfer the first response content to the second device. The second device then receives the first response content transferred from the first device and outputs the remaining content in the first response content that was not output by the first device.

[0169] The second device determines whether the second user meets the conditions for releasing the interaction binding, including the following:

[0170] If, during the response period for the second device to the second user, the second device forwards the interactive content output by the second device in response to the second user's voice interaction information to another device, then the second device will begin detecting whether the interactive content forwarded to the other device has been forwarded back to the second device within a second preset time period. If the second device does not detect whether the interactive content forwarded to the other device has been forwarded back to the second device within the second preset time period, it is determined that the second user meets the interaction binding release condition. If the second device detects whether the interactive content forwarded to the other device has been forwarded back to the second device within the second preset time period, it is determined that the second user does not meet the interaction binding release condition.

[0171] If, during the response period for the second device to the second user, the second device does not transfer the interactive content output in response to the second user's voice interaction information to other devices, then, at the moment the interactive content output is completed, the second device will detect whether the second user has issued new voice interaction information within a second preset time period. If the second device does not detect the second user issuing new voice interaction information within the second preset time period, it determines that the second user meets the interaction binding release condition. If the second device detects the second user issuing new voice interaction information within the second preset time period, it determines that the second user meets the interaction binding release condition.

[0172] If the second device is not currently responding to the second user, it will immediately begin detecting whether the second user has issued any new voice interaction information within a second preset time period. If the second device does not detect any new voice interaction information from the second user within the second preset time period, it will determine that the second user meets the conditions for releasing the interaction binding. If the second device detects that the second user has issued new voice interaction information within the second preset time period, it will determine that the second user meets the conditions for releasing the interaction binding.

[0173] In this embodiment of the disclosure, the second preset duration is the largest device response time difference among the first N device response time differences of the second device. Here, N is a positive integer, and the value of N can be determined according to actual needs; for example, N can be 3, 4, 5, etc.

[0174] The cross-device user interaction method provided in this disclosure requires obtaining the user's facial information and voiceprint information when binding with a device. This ensures the uniqueness of the user in a single interaction. Furthermore, by comparing the user's facial information and voiceprint information, the method achieves the purpose of transferring interactive content across devices. This not only improves the security of the interaction but also enhances its convenience.

[0175] Combination Figure 7 As shown, this disclosure provides a first device 700, which can be any device that can be bound to user interaction in cross-device user interaction. The first device 700 includes a processor 701 and a memory 702. Optionally, the first device 700 may further include a communication interface 703 and a bus 704. The processor 701, communication interface 703, and memory 702 can communicate with each other via the bus 704. The communication interface 703 can be used for information transmission. The processor 701 can call logical instructions in the memory 702 to execute the cross-device user interaction method of the above embodiments.

[0176] Furthermore, the logic instructions in the aforementioned memory 702 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.

[0177] The memory 702, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this disclosure. The processor 701 executes functional applications and data processing by running the program instructions / modules stored in the memory 702, thereby realizing the cross-device user interaction method in the above embodiments.

[0178] The memory 702 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 702 may include high-speed random access memory and may also include non-volatile memory.

[0179] Combination Figure 8As shown, this disclosure provides a second device 800, which can be any device that can be bound to user interaction in cross-device user interaction. The second device 800 includes a processor 801 and a memory 802. Optionally, the second device 800 may further include a communication interface 803 and a bus 804. The processor 801, communication interface 803, and memory 802 can communicate with each other via the bus 804. The communication interface 803 can be used for information transmission. The processor 801 can call logical instructions in the memory 802 to execute the cross-device user interaction method of the above embodiments.

[0180] Furthermore, the logic instructions in the aforementioned memory 802 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.

[0181] The memory 802, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this disclosure. The processor 801 executes functional applications and data processing by running the program instructions / modules stored in the memory 802, thereby realizing the cross-device user interaction method in the above embodiments.

[0182] The memory 802 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 802 may include high-speed random access memory and may also include non-volatile memory.

[0183] This disclosure provides a computer-readable storage medium storing computer-executable instructions configured to perform the aforementioned cross-device user interaction method.

[0184] The technical solutions of this disclosure can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in this disclosure. The aforementioned storage medium can be a non-transitory storage medium, such as a USB flash drive, external hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc., and other media capable of storing program code.

[0185] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operation may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the claims. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or” as used in this application means including one or more of the associated listed items and all possible combinations thereof. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.

[0186] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this disclosure. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0187] The methods and products (including but not limited to devices and equipment) disclosed in the embodiments herein can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units may be merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed units may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to implement this embodiment according to actual needs. Furthermore, the functional units in the embodiments of this disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0188] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

Claims

1. A cross-device user interaction method, applied to a first device, wherein the first device and a first user are in an interaction-bound state, characterized in that, include: In response to the first voice interaction message from the first user, output the first response content; If, during the duration of the response, it is determined that the first user has exceeded its own valid scope, it is determined whether a content transfer notification has been received; Upon receiving a content transfer notification, the system unbinds the interaction with the first user and stops outputting the first response content, and transfers the first response content to the second device so as to output the remaining content in the first response content that was not output by the first device. The first user is currently in an interactive binding state with the second device. If no content transfer notification is received, continue to output the first response content.

2. The user interaction method according to claim 1, characterized in that, If no content transfer notification is received, continue outputting the first response content, including: In the absence of a content transfer notification, the system detects whether the first user sends a second voice interaction message and whether other users have entered its effective range. If it is determined that the first user has not sent a second voice interaction message and other users have not entered its effective range, the first response content will continue to be output.

3. The user interaction method according to claim 2, characterized in that, Also includes: If it is determined that the first user has not issued a second voice interaction message and other users have been detected entering its effective range, the first response content will continue to be output. Upon completion of the first response, the system continues to detect whether the first user has issued a second voice interaction message. If no second voice interaction information is detected from the first user within the first preset time period, the system will interact and bind with the other user.

4. The user interaction method according to claim 3, characterized in that, After interacting and binding with the other user, the function also includes: receiving response content from other devices that have previously interacted and bound with the other user, and outputting the remaining content in the response content that was not output by the other device.

5. The user interaction method according to claim 2, characterized in that, Also includes: If it is determined that the first user has issued the second voice interaction information, the type of resource requested by the second voice interaction information shall be determined. When the resource type is a streaming resource, do not respond to the second voice interaction information; If the resource type is a non-streaming resource, determine whether it is the device with the smallest distance to the first user among all devices; When it determines that it is the device with the smallest distance from the first user among all devices, it responds to the second voice interaction information and outputs the second response content.

6. A cross-device user interaction method, applied to a second device, characterized in that, include: When the first user is detected to have entered its effective range, the feature information of the first user is obtained, and the first user is determined to meet the interaction binding conditions based on the feature information of the first user, wherein the first user is already in the interaction binding state with the first device. When the first user meets the interaction binding conditions, the device interacts with the first user and sends a content transfer notification so that after receiving the content transfer notification, the first device unbinds from the first user and stops outputting the first response content, and transfers the first response content to the second device. Receive the first response content from the first device and output the remaining content in the first response content that was not output by the first device.

7. The user interaction method according to claim 6, characterized in that, Obtain the characteristic information of the first user, and determine whether the first user meets the interaction binding conditions based on the characteristic information of the first user, including: Determine whether the first user has entered its own face recognition range, where the face recognition range is smaller than the effective range; Once it is determined that the first user has entered the facial recognition range, the facial information of the first user is obtained and compared with the facial information in the database. If the obtained facial information of the first user is consistent with the facial information of the first user in the database, the first user is determined to meet the interaction binding conditions.

8. The user interaction method according to claim 7, characterized in that, Also includes: If the first user's facial information does not exist in the database, or if it is determined that the obtained first user's facial information is inconsistent with the first user's facial information in the database, then obtain the first user's voiceprint information. Once the voiceprint information of the first user is obtained, it is compared with the voiceprint information in the database. If the obtained voiceprint information of the first user is consistent with the voiceprint information of the first user in the database, the first user is determined to meet the interaction binding conditions.

9. The user interaction method according to claim 7, characterized in that, Also includes: If it is determined that the first user has not entered the facial recognition range, obtain the voiceprint information of the first user; Once the voiceprint information of the first user is obtained, it is compared with the voiceprint information in the database. If the obtained voiceprint information of the first user is consistent with the voiceprint information of the first user in the database, the first user is determined to meet the interaction binding conditions.

10. A first device, comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to execute the cross-device user interaction method as described in any one of claims 1 to 5 when running the program instructions.

11. A second device comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to execute the cross-device user interaction method as described in any one of claims 6 to 9 when running the program instructions.

12. A cross-device interaction system, characterized in that, Includes the first device as described in claim 10 and the second device as described in claim 11.