Interaction device and control method thereof
By designing an interactive device including an image capture device, a communication interface and a processor, and using the identification model to identify scene objects and control it, the application problem of the lack of interaction with entity objects in the prior art is solved, and a wider interactive function is achieved.
Patent Information
- Application Number
- CN202311871580.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
Existing interactive devices of virtual reality, amplified reality or hybrid reality are mainly used for the interaction of virtual objects, and lack the application of using interactive devices to interact with physical objects.
An interactive device is designed, including a first image capture device, a second image capture device, a communication interface and a processor. The device uses the identification model to identify scene objects in the scene image and scene image to capture the user's face image and to transmit control requests to control these objects.
It realizes that the user interacts with the entity object through the interactive device, which enhances the practicality and application scope of the device.
Smart Images

Figure CN120233869A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an interaction device and method, and more particularly to an interaction device and method for remotely controlling other devices. Background Art
[0002] Most of the current interaction devices for virtual reality, augmented reality or mixed reality are used for allowing users to interact with virtual objects, and lack applications for interacting with objects in the real world using an interaction device (e.g., augmented reality glasses).
[0003] In view of this, how to provide a technology for users to interact with physical objects using an interaction device is an objective that the industry urgently needs to strive for. Summary of the Invention
[0004] To solve the above problems, the present disclosure provides an interaction device, including a first image capturing device, a second image capturing device, a communication interface, and a processor. The first image capturing device is used to capture a face image of a user. The second image capturing device is used to capture a scene image of a scene. The face in the present disclosure is not limited to the complete face of the user, and may also be a part of the face of the user. The processor is electrically connected to the first image capturing device, the second image capturing device, and the communication interface. The interaction device is used to perform the following operations: the processor determines whether the user is in a gazing state based on the face image; in response to the user being in the gazing state, the processor uses an identification model to identify at least one scene object in the scene image; and in response to the processor identifying the at least one scene object in the scene image, the communication interface transmits a control request to the at least one scene object to control the at least one scene object.
[0005] In an embodiment of the present invention, it further includes a memory, and the memory is electrically connected to the processor, wherein the interaction device is further used to perform the following operations: the processor obtains at least one identification information corresponding to the at least one scene object; the communication interface receives at least one credential corresponding to the at least one scene object from a server, where the at least one credential is generated by the server in response to the at least one identification information received from the interaction device; and the memory stores the at least one credential; wherein the control request transmitted by the communication interface further includes the at least one credential corresponding to the at least one scene object.
[0006] In an embodiment of the present invention, the operation of the processor determining whether the user is in the gazing state further includes: the processor calculates a plurality of pupil sizes and a plurality of line-of-sight angles based on a plurality of eye images in the face image; and the processor determines whether the user is in the gazing state based on the pupil sizes and the line-of-sight angles.
[0007] In an embodiment of the present invention, the recognition model further includes a feature extraction layer and a classification layer. The feature extraction layer is used to output corresponding multiple feature vectors based on the scene image, and the classification layer is used to recognize the at least one scene object in the scene image based on the feature vectors output by the feature extraction layer.
[0008] In an embodiment of the present invention, the recognition model further includes a feature extraction layer, and the interaction device is further configured to perform the following operations: The processor trains the feature extraction layer based on at least one object image corresponding to the at least one scene object in the scene image.
[0009] In an embodiment of the present invention, the recognition model further includes a feature extraction layer, and the interaction device is further configured to perform the following operations: The communication interface receives at least one preset image corresponding to the at least one scene object from the server; and the processor trains the feature extraction layer based on the at least one preset image.
[0010] In an embodiment of the present invention, the interaction device is further configured to perform the following operations: The communication interface receives update parameters from the server; and the processor updates the recognition model based on the update parameters.
[0011] In an embodiment of the present invention, the at least one scene object includes a first scene object and a second scene object, and the operation of the communication interface transmitting the control request to the at least one scene object further includes: The processor calculates the left eye line of sight and the right eye line of sight based on the face image; the processor calculates the projection point of the intersection of the left eye line of sight and the right eye line of sight on the plane based on the intersection; the processor selects one of the first scene object and the second scene object based on the projection point; and the communication interface transmits the control request to the selected one of the first scene object and the second scene object to control the selected one of the first scene object and the second scene object.
[0012] In an embodiment of the present invention, the communication interface further includes a first antenna and a second antenna, and the operation of the processor selecting one of the first scene object and the second scene object further includes: The first antenna and the second antenna respectively receive multiple positioning signals from the first scene object and the second scene object; the processor calculates the first projection point and the second projection point of the first scene object and the second scene object on the plane based on the positioning signals; and the processor selects one of the first scene object and the second scene object based on the projection point, the first projection point, and the second projection point.
[0013] The present disclosure also provides an interaction device control method applicable to an electronic device. The interaction device control method includes: the electronic device captures a face image of a user and a scene image in a scene; the electronic device determines whether the user is in a gazing state based on the face image; in response to the user being in the gazing state, the electronic device uses an identification model to identify at least one scene object in the scene image; and in response to identifying the at least one scene object in the scene image, the electronic device transmits a control request to the at least one scene object to control the at least one scene object.
[0014] It should be understood that the foregoing general description and the following detailed description are merely exemplary and explanatory and are intended to provide further explanation of the present disclosure as claimed. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To make the above and other objects, features, advantages and embodiments of the present disclosure more apparent and understandable, the descriptions of the accompanying drawings are as follows:
[0016] Figure 1 Schematic diagram of an interaction device in the first embodiment of the present disclosure;
[0017] Figure 2 Flowchart of an interaction device logging in to a scene object in some embodiments of the present disclosure;
[0018] Figure 3 Flowchart of an interaction device controlling a scene object in a scene in some embodiments of the present disclosure;
[0019] Figure 4 Schematic diagram of an identification model in some embodiments of the present disclosure;
[0020] Figure 5 Schematic diagram of an interaction device calculating a projection point corresponding to a user's line of sight in some embodiments of the present disclosure;
[0021] Figure 6 Schematic diagram of an interaction device calculating a projection point corresponding to a scene object in some embodiments of the present disclosure;
[0022] Figure 7 Flowchart of an interaction device controlling a scene object in a scene in another embodiment of the present disclosure;
[0023] Figure 8 Schematic diagram of a user interface and a menu in some embodiments of the present disclosure;
[0024] Figure 9 Flowchart of an interaction device control method in the second embodiment of the present disclosure;
[0025] Figure 10 Partial flowchart of an interaction device control method in some embodiments of the present disclosure;
[0026] Figure 11 is another part of the flowchart of the interactive device control method in some embodiments of the present disclosure; and
[0027] Figures 12 to 15 is the flowchart of the details of some steps of the interactive device control method in some embodiments of the present disclosure. Detailed implementation manners
[0028] To make the description of the present disclosure more detailed and complete, reference may be made to the accompanying drawings and the following various embodiments, in which the same numbers represent the same or similar elements.
[0029] Please refer to Figure 1 , which is a schematic diagram of the interactive device 1 in the first implementation manner of the present disclosure. The interactive device 1 includes a processor 11, a communication interface 12, a second image capturing device 13, and a first image capturing device 14, wherein the processor 11 is electrically connected to the communication interface 12, the second image capturing device 13, and the first image capturing device 14 respectively. The interactive device 1 is used to enable a user to interact with scene objects in the environment. For example, in the application field of smart home appliances, the user can operate the interactive device 1 to connect and control nearby home appliances (such as speakers, table lamps, air conditioners); in the application field of intelligent manufacturing, the user can operate the interactive device 1 to connect and control instruments in the factory (such as exhaust fans, machine tools). In some embodiments, the interactive device 1 can be a virtual reality glasses, an augmented reality glasses, or a mixed reality glasses.
[0030] The processor 11 is used to perform arithmetic functions. In some embodiments, the processor 11 may include a central processing unit (CPU), a graphics processing unit (GPU), a multi-processor, a distributed processing system, an application specific integrated circuit (ASIC), and / or a suitable arithmetic unit.
[0031] The communication interface 12 is used to transmit and / or receive information to / from other devices. In some embodiments, the communication interface 12 may include Bluetooth, Wi-Fi, and / or other data transceiver interfaces.
[0032] The second image capturing device 13 is used to capture scene images in the scene where the interactive device 1 is located. In some embodiments, the second image capturing device 13 may include one or more cameras disposed on the glasses and facing the outside of the glasses (i.e., the side not facing the user) for shooting.
[0033] The first image capturing device 14 is used to capture the facial image of the user. In some embodiments, the second image capturing device 13 may include one or more cameras disposed on the glasses and facing the user for shooting.
[0034] In some embodiments, the interaction device 1 further includes a memory, and the memory is electrically connected to the processor 11. Further, before the interaction device 1 connects to and controls the scene object, the interaction device 1 and the scene object may exchange credentials to log in to each other's information first, and then confirm the permissions of the connection objects with each other when connecting in the future.
[0035] Specifically, the interaction device 1 may log in to the scene object through the following operations: the processor 11 obtains at least one identification information corresponding to the at least one scene object; and the communication interface 12 receives at least one credential corresponding to the at least one scene object from the server, where the at least one credential is generated by the server in response to the at least one identification information received from the interaction device; and the control request transmitted by the communication interface further includes the at least one credential corresponding to the at least one scene object.
[0036] Regarding the operation of the interaction device 1 logging in to the scene object, please refer to Figure 2 , which is the flowchart of the interaction device 1 logging in to the scene object in some embodiments of the present disclosure.
[0037] First, the interaction device 1 executes operation OP11 to read the identification information from the scene object SD. Specifically, the interaction device 1 may obtain the identification information by means such as scanning a two-dimensional barcode on the scene object SD, downloading from a gateway, etc., where the identification information may be a unique string or other data formats for identification.
[0038] Next, the interaction device 1 executes operation OP12 to transmit the identification information to the server SV. Correspondingly, after receiving the identification information, the server SV may execute operation OP13 to confirm device authorization based on the identification information.
[0039] Specifically, the server SV may identify the scene object SD based on the identification information and confirm the status of the scene object SD. For example: if the scene object SD has not been logged in by other interaction devices, the server SV may determine that the interaction device 1 is the owner of the scene object SD, and the interaction device 1 has the control permission of the scene object SD; on the other hand, if the scene object SD has been logged in by other interaction devices in the past, the server SV may notify the owner of the scene object SD (i.e., the first interaction device that logged in to the scene object SD), and confirm with the owner whether to agree to authorize the interaction device 1 the control permission of the scene object SD.
[0040] After confirming that the interaction device 1 has the control authority over the scene object SD, the server SV can then execute operations OP14 and OP15, transmitting the device credential to the interaction device 1 and transmitting the user credential to the scene object SD, where the device credential is the authority credential of the scene object SD and the user credential is the authority credential of the interaction device 1. Correspondingly, after receiving the device credential, the interaction device 1 can execute operation OP16, storing the device credential in the memory to log in to the scene object SD; after receiving the user credential, the scene object SD can then execute operation OP17, storing the user credential to log in to the interaction device 1. It should be noted that the present disclosure does not limit the execution order of operations OP14 and OP15.
[0041] In this way, when the interaction device 1 and the scene object SD are connected in the future, they can exchange the device credential and the user credential, and each verify the received credential to confirm the identity and control authority of the connection object.
[0042] In some embodiments, the memory of the interaction device 1 may include semiconductor or solid-state memory, magnetic tape, removable computer disks, random access memory (RAM), read-only memory (ROM), hard disks, and / or optical disks.
[0043] Regarding the operations of the interaction device 1 interacting with the scene object in the scene, please refer to Figure 3 , the interaction device 1 is used to execute operations OP21 to OP23 to control the scene object in the scene.
[0044] First, the interaction device 1 executes operation OP21, and the processor 11 determines whether the user is gazing at an object. Specifically, the processor 11 determines whether the user is in a gazing state based on the facial image captured by the first image capturing device 14.
[0045] In some embodiments, the processor 11 calculates a plurality of pupil sizes and a plurality of line-of-sight angles based on a plurality of eye images in the facial image; and the processor 11 determines whether the user is in the gazing state based on the pupil sizes and the line-of-sight angles.
[0046] For example, when the user desires to control a certain scene object, their expression will have specific characteristics due to gazing at the scene object, such as the line of sight converging, the pupils shrinking, frowning, etc. due to the line of sight focusing on a specific object. Correspondingly, the processor 11 can then determine whether the user is gazing at a specific object based on the size change of the pupils presented in the facial image, the line-of-sight angle, and / or the expression characteristics.
[0047] Further, if the processor 11 determines that the user does not belong to the gazing state, the interaction device 1 can continuously execute operation OP21 without executing the next operation. In other words, the interaction device 1 will continuously determine whether the user belongs to the gazing state until the determination is successful and then execute operation OP22.
[0048] On the other hand, if the processor 11 determines that the user belongs to the gazing state, the interaction device 1 executes operation OP22, and the processor 11 determines whether there is a scene object in the scene image. Specifically, the processor 11 uses the recognition model to recognize at least one scene object in the scene image captured by the second image capturing device 13.
[0049] For example, the processor 11 can input the scene image into the recognition model to determine whether the scene image contains a controllable device (e.g., home appliance, Internet of Things device).
[0050] It should be noted that the recognition model can be a trained image recognition machine learning model. In some embodiments, the appearance images of the devices with which the interaction device 1 can interact can be used as training data, and an image recognition machine learning model is trained based on the training data, and the trained machine learning model is used as the recognition model.
[0051] In some embodiments, please refer to Figure 4 , the recognition model RM further includes a feature extraction layer FEL and a classification layer CL. The feature extraction layer FEL is used to output corresponding multiple feature vectors FV based on the scene image SI, and the classification layer CL is used to recognize at least one scene object SO in the scene image SI based on the feature vectors FV output by the feature extraction layer FEL.
[0052] In this embodiment, the recognition model RM includes a feature extraction layer FEL with relatively fewer parameters and a classification layer CL with relatively more parameters. In this way, when the recognition model RM needs to be updated, different update methods can be adopted for the different functions of the feature extraction layer FEL and the classification layer CL respectively.
[0053] In some embodiments, since the feature extraction layer FEL has relatively few parameters and low training computational complexity, the processor 11 can further be used to train the feature extraction layer FEL based on at least one object image corresponding to at least one scene object SO in the scene image SI. Since even for devices of the same model, users may customize and decorate the scene objects according to their own preferences (e.g., change the color, add patterns on the surface), the interaction device 1 can train the feature extraction layer FEL with the images of the scene object SO it captures to improve the accuracy of future image recognition.
[0054] In some embodiments, the communication interface 12 of the interaction device 1 may also receive at least one preset image of the corresponding scene object SO from the server; and the processor 11 trains the feature extraction layer FEL based on the at least one preset image. In this way, the interaction device 1 can use the preset images downloaded from the server to train the feature extraction layer FEL to improve the recognition accuracy of the scene object, where the preset images may include the appearance images of existing scene objects and may also include the appearance images of newly added scene objects.
[0055] In some embodiments, the communication interface 12 of the interaction device 1 may also receive update parameters from the server; and the processor 11 updates the recognition model RM based on the update parameters. As described above, the recognition model RM may include a classification layer CL with more parameters, and if the classification layer CL needs to be trained by the interaction device 1, it may consume a long time and energy. Therefore, the server can train the recognition model RM and transmit the trained update parameters to the interaction device 1 to update the recognition model RM in the interaction device 1, where the update parameters can be used to update the feature extraction layer FEL and / or the classification layer CL in the recognition model RM.
[0056] It should be noted that the aforementioned server may be a cloud server, and the interaction device 1 can connect to the server through the communication interface 12 (such as a data transceiver interface or other data transmission interfaces) to transmit data.
[0057] In some embodiments, when the interaction device 1 identifies multiple scene objects SO in the scene image SI, the interaction device 1 may also determine the scene object that the user desires to control based on the direction of the user's eye gaze.
[0058] Specifically, the processor 11 calculates the left-eye line of sight and the right-eye line of sight based on the face image; the processor 11 calculates the projection point of the intersection of the left-eye line of sight and the right-eye line of sight on the plane based on the intersection; the processor 11 selects one of the first scene object and the second scene object based on the projection point; and the communication interface 12 transmits the control request to the one of the first scene object and the second scene object to control the one of the first scene object and the second scene object.
[0059] Please refer to Figure 5 , which is a schematic diagram of the interaction device 1 calculating the projection point α corresponding to the user's line of sight. As Figure 5As shown, the processor 11 can calculate the line of sight LS of the user's left eye LE and the line of sight RS of the right eye RE based on the pupil images in the face image. Further, the processor 11 calculates the intersection point F of the line of sight LS and the line of sight RS, and then projects the intersection point F onto the projection point α on the virtual plane VP established by the processor 11, where the virtual plane VP can be a plane perpendicular to the direction of the user's face orientation and / or a plane parallel to the scene image.
[0060] Accordingly, after the interaction device 1 calculates the projection point, it can calculate the distances between the projection point and multiple scene objects in the scene image, and further transmit a control request to the scene object closest in distance.
[0061] Further, in some embodiments, the interaction device 1 can also perform multi-antenna positioning on multiple scene objects by multiple antennas respectively, and determine the object that the user desires to control based on the positioning results.
[0062] Specifically, the communication interface 11 further includes a first antenna and a second antenna, and the operation of the processor to select one of the first scene object and the second scene object further includes: the first antenna and the second antenna respectively receive multiple positioning signals from the first scene object and the second scene object; the processor calculates the first projection point and the second projection point of the first scene object and the second scene object respectively on the plane based on the positioning signals; and selects one of the first scene object and the second scene object based on the projection point, the first projection point, and the second projection point.
[0063] Please refer to Figure 6 , which is a schematic diagram of the interaction device 1 calculating the projection points β1 and β2 corresponding to the scene objects. As Figure 6 shown, the communication interface 11 of the interaction device 1 includes antennas LA and RA. The interaction device 1 can respectively locate the positions of the scene objects SD1 and SD2 in the three-dimensional space by the antennas LA and RA through means such as the angle of departure (AoD), and the processor 11 calculates the projection point β1 corresponding to the scene object SD1 and the projection point β2 corresponding to the scene object SD2 on the virtual plane VP based on the positions of the scene objects SD1 and SD2. In some embodiments, the antennas LA and RA can perform multi-antenna positioning through Bluetooth, Wi-Fi, or other wireless communication protocols.
[0064] Further, after the interaction device 1 obtains the projection points β1 and β2, it can select the projection point closer to the projection point α based on the projection points β1 and β2 and control the corresponding scene object.
[0065] In some embodiments, the interaction device 1 may further include an input interface. If the interaction device 1 is connected to an incorrect scene object, the user can input commands to the input interface in ways such as eye movements, gestures, voice, etc. to cancel the connection. The input interface may include a camera, a microphone, buttons, a remote control device, a joystick, and / or other input interfaces.
[0066] In some embodiments, when the interaction device 1 identifies multiple scene objects SO in the scene image SI and does not select one of the scene objects due to a judgment failure, user cancellation of the connection, or other reasons, the interaction device 1 may also provide a user menu to select the scene object to be controlled.
[0067] Please refer to Figure 7 , the interaction device 1 can generate a menu by operating OP31 to OP38 to provide the user with the option to select the correct scene object, where the operation of OP31 is the same as Figure 3 the illustrated operation of OP21, the operation of OP32 is the same as Figure 3 the illustrated operation of OP22, and the operation of OP33 is the same as Figure 3 the illustrated operation of OP23, so it will not be elaborated here.
[0068] And Figure 3 different from the illustrated operation process, in the operation of OP32, if the processor 11 does not identify a scene object in the scene image, the interaction device 1 can execute the operation of OP36, and the processor 11 generates a menu based on the credentials stored in the memory, and the menu lists the previously logged-in scene objects for the user to select.
[0069] On the other hand, after the interaction device 1 executes the operation of OP33, if the interaction device 1 receives a cancellation command input by the user from the input interface, the interaction device 1 can execute the operation of OP37, and the processor 11 generates a menu based on the scene object identified in the operation of OP32, and the menu lists the scene objects in the scene image for the user to select. Conversely, if the input interface does not receive a cancellation command, the interaction device 1 can execute the operation of OP35 and continue to interact with the scene object.
[0070] After the interaction device 1 executes the operation of OP36 and / or OP37, the interaction device 1 can execute the operation of OP38, and the processor 11 selects the scene object based on the input command received from the input interface, that is, selects the interaction object according to the scene object selected by the user in the menu.
[0071] In some embodiments, if the menu provided when the interaction device 1 executes the operation of OP37 does not meet the user's needs, the interaction device 1 can also return to the operation of OP36 to generate a menu based on the credentials for the user to select.
[0072] Please refer to the menu generated by the interaction device 1 when performing operation OP37 Figure 8 , which is a schematic diagram of the user interface UI in some embodiments of the present disclosure. As Figure 8 shown, there are three scene objects in the scene presented in the user interface UI: air conditioner A, table lamp B, and switch C. After the interaction device 1 identifies the three scene objects, the interaction device 1 can provide a menu MU in the user interface UI to list the three scene objects for the user to select the scene object to be controlled. In this way, the user can select the scene object to be controlled through the input interface from the menu MU.
[0073] Similarly, the menu generated by the interaction device 1 when performing operation OP36 can also be presented in a manner similar to the menu MU, with the only difference being the generation method of the options in the menu, so it will not be elaborated here.
[0074] Finally, please return to Figure 3 , after confirming the scene object (for example: selecting the scene object SD1), the interaction device 1 can perform operation OP23, and the communication interface 12 transmits a control request to the scene object SD1 to control the scene object SD1.
[0075] For example, the communication interface 12 can connect to the scene object SD1 through an exchange certificate or a secure communication protocol (Secure Sockets Layer; SSL) between Internet of Things (IoT) devices.
[0076] After establishing the connection, the user can control the scene object SD1 through the interaction device 1 to perform corresponding functions, such as: adjusting the brightness of the table lamp, turning the table lamp on / off, etc. It should be noted that the functions that the scene object can perform depend on its device type. For example, if the scene object is an exhaust fan, the interaction device 1 can control its operating power; if the scene object is an air conditioner, the interaction device 1 can control its air outlet temperature, and the technology of the present disclosure is not limited to the foregoing examples.
[0077] In some embodiments, the scene object SD1 can confirm the control authority of the interaction device 1 through the device certificate transmitted by the communication interface 12, for example: having full or only partial control of the functions. Conversely, the interaction device 1 can also identify the scene object SD1 through the user certificate transmitted by the scene object SD1 to confirm whether it is connected to the correct scene object.
[0078] In summary, the interactive device 1 proposed in this disclosure can determine whether a user desires to control other scene objects based on the user's face image, then use an identification model to identify the scene objects in the environment, and further connect and control the scene objects. In addition, the interactive device 1 can pre-register the credentials corresponding to the scene objects, and confirm whether it is connected to the correct scene object during connection. Further, the interactive device 1 can also determine the line of sight direction based on the user's eye image, and combine multi-antenna positioning to locate the position of the scene object, and then connect to the scene object stared at by the user. When the identification model needs to be updated, the interactive device 1 can use the captured object image or the downloaded preset image to locally update the feature extraction layer in the identification model, and directly download the updated parameters to update the classification layer in the identification model to increase the efficiency of the interactive device 1 in updating the identification model.
[0079] Please refer to Figure 9 , which is a flowchart of the interactive device control method 20 in the second embodiment of this disclosure. The interactive device control method 20 includes steps S21 to S23. The interactive device control method 20 is used to enable a user to interact with scene objects in the environment. The interactive device control method 20 can be executed by an electronic device (for example: Figure 1 the interactive device 1 shown).
[0080] In some embodiments, the electronic device includes a processor (for example: Figure 1 the processor 11 shown), a communication interface (for example: Figure 1 the communication interface 12 shown), a second image capture device (for example: Figure 1 the second image capture device 13 shown), and a first image capture device (for example: Figure 1 the first image capture device 14 shown), wherein the first image capture device is used to capture the user's face image, the second image capture device is used to capture the scene image in the scene, and the processor is electrically connected to the first image capture device, the second image capture device, and the communication interface.
[0081] First, in step S21, the electronic device captures the user's face image and the scene image in the scene.
[0082] Next, in step S22, the electronic device determines whether the user is in a staring state based on the face image.
[0083] Then, in step S23, in response to the user being in the staring state, the electronic device uses an identification model to identify at least one scene object in the scene image.
[0084] Finally, in step S24, in response to identifying the at least one scene object in the scene image, the electronic device transmits a control request to the at least one scene object to control the at least one scene object.
[0085] In some embodiments, the interactive device control method 20 further includes steps S21A and S23A as Figure 10 illustrated.
[0086] In step S21A, the electronic device obtains at least one identification information corresponding to the at least one scene object.
[0087] In step S22A, the electronic device receives at least one credential corresponding to the at least one scene object from the server, where the at least one credential is generated by the server in response to the at least one identification information received from the electronic device; wherein the control request transmitted by the communication interface further includes the at least one credential corresponding to the at least one scene object.
[0088] In step S23A, the electronic device stores the at least one credential.
[0089] Wherein the control request transmitted by the electronic device further includes the at least one credential corresponding to the at least one scene object.
[0090] In some embodiments, the electronic device further includes a memory (e.g., the memory of the interactive device 1 in the first embodiment), the memory is electrically connected to the processor, and is used to store the at least one credential.
[0091] In some embodiments, the interactive device control method 20 further includes the electronic device receiving an acknowledgment response from the at least one scene object to establish a connection, where the acknowledgment response is generated after the at least one scene object verifies the at least one credential of the control request.
[0092] In some embodiments, the interactive device control method 20 further includes steps S21B to S23B as Figure 11 illustrated.
[0093] In step S21B, in response to not identifying the at least one scene object in the scene image, the electronic device generates a menu based on the at least one credential, where the menu includes the at least one scene object corresponding to the at least one credential.
[0094] In step S22B, the electronic device selects one of the at least one scene object based on the received input instruction.
[0095] In step S23B, the electronic device transmits the control request to the one of the at least one scene object.
[0096] In some embodiments, step S22 further includes asFigure 12 The illustrated steps S221 and S222.
[0097] In step S221, the electronic device calculates a plurality of pupil sizes and a plurality of line-of-sight angles based on a plurality of eye images in the face image.
[0098] In step S222, the electronic device determines whether the user is in the gazing state based on the pupil sizes and the line-of-sight angles.
[0099] In some embodiments, the recognition model further includes a feature extraction layer and a classification layer. The feature extraction layer is configured to output corresponding feature vectors based on the scene image, and the classification layer is configured to identify the at least one scene object in the scene image based on the feature vectors output by the feature extraction layer.
[0100] In some embodiments, the recognition model further includes a feature extraction layer, and the interaction device control method 20 further includes the electronic device training the feature extraction layer based on at least one object image corresponding to the at least one scene object in the scene image.
[0101] In some embodiments, the recognition model further includes a feature extraction layer, and the interaction device control method 20 further includes the electronic device receiving at least one preset image corresponding to the at least one scene object from a server; and the electronic device training the feature extraction layer based on the at least one preset image.
[0102] In some embodiments, the interaction device control method 20 further includes the electronic device receiving updated parameters from a server; and the electronic device updating the recognition model based on the updated parameters.
[0103] In some embodiments, the at least one scene object includes a first scene object and a second scene object, and step S24 further includes steps S241 to S244 as Figure 13 illustrated.
[0104] In step S241, the electronic device calculates a left-eye line of sight and a right-eye line of sight based on the face image.
[0105] In step S242, the electronic device calculates a projection point of the intersection on a plane based on the intersection of the left-eye line of sight and the right-eye line of sight.
[0106] In step S243, the electronic device selects one of the first scene object and the second scene object based on the projection point.
[0107] In step S244, the electronic device transmits the control request to the selected one of the first scene object and the second scene object to control the selected one of the first scene object and the second scene object.
[0108] In some embodiments, the electronic device further includes a first antenna and a second antenna, and step S243 further includes steps S2431 to S2433 as Figure 14 illustrated.
[0109] In step S2431, the first antenna and the second antenna respectively receive a plurality of positioning signals from the first scene object and the second scene object.
[0110] In step S2432, the electronic device calculates a first projection point and a second projection point of the first scene object and the second scene object on the plane respectively based on the positioning signals.
[0111] In step S2433, the electronic device selects one of the first scene object and the second scene object based on the projection point, the first projection point, and the second projection point.
[0112] In some embodiments, the interactive device control method 20 further includes that in response to identifying the at least one scene object in the scene image, the electronic device generates a menu based on the at least one scene object; and the electronic device transmits the control request to one of the at least one scene object based on the option selected by the user in the menu.
[0113] In some embodiments, step S24 further includes steps S245 to S247 as Figure 15 illustrated.
[0114] In step S245, in response to receiving a cancellation instruction, the electronic device generates a menu, where the menu includes the at least one scene object.
[0115] In step S246, the electronic device selects one of the at least one scene object based on the received input instruction.
[0116] In step S247, the electronic device transmits the control request to the one of the at least one scene object.
[0117] In some embodiments, the electronic device further includes an input interface for receiving instructions from the user, such as the cancellation instruction and / or the input instruction.
[0118] In summary, the interactive device control method 20 proposed in this disclosure can determine whether a user desires to control other scene objects based on the user's facial image, then use an identification model to identify the scene objects in the environment, and further connect to and control the scene objects. In addition, the interactive device control method 20 can pre-register the credentials corresponding to the scene objects, and confirm whether it is connected to the correct scene object during connection. Further, the interactive device control method 20 can also determine the line of sight direction based on the user's eye image, and combine multi-antenna to locate the position of the scene object, and then connect to the scene object stared at by the user. When the identification model needs to be updated, the interactive device control method 20 can locally update the feature extraction layer in the identification model using the captured object image or the downloaded preset image, and directly download the updated parameters to update the classification layer in the identification model to increase the efficiency of updating the identification model by the interactive device control method 20.
[0119] Although several embodiments are described in detail above as examples, the interactive device and method proposed in this disclosure can also be implemented by other systems, hardware, software, storage media, or combinations thereof. Therefore, the protection scope of this disclosure should not be limited to the specific implementation manners described in the embodiments of this disclosure, and should be subject to what is defined by the appended claims.
[0120] It is obvious to those skilled in the art to which this disclosure pertains that various modifications and changes can be made to the structure of this disclosure without departing from the scope or spirit of this disclosure. In view of the foregoing, the protection scope of this disclosure also covers the modifications and changes made within the appended claims.
[0121]
Symbol Description
[0122] 1: Interactive device
[0123] 11: Processor
[0124] 12: Communication interface
[0125] 13: Second image capture device
[0126] 14: First image capture device
[0127] SD, SD1, SD2: Scene object
[0128] SV: Server
[0129] OP11~OP17, OP21~OP23, OP31~OP38: Operation
[0130] RM: Identification model
[0131] SI: Scene image
[0132] FEL: Feature extraction layer
[0133] FV: Feature Vector
[0134] CL: Classification Layer
[0135] SO: Scene Object
[0136] LE: Left Eye
[0137] RE: Right Eye
[0138] LS, RS: Line of Sight
[0139] F: Intersection Point
[0140] VP: Virtual Plane
[0141] α, β1, β2: Projection Point
[0142] LA, RA: Antenna
[0143] UI: User Interface
[0144] MU: Menu
[0145] 20: Interactive Device Control Method
[0146] S21~S24, S21A~S23A, S21B~S23B, S221, S222, S241~S247, S2431~S2433: Steps
Claims
1. An interaction device, characterized in that, Comprising: A first image capturing device for capturing a facial image of a user; A second image capturing device for capturing a scene image in a scene; A communication interface; and A processor electrically connected to the first image capturing device, the second image capturing device, and the communication interface; Wherein the interaction device is used to perform the following operations: The processor determines whether the user is in a gazing state based on the facial image; In response to the user being in the gazing state, the processor uses an identification model to identify at least one scene object in the scene image; And In response to the processor identifying the at least one scene object in the scene image, the communication interface transmits a control request to the at least one scene object to control the at least one scene object.
2. The interactive device according to claim 1, wherein Further comprising a memory electrically connected to the processor, wherein the interaction device is further used to perform the following operations: The processor obtains at least one identification information corresponding to the at least one scene object; The communication interface receives at least one credential corresponding to the at least one scene object from a server, wherein the at least one credential is generated by the server in response to the at least one identification information received from the interaction device; and The memory stores the at least one credential; Wherein the control request transmitted by the communication interface further includes the at least one credential corresponding to the at least one scene object.
3. The interactive device according to claim 1, characterized in that Wherein the operation of the processor to determine whether the user is in the gazing state further includes: The processor calculates a plurality of pupil sizes and a plurality of line-of-sight angles based on a plurality of eye images in the facial image; and The processor determines whether the user is in the gazing state based on the pupil sizes and the line-of-sight angles.
4. The interactive device according to claim 1, wherein Wherein the identification model further includes a feature extraction layer and a classification layer, the feature extraction layer is used to output corresponding feature vectors based on the scene image, and the classification layer is used to identify the at least one scene object in the scene image based on the feature vectors output by the feature extraction layer.
5. The interactive device according to claim 1, wherein Wherein the identification model further includes a feature extraction layer, and the interaction device is further used to perform the following operations: The processor trains the feature extraction layer based on at least one object image corresponding to the at least one scene object in the scene image.
6. The interactive device according to claim 1, wherein Wherein the identification model further includes a feature extraction layer, and the interaction device is further used to perform the following operations: The communication interface receives at least one preset image corresponding to the at least one scene object from a server; and The processor trains the feature extraction layer based on the at least one preset image.
7. The interactive device according to claim 1, wherein Wherein the interaction device is further used to perform the following operations: The communication interface receives update parameters from a server; and The processor updates the identification model based on the update parameters.
8. The interactive device according to claim 1, characterized in that Wherein the at least one scene object includes a first scene object and a second scene object, and the operation of the communication interface to transmit the control request to the at least one scene object further includes: The processor calculates a left-eye line of sight and a right-eye line of sight based on the facial image; The processor calculates a projection point of the intersection of the left-eye line of sight and the right-eye line of sight on a plane; The processor selects one of the first scene object and the second scene object based on the projection point; and The communication interface transmits the control request to one of the first scene object and the second scene object to control one of the first scene object and the second scene object.
9. The interactive device according to claim 8, wherein Wherein the communication interface further includes a first antenna and a second antenna, and the operation of the processor to select one of the first scene object and the second scene object further includes: The first antenna and the second antenna respectively receive a plurality of positioning signals from the first scene object and the second scene object; The processor calculates a first projection point and a second projection point of the first scene object and the second scene object on the plane respectively based on the positioning signals; and The processor selects one of the first scene object and the second scene object based on the projection point, the first projection point and the second projection point.
10. A method for controlling an interaction device, characterized in that, Applicable to an electronic device, wherein the interactive device control method includes: The electronic device captures a face image of the user and a scene image in the scene; The electronic device determines whether the user is in a gazing state based on the face image; In response to the user being in the gazing state, the electronic device uses an identification model to identify at least one scene object in the scene image; and In response to identifying the at least one scene object in the scene image, the electronic device transmits a control request to the at least one scene object to control the at least one scene object.