Intelligent device control method, apparatus, device, and medium

By identifying the first and second parts of the target object, performing identity authentication and granting permissions, the problem of intelligent devices being unable to be accurately controlled in multi-person interaction scenarios is solved, ensuring that only authorized objects can trigger device operations.

CN114581974BActive Publication Date: 2025-10-10SHENZHEN LUMIUNITED TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210134001.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-14
Publication Date
2025-10-10
Estimated Expiration
2042-02-14

AI Technical Summary

Technical Problem

In multi-person interaction scenarios, smart devices cannot implement precise control, resulting in gestures from unauthorized personnel triggering device operations.

Method used

By acquiring the image to be identified, identifying the first part and the second part of the target object, performing identity authentication, determining whether it has control authority, and triggering the target device to perform operations based on the association relationship.

Benefits of technology

It achieves precise control of smart devices, ensuring that only authorized objects can trigger corresponding operations and preventing unauthorized objects from performing incorrect operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114581974B_ABST
    Figure CN114581974B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a kind of intelligent device control method, device, electronic equipment and storage medium, it is related to Internet of Things technical field.The method comprises: obtaining image to be identified;If at least one target object is identified in image to be identified, then based on the first part and the second part of at least one target object in image to be identified, the identity authentication of target object is carried out, identity authentication is used to indicate whether target object grants control authority to target device, control authority is used to indicate that target device responds to the corresponding operation of target action of target object is executed;If target object passes identity authentication, determine the target action with the first part of target object has association relationship;Based on the target action of target object, trigger the corresponding operation of target device executed in the control authority granted by target object.The embodiments of the present application solve the problem that intelligent device cannot be implemented in related technologies Precise control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet of Things technology, and more specifically, to a method, apparatus, device, and medium for controlling an intelligent device. Background Art

[0002] With the development of IoT technology, the use of motion-based control of smart devices, such as gesture control, is becoming increasingly common in interactive scenarios, such as smart homes. Typically, based on images captured in the interactive scene, as long as a gesture can be recognized in the image, the smart device deployed in the interactive scene will respond to the gesture and perform the corresponding operation.

[0003] However, in some interactive scenarios with multiple people, such as museums, exhibition halls, conference venues, etc., only tour guides are allowed to control smart devices using gestures, and similar gestures by visitors or participants are not expected to trigger smart devices to perform corresponding operations.

[0004] From the above, it can be seen that the relevant technology still has the defect that smart devices cannot implement precise control. Summary of the Invention

[0005] The embodiments of the present application provide a method, apparatus, electronic device, and storage medium for controlling an intelligent device, which can solve the problem in related technologies that intelligent devices cannot be accurately controlled. The technical solution is as follows:

[0006] According to one aspect of an embodiment of the present application, a method for controlling a smart device includes: acquiring an image to be identified; if at least one target object is identified in the image to be identified, authenticating the target object based on a first part and a second part of the at least one target object in the image to be identified, wherein the authentication is used to indicate whether the target object grants control authority for a target device, and the control authority is used to indicate that the target device performs a corresponding operation in response to a target action of the target object, and the target action is performed by the second part of the target object; if the target object passes the authentication, determining a target action associated with the first part of the target object; and based on the target action of the target object, triggering the target device in the control authority granted by the target object to perform a corresponding operation.

[0007] According to one aspect of an embodiment of the present application, a smart device control device includes: an image acquisition module for acquiring an image to be identified; an identity authentication module for authenticating the target object based on a first part and a second part of at least one target object in the image to be identified if at least one target object is identified in the image to be identified, the identity authentication being used to indicate whether the target object grants control authority for a target device, the control authority being used to instruct the target device to perform a corresponding operation in response to a target action of the target object, the target action being performed by the second part of the target object; an action determination module for determining a target action associated with the first part of the target object if the target object passes the identity authentication; and a control module for triggering the target device in the control authority granted by the target object to perform a corresponding operation based on the target action of the target object.

[0008] According to one aspect of an embodiment of the present application, an electronic device includes: at least one processor, at least one memory, and at least one communication bus, wherein a computer program is stored on the memory, and the processor reads the computer program in the memory through the communication bus; the computer program is executed by the processor to perform the following steps: acquiring an image to be identified; if at least one target object is identified in the image to be identified, then authenticating the target object based on a first part and a second part of at least one target object in the image to be identified, the authentication is used to indicate whether the target object grants control authority for a target device, the control authority is used to indicate that the target device performs a corresponding operation in response to a target action of the target object, and the target action is performed by the second part of the target object; if the target object passes the authentication, determining a target action associated with the first part of the target object; based on the target action of the target object, triggering the target device in the control authority granted by the target object to perform a corresponding operation.

[0009] In an exemplary embodiment, the processor is further configured to perform the following step: configuring control permissions of the target device for the target object based on the image of the target object.

[0010] In an exemplary embodiment, the processor is also used to perform the following steps: acquiring an image of the target object; performing biometric recognition on a first part of the target object in the image to obtain a first recognition result, and performing action recognition on an action performed by a second part of the target object in the image to obtain a second recognition result; and configuring control permissions of a target device for the target object based on an association between the first part of the target object indicated by the first recognition result and the action performed by the second part of the target object indicated by the second recognition result.

[0011] In an exemplary embodiment, the first part includes a face; the processor is further used to perform the following steps: determining a face area of ​​the target object's face in the image; performing face category prediction on the target object's face in the face area to obtain the first recognition result, and the first recognition result is used to indicate the target object's face.

[0012] In an exemplary embodiment, the second part includes a hand, and the action performed by the second part includes a gesture; the processor is further used to perform the following steps: determining a gesture area of ​​the target object's gesture in the image; performing gesture category prediction on the target object's gesture in the gesture area to obtain the second recognition result, and the second recognition result is used to indicate the gesture of the target object.

[0013] In an exemplary embodiment, the processor is also used to perform the following steps: establishing an association relationship between the face and gesture of the target object based on the face of the target object indicated by the first recognition result and the gesture of the target object indicated by the second recognition result; and enabling control permissions of the target device for the target object based on the association relationship.

[0014] In an exemplary embodiment, the processor is also used to perform the following steps: extract key points of at least one target object in the image to be identified to obtain skeleton information of at least one corresponding target object; perform biometric identification on a first part of at least one target object in the image to be identified to obtain first part information of at least one corresponding target object; perform action recognition on an action performed on a second part of at least one target object in the image to be identified to obtain second part information of at least one corresponding target object; based on at least one skeleton information, match target objects between at least one first part information and at least one second part information to determine the first part information and second part information of the same target object; and determine whether the target object has passed the identity authentication based on the first part information and second part information of the same target object.

[0015] In an exemplary embodiment, the processor is also used to perform the following steps: locating the human body key points of at least one target object in the image to be identified to obtain multiple heat maps, each heat map being used to mark at least one type of human body key point of a target object; based on the human body key points of the target object marked in the multiple heat maps, establishing connections between the human body key points of the same target object to obtain skeleton information of at least one corresponding target object.

[0016] In an exemplary embodiment, the processor is also used to perform the following steps: matching the target object between at least one skeleton information and at least one first part information to determine the skeleton information and first part information of the same target object; matching the target object between at least one skeleton information and at least one second part information to determine the skeleton information and second part information of the same target object; and determining the first part information and second part information of the same target object based on the target object to which each skeleton information belongs.

[0017] In an exemplary embodiment, the first part is a face; the processor is further used to perform the following steps: for each skeleton information, determine the facial key points associated with the skeleton information, and for each first part information, determine the facial area associated with the first part information; for each facial area, detect the number of facial key points in the facial area; if it is detected that there is only one facial key point of a target object in the facial area, then the first part information associated with the facial area and the skeleton information associated with the facial key point in the facial area are determined to belong to the same target object.

[0018] In an exemplary embodiment, the processor is also used to perform the following steps: if facial key points of multiple target objects are detected in the facial area, determining the center point of the facial area; respectively calculating the distances between the facial key points of the multiple target objects and the center point of the facial area, and determining the first part information associated with the facial area and the skeleton information associated with the facial key point with the smallest distance from the center point of the facial area as belonging to the same target object.

[0019] In an exemplary embodiment, the second part is a hand; the processor is further used to perform the following steps: for each piece of skeleton information, determine the hand key points associated with the skeleton information, and for each piece of second part information, determine the gesture area associated with the second part information; for each gesture area, detect the number of hand key points in the gesture area; if it is detected that there is only one hand key point of a target object in the gesture area, then the second part information associated with the gesture area and the skeleton information associated with the hand key point in the gesture area are determined to belong to the same target object.

[0020] In an example embodiment, the processor is further configured to perform the following steps: if it is detected that there are multiple target object hand key points in the gesture region, determining a center point of the gesture region; calculating distances between the multiple target object hand key points and the center point of the gesture region respectively, determining the second part information associated with the gesture region, and the skeleton information associated with the hand key point with the smallest distance between the hand key point and the center point of the gesture region, as belonging to the same target object.

[0021] According to an aspect of the embodiments of the present application, a storage medium has a computer program stored thereon, and the computer program is executed by a processor to implement the intelligent device control method described above.

[0022] According to an aspect of the embodiments of the present application, a computer program product includes a computer program stored in a storage medium, and a processor of a computer device reads the computer program from the storage medium, and the processor executes the computer program, so that the computer device is implemented to execute the intelligent device control method described above.

[0023] The technical scheme provided by the present application has the beneficial effects that:

[0024] In the above technical scheme, if at least one target object is recognized in the to-be-recognized image, the target object is identity verified based on the first part and the second part of the at least one target object in the to-be-recognized image, if the target object passes the identity verification, it indicates that the target object is granted the control right, and it can be determined that the first part of the target object has a correlation with the target action performed by the second part, and then the target action of the target object can be used to trigger the target device in the control right granted by the target object to perform the corresponding operation, that is, through the authorization of the control right of the target device to different objects, the action performed by the authorized object can trigger the target device to perform the corresponding operation, and the non-authorized object without the control right of the target device cannot control the target device by using similar actions, thereby effectively solving the problem that the intelligent device cannot be precisely controlled in the related art. BRIEF DESCRIPTION OF DRAWINGS

[0025] In order to more clearly illustrate the technical schemes in the embodiments of the present application, the drawings needed in the description of the embodiments of the present application will be briefly introduced.

[0026] Figure 1 is a schematic diagram according to the implementation environment involved in the present application;

[0027] Figure 2 is a flowchart of an intelligent device control method according to an example embodiment;

[0028] Figure 3 is a flow chart showing another method for controlling an intelligent device according to an exemplary embodiment;

[0029] Figure 4 yes Figure 3 A flowchart of an embodiment corresponding to step 450 in one embodiment;

[0030] Figure 5 is a schematic diagram illustrating the extraction of key points of a human body through multiple heat maps according to an exemplary embodiment;

[0031] Figure 6 is a flow chart of a method for matching a target object based on skeleton information and face information according to an exemplary embodiment;

[0032] Figure 7 yes Figure 6 A schematic diagram of determining the positional relationship between key points of a human body and a facial area according to the corresponding embodiment;

[0033] Figure 8 is a flow chart of a method for matching a target object based on skeleton information and gesture information according to an exemplary embodiment;

[0034] Figure 9 yes Figure 8 A schematic diagram of determining the positional relationship between the hand key points and the gesture area shown in the corresponding embodiment;

[0035] Figure 10 is a structural block diagram of a smart device control device according to an exemplary embodiment;

[0036] Figure 11 is a hardware structure diagram of an electronic device according to an exemplary embodiment;

[0037] Figure 12 The figure is a structural block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0038] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present application.

[0039] It is to be understood that the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It is further understood that the terms "comprise" and "comprising" and the like, when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It is further understood that when an element or layer is referred to as being "on" or "connected to" another element or layer, it can be directly on or connected to the other element or layer or intervening elements or layers can be present. In addition, the term "connected" or "coupled" as used herein refers to any connection or coupling, either direct or indirect, between otherwise isolated components. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0040] The following is an introduction and explanation of several terms related to the present application:

[0041] Recognition: such as face recognition, gesture recognition, etc., recognition usually includes two main tasks: target positioning and target classification, for example, the target can refer to a face, a gesture. Recognition can be achieved through two different modes of models, which are one-stage mode and two-stage mode. Among them, one-stage mode refers to only performing one of the two main tasks of recognition, target positioning or target classification; two-stage model refers to performing the two main tasks of recognition, target positioning and target classification.

[0042] Deep learning model: a machine learning model, refers to using supervised and unsupervised learning methods to train a deep neural network, which is a neural network structure containing multiple hidden layers.

[0043] In order to make the purpose, technical scheme and advantages of the present application more clear, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0044] Figure 1 A schematic diagram of an implementation environment related to a smart device control method. The implementation environment includes a user terminal 110, a smart device 130, a gateway 150, a server end 170, and a router 190.

[0045] Specifically, the user terminal 110, which can also be referred to as a user end or a terminal, can deploy a client associated with the smart device 130. The user terminal 110 can be an electronic device such as a smartphone, a tablet computer, a notebook computer, a desktop computer, etc., which is not limited herein.

[0046] Among them, the client is associated with the smart device 130, so that when the client is run in the user terminal 110, it can provide the user with functions such as authorization of control permissions for the smart device 130. This client can be in the form of an application or a web page. Accordingly, the user interface for the client to access data configuration can be in the form of a program window or a web page, which is not limited here.

[0047] The smart device 130 is deployed in the gateway 150 and communicates with the gateway 150 through its own configured communication module, and is thereby controlled by the gateway 150. In one application scenario, the smart device 130 accesses the gateway 150 via a local area network and is thus deployed in the gateway 150. The process of the smart device 130 accessing the gateway 150 via the local area network includes: the gateway 150 first establishes a local area network, and the smart device 130 joins the local area network established by the gateway 150 by connecting to the gateway 150. This local area network includes, but is not limited to, ZIGBEE or Bluetooth. Among them, the smart device 130 can be a smart printer, a smart fax machine, a smart camera, a smart air conditioner, a smart door lock, a smart light, or an electronic device equipped with a communication module, such as a human body sensor, door and window sensor, temperature and humidity sensor, water immersion sensor, natural gas alarm, smoke alarm, wall switch, wall socket, wireless switch, wireless wall switch, magic cube controller, curtain motor, etc.

[0048] The interaction between the user terminal 110 and the smart device 130 can be achieved through a local area network or a wide area network. In one application scenario, the user terminal 110 establishes a communication connection between the router 190 and the gateway 150 by means of a wired or wireless method. For example, the wired or wireless method includes but is not limited to WIFI, so that the user terminal 110 and the gateway 150 are deployed in the same local area network, thereby enabling the user terminal 110 to interact with the smart device 130 through the local area network path. In another application scenario, the user terminal 110 establishes a communication connection between the server 170 and the gateway 150 by means of a wired or wireless method. For example, the wired or wireless method includes but is not limited to 2G, 3G, 4G, 5G, WIFI, etc., so that the user terminal 110 and the gateway 150 are deployed in the same wide area network, thereby enabling the user terminal 110 to interact with the smart device 130 through the wide area network path.

[0049] The server side 170 can also be considered as a cloud, cloud platform, platform side, service side, etc. The server side 170 can be a single server, a server cluster composed of multiple servers, or a cloud computing center composed of multiple servers, so as to better provide background services to a large number of user terminals 110. For example, the background services include but are not limited to biometric recognition, motion recognition, model training, etc.

[0050] In an application scenario, through the user terminal 110 (such as a smart phone), the control authority of the target device is configured for the target object. Specifically: based on the image of the target object captured and collected by the built-in camera of the smart phone, biometric recognition and motion recognition are performed respectively, and based on the association between the actions performed by the first part of the target object indicated by the first recognition result and the second part of the target object indicated by the second recognition result, the control authority of the smart device 130 is configured for the target object to notify the gateway 150, so that the gateway 150 can trigger the smart device 130 to perform corresponding operations based on the gesture of the target object.

[0051] In an application scenario, the control authority of the target device is configured for the target object through the gateway 150 (for example, a gateway-type camera). Specifically: in response to the authorization request initiated by the user terminal 110, biometric recognition and motion recognition are performed based on the image of the target object captured and collected by the gateway-type camera, and based on the association between the actions performed by the first part of the target object indicated by the first recognition result and the second part of the target object indicated by the second recognition result, the control authority of the smart device 130 is configured for the target object to trigger the smart device 130 to perform corresponding operations based on the gesture of the target object.

[0052] In one application scenario, the server 170 (e.g., a server) configures control permissions for a target device for a target object. Specifically, in response to an authorization request initiated by the user terminal 110, an image of the target object is captured and collected using the image acquisition module and transmitted to the server. Biometric recognition and action recognition are then performed on the image of the target object. Based on the association between the first part of the target object indicated by the first recognition result and the action performed by the second part of the target object indicated by the second recognition result, the target object is configured with control permissions for the smart device 130. The gateway 150 is then notified, enabling the gateway 150 to trigger the smart device 130 to perform a corresponding action based on the gesture of the target object. The image acquisition module can be a video camera, a camera, or an electronic device equipped with a camera (e.g., a smartphone), and can interact with the server via the gateway 150 / router 190.

[0053] See also Figure 2The embodiment of the application provides a control authority authorization process in a smart device control method.

[0054] As shown in the implementation environment shown in Figure 1 , the electronic device can be Figure 1 a user terminal 110 in , can also be Figure 1 a gateway 150 in , and can also be Figure 1 a server end 170 in , which does not constitute a specific limitation here.

[0055] As shown in Figure 2 , the method can include the following steps:

[0056] Step 310: Obtain an image of a target object.

[0057] The target object, which can also be regarded as an object to be authorized in the control authority authorization process, is an object that requests to control a smart device by an action. In contrast, the authorized object is an object that is allowed to control the smart device by the action. The action can be a hand action (i.e., a gesture), a leg action, or a full-body action, which is not limited here. It can be understood that the object to be authorized and the authorized object can be transformed, that is, the object to be authorized is granted the control authority of the smart device and then is transformed into the authorized object. For example, the object can be at least one of the holder of the user terminal 110 in the implementation environment shown in Figure 1 , or a person, a robot, an animal, etc. that can appear in the interaction scene.

[0058] The image of the target object is a static image of the target object, which can be derived from one frame of image in a video containing multiple frames of image, or from a unique frame of image in a photo, picture, etc. containing one frame of image. Based on this, the control authority authorization in the embodiment is performed in units of frames.

[0059] The image of the target object is obtained by using an image acquisition module to photograph and acquire the target object. The image acquisition module can be a video camera, a camera, or an electronic device configured with a camera, for example, the camera configured for the user terminal 110 in the implementation environment shown in Figure 1 , which can be regarded as the image acquisition module, so as to configure the control authority of the smart device for the holder of the user terminal 110 based on the image of the holder.

[0060] Step 330: Perform biometric recognition on a first part of the target object in the image to obtain a first recognition result, and perform action recognition on an action performed by a second part of the target object in the image to obtain a second recognition result.

[0061] First, biometric recognition is performed on the biometric features of the target object in the image. Biometric features include, but are not limited to, facial features, fingerprints, palm prints, irises, and so on. In one embodiment, biometric recognition may include the following steps: determining a first part of the target object within a first region of the image; and performing a category prediction on the first part of the target object within the first region to obtain a first recognition result. This first recognition result includes at least the category to which the first part of the target object belongs. Alternatively, this first recognition result can be used to indicate the first part of the target object. Depending on the biometric feature, the first part can be a face, finger, palm, eye, and so on, without limitation.

[0062] Specifically, assuming that the biometric feature is a facial feature and the first part is the face, face recognition includes two tasks: face location and face classification. Face location refers to determining the facial region within an image where the face of the target object is located; face classification refers to determining the category to which the face of the target object within the facial region belongs. In one embodiment, face recognition is achieved using a face recognition algorithm. This face recognition algorithm includes at least the RetinaFace algorithm, the ArcFace algorithm, etc. In another embodiment, face recognition is achieved using a two-stage face recognition model. This face recognition model is generated through model training using a machine learning model constructed using the face recognition algorithm. This machine learning model includes at least a deep learning model such as a convolutional neural network. Thus, by performing face recognition on the face of the target object in the image, a first recognition result can be obtained. The first recognition result includes at least the category to which the face of the target object belongs. In other words, the first recognition result indicates the face of the target object, thereby uniquely identifying the identity of the target object.

[0063] Secondly, action recognition is performed on the action performed by the second part of the target object in the image. The second part includes, but is not limited to, hands, legs, limbs, etc. Correspondingly, the action performed by the second part includes, but is not limited to, gestures, leg movements, whole-body movements, etc. In one embodiment, action recognition may include the following steps: determining a second part region of the second part of the target object in the image; performing a category prediction on the action performed by the second part of the target object in the second part region to obtain a second recognition result. The second recognition result includes at least the category to which the action performed by the second part of the target object belongs. It can also be considered that the second recognition result is used to indicate the action performed by the second part of the target object.

[0064] Specifically, assuming the second part is a hand and the action performed by the second part is a gesture, gesture recognition includes two tasks: gesture localization and gesture classification. Gesture localization refers to determining the gesture region within the image where the gesture of the target object is located; gesture classification refers to determining the category to which the gesture of the target object belongs within the gesture region. In one embodiment, gesture recognition is implemented using a gesture recognition algorithm. This gesture recognition algorithm includes at least the SSD algorithm, the MobileNet algorithm, etc. In another embodiment, gesture recognition is implemented using a two-stage gesture recognition model. This gesture recognition model is generated through model training using a machine learning model constructed using the gesture recognition algorithm. This machine learning model includes at least a deep learning model such as a convolutional neural network. Thus, by performing gesture recognition on the gesture of the target object in the image, a second recognition result can be obtained. The second recognition result includes at least the category to which the gesture of the target object belongs; that is, the second recognition result is used to indicate the gesture of the target object.

[0065] It is worth mentioning that the above-mentioned recognition models (such as face recognition models, gesture recognition models, etc.) can be generated by offline training machine learning models on the server and sent to electronic devices, so that the electronic devices can use the above-mentioned recognition models online to perform biometric recognition and / or motion recognition respectively, thereby fully ensuring the operating efficiency of the electronic devices, and thus helping to save the response time of using motion to control smart devices, thereby improving the user experience.

[0066] Step 350 : configuring control authority of a target device for the target object based on the association between the action performed by the first part of the target object indicated by the first recognition result and the action performed by the second part of the target object indicated by the second recognition result.

[0067] That is, after configuring the control authority of the target device for the target object, the target object becomes the authorized object to obtain the control authority of the target device, wherein the control authority is used to instruct the target device to perform a corresponding operation in response to the gesture of the target object.

[0068] In one embodiment, the configuration of control authority is essentially to establish an association relationship between the actions performed by the first part and the second part of the target object in the electronic device, so that based on the association relationship, it is considered that the target object is granted control authority for the target device, that is, the target object is changed to an authorized object.

[0069] In other words, based on the association relationship between the actions performed by the first part and the second part of several authorized objects stored in the electronic device, if the association relationship of the target object can be found to match the association relationship of the authorized object, that is, the target object and the authorized object are consistent in the actions performed by the first part and the second part, then the target object is deemed to be granted control authority over the target device, so that the action of the target object can trigger the target device to perform corresponding operations.

[0070] It is worth mentioning that the target device can refer to all smart devices of the same type in the interactive scene. For example, if the control permissions of all cameras in a museum are configured for the tour guide, the target device refers to all cameras in the museum. It can also be targeted at a specific smart device in the interactive scene. For example, if the control permissions of the projector are configured for the tour guide in the conference room, the target device refers to the projector in the conference room.

[0071] Through the above process, the control authority of the target device is authorized to different objects, so that the action performed by the authorized object can trigger the target device to perform the corresponding operation, while the unauthorized object that has not obtained the control authority of the target device cannot use similar actions to control the target device, thereby effectively solving the problem of the inability of smart devices to implement precise control in related technologies.

[0072] See also Figure 3 , an embodiment of the present application provides a method for controlling an intelligent device, which is described by taking the application of the method to an electronic device as an example.

[0073] Among them, such as Figure 1 In the implementation environment shown, the electronic device may be Figure 1 The user terminal 110 in Figure 1 The gateway 150 in the Figure 1 The server side 170 in the description does not constitute a specific limitation here.

[0074] like Figure 3 As shown, the method may include the following steps:

[0075] Step 410: Acquire an image to be recognized.

[0076] First, the image to be identified is obtained by photographing and collecting the environment where the target object is located. The environment where the target object is located can be some interactive scenes, such as smart home scenes, museums, exhibition halls, conference halls, etc. With the image acquisition module deployed in the above-mentioned interactive scenes, such as a corner of the ceiling of a room, a column in a museum, the body of a robot in a conference hall, etc., it is possible to photograph these interactive scenes and collect images to be identified, for example, an image to be identified of the target object making a gesture. The image acquisition module can be a video camera, a camera, or an electronic device equipped with a camera, for example, Figure 1 The gateway 150 in the illustrated implementation environment (which may be a gateway-type camera) is used to trigger the smart device to perform corresponding operations based on the image to be recognized captured and collected by the gateway-type camera.

[0077] It is understood that shooting can be a single shot or continuous shooting. For the same target object, continuous shooting can produce a video, and the image to be identified can be from one frame of the video. A single shot can produce multiple photos, and the image to be identified can be from one of the photos. In other words, the smart device control in this embodiment can be based on dynamic images, such as a video, or static images, such as a photo.

[0078] Regarding the acquisition of the image to be recognized, the image to be recognized can be an image to be recognized captured in real time by the image acquisition module, or it can be an image to be recognized captured by the image acquisition module during a historical time period pre-stored on the server. In other words, after the image acquisition module captures the image to be recognized, the electronic device can process the image to be recognized in real time, or it can pre-store it for processing, for example, processing it according to a time specified by the operator. Therefore, for the electronic device, it can obtain an image to be recognized captured in real time, or it can obtain an image to be recognized pre-stored, that is, by retrieving an image to be recognized captured during a historical time period. This embodiment does not limit this.

[0079] Secondly, the target object refers to the object in the image to be identified, and whether it has been granted control permissions for the target device. This object may be an authorized object, such as a museum guide, or an unauthorized object, such as a participant in a conference.

[0080] The image to be identified may or may not contain at least one target object. Therefore, it is necessary to determine whether at least one target object can be identified in the image to be identified. In one embodiment, a human body sensor is used to detect human bodies in the environment where the target object is located. If the target object is detected, the image acquisition module is controlled to capture the environment where the target object is located and obtain an image to be identified, which contains at least one target object. In one embodiment, target object recognition is performed on the image to be identified to determine whether the image to be identified contains at least one target object. For example, target object recognition can be implemented using a deep learning model.

[0081] If at least one target object is recognized in the image to be recognized, step 430 is executed; otherwise, step 410 is executed to reacquire the image to be recognized until at least one target object is recognized in the acquired image to be recognized.

[0082] Step 430: If at least one target object is identified in the image to be identified, identity verification is performed on the target object based on the first part and the second part of the at least one target object in the image to be identified.

[0083] Authentication is used to indicate whether the target object has granted control authority to the target device. Control authority is used to instruct the target device to perform a corresponding operation in response to the target object's target action, where the target action is performed by the second part of the target object. Specifically, based on a plurality of associations between actions performed by the first part and the second part of the authorization object stored in the electronic device, an association of the authorization object that matches the association of the target object is searched for, thereby verifying whether the target object has passed the authentication for the control authority authorization.

[0084] If an association relationship of the authorization object that matches the association relationship of the target object can be searched, that is, the target object and the authorization object are consistent in the actions performed in the first part and the second part, then it is considered that the target object has passed the identity authentication in the control authority authorization, that is, the target object is the authorization object, and jump to execute step 450.

[0085] Otherwise, if no association relationship of the authorized object that matches the association relationship of the target object is found, the target object is considered to be an unauthorized object.

[0086] Furthermore, if there are other target objects in the image to be identified, the authentication of the control permission authorization for the other target objects is continued, that is, the process returns to step 430 until all target objects in the image to be identified have completed authentication, and the process returns to step 410.

[0087] Step 450: If the target object passes the identity authentication, a target action associated with the first part of the target object is determined.

[0088] When the target object passes the authentication, it is considered that the target object is granted the control permission of the target device, that is, the target object is allowed to control the target device using the target action.

[0089] Since the target object and the authorized object are consistent in the actions performed on the first part and the second part, the target action associated with the first part of the target object can be the action associated with the first part of the authorized object stored by the electronic device during the control authority authorization, or it can be the target action performed by the second part of the identified target object, which is not limited here.

[0090] Step 470 : Based on the target action of the target object, trigger the target device in the control authority granted by the target object to perform a corresponding operation.

[0091] Through the above process, the target object's control authority over the smart device is accurately identified, thereby fully ensuring the precise control of the smart device. That is, only actions performed by authorized objects can trigger the smart device to perform corresponding operations, while unauthorized objects cannot use similar actions to control the smart device.

[0092] See also Figure 4 , an embodiment of the present application provides a possible implementation method, step 450 may include the following steps:

[0093] Step 451 : extract key points of at least one target object in the image to be identified, and obtain skeleton information of at least one corresponding target object.

[0094] Keypoint extraction is performed on the human body of the target object in the image to be identified. Specifically, keypoint extraction can be achieved using the OpenPose algorithm, or by training a machine learning model built using the OpenPose algorithm to generate a keypoint extraction model. This machine learning model includes at least a deep learning model such as a convolutional neural network.

[0095] In one embodiment, the key point extraction process can include the following steps: human key point positioning on the human body of at least one target object in the to-be-identified image to obtain a plurality of heat maps; and based on the human key points of the target object marked in the plurality of heat maps, connection is established between the human key points of the same target object to obtain the skeleton information of at least one corresponding target object. Each heat map is used to mark at least one type of human key point of a target object, and the type of the human key point includes but is not limited to: eye key point, nose key point, ear key point, elbow key point, wrist key point, neck key point, shoulder key point, hip key point, knee key point, ankle key point, etc.

[0096] Figure 5 A schematic diagram of human key point extraction through a plurality of heat maps is shown, Figure 5 (a) is a to-be-identified image containing a target object, Figure 5 (b) is a heat map containing a nose key point 401 of the target object, Figure 5 (c) is a heat map containing wrist key points 402, 403 of the target object, Figure 5 (d) is a heat map containing ankle key points 404, 405 of the target object, and by Figure 5 (b) to Figure 5 (d), the human key points of the target object can be connected to form a skeleton 406 of the target object, as shown in Figure 5 (e).

[0097] Thus, the skeleton information describes the positions of the human key points of the corresponding target object in the to-be-identified image, and can also be understood as indicating the human key points of the corresponding target object. In one embodiment, the human key points are represented by the coordinate values of the pixel points in the to-be-identified image. It should be understood that different skeleton information is associated with the human key points of different target objects, that is, each skeleton information uniquely identifies the human key points of the corresponding target object in the to-be-identified image.

[0098] After determining the skeleton information of the target object, the first part key point (for example, the nose key point) and the second part key point (for example, the wrist key point) of the same target object can be determined in the to-be-identified image based on the human key points identified by the skeleton information, and then the first part and the second part of the same target object are matched from the first part of at least one target object and the second part of at least one target object in the to-be-identified image. In this embodiment, the first part of the target object is uniquely identified by the first part information, and the first part information is obtained by step 453; and the gesture of the target object is uniquely identified by the second part information, and the second part information is obtained by step 455.

[0099] Step 453 : Perform biometric recognition on a first part of at least one target object in the image to be recognized, and obtain information of the first part of at least one corresponding target object.

[0100] The inventors realized that since the target object is matched based on the key points of the human body identified by the skeleton information, the first part information only needs to reflect the key points of the first part, regardless of the category of the first part. In other words, the biometric recognition here does not focus on the first part classification task; only the first part positioning task can be used.

[0101] In one embodiment, biometric feature recognition includes at least one task: first part positioning, which refers to determining a first part area where a first part of at least one target object in an image to be recognized is located.

[0102] The first part information describes the location of the first part of the corresponding target object in the image to be recognized. It can also be understood that the first part information is used to indicate the first part region of the corresponding target object. It should be understood that different first part information is associated with different first part regions of the target object, so that the first part information can uniquely identify the first part of the corresponding target object in the image to be recognized.

[0103] It should be noted here that, unlike the two-stage mode recognition model, which fully guarantees the recognition accuracy, the one-stage mode recognition model has higher recognition efficiency and does not significantly increase resources, which is conducive to the realization of a lightweight edge deployment solution.

[0104] Step 455 : Perform action recognition on the action performed on the second part of at least one target object in the image to be recognized, and obtain information of the second part of at least one corresponding target object.

[0105] Similar to the first part information, the second part information describes the location of the second part of the corresponding target object in the image to be recognized. Alternatively, the second part information can be understood as indicating the second part region of the corresponding target object. As will be readily understood, different second part information is associated with different second part regions of the target object, enabling the second part information to uniquely identify the second part of the corresponding target object in the image to be recognized.

[0106] In one embodiment, second part recognition includes at least one task: second part positioning, which refers to determining a second part region where a second part of at least one target object in the image to be recognized is located.

[0107] It is worth mentioning that the order of the above steps 451 to 455 is not limited to the sequential execution in this embodiment. Step 453 and step 455 can also be executed simultaneously before step 451, or step 453 and step 455 can be executed in sequence before step 451. Step 451, step 453, and step 455 can also be executed in parallel. This can save the response time of using actions to control smart devices, which is beneficial to improving the user experience. This does not constitute a specific limitation here.

[0108] Step 457 : Based on the at least one skeleton information, match the target object between the at least one first part information and the at least one second part information to determine the first part information and the second part information of the same target object.

[0109] The inventors recognized that since the skeleton information, first part information, and second part information are obtained through different recognition tasks, it is necessary to use a target object matching method to determine the first and second part information of the same target object, and then determine whether the target object has passed identity verification. In this embodiment, the target object matching is achieved based on the relationship between the first part information and the skeleton information, and the relationship between the second part information and the skeleton information.

[0110] Specifically, target object matching is performed between at least one skeleton information and at least one first part information to determine the skeleton information and first part information of the same target object.

[0111] Matching of the target object is performed between the at least one skeleton information and the at least one second part information to determine the skeleton information and the second part information of the same target object.

[0112] Based on the target object to which each skeleton information belongs, first part information and second part information of the same target object are determined.

[0113] For example, assume that the image to be identified contains three target objects A, B, and C.

[0114] Then, through human body key point extraction, we can obtain the skeleton information A1 of target object A, the skeleton information B1 of target object B, and the skeleton information C1 of target object C; through biometric recognition, we can obtain the first part information A2 of target object A, the first part information B2 of target object B, and the first part information C2 of target object C; through motion recognition, we can obtain the second part information A3 of target object A, the second part information B3 of target object B, and the second part information C3 of target object C.

[0115] Based on this, target objects are matched between the skeleton information A1, B1, C1 and the first part information A2, B2, C2 to determine that the skeleton information A1 and the first part information A2 belong to the same target object A, the skeleton information B1 and the first part information B2 belong to the same target object B, and the skeleton information C1 and the first part information C2 belong to the same target object C.

[0116] Match the target objects between the skeleton information A1, B1, C1 and the second part information A3, B3, C3 to determine that the skeleton information A1 and the second part information A3 belong to the same target object A, the skeleton information B1 and the second part information B3 belong to the same target object B, and the skeleton information C1 and the second part information C3 belong to the same target object C.

[0117] From the above, we can see that skeleton information A1, first part information A2, and second part information A3 belong to the same target object A. Therefore, we can confirm that first part information A2 and second part information A3 belong to the same target object A. Similarly, we can confirm that first part information B2 and second part information B3 belong to the same target object B, and first part information C2 and second part information C3 belong to the same target object C.

[0118] Step 459: Determine whether the target object passes identity authentication based on the first part information and the second part information of the same target object.

[0119] Still using the above example for explanation, it is assumed that the electronic device stores an association relationship between the first part of the target object A and the action performed by the second part, wherein the first part of the target object A and the second part performing the action are uniquely identified by the first part information A2 and the second part information A3 respectively.

[0120] Based on this, it can be determined that target object A passes identity authentication and is considered an authorized object; while target objects B and C fail identity authentication and are considered unauthorized objects.

[0121] Under the influence of the above-mentioned embodiment, identity authentication of the target object in the control authority authorization is realized, which serves as the basis for the accurate identification of the target object's control authority over the smart device, thereby facilitating the accurate control of the smart device, that is, the actions performed by the authorized object can trigger the smart device to perform the corresponding operation, while the unauthorized object cannot use similar actions to control the smart device.

[0122] See also Figure 6 In an embodiment of the present application, a possible implementation method is provided for matching a target object based on skeleton information and information of a first part (taking a face as an example), which is achieved by determining the positional relationship between facial key points and facial regions. Specifically, the following steps may be included:

[0123] Step 510 : for each skeleton information, determine the facial key points associated with the skeleton information, and for each first part information, determine the facial area associated with the first part information.

[0124] In one embodiment, the facial landmark is a nose landmark.

[0125] In one embodiment, the face region is identified by a positioning frame in the image to be recognized. It should be noted that the shape of the positioning frame can be rectangular, circular, triangular, etc., and the position is represented by the coordinate values ​​of the pixel points in the image to be recognized.

[0126] Step 530: For each face region, detect the number of facial key points in the face region.

[0127] If it is detected that there is only one facial key point of the target object in the face region, step 550 is executed.

[0128] Otherwise, if multiple facial key points of the target object are detected in the face region, step 570 is executed.

[0129] Step 550: Determine that the first part information associated with the face region and the skeleton information associated with the facial key points in the face region belong to the same target object.

[0130] like Figure 7 As shown in (a), for target object 501, first part information is obtained through face recognition. This first part information is used to indicate a facial region 502 of target object 501. Skeleton information is obtained through body key point extraction. This skeleton information at least indicates a nose key point 503 of target object 501. Then, the number of facial key points is detected for facial region 502. Since there is only one nose key point 503 in facial region 502, it is assumed that the first part information associated with facial region 502 and the skeleton information associated with nose key point 503 in facial region 502 belong to the same target object, namely, target object 501.

[0131] Step 570, respectively calculate the distances between the facial key points of multiple target objects and the center point of the facial area, and determine the first part information associated with the facial area and the skeleton information associated with the facial key point with the smallest distance from the center point of the facial area as belonging to the same target object.

[0132] like Figure 7As shown in (b), for the target object 501, the first part information is obtained through face recognition, and this first part information is used to indicate the face area 502 of the target object 501. The skeleton information is obtained through human body key point extraction, and this skeleton information at least indicates the nose key point 503 of the target object 501. Similarly, for another target object (not shown in the figure), the first part information is obtained through face recognition, and this first part information is used to indicate the face area 504 of the other target object. The skeleton information is obtained through human body key point extraction, and this skeleton information at least indicates the nose key point 505 of the target object 504. Then, the number of facial key points is detected for face area 502 and face area 504. There are nose key points 503 and nose key points 505 in face area 502, and there are nose key points 503 and nose key points 505 in face area 504. Since the distance between nose key point 503 and the center point of face area 502 is the smallest, and the distance between nose key point 505 and the center point of face area 504 is the smallest, it is considered that the first part information associated with face area 502 and the skeleton information associated with nose key point 503 in face area 502 belong to the same target object, that is, target object 501; while the first part information associated with face area 504 and the skeleton information associated with nose key point 505 in face area 504 belong to another target object.

[0133] In the above process, the skeleton information and the first part information are identified as belonging to the same target object, which is used as the basis for identifying that the first part information and the second part information belong to the same target object, thereby facilitating the identity authentication of the target object in the control authority authorization.

[0134] See also Figure 8 In an embodiment of the present application, a possible implementation method is provided for matching a target object based on skeleton information and information of a second part (e.g., a hand), which is achieved by determining the positional relationship between the key points of the hand and the gesture area. Specifically, the following steps may be included:

[0135] Step 610 : for each piece of skeleton information, determine the hand key points associated with the skeleton information, and for each piece of second part information, determine the gesture area associated with the second part information.

[0136] In one embodiment, the hand keypoints are wrist keypoints.

[0137] In one embodiment, the gesture area identifies the hand of the target object performing the gesture by using a positioning frame in the image to be recognized. Alternatively, the positioning frame can be shaped like a rectangle, a circle, a triangle, etc., and its position is represented by the coordinates of pixels in the image to be recognized.

[0138] It is worth mentioning that the hand of the target object marked by the positioning frame that performs the gesture is actually the palm, and the hand key point associated with the skeleton information is actually the wrist. Then, considering the positions of the palm and wrist in the hand under different gestures, it is possible that the wrist key point exists in the gesture area, and it is also possible that the wrist key point cannot exist in the gesture area. For this reason, the inventor designed to extend the connecting line between the wrist key point and the elbow key point in the direction pointing to the gesture area by a set length (for example, 1 / 3 of the length of the connecting line), and use the end point of this connecting line as an auxiliary wrist key point located in the gesture area, thereby ensuring that the wrist key point can exist in the gesture area, so as to facilitate subsequent target object matching. In other words, in this embodiment, the hand key point can refer to the actual wrist key point obtained by extracting the human body key point, or it can refer to the auxiliary wrist key point formed by extending the connecting line.

[0139] Step 630: For each gesture area, detect the number of hand key points in the gesture area.

[0140] If it is detected that there is only one hand key point of the target object in the gesture area, step 650 is executed.

[0141] Otherwise, if it is detected that there are multiple hand key points of the target object in the gesture area, step 670 is executed.

[0142] Step 650 : Determine that the second part information associated with the gesture area and the skeleton information associated with the hand key points in the gesture area belong to the same target object.

[0143] like Figure 9 As shown in (a), for target object 601, gesture recognition yields second part information, which indicates gesture region 602 of target object 601. Human body key point extraction yields skeleton information, which indicates at least wrist key point 603 of target object 601. Then, the number of hand key points is detected for gesture region 602. Since there is only one wrist key point 603 in gesture region 602, the second part information associated with gesture region 602 and the skeleton information associated with wrist key point 603 in gesture region 602 are assumed to belong to the same target object, namely, target object 601.

[0144] Step 670, respectively calculate the distances between the hand key points of multiple target objects and the center point of the gesture area, and determine that the second part information associated with the gesture area and the skeleton information associated with the hand key point with the smallest distance from the center point of the gesture area belong to the same target object.

[0145] like Figure 9As shown in (b), for the target object 601, the second part information is obtained through gesture recognition, and this second part information is used to indicate the gesture area 602 of the target object 601. The skeleton information is obtained through human body key point extraction, and this skeleton information at least indicates the wrist key point 603 of the target object 601. Similarly, for another target object (not shown in the figure), the second part information is obtained through gesture recognition, and this second part information is used to indicate the gesture area 604 of the other target object. The skeleton information is obtained through human body key point extraction, and this skeleton information at least indicates the wrist key point 605 of the target object 604. Then, the number of hand key points is detected for gesture area 602 and gesture area 604. There are wrist key point 603 and wrist key point 605 in gesture area 602, and there is only one wrist key point 605 in gesture area 604. Since the distance between wrist key point 603 and the center point of gesture area 602 is the smallest, it is considered that the second part information associated with gesture area 602 and the skeleton information associated with wrist key point 603 in gesture area 602 belong to the same target object, that is, target object 601; and the second part information associated with gesture area 604 and the skeleton information associated with wrist key point 605 in gesture area 604 belong to another target object.

[0146] Under the effect of the above-mentioned embodiments, on the one hand, the identification that the skeleton information and the second part information belong to the same target object is realized, which serves as the basis for identifying that the first part information and the second part information belong to the same target object, thereby facilitating the identity verification of the target object in the authorization of control authority; on the other hand, by associating the second part information and the first part information through the skeleton information, the identity verification of the target object in the authorization of control authority can be realized simply and efficiently, and at the same time it is conducive to reducing development costs.

[0147] The following are embodiments of the apparatus of the present application, which can be used to execute the intelligent device control method involved in the present application. For details not disclosed in the apparatus embodiments of the present application, please refer to the method embodiments of the intelligent device control method involved in the present application.

[0148] See also Figure 10 In an embodiment of the present application, a smart device control apparatus 900 is provided, including but not limited to: an image acquisition module 910 , an identity authentication module 930 , an action determination module 950 and a control module 970 .

[0149] The image acquisition module 910 is used to acquire an image to be identified.

[0150] The identity authentication module 930 is used to authenticate the target object based on the first part and the second part of the at least one target object in the image to be identified if at least one target object is identified in the image to be identified. The identity authentication is used to indicate whether the target object is granted control authority for the target device. The control authority is used to indicate that the target device performs a corresponding operation in response to the target action of the target object, and the target action is performed by the second part of the target object.

[0151] The action determination module 950 is configured to determine a target action associated with the first part of the target object if the target object passes the identity authentication.

[0152] The control module 970 is configured to trigger the target device in the control authority granted by the target object to perform a corresponding operation based on the target action of the target object.

[0153] In an exemplary embodiment, the intelligent device control apparatus 900 further includes: an authorization module configured to configure control authority of a target device for a target object based on an image of the target object.

[0154] In an exemplary embodiment, the authorization module includes: an image acquisition unit for acquiring an image of a target object; an identification unit for performing biometric identification on a first part of the target object in the image to obtain a first identification result, and performing action identification on an action performed by a second part of the target object in the image to obtain a second identification result; and a permission configuration module for configuring control permissions of a target device for the target object based on the association between the biometrics of the target object indicated by the first identification result and the action performed by the second part of the target object indicated by the second identification result.

[0155] In an exemplary embodiment, the first part includes a face; the recognition unit includes: a face area determination subunit, used to determine the face area of ​​the target object's face in the image; an identity prediction subunit, used to predict the face category of the target object's face in the face area to obtain a first recognition result, and the first recognition result is used to indicate the facial features of the target object.

[0156] In an exemplary embodiment, the second part includes a hand, and the action performed by the second part includes a gesture; the recognition unit also includes: a gesture area determination subunit, used to determine the gesture area of ​​the target object's gesture in the image; a category prediction subunit, used to predict the gesture category of the target object's gesture in the gesture area to obtain a second recognition result, and the second recognition result is used to indicate the gesture of the target object.

[0157] In an exemplary embodiment, the identity authentication module 930 includes: an extraction unit for extracting key points of at least one target object in the image to be identified, and obtaining skeleton information of at least one corresponding target object; a feature recognition unit for performing biometric feature recognition on a first part of at least one target object in the image to be identified, and obtaining first part information of at least one corresponding target object; an action unit for performing action recognition on an action performed on a second part of at least one target object in the image to be identified, and obtaining second part information of at least one corresponding target object; a matching unit for matching target objects between at least one first part information and at least one second part information based on at least one skeleton information, and determining the first part information and second part information of the same target object; and a verification unit for determining whether the target object has passed the identity authentication based on the first part information and second part information of the same target object.

[0158] In an exemplary embodiment, the extraction unit includes: a positioning unit, used to locate the human body key points of at least one target object in the image to be identified, and obtain multiple heat maps, each heat map is used to mark at least one type of human body key point of a target object; a connection unit, used to establish connections between the human body key points of the same target object based on the human body key points of the target objects marked in multiple heat maps, and obtain skeleton information of at least one corresponding target object.

[0159] In an exemplary embodiment, the matching unit includes: a first matching subunit, used to match the target object between at least one skeleton information and at least one first part information, and determine the skeleton information and first part information of the same target object; a second matching subunit, used to match the target object between at least one skeleton information and at least one second part information, and determine the skeleton information and second part information of the same target object; a matching determination subunit, used to determine the first part information and second part information of the same target object based on the target object to which each skeleton information belongs.

[0160] In an exemplary embodiment, the first part is a face; the first matching subunit includes: a key point and face area determination subunit, which is used to determine the face key points associated with the skeleton information for each piece of skeleton information, and to determine the face area associated with the first part information for each piece of first part information; a face key point number detection subunit, which is used to detect the number of face key points in the face area for each face area; and a first determination subunit, which is used to determine the first part information associated with the face area and the skeleton information associated with the face key points in the face area as belonging to the same target object if it is detected that there is only one face key point of a target object in the face area.

[0161] In an exemplary embodiment, the first matching subunit also includes: a face center point determination subunit, which is used to determine the center point of the face area if multiple face key points of target objects are detected in the face area; a second determination subunit, which is used to respectively calculate the distances between the face key points of multiple target objects and the center point of the face area, and determine the first part information associated with the face area and the skeleton information associated with the face key point with the smallest distance to the center point of the face area as belonging to the same target object.

[0162] In an exemplary embodiment, the second part is a hand; the second matching subunit includes: a key point and gesture area determination subunit, which is used to determine the hand key points associated with each skeleton information, and to determine the gesture area associated with each second part information; a hand key point number detection subunit, which is used to detect the number of hand key points in the gesture area for each gesture area; and a third determination subunit, which is used to determine the second part information associated with the gesture area and the skeleton information associated with the hand key points in the gesture area as belonging to the same target object if it is detected that there is only one hand key point of a target object in the gesture area.

[0163] In an exemplary embodiment, the hand center point determination subunit is used to determine the center point of the gesture area if multiple hand key points of target objects are detected in the gesture area; the fourth determination subunit is used to respectively calculate the distances between the hand key points of the multiple target objects and the center point of the gesture area, and determine the second part information associated with the gesture area and the skeleton information associated with the hand key point with the smallest distance from the center point of the gesture area as belonging to the same target object.

[0164] It should be noted that the intelligent device control device provided in the above embodiment only uses the division of the above-mentioned functional modules as an example when performing intelligent device control. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the intelligent device control device will be divided into different functional modules to complete all or part of the functions described above.

[0165] In addition, the embodiments of the intelligent device control apparatus and the intelligent device control method provided in the above embodiments belong to the same concept, wherein the specific manner in which each module performs operations has been described in detail in the method embodiments and will not be repeated here.

[0166] Figure 11 According to an exemplary embodiment, a structural diagram of an electronic device is shown. The electronic device is suitable for Figure 1 The user terminal, gateway, and server in the implementation environment are shown.

[0167] It should be noted that the electronic device is only an example adapted for this application and cannot be considered to provide any limitation on the scope of use of this application. The electronic device cannot be interpreted as needing to rely on or must have Figure 11 One or more components of exemplary electronic device 2000 are shown.

[0168] The hardware structure of the electronic device 2000 may vary greatly due to different configurations or performances, such as Figure 11 As shown, the electronic device 2000 includes a power supply 210 , an interface 230 , at least one memory 250 , and at least one central processing unit (CPU) 270 .

[0169] Specifically, the power supply 210 is used to provide operating voltage for various hardware devices on the electronic device 2000 .

[0170] The interface 230 includes at least one wired or wireless network interface for interacting with external devices. Figure 1 The interaction between the terminal 100 and the electronic device 200 in the implementation environment is shown.

[0171] Of course, in other examples adapted by this application, the interface 230 may further include at least one serial-to-parallel conversion interface 233, at least one input-output interface 235, and at least one USB interface 237, etc. Figure 11 As shown, this does not constitute a specific limitation.

[0172] The memory 250 serves as a carrier for resource storage and can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon include an operating system 251, application 253 and data 255, etc. The storage method can be temporary storage or permanent storage.

[0173] Among them, the operating system 251 is used to manage and control the various hardware devices and application programs 253 on the electronic device 200 to enable the central processing unit 270 to calculate and process the massive data 255 in the memory 250. It can be WindowsServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0174] The application program 253 is a computer program that performs at least one specific task based on the operating system 251 and may include at least one module ( Figure 11 (not shown), each module can respectively include a computer program for the electronic device 2000. For example, the intelligent device control apparatus can be regarded as an application 253 deployed on the electronic device 2000.

[0175] The data 255 may be videos, images, etc. stored in a disk, or various types of recognition models (such as face recognition models, gesture recognition models), etc., stored in the memory 250 .

[0176] The central processing unit 270 may include one or more processors and is configured to communicate with the memory 250 via at least one communication bus to read computer programs stored in the memory 250, thereby performing operations and processing on the massive amount of data 255 in the memory 250. For example, the intelligent device control method may be implemented by the central processing unit 270 reading a series of computer programs stored in the memory 250.

[0177] In addition, the present application can also be implemented through hardware circuits or hardware circuits combined with software. Therefore, the implementation of the present application is not limited to any specific hardware circuits, software, or a combination of the two.

[0178] See also Figure 12 In an embodiment of the present application, an electronic device 4000 is provided. The electronic device 4000 may include: a gateway-type camera, a gateway, a smart phone, a tablet computer, a server, etc.

[0179] exist Figure 12 In the embodiment, the electronic device 4000 includes at least one processor 4001, at least one communication bus 4002 and at least one memory 4003.

[0180] The processor 4001 and the memory 4003 are connected, for example, via a communication bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which may be used for data exchange between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.

[0181] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0182] The communication bus 4002 may include a path for transmitting information between the above components. The communication bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The communication bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 12 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0183] The memory 4003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.

[0184] The memory 4003 stores a computer program, and the processor 4001 reads the computer program stored in the memory 4003 through the communication bus 4002 .

[0185] When the computer program is executed by the processor 4001 , the smart device control method in the above-mentioned embodiments is implemented.

[0186] In addition, an embodiment of the present application provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the smart device control method in the above embodiments is implemented.

[0187] In an embodiment of the present application, a computer program product is provided, which includes a computer program stored in a storage medium. A processor of a computer device reads the computer program from the storage medium and executes the computer program, causing the computer device to perform the intelligent device control method of each of the above embodiments.

[0188] Compared with related technologies, the control authority of the target device is authorized to different objects, so that the action performed by the authorized object can trigger the target device to perform corresponding operations, while the unauthorized object that has not obtained the control authority of the target device cannot use similar actions to control the target device, thereby effectively solving the problem that smart devices in related technologies cannot implement precise control.

[0189] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0190] The above description is only a partial implementation method of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for controlling an intelligent device, characterized in that: The method comprises: Obtain the image to be recognized; If at least one target object is identified in the image to be identified, matching is performed with the same target object based on skeleton information of at least one target object, information of a first part of at least one target object, and information of a second part of at least one target object, determining the first part and the second part of the same target object in the image to be identified, and performing identity authentication on the target object based on the first part and the second part of the same target object in the image to be identified, the identity authentication being used to indicate whether the target object is granted control authority for a target device, the control authority being used to instruct the target device to perform a corresponding operation in response to a target action of the target object, the target action being performed by the second part of the target object; If the target object passes the identity authentication, determining a target action associated with the first part of the target object; Based on the target action of the target object, the target device in the control authority granted by the target object is triggered to perform a corresponding operation.

2. The method according to claim 1, wherein If the target object passes the identity authentication, before determining the target action associated with the first part of the target object, the method further includes: Based on the image of the target object, configuring control permissions of the target device for the target object; The configuring the control permission of the target device for the target object based on the image of the target object includes: Acquiring an image of the target object; performing biometric recognition on a first part of the target object in the image to obtain a first recognition result, and performing action recognition on an action performed by a second part of the target object in the image to obtain a second recognition result; Based on the association between the action performed by the first part of the target object indicated by the first recognition result and the action performed by the second part of the target object indicated by the second recognition result, the control authority of the target device is configured for the target object.

3. The method according to claim 2, wherein The first part includes a human face; The performing biometric recognition on the first part of the target object in the image to obtain a first recognition result includes: Determine a face area of ​​the target object's face in the image; A face category prediction is performed on the face of the target object in the face area to obtain the first recognition result, where the first recognition result is used to indicate the face of the target object.

4. The method according to claim 3, wherein The second part includes a hand, and the action performed by the second part includes a gesture; The performing action recognition on the action performed on the second part of the target object in the image to obtain a second recognition result includes: determining a gesture area of ​​the target object's gesture in the image; A gesture category prediction is performed on the gesture of the target object in the gesture area to obtain a second recognition result, where the second recognition result is used to indicate the gesture of the target object.

5. The method according to any one of claims 1 to 4, characterized in that The matching of the same target object based on the skeleton information of at least one target object, the first part information of at least one target object, and the second part information of at least one target object, determining the first part and the second part of the same target object in the image to be identified, and authenticating the target object based on the first part and the second part of the same target object in the image to be identified, includes: Extracting key points of at least one target object in the image to be identified to obtain skeleton information of at least one corresponding target object; Performing biometric recognition on a first part of at least one target object in the image to be recognized to obtain at least one first part information corresponding to the target object; performing action recognition on an action performed on a second part of at least one target object in the image to be recognized, and obtaining information on the second part of at least one corresponding target object; Based on at least one skeleton information, matching the target object between at least one first part information and at least one second part information to determine the first part information and the second part information of the same target object; Determine whether the target object passes the identity authentication according to the first part information and the second part information of the same target object.

6. The method according to claim 5, wherein The step of extracting key points of at least one target object in the image to be identified to obtain skeleton information of at least one corresponding target object includes: Positioning key points of a human body of at least one target object in the image to be identified to obtain a plurality of heat maps, each heat map being used to mark at least one type of key point of a human body of the target object; Based on the human body key points of the target objects marked in multiple heat maps, connections are established between the human body key points of the same target object to obtain skeleton information of at least one corresponding target object.

7. The method according to claim 5, wherein The matching of the target object between at least one first part information and at least one second part information based on at least one skeleton information to determine the first part information and the second part information of the same target object includes: Matching the target object between at least one skeleton information and at least one first part information to determine the skeleton information and first part information of the same target object; Matching the target object between at least one skeleton information and at least one second part information to determine the skeleton information and the second part information of the same target object; Based on the target object to which each skeleton information belongs, first part information and second part information of the same target object are determined.

8. The method according to claim 7, wherein The first part is a human face; The matching of the target object between at least one skeleton information and at least one first part information to determine the skeleton information and first part information of the same target object includes: For each skeleton information, determining a facial key point associated with the skeleton information, and for each first part information, determining a facial region associated with the first part information; For each face region, detecting the number of facial key points in the face region; If it is detected that there is only one facial key point of the target object in the face region, the first part information associated with the face region and the skeleton information associated with the facial key point in the face region are determined to belong to the same target object.

9. The method according to claim 8, wherein The matching of the target object between the at least one skeleton information and the at least one first part information to determine the skeleton information and the first part information of the same target object further includes: If it is detected that there are multiple facial key points of the target object in the face area, determining the center point of the face area; The distances between the facial key points of multiple target objects and the center point of the facial area are calculated respectively, and the first part information associated with the facial area and the skeleton information associated with the facial key point with the smallest distance from the center point of the facial area are determined to belong to the same target object.

10. The method according to claim 7, wherein: The second part is a hand; The matching of the target object between at least one skeleton information and at least one second part information to determine the skeleton information and the second part information of the same target object includes: For each piece of skeleton information, determining a hand key point associated with the skeleton information, and for each piece of second part information, determining a gesture area associated with the second part information; For each gesture area, detecting the number of hand key points in the gesture area; If it is detected that there is only one hand key point of the target object in the gesture area, the second part information associated with the gesture area and the skeleton information associated with the hand key point in the gesture area are determined to belong to the same target object.

11. The method according to claim 10, wherein The method further comprises: If multiple hand key points of the target object are detected in the gesture area, determining the center point of the gesture area; The distances between the hand key points of multiple target objects and the center point of the gesture area are calculated respectively, and the second part information associated with the gesture area and the skeleton information associated with the hand key point with the smallest distance from the center point of the gesture area are determined to belong to the same target object.

12. A smart device control device, characterized in that: include: An image acquisition module, used to acquire an image to be identified; an identity authentication module, configured to, if at least one target object is identified in the image to be identified, match the target object with the same target object based on skeleton information of the at least one target object, information about a first part of the at least one target object, and information about a second part of the at least one target object, determine the first part and the second part of the same target object in the image to be identified, and authenticate the target object based on the first part and the second part of the same target object in the image to be identified, wherein the identity authentication indicates whether the target object is granted control authority for a target device, and the control authority indicates that the target device performs a corresponding operation in response to a target action of the target object, wherein the target action is performed by the second part of the target object; an action determination module, configured to determine a target action associated with a first part of the target object if the target object passes the identity authentication; The control module is configured to trigger the target device in the control authority granted by the target object to perform a corresponding operation based on the target action of the target object.

13. An electronic device, characterized in that: include: at least one processor, at least one memory, and at least one communication bus, wherein: The memory stores a computer program, and the processor reads the computer program in the memory through the communication bus; When the computer program is executed by the processor, the smart device control method according to any one of claims 1 to 11 is implemented.

14. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the smart device control method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Virtual data acquisition and transmission system taking both hands of humankind as carrier

    CN103309446A

  • Electronic device controlling and user registration method

    CN106462681A

  • Information sending method, information display method and mobile terminal

    CN107181852A

  • Image recognition method and device, medium and electronic equipment thereof

    CN112580544A

  • Association method and system based on key points and medium

    CN112800825A