Target object recognition method and electronic device

By calculating the overlap between the target object and the background object in the target image frame in the electronic device, the misidentification problem is solved, and the recognition accuracy and user experience are improved.

CN118470311BActive Publication Date: 2025-06-06HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311402484.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-26
Publication Date
2025-06-06
Estimated Expiration
2043-10-26

AI Technical Summary

Technical Problem

In electronic devices, images collected by the camera may contain non-target objects (such as faces), resulting in misidentification of the target objects (such as user's hands), thereby performing incorrect operations and affecting the user's user experience.

Method used

By obtaining the target image frame, performing object detection, obtaining the detection result, and when the detection result includes two recognition objects, the overlap between the two recognition objects is calculated. If the overlap degree is less than or equal to the target overlap degree threshold, a target operation corresponding to the first identified object is performed.

Benefits of technology

It reduces the probability of false detection of electronic devices, realizes filtering of background objects, improves the recognition accuracy of target objects, reduces errors caused by target objects recognition errors, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118470311B_ABST
    Figure CN118470311B_ABST
Patent Text Reader

Abstract

The embodiment of the present application is applied to the field of image processing technology, and provides a target object recognition method and an electronic device. The electronic device obtains a target image frame, performs target detection on the target image frame, and obtains a detection result. Afterwards, when the detection result indicates that the target image frame includes two recognition objects, the electronic device determines the overlap between the two recognition objects, wherein the two recognition objects include a first recognition object, and the overlap is the overlap degree of the overlapping area between the two detection frames compared to the area of ​​the detection frame corresponding to the first recognition object. Afterwards, when the overlap is less than or equal to the target overlap threshold, the electronic device performs a target operation corresponding to the first recognition object. In the present application, the recognition accuracy of the target object can be improved, thereby improving the accuracy of the electronic device in performing the target operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to image processing technology, and in particular to a target object recognition method and electronic device. Background Art

[0002] With the continuous development of electronic devices, electronic devices are equipped with more and more functions (such as air gestures, scanning codes by turning the wrist, etc.) for users to use. Specifically, if the camera of an electronic device (such as a mobile phone) captures an image, the electronic device can recognize the captured image. If it is recognized that the image includes a target object (such as a user's hand or a QR code, etc.), the electronic device can perform corresponding operations based on the recognized target object.

[0003] However, in some scenarios, such as air gesture scenarios, if the image captured by the camera includes non-target objects (such as faces), in the process of the electronic device identifying the target object (such as the user's hand), since both the face and the user's hand contain multiple feature points, the electronic device may mistakenly identify the face as a gesture, causing the electronic device to perform incorrect operations, thereby affecting the user's experience. Summary of the invention

[0004] The embodiments of the present application provide a target object recognition method and an electronic device for improving the recognition accuracy of a target object.

[0005] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:

[0006] In a first aspect, a target object recognition method is provided, in which an electronic device acquires a target image frame, wherein the target image frame is an image frame currently collected by the electronic device. Afterwards, the electronic device performs target detection on the target image frame to obtain a detection result, wherein the detection result is used to indicate the recognition object contained in the target image frame. Afterwards, when the detection result indicates that the target image frame includes two recognition objects, the electronic device determines the overlap between the two recognition objects, wherein the two recognition objects include a first recognition object and a second recognition object, and the overlap refers to the overlap between the overlapping areas of the detection frames corresponding to the two recognition objects compared to the area of ​​the detection frame corresponding to the first recognition object. Afterwards, when the overlap is less than or equal to the target overlap threshold, the electronic device performs a target operation corresponding to the first recognition object.

[0007] In the present application, since the degree of overlap is determined based on the overlapping area between the two detection frames, that is, the larger the degree of overlap, the higher the probability that the second recognition object is mistaken for the first recognition object. Therefore, when the target image frame includes two recognition objects, the electronic device can determine whether the electronic device needs to perform the target operation based on the degree of overlap between the detection frames corresponding to the two recognition objects. In this way, the probability of false detection of the electronic device can be reduced, the filtering of background objects (the second recognition object) can be achieved, the recognition accuracy of the target object (the first recognition object) can be improved, the occurrence of electronic devices performing erroneous operations due to incorrect target object recognition can be reduced, the accuracy of the electronic device in performing target operations can be improved, and the user experience can be improved.

[0008] In a possible implementation manner of the first aspect, the method further includes: when the overlap degree is greater than a target overlap degree threshold, the electronic device does not perform a target operation corresponding to the first recognition object.

[0009] In the present application, if the overlap is greater than the target overlap threshold, it means that the overlap between the two detection frames is greater, that is, the probability that the second identified object is misidentified as the first identified object is higher. Therefore, the electronic device does not perform the target operation corresponding to the first identified object. In this way, it can avoid incorrect identification of the target object, improve the execution accuracy of the electronic device, and thereby improve the user experience.

[0010] In a possible implementation of the first aspect, when the first recognition object is a hand and the second recognition object is a face, the detection process of the target image frame may specifically include: the electronic device performs hand detection on the target image frame based on a hand detection model to obtain a first detection result, wherein the first detection result is used to indicate whether the target image frame includes a hand. Afterwards, the electronic device may perform face detection on the target image frame based on a face detection model to obtain a second detection result, wherein the second detection result is used to indicate whether the target image frame includes a face.

[0011] In the present application, two detection models (hand detection model and face detection model) are used to perform target detection on the target image frame. In this way, background objects such as faces can be filtered out, the probability of faces being recognized as hands can be reduced, and the occurrence of electronic devices performing erroneous operations due to incorrect target object recognition can be reduced, thereby improving the execution accuracy of electronic devices and thereby improving the user experience.

[0012] In a possible implementation of the first aspect, when the first detection result indicates that the target image frame includes a hand, if the confidence of the hand detection frame is greater than the second confidence threshold, and the confidence of the hand detection frame is less than or equal to the first confidence threshold, the electronic device may perform face detection on the target image frame based on the face detection model to obtain a second detection result. Thereafter, when the second detection result indicates that the target image frame includes a face, if the confidence of the face detection frame is greater than the third confidence threshold, the electronic device determines the overlap between the face detection frame and the hand detection frame.

[0013] In the present application, the electronic device performs face detection on the target image frame only when it is determined that the confidence of the hand detection frame is between the first confidence threshold and the second confidence threshold. In this way, the waste of resources caused by unnecessary face detection can be reduced, and unnecessary power consumption loss can be reduced.

[0014] In a possible implementation of the first aspect, when the detection results include a face detection frame and a hand detection frame, the electronic device determines the degree of overlap between the face detection frame and the hand detection frame, and when the degree of overlap is less than or equal to a target overlap threshold, the electronic device performs a gesture recognition operation corresponding to the hand.

[0015] In the present application, for the air gesture scenario, if the overlap between the face detection frame and the hand detection frame is less than or equal to the target overlap threshold, it means that the overlap between the hand detection frame and the face detection frame is smaller, that is, the probability of the face being misjudged as a hand is lower. Therefore, the electronic device can perform gesture recognition operations. In this way, the probability of the electronic device mistakenly triggering the gesture recognition operation can be greatly reduced without affecting the recall rate and accuracy of the target object, thereby improving the execution accuracy of the electronic device.

[0016] In a possible implementation manner of the first aspect, the method further includes: when the overlap between the face detection frame and the hand detection frame is greater than a target overlap threshold, the electronic device does not perform the gesture recognition operation.

[0017] In the present application, if the overlap between the face detection frame and the hand detection frame is greater than the target overlap threshold, it means that the overlap between the hand detection frame and the face detection frame is greater, that is, the probability that the face is misjudged as a hand is higher. Therefore, in order to avoid the face being misjudged as a hand, the electronic device may directly not perform the gesture recognition operation. In this way, the occurrence of the electronic device performing incorrect operations due to the face being misjudged as a hand can be reduced, thereby improving the execution accuracy of the electronic device.

[0018] In a possible implementation manner of the first aspect, the target overlap threshold is an overlap threshold corresponding to a confidence threshold interval to which the confidence of the hand detection frame belongs.

[0019] In the present application, the target overlap threshold is determined based on the confidence threshold interval to which the confidence of the hand detection frame belongs. That is to say, different confidence levels of hand detection frames may result in different target overlap thresholds. In this way, accurate filtering of faces can be achieved and the execution accuracy of electronic devices can be improved.

[0020] In a possible implementation of the first aspect, when the detection result also includes the confidence of the face detection frame and the confidence of the hand detection frame, the above process of determining the overlap may specifically include: the electronic device determines the relationship between the confidence of the hand detection frame and the first confidence threshold and the second confidence threshold, wherein the first confidence threshold is greater than the second confidence threshold. If the confidence of the hand detection frame is greater than the second confidence threshold, but not greater than the first confidence threshold, the electronic device may determine whether the confidence of the face detection frame is greater than the third confidence threshold. If the confidence of the face detection frame is greater than the third confidence threshold, the electronic device may calculate the overlap between the face detection frame and the hand detection frame based on the overlap algorithm, wherein the overlap is the ratio between the overlapping area between the face detection frame and the hand detection frame and the area of ​​the hand detection frame.

[0021] In the present application, if the confidence of the hand detection frame is greater than the second confidence threshold and not greater than the first confidence threshold, it means that although the accuracy of the above-mentioned first detection result is not low, there is still a situation where the face is misjudged as a hand. Therefore, in order to reduce the probability of the face being misjudged as a hand, the electronic device can calculate the overlap between the face detection frame and the hand detection frame. However, in order to avoid the situation where the detection result is inaccurate due to the low confidence of the face detection frame, the electronic device can calculate the overlap after determining that the confidence of the face detection frame is greater than the third confidence threshold. In this way, not only the waste of computing resources can be reduced, but also the execution efficiency of the electronic device can be improved.

[0022] In a possible implementation manner of the first aspect, the method further includes: when the confidence of the hand detection frame is greater than a first confidence threshold, the electronic device performs a gesture recognition operation.

[0023] In the present application, if the confidence of the hand detection frame is greater than the first confidence threshold, it means that the accuracy of the above-mentioned first detection result is relatively high, that is, the probability of the face being misjudged as a hand is relatively low. Therefore, the electronic device can directly perform gesture recognition operations. In this way, not only can unnecessary waste of resources be reduced, but also the execution efficiency of the electronic device can be improved.

[0024] In a possible implementation manner of the first aspect, the method further includes: when the confidence of the hand detection frame is less than or equal to a second confidence threshold, the electronic device does not perform the gesture recognition operation.

[0025] In the present application, if the confidence of the hand detection frame is less than or equal to the second confidence threshold, it means that the confidence level of the above-mentioned first detection result is low. Therefore, the electronic device may not perform the gesture recognition operation, thereby reducing unnecessary waste of resources.

[0026] In a possible implementation manner of the first aspect, the method further includes: when the confidence of the face detection frame is less than or equal to a third confidence threshold, the electronic device performs a gesture recognition operation.

[0027] In the present application, if the confidence of the face detection frame is less than or equal to the third confidence threshold, it means that the confidence level of the above-mentioned second detection result is low. Therefore, in order to reduce unnecessary power consumption loss, the electronic device can ignore the second detection result, that is, determine that there are only hands in the target image frame, and there is no face, that is, the electronic device can directly perform gesture recognition operations.

[0028] In a possible implementation of the first aspect, before the electronic device performs the gesture recognition operation, the method further includes: the electronic device obtains a plurality of hand feature points and a plurality of facial feature points in the target image frame, wherein the facial feature points include two eye feature points. Afterwards, the electronic device may determine whether the distance between the plurality of hand feature points and the first target height is less than a first preset distance, wherein the first target height is the height corresponding to the line between the two eye feature points. If the distances between the plurality of hand feature points and the first target height are all greater than or equal to the first preset distance, the electronic device may perform the gesture recognition operation.

[0029] In the present application, if the distances between multiple hand feature points and the first target height are not less than the first preset distance, it means that the user's hand is not touching the glasses. Therefore, the electronic device can directly perform gesture recognition operations. In this way, the scene of touching the glasses can be filtered, the probability of the electronic device misrecognizing gestures can be reduced, and the execution accuracy of the electronic device in performing air gesture services can be improved.

[0030] In a possible implementation manner of the first aspect, the method further includes: if there is a feature point among the multiple hand feature points whose distance to the first target height is less than a first preset distance, the electronic device may not perform the gesture recognition operation.

[0031] In the present application, if there is a feature point among multiple hand feature points whose distance from the first target height is less than the first preset distance, it means that the above-mentioned target user may be in the process of touching the glasses with his hand, and the target image frame in the scene is captured by the electronic device. Therefore, the electronic device may not perform the gesture recognition operation. In this way, the incorrect execution of the gesture recognition operation can be avoided, and the execution accuracy of the electronic device in performing the air gesture service can be improved.

[0032] In a possible implementation of the first aspect, before the electronic device performs the gesture recognition operation, the method further includes: the electronic device acquires a plurality of hand feature points and a plurality of facial feature points in the target image frame, wherein the facial feature points include forehead feature points. Afterwards, the electronic device may determine whether the distance between the plurality of hand feature points and the second target height is less than a second preset distance, wherein the second target height is the height corresponding to the horizontal line of the forehead feature points. If the distance between the plurality of hand feature points and the second target height is greater than or equal to the second preset distance, the electronic device may perform the gesture recognition operation.

[0033] In the present application, if the distances between multiple hand feature points and the second target height are not less than the second preset distance, it means that the user's hand is not touching the hair. Therefore, the electronic device can directly perform the gesture recognition operation. In this way, the hand touching hair scene can be filtered, the probability of the electronic device misrecognizing gestures can be reduced, and the execution accuracy of the electronic device in performing air gesture services can be improved.

[0034] In a possible implementation manner of the first aspect, the method further includes: if there is a feature point among the multiple hand feature points whose distance to the second target height is less than a second preset distance, the electronic device may not perform the gesture recognition operation.

[0035] In the present application, if there is a feature point among multiple hand feature points whose distance from the second target height is less than the second preset distance, it means that the above-mentioned target user may be in the process of touching his hair, and the target image frame in this scene is captured by the electronic device. Therefore, the electronic device may not perform the gesture recognition operation. In this way, the incorrect execution of the gesture recognition operation can be avoided, and the execution accuracy of the electronic device in performing the air gesture service can be improved.

[0036] In a possible implementation of the first aspect, when the first recognition object is a face and the second recognition object is a hand, and the detection result includes a face detection frame and a hand detection frame, the electronic device determines the overlap between the face detection frame and the hand detection frame, and when the overlap is less than or equal to the target overlap threshold, the electronic device performs a facial recognition operation corresponding to the face.

[0037] In the present application, for face recognition scenarios (such as smart gaze scenarios), if the overlap between the face detection frame and the hand detection frame is less than or equal to the target overlap threshold, it means that the overlap between the face detection frame and the hand detection frame is smaller, that is, the probability of the hand being misjudged as a face is lower. Therefore, the electronic device can perform face recognition operations. In this way, the probability of the electronic device falsely triggering the face recognition operation can be greatly reduced without affecting the recall rate and accuracy of the target object, thereby improving the execution accuracy of the electronic device.

[0038] In a possible implementation of the first aspect, when the first recognition object is a QR code and the second recognition object is a background object, and the detection result includes a QR code detection frame and a background object detection frame, the electronic device determines the degree of overlap between the QR code detection frame and the background object detection frame, and when the degree of overlap is less than or equal to the target overlap threshold, the electronic device performs a QR code recognition operation corresponding to the QR code.

[0039] In the present application, for the QR code scanning scenario (such as the wrist scanning scenario), if the overlap between the QR code detection frame and the background object detection frame is less than or equal to the target overlap threshold, it means that the smaller the overlap between the QR code detection frame and the background object detection frame, the lower the probability of the background object being misjudged as a QR code. Therefore, the electronic device can perform a QR code recognition operation. In this way, the probability of the electronic device erroneously triggering the QR code recognition operation can be greatly reduced without affecting the recall rate and accuracy of the target object, thereby improving the execution accuracy of the electronic device.

[0040] In a possible implementation of the first aspect, the above method also includes: when the detection result indicates that the target image frame includes an identification object, and the identification object is a first identification object, the electronic device determines whether the confidence of the detection box corresponding to the first identification object is greater than a fourth confidence threshold; if the confidence of the detection box corresponding to the first identification object is greater than the fourth confidence threshold, the electronic device can perform the target operation corresponding to the first identification object.

[0041] In the present application, if the target image frame only includes a first identification object, and the confidence of the detection box corresponding to the first identification object is greater than the fourth confidence threshold, it means that the detection result is relatively accurate. Therefore, the electronic device can directly execute the target operation corresponding to the first identification object. In this way, not only can unnecessary waste of resources be reduced, but also the recognition efficiency of the electronic device can be improved.

[0042] In a possible implementation manner of the first aspect, the method further includes: if the confidence of the detection box corresponding to the first recognized object is less than or equal to a fourth confidence threshold, the electronic device does not perform the target operation corresponding to the first recognized object.

[0043] In the present application, if the target image frame only includes a first recognition object, and the confidence of the detection box corresponding to the first recognition object is less than or equal to the fourth confidence threshold, it means that the accuracy of the detection result is low. Therefore, the electronic device does not perform the target operation corresponding to the first recognition object. In this way, unnecessary waste of resources caused by the electronic device continuing to identify can be avoided, thereby improving the utilization of computing resources.

[0044] In a possible implementation manner of the first aspect, the method further includes: when the detection result indicates that the target image frame includes an identification object, and the identification object is the second identification object, the electronic device does not perform the target operation corresponding to the first identification object.

[0045] In the present application, if the target image frame only includes the second identification object, that is, the target image frame does not include the first identification object, the electronic device may not perform the gesture recognition operation. In this way, the occurrence of erroneous operations performed by the electronic device can be reduced, and the accuracy of the electronic device in performing the target operation can be improved, thereby improving the user experience.

[0046] In a possible implementation of the first aspect, the method further includes: when the detection result indicates that the target image frame includes at least three recognition objects, and the at least three recognition objects include at least one first recognition object, the electronic device may determine, for each first recognition object, a degree of overlap between the first recognition object and other recognition objects other than the first recognition object. Thereafter, the electronic device may determine, based on the multiple degrees of overlap, whether the electronic device can perform a target operation corresponding to the first recognition object.

[0047] In a possible implementation of the first aspect, the above process of determining whether the electronic device performs the target operation may specifically include: the electronic device determines whether the multiple overlap degrees are greater than the target overlap degree threshold; if the multiple overlap degrees are all greater than the target overlap degree threshold, the electronic device does not perform the target operation corresponding to the first identified object.

[0048] In the present application, if multiple overlaps are greater than the target overlap threshold, it means that the overlap between different identification objects is large. Therefore, the electronic device may not perform the target operation. In this way, the occurrence of erroneous operations performed by the electronic device can be reduced, and the accuracy of the electronic device in performing the target operation can be improved, thereby improving the user experience.

[0049] In a possible implementation of the first aspect, the method further includes: if there is an overlap degree less than or equal to the target overlap degree threshold among the multiple overlap degrees, the electronic device may take the first identified object whose overlap degree is less than or equal to the target overlap degree threshold as the target object, and perform the target operation corresponding to the target object.

[0050] In the present application, if there is an overlap degree less than or equal to the target overlap degree threshold among multiple overlap degrees, it means that there are two identification objects with a small degree of overlap. Therefore, the electronic device can perform the target operation. In this way, the accuracy of the electronic device in performing the target operation can be improved, thereby improving the user experience.

[0051] In a second aspect, the present application provides an electronic device, comprising a camera, a memory and one or more processors; the camera, the memory and the processor are coupled; the camera is used to capture images, the memory is used to store computer program codes, and the computer program codes include computer instructions; when the processor executes the computer instructions, the electronic device executes the method described above.

[0052] In a third aspect, the present application provides a computer-readable storage medium, comprising computer instructions, which, when executed on an electronic device, enables the electronic device to execute the method described above.

[0053] In a fourth aspect, the present application provides a computer program product, which, when executed on an electronic device, enables the electronic device to execute the method described above.

[0054] In a fifth aspect, a chip is provided, comprising: an input interface, an output interface, a processor and a memory, wherein the input interface, the output interface, the processor and the memory are connected via an internal connection path, and the processor is used to execute the code in the memory. When the code is executed, the processor is used to execute the method as described above.

[0055] Among them, the beneficial effects that can be achieved by the electronic device described in the second aspect, the computer-readable storage medium described in the third aspect, the computer program product described in the fourth aspect, and the chip described in the fifth aspect provided above can refer to the beneficial effects in the first aspect and any possible design method thereof, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 A schematic diagram of a scene photographed by an electronic device provided in an embodiment of the present application;

[0057] Figure 2 A schematic diagram of a scenario of air sliding video content provided in an embodiment of the present application;

[0058] Figure 3 A schematic diagram of a scene for zooming in on an image in the air provided in an embodiment of the present application;

[0059] Figure 4 A schematic diagram of a scene of sliding an image in the air provided in an embodiment of the present application;

[0060] Figure 5 A schematic diagram of an interface for scanning a QR code with a mobile phone provided in an embodiment of the present application;

[0061] Figure 6 A schematic diagram of a wrist-flipping code scanning scenario provided in an embodiment of the present application;

[0062] Figure 7 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;

[0063] Figure 8 A schematic diagram showing a front camera and a camera of a mobile phone provided in an embodiment of the present application;

[0064] Fig. 9 A schematic diagram of the software structure of an electronic device provided in an embodiment of the present application;

[0065] Fig.10 A flowchart of a target object recognition method provided in an embodiment of the present application;

[0066] Fig.11 A schematic diagram of the function settings of a method for air gesture recognition provided by an embodiment of the present application;

[0067] Fig.12 A schematic diagram of a mobile phone performing hand detection provided in an embodiment of the present application;

[0068] Fig.13 A schematic diagram of a mobile phone performing face detection provided in an embodiment of the present application;

[0069] Fig.14 A schematic diagram of an interface for a remote screenshot scenario provided in an embodiment of the present application;

[0070] Fig.15 A schematic diagram of extracting a face detection frame and a hand detection frame provided in an embodiment of the present application;

[0071] Fig.16 A flowchart of another target object recognition method provided in an embodiment of the present application;

[0072] Fig.17 A schematic diagram showing hand feature points provided in an embodiment of the present application;

[0073] Fig.18 A schematic diagram showing facial feature points provided in an embodiment of the present application;

[0074] Fig.19 Another schematic diagram showing facial feature points provided in an embodiment of the present application;

[0075] Fig. 20 A schematic diagram of a mobile phone performing hand feature point determination in a hand-touching glasses scenario provided in an embodiment of the present application;

[0076] Fig.21 A schematic diagram of a mobile phone performing hand feature point judgment in a hand touching hair scenario provided in an embodiment of the present application. DETAILED DESCRIPTION

[0077] The technical scheme in the embodiment of the present application will be described below in conjunction with the accompanying drawings in the embodiment of the present application. Wherein, in the description of the present application, unless otherwise specified, the "and / or" in the present application is only a kind of association relationship describing the associated object, indicating that there can be three kinds of relationships, for example, A and / or B, which can be represented by: A exists alone, A and B exist at the same time, and B exists alone, wherein A, B can be singular or plural. And, in the description of the present application, unless otherwise specified, "multiple" refers to two or more than two. "At least one of the following (individuals)" or its similar expressions refers to any combination of these items, including any combination of single items (individuals) or plural items (individuals). For example, at least one of a, b, or c (individuals) can be represented by: a, b, c, ab, ac, bc, or abc, wherein a, b, c can be single or multiple. In addition, in order to facilitate the clear description of the technical scheme of the embodiment of the present application, in the embodiment of the present application, the words "first", "second" and the like are used to distinguish the same items or similar items with substantially the same functions and effects. Those skilled in the art will appreciate that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit the difference. At the same time, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way for easy understanding.

[0078] In some embodiments, after the camera of an electronic device (such as a mobile phone) captures an image, the electronic device can identify the target object of the image. If the electronic device recognizes that the image includes a target object (such as a user's hand or a QR code, etc.), the electronic device can perform a corresponding operation according to the content displayed by the target object. However, if the electronic device incorrectly identifies the target object, it may mistake a non-target object (or background object) in the image for the target object, which may cause the electronic device to perform an incorrect operation, thereby affecting the user's experience.

[0079] In one implementation, taking the air gesture scenario as an example, when a user is using an electronic device and is unable to touch the screen for some reason (such as wearing gloves or having dirty hands, etc.), the user can control the electronic device to perform corresponding operations through air gestures. Specifically, if the camera captures the target image frame, the mobile phone can perform hand recognition on the target image frame to obtain a recognition result, wherein the target image frame is the image corresponding to when the user triggers the air gesture service. Afterwards, when the recognition result indicates that the target image frame includes the user's hand, the mobile phone can further perform gesture recognition on the user's hand to perform corresponding operations according to the gesture content.

[0080] However, in some cases, the electronic device may mistake some background objects (such as the user's face) for the user's hands, that is, the electronic device mistakenly believes that the user wants to trigger the air gesture service. Therefore, if the electronic device continues to perform gesture recognition based on the background object, it will cause the electronic device to perform incorrect operations. For example, see Figure 1 The image 10 captured by the electronic device is an image taken when the user is eating. It can be seen that the image 10 does not include the user's hand, but only includes the face of user A. The detection result obtained after the electronic device performs target detection on the image 10 includes a hand detection frame 11. That is to say, the electronic device mistakenly recognizes the face of user A as the user's hand, which will cause the electronic device to continue to perform gesture recognition according to the feature points of the face, thereby causing the electronic device to perform incorrect operations, ultimately affecting the user's experience.

[0081] It should be noted that the above-mentioned air gesture service means that when the user does not touch the electronic device, that is, when there is a short distance between the user and the electronic device (for example, the distance between the user's hand and the electronic device is 20 cm), the electronic device can combine the gesture made by the user's hand to perform the operation corresponding to the gesture content. Among them, the air gesture can be a dynamic air gesture or a static air gesture, and there is no specific limitation. Exemplarily, the dynamic air gesture may include the index finger moving in the longitudinal or transverse direction, the index finger and thumb pinching and moving in the longitudinal or transverse direction, etc. The static air gesture may include a V gesture, an OK gesture, a pinch gesture, and a fist gesture, etc.

[0082] In one example, the electronic device is a mobile phone, and the interface currently displayed on the mobile phone is a short video playback interface. If the user cannot touch the screen for some reason (such as wearing gloves or dirty hands) while watching the short video, the user can control the electronic device to perform the corresponding operation by moving his fingers in the air. For example, see Figure 2 If the user's index finger makes an upward gesture through the air, mobile phone E can display the corresponding content of the next video; or, if the user's index finger makes a downward gesture through the air, the mobile phone can display the corresponding content of the previous video.

[0083] In another example, the electronic device is a mobile phone, and the interface currently displayed on the mobile phone is a photo browsing interface. If the user cannot touch the screen during the photo browsing process due to some reasons (such as wearing gloves or dirty hands, etc.), the user can control the electronic device to perform corresponding operations by moving fingers in the air. For example, see Figure 3 If the user's index finger and thumb make a zoom-in gesture in the vertical direction, that is, the index finger moves upward while the thumb moves downward, the mobile phone F can zoom in on the current photo content; or, if the user's index finger and thumb make a zoom-out gesture in the vertical direction, that is, the index finger moves downward while the thumb moves upward, the mobile phone can zoom out on the current photo content. For another example, see Figure 4 If the user's index finger makes a gesture of moving to the left in the air, the mobile phone G can display the next photo; or, if the user's index finger makes a gesture of moving to the right in the air, the mobile phone can display the previous photo.

[0084] In another implementation, taking the example of scanning a QR code by turning the wrist, when a user encounters a situation where a QR code needs to be scanned (such as scanning a code to order food) while using an electronic device, the user can turn the wrist to face the screen of the electronic device to the QR code information (such as a QR code icon). After that, the electronic device can display a code scanning interface to automatically scan the QR code information. After that, when the QR code information is scanned, the electronic device can display relevant information for the user to browse.

[0085] However, in some cases, for example, if the user only performs a small wrist-turning operation (such as turning the wrist 10 degrees), the electronic device may mistakenly believe that the user has triggered the wrist-turning scanning service to further identify whether there is QR code information within the current shooting range. However, in the process of identifying QR code information, the electronic device may mistake some background objects (such as ceilings, wall tiles, etc.) as QR code information. If the electronic device continues to perform QR code recognition based on the background objects, it will cause the electronic device to display incorrect information. For example, Figure 5 As shown, the interface currently displayed on the mobile phone is a QR code scanning interface. It can be seen that the content scanned by the mobile phone is a ceiling 50 in an indoor place, and the ceiling 50 includes a plurality of wooden boards, that is, the QR code scanning interface does not include QR code information, but only includes the ceiling 50. That is, if the mobile phone mistakenly recognizes the ceiling 50 as a QR code, it will cause the mobile phone to continue to perform QR code recognition according to the feature points of the ceiling 50, which will cause the mobile phone to display wrong information, thereby affecting the user experience.

[0086] It should be noted that the wrist-flipping scanning service means that when the user performs a wrist-flipping operation, the electronic device can display the content corresponding to the QR code information in combination with the QR code information within the current shooting range. The wrist-flipping operation refers to the operation of switching the screen of the electronic device from facing the user to facing away from the user. For example, Figure 6 As shown, when the screen of the mobile phone is displaying the content of the e-book, if the user holds the mobile phone and turns the wrist, the screen of the mobile phone will face the QR code information of the ordering dish, and the mobile phone will display the code scanning interface to scan the QR code information. Afterwards, if the mobile phone scans successfully, the mobile phone can display the menu content corresponding to the QR code information for the user to order the dish.

[0087] Therefore, in order to avoid misidentification of the target object, an embodiment of the present application provides a target object recognition method. In this method, an electronic device acquires a target image frame. Afterwards, the electronic device performs target detection on the target image frame to obtain a detection result, wherein the detection result is used to indicate the recognition object included in the target image frame. Afterwards, when the detection result indicates that the target image frame includes two recognition objects, the electronic device determines the overlap between the detection frames corresponding to the two recognition objects, wherein the two recognition objects include a first recognition object (or referred to as a target object) and a second recognition object (or referred to as a background object), and the overlap refers to the overlap between the overlapping areas of the two detection frames compared to the area of ​​the detection frame corresponding to the first recognition object. Afterwards, when the overlap between the two detection frames is not greater than the target overlap threshold, the electronic device performs a target operation corresponding to the first recognition object.

[0088] In the embodiment of the present application, since the degree of overlap is determined based on the overlapping area between the two detection frames, that is, the larger the degree of overlap, the higher the probability that the second identified object is mistaken for the first identified object. Therefore, when the target image frame includes two identified objects, the electronic device can determine whether the electronic device needs to perform the target operation based on the degree of overlap between the detection frames corresponding to the two identified objects. In this way, the probability of false detection of the electronic device can be reduced, background objects can be filtered, the recognition accuracy of the target object can be improved, the occurrence of electronic devices performing erroneous operations due to incorrect target object recognition can be reduced, the accuracy of the electronic device in performing target operations can be improved, and the user experience can be improved.

[0089] In some embodiments, if the current scene is an air gesture scene, and the target object is a hand, and the background object is a face, then after the electronic device performs target detection on the target image frame, the detection result obtained is used to indicate whether the target image frame includes a face and / or a hand, that is, whether the detection result includes a face detection frame and / or a hand detection frame. Afterwards, when the detection result indicates that the target image frame includes a face and a hand, the electronic device calculates the overlap between the face detection frame and the hand detection frame, wherein the overlap is the ratio of the overlapping area between the face detection frame and the hand detection frame to the area of ​​the hand detection frame. Afterwards, when the overlap is not greater than the target overlap threshold, the electronic device performs a gesture recognition operation.

[0090] In some embodiments, if the current scene is an intelligent gaze scene, and the target object is a face and the background object is a hand, then after the electronic device performs target detection on the target image frame, the detection result obtained is used to indicate whether the target image frame includes a face and / or a hand, that is, whether the detection result includes a face detection frame and / or a hand detection frame. Afterwards, when the detection result indicates that the target image frame includes a face and a hand, the electronic device calculates the overlap between the face detection frame and the hand detection frame, wherein the overlap is the ratio of the overlapping area between the face detection frame and the hand detection frame to the area of ​​the hand detection frame. Afterwards, when the overlap is not greater than the target overlap threshold, the electronic device performs a face recognition operation.

[0091] It should be noted that the above-mentioned intelligent watching scenarios include watching without turning off the screen and watching to lower the volume. The specific scene introduction will be described in detail later.

[0092] In other embodiments, if the current scene is a QR code scanning scene (such as a scene of scanning by turning the wrist or a scene of triggering scanning by clicking a scanning control), and the above-mentioned target object is a QR code, and the background object is a background object (such as a ceiling), then after the electronic device performs target detection on the above-mentioned target image frame, the detection result obtained is used to indicate whether the target image frame includes a QR code and / or a ceiling, that is, whether the detection result includes a QR code detection frame and / or a ceiling detection frame. Afterwards, when the detection result indicates that the target image frame includes a QR code and a ceiling, the electronic device calculates the overlap between the QR code detection frame and the ceiling detection frame, wherein the overlap is the ratio of the overlapping area between the QR code detection frame and the ceiling detection frame to the area of ​​the QR code detection frame. Afterwards, when the overlap is not greater than the target overlap threshold, the electronic device performs a QR code recognition operation.

[0093] It should be noted that the electronic device in the embodiments of the present application may be a mobile phone, a tablet computer, a smart watch, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, as well as a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) device, a virtual reality (VR) device, and other devices containing a camera. The embodiments of the present application do not impose any special restrictions on the specific form of the electronic device.

[0094] For example, Figure 7 FIG. 2 shows a schematic diagram of the structure of the electronic device 200. Figure 7As shown, the electronic device 200 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 211, a power management module 212, a battery 213, an antenna 1, an antenna 2, a mobile communication module 240, a wireless communication module 250, an audio module 270, a sensor module 280, a button 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.

[0095] It is to be understood that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0096] The processor 210 may include one or more processing units, for example, the processor 210 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0097] The controller may be the nerve center and command center of the electronic device 200. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0098] The processor 210 may also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. The memory may store instructions or data that the processor 210 has just used or cyclically used. If the processor 210 needs to use the instruction or data again, it may be directly called from the memory. This avoids repeated access, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.

[0099] In some embodiments, the processor 210 may include one or more interfaces. The interface may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0100] It is understandable that the interface connection relationship between the modules illustrated in the embodiment of the present invention is only a schematic illustration and does not constitute a structural limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0101] The charging management module 211 is used to receive charging input from a charger. While the charging management module 211 is charging the battery 213 , it can also power the electronic device through the power management module 212 .

[0102] The wireless communication function of the electronic device 200 can be implemented through the antenna 1, the antenna 2, the mobile communication module 240, the wireless communication module 250, the modem processor and the baseband processor.

[0103] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 200 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve the utilization of the antennas. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0104] The mobile communication module 240 may provide solutions for wireless communications including 2G / 3G / 4G / 5G etc. applied to the electronic device 200. The modem processor may include a modulator and a demodulator.

[0105] The wireless communication module 250 can provide wireless communication solutions for application in the electronic device 200, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication technology (NFC), infrared technology (IR), etc.

[0106] The electronic device 200 implements the display function through a GPU, a display screen 294, and an application processor. The GPU is a microprocessor for image processing, which connects the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or change display information.

[0107] The display screen (or screen) 294 is used to display images, videos, etc. The display screen 294 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Miniled, MicroLed, Micro-oLed, a quantum dot light emitting diode (QLED), etc. In some embodiments, the electronic device 200 may include 1 or N display screens 294, where N is a positive integer greater than 1.

[0108] The electronic device 200 can realize the shooting function through ISP, camera 293, video codec, GPU, display screen 294 and application processor.

[0109] ISP is used to process the data fed back by camera 293. For example, when an electronic device takes a photo, the shutter is opened, and light is transmitted to the camera photosensitive element (or image sensor) through the lens. The light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to ISP for processing and converts it into an image visible to the naked eye. ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. ISP can also optimize the exposure, color temperature and other parameters of the shooting scene. In some embodiments, ISP can be set in camera 293. In some embodiments, camera 293 includes a shutter. The shutter is a device in the camera used to control the time that light irradiates the photosensitive element.

[0110] The camera 293 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then passes the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 200 may include 1 or N cameras 293, where N is a positive integer greater than 1.

[0111] In some embodiments, the camera 293 may include a lens, which is an optical component for generating an image.

[0112] Exemplarily, the N cameras 293 may include: one or more front cameras and one or more rear cameras. Figure 8 , taking the above-mentioned electronic device 200 as a mobile phone as an example. Figure 8 The (a) interface shows a front camera, such as front camera 20. Figure 8 The interface (b) in FIG. 1 shows three rear cameras, such as rear cameras 21, 22 and 23. Of course, the number of cameras in the mobile phone includes but is not limited to the number described in the above embodiment.

[0113] Among them, the above-mentioned N cameras 293 may include one or more of the following cameras: a main camera, a telephoto camera, a wide-angle camera, an ultra-wide-angle camera, a macro camera, a fisheye camera, an infrared camera, a depth camera and a black and white camera.

[0114] In this embodiment, the front camera is a camera with an always-on camera (AON) function. Specifically, when the electronic device uses the AON function, the front camera of the electronic device is in an always-on state, and can collect images in real time. The electronic device performs gesture recognition through image analysis, and can respond to user gestures to control the screen, so that the screen of the electronic device can be controlled without the user touching the electronic device.

[0115] The digital signal processor is used to process digital signals, and can process not only digital image signals but also other digital signals. For example, when the electronic device 200 is selecting a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.

[0116] The external memory interface 220 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 200.

[0117] The internal memory 221 can be used to store computer executable program codes, which include instructions. The processor 210 executes various functional applications and data processing of the electronic device 200 by running the instructions stored in the internal memory 221. The internal memory 221 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the electronic device 200 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 221 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0118] The electronic device 200 can implement audio functions through the audio module 270 and the application processor, such as music playing and recording, etc. The audio module 270 may include a speaker, a receiver, a microphone, and an earphone interface, etc.

[0119] The buttons 290 include a power button, a volume button, etc. The indicator 292 may be an indicator light.

[0120] The sensor module 280 may include a pressure sensor, a gyro sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, and the like.

[0121] The gyro sensor may be used to determine the motion posture of the electronic device 200. In some embodiments, the angular velocity of the electronic device 200 around three axes (ie, x, y, and z axes) may be determined by the gyro sensor.

[0122] The acceleration sensor can detect the magnitude of the acceleration of the electronic device 200 in all directions (generally three axes). When the electronic device 200 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers.

[0123] The software system of the electronic device 200 may adopt a layered architecture, an event-driven architecture, a micro-core architecture, a micro-service architecture, or a cloud architecture. The present application embodiment takes the Android system of the layered architecture as an example to exemplify the software structure of the electronic device 200.

[0124] Fig. 9 It is a software structure block diagram of the electronic device 200 of the embodiment of the present application. The embodiments of the present application will be discussed based on the following technical architecture. It should be noted that, in order to facilitate the explanation of logic, only the business logic relationship is illustrated by a schematic block diagram, and the specific location of the technical architecture where each business is located is not strictly expressed. In addition, the naming of each module in the software architecture diagram is an exemplary example. The embodiment of the present application does not limit the naming of each module in the software architecture diagram. In actual implementation, the specific naming of the module can be determined according to actual needs.

[0125] The layered architecture divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system includes four layers, from top to bottom: application layer (applications), application framework layer (application framework), hardware abstraction layer (HAL), and kernel layer (kernel).

[0126] The application layer may include a series of application packages. For example, the application layer may include applications such as camera, smart perception, etc. (applications may be referred to as applications for short), and the present application embodiment does not impose any restrictions on this.

[0127] The smart perception application provided in the embodiment of the present application supports various services. For example, the services supported by the smart perception application may include air gesture services, smart code scanning services, gaze-on-screen services, gaze-down volume services, smart horizontal and vertical screen services, auxiliary photo services, and smart always on display (AOD) services. These services may be collectively referred to as smart perception services.

[0128] It should be noted that the implementation of these services supported by the smart perception application depends on the AON camera (front camera) of the electronic device being in a normally open state, collecting images in real time, and obtaining data related to the smart perception service. When the smart perception application monitors the service-related data, the smart perception application sends the service-related data to the smart perception algorithm platform, which analyzes the collected images and determines whether to execute the relevant services of the smart perception application or which specific service to execute based on the analysis results.

[0129] Among them, the air gesture service refers to the service in which the electronic device recognizes user gestures and responds according to the preset strategy corresponding to the gesture to achieve human-computer interaction. It can be understood that the user gesture is an action made when the user's hand is within a preset distance of the electronic device. For example, the preset distance can be 20 cm. When the electronic device is in the screen-on state, the electronic device can support the recognition of various preset air gestures. Different gestures can correspond to different preset strategies, such as: air grab gesture → screenshot, air up / down gesture → page turning, air press gesture → answering a call.

[0130] For example, taking the air grabbing gesture as an example, the user changes from an extended palm state to a clenched fist state, and the control method corresponding to the gesture is preset to take a screenshot of the display interface of the electronic device. When the electronic device is in the screen-on state, the electronic device will automatically perform a screenshot operation when the air grabbing gesture is detected through the AON camera (front camera). Among them, the air grabbing gesture can also be called an air screenshot gesture, and the scene can be called an air screenshot scene.

[0131] For example, taking the air up / down gesture as an example, the fingers change from a close and spread state to a downward bent state (upward sliding gesture), and the control method corresponding to the gesture is preset to turn the page up or slide the screen upward, and the fingers change from a close and bent state to an upward spread state (downward sliding gesture), and the control method corresponding to the gesture is preset to turn the page down or slide the screen downward. When the electronic device is in the screen-on state, the electronic device will automatically perform the operation of turning the page up when the upward sliding gesture is detected through the AON camera (front camera), and will automatically perform the operation of turning the page down when the downward sliding gesture is detected, and the human-computer interaction can be completed without the user touching the electronic device. Among them, the air up / down gesture can also be called the air sliding screen gesture. Taking the upward or downward sliding gesture as an example for explanation, it can be understood that in actual implementation, the air sliding screen gesture can also be a left or right sliding gesture.

[0132] For another example, taking the air press gesture as an example, the user's palm approaches the screen of the electronic device from far to near, and the control mode corresponding to the gesture is preset to answer the call. When the electronic device receives the incoming call signal and displays the incoming call interface, the electronic device will automatically answer the call when it recognizes the air press gesture through the AON camera (front camera). This scenario can be called the air press scenario.

[0133] In addition, in an embodiment of the present application, the air gesture service also supports user-defined operations, for example, the user cancels the "air press gesture → answer the call" and resets it to "air press gesture → jump to the payment code interface".

[0134] Exemplarily, when the electronic device is in the screen-on state and the electronic device displays the main desktop, the user can trigger a predefined quick service through an air press gesture, such as quickly jumping from the main desktop to the payment code interface. In this scenario, the payment interface can be quickly called up through an air press gesture, so this scenario can be called a smart payment scenario. It should be noted that the user can also reset it to "air press gesture → jump to the scan interface" or "air press gesture → jump to the ride code interface" and so on.

[0135] It should be noted that the various services supported by the above-mentioned smart sensing application are illustrative examples. It can be understood that in actual implementation, the smart sensing application in the embodiment of the present application can also support other possible services, which can be determined based on actual usage requirements and are not limited by the embodiment of the present application.

[0136] The application framework layer provides application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions. Fig. 9 As shown, the application framework layer may include AON service, camera service, view system, content provider, etc.

[0137] The AON service is used to collect images in real time. The camera service is used to collect images. Specifically, when the camera application in the electronic device is turned on, the camera of the electronic device can perform corresponding operations according to the user's touch operation. For example, if the shooting control in the camera application is clicked by the user, the camera of the electronic device can perform a shooting operation to collect images.

[0138] The above view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a text notification icon can include a view for displaying text and a view for displaying pictures.

[0139] The content provider is used to store and retrieve data and make it accessible to applications. The data may include video, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.

[0140] The hardware abstraction layer is an encapsulation of the Linux kernel driver, providing an interface to the upper layer, which hides the hardware interface details of a specific platform and provides a virtual hardware platform for the operating system. In the embodiment of the present application, the hardware abstraction layer includes modules such as the camera HAL and the imaging science foundation (ISF) HAL.

[0141] The kernel layer is the layer between hardware and software. The kernel layer includes at least Qualcomm communication interface (QMI), air gesture module, smart code scanning module, display driver, camera driver and sensor driver.

[0142] The Qualcomm communication interface is a multi-processor inter-process communication functional interface provided by Qualcomm and is used to call functions, read data, etc.

[0143] The hardware layer provides various hardware devices. For example, the hardware devices involved in the embodiments of the present application include AON ISP, camera sensors, and physical sensors. Among them, the camera sensor is used to capture images in real time, the AON ISP is used to process image signals of the captured images, and the physical sensor may include an acceleration sensor and a gyroscope sensor. In the embodiments of the present application, the acceleration sensor is used to collect acceleration data of the terminal device during the user's wrist turning process; the gyroscope sensor is used to collect angular acceleration data of the terminal device during the user's wrist turning process. The gyroscope sensor and the acceleration sensor can be used together to detect the user's wrist turning action.

[0144] It can be understood that the above-mentioned AON ISP and camera sensor provide hardware support for the above-mentioned air gesture module, and the above-mentioned physical sensor provides hardware support for the above-mentioned smart code scanning module.

[0145] Understandably, Fig. 9 The layers in the structure shown and the components contained in each layer do not constitute a specific limitation on the electronic device 200, i.e., the folding screen device. In other embodiments of the present application, the structure may include more or fewer layers than shown in the figure, and each layer may include more or fewer components, which is not limited in the present application.

[0146] Based on the electronic device described above, an embodiment of the present application provides a method for identifying a target object. The method can be applied to any intelligent perception scenario, such as a gesture scene, a wrist scan scene, etc. The following takes the electronic device as a mobile phone and the application scenario as an air gesture scene as an example to illustrate the method of the embodiment of the present application. Specifically, Fig.10 As shown, the target object recognition method may include S1001 to S1010.

[0147] S1001, the mobile phone collects target image frames.

[0148] The target image frame is the image frame currently captured by the mobile phone, that is, the image frame within the field of view that can be captured by the front camera of the mobile phone. The front camera is a camera with AON function, that is, the front camera can capture the target image frame in real time.

[0149] It should be noted that the above-mentioned target image frame can be used to identify user gestures and / or faces so that the mobile phone can perform subsequent operations. In order to protect the user's privacy, it is necessary to determine whether the user has turned on the corresponding perception function of the smart perception application before the mobile phone collects the target image frame. When any perception function in the smart perception application is turned on, that is, when the mobile phone receives the user's turn-on operation for any perception function in the smart perception application, the mobile phone can collect the target image frame to determine the detection result of the target image frame.

[0150] For example, Fig.11 As shown, when the setting control in the initial interface is clicked by the user, the mobile phone can display the setting interface, wherein the setting interface includes setting items such as WLAN, Bluetooth, mobile network, desktop and wallpaper, smart assistant and smart perception. Afterwards, when the smart perception setting item in the setting interface is clicked by the user, the mobile phone can display the smart perception setting interface. Among them, the smart perception setting interface includes an air gesture setting bar and other setting bars. The air gesture setting bar includes air sliding screen, air screenshot, air pressing and smart payment. Other setting bars include smart watching, smart screen off display, smart horizontal and vertical screen and assisted photography. Smart watching includes watching the screen without turning off the screen and watching the screen to reduce the volume.

[0151] After that, if any sensing function in the air gesture setting bar in the smart sensing setting interface is clicked, the phone can control the clicked sensing function to be turned on or off. For example, see Fig.11 If the air sliding screen is clicked by the user, it means that the user wants to turn on the air sliding screen function, so the mobile phone can turn on the air sliding screen function. For another example, please refer to Fig.11If the air screenshot is clicked by the user, it means that the user wants to turn off the air sliding screen function. Therefore, the mobile phone can turn off the air screenshot function.

[0152] It should be noted that the above Fig.11 The interface for setting the perception function of the smart perception application is only an example, and the user can also control the perception function to be turned on or off in other ways. For example, the mobile phone can directly display the air gesture option so that the user can agree to control the turning on or off of the air gesture function. That is to say, if the air gesture option is clicked by the user, all functions included in the air gesture option are turned on or off.

[0153] S1002, the mobile phone performs target detection on the target image frame to obtain a detection result, wherein the detection result includes a face detection frame and / or a hand detection frame.

[0154] Specifically, after acquiring the target image frame, the mobile phone can perform hand detection on the target image frame to determine whether the target image frame contains a hand, and the mobile phone can perform face detection on the target image frame to determine whether the target image frame contains a face. It can be understood that when the above detection result includes a hand detection frame, it means that the target image frame includes a hand, or, when the above detection result includes a face detection frame, it means that the target image frame includes a face, or, when the above detection frame result includes a face detection frame and a hand detection frame, it means that the target image frame includes a face and a hand.

[0155] In some embodiments, the above detection results may also include the confidence of the detection frame. The confidence of the detection frame refers to a confidence score of the output result. For example, if the acquired target image frame may not be clear, the detection and output results of the mobile phone will be inaccurate, and the confidence score generated at this time will be lower. Exemplarily, if the target image frame contains a hand, the detection result may include a hand detection frame and the confidence corresponding to the hand detection frame; if the target image frame contains a face, the detection result may include a face detection frame and the confidence corresponding to the face detection frame; if the target image frame contains a hand and a face, the detection result may include a face detection frame, a hand detection frame, the confidence of the face detection frame, and the confidence of the hand detection frame.

[0156] In one implementation, the mobile phone can perform hand detection on the target image frame according to the hand detection model to obtain a first detection result. The first detection result is used to indicate whether the target image frame includes a hand. It can be understood that when the first detection result includes a hand detection frame and the confidence corresponding to the hand detection frame, it means that the target image frame includes a hand; when the first detection result does not include a hand detection frame, it means that the target image frame does not include a hand. The confidence corresponding to the hand detection frame is used to indicate the confidence level of the hand detection frame. The higher the confidence of the hand detection frame, the more accurate the first detection result, that is, the more precise the hand detection frame. Exemplarily, the mobile phone can output the confidence in the form of a numerical value, for example, the confidence is 0.88.

[0157] For example, Fig.12 As shown, after the front camera 32 of the mobile phone captures the image, the mobile phone can input the image into the trained hand detection model for hand detection to obtain a first detection result. It can be seen that the image includes the user's hand (such as a gesture of an open palm). Therefore, after the mobile phone inputs the image with the user's hand into the hand detection model, the first detection result obtained includes a hand detection frame T and a confidence level of 0.87 corresponding to the hand detection frame T. Among them, the hand detection frame is used to indicate the position information of the user's hand. For example, the mobile phone can mark the user's hand in the target image frame through the detection frame, and the position information of the user's hand is represented according to the coordinates corresponding to the four vertex corners of the detection frame.

[0158] It should be noted that the above-mentioned hand detection model can be a method of taking the first image sample as the input of the hand detection model to be trained, outputting the detection frame of the hand in the first image sample through learning and prediction of the hand detection model to be trained, and adjusting the parameters of the hand detection model to be trained based on the output hand detection frame and the real annotation frame of the hand in the first image sample until a trained hand detection model is obtained. The real annotation frame is used to indicate the annotation position of the hand in the first image sample, and the real annotation frame is a pre-annotated rectangular frame.

[0159] In another implementation, the mobile phone can perform face detection on the target image frame according to the face detection model to obtain a second detection result. The second detection result is used to indicate whether the target image frame includes a face. It can be understood that when the second detection result includes a face detection frame and the confidence level corresponding to the face detection frame, it means that the target image frame includes a face; when the first detection result does not include a face detection frame, it means that the target image frame does not include a face. The confidence level corresponding to the face detection frame is used to indicate the confidence level of the face detection frame. The higher the confidence level of the face detection frame, the more accurate the second detection result, that is, the more precise the face detection frame.

[0160] For example, Fig.13 As shown, after the front camera 42 of the mobile phone captures the image 43, the mobile phone can input the image 43 into the trained face detection model for face detection to obtain a second detection result. It can be seen that the image 43 includes a face. Therefore, after the mobile phone inputs the image 43 with the face into the face detection model, the second detection result obtained includes a face detection frame Y and a confidence level of 0.97 corresponding to the face detection frame Y. Among them, the face detection frame is used to indicate the location information of the user's face. For example, the mobile phone can mark the face in the target image frame through the face detection frame, and the coordinates corresponding to the four top corners of the face detection frame represent the location information of the face.

[0161] It should be noted that the above-mentioned face detection model can be to use the second image sample as the input of the face detection model to be trained, and output the detection frame of the face in the second image sample through learning and prediction of the face detection model to be trained, and adjust the parameters of the face detection model to be trained based on the output face detection frame and the real annotation frame of the face in the second image sample, until a trained face detection model is obtained. The second image sample can be the same as or different from the above-mentioned first image sample. If the second image sample is the same as the first image sample, the second image sample and the first image sample can contain both the face and the hand.

[0162] In some embodiments, the process of the mobile phone performing target detection on the target image frame may be to detect the target image frame simultaneously or in a preset order, without specific limitation. For example, the mobile phone may first perform hand detection on the target image frame, and then perform face detection on the target image frame. For another example, the mobile phone may first perform face detection on the target image frame, and then perform hand detection on the target image frame.

[0163] It can be understood that the detection result obtained by the mobile phone when performing target detection on the above-mentioned target image frame may include a first detection result and a second detection result, that is, the detection result includes at least one detection frame and the confidence of the detection frame, and the at least one detection frame includes a face detection frame and / or a hand detection frame.

[0164] In one implementation, if the above detection results include a hand detection frame, a face detection frame, the confidence of the hand detection frame, and the confidence of the face detection frame, the mobile phone can execute S1003 to further determine whether a gesture recognition operation needs to be performed. In this way, the accuracy of face detection can be improved, and the occurrence of incorrect operations performed by the mobile phone due to the face being misdetected as a hand can be reduced, thereby improving the execution accuracy of the mobile phone and thus improving the user experience.

[0165] In one implementation, if the above detection result only includes the hand detection frame and the confidence of the hand detection frame, the mobile phone can determine whether the confidence of the hand detection frame is greater than the fourth confidence threshold. If the confidence of the hand detection frame is greater than the fourth confidence threshold, it means that the detection result is relatively accurate. Therefore, the mobile phone can directly execute S1009 to reduce unnecessary waste of resources. If the confidence of the hand detection frame is not greater than the fourth confidence threshold, it means that the accuracy of the detection result is low. Therefore, the mobile phone can directly execute S1008 to reduce unnecessary waste of resources.

[0166] Among them, the fourth confidence threshold is a pre-set confidence threshold. In this embodiment, the fourth confidence threshold is 0.90. In other embodiments, the fourth confidence threshold can also be the same as the first confidence threshold, the second confidence threshold or the third confidence threshold, for example, the fourth confidence threshold can be 0.95, etc., which is not specifically limited.

[0167] It should be noted that the first confidence threshold, the second confidence threshold and the third confidence threshold will be described in detail later.

[0168] In another implementation, if the detection result only includes the face detection frame and the confidence of the face detection frame, the mobile phone can determine that the target image frame does not include the hand detection frame. Therefore, the mobile phone may not perform gesture recognition operations, but perform face-related recognition operations, for example, the mobile phone may perform face recognition operations. In this way, the occurrence of wrong operations performed by the mobile phone can be reduced, the execution accuracy of the mobile phone can be improved, and the user experience can be improved.

[0169] S1003: When the detection result includes a hand detection frame and a face detection frame, the mobile phone determines a relationship between the confidence of the hand detection frame and a first confidence threshold and a second confidence threshold.

[0170] In some embodiments, after determining that the above detection results include a hand detection frame and a face detection frame, the mobile phone can preferentially determine the relationship between the confidence of the hand detection frame and the first confidence threshold and the second confidence threshold. If the confidence of the hand detection frame is greater than the first confidence threshold, it means that the accuracy of the above first detection result is high, that is, the probability of the face being misjudged as a hand is low, so the mobile phone can execute S1004. If the confidence of the hand detection frame is greater than the second confidence threshold and not greater than the first confidence threshold, it means that although the accuracy of the above first detection result is not low, there is still a situation where the face is misjudged as a hand. Therefore, the mobile phone can execute S1005 to further determine whether the above target image frame contains a hand. If the confidence of the hand detection frame is not greater than the second confidence threshold, it means that the confidence level of the above first detection result is low, so the mobile phone can execute S1011.

[0171] Among them, the above-mentioned first confidence threshold is greater than the second confidence threshold. The first confidence threshold and the second confidence threshold are pre-set confidence thresholds. In this embodiment, the first confidence threshold is 0.95 and the second confidence threshold is 0.7. In other embodiments, the first confidence threshold can also be 0.97, 0.93, etc., and the second confidence threshold can also be 0.75, 0.67, etc., as long as the first confidence threshold is greater than the second confidence threshold, and there is no specific limitation.

[0172] S1004: When the confidence level of the hand detection frame is greater than a first confidence level threshold, the mobile phone performs a gesture recognition operation.

[0173] Specifically, after determining that the confidence of the hand detection frame is greater than a first confidence threshold (such as 0.95), the mobile phone can perform a gesture recognition operation, thereby avoiding unnecessary resource waste due to subsequent overlap judgment and saving computing resources.

[0174] The mobile phone performing a gesture recognition operation means that the mobile phone can perform gesture recognition on a hand in a target image frame to obtain a target gesture, and perform a corresponding operation based on the target gesture.

[0175] Specifically, the mobile phone can obtain multiple consecutive image frames after the target image frame. The multiple consecutive image frames can be a preset number of image frames, or all image frames acquired by the mobile phone from the time when the hand exists in the image frame to the time when the hand exists in the last image frame, that is, the multiple consecutive image frames include the hand. Afterwards, the mobile phone can perform gesture recognition on the target image frame and the hands in the multiple image frames based on the gesture recognition model to obtain the target gesture action. The gesture recognition model is a model for performing gesture analysis based on the shape and motion trajectory of the hand in the image. Afterwards, when the target gesture action is any of the preset gesture actions, the mobile phone can perform an operation corresponding to the target gesture action.

[0176] For example, Fig.14 As shown, Fig.14 The interface (a) in the figure is a scene corresponding to an image frame of a user's unfolded palm captured by the front camera 52 of the mobile phone. Fig.14 The (b) interface in the figure is a scene corresponding to the image frame of the user's closed palm captured by the front camera 52 of the mobile phone. Specifically, after the mobile phone captures the image frame of the user's open palm and the image frame of the user's closed palm, it can perform gesture recognition on the image frame of the user's open palm and the image frame of the user's closed palm based on the gesture recognition model to obtain the target gesture action, wherein the target gesture action is the gesture action of the user's palm from open to closed, that is, the gesture action of the palm grasping. Afterwards, the mobile phone can determine whether there is a gesture action that is the same as the target gesture action in the preset gesture actions. If there is a gesture action that is the same as the target gesture action in the preset gesture actions, the mobile phone can perform the operation corresponding to the target gesture action, that is, the mobile phone can perform the remote screenshot operation corresponding to the palm grasping action. Afterwards, the mobile phone can display a prompt message of "Screenshot is being taken..." and perform an operation on the current interface (such as Fig.14 Use the (b) interface in the figure to take a screenshot.

[0177] In some embodiments, after acquiring the target image frame, the mobile phone may first perform hand detection on the target image frame to obtain a first detection result. Afterwards, when the first detection result includes a hand detection frame and the confidence of the hand detection frame, and the confidence of the hand detection frame is greater than the first confidence threshold, it indicates that the accuracy of the first detection result is high, that is, the probability of mistaking a face for a hand is low, and therefore, the mobile phone can directly perform a gesture recognition operation, that is, the mobile phone no longer performs face detection on the target image frame, thereby reducing unnecessary power consumption losses.

[0178] S1005, when the confidence of the hand detection frame is greater than the second confidence threshold but not greater than the first confidence threshold, the mobile phone determines whether the confidence of the face detection frame is greater than the third confidence threshold.

[0179] In some embodiments, after determining that the confidence of the hand detection frame is not greater than the first confidence threshold, but the confidence of the hand detection frame is greater than the second confidence threshold, the mobile phone can further determine whether the confidence of the face detection frame is greater than the third confidence threshold. If the confidence of the face detection frame is greater than the third confidence threshold, it means that the target image frame includes a face, that is, there is a certain probability that the mobile phone will misjudge the face as a hand. Therefore, in order to reduce the probability that the mobile phone misjudges the face as a hand, the mobile phone can execute S1006 to further determine whether the target image frame contains a hand. If the confidence of the face detection frame is not greater than the third confidence threshold, it means that the confidence level of the second detection result is low. Therefore, in order to reduce unnecessary power consumption loss, the mobile phone can ignore the second detection result, that is, it is determined that there are only hands in the target image frame, and there is no face, that is, the mobile phone can directly execute S1009.

[0180] In one implementation, if the confidence of the hand detection frame is not greater than the first confidence threshold, but greater than the second confidence threshold, the mobile phone can further determine the confidence threshold interval corresponding to the confidence of the hand detection frame, so as to provide a basis for the subsequent mobile phone to determine the target overlap threshold. Among them, the confidence threshold interval can be a pre-set interval. In this implementation, the confidence threshold interval can include 0.80-0.95 and 0.70-0.80. In other implementations, the confidence threshold interval can also include 0.60-0.70, etc., which are not specifically limited.

[0181] The third confidence threshold is a preset confidence threshold. In this embodiment, the third confidence threshold is 0.80. In other embodiments, the third confidence threshold may be the same as the first confidence threshold or the second confidence threshold, for example, the third confidence threshold may be 0.95, etc., without specific limitation.

[0182] In some embodiments, after acquiring the target image frame, the mobile phone may first perform hand detection on the target image frame to obtain a first detection result. Afterwards, when the first detection result includes a hand detection frame and the confidence of the hand detection frame, and the confidence of the hand detection frame is not greater than the first confidence threshold, but greater than the second confidence threshold, it means that the first detection result is relatively accurate, but there is still a probability of misjudging a face as a hand. Therefore, the mobile phone may perform face detection on the target image frame again to further determine whether the target image frame contains a face, thereby reducing the probability of misjudging a face as a hand and improving the accuracy of hand detection.

[0183] S1006: The mobile phone determines the overlap between the hand detection frame and the face detection frame based on an overlap algorithm.

[0184] The overlap is determined based on the overlap area between the hand detection frame and the face detection frame and the area of ​​the hand detection frame. Exemplarily, the overlap is the ratio of the overlap area between the hand detection frame and the face detection frame to the area of ​​the hand detection frame.

[0185] Specifically, after determining that the confidence of the face detection frame is greater than the third confidence threshold, the mobile phone can calculate the overlap between the hand detection frame and the face detection frame according to the overlap algorithm. The overlap algorithm is used to calculate the degree of overlap between the hand detection frame and the face detection frame. It can be understood that the greater the degree of overlap between the hand detection frame and the face detection frame, the greater the probability that the face is mistaken for a hand. Therefore, in order to improve the detection accuracy, the mobile phone may not perform the gesture recognition operation to avoid the mobile phone performing an erroneous operation.

[0186] In some embodiments, the mobile phone can extract the target area from the above-mentioned image frame with the face detection frame and the hand detection frame. The target area includes the face area corresponding to the face detection frame and the hand area corresponding to the hand detection frame. Afterwards, the mobile phone can determine the overlap between the hand detection frame and the face detection frame based on the face area and the hand area in combination with the above-mentioned overlap algorithm. Exemplarily, extracting the target area can be that the mobile phone crops the target image frame according to the target area to obtain a target area image. The target area image includes the face area corresponding to the face detection frame and the hand area corresponding to the hand detection frame in the target image frame.

[0187] For example, Fig.15 As shown, image 15 is a target image frame of the user's unfolded palm captured by the front camera of the mobile phone. It can be understood that image 15 is an image frame including a face detection frame H and a hand detection frame N after the mobile phone performs target detection. After obtaining image 15, the mobile phone will extract the face area corresponding to the face detection frame H and the hand area corresponding to the hand detection frame N from image 15. Afterwards, the mobile phone can calculate the ratio between the overlapping area between the face detection frame H and the hand detection frame N and the area of ​​the hand detection frame N. Among them, the overlapping area is the area corresponding to the overlapping area between the face detection frame H and the hand detection frame N.

[0188] S1007: The mobile phone determines whether the above overlap is greater than a target overlap threshold.

[0189] Specifically, after determining the overlap between the hand detection frame and the face detection frame, the mobile phone can determine whether the overlap is greater than the target overlap threshold. If the overlap is greater than the target overlap threshold, it means that the overlap between the hand detection frame and the face detection frame is greater, that is, the probability that the face is misjudged as a hand is higher. Therefore, in order to reduce the occurrence of the mobile phone performing an erroneous operation due to the face being misjudged as a hand, the mobile phone can execute S1008. If the overlap is not greater than the target overlap threshold, it means that the overlap between the hand detection frame and the face detection frame is smaller, that is, the probability that the face is misjudged as a hand is lower. Therefore, the mobile phone can execute S1009.

[0190] In some embodiments, the target overlap threshold may be a preset threshold, for example, the target overlap threshold is 80%, 85%, etc., and is not specifically limited.

[0191] In other embodiments, the target overlap threshold may also be determined based on the confidence threshold interval corresponding to the confidence of the hand detection frame, and the mapping relationship between the confidence threshold interval and the overlap threshold, that is, the target overlap threshold is the overlap threshold corresponding to the confidence threshold interval to which the confidence of the hand detection frame belongs. Specifically, based on the mapping relationship between the confidence threshold interval and the overlap threshold, the mobile phone may search for the overlap threshold corresponding to the confidence threshold interval in which the confidence of the hand detection frame is located, and determine the overlap threshold corresponding to the confidence threshold interval in which the confidence of the hand detection frame is located as the target overlap threshold.

[0192] For example, as shown in Table 1, if the confidence of the hand detection frame is 0.88, that is, the confidence threshold interval corresponding to the confidence of the hand detection frame is 0.80-0.95, then the overlap threshold is 80%. If the confidence of the hand detection frame is 0.76, that is, the confidence threshold interval corresponding to the confidence of the hand detection frame is 0.70-0.80, then the overlap threshold is 70%.

[0193] Table 1

[0194] Confidence threshold interval Overlap Threshold 0.80-0.95 80% 0.70-0.80 70% 0.60-0.70 60%

[0195] It can be understood that if the confidence of the hand detection frame is large, that is, the confidence of the hand detection frame is greater than 0.95, it means that the confidence level of the above-mentioned first detection result is high. Therefore, the mobile phone can directly perform the gesture recognition operation, that is, there is no need to perform subsequent overlap calculations, so that unnecessary waste of resources can be reduced. If the confidence of the hand detection frame is small, that is, the confidence of the hand detection frame is less than 0.60, it means that the confidence level of the above-mentioned first detection result is low, that is, the target image frame may not be clearly captured by the front camera of the mobile phone. If the mobile phone continues to calculate the overlap based on the hand detection frame and the face detection frame, it may cause unnecessary waste of resources. Therefore, in order to reduce the waste of computing resources, the mobile phone may not perform subsequent overlap calculations for hand detection frames with lower confidence.

[0196] It should be noted that the mapping relationship between the confidence threshold interval and the overlap threshold shown in the above Table 1 is only an example, that is, the mapping relationship between the confidence threshold interval and the overlap threshold may change. For example, the overlap threshold corresponding to the confidence threshold interval 0.70-0.80 can be 65%, and this application does not limit it.

[0197] S1008: The mobile phone does not perform the gesture recognition operation.

[0198] In some embodiments, after determining that the overlap between the hand detection frame and the face detection frame is greater than the target overlap threshold, the mobile phone may not perform the gesture recognition operation. In this way, the occurrence of the mobile phone performing erroneous operations due to misjudging the face as a hand can be reduced, thereby improving the execution accuracy of the mobile phone and thereby improving the user experience.

[0199] S1009, the mobile phone performs a gesture recognition operation.

[0200] In some embodiments, after determining that the overlap between the hand detection frame and the face detection frame is not greater than the target overlap threshold, the mobile phone can continue to perform gesture recognition operations. In this way, the gesture detection accuracy can be improved, and the probability of the face being misjudged as a hand can be reduced, thereby improving the execution accuracy of the mobile phone.

[0201] In other embodiments, after determining that the confidence of the face detection frame is not greater than the third confidence threshold, the mobile phone can perform a gesture recognition operation. In this way, the waste of computing resources caused by continuing to calculate the overlap can be reduced, and unnecessary power consumption loss can be avoided.

[0202] S1010: When the confidence level of the hand detection frame is not greater than a second confidence level threshold, the mobile phone does not perform a gesture recognition operation.

[0203] Specifically, after determining that the confidence of the hand detection frame is not greater than the second confidence threshold, the mobile phone may not perform the gesture recognition operation. It can be understood that if the confidence of the hand detection frame is not greater than the second confidence threshold, it means that the confidence level of the first detection result is low, that is, the target image frame may not be clearly captured by the front camera of the mobile phone. If the mobile phone continues to perform gesture recognition, it may cause unnecessary waste of resources. Therefore, in order to reduce the waste of computing resources, the mobile phone may not perform gesture recognition operations, so as to reduce the occurrence of wrong operations performed by the mobile phone due to misrecognition of the hand.

[0204] In one implementation, after acquiring the target image frame, the mobile phone may first perform hand detection on the target image frame to obtain a first detection result. Afterwards, when the first detection result includes a hand detection frame and the confidence of the hand detection frame, and the confidence of the hand detection frame is not greater than the second confidence threshold, it indicates that the accuracy of the first detection result is low, that is, the probability of mistaking a face for a hand is high. Therefore, the mobile phone may not directly perform the gesture recognition operation, that is, the mobile phone may no longer perform face detection on the target image frame, thereby reducing unnecessary power consumption losses.

[0205] It can be understood that the above detection results only include one hand detection frame and one face detection frame, that is, the detection result indicates that the target image frame includes two recognition objects, namely a face and a hand. However, in some scenarios, the detection result may also indicate that the target image frame includes at least three recognition objects, such as two hands and a face. If the detection result includes detection frames for at least three recognition objects, and the at least three recognition objects include at least one hand, the mobile phone can determine the overlap between the hand detection frame and each face detection frame for each hand detection frame. Afterwards, the mobile phone can determine whether the mobile phone performs a gesture recognition operation based on multiple overlaps.

[0206] Specifically, the mobile phone determines whether the multiple overlaps are greater than the target overlap threshold. If the multiple overlaps are greater than the target overlap threshold, it means that the overlap between different recognition objects is relatively large. Therefore, the mobile phone may not perform gesture recognition operations. In this way, the occurrence of erroneous operations performed by electronic devices can be reduced, and the accuracy of electronic devices in performing target operations can be improved, thereby improving the user experience. If there is an overlap less than or equal to the target overlap threshold among the multiple overlaps, it means that there are two recognition objects with relatively small overlaps. Therefore, the mobile phone can perform gesture recognition operations. In this way, the accuracy of electronic devices in performing target operations can be improved, thereby improving the user experience.

[0207] The following is a detailed description using an example where the detection result includes three detection frames of the identified objects.

[0208] In one example, if the detection result includes two hand detection frames and a face detection frame, and the two hand detection frames include a first hand detection frame and a second hand detection frame, the mobile phone can determine the overlap between the first hand detection frame and the face detection frame (or referred to as the first overlap), and the overlap between the second hand detection frame and the face detection frame (or referred to as the second overlap). Afterwards, if both the first overlap and the second overlap are greater than the target overlap threshold, it means that the overlap between the two hands and the user's face is large, so the mobile phone does not perform the gesture recognition operation; if the first overlap is greater than the target overlap threshold, and the second overlap is not greater than the target overlap threshold, the mobile phone can perform the gesture recognition operation based on the hand corresponding to the first overlap; if the second overlap is greater than the target overlap threshold, and the first overlap is not greater than the target overlap threshold, the mobile phone can perform the gesture recognition operation based on the hand corresponding to the second overlap; if both the first overlap and the second overlap are not greater than the target overlap threshold, the mobile phone can not perform the gesture recognition operation, or the mobile phone can select the hand corresponding to the hand detection frame with the largest area to perform the gesture recognition operation.

[0209] In another example, if the detection result includes a hand detection frame and two face detection frames, and the two face detection frames include a first face detection frame and a second face detection frame, the mobile phone can determine the overlap between the first face detection frame and the hand detection frame (or the third overlap), and the overlap between the second face detection frame and the hand detection frame (or the fourth overlap). Afterwards, if the third overlap and the fourth overlap are both less than or equal to the target overlap threshold, the mobile phone can perform a gesture recognition operation; if the third overlap and / or the fourth overlap is greater than the target overlap threshold, the mobile phone does not perform a gesture recognition operation.

[0210] In some embodiments, for some scenarios (such as the user's hand touching the eyes, the user's hand touching the hair, etc.), the mobile phone will still detect the target image frame in the scenario to determine whether a gesture recognition operation needs to be performed. It can be understood that in this scenario, the user is just making an action and does not want to trigger the gesture recognition service. In other words, if the mobile phone continues to detect the target image frame in the scenario, it will mistakenly detect the user's current action as the starting action of the air gesture, causing the mobile phone to mistakenly perform the gesture recognition operation, thereby affecting the user's experience. Therefore, in order to avoid the mobile phone from mistakenly performing the gesture recognition operation, the mobile phone can filter the target image frame in the scenario to improve the user's experience.

[0211] Exemplarily, after the mobile phone determines that the overlap between the hand detection frame and the face detection frame is not greater than the target overlap threshold, the mobile phone can filter the scene of the user's hand touching the eyes by judging the height difference between the hand feature point and the user's eye position. In addition, the mobile phone can also filter the scene of the user's hand touching the hair by judging the height difference between the hand feature point and the upper side length of the face detection frame. The process of filtering based on the target image frame may include the following steps: Fig.16 S1011 to S1013 shown:

[0212] S1011, the mobile phone obtains multiple hand feature points and multiple facial feature points contained in the target image frame.

[0213] Specifically, after determining that the overlap between the hand detection frame and the face detection frame is not greater than the target overlap threshold, the mobile phone can extract the hand feature points and face feature points appearing in the target image frame to obtain multiple hand feature points and multiple face feature points of the target user. The target user refers to the user who is using the mobile phone, that is, the user to whom the face and hand in the target image frame belong.

[0214] In one implementation, the hand feature points may be feature points at the ends of fingers, feature points at joints, feature points at wrists, etc., without limitation. Facial feature points may include eye feature points, nose feature points, mouth feature points, ear feature points, head feature points (or forehead feature points), etc.

[0215] In one example, if Fig.17 As shown, the Fig.17 All hand feature points included in the user's hand are displayed, and the user's hand includes a total of 21 hand feature points. Among them, hand feature point a is a feature point located at the end of the finger, hand feature point b is a feature point located at the connection of the finger joint, and hand feature point h is a feature point located at the wrist position. It should be noted that in this embodiment, only the 21 hand feature points are used as all the feature points of the user's hand. In other embodiments, the number of hand feature points can also be other numbers, such as 26 feature points, etc., which is not specifically limited.

[0216] In another example, if Fig.18 As shown, the Fig.18All facial feature points included in the user's face are displayed, and the user's face includes a total of 10 facial feature points. Among them, facial feature point c is a feature point located at the top of the head, facial feature point d is a feature point located at the eye, facial feature point e is a feature point located at the ear, facial feature point f is a feature point located at the corner of the mouth, and facial feature point g is a feature point located at the tip of the nose. It should be noted that in this embodiment, only the 10 hand feature points are used as all the feature points of the user's hands. In other embodiments, the number of hand feature points can also be other numbers, such as 50 feature points, etc., which are not specifically limited.

[0217] In another implementation, the facial feature points may also include eyebrow feature points, cheek feature points, etc. For example, Fig.19 As shown, the Fig.19 All facial feature points included in the user's face are displayed, and the user's face includes a total of 68 facial feature points, wherein facial feature point i is a feature point located at the eyebrow position, and facial feature point h is a feature point located at the cheek position.

[0218] S1012, the mobile phone determines whether there is a feature point among the above-mentioned multiple hand feature points, the distance between which and the first target height is less than a first preset distance.

[0219] Specifically, after obtaining multiple hand feature points and multiple facial feature points, the mobile phone can determine whether there is a feature point among the multiple hand feature points whose distance from the first target height is less than the first preset distance. The first target height is the height corresponding to the line between the two eye feature points. The first preset distance is a pre-set distance. In this embodiment, the first preset distance is 10 pixels. In other embodiments, the first preset distance can also be 8 pixels, 13 pixels, etc., which are not specifically limited.

[0220] In some embodiments, if there is a feature point whose distance from the first target height is less than the first preset distance among the above-mentioned multiple hand feature points, it means that the above-mentioned target user may be in the process of touching the glasses with his hand, and the target image frame in this scene is captured by the mobile phone. Therefore, in order to avoid the mobile phone from mistakenly performing the gesture recognition operation, the mobile phone can execute S1008. If there is no feature point whose distance from the first target height is less than the first preset distance among the above-mentioned multiple hand feature points, it means that the distances between the above-mentioned multiple hand feature points and the first target height are not less than the first preset distance, that is, the target user's hand is not touching the glasses with his hand. Therefore, the mobile phone can continue to execute S1013 to further determine whether the target user is touching his hair with his hand.

[0221] In one implementation, the mobile phone may only determine whether there is a feature point located at the end of a finger among the above-mentioned multiple hand feature points, the distance between which and the first target height is less than a first preset distance. In other words, the mobile phone does not need to make a distance judgment between each hand feature point and the first target height. In this way, unnecessary power consumption loss can be reduced and detection efficiency can be improved.

[0222] Specifically, after acquiring multiple hand feature points and multiple facial feature points, the mobile phone can determine the first target height according to the relative positions of the two eye feature points. Afterwards, the mobile phone can determine whether the distance between each feature point at the end of the finger and the first target height is less than the first preset distance. If the distance between any feature point at the end of the finger and the first target height is less than the first preset distance, it means that the target user may not have triggered the gesture recognition service, so the mobile phone does not need to perform the gesture recognition operation.

[0223] For example, Fig. 20 As shown, the target image frame P is the image captured by the mobile phone when the user touches the glasses. The mobile phone can connect the eye feature point d1 with the eye feature point d2 to form a line segment L1 (or called the first target height). Afterwards, the mobile phone can respectively determine the distances between the hand feature point a1, the hand feature point a2, and the hand feature point a3 and the line segment L1, that is, determine the distance S1, the distance S2, and the distance S3. Afterwards, the mobile phone can determine whether the distance S1, the distance S2, and the distance S3 are less than the first preset distance (such as 10 pixels). By Fig. 20 It can be seen that the target user touches the glasses with the index finger, which means that the distance S2 between the hand feature point a2 and the line segment L1 will be less than the first preset distance. At the same time, the hand feature point a3 corresponding to the middle finger of the target user is located above the hand feature point a2 and below the eye feature point d2. Therefore, the mobile phone can determine that the distance S3 between the hand feature point a3 and the line segment L1 is less than the first preset distance. In addition, the hand feature point a1 corresponding to the thumb of the target user is located below the hand feature point a2. Therefore, the mobile phone can determine that the distance S1 between the hand feature point a1 and the line segment L1 is greater than the first preset distance.

[0224] S1013, the mobile phone determines whether there is a feature point among the above-mentioned multiple hand feature points, the distance between which and the second target height is less than the second preset distance.

[0225] Specifically, after determining that there is no feature point in the above-mentioned multiple hand feature points whose distance to the first target height is less than the first preset distance, the mobile phone can further determine whether there is a feature point in the multiple hand feature points whose distance to the second target height is less than the second preset distance. Among them, the second target height is the height corresponding to the horizontal line of the head feature point. In this embodiment, the second target height is the height corresponding to the upper side length of the above-mentioned face detection frame. The second preset distance is a pre-set distance, and the second preset distance is greater than the above-mentioned first preset distance. Exemplarily, the second preset distance can be 20 pixels, 15 pixels, etc., and is not specifically limited.

[0226] In some embodiments, if there is a feature point whose distance from the second target height is less than the second preset distance among the above-mentioned multiple hand feature points, it means that the above-mentioned target user may be in the process of touching the hair with his hand, and the target image frame in this scene is captured by the mobile phone. Therefore, in order to avoid the mobile phone from mistakenly performing the gesture recognition operation, the mobile phone can execute S1008. If there is no feature point whose distance from the second target height is less than the second preset distance among the above-mentioned multiple hand feature points, it means that the distances between the above-mentioned multiple hand feature points and the second target height are not less than the second preset distance, that is, the target user's hand is not touching the hair, so the mobile phone can continue to execute S1009 to further recognize the user gesture.

[0227] In one implementation, the mobile phone may only determine whether there is a feature point located at the end of a finger among the above-mentioned multiple hand feature points, the distance between which and the second target height is less than the second preset distance. In other words, the mobile phone does not need to make a distance judgment between each hand feature point and the second target height. In this way, unnecessary power consumption loss can be reduced and detection efficiency can be improved.

[0228] Specifically, after determining that there is no feature point among the above-mentioned multiple hand feature points whose distance from the first target height is less than the first preset distance, the mobile phone can use the height corresponding to the upper side length of the above-mentioned face detection frame as the second target height. Afterwards, the mobile phone can determine whether the distance between each hand feature point located at the end of the finger and the second target height is less than the second preset distance. If there is any feature point located at the end of the finger whose distance from the second target height is less than the second preset distance, it means that the target user may not have triggered the gesture recognition service, and therefore, the mobile phone does not need to perform the gesture recognition operation.

[0229] For example, Fig.21As shown, the target image frame K is the image captured by the mobile phone when the user touches his hair. The mobile phone can use the height corresponding to the upper length of the face detection frame J as the second target height L2. Afterwards, the mobile phone can respectively determine the distances between the hand feature point a4, the hand feature point a5, and the hand feature point a6 and the second target height L2, that is, determine the distance S4, the distance S5, and the distance S6. Afterwards, the mobile phone can determine whether the distance S4, the distance S5, and the distance S6 are less than the second preset distance (such as 20 pixels). By Fig.21 It can be seen that the target user touches the hair with the index finger, which means that the distance S5 between the hand feature point a5 and the second target height L2 will be less than the second preset distance. At the same time, the hand feature point a6 corresponding to the middle finger of the target user is located above the hand feature point a5 and below the upper side length of the face detection frame J. Therefore, the mobile phone can determine that the distance S6 between the hand feature point a6 and the second target height L2 is less than the second preset distance. In addition, the hand feature point a4 corresponding to the thumb of the target user is located below the hand feature point a5. Therefore, the mobile phone can determine that the distance S4 between the hand feature point a5 and the second target height L2 is greater than the second preset distance.

[0230] It should be noted that the execution order of the above-mentioned filtering process for the scene of hand touching glasses and the filtering process for the scene of hand touching hair is not limited. For example, the mobile phone can first execute the above-mentioned step S1012 and then execute the step S1013, or the mobile phone can first execute the step S1013 and then execute the above-mentioned step S1012, or the mobile phone can execute the above-mentioned steps S1012 and S1013 at the same time. There is no specific limitation.

[0231] In some embodiments, the present application provides a computer storage medium including computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the method for adjusting usage parameters as described above.

[0232] In some embodiments, the present application provides a computer program product, which, when executed on an electronic device, enables the electronic device to execute the method for adjusting usage parameters as described above.

[0233] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0234] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0235] The units described as separate components may or may not be physically separated, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple different places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0236] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0237] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.

[0238] The above contents are only specific implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application shall be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A target object recognition method, It is characterized in that include: The electronic device acquires a target image frame; wherein the target image frame is an image frame currently acquired by the electronic device; The electronic device performs target detection on the target image frame to obtain a detection result, wherein the detection result is used to indicate an identification object included in the target image frame; In a case where the detection result indicates that the target image frame includes two recognition objects, the electronic device determines the overlap between the two recognition objects; wherein the two recognition objects include a first recognition object, and the overlap refers to the ratio between the overlapping areas between the detection frames corresponding to the two recognition objects and the area of ​​the detection frame corresponding to the first recognition object; When the overlap is less than or equal to a target overlap threshold, the electronic device performs a target operation corresponding to the first identified object; wherein the target overlap threshold is positively correlated with a confidence threshold interval to which the confidence of the detection box corresponding to the first identified object belongs.

2. The method according to claim 1, It is characterized in that The two recognition objects are face and hand, the first recognition object is the hand, and the detection result includes a face detection frame and a hand detection frame; The electronic device determines the overlap between the detection frames corresponding to the two recognition objects, including: The electronic device determines the degree of overlap between the face detection frame and the hand detection frame; The electronic device performs a target operation corresponding to the first recognition object, including: The electronic device performs a gesture recognition operation corresponding to the hand.

3. The method according to claim 2, It is characterized in that The detection result also includes the confidence of the face detection frame and the confidence of the hand detection frame, and the electronic device determines the overlap between the face detection frame and the hand detection frame, including: When the confidence of the hand detection frame is greater than the second confidence threshold and less than or equal to the first confidence threshold, if the confidence of the face detection frame is greater than the third confidence threshold, the electronic device calculates the overlap between the face detection frame and the hand detection frame based on an overlap algorithm; The first confidence threshold is greater than the second confidence threshold, and the overlap is the ratio of the overlapping area between the face detection frame and the hand detection frame to the area of ​​the hand detection frame.

4. The method according to claim 2, It is characterized in that The detection result also includes the confidence of the face detection frame and the confidence of the hand detection frame. The method also includes: When the confidence level of the hand detection frame is greater than a first confidence level threshold, the electronic device performs a gesture recognition operation; or, When the confidence level of the hand detection frame is less than or equal to a second confidence level threshold, the electronic device does not perform a gesture recognition operation.

5. The method according to any one of claims 2 to 4, It is characterized in that The electronic device performs a gesture recognition operation, including: The electronic device acquires a plurality of hand feature points and a plurality of facial feature points in the target image frame; wherein the facial feature points include two eye feature points; When the distances between the multiple hand feature points and the first target height are all greater than or equal to the first preset distance, the electronic device performs a gesture recognition operation; wherein the first target height is the height corresponding to the line between the two eye feature points.

6. The method according to any one of claims 2 to 4, It is characterized in that The electronic device performs a gesture recognition operation, including: The electronic device acquires a plurality of hand feature points and a plurality of facial feature points in the target image frame; wherein the facial feature points include forehead feature points; When the distances between the multiple hand feature points and the second target height are all greater than or equal to the second preset distance, the electronic device performs a gesture recognition operation; wherein the second target height is the height corresponding to the horizontal line of the forehead feature points.

7. The method according to any one of claims 2 to 4, It is characterized in that The target overlap threshold is an overlap threshold corresponding to the confidence threshold interval to which the confidence of the hand detection frame belongs.

8. The method according to claim 2, It is characterized in that The electronic device performs target detection on the target image frame to obtain a detection result, including: The electronic device performs hand detection on the target image frame based on the hand detection model to obtain a first detection result; wherein the first detection result is used to indicate whether the target image frame includes a hand; The electronic device performs face detection on the target image frame based on a face detection model to obtain a second detection result; wherein the second detection result is used to indicate whether the target image frame includes a face.

9. The method according to claim 8, It is characterized in that The electronic device performs face detection on the target image frame based on the face detection model to obtain a second detection result, including: In a case where the first detection result indicates that the target image frame includes a hand, if the confidence of the hand detection frame is greater than a second confidence threshold and less than or equal to the first confidence threshold, the electronic device performs face detection on the target image frame based on the face detection model to obtain the second detection result; When the detection result indicates that the target image frame includes two recognition objects, the electronic device determines the overlap between the detection frames corresponding to the two recognition objects, including: In a case where the second detection result indicates that the target image frame includes a face, if the confidence of the face detection frame is greater than a third confidence threshold, the electronic device determines the overlap between the face detection frame and the hand detection frame.

10. The method according to claim 1, It is characterized in that The two recognition objects are a two-dimensional code and a background object, the first recognition object is a two-dimensional code, and the detection result includes a two-dimensional code detection frame and a background object detection frame; The electronic device determines the overlap between the detection frames corresponding to the two recognition objects, including: The electronic device determines the degree of overlap between the two-dimensional code detection frame and the background object detection frame; The electronic device performs a target operation corresponding to the first recognition object, including: The electronic device performs a two-dimensional code recognition operation corresponding to the two-dimensional code.

11. An electronic device, It is characterized in that The electronic device includes a camera, a memory and one or more processors; the camera, the memory and the processor are coupled; the camera is used to capture images, the memory is used to store computer program codes, and the computer program codes include computer instructions; when the processor executes the computer instructions, the electronic device executes the method as described in any one of claims 1 to 10.

12. A computer-readable storage medium, It is characterized in that The method comprises computer instructions, which, when executed on an electronic device, cause the electronic device to execute the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Gesture recognition method taking face as reference

    CN106971130A

  • Method and device for determining interaction gesture and electronic equipment

    CN114816045A