Target detection method, electronic device and gesture detection system

By setting an effective judgment range, gesture recognition is performed only when the hand position is within the effective range, which solves the problems of gesture recognition misjudgment and resource waste in the existing technology and achieves efficient gesture recognition.

CN120823640APending Publication Date: 2025-10-21COOL BOLE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410435063.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-11
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing gesture recognition methods cannot distinguish meaningful gestures from unconscious finger movements, resulting in misjudgments and consuming a large amount of computing resources.

Method used

By setting the effective judgment range, gesture recognition is performed only when the hand position is within the effective judgment range. The processor is used to execute the target detection module to obtain the head, body and hand position information, and the gesture recognition module is executed after the effective judgment range is set.

Benefits of technology

It reduces misjudgment of gesture recognition, saves computing resources and improves recognition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823640A_ABST
    Figure CN120823640A_ABST
Patent Text Reader

Abstract

The invention provides a target detection method, an electronic device and a gesture detection system. A processor is utilized to implement the following steps: executing a target detection module to detect an original image, and acquiring first position information, second position information and third position information related to the same portrait object from the original image through the target detection module; setting an effective determination range based on at least one of the first position information and the second position information; obtaining a hand position in the original image based on the third position information; executing a gesture recognition module in response to the fact that the hand position is located in the effective judgment range; and responding to the fact that the hand position is not in the effective judgment range, not executing the gesture recognition module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image recognition mechanism, and more particularly to a target detection method, an electronic device, and a gesture detection system. Background Art

[0002] Mediapipe Holistic combines three models and related algorithms: human pose, facial landmarks, and hand tracking. It can detect body poses, facial meshes, and hand movements. Complete detection generates 543 detection nodes (33 pose nodes, 468 face nodes, and 21 hand nodes for each hand). However, this computational approach consumes significant computing resources, making it difficult to scale.

[0003] Furthermore, in practical applications, current gesture recognition methods cannot distinguish between intentional gestures and unconscious finger movements. Whether intentional gestures directed at the camera or specific objects, or unconscious finger movements, all are detected and recognized as gestures. Consequently, existing gesture recognition methods are prone to misjudgment and waste computing resources. Summary of the Invention

[0004] The present invention provides a target detection method, an electronic device, and a gesture detection system, which can reduce misjudgment of gesture recognition and save computing resources.

[0005] The target detection method of the present invention uses a processor to implement the following steps, including: executing a target detection module to detect an original image, and obtaining first position information, second position information and third position information related to the same portrait object from the original image through the target detection module, wherein the first position information corresponds to the head area, the second position information corresponds to the body area, and the third position information corresponds to the hand area; setting an effective judgment range based on at least one of the first position information and the second position information; obtaining the hand position in the original image based on the third position information; executing a gesture recognition module in response to the hand position being within the effective judgment range; and not executing the gesture recognition module in response to the hand position not being within the effective judgment range.

[0006] In one embodiment of the present invention, the step of setting the effective determination range includes: calculating the facial width and facial area based on first position information corresponding to the head area; setting a threshold value based on the facial width; and setting the effective determination range as a circular range with a center point of the facial area as the center and the threshold value as the radius.

[0007] In one embodiment of the present invention, the step of calculating the facial width includes: calculating the height of the head region in the vertical direction and the width of the head region in the horizontal direction based on first position information corresponding to the head region; judging whether the obtained facial region is a front face or a profile face based on a ratio of the height and the width; if the facial region is determined to be a front face, using the width as the facial width; and if the facial region is determined to be a profile face, not calculating the facial width, and thus not executing the gesture recognition module.

[0008] In one embodiment of the present invention, the step of setting the effective judgment range includes: obtaining the body length range in the vertical direction based on the second position information corresponding to the body area; and setting the effective judgment range in the body length range according to a preset ratio to determine whether the position of the hand in the vertical direction is within the effective judgment range.

[0009] In one embodiment of the present invention, the preset ratio includes a first ratio and a second ratio, and the step of setting the effective judgment range in the body length range according to the preset ratio includes: when the portrait object is determined to be in a half-body state, setting the effective judgment range in the body length range according to the first ratio; and when the portrait object is determined to be in a full-body state, setting the effective judgment range in the body length range according to the second ratio.

[0010] In one embodiment of the present invention, the target detection method further includes: obtaining the head size of the head area in the vertical direction with reference to the first position information; obtaining the length and width of the body area with reference to the second position information; and determining whether the portrait object is in a half-body state or a full-body state based on the length, width and head size.

[0011] In one embodiment of the present invention, the step of setting the effective judgment range includes: obtaining a width range in the horizontal direction based on the second position information corresponding to the body area; calculating the top of the head position based on the first position information corresponding to the head area; and setting the effective judgment range based on the upper area and the width range of the top of the head position.

[0012] In one embodiment of the present invention, in response to the hand position being within the valid determination range, a gesture recognition module is executed to obtain a gesture recognition result; and based on the gesture recognition result, a corresponding operation is performed.

[0013] In one embodiment of the present invention, the operation includes at least one of controlling the actuation of a physical device and controlling the adjustment of parameter settings of an electronic device having the processor.

[0014] The electronic device of the present invention includes: a communication interface configured to receive an original image; and a processor coupled to the communication interface and configured to execute the target detection method.

[0015] The gesture detection system of the present invention includes: an imaging device for obtaining an original image and the electronic device.

[0016] Based on the above, the present disclosure can pre-filter out the user's unconscious or meaningless gesture activities by setting an effective determination range, thereby reducing misjudgment of gesture recognition and saving computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a block diagram of a gesture detection system according to an embodiment of the present invention.

[0018] Figure 2 is a flow chart of a target detection method according to an embodiment of the present invention.

[0019] Figure 3 is a schematic diagram of an original image according to an embodiment of the present invention.

[0020] Figure 4 FIG. 1 is a schematic diagram of a first application example of an effective determination range according to an embodiment of the present invention.

[0021] Figure 5 FIG. 1 is a schematic diagram of a second application example of the effective determination range according to an embodiment of the present invention.

[0022] Figure 6 FIG. 1 is a schematic diagram of a third application example of the effective determination range according to an embodiment of the present invention.

[0023] Figure 7 FIG. 4 is a schematic diagram of a fourth application example of the effective determination range according to an embodiment of the present invention.

[0024] Description of Reference Numerals

[0025] 10: Gesture detection system

[0026] 100: Electronic devices

[0027] 110: Processor

[0028] 120: Communication interface

[0029] 130: Imaging device

[0030] 41, 51, 61, 71: Valid judgment range

[0031] b310, b320, b330, b332: bounding box

[0032] b410, b420, b430, b432: bounding box

[0033] b510, b520, b530, b532: bounding box

[0034] b610, b620, b631, b632: bounding box

[0035] b710, b720, b730, b732: bounding box

[0036] h1, h2: height range

[0037] I300, I400, I500, I600, I700: Original image

[0038] S205~S230: Steps DETAILED DESCRIPTION

[0039] Figure 1 is a block diagram of a gesture detection system according to an embodiment of the present invention. Figure 1 The gesture detection system 10 includes an electronic device 100 and an imaging device 130. The electronic device 100 is, for example, an electronic device with computing functions such as a smart phone, a tablet computer, a laptop computer, a personal computer, etc. The imaging device 130 is a camera or a still camera that uses a charge coupled device (CCD) lens, a complementary metal oxide semiconductor transistor (CMOS) lens, etc. The imaging device 130 can be connected to the electronic device 1001 for communication via a wired or wireless method. The electronic device 100 includes a processor 110 and a communication interface 120. The processor 110 is coupled to the communication interface 120. The communication interface 120 is used to receive the original image from the imaging device 130.

[0040] The processor 110 is, for example, a central processing unit (CPU), a graphics processing unit (GPU), a physical processing unit (PPU), a programmable microprocessor, an embedded control chip, a digital signal processor (DSP), an application specific integrated circuit (ASIC), or other similar devices.

[0041] The communication interface 120 is used to communicate with other devices or communication networks. The communication network can be an Ethernet network, a Radio Access Network (RAN), or a wireless local area network (WLAN). The communication interface 120 can be a wired communication interface or a wireless communication interface. Specifically, the communication interface 120 can be an Ethernet interface, a Fast Ethernet (FE) interface, a Gigabit Ethernet (GE) interface, an Asynchronous Transfer Mode (ATM) interface, a Wireless Local Area Network (WLAN) interface, a cellular network communication interface, or a combination thereof. The Ethernet interface can be an optical interface, an electrical interface, or a combination thereof. The communication interface 120 can be used to communicate with network devices and other devices.

[0042] In another embodiment, the communication interface 120 is a wired / wireless signal transceiver such as a network interface card, a high frequency circuit (RF circuit), a Bluetooth signal transceiver, or an infrared signal transceiver.

[0043] The electronic device 100 also includes a memory. The memory can be implemented as any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, a hard disk, or other similar device, or a combination of these devices. The memory includes one or more code segments, which, after being installed, are executed by the processor 110 to implement the target detection method described below.

[0044] Figure 2 This is a flow chart of a target detection method according to an embodiment of the present invention. Figure 1 and Figure 2In step S205, the processor 110 executes the target detection module to detect the original image, and obtains the first position information, the second position information, and the third position information related to the same portrait object from the original image through the target detection module. Here, the first position information corresponds to the head area, the second position information corresponds to the body area, and the third position information corresponds to the hand area. The processor 110 inputs the original image into the target detection module, and after recognition by the target detection module, obtains the first position information, the second position information, and the third position information corresponding to the head area, the body area, and the hand area, respectively. The processor 110 can further mark the bounding boxes corresponding to the head area, the body area, and the hand area, respectively, on the original image based on the first position information, the second position information, and the third position information.

[0045] In one embodiment, the object detection module is pre-trained through a large number of sample images to understand general image knowledge. The object detection module is based on a convolutional neural network (CNN) architecture and is divided into three parts: a backbone network, a link layer (neck), and a detection head (head). The backbone network is responsible for extracting features from the original image. For example, the backbone network can extract multiple initial feature layers (initial feature layers) of different scales from the original image from the bottom up. The backbone network can adopt models such as ResNet-18, MobileNetV2-100, and ShuffleNetV2.

[0046] The link layer is used to reprocess and rationally utilize important features extracted by the backbone network. For example, it can simultaneously extract features at different stages to facilitate task-specific learning for the detection head. The link layer can include both top-down and bottom-up paths. The link layer can adopt a Feature Pyramid Network (FPN) structure or a bidirectional FPN structure.

[0047] The detection head generates specific outputs based on different detection targets (for example, body area, head area, and hand area). The detection head is responsible for redrawing the features extracted by the backbone network into several fixed-size grids (grids), such as 64×64, 32×32, or 16×16, and then predicting the probability of the center of the object appearing in each grid, the size, position, and category of the anchor. For example, the detection head includes a classification branch and a bounding box regression branch. The classification branch is used to obtain the classification probability distribution. The bounding box regression branch is used to obtain the bounding box position probability distribution.

[0048] In this embodiment, during the training stage, multiple bounding boxes corresponding to the body area, head area, and hand area of ​​the same portrait object are marked in each training image, and these data are input into the target detection module for training, and the detection head is further set to output position information corresponding to the body area, head area, and hand area.

[0049] After the target detection module completes training, the processor 110 inputs an original image to be identified into the target detection module, and the target detection module can output position information corresponding to the body area, head area, and hand area.

[0050] Figure 3 is a schematic diagram of an original image according to an embodiment of the present invention. Figure 3 In this embodiment, the detection targets of the object detection module include the head region, the body region, and the hand region. The processor 110 inputs the original image I300 into the object detection module. After the object detection module detects (identifies the head region, body region, and hand region of the same portrait object), it obtains first position information corresponding to the head region, second position information corresponding to the body region, and third position information corresponding to the hand region. Then, based on the first position information, the second position information, and the third position information, the processor 110 annotates the bounding box b320 corresponding to the head region, the bounding box b310 corresponding to the body region, and the bounding boxes b330 and b332 corresponding to the hand region (assuming both hands are recognized) in the original image I300. For example, the first position information includes the upper left and lower right coordinates of the bounding box b320, the second position information includes the upper left and lower right coordinates of the bounding box b310, and the third position information includes the upper left and lower right coordinates of the bounding box b330, as well as the upper left and lower right coordinates of the bounding box b332.

[0051] In this embodiment, the range of the body region (ie, the range enclosed by the bounding box b310 ) covers the entire portrait object.

[0052] return Figure 2 After obtaining the first, second, and third location information, in step S210, the processor 110 sets an effective determination range based on at least one of the first and second location information. The effective determination range is used to determine whether to execute the gesture recognition module. In one embodiment, the effective determination range can be set for different usage scenarios. Examples of usage scenarios include, but are not limited to, remote monitoring, gaming, and conference scenarios.

[0053] Next, in step S215, processor 110 obtains the hand positions in the original image based on the third position information. Taking original image I300 as an example, based on the third position information, bounding boxes b330 and b332 are obtained. The center points of bounding boxes b330 and b332 are then determined to represent the hand positions of the left and right hands. In other embodiments, arbitrary reference points can be used within bounding boxes b330 and b332 to represent the hand positions of the left and right hands.

[0054] In step S220, processor 110 determines whether the hand position is within the valid determination range. In response to the hand position being within the valid determination range, processor 110 executes the gesture recognition module in step S225. In response to the hand position not being within the valid determination range, processor 110 does not execute the gesture recognition module in step S230.

[0055] In step S225, the processor 110 executes the gesture recognition module and obtains the gesture recognition result, and then performs the corresponding operation based on the gesture recognition result. The operation includes at least one of controlling the actuation of the physical device and controlling the adjustment of the parameter settings of the electronic device 100. For example, the processor 110 decides to control the imaging device 130 to start or stop recording a video based on the gesture recognition result. Alternatively, the processor 110 decides to control the sound receiving device (microphone) to start or stop receiving sound based on the gesture recognition result. The sound receiving device can be built into the electronic device 100, or it can be externally connected to the electronic device 100 via a wired or wireless method. Alternatively, the processor 110 determines whether the speaker of the electronic device 100 is turned on or not based on the gesture recognition result. Alternatively, the processor 110 determines the brightness parameters of the display of the electronic device 100 based on the gesture recognition result.

[0056] The following examples illustrate the setting of the effective determination range.

[0057] Figure 4This is a schematic diagram illustrating a first application example of the effective determination range according to an embodiment of the present invention. Generally speaking, when performing meaningful gestures, the hand must be within a certain distance from the head. Therefore, in the usage scenario of this embodiment, the gesture recognition module will only be executed when the distance between the hand position and the facial area is less than a certain distance.

[0058] Please refer to Figure 4 In this embodiment, the object detection module marks the bounding box b410 corresponding to the body area, the bounding box b420 corresponding to the head area, and the bounding boxes b430 and b432 corresponding to the hand area in the original image I400.

[0059] Specifically, the processor 110 calculates the face width and the face area based on the first position information corresponding to the head area. For example, the range of the face area relative to the head area can be obtained based on a preset ratio obtained by statistics, and then the face width can be obtained from the widest point of the face area in the horizontal direction. Then, the processor 110 sets a threshold value based on the face width. For example, 1.5 times the face width is used as the threshold value. Then, the processor 110 sets the effective judgment range 41 as a circular range with the center point of the face area as the center and the threshold value as the radius. Figure 4 In the illustrated embodiment, the processor 110 determines that the hand position of one of the hands (the center point position of the bounding box b432 ) is located within the valid determination range 41 .

[0060] Furthermore, in human behavior, meaningful gestures are always performed with the face facing forward; the gesture cannot be accurately determined if the face is facing sideways. Accordingly, in another embodiment, after obtaining an output result through detection by the object detection module, the processor 110 can further determine whether the facial region is frontal or in profile. Specifically, based on the first position information corresponding to the head region, the processor 110 calculates the vertical height and horizontal width of the head region, and determines whether the obtained facial region is frontal or in profile based on the ratio of the height to width. For example, if the ratio of height to width is greater than or equal to 2, the facial region is determined to be in profile; if the ratio of height to width is less than 2, the facial region is determined to be frontal.

[0061] If the face region is determined to be frontal, the processor 110 uses the width as the face width. If the face region is determined to be profile, the face width is not calculated, and the gesture recognition module is not executed.

[0062] In other usage scenarios, you can also set the effective judgment range based on standing or sitting posture. Figure 5 To illustrate the standing posture, Figure 6 To illustrate the sitting posture.

[0063] Figure 5 is a schematic diagram of a second application example of the effective determination range according to an embodiment of the present invention. Figure 5 The object detection module annotates the original image I500 with a bounding box b510 corresponding to the body region, a bounding box b520 corresponding to the head region, and bounding boxes b530 and b532 corresponding to the hand regions. This embodiment is suitable for use in scenarios where the imaging device 130 is far away from the person being photographed, such as in object detection applications in factories.

[0064] Specifically, the processor 110 obtains the body length range h1 in the vertical direction based on the second position information corresponding to the body area, and sets the effective judgment range 51 in the body length range h1 according to a preset ratio to determine whether the hand position in the vertical direction is within the effective judgment range.

[0065] In the case where the portrait object is in a full-body state, it is assumed that the second position information of the body area includes the upper left coordinate point (x1, y1) and the lower right coordinate point (x2, y2), and the body length range h1 is set to y1~y2. In addition, it is assumed that the preset ratio (second ratio) obtained based on statistical data includes 1 / 4 and 1 / 2. The two positions α1 and β1 in the vertical direction are calculated based on the preset ratio, and the range between the two positions α1 and β1 is set as the effective judgment range 51. That is, α1=y1+(y2-y1)×(1 / 4), β1=y1+(y2-y1)×(1 / 2). In Figure 5 In the illustrated embodiment, the processor 110 determines that the hand position of one of the hands (the center point position of the bounding box b532 ) is located within the valid determination range 51 .

[0066] The processor 110 obtains the vertical head size of the head region by referring to the first position information corresponding to the head region, and obtains the body length and width of the body region by referring to the second position information corresponding to the body region. The processor 110 then determines whether the portrait object is in a half-body or full-body state based on the body length, width, and head size of the body region. For example, if the ratio of the body length to the body width of the body region is less than a first preset value, and the difference between the body length and the head size of the body region is greater than a second preset value, the portrait object is determined to be in a half-body state. If the ratio of the body length to the body width of the body region is not less than the first preset value, and the difference between the body length and the head size of the body region is not greater than a second preset value, the portrait object is determined to be in a full-body state.

[0067] Figure 6 FIG is a schematic diagram of a third application example of an effective determination range according to an embodiment of the present invention. Figure 6In this embodiment, the object detection module annotates the original image I600 with a bounding box b610 corresponding to the body region, a bounding box b620 corresponding to the head region, and bounding boxes b631 and b632 corresponding to the hand regions. This embodiment is suitable for use in scenarios where the imaging device 130 is relatively close to the subject, such as when performing object detection in a meeting.

[0068] Specifically, when the portrait object is in a half-body state, it is assumed that the second position information of the body area includes the upper left coordinate point (x3, y3) and the lower right coordinate point (x4, y4), and the body length range h2 is set to y3~y4. In addition, it is assumed that the preset ratio (first ratio) obtained based on statistical data includes 1 / 2 and 1 / 1. The two positions α2 and β2 in the vertical direction are calculated based on the preset ratio, and the range between the two positions α2 and β2 is set as the effective judgment range 61. That is, α2=y3+(y4-y3)×(1 / 2), β2=y3+(y4-y3)×1. In Figure 6 In the illustrated embodiment, the processor 110 determines that the hand position of one of the hands (the center point position of the bounding box b 632 ) is located within the valid determination range 61 .

[0069] Figure 7 FIG is a schematic diagram of a fourth application example of an effective determination range according to an embodiment of the present invention. Figure 7 In this embodiment, the object detection module marks the bounding box b710 corresponding to the body area, the bounding box b720 corresponding to the head area, and the bounding boxes b730 and b732 corresponding to the hand area in the original image I700.

[0070] Specifically, the processor 110 obtains the width range in the horizontal direction based on the second position information corresponding to the body area, calculates the top of the head position based on the first position information corresponding to the head area, and then sets the effective judgment range 71 based on the upper area of ​​the top of the head position and the width range.

[0071] For example, assuming that the second position information of the body area includes the upper left coordinate point (x5, y5) and the lower right coordinate point (x6, y6), the body width range is set to x5 to x6. Assuming that the top of the head is at y0, the area between x5 and x6 and above y0 is set as the effective judgment range 71. Figure 7 In the illustrated embodiment, the processor 110 determines that the hand position of one of the hands (the center point position of the bounding box b730 ) is located within the valid determination range 71 .

[0072] In addition, in another embodiment, it can also be set as follows: when the bounding box (b730 or b732) of the hand area is located above the head, and the center point of the bounding box (b730 or b732) of the hand area falls within the width range of the bounding box b710 corresponding to the body, the gesture recognition module is executed. And, when the recognized gesture meets the preset gesture (such as fist, open palm), the corresponding operation is performed. For example, when the processor 110 determines through the object detection module and the gesture recognition module that the user makes a fist and raises one hand above the head, the imaging device 130 is controlled to start recording the video. When the processor 110 determines through the object detection module and the gesture recognition module that the user spreads the palm of one hand and raises it above the head, the imaging device 130 is controlled to stop recording the video.

[0073] In summary, the present disclosure can pre-filter out the user's unconscious or meaningless gesture activities by setting an effective determination range, thereby reducing misjudgment of gesture recognition and saving computing resources.

[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A target detection method, characterized in that: The processor is used to implement the following steps, including: executing an object detection module to detect an original image, and obtaining, by the object detection module, first position information, second position information, and third position information associated with a same portrait object from the original image, wherein the first position information corresponds to a head region, the second position information corresponds to a body region, and the third position information corresponds to a hand region; setting a valid determination range based on at least one of the first location information and the second location information; Based on the third position information, obtaining a hand position in the original image; In response to the hand position being within the valid determination range, executing a gesture recognition module; and In response to the hand position not being within the valid determination range, the gesture recognition module is not executed.

2. The target detection method according to claim 1, wherein: The step of setting the effective determination range based on at least one of the first location information and the second location information includes: calculating a face width and a face area based on the first position information corresponding to the head area; setting a threshold value based on the face width; and The effective determination range is set as a circular range with the center point of the face area as the center and the threshold value as the radius.

3. The target detection method according to claim 2, characterized in that: The step of calculating the face width based on the first position information corresponding to the head region includes: calculating a height in a vertical direction and a width in a horizontal direction of the head region based on the first position information corresponding to the head region; Determining whether the obtained facial region is a front face or a side face based on a ratio of the height to the width; In the case where it is determined that the face area is frontal, using the width as the face width; and When it is determined that the facial region is a side face, the facial width is not calculated, and the gesture recognition module is not executed.

4. The target detection method according to claim 1, wherein: The step of setting the effective determination range based on at least one of the first location information and the second location information includes: obtaining a body length range in a vertical direction based on the second position information corresponding to the body region; and The effective determination range is set according to a preset ratio within the body length range to determine whether the position of the hand in the vertical direction is within the effective determination range.

5. The target detection method according to claim 4, characterized in that: The preset ratio includes a first ratio and a second ratio, and the step of setting the effective determination range in the body length range according to the preset ratio includes: When it is determined that the portrait object is in a half-body state, setting the effective determination range within the body length range according to the first ratio; and When it is determined that the human-like object is in a full-body state, the effective determination range is set in the body length range according to the second ratio.

6. The target detection method according to claim 5, further comprising: obtaining a head size of the head region in the vertical direction by referring to the first position information; obtaining the length and width of the body region by referring to the second position information; as well as The human-like object is judged to be in the half-body state or the full-body state based on the body length, the body width, and the head size.

7. The target detection method according to claim 1, wherein: The step of setting the effective determination range based on at least one of the first location information and the second location information includes: obtaining a width range in a horizontal direction based on the second position information corresponding to the body region; calculating a top-of-the-head position based on the first position information corresponding to the head region; and The effective determination range is set based on the area above the top of the head position and the width range.

8. The target detection method according to claim 1, wherein: In response to the hand position being within the valid determination range, the method further comprises: Executing the gesture recognition module and obtaining a gesture recognition result; and Based on the gesture recognition result, a corresponding operation is performed.

9. The target detection method according to claim 8, characterized in that: The operation includes at least one of controlling the actuation of a physical device and controlling the adjustment of parameter settings of an electronic device having the processor.

10. An electronic device, characterized in that: include: a communication interface configured to receive a raw image; as well as a processor coupled to the communication interface and configured to: executing an object detection module to detect the original image, and obtaining, by the object detection module, first position information, second position information, and third position information associated with a same portrait object from the original image, wherein the first position information corresponds to a head region, the second position information corresponds to a body region, and the third position information corresponds to a hand region; setting a valid determination range based on at least one of the first location information and the second location information; Based on the third position information, obtaining a hand position in the original image; In response to the hand position being within the valid determination range, executing a gesture recognition module; as well as In response to the hand position not being within the valid determination range, the gesture recognition module is not executed.

11. The electronic device of claim 10, wherein the processor is configured to: calculating a face width and a face area based on the first position information corresponding to the head area; setting a threshold value based on the face width; and The effective determination range is set as a circular range with the center point of the face area as the center and the threshold value as the radius.

12. The electronic device of claim 11, wherein the processor is configured to: calculating a height in a vertical direction and a width in a horizontal direction of the head region based on the first position information corresponding to the head region; Determining whether the obtained facial region is a front face or a side face based on a ratio of the height to the width; When it is determined that the face area is frontal, the width is used as the face width; as well as When it is determined that the facial region is a side face, the facial width is not calculated, and the gesture recognition module is not executed.

13. The electronic device according to claim 10, wherein: The processor is configured to: obtaining a body length range in a vertical direction based on the second position information corresponding to the body region; and The effective determination range is set according to a preset ratio within the body length range to determine whether the position of the hand in the vertical direction is within the effective determination range.

14. The electronic device according to claim 13, wherein: The preset ratio includes a first ratio and a second ratio, and the processor is configured to: When it is determined that the portrait object is in a half-body state, setting the effective determination range within the body length range according to the first ratio; as well as When it is determined that the human-like object is in a full-body state, the effective determination range is set in the body length range according to the second ratio.

15. The electronic device according to claim 14, wherein: The processor is configured to: obtaining a head size of the head region in the vertical direction by referring to the first position information; obtaining the length and width of the body region by referring to the second position information; as well as The human-like object is judged to be in the half-body state or the full-body state based on the body length, the body width, and the head size.

16. The electronic device according to claim 10, wherein: The processor is configured to: obtaining a width range in a horizontal direction based on the second position information corresponding to the body region; calculating a top-of-the-head position based on the first position information corresponding to the head region; as well as The effective determination range is set based on the area above the top of the head position and the width range.

17. The electronic device according to claim 10, wherein: The processor is configured to: Executing the gesture recognition module and obtaining a gesture recognition result; and Based on the gesture recognition result, a corresponding operation is performed.

18. The electronic device according to claim 17, wherein: The operation includes at least one of controlling the actuation of a physical device and controlling the adjustment of parameter settings of the electronic device.

19. A gesture detection system, characterized in that: include: an imaging device configured to obtain an original image; as well as Electronic device, including: a communication interface configured to receive the original image from the imaging device; and a processor coupled to the communication interface and configured to: executing an object detection module to detect the original image, and obtaining, by the object detection module, first position information, second position information, and third position information associated with a same portrait object from the original image, wherein the first position information corresponds to a head region, the second position information corresponds to a body region, and the third position information corresponds to a hand region; setting a valid determination range based on at least one of the first location information and the second location information; Based on the third position information, obtaining a hand position in the original image; In response to the hand position being within the valid determination range, executing a gesture recognition module; and In response to the hand position not being within the valid determination range, the gesture recognition module is not executed.