Gesture recognition prompting method and device, electronic equipment and storage medium

By judging the overlap between the gesture recognition frame and the image acquisition frame and the sound zone positioning of the voice control command, an out-of-bounds prompt is provided, which solves the problem of users being unable to recognize out-of-bounds gestures and improves the interactive experience of gesture recognition.

CN120848716APending Publication Date: 2025-10-28BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410526533.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-28
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Users cannot recognize whether the gesture is out of bounds in time, resulting in the need to adjust the gesture position multiple times, affecting the gesture recognition experience.

Method used

By determining whether the gesture recognition frame boundary coincides with the image acquisition frame boundary, combined with the voice zone positioning of the voice control command, a prompt message is output indicating that the target gesture is out of bounds, so that the user can adjust the gesture position in time.

Benefits of technology

It improves the interactive experience of gesture recognition and reduces the number of times users need to adjust the gesture position.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120848716A_ABST
    Figure CN120848716A_ABST
Patent Text Reader

Abstract

The invention discloses a gesture recognition prompting method and device, electronic equipment and a storage medium, and the method comprises the steps: determining a first boundary of a recognition frame of a target gesture after recognizing that the gesture in an image collection frame is the target gesture, and determining a first region corresponding to the target gesture; determining whether the target gesture is out of bound or not based on whether the first boundary and a second boundary of the image acquisition frame have an overlapping area or not; performing voice area positioning on the received voice control instruction, determining a second area corresponding to the voice control instruction, and cooperatively triggering the voice control instruction and the target gesture; and when it is determined that the first region is the same as the second region, outputting prompt information that the target gesture is out of the bound. After the target gesture is judged to be out of the bound, the prompt information of the out-of-bound target gesture is output, so that the user corresponding to the target gesture can adjust the position of the target gesture in time to trigger the corresponding operation, the frequency of adjusting the position of the target gesture by the user is reduced, and the interaction experience feeling of gesture recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing, and more particularly to a gesture recognition prompting method, apparatus, electronic device, and storage medium. Background Technology

[0002] Gesture recognition is a technology that uses computer vision to recognize and understand human hand gestures. It translates human hand movements into a form that computers can understand, thus enabling interaction with computers. Gesture recognition can be applied in multiple fields, including human-computer interaction, virtual reality, gaming, and smart homes. In human-computer interaction, gesture recognition can replace traditional mouse and keyboard input methods, allowing users to control computers or devices through gestures.

[0003] During gesture recognition, due to factors such as camera angle, hardware limitations, or user operating habits, the target gesture or part of the target gesture may not be displayed in the camera's image acquisition area, causing the target gesture to go out of bounds. When the target gesture goes out of bounds, the target device cannot perform the corresponding operation. However, since the user who made the target gesture cannot be unaware that their gesture is out of bounds, the user may need to adjust the position of the target gesture multiple times to trigger the operation, resulting in a poor user experience for gesture recognition. Summary of the Invention

[0004] This disclosure provides a gesture recognition prompting method, apparatus, electronic device, and storage medium. Its main purpose is to solve the problem that users making a target gesture cannot be unaware that their gesture is outside the bounds, and users may need to adjust the position of the target gesture multiple times to trigger the operation, resulting in a poor gesture recognition experience.

[0005] According to a first aspect of this disclosure, a gesture recognition prompting method is provided, comprising:

[0006] After recognizing the gesture within the image acquisition frame as the target gesture, the first boundary of the recognition frame of the target gesture is determined, and the first region corresponding to the recognized target gesture is determined, wherein the target gesture matches a preset gesture;

[0007] Based on whether there is an overlapping area between the first boundary and the second boundary of the image acquisition frame, it is determined whether the target gesture has gone out of bounds;

[0008] The received voice control command is localized to determine the second region corresponding to the voice control command, wherein the voice control command is triggered in conjunction with the target gesture;

[0009] If the first region is determined to be the same as the second region, a prompt message indicating that the target gesture has gone out of bounds is output.

[0010] Optionally, the process of recognizing the target gesture includes:

[0011] Determine whether the arm posture within the image acquisition frame is consistent with the preset arm posture;

[0012] If it is determined that the arm posture within the image acquisition frame is consistent with the preset arm posture, then the hand gesture corresponding to the arm posture within the image acquisition frame is subjected to hand gesture feature recognition to obtain the recognition result of whether the hand gesture to be recognized matches the preset hand gesture.

[0013] If the recognition result indicates that the gesture to be recognized matches a preset gesture, then the gesture to be recognized is identified as the target gesture.

[0014] Optionally, before determining whether the arm posture corresponding to the gesture within the image acquisition frame is consistent with a preset arm posture, the method includes:

[0015] The relationship between the gesture to be recognized, the arm corresponding to the gesture to be recognized, and the face of the target being recognized is determined by the pixel distance between the gesture to be recognized in the image acquisition frame, the arm corresponding to the gesture to be recognized, and the face of the target being recognized.

[0016] Based on the association, it is determined whether the gesture to be identified, the arm corresponding to the gesture to be identified, and the face of the target to be identified belong to the same target;

[0017] If it is determined that the gesture to be recognized, the arm corresponding to the gesture to be recognized, and the face of the recognition target belong to the same recognition target, the area where the recognition target is located is determined as the first area corresponding to the target gesture.

[0018] Optionally, the out-of-bounds type of the target gesture includes arm out-of-bounds;

[0019] The process of determining that the target gesture is out of bounds includes:

[0020] Determine whether there is an overlapping area between the first boundary and the second boundary of the image acquisition frame;

[0021] If it is determined that the first boundary and the second boundary do not overlap, then it is determined that the arm of the target gesture is out of bounds.

[0022] Optionally, the out-of-bounds type of the target gesture includes hand out-of-bounds movement;

[0023] The process of determining that the target gesture is out of bounds also includes:

[0024] If it is determined that the first boundary and the second boundary have an overlapping area, determine whether the first boundary and the second boundary completely overlap;

[0025] If it is determined that the first boundary and the second boundary do not completely coincide, it is determined that the hand of the target gesture is out of bounds.

[0026] Optionally, the output of the prompt message indicating that the target gesture has gone out of bounds includes:

[0027] Output the type of the target gesture going out of bounds, the identification information of the first region or the second region, and / or,

[0028] Based on the boundary-crossing type of the target gesture, determine the direction and distance of the target gesture within the second boundary;

[0029] The direction and distance of the adjusted target gesture are displayed on the display device.

[0030] According to a second aspect of this disclosure, a gesture recognition prompting device is provided, comprising:

[0031] The first determining unit is used to determine the first boundary of the recognition frame of the target gesture and the first region corresponding to the recognized target gesture after recognizing the gesture within the image acquisition frame as the target gesture, wherein the target gesture matches a preset gesture.

[0032] The second determining unit is used to determine whether the target gesture has gone out of bounds based on whether there is an overlapping area between the first boundary and the second boundary of the image acquisition frame;

[0033] A positioning unit is used to perform sound region positioning on the received voice control command and determine the second region corresponding to the voice control command, wherein the voice control command is triggered in conjunction with the target gesture;

[0034] The output unit is used to output a prompt message indicating that the target gesture has gone out of bounds when it is determined that the first region and the second region are the same.

[0035] Optionally, the device further includes an identification unit, the identification unit comprising:

[0036] The judgment module is used to determine whether the arm posture within the image acquisition frame is consistent with the preset arm posture;

[0037] The recognition module is used to perform gesture feature recognition on the gesture to be recognized corresponding to the arm posture in the image acquisition frame if it is determined that the arm posture in the image acquisition frame is consistent with the preset arm posture, and to obtain the recognition result of whether the gesture to be recognized matches the preset gesture.

[0038] The first determining module is used to determine the gesture to be recognized as the target gesture if the recognition result is determined to be a match between the gesture to be recognized and a preset gesture.

[0039] Optionally, the apparatus further includes a third determining unit, the third determining unit comprising:

[0040] The second determining module is used to determine the association between the hand gesture to be recognized, the arm corresponding to the hand gesture to be recognized, and the passenger's face by means of the pixel distance between the hand gesture to be recognized in the image acquisition frame, the arm corresponding to the hand gesture to be recognized, and the face of the recognition target before determining whether the arm gesture corresponding to the hand gesture in the image acquisition frame is consistent with the preset arm gesture.

[0041] The third determining module is used to determine, based on the association relationship, whether the gesture to be identified, the arm corresponding to the gesture to be identified, and the face of the target to be identified belong to the same target to be identified;

[0042] The fourth determining module is used to determine the area where the target is located as the first area corresponding to the target gesture when it is determined that the gesture to be recognized, the arm corresponding to the gesture to be recognized, and the face of the target to be recognized belong to the same target to be recognized.

[0043] Optionally, the out-of-bounds type of the target gesture includes arm out-of-bounds;

[0044] Optionally, the second determining unit includes:

[0045] The judgment module is used to determine whether there is an overlapping area between the first boundary and the second boundary of the image acquisition frame;

[0046] The fifth determining module is used to determine that the arm of the target gesture is out of bounds if it is determined that the first boundary and the second boundary do not overlap.

[0047] Optionally, the out-of-bounds type of the target gesture includes hand out-of-bounds movement;

[0048] Optionally, the second determining unit further includes:

[0049] The sixth determining module is used to determine whether the first boundary and the second boundary completely overlap if it is determined that there is an overlapping area between the first boundary and the second boundary;

[0050] The seventh determining module is used to determine that the hand part of the target gesture is out of bounds if it is determined that the first boundary and the second boundary do not completely coincide.

[0051] Optionally, the output unit includes:

[0052] The output module is used to output the type of the target gesture going out of bounds, the identification information of the first region or the second region, and / or,

[0053] The eighth determining module is used to determine the direction and distance of adjusting the target gesture within the second boundary based on the boundary type of the target gesture;

[0054] The display module is used to display the direction and distance of the adjusted target gesture on a display device.

[0055] According to a third aspect of this disclosure, an electronic device is provided, comprising:

[0056] At least one processor; and

[0057] A memory communicatively connected to the at least one processor; wherein,

[0058] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect above.

[0059] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method described in the first aspect above.

[0060] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described in the first aspect above.

[0061] The gesture recognition prompting method, apparatus, electronic device, and storage medium provided in this disclosure, after recognizing a gesture within an image acquisition frame as a target gesture, determine a first boundary of the recognition frame of the target gesture and a first region corresponding to the recognized target gesture, wherein the target gesture matches a preset gesture; determine whether the target gesture has exceeded the boundary based on whether there is an overlapping area between the first boundary and the second boundary of the image acquisition frame; perform voice region localization on the received voice control command to determine a second region corresponding to the voice control command, wherein the voice control command and the target gesture are triggered collaboratively; and output a prompt message indicating that the target gesture has exceeded the boundary when the first region and the second region are determined to be the same. This disclosure determines whether the target gesture has exceeded the boundary by whether there is an overlapping area between the first boundary and the second boundary of the image acquisition frame. When the target gesture is determined to be outside the boundary, it combines the voice region localization of the voice control command to determine that the first region and the second region are the same, and then outputs a prompt message indicating that the target gesture has exceeded the boundary corresponding to the first region or the second region. This allows the user corresponding to the target gesture to adjust the position of the target gesture in a timely manner to trigger the corresponding operation, reducing the number of times the user needs to adjust the position of the target gesture and improving the interactive experience of gesture recognition.

[0062] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0063] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0064] Figure 1 This is a flowchart illustrating a gesture recognition prompting method provided in an embodiment of the present disclosure;

[0065] Figure 2 A schematic diagram illustrating a gesture going out of bounds, provided as an embodiment of this disclosure;

[0066] Figure 3 A flowchart illustrating a gesture out-of-bounds judgment method provided in an embodiment of this disclosure;

[0067] Figure 4 A schematic diagram illustrating another gesture going out of bounds, provided as an embodiment of this disclosure;

[0068] Figure 5 A flowchart illustrating another gesture out-of-bounds judgment method provided in this embodiment of the present disclosure;

[0069] Figure 6 This is a schematic diagram of the structure of a gesture recognition prompting device provided in an embodiment of the present disclosure;

[0070] Figure 7 This is a schematic diagram of another gesture recognition prompting device provided in an embodiment of the present disclosure;

[0071] Figure 8 A schematic block diagram of an example electronic device 400 provided for embodiments of this disclosure. Detailed Implementation

[0072] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0073] The gesture recognition prompting method, apparatus, electronic device, and storage medium of this disclosure are described below with reference to the accompanying drawings.

[0074] Figure 1 This is a flowchart illustrating a gesture recognition prompting method provided in an embodiment of this disclosure. The method is applied in a vehicle, and the control method can be executed by an information prompting control device or equipment. This device or equipment can be configured in a server, processor, or main control chip; for example, it can be deployed on the vehicle's infotainment system side, domain controller side, etc. Figure 1 As shown, the method includes the following steps:

[0075] Step 101: After recognizing the gesture within the image acquisition frame as the target gesture, determine the first boundary of the recognition frame of the target gesture and determine the first region corresponding to the recognized target gesture, wherein the target gesture matches a preset gesture.

[0076] In one embodiment provided in this disclosure, the image acquisition frame is the visual range within which the camera acquires images of the vehicle interior. The target gesture is a gesture within the image acquisition frame that is recognized as matching a preset gesture. The preset gesture is a pre-set pointing gesture or other specific gesture; for example, the preset gesture could be pointing to the left window, pointing to the seat, etc. The recognition frame for the target gesture is the bounding box within the image acquisition frame that recognizes the target gesture, such as... Figure 2 As shown. The first area is the area where the passenger corresponding to the target gesture is located. For example, the first area can be the driver's area in the smart cockpit, or the rear left side area in the smart cockpit. The specific embodiments of this application are not limited.

[0077] To better understand the image acquisition frame and the target gesture recognition frame, Figure 2This is a schematic diagram illustrating the principle of gesture going out of bounds according to an embodiment of the present disclosure, such as... Figure 2 As shown, the image acquisition box is box b, and the recognition box is box a.

[0078] Step 102: Determine whether the target gesture has gone out of bounds based on whether there is an overlapping area between the first boundary and the second boundary of the image acquisition frame.

[0079] In practical applications, because the in-vehicle camera is installed at the front of the cabin, passengers in the first row often make gestures that match preset hand signals, such as pointing to the left or right front, which can easily fall outside the camera's recognition range. Furthermore, due to the close proximity of the passenger to the camera, the recognition frame for the target gesture is usually near the lower boundary of the image capture frame. Please refer to [further details omitted]. Figure 2 To address the issue of determining whether a target gesture has crossed the boundary, one approach, but not limited to this one, is to determine whether the target gesture has crossed the boundary by judging whether there is an overlapping area between the first boundary of the recognition box and the second boundary of the image acquisition box.

[0080] In some embodiments, to determine whether the first boundary and the second boundary coincide, a first distance threshold between the upper boundary of the first boundary and the lower boundary of the second boundary, and a second distance threshold between the left and right boundaries of the first boundary and the left and right boundaries of the second boundary can be preset. The coincidence of the first boundary and the second boundary can be determined by judging whether the distances between each boundary of the first boundary and each boundary of the second boundary satisfy the first distance threshold and the second distance threshold. For example, please refer to... Figure 2 If the first distance threshold between the upper boundary of the first boundary and the lower boundary of the second boundary is set to be less than 0 cm, then when the distance between the upper boundary of the first boundary and the lower boundary of the second boundary is -5 cm, it is determined that the first boundary and the second boundary do not coincide. In this case, the first boundary is completely outside the second boundary. Figure 4 As shown. Conversely, if the distance between the upper boundary of the first boundary and the lower boundary of the second boundary is 10cm, then the first boundary and the second boundary are determined to coincide.

[0081] Step 103: Perform sound region localization on the received voice control command to determine the second region corresponding to the voice control command, wherein the voice control command is triggered in conjunction with the target gesture.

[0082] In one embodiment provided in this disclosure, since the voice control command and the target gesture are triggered collaboratively, it is necessary to determine whether the voice control command and the target gesture were issued by the same passenger. Specifically, it is necessary to determine whether the first region corresponding to the target gesture and the second region corresponding to the voice control command are the same, where the second region is the sound region from which the voice control command was issued. Sound region localization refers to the process of segmenting sound pickup when a scene contains multiple sound sources in order to extract the signal of a specific sound source. For example, in a vehicle interior scenario, due to the low reverberation, small noise distribution range, and close proximity of the interior environment, it is suitable to segment sound pickup of sound sources within the vehicle. This involves speech enhancement targeting specific sound regions, separating the speaker's voice from those specific regions to meet different practical application needs and suppressing engine noise, tire noise, music noise, etc. The region corresponding to the voice control command is determined through sound region localization.

[0083] To better understand the received voice control commands by locating the voice region, one approach can be adopted, but is not limited to, first determining the voice wake-up region of the voice control quality using a preset sound sensor, and then determining the second region corresponding to the voice wake-up region.

[0084] Step 104: If it is determined that the first region and the second region are the same, output a prompt message that the target gesture has gone out of bounds.

[0085] When the first region and the second region are determined to be the same, it indicates that the passenger making the target gesture and the passenger issuing the voice control command are the same person, and it can be determined that the passenger intends to operate a certain device in the vehicle. However, because the target gesture is out of bounds, the operation corresponding to the target gesture cannot be triggered. Therefore, when the first region and the second region are determined to be the same, a prompt message indicating that the target gesture is out of bounds is output. The prompt message may take, but is not limited to, the following forms: displaying a gesture out of bounds prompt animation on the in-vehicle display terminal, and simultaneously outputting a voice prompt message indicating that the gesture is out of bounds using Text-to-Speech (TTS) speech synthesis technology. The prompt message includes the region corresponding to the target gesture, so as to prompt the user to move the target gesture into the image acquisition frame, so that the target gesture can be fully recognized and thus trigger the corresponding operation.

[0086] The gesture recognition prompting method provided in this disclosure, after recognizing a gesture within an image acquisition frame as a target gesture, determines a first boundary of the recognition frame of the target gesture and a first region corresponding to the recognized target gesture, wherein the target gesture matches a preset gesture; based on whether there is an overlapping region between the first boundary and the second boundary of the image acquisition frame, it determines whether the target gesture has exceeded the boundary; performs voice region localization on the received voice control command to determine a second region corresponding to the voice control command, wherein the voice control command and the target gesture are triggered collaboratively; if it is determined that the first region and the second region are the same, a prompt message indicating that the target gesture has exceeded the boundary is output. This disclosure determines whether the target gesture has exceeded the boundary by whether there is an overlapping region between the first boundary and the second boundary of the image acquisition frame. If the target gesture is determined to be outside the boundary, it combines the voice region localization of the voice control command to determine that the first region and the second region are the same, and then outputs a prompt message indicating that the target gesture has exceeded the boundary corresponding to the first region or the second region. This allows the user corresponding to the target gesture to adjust the position of the target gesture in a timely manner to trigger the corresponding operation, reducing the number of times the user needs to adjust the position of the target gesture and improving the interactive experience of gesture recognition.

[0087] In one embodiment provided in this disclosure, recognizing all gestures appearing inside the vehicle would consume the computational resources of the gesture recognition system. To reduce the consumption of computational resources, before performing gesture recognition, it can be determined whether the passenger's arm posture is consistent with a preset posture, thereby reducing the recognition of invalid gestures and reducing resource consumption. Specific methods for determining the arm posture can be, but are not limited to, using a three-dimensional human posture estimation model to estimate the three-dimensional posture and angle information of the passenger's arm within the image acquisition frame. By judging the angle of the arm, it can be determined whether the passenger's arm posture is consistent with a preset arm posture. The preset arm posture can be set according to the user's arm-raising habits, such as an upright arm or other arm postures at different angles. For example, some passengers prefer to raise their arms to make a target gesture, so the preset arm posture can be set such that the angle between the arm and the horizontal line is between 75° and 105°, etc. If the arm posture within the image acquisition frame is determined to be consistent with the preset arm posture, gesture recognition is performed on the gesture within the image acquisition frame to obtain a recognition result indicating whether the gesture within the image acquisition frame matches the preset gesture. The gesture recognition can be performed, but is not limited to, by using a preset gesture detection model to recognize the gesture to be recognized within the image acquisition frame, or by using gesture key point capture, etc. The specific gesture recognition method is not limited in this application embodiment.

[0088] In some embodiments, in order to determine the first region corresponding to the target gesture, it is necessary to determine the association between the gesture to be recognized, the arm corresponding to the gesture to be recognized, and the face of the target being recognized by using the pixel distance between the gesture to be recognized in the image acquisition frame, the arm corresponding to the gesture to be recognized, and the face of the target being recognized. Here, the target being recognized refers to the user who makes the target gesture inside the vehicle. Then, based on the association, it is determined whether the gesture to be recognized, the arm corresponding to the gesture to be recognized, and the face of the target being recognized belong to the same target being recognized. If it is determined that the gesture to be recognized, the arm corresponding to the gesture to be recognized, and the face of the target being recognized belong to the same target being recognized, the region where the target being recognized is located is determined as the first region corresponding to the target gesture.

[0089] In one embodiment provided in this disclosure, the preset gesture is a gesture pre-set in the gesture recognition system. If the recognition result is determined to be a match between the gesture to be recognized within the image acquisition frame and the preset gesture, then the gesture to be recognized is determined to be the target gesture. Further determination is needed to determine whether the target gesture is a valid gesture. A valid gesture refers to a gesture that does not go out of bounds. If the gesture to be recognized is a valid gesture, it is triggered in conjunction with a voice control command to perform the operation corresponding to the target gesture. For example, if the gesture voice control command is "open the window," and the target gesture is "point to the right," and the target gesture does not go out of bounds, then the command "open the right window" is executed directly. If the recognition result is determined to be that the target gesture goes out of bounds, for example, when the passenger in the front seat says "turn on the seat heating," they simultaneously point to their seat. However, because the user's gesture is not captured by the camera, or the camera only captures a portion of the gesture (e.g., the camera only captures the image of the palm and not the pointing image of the fingers), the system cannot determine which area of ​​the seat the user needs to heat. Therefore, the system cannot obtain an accurate control command and cannot execute the next operation.

[0090] In some embodiments, after recognizing the gesture within the image acquisition frame as the target gesture, the position information of the recognition frame of the target gesture can be obtained by, but is not limited to, the following methods: based on the installation angle of the in-vehicle camera and the three-dimensional objects in the vehicle cabin, the three-dimensional spatial position of the target gesture in the vehicle is estimated by a monocular, binocular, or time-of-flight (TOF) camera, and the position information of the recognition frame of the target gesture is output.

[0091] In some embodiments, the target gesture going out of bounds can be divided into a hand part going out of bounds state and an arm going out of bounds state. To better understand the determination of the target gesture going out of bounds, this disclosure provides a flowchart illustrating a method for judging gesture going out of bounds, as follows: Figure 3As shown, the method includes the following steps:

[0092] Step 201: Determine whether there is an overlapping area between the first boundary and the second boundary of the image acquisition frame.

[0093] In one embodiment provided in this disclosure, for a better explanation of the overlapping region, please refer to [the relevant documentation / reference needed]. Figure 2 and Figure 4 ,exist Figure 4 In, the first boundary ( Figure 2 (a) and the second boundary ( Figure 2 The middle box (b) does not overlap, therefore there is no overlapping area. Figure 2 In, the first boundary ( Figure 2 (a) and the second boundary ( Figure 2 The middle frame (b) overlaps, and the overlapping area is as follows: Figure 2 As shown.

[0094] Step 202: If it is determined that the first boundary and the second boundary do not overlap, then it is determined that the arm of the target gesture is out of bounds.

[0095] After determining that the first boundary and the second boundary do not have the overlapping area, as Figure 4 As shown, this indicates that no target gesture was detected in the image acquisition frame, but the presence of the target gesture could be sensed by other sensors inside the vehicle. This suggests that an arm has gone out of bounds. To further determine whether the arm has indeed gone out of bounds and to output accurate prompt information, it is also necessary to determine the arm point of the target gesture. Figure 4 The distance between point B in the image and the left and right boundaries of the image acquisition frame is used to determine whether the arm has gone out of bounds.

[0096] Step 203: If it is determined that the first boundary and the second boundary have an overlapping area, determine whether the first boundary and the second boundary are completely overlapping.

[0097] In some embodiments, if it is determined that the first boundary and the second boundary have an overlapping area, it indicates that there are two situations: the hand goes out of bounds and the target gesture does not go out of bounds. If it is further determined that the first boundary and the second boundary completely overlap, it indicates that the target gesture does not go out of bounds, and the operation corresponding to the target gesture can be directly executed. If the hand goes out of bounds, the subsequent operation cannot be executed.

[0098] Step 204: If it is determined that the first boundary and the second boundary do not completely coincide, the hand of the target gesture is determined to be out of bounds.

[0099] In practical applications, when it is determined that the first boundary and the second boundary do not completely coincide, the hand of the target gesture is determined to be out of bounds, such as... Figure 2 As shown, to further determine whether the first boundary and the second boundary completely coincide, a first distance threshold between the upper boundary of the first boundary and the lower boundary of the second boundary, and a second distance threshold between the left and right boundaries of the first boundary and the left and right boundaries of the second boundary can be preset. By determining whether the distances between each boundary of the first boundary and each boundary of the second boundary satisfy the first distance threshold and the second distance threshold, it can be determined whether the first boundary and the second boundary completely coincide. For example, please refer to [reference needed]. Figure 2 If the first distance threshold between the upper boundary of the first boundary and the lower boundary of the second boundary is set to be less than 0cm, and the distance between the upper boundary of the first boundary and the lower boundary of the second boundary is 10cm, then the first boundary and the second boundary are determined to coincide.

[0100] To better understand the gesture out-of-bounds judgment method provided in the above embodiments Figure 5 This is a flowchart illustrating another method for determining hand gesture boundaries according to an embodiment of this application. It includes steps for gesture recognition and determining the associated relationship.

[0101] The target gesture out-of-bounds judgment method provided in the above embodiments can accurately judge the out-of-bounds state of the target gesture, output accurate prompt information so that the user can adjust the position of the target gesture, and improve the user's interactive experience of gesture recognition.

[0102] In summary, the embodiments disclosed herein achieve the following effects:

[0103] 1. It accurately judges the out-of-bounds state of the target gesture and outputs accurate prompts to help users adjust the position of the target gesture, thereby improving the user's interactive experience of gesture recognition.

[0104] 2. Reduce the consumption of computing resources and improve the efficiency of gesture recognition.

[0105] Corresponding to the gesture recognition prompting method described above, this invention also proposes a gesture recognition prompting device. Since the device embodiments of this invention correspond to the method embodiments described above, details not disclosed in the device embodiments can be referred to in the method embodiments described above, and will not be repeated here.

[0106] Figure 6 This is a schematic diagram of the structure of a gesture recognition prompting device provided in an embodiment of this disclosure, as shown below. Figure 6 As shown, it includes: a first determining unit 31, a second determining unit 32, a positioning unit 33, and an output unit 34.

[0107] The first determining unit 31 is used to determine the first boundary of the recognition frame of the target gesture and the first region corresponding to the recognized target gesture after recognizing the gesture within the image acquisition frame as the target gesture, wherein the target gesture matches a preset gesture.

[0108] The second determining unit 32 is used to determine whether the target gesture has gone out of bounds based on whether there is an overlapping area between the first boundary and the second boundary of the image acquisition frame;

[0109] The positioning unit 33 is used to perform sound region positioning on the received voice control command and determine the second region corresponding to the voice control command, wherein the voice control command is triggered in conjunction with the target gesture;

[0110] The output unit 34 is used to output a prompt message indicating that the target gesture has gone out of bounds when it is determined that the first region and the second region are the same.

[0111] The gesture recognition prompting device provided in this disclosure, after recognizing a gesture within an image acquisition frame as a target gesture, determines the first boundary of the recognition frame of the target gesture and the first region corresponding to the recognized target gesture, wherein the target gesture matches a preset gesture; based on whether there is an overlapping area between the first boundary and the second boundary of the image acquisition frame, it determines whether the target gesture has gone out of bounds; it performs voice region localization on the received voice control command to determine the second region corresponding to the voice control command, wherein the voice control command and the target gesture are triggered collaboratively; if it is determined that the first region and the second region are the same, it outputs a prompt message indicating that the target gesture has gone out of bounds. This disclosure determines whether the target gesture has gone out of bounds by whether there is an overlapping area between the first boundary and the second boundary of the image acquisition frame. If the target gesture has gone out of bounds, it combines the voice region localization of the voice control command to determine that the first region and the second region are the same, and then outputs a prompt message indicating that the target gesture has gone out of bounds corresponding to the first region or the second region, so that the user corresponding to the target gesture can adjust the position of the target gesture in time to trigger the corresponding operation, reducing the number of times the user adjusts the position of the target gesture and improving the interactive experience of gesture recognition.

[0112] Furthermore, in one possible implementation of this embodiment, such as Figure 7 As shown, the device further includes an identification unit 35, which includes:

[0113] The judgment module 351 is used to determine whether the arm posture within the image acquisition frame is consistent with the preset arm posture;

[0114] The recognition module 352 is used to perform gesture feature recognition on the gesture to be recognized corresponding to the arm posture in the image acquisition frame if it is determined that the arm posture in the image acquisition frame is consistent with the preset arm posture, and to obtain the recognition result of whether the gesture to be recognized matches the preset gesture.

[0115] The first determining module 353 is used to determine the gesture to be recognized as the target gesture if the recognition result is determined to be a match between the gesture to be recognized and a preset gesture.

[0116] Furthermore, in one possible implementation of this embodiment, such as Figure 7 As shown, the device further includes a third determining unit 36, which includes:

[0117] The second determining module 361 is used to determine the association between the hand gesture to be recognized, the arm corresponding to the hand gesture to be recognized, and the passenger's face by means of the pixel distance between the hand gesture to be recognized in the image acquisition frame, the arm corresponding to the hand gesture to be recognized, and the face of the recognition target before determining whether the arm gesture corresponding to the hand gesture in the image acquisition frame is consistent with the preset arm gesture.

[0118] The third determining module 362 is used to determine, based on the association relationship, whether the gesture to be identified, the arm corresponding to the gesture to be identified, and the face of the target to be identified belong to the same target to be identified;

[0119] The fourth determining module 363 is used to determine the area where the target is located as the first area corresponding to the target gesture when it is determined that the gesture to be recognized, the arm corresponding to the gesture to be recognized, and the face of the target to be recognized belong to the same target to be recognized.

[0120] Optionally, the out-of-bounds type of the target gesture includes arm out-of-bounds;

[0121] Furthermore, in one possible implementation of this embodiment, such as Figure 7 As shown, the second determining unit 32 includes:

[0122] The judgment module 321 is used to determine whether there is an overlapping area between the first boundary and the second boundary of the image acquisition frame;

[0123] The fifth determining module 322 is used to determine that the arm of the target gesture is out of bounds if it is determined that the first boundary and the second boundary do not overlap.

[0124] Optionally, the out-of-bounds type of the target gesture includes hand out-of-bounds movement;

[0125] Furthermore, in one possible implementation of this embodiment, such as Figure 7 As shown, the second determining unit 32 further includes:

[0126] The sixth determining module 323 is used to determine whether the first boundary and the second boundary completely overlap if it is determined that there is an overlapping area between the first boundary and the second boundary;

[0127] The seventh determining module 324 is used to determine that the hand part of the target gesture is out of bounds if it is determined that the first boundary and the second boundary are not completely coincident.

[0128] Furthermore, in one possible implementation of this embodiment, such as Figure 7 As shown, the output unit 34 includes:

[0129] Output module 341 is used to output the type of the target gesture going out of bounds, the identification information of the first area or the second area, and / or,

[0130] The eighth determining module 342 is used to determine the direction and distance of adjusting the target gesture within the second boundary based on the boundary type of the target gesture;

[0131] The display module 343 is used to display the direction and distance of the adjusted target gesture on a display device.

[0132] It should be noted that the foregoing explanation of the method embodiments also applies to the apparatus of this embodiment, and the principle is the same, so it is not limited in this embodiment.

[0133] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0134] Figure 8 A schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0135] like Figure 8As shown, device 400 includes a computing unit 401, which can perform various appropriate actions and processes based on a computer program stored in ROM (Read-Only Memory) 402 or a computer program loaded from storage unit 408 into RAM (Random Access Memory) 403. RAM 403 may also store various programs and data required for the operation of device 400. The computing unit 401, ROM 402, and RAM 403 are interconnected via bus 404. I / O (Input / Output) interface 405 is also connected to bus 404.

[0136] Multiple components in device 400 are connected to I / O interface 405, including: input unit 406, such as keyboard, mouse, etc.; output unit 407, such as various types of monitors, speakers, etc.; storage unit 408, such as disk, optical disk, etc.; and communication unit 409, such as network card, modem, wireless transceiver, etc. Communication unit 409 allows device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0137] The computing unit 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphics Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any suitable processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above, such as the gesture recognition prompting method. For example, in some embodiments, the gesture recognition prompting method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on device 400 via ROM 402 and / or communication unit 409. When the computer program is loaded into RAM 403 and executed by the computing unit 401, one or more steps of the methods described above may be performed. Alternatively, in other embodiments, the computing unit 401 may be configured to perform the aforementioned gesture recognition prompting method by any other suitable means (e.g., by means of firmware).

[0138] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System-on-Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0139] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0140] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0141] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0142] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include LANs (Local Area Networks), WANs (Wide Area Networks), the Internet, and blockchain networks.

[0143] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0144] It's important to note that artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0145] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0146] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A gesture recognition prompting method, characterized in that, include: After recognizing the gesture within the image acquisition frame as the target gesture, the first boundary of the recognition frame of the target gesture is determined, and the first region corresponding to the recognized target gesture is determined, wherein the target gesture matches a preset gesture; Based on whether there is an overlapping area between the first boundary and the second boundary of the image acquisition frame, it is determined whether the target gesture has gone out of bounds; The received voice control command is localized to determine the second region corresponding to the voice control command, wherein the voice control command is triggered in conjunction with the target gesture; If the first region is determined to be the same as the second region, a prompt message indicating that the target gesture has gone out of bounds is output.

2. The method according to claim 1, characterized in that, The process of recognizing the target gesture includes: Determine whether the arm posture within the image acquisition frame is consistent with the preset arm posture; If it is determined that the arm posture within the image acquisition frame is consistent with the preset arm posture, then the hand gesture corresponding to the arm posture within the image acquisition frame is subjected to hand gesture feature recognition to obtain the recognition result of whether the hand gesture to be recognized matches the preset hand gesture. If the recognition result indicates that the gesture to be recognized matches a preset gesture, then the gesture to be recognized is identified as the target gesture.

3. The method according to claim 2, characterized in that, Before determining whether the arm posture corresponding to the gesture within the image acquisition frame is consistent with a preset arm posture, the method includes: The relationship between the gesture to be recognized, the arm corresponding to the gesture to be recognized, and the face of the target to be recognized is determined by the pixel distance between the gesture to be recognized in the image acquisition frame, the arm corresponding to the gesture to be recognized, and the face of the target to be recognized. Based on the association, it is determined whether the gesture to be identified, the arm corresponding to the gesture to be identified, and the face of the target to be identified belong to the same target; If it is determined that the gesture to be recognized, the arm corresponding to the gesture to be recognized, and the face of the recognition target belong to the same recognition target, the area where the recognition target is located is determined as the first area corresponding to the target gesture.

4. The method according to claim 1, characterized in that, The out-of-bounds types of the target gesture include arm out of bounds; The process of determining that the target gesture is out of bounds includes: Determine whether there is an overlapping area between the first boundary and the second boundary of the image acquisition frame; If it is determined that the first boundary and the second boundary do not overlap, then it is determined that the arm of the target gesture is out of bounds.

5. The method according to claim 4, characterized in that, The out-of-bounds types of the target gesture include hand out-of-bounds; The process of determining that the target gesture is out of bounds also includes: If it is determined that the first boundary and the second boundary have an overlapping area, determine whether the first boundary and the second boundary completely overlap; If it is determined that the first boundary and the second boundary do not completely coincide, it is determined that the hand of the target gesture is out of bounds.

6. The method according to claim 5, characterized in that, The output of the prompt message indicating that the target gesture has gone out of bounds includes: Output the type of the target gesture going out of bounds, the identification information of the first region or the second region, and / or, Based on the boundary-crossing type of the target gesture, determine the direction and distance of the target gesture within the second boundary; The direction and distance of the adjusted target gesture are displayed on the display device.

7. A gesture recognition prompting device, characterized in that, include: The first determining unit is used to determine the first boundary of the recognition frame of the target gesture and the first region corresponding to the recognized target gesture after recognizing the gesture within the image acquisition frame as the target gesture, wherein the target gesture matches a preset gesture. The second determining unit is used to determine whether the target gesture has gone out of bounds based on whether there is an overlapping area between the first boundary and the second boundary of the image acquisition frame; A positioning unit is used to perform sound region positioning on the received voice control command and determine the second region corresponding to the voice control command, wherein the voice control command is triggered in conjunction with the target gesture; The output unit is used to output a prompt message indicating that the target gesture has gone out of bounds when it is determined that the first region and the second region are the same.

8. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.

10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-6.