Gesture recognition-based permission verification and device control method in smart cockpit scenarios

By setting the palm to stretched as the initial state in the car cabin, learning the association between gesture travel and curvature data, and using the fitting curve and grid number of the thumb fingertip point, the gesture recognition process is simplified, solving the problem of low accuracy in recognizing subtle gesture states, and achieving efficient device control and permission verification.

CN119357937BActive Publication Date: 2025-09-26HANGZHOU YUANSHITAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411395765.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-08
Publication Date
2025-09-26
Estimated Expiration
2044-10-08

AI Technical Summary

Technical Problem

In a car cockpit, when the distinction between subtle gesture states is low, the recognition accuracy of image recognition is greatly reduced. In addition, the complex neural network gesture recognition algorithm is time-consuming and has high performance requirements, making it difficult to improve device control accuracy while ensuring gesture recognition accuracy and efficiency.

Method used

By setting the palm stretched as the initial state, learning the correlation between gesture stroke and curvature data under gesture commands, and using the stroke fitting curve of the thumb fingertip point and the number of grid number jumps, the gesture recognition process is simplified, the use of complex algorithms is reduced, and accurate permission verification and device control are achieved.

Benefits of technology

It improves the accuracy and efficiency of gesture recognition, reduces the possibility of misoperation, and ensures the accuracy and response speed of device control without using complex algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119357937B_ABST
    Figure CN119357937B_ABST
Patent Text Reader

Abstract

The present invention discloses a permission verification and device control method based on gesture recognition in a smart cockpit scenario. By constructing a correlation relationship between thumb travel and gesture curvature in each thumb travel segment, the correlation relationship is used to accurately characterize the subtle feature differences between different gesture states corresponding to each captured gesture instruction. Based on these subtle feature differences, the suspected gesture instructions corresponding to the gesture state data collected in real time can be accurately identified, and the similarity between the first change state and the second change state is verified through step S5, thereby improving the accuracy of gesture state recognition based on these subtle feature differences, making it possible to avoid the use of complex gesture recognition algorithms such as neural network model learning, and at the same time improving the efficiency of gesture recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart cockpits, and in particular to a method for authority verification and device control based on gesture recognition in a smart cockpit scenario. Background Art

[0002] Controlling devices to execute corresponding commands using different gesture states has been widely used in various fields. For example, in automotive cockpits, image recognition is used to identify different gesture states and then control the device to execute the corresponding commands. However, to ensure gesture recognition accuracy, image recognition methods must have high differentiation between different gesture states corresponding to different commands. For example, a "clenched fist" gesture can control the device to execute one command, while an "open palm" gesture can control the device to execute a second command. However, when the differentiation between different gesture states is small, such as controlling the device to execute the first command with a "clenched thumb and index finger" gesture in an open palm state, and the second command with a "clenched thumb and middle finger" gesture in an open palm state, the accuracy of image recognition methods will be greatly reduced. To improve the recognition accuracy of these subtle gesture states, gesture state model learning methods such as neural networks have been studied in the field. Although recognition accuracy has improved, the results are still limited. Moreover, the more complex the gesture state recognition algorithm, the more time-consuming the recognition process is, and the higher the performance requirements for the vehicle computer system are, which is difficult to accept in the field. Therefore, in the gesture command control scenario where the initial state of the gesture is controlled by stretching the palm, and the gesture change state is controlled by the clenched state of one or more fingers, ensuring the recognition accuracy and efficiency of the gesture state while reducing the use of complex algorithms has become a technical problem that needs to be urgently solved in this field. Summary of the Invention

[0003] The present invention reduces the use of complex gesture recognition algorithms in gesture command control scenarios where the initial state of a gesture is controlled by stretching the palm and the state of a gesture change is controlled by clenching one or more fingers. This improves the efficiency of gesture state recognition while ensuring the accuracy of gesture state recognition, and aims to improve the accuracy and response speed of device control. A method for permission verification and device control based on gesture recognition in a smart cockpit scenario is provided.

[0004] To achieve this object, the present invention adopts the following technical solutions:

[0005] Provided is a method for permission verification and device control based on gesture recognition in a smart cockpit scenario, comprising the following steps:

[0006] S1, setting the initial state of the slide to an extended state to cater to the user's extended palm, and then learning the user's gesture state characteristics under different gesture commands, the learning content including: gesture stroke data and gesture curvature data under each gesture command, and the correlation between the gesture stroke data and the gesture curvature data generated in the same stroke segment;

[0007] S2, identifying suspected gesture commands based on gesture state data collected and processed in real time from the user, and then obtaining gesture learning data learned by the user for the suspected gesture commands, including a stroke learning fitting curve pre-drawn based on the gesture stroke data learned for the suspected gesture commands;

[0008] S3, drawing a travel fitting curve of the user's thumb tip based on the gesture state data collected in real time;

[0009] S4, determining whether the travel fitting curve is similar to the travel learning fitting curve,

[0010] If yes, go to step S5;

[0011] If not, an alarm is issued and the user is prompted that he or she does not have the authority to control the corresponding device with the suspected gesture instruction, and then the process returns to step S2;

[0012] S5, determining whether a first change state of the user's gesture recorded in the gesture state data in each first travel segment of the thumb tip point is similar to a second change state of the gesture recorded in the gesture learning data in the corresponding second travel segment,

[0013] If so, the suspected gesture command is used to control the operation of the corresponding device;

[0014] If not, the user is prompted that the control gesture is invalid, and then the process returns to step S2.

[0015] Preferably, in step S1, the method of setting the initial state of the slide to the stretched state comprises the steps of:

[0016] A1, the user presses the thumb button in the initial position on the gesture learning device with the tip of the thumb, so that the thumb button is fixed to the slide;

[0017] A2, the thumb tip pushes the thumb button to drive the slide along the slide rail provided in the slide until the user's middle finger tip shifts to the press button provided on the gesture learning device when the palm is stretched;

[0018] A3, press the push button to fix the slide relative to the slide rail, and then press the thumb button a second time to allow the thumb button to slide freely in the slide.

[0019] Preferably, the slide covers the thumb button with a soft material, and the thumb button is provided with a protrusion for pressing against the thumb.

[0020] Preferably, in step S1, the method for learning the user's gesture state under different gesture instructions includes the steps of:

[0021] B1, after selecting the minimum rectangle of the hand key points in the palm image in the stretched state, the selected rectangle is discretized into grids and the discrete grids are numbered to obtain the numbered grid where each hand key point is located;

[0022] B2. Collect each frame of gesture state image under each gesture instruction at a specified frequency, and calculate the position change data of the gesture state in each frame image under the corresponding travel segment, wherein the first position change data of the thumb fingertip point in each travel segment constitutes the gesture travel data learned under the gesture instruction, and the second position change data of each hand key point in each travel segment constitutes the gesture curvature data learned under the gesture instruction.

[0023] As a preference, let the sequence s consisting of several frames of gesture state images collected for the specified gesture instruction i in step B2 be i ={p1, p2, ..., p j ,…,p m}, p j represents the j-th frame gesture state image, j=1, 2, ..., m, where m represents the image frame number, and the first position change data includes: a first relative coordinate position of the thumb tip point and the lower right vertex of the initial numbered grid where the thumb tip point is located in the stretched state at the j+1 time point at the end of the stroke segment, and a first numbered jump number of the first fingertip point numbered grid where the thumb tip point is located at the j+1 time point at the end of the stroke segment relative to the second fingertip point numbered grid where the thumb tip point is located at the j time point at the start of the stroke segment;

[0024] The second position change data includes: the second relative coordinate position of the hand key point and the upper left vertex of the starting number grid where the hand key point is located in the stretched state at the j+1 time point at the end of the stroke segment, and the second number jump number of the first hand key point number grid where the hand key point is located at the j+1 time point at the end of the stroke segment relative to the second hand key point number grid where the hand key point is located at the j time point at the start of the stroke segment.

[0025] Preferably, in step S2, the method for collecting the gesture state data of the user in real time includes the steps of:

[0026] C1, using the selected rectangle formed by the grid discretization and numbering in step B1 as a virtual projection tracking to obtain the first frame of gesture image captured in quasi-real time;

[0027] C2, starting from tracking the first frame of the gesture image, capturing a sequence of continuous frame images in real time at a preset frequency until a continuous capture termination condition is reached, and using the position change data of the gesture state in each frame of the continuous frame image sequence obtained after the capture in the corresponding travel segment as the gesture state data collected in real time from the user;

[0028] The continuous acquisition termination condition is that the relative distance between the thumb tip point in each frame in the continuous frame image sequence and the lower right vertex of the initial numbered grid in the selection rectangle decreases as the number of acquisition frames increases and reaches a minimum relative distance.

[0029] Preferably, in step C1, the method for tracking the first frame of gesture image collected in quasi-real time comprises the steps of:

[0030] C11, adjusting the projection posture of the selection rectangle so that the thumb tip point recognized in real time falls into the initial numbered grid in the minimum rectangular frame projected parallel to the gesture;

[0031] C12, determining whether the middle fingertip point recognized in real time falls within the middle fingertip landing area grid in the projected frame selection rectangle,

[0032] If so, it is determined that the first frame of the gesture image is successfully tracked;

[0033] If not, return to step C11.

[0034] Preferably, in step S2, the method for identifying the suspected gesture instruction includes the steps of:

[0035] D1, calculating the position change of the thumb tip point in the last frame relative to the first frame in the continuous frame image sequence, and calculating the number of grid number jumps of each hand key point of the user in the last frame number grid relative to the first frame number grid in the first frame, and then solving the average number of grid number jumps corresponding to all the hand key points;

[0036] D2, from the gesture instruction library that has completed gesture state learning, match whether there is a gesture instruction whose first deviation between the position change scalar and the position change is less than a preset first deviation threshold, and whose second deviation between the standard mean of the number of beats and the mean of the number of beats is less than a preset second deviation threshold,

[0037] If yes, identifying the matched gesture instruction as the suspected gesture instruction intended to be represented by the continuous frame image sequence;

[0038] If not, it is determined that the recognition of the suspected gesture instruction of the continuous frame image sequence has failed, and the user is prompted to learn the instruction.

[0039] Preferably, the continuous frame image sequence is divided into a plurality of first run segments, and a continuous frame image subsequence corresponding to each first run segment is obtained. In step S5, the first change state corresponding to the first run segment includes: a first grid number jump number and a first position change amount of the position of each hand key point recognized by the user in real time in the last frame relative to the position of the first frame in the continuous frame image subsequence; the second change state corresponding to the second run segment having a corresponding relationship with the first run segment includes: a second grid number jump number and a second position change amount of the position of each hand key point learned by the user for the second run segment in the suspected finger instruction in the last frame relative to the position of the first frame in the continuous frame learning subsequence;

[0040] The corresponding relationship between the first travel segment and the second travel segment is: the first travel segment and the second travel segment have the same data collection start frequency and the same data collection end frequency.

[0041] Preferably, the method for determining similarity in step S5 includes the steps of:

[0042] D1, calculating a first mean of the number of jitters of the first grid numbers of each of the hand key points in each of the first travel segments, and then calculating a second mean of the first mean values ​​associated with each of the first travel segments; and calculating a third mean of the number of jitters of the second grid numbers in each of the second travel segments, and then calculating a fourth mean of the third mean values ​​associated with each of the second travel segments;

[0043] D2, determining whether the deviation between the second mean and the fourth mean is less than a preset third deviation threshold,

[0044] If yes, go to step D3;

[0045] If not, determining that the first change state and the second change state are not similar;

[0046] D3, calculating a fifth mean of the first position changes of each of the hand key points in each of the first travel segments, and then calculating a sixth mean of the fifth mean values ​​associated with each of the first travel segments; and calculating a seventh mean of the second position changes of each of the second travel segments, and then calculating an eighth mean of the seventh mean values ​​associated with each of the second travel segments;

[0047] D4, determining whether the deviation between the sixth mean and the eighth mean is less than a preset fourth deviation threshold,

[0048] If so, determining that the first change state and the second change state have similarity;

[0049] If not, it is determined that the first change state and the second change state have no similarity.

[0050] The present invention has the following beneficial effects:

[0051] 1. By establishing a correlation between thumb travel and gesture curvature in each thumb travel segment, this correlation accurately characterizes the subtle characteristic differences between different gesture states corresponding to each captured gesture instruction. Based on these subtle characteristic differences, the suspected gesture instruction corresponding to the real-time gesture state data can be accurately identified. In addition, the similarity between the first change state and the second change state is verified in step S5. This improves the accuracy of gesture recognition based on these subtle characteristic differences, makes it possible to avoid the use of complex gesture recognition algorithms such as neural network model learning, and improves the efficiency of gesture recognition.

[0052] 2. By utilizing the association between thumb travel and gesture curvature in each thumb travel segment, and through steps S2-S4, whether the user has the authority to control the corresponding device with the suspected gesture instruction is accurately verified, thereby reducing the possibility of other users misoperating the device.

[0053] 3. Verify the similarity between the first change state and the second change state by judging the number of grid number jumps and the change in curvature stroke. Specifically, based on the number of grid number jumps and the change in curvature stroke, in each grid of the selection rectangle learned in steps B1-B2, find the data feature differences between the real-time collected gesture state data and the gesture learning data. This greatly improves the accuracy of verifying that the suspected gesture command identified in step S2 is a real gesture command, ensuring the accuracy of the gesture command generated based on the identified gesture state without using a complex gesture recognition algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.

[0055] Figure 1 This is a diagram illustrating the steps for implementing a method for permission verification and device control based on gesture recognition in a smart cockpit scenario provided by an embodiment of the present invention;

[0056] Figure 2 It is a schematic diagram of the key points of the hand in the palm;

[0057] Figure 3 is a structural diagram of the gesture learning device provided in this embodiment;

[0058] Figure 4 It is a schematic diagram of the thumb button set on the slide;

[0059] Figure 5 This is an example diagram of a selection rectangle generated by the gesture learning device in this embodiment;

[0060] Figure 6 This is an example diagram in which the stroke learning fitting curve and the stroke fitting curve are plotted in the same xy-axis coordinate system. DETAILED DESCRIPTION

[0061] The technical solution of the present invention will be further described below with reference to the accompanying drawings and through specific implementation methods.

[0062] Among them, the drawings are only used for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting this patent; in order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0063] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "inside", "outside" and the like indicate an orientation or position relationship based on the orientation or position relationship shown in the drawings, it is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting this patent. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0064] In the description of the present invention, unless otherwise expressly specified or limited, when the term "connection" or the like appears to indicate a connection relationship between components, such term should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be internal communication between two components or an interaction between two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood in specific circumstances.

[0065] The embodiment of the present invention provides a method for authority verification and device control based on gesture recognition in a smart cockpit scenario, such as Figure 1 As shown, the steps include:

[0066] S1: Set the initial state of the slide to an extended state to cater to the user's extended palm, and then learn the user's gesture state characteristics under different gesture commands. The learning content includes: gesture stroke data and gesture curvature data under each gesture command, and the correlation between gesture stroke data and gesture curvature data generated in the same stroke segment;

[0067] The following combination Figure 2 、 Figure 3 、 Figure 4 The method of setting the initial state of the slide to the stretched state of the user's palm in step S1 is described below. The steps specifically include:

[0068] A1, the user presses with the tip of the thumb Figure 3 The thumb button 10 on the gesture learning device is in the initial position, so that the thumb button 10 is fixed to the slide 20; the initial position of the thumb button 10 is explained as follows: when the thumb button 10 and the slide 20 are in a state where they can slide relative to each other, the thumb button 10 has the maximum sliding stroke and the distance Figure 5 As shown in FIG, the middle fingertip landing area 100 is the farthest, and the thumb button 10 is at the initial position;

[0069] A2, the thumb pushes the thumb button 10 to drive the slide 20 along the slide rail 30 ( Figure 3 ) until your palm is stretched (five fingers naturally spread apart to Figure 5 In the stretched state shown in FIG), the user's middle fingertip moves to the position set at Figure 3 On the gesture learning device shown, press button 40;

[0070] A3, press the push button 40 to fix the slide 20 relative to the slide rail 30, and then press the thumb button 10 a second time to allow the thumb button 10 to slide freely in the slide 20.

[0071] It should be noted that the fixing and separation of the thumb button 10 and the slide 20 can be achieved using a conventional structure, such as using a spring device, pressing the thumb button 10, Figure 3 The spring 60 shown in the figure is compressed, so that the thumb button 10 is tightly fixed to the slide 20. At this time, when the thumb button 10 is pushed, the slide 20 is also moved. After pressing the spring again, the thumb button 10 is separated from the slide 20, and the thumb button 10 can slide relative to the slide 20.

[0072] In order to facilitate the thumb button to slide freely in the slide and not to be easily separated, preferably, the slide 20 covers the thumb button 10 with a soft material (such as silicone), and the thumb button 10 is provided with a member such as a thumb member for pressing against the thumb. Figure 4 The protrusion 50 is shown in FIG.

[0073] In step S1, the method for learning the user's gesture state under different gesture instructions includes the following steps:

[0074] B1: The gesture learning device selects the smallest rectangle of the hand key points in the stretched palm image, discretizes the selected rectangle into grids, and numbers the discrete grids to obtain the numbered grid where each hand key point is located;

[0075] Specifically, through the above steps A1-A3, the user's middle fingertip falls to Figure 5 After the tip of the middle finger on the gesture learning device shown in the figure falls on area 100, the gesture learning device starts to collect the palm image of the user in the stretched palm state, and then detects each hand key point in the palm image through the existing algorithm, and then selects the palm area image in the palm image using the existing minimum rectangular frame selection method. In this embodiment, the rectangular frame that selects the palm area image is defined as the frame selection rectangle.

[0076] Then Figure 5 The selected rectangle 200 is discretized into grids and the discretized grids are numbered, for example, 1, 2, 3, ... from left to right and from top to bottom, thereby obtaining the numbered grid where each key point of the hand is located. Figure 5 In the figure, the key point k12 (i.e. the tip of the middle finger) of the hand is in the numbered grid 6. It should be noted that each numbered grid is preferably a square with the same size, and the side length is preferably the longitudinal axis between two key points of the hand in the same palm (with Figure 5 The horizontal side of the selection rectangle 200 shown in FIG is the horizontal axis of the xy axis coordinate system, and the vertical side is the vertical axis) is the minimum value of the distance.

[0077] B2. Collect each frame of gesture state image under each gesture instruction at a specified frequency, and calculate the position change data of the gesture state in each frame image under the corresponding travel segment, wherein the first position change data of the thumb fingertip point in each travel segment constitutes the gesture travel data learned under the gesture instruction, and the second position change data of each hand key point in each travel segment constitutes the gesture curvature data learned under the gesture instruction.

[0078] For example, suppose gesture command i1 is to control a device to perform action a. The gesture state corresponding to activating gesture command i1 is: the palm begins outstretched, the thumb and index finger gradually move closer together, and eventually touch, with the other fingers naturally flexing during this process. In step B2, each frame of gesture state image of the user performing the aforementioned gesture state under gesture command i1 is captured at a specified frequency, and the position change data of the gesture state in each frame of image is calculated within the corresponding travel segment.

[0079] The first position change data and the second position change data are described in detail below:

[0080] Let the sequence of several frames of gesture state images collected for the specified gesture instruction, such as gesture instruction i1, in step B2 be s i1 ={p1, p2, ..., p j ,…,p m}, p j represents the jth frame of the gesture state image, where j = 1, 2, …, m, where m represents the image frame number. The first position change data includes: the first relative coordinate position of the thumb tip point relative to the lower right vertex of the initial numbered grid where it is located in the stretched state at the j+1 point in time that forms the end of the stroke segment, and the first number jump of the first fingertip point numbered grid where the thumb tip point is located at the j+1 point in time that forms the end of the stroke segment relative to the second fingertip point numbered grid where the thumb tip point is located at the jth point in time that forms the start of the stroke segment.

[0081] For example, assuming that each frame is a travel segment, when j=1, it is the starting time point of the travel segment when the thumb tip slides in the travel segment, and when j+1 is the end time point of the travel segment when the thumb tip slides in the travel segment. At the starting time point of the travel segment, assuming that the thumb tip is at Figure 5 In the numbered grid "108" shown, the numbered grid 108 is assumed to be the initial numbered grid where the thumb tip is in the stretched state of the palm. The lower right vertex of the numbered grid 108 is Figure 5 Assume that at time j+1, the thumb tip point k4 falls into Figure 5For grid number 95, the first relative coordinate position of the thumb tip in this travel segment is the relative coordinate position between the landing point of thumb tip k4 in grid number 95 and the lower right vertex p1 in grid number 108 (defined as the first relative coordinate position). The "relativity" of the coordinate position here is reflected in the relative distance and relative direction between thumb tip k4 and lower right vertex p1. The first number jump generated in this travel segment is the number jump between grid numbers 95 and 108 (defined as the first number jump), that is, 108 - 95 = 13.

[0082] The second position change data includes: the second relative coordinate position of the hand key point and the upper left vertex of the starting number grid where the hand key point is located in the stretched state at the j+1 time point when the stroke segment ends, and the second number jump number of the first hand key point number grid where the hand key point is located at the j+1 time point when the stroke segment ends relative to the second hand key point number grid where the hand key point is located at the j time point when the stroke segment starts.

[0083] For example, Figure 5 In the segment where the thumb tip point k4 moves from the grid number 108 to the grid number 95, assuming Figure 5 If the index fingertip key point k8 moves from grid number 21 to grid number 34, the second relative coordinate position of the index fingertip key point k8 in this travel segment is: the relative coordinate position (defined as the second relative coordinate position) between the landing point of the index fingertip key point k8 in grid number 34 at the end of the travel segment j+1 and the upper left vertex p2 of the starting grid number (grid number 21) where the index fingertip key point k8 is located when the palm is stretched. The second numbered jump number corresponding to the index fingertip key point k8 in this travel segment is 34-21=13.

[0084] In the above example, it should be noted that the first relative coordinate position and the second relative coordinate position can be calculated using the existing algorithm, because Figure 5 The coordinates of each vertex of each numbered grid in the shown selection rectangle are known, and the distance between the hand key point and each vertex of the numbered grid can be calculated using an existing algorithm. Therefore, the coordinates of the landing point of the hand key point in the numbered grid can be derived and calculated, and finally the above-mentioned relative coordinate position is calculated.

[0085] In addition, in step S1, the method for associating the gesture travel data and the gesture curvature data generated in the same travel segment is, for example:

[0086] A unique code is assigned to each stroke segment under each gesture command, and the gesture stroke data and gesture curvature data generated under the same stroke segment are assigned a unique code of the corresponding stroke segment. In this way, an association relationship marked by the unique stroke segment code is established between the gesture stroke data and the gesture curvature data.

[0087] After completing the learning of the gesture state characteristics of each type of gesture command in step S1, the permission verification and device control method based on gesture recognition in the smart cockpit scenario provided by this embodiment proceeds to step:

[0088] S2, identifying a suspected gesture command based on the gesture state data collected from the user in real time, and then obtaining gesture learning data learned by the user for the suspected gesture command, including a stroke learning fitting curve pre-drawn based on the gesture stroke data learned for the suspected gesture command;

[0089] Specifically, in step S2, the method for collecting the user's gesture state data in real time includes the following steps:

[0090] C1 is formed by discretizing and numbering the grid in step B1. Figure 5 The selected rectangle shown is the first frame of gesture image acquired by virtual projection tracking in quasi-real-time capture, which specifically includes the following steps:

[0091] C11, Adjustment Figure 5 The projection posture of the selection rectangle shown in FIG. 1 makes the recognized thumb tip point fall into the initial numbered grid of the selection rectangle projected parallel to the gesture; for example, after the camera recognizes the key points of the hand, the recognized thumb tip point is used as the alignment point for virtually projecting the selection rectangle onto the recognized gesture, and the corresponding size generated in step B1 is as follows Figure 5 The initial numbered grid in the selection rectangle shown (i.e. Figure 5 108 grids in the figure) are projected onto the identified thumb tip point; it should be noted here that the projection of the selection rectangle onto the gesture is a parallel projection, which means that the identified hand key points fall into the projected selection rectangle.

[0092] C12, determine whether the real-time identified end fingertip point falls within the middle fingertip landing area in the projected frame selection rectangle (i.e. Figure 5 The middle fingertip landing area 100),

[0093] If yes, then determine the first frame of gesture image that is successfully tracked;

[0094] If not, the user is prompted to adjust the gesture posture to make the palm of the hand stretch out, and then the process returns to step C11.

[0095] Steps C11-C12 ensure the accuracy of gesture recognition by partially pre-aligning the thumb and middle fingertips through the subsequent recognition algorithm. The gesture travel data and curvature data for each travel segment are calculated for each frame of real-time image capture using a gridded and numbered minimum rectangular frame in a projection manner. This greatly simplifies the subsequent gesture recognition process and helps significantly improve gesture recognition accuracy.

[0096] After the first frame of gesture image to be collected in real time is tracked in step C1, in step S2, the method for collecting the user's gesture state data in real time proceeds to the following steps:

[0097] C2, starting from the first frame of gesture image tracking, a continuous frame image sequence is captured in real time at a preset frequency until the continuous capture termination condition is reached. The position change data of the gesture state in each frame of the continuous frame image sequence obtained after the acquisition is completed in the corresponding travel segment is used as the real-time gesture state data of the user;

[0098] The termination condition of continuous acquisition is: the relative distance between the thumb tip point in each frame of the continuous frame image sequence and the lower right vertex of the initial numbered grid in the selection rectangle decreases as the number of acquisition frames increases and reaches the minimum relative distance. Figure 5 As shown, the relative distance of the lower right vertex p1 of the initial numbered grid (grid numbered 108) in the frame selection rectangle decreases with the increase of the number of acquisition frames and reaches the minimum relative distance.

[0099] It's also important to note that in step C1, tracking the first frame of the quasi-real-time gesture image capture not only provides a continuous frame sequence for verifying whether the suspected gesture command is truly the corresponding gesture command, but also provides preliminary verification of whether the user has permission to control the device using gesture commands. If the system fails to track the first frame of the gesture image, it will determine that the user does not have permission to control the device using gesture commands and will issue an alert, such as prompting the user to learn gesture commands.

[0100] In step S2, the method for identifying suspected gesture instructions based on the gesture state data collected in real time specifically includes the following steps:

[0101] D1: Calculate the position change of the thumb tip in the last frame relative to the first frame in the continuous frame image sequence, and calculate the number of grid number jumps of each hand key point in the last frame relative to the first frame grid number in the first frame, and then solve the average number of grid number jumps corresponding to all hand key points;

[0102] D2: Check whether there is a gesture instruction in the gesture instruction library that has completed gesture state learning, and whether there is a gesture instruction in which the first deviation between the position change scalar and the position change (the absolute value of the difference between the position change scalar and the position change) is less than a preset first deviation threshold, and the second deviation between the standard mean of the number of beats and the mean number of beats (the absolute value of the difference between the standard mean of the number of beats and the mean number of beats) is less than a preset second deviation threshold.

[0103] If so, the matched gesture instruction is identified as a suspected gesture instruction to be represented by the continuous frame image sequence;

[0104] If not, it is determined that the recognition of the suspected gesture instruction of the continuous frame image sequence has failed, and the user is prompted to learn the suspected gesture instruction.

[0105] The following is an example of implementing steps D1-D2:

[0106] In step D1, it is assumed that in the first frame of the continuous frame image sequence, the thumb tip point falls on Figure 5 The distance between the 108th grid and the lower right vertex p1 of the 108th grid is assumed to be L1; in the last frame, the thumb tip is assumed to fall into Figure 5 For grid No. 57 in the image, the distance between it and the lower right vertex p1 is assumed to be L2. Then the position change of the thumb tip point in the last frame relative to the first frame in the continuous frame image sequence is L2-L1.

[0107] Taking the number of grid number jumps of the thumb fingertip as an example, the number of grid number jumps of each hand key point of the user in the last frame number grid relative to the first frame number grid in the first frame is explained. For example, in the above example, the thumb fingertip jumps from grid number 108 to grid number 57, so the number of grid number jumps of the thumb fingertip in the first frame relative to the last frame is 108-57=51. The mean jump number is the mean of the number of grid number jumps corresponding to all hand key points. The grid number jump number is the absolute value of the difference between the grid number jump numbers of the first frame number grid and the last frame number grid.

[0108] In step D2, the calculation method for the position change scalar is the same as the calculation method for the position change amount, and the calculation methods for the standard mean of the number of jumps and the mean number of jumps are also the same. The difference is that the calculation basis for the position change scalar and the standard mean of the number of jumps is the gesture state feature learning data of the suspected gesture command.

[0109] In step S2, the method for drawing the travel learning fitting curve is briefly described as follows:

[0110] like Figure 6In the stroke fitting curve 300 and the stroke learning fitting curve 400 , the x-axis coordinates of the fitting points are the corresponding acquisition frequencies in the continuous frame image sequence, and the y-axis coordinates are the distance of the thumb tip from the lower right vertex p1 of the initial numbered grid at that acquisition frequency. The fitting points corresponding to each frame in the continuous frame image sequence constitute the data points used to fit the fitting curve 300 . A quadratic function or other higher-order equation is then used to fit each data point to obtain the stroke fitting curve. The fitting method for the stroke learning fitting curve is the same as that for the stroke fitting curve, except that the fitting data is the learning data of the gesture state features.

[0111] After obtaining the gesture learning data learned by the user for the suspected gesture command in step S2, the permission verification and device control method based on gesture recognition in the smart cockpit scenario provided by this embodiment proceeds to step:

[0112] S3, based on the real-time collected gesture state data, draw a travel fitting curve of the user's thumb fingertip point (such as Figure 6 (see the example of “300” in the accompanying figure);

[0113] S4, judging whether the travel fitting curve and the travel learning fitting curve are similar.

[0114] If yes, go to step S5;

[0115] If not, an alarm is issued and the user is prompted that he or she does not have the authority to control the corresponding device with the suspected gesture instruction, and then the process returns to step S2;

[0116] If it is necessary to explain here, Figure 6 There are many existing methods for similarity matching of the two curves in the example, such as determining whether the overlap of the two curves exceeds an overlap threshold. If so, it is determined that the two curves are similar, so this will not be elaborated in detail.

[0117] The purpose of step S4 is to identify irregular control gestures on the premise that it is determined in step S2 that the user has the authority to control the corresponding device with suspected gesture instructions, so as to reduce the possibility of false triggering of instructions, especially when there are a large number of instructions controlled by the same palm and the distinguishability of the gesture state characteristics of each instruction is low, it can effectively identify irregular control fingers and reduce the possibility of false triggering of instructions.

[0118] After step S4, Figure 1 As shown, the permission verification and device control method based on gesture recognition in the smart cockpit scenario provided by this embodiment enters the following steps:

[0119] S5, determining whether a first change state of the user's gesture recorded in the real-time collected and processed gesture state data at the thumb tip in each first travel segment is similar to a second change state of the gesture recorded in the gesture learning data of the suspected gesture instruction in the corresponding second travel segment,

[0120] If so, the suspected gesture command is used to control the operation of the corresponding device;

[0121] If not, the user is prompted that the control gesture is invalid, and then the process returns to step S2.

[0122] The following explains the data contents of the first change state and the second change state:

[0123] The continuous frame image sequence is divided into a number of first run segments to obtain a continuous frame image subsequence corresponding to each first run segment. For example, for the expression s i ={p1, p2, ..., p j ,…,p m}, assuming that 3 frames of images constitute a first run segment, the continuous frame image subsequence corresponding to the first run segment is expressed as {p j , p j+1 , p j+2}.

[0124] The first change state corresponding to the first travel segment includes: the first grid number jump number and the first position change amount of each hand key point recognized in real time by the user in the last frame relative to the first frame of the continuous frame image subsequence. The first grid number jump number is, for example: assuming that the last frame of the thumb tip in a first travel segment falls into Figure 5 The first frame of the grid No. 82 shown falls on Figure 5 The number of grids with the number 108 shown in the figure, the number of jumps of the first grid number of the thumb tip point in the first stroke segment is 108-82=26. The first change is: the number of grids with the number 82 after the thumb tip point falls into the last frame. Figure 5 The absolute value of the difference between the distance L3 from the lower right vertex p1 in the first frame and the distance L4 from the lower right vertex p1 when it falls into grid number 108 in the first frame. The method for calculating the first grid number jump count and first change amount for other hand key points in each first travel segment is the same as the method for calculating the first grid number jump count and first change amount for the thumb fingertip point in the above example. The difference is that when calculating the first change amount for other hand key points, the relative distance between the hand key point and the upper left vertex of the grid number it falls into in the first frame of the travel segment is calculated.

[0125] The calculation principle of the second grid number jump number and the second position change is the same as the calculation principle of the first grid number jump number and the first position variable. The difference is that gesture learning data for suspected gesture instructions is used.

[0126] The corresponding relationship between the first stroke segment and the second stroke segment is described as follows:

[0127] For example, for suspected gesture instructions, by expressing s i-学习 ={p 1-学习 , p 2-学习 ,…,p j-学习 ,…,p m-学习 The gesture state learning is performed on the continuous frame image sequence of}, and the continuous frame image sequence is divided into a number of second journey segments arranged in sequence with every 3 frames as a second journey segment. The gesture state data collected in real time is expressed as s i ={p1, p2, ..., p j ,…,p m}, and every 3 frames is a first run segment. i The first travel segments and the second travel segments that are arranged in the same order have a corresponding relationship.

[0128] Step S5 specifically includes the following steps:

[0129] D1, calculate the mean of the first grid number jitter counts of each hand key point in each first travel segment (defined as the first mean), then calculate the mean of the first mean values ​​associated with each first travel segment (defined as the second mean); calculate the mean of the second grid number jitter counts in each second travel segment (defined as the third mean), then calculate the mean of the third mean values ​​associated with each second travel segment (defined as the fourth mean);

[0130] D2, determine whether the deviation (absolute value of the difference) between the second mean and the fourth mean is less than a preset third deviation threshold,

[0131] If yes, go to step D3;

[0132] If not, it is determined that the first change state and the second change state are not similar;

[0133] D3: Calculate the mean of the first position changes of each hand key point in each first travel segment (defined as the fifth mean), then calculate the mean of the fifth means associated with each first travel segment (defined as the sixth mean); and calculate the mean of the second changes in each second travel segment (defined as the seventh mean), then calculate the mean of the seventh means associated with each second travel segment (defined as the eighth mean);

[0134] D4, judging whether the deviation (absolute value of the difference) between the sixth mean and the eighth mean is less than a preset fourth deviation threshold,

[0135] If so, it is determined that the first change state and the second change state are similar;

[0136] If not, it is determined that the first change state and the second change state have no similarity.

[0137] It should be noted that as more gesture commands can be performed by the same hand, the characteristic differences between different gesture states generally decrease. This may result in the same hand key point having the same number of grid number jumps but different position changes in the first and second corresponding travel segments, or the same position change but different number of grid number jumps. To reduce the possibility of gesture command misrecognition, this embodiment improves the accuracy of gesture command recognition through dual judgment in steps D2 and D4.

[0138] In summary, the present invention constructs a correlation relationship between the thumb stroke and the gesture curvature in each thumb stroke segment, and uses this correlation relationship to accurately characterize the subtle feature differences between different gesture states corresponding to each captured gesture instruction. Based on these subtle feature differences, the suspected gesture instructions corresponding to the gesture state data collected in real time can be accurately identified, and the similarity between the first change state and the second change state is verified through step S5, thereby improving the accuracy of gesture state recognition based on these subtle feature differences, making it possible to avoid the use of complex gesture recognition algorithms such as neural network model learning, and at the same time improving the efficiency of gesture recognition.

[0139] It should be noted that the above-described specific embodiments are merely preferred embodiments of the present invention and the technical principles employed. Those skilled in the art will appreciate that various modifications, equivalent substitutions, and variations may be made to the present invention. However, as long as these modifications do not depart from the spirit of the present invention, they are intended to be within the scope of protection of the present invention. Furthermore, certain terms used in the specification and claims of this application are not intended to be limiting and are provided solely for ease of description.

Claims

1. A method for authority verification and device control based on gesture recognition in a smart cockpit scenario, characterized in that the steps include: S1, setting the initial state of the slide to an extended state to cater to the user's extended palm, and then learning the user's gesture state characteristics under different gesture commands, the learning content including: gesture stroke data and gesture curvature data under each gesture command, and the correlation between the gesture stroke data and the gesture curvature data generated in the same stroke segment; S2, identifying suspected gesture commands based on gesture state data collected and processed in real time from the user, and then obtaining gesture learning data learned by the user for the suspected gesture commands, including a stroke learning fitting curve pre-drawn based on the gesture stroke data learned for the suspected gesture commands; S3, drawing a travel fitting curve of the user's thumb tip based on the gesture state data collected in real time; S4, determining whether the travel fitting curve is similar to the travel learning fitting curve, If yes, go to step S5; If not, an alarm is issued and the user is prompted that he or she does not have the authority to control the corresponding device with the suspected gesture instruction, and then the process returns to step S2; S5, determining whether a first change state of the user's gesture recorded in the gesture state data in each first travel segment of the thumb tip point is similar to a second change state of the gesture recorded in the gesture learning data in the corresponding second travel segment, If so, the suspected gesture command is used to control the operation of the corresponding device; If not, the user is prompted that the control gesture is invalid, and then the process returns to step S2.

2. The method for permission verification and device control based on gesture recognition in a smart cockpit scenario according to claim 1 is characterized in that: In step S1, the method for setting the initial state of the slide to the stretched state includes the steps of: A1, the user presses the thumb button in the initial position on the gesture learning device with the tip of the thumb, so that the thumb button is fixed to the slide; A2, the thumb tip pushes the thumb button to drive the slide along the slide rail provided in the slide until the user's middle finger tip shifts to the press button provided on the gesture learning device when the palm is stretched; A3, press the push button to fix the slide relative to the slide rail, and then press the thumb button a second time to allow the thumb button to slide freely in the slide.

3. The method for permission verification and device control based on gesture recognition in a smart cockpit scenario according to claim 2 is characterized in that: The slideway covers the thumb button with soft material, and the thumb button is provided with a protrusion for pressing against the thumb.

4. The method for permission verification and device control based on gesture recognition in a smart cockpit scenario according to claim 1 is characterized in that: In step S1, the method for learning the user's gesture state under different gesture instructions includes the following steps: B1, after selecting the minimum rectangle of the hand key points in the palm image in the stretched state, the selected rectangle is discretized into grids and the discrete grids are numbered to obtain the numbered grid where each hand key point is located; B2. Collect each frame of gesture state image under each gesture instruction at a specified frequency, and calculate the position change data of the gesture state in each frame image under the corresponding travel segment, wherein the first position change data of the thumb fingertip point in each travel segment constitutes the gesture travel data learned under the gesture instruction, and the second position change data of each hand key point in each travel segment constitutes the gesture curvature data learned under the gesture instruction.

5. The method for permission verification and device control based on gesture recognition in a smart cockpit scenario according to claim 4 is characterized in that: In step B2, the specified gesture instruction A sequence of several frames of gesture state images collected , Indicates the Frame gesture state image, , Indicates the number of image frames, the first position change data includes: the end of the travel segment forming the travel segment At the time point, the first relative coordinate position of the thumb tip and the lower right vertex of the initial numbered grid in the stretched state, and ... At the time point, the first fingertip point number grid where the thumb fingertip point is located is relative to the starting point of the stroke segment forming the stroke segment. The number of beats of the first number of the second fingertip point number grid at the time point; The second position change data includes: the second relative coordinate position of the hand key point and the upper left vertex of the starting number grid in the stretched state at the end j+1 point of the stroke segment, and the second relative coordinate position of the hand key point and the upper left vertex of the starting number grid in the stretched state at the end of the stroke segment. At the time point, the first hand key point number grid where the hand key point is located is relative to the hand key point at the start of the stroke segment The second number jump number of the second hand key point number grid at the time point.

6. The method for authority verification and device control based on gesture recognition in a smart cockpit scenario according to claim 4 is characterized in that: In step S2, the method for collecting the gesture state data of the user in real time includes the following steps: C1, using the selected rectangle formed by the grid discretization and numbering in step B1 as a virtual projection tracking to obtain the first frame of gesture image captured in quasi-real time; C2, starting from tracking the first frame of the gesture image, capturing a sequence of continuous frame images in real time at a preset frequency until a continuous capture termination condition is reached, and using the position change data of the gesture state in each frame of the continuous frame image sequence obtained after the capture in the corresponding travel segment as the gesture state data collected in real time from the user; The continuous acquisition termination condition is that the relative distance between the thumb tip point in each frame in the continuous frame image sequence and the lower right vertex of the initial numbered grid in the selection rectangle decreases as the number of acquisition frames increases and reaches a minimum relative distance.

7. The method for authority verification and device control based on gesture recognition in a smart cockpit scenario according to claim 6 is characterized in that: In step C1, the method for tracking the first frame of gesture image collected in real time includes the following steps: C11, adjusting the projection posture of the selection rectangle so that the thumb tip point recognized in real time falls into the initial numbered grid in the minimum rectangular frame projected parallel to the gesture; C12, determining whether the middle fingertip point recognized in real time falls within the middle fingertip landing area grid in the projected frame selection rectangle, If so, it is determined that the first frame of the gesture image is successfully tracked; If not, return to step C11.

8. The method for authority verification and device control based on gesture recognition in a smart cockpit scenario according to claim 6 is characterized in that: In step S2, the method for identifying the suspected gesture instruction includes the following steps: D1, calculating the position change of the thumb tip point in the last frame relative to the first frame in the continuous frame image sequence, and calculating the number of grid number jumps of each hand key point of the user in the last frame number grid relative to the first frame number grid in the first frame, and then solving the average number of grid number jumps corresponding to all the hand key points; D2, from the gesture instruction library that has completed gesture state learning, match whether there is a gesture instruction whose first deviation between the position change scalar and the position change is less than a preset first deviation threshold, and whose second deviation between the standard mean of the number of beats and the mean of the number of beats is less than a preset second deviation threshold, If yes, identifying the matched gesture instruction as the suspected gesture instruction intended to be represented by the continuous frame image sequence; If not, it is determined that the recognition of the suspected gesture instruction of the continuous frame image sequence has failed, and the user is prompted to learn the instruction.

9. The method for authority verification and device control based on gesture recognition in a smart cockpit scenario according to claim 6 is characterized in that: The continuous frame image sequence is divided into a plurality of first run segments, and a continuous frame image subsequence corresponding to each first run segment is obtained. In step S5, the first change state corresponding to the first run segment includes: the number of first grid number jumps and the first position change of the position of each hand key point recognized by the user in real time in the last frame relative to the position of the first frame in the continuous frame image subsequence; the second change state corresponding to the second run segment having a corresponding relationship with the first run segment includes: the number of second grid number jumps and the second position change of the position of each hand key point learned by the user for the second run segment in the suspected gesture instruction in the last frame of the continuous frame learning subsequence relative to the position of the first frame; The corresponding relationship between the first travel segment and the second travel segment is: the first travel segment and the second travel segment have the same data collection start frequency and the same data collection end frequency.

10. The method for authority verification and device control based on gesture recognition in a smart cockpit scenario according to claim 9 is characterized in that: The method for determining similarity in step S5 includes the following steps: D1, calculating a first mean of the number of jitters of the first grid numbers of each of the hand key points in each of the first travel segments, and then calculating a second mean of the first mean values ​​associated with each of the first travel segments; and calculating a third mean of the number of jitters of the second grid numbers in each of the second travel segments, and then calculating a fourth mean of the third mean values ​​associated with each of the second travel segments; D2, determining whether the deviation between the second mean and the fourth mean is less than a preset third deviation threshold, If yes, go to step D3; If not, determining that the first change state and the second change state are not similar; D3, calculating a fifth mean of the first position changes of each of the hand key points in each of the first travel segments, and then calculating a sixth mean of the fifth mean values ​​associated with each of the first travel segments; and calculating a seventh mean of the second position changes of each of the second travel segments, and then calculating an eighth mean of the seventh mean values ​​associated with each of the second travel segments; D4, determining whether the deviation between the sixth mean and the eighth mean is less than a preset fourth deviation threshold, If so, determining that the first change state and the second change state have similarity; If not, it is determined that the first change state and the second change state have no similarity.

Citation Information

Patent Citations

  • Gesture action recognition method and device, terminal equipment and storage medium

    CN114648807A

  • Real-time dynamic gesture recognition and personalized customization method and system

    CN118537914A