Screen control method and device, equipment, storage medium and vehicle
By obtaining palm skeleton information and key point information of continuous t-frame gesture images, and using dynamic gesture recognition model for feature extraction and recognition, the problem of low gesture recognition accuracy in the prior art is solved, and higher screen control accuracy and user experience are achieved.
Patent Information
- Application Number
- CN202311489771.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-09
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, the dynamic gesture recognition process is based only on the hand key point information corresponding to each of the multi-frame gesture images, resulting in low accuracy of gesture recognition results, thereby reducing the accuracy of screen control and user interaction experience.
By obtaining the palm skeleton information and the corresponding hand key point information of the continuous t-frame gesture images, the dynamic gesture recognition model is used to extract the information, obtain the gesture feature information, and dynamic gesture recognition to obtain the target gesture type, and finally execute the corresponding control instructions according to the target gesture type.
It improves the accuracy of gesture recognition, thereby improving the accuracy of screen control and improving the user's interactive experience.
Smart Images

Figure CN119960592A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of screen control technology, and in particular, relates to a screen control method, device, equipment, storage medium and vehicle. Background Art
[0002] At present, a second row of entertainment screens can be set on the top of the vehicle, so that users sitting in the second row of the vehicle can watch the screens of the second row of entertainment screens for leisure and entertainment. Since the distance between the screen and the user is far, the screen can be controlled by gestures for the convenience of interaction. Specifically, controlling the screen by gestures can include first recognizing the user's gesture by dynamic gesture recognition technology, obtaining a gesture recognition result, and then executing a control instruction corresponding to the gesture recognition result.
[0003] In the prior art, the dynamic gesture recognition process is usually implemented only based on the hand key point information corresponding to each of the multiple frames of gesture images. The accuracy of the gesture recognition result is low, resulting in low accuracy of screen control and reduced user interaction experience. Summary of the invention
[0004] The embodiments of the present application provide a screen control method, device, equipment, storage medium and vehicle, which can improve the accuracy of gesture recognition, improve the accuracy of screen control, and enhance the user's interactive experience.
[0005] In a first aspect, an embodiment of the present application provides a screen control method, the method comprising:
[0006] Obtain the palm skeleton information and the hand key point information corresponding to each of the consecutive t-frame gesture images, where t is a positive integer;
[0007] Using a dynamic gesture recognition model to extract features from the palm skeleton information and the t key point information of the hand to obtain gesture feature information;
[0008] Performing dynamic gesture recognition on the gesture feature information to obtain a target gesture type;
[0009] The screen is controlled according to the control instruction corresponding to the target gesture type.
[0010] In a possible implementation, the dynamic gesture recognition model includes a graph convolutional neural network layer, and the use of the dynamic gesture recognition model to extract features from the palm skeleton information and t hand key point information to obtain gesture feature information includes:
[0011] The graph convolutional neural network layer is used to extract features of the palm skeleton information and t hand key point information to obtain gesture feature information.
[0012] In a possible implementation, performing dynamic gesture recognition on the gesture feature information to obtain a target gesture type includes:
[0013] Performing dynamic gesture recognition on the gesture feature information to obtain a plurality of preset gesture types and probabilities corresponding to the plurality of preset gesture types;
[0014] Determine the probability with the largest value among the multiple probabilities as the target probability;
[0015] The preset gesture type corresponding to the target probability is determined as the target gesture type.
[0016] In a possible implementation, the hand key point information includes first coordinates of m hand key points, where m is a positive integer; and the feature extraction of the palm skeleton information and the t hand key point information to obtain gesture feature information includes:
[0017] Subtracting the t×m first coordinates from the first coordinates of the target hand key point in the target image respectively to obtain t×m second coordinates, wherein the target image is any one of the t frames of gesture images, and the target hand key point is any one of the m hand key points;
[0018] Feature extraction is performed on the palm skeleton information and t×m second coordinates to obtain gesture feature information.
[0019] In a possible implementation, before extracting features from the palm skeleton information and the t hand key point information using the dynamic gesture recognition model to obtain gesture feature information, the method further includes:
[0020] Acquire the palm skeleton information, hand key point sample information corresponding to each of the consecutive frames of gesture sample images, and a first gesture type corresponding to a plurality of the hand key point sample information, wherein the first gesture type is any one of a plurality of preset gesture types, and the first gesture type is obtained by manually labeling the plurality of the hand key point sample information;
[0021] Using an initial dynamic gesture recognition model, feature extraction is performed on the palm skeleton information and a plurality of hand key point sample information to obtain predicted gesture feature information;
[0022] Performing dynamic gesture recognition on the predicted gesture feature information to obtain a second gesture type;
[0023] determining a loss function value according to a similarity between the first gesture type and the second gesture type;
[0024] The model parameters of the initial dynamic gesture recognition model are adjusted according to the loss function value, and the dynamic gesture recognition model is obtained by training.
[0025] In a possible implementation, obtaining the hand key point information corresponding to each of the consecutive t-frame gesture images includes:
[0026] Get t consecutive frames of gesture images;
[0027] For each frame of the gesture image, performing hand detection on the gesture image to obtain hand position information;
[0028] Perform hand key point positioning on the hand corresponding to the hand position information to obtain the hand key point information.
[0029] In a second aspect, an embodiment of the present application provides a screen control device, the device comprising:
[0030] The first acquisition module is used to obtain palm skeleton information and hand key point information corresponding to each of t consecutive frames of gesture images, where t is a positive integer;
[0031] A first extraction module is used to extract features of the palm skeleton information and the t hand key point information using a dynamic gesture recognition model to obtain gesture feature information;
[0032] A first recognition module, used to perform dynamic gesture recognition on the gesture feature information to obtain a target gesture type;
[0033] The control module is used to control the screen according to the control instruction corresponding to the target gesture type.
[0034] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device comprising: a processor and a memory storing computer program instructions;
[0035] When the processor executes the computer program instructions, the processor implements any possible implementation method of the first aspect described above.
[0036] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implements a method in any possible implementation method of the first aspect described above.
[0037] In a fifth aspect, an embodiment of the present application provides a vehicle, the vehicle comprising at least one of the following:
[0038] A screen control device as in any one of the embodiments of the second aspect;
[0039] An electronic device as in any one of the embodiments of the third aspect;
[0040] A computer-readable storage medium as in any embodiment of the fourth aspect.
[0041] In the screen control method, device, equipment, storage medium and vehicle of the embodiment of the present application, the hand key point information corresponding to each of the continuous t-frame gesture images provides the association information between multiple key points between frames, and the palm skeleton information provides the association information between multiple key points in a single frame. Based on this, by using the dynamic gesture recognition model to extract features from the palm skeleton information and t hand key point information, gesture feature information is obtained, and gesture feature information including palm skeleton information can be obtained. By performing dynamic gesture recognition on the gesture feature information including the palm skeleton information, the model recognizes the dynamic change process of multiple key points between frames on the basis of paying attention to the association information between multiple key points in a single frame, which can improve the accuracy of the target gesture type recognized. Then, by controlling the screen according to the control instructions corresponding to the target gesture type, the accuracy of screen control can be improved, and the user's interactive experience can be enhanced. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solution of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0043] Figure 1 It is a flowchart of a screen control method provided in an embodiment of the present application;
[0044] Figure 2 is a schematic diagram of palm skeleton information provided by an embodiment of the present application;
[0045] Figure 3 is a schematic diagram of performing dynamic gesture recognition using a dynamic gesture recognition model provided by an embodiment of the present application;
[0046] Figure 4 is a flowchart of another screen control method provided in an embodiment of the present application;
[0047] Figure 5 is a structural schematic diagram of a screen control device provided in an embodiment of the present application;
[0048] Figure 6 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0049] In order to more clearly understand the above-mentioned purposes, features and advantages of the present application, the scheme of the present application will be further described below. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.
[0050] In the following description, many specific details are set forth to facilitate a full understanding of the present application, but the present application may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only part of the embodiments of the present application, rather than all of the embodiments.
[0051] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprises" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0052] As described in the background technology section, in order to solve the existing technical problems, the embodiments of the present application provide a screen control method, device, equipment, storage medium and vehicle. Among them, the screen control method can be applied to the scenario of controlling the second row entertainment screen in the vehicle.
[0053] The following first introduces the screen control method provided in the embodiment of the present application.
[0054] Figure 1 FIG. 1 is a flow chart of a screen control method provided by an embodiment of the present application. The screen control method can be executed by any module including a dynamic gesture recognition model. Figure 1 As shown, the screen control method provided in the embodiment of the present application includes the following steps:
[0055] S110, obtaining palm skeleton information and hand key point information corresponding to each of t consecutive frames of gesture images, where t is a positive integer;
[0056] S120, extracting features from the palm skeleton information and t hand key point information using a dynamic gesture recognition model to obtain gesture feature information;
[0057] S130, performing dynamic gesture recognition on the gesture feature information to obtain a target gesture type;
[0058] S140: Control the screen according to the control instruction corresponding to the target gesture type.
[0059] In the screen control method of the embodiment of the present application, the hand key point information corresponding to each of the continuous t-frame gesture images provides the association information between multiple key points between frames, and the palm skeleton information provides the association information between multiple key points in a single frame. Based on this, by using the dynamic gesture recognition model to extract features from the palm skeleton information and t hand key point information, gesture feature information is obtained, and gesture feature information including palm skeleton information can be obtained. By performing dynamic gesture recognition on the gesture feature information including the palm skeleton information, the model recognizes the dynamic change process of multiple key points between frames on the basis of paying attention to the association information between multiple key points in a single frame, which can improve the accuracy of the target gesture type recognized. Then, by controlling the screen according to the control instructions corresponding to the target gesture type, the accuracy of screen control can be improved, and the user's interactive experience can be enhanced.
[0060] The specific implementation methods of the above steps are introduced below.
[0061] In some embodiments, in S110, the palm skeleton information may include m key points in the palm and association information between the m key points. That is, the palm skeleton information may actually be topological structure information between the m key points in the palm. In the case where the palm includes 21 key points, the schematic diagram of the palm skeleton information may be as follows: Figure 2 shown.
[0062] As an example, by encoding the topological structure information between m key points, an adjacency matrix can be obtained. That is, the palm skeleton information can exist in the form of an adjacency matrix.
[0063] In addition, the hand key point information corresponding to the gesture image may include coordinate information of m hand key points, wherein the m hand key points in the gesture image and the m key points in the palm skeleton information may be the same.
[0064] As an example, if the hand key point information includes 21 key points, the hand key point information can be represented by a first matrix with a dimension of (t, 42).
[0065] In addition, the continuous t-frame gesture image can be the continuous t-frame image captured by the image acquisition device. The image acquisition device can be, for example, a camera. After many tests, it is known that if the camera can capture 15 frames of images per second, each type of gesture can complete the entire action within 10 frames. Therefore, when the image acquisition device is a 15FPS camera, t can be 10. When the acquisition frame rate of the camera changes, t can change accordingly.
[0066] Based on this, in some embodiments, the above-mentioned obtaining of the hand key point information corresponding to each of the consecutive t-frame gesture images may specifically include:
[0067] Get t consecutive frames of gesture images;
[0068] For each frame of gesture image, hand detection is performed on the gesture image to obtain hand position information;
[0069] The hand key points of the hand corresponding to the hand position information are located to obtain the hand key point information.
[0070] Here, after determining the value of t, a video frame buffer queue with a length of t can be established. Specifically, an infrared camera can be installed above the screen of the second row of display screens. The infrared camera can shoot users sitting in the second row of the vehicle to obtain a video stream. Every time a frame of image (i.e., video frame) in the video stream is obtained, the image can be put into the video frame buffer queue. After the image in the queue is full of t frames, the t-frame images in the queue can be determined as continuous t-frame gesture images. When a new video frame enters the end of the queue, the video frame at the head of the queue can be discarded to obtain a new queue, and the t-frame images in the new queue can be determined as continuous t-frame gesture images.
[0071] After obtaining t consecutive frames of gesture images, the hand key point detection can be performed on each frame of the gesture image to obtain the hand key point information. Specifically, for each frame of the image, the hand detection model can be used to perform hand detection on the gesture image to obtain the hand position information. The hand position information may include the coordinate information corresponding to the hand detection frame. After determining the hand position information, the key point detection model can be used to locate the hand key points of the hand corresponding to the hand position information to obtain the hand key point information.
[0072] In some embodiments, in S120, the dynamic gesture recognition model can identify the complete dynamic gesture types contained in the continuous multi-frame image through a deep learning algorithm. After obtaining the palm skeleton information and the continuous t hand key point information, the palm skeleton information and the continuous t hand key point information can be input into the dynamic gesture recognition model together. Among them, the palm skeleton information can be input in the form of an adjacency matrix to provide the model with association information between m key points in a single frame. The continuous t hand key point information can be input in the form of a first matrix to provide the model with association information between m key points between frames. In the case where 21 key points are included in the hand key point information, the dimension of the first matrix can be (t, 42). Based on this, the gesture feature information can be a second matrix with a dimension of (t, 42).
[0073] As an example, after receiving the palm skeleton information and t consecutive hand key point information, the dynamic gesture recognition model can first perform feature extraction on the palm skeleton information and the t hand key point information to obtain gesture feature information including the palm skeleton information. Since the model pays attention to the association information between the m key points in a single frame when performing feature extraction, the accuracy of feature extraction can be improved.
[0074] As an example, the dynamic gesture recognition model may include a graph convolutional neural network layer (Graph Convolutional Network, GCN). Based on this, the above S120 may specifically include:
[0075] The graph convolutional neural network layer is used to extract the palm skeleton information and t hand key point information to obtain the gesture feature information.
[0076] Here, the graph convolutional neural network layer can extract the correlation information between t key points of the hand in combination with the palm skeleton information, give each key point in each frame of the gesture image a different weight ratio, and obtain a second matrix with a dimension of (t, 42). Among them, the process of feature extraction using the graph convolutional neural network layer can refer to the existing process, and will not be described in detail here. Since the graph convolutional neural network layer pays attention to the correlation information between multiple key points in a single frame when performing feature extraction, it can more effectively extract gesture feature information.
[0077] As can be seen from the above, the hand key point information may include coordinate information of m hand key points. Among them, the coordinate information may include a first coordinate. The first coordinate may be a coordinate directly obtained by locating the hand key points of the gesture image. Since the coordinates obtained by locating are relatively complex, the complexity of model processing may be increased. Therefore, in order to simplify the complexity of model processing, improve the model processing efficiency, and thus improve the gesture recognition efficiency, in some embodiments, the palm skeleton information and the t hand key point information are subjected to feature extraction to obtain gesture feature information, which may specifically include:
[0078] Subtract the t×m first coordinates from the first coordinates of the target hand key points in the target image to obtain t×m second coordinates, where the target image is any frame in the t-frame gesture image, and the target hand key point is any one of the m hand key points;
[0079] Feature extraction is performed on the palm skeleton information and t×m second coordinates to obtain gesture feature information.
[0080] Here, the target image may be, for example, the first frame image in the t-frame gesture image. The target hand key point may be, for example, the wrist key point. In addition, by processing the t×m second coordinates, a third matrix of (t, 42) dimensions may be obtained. Among them, the complexity of the third matrix may be less than the complexity of the first matrix. Therefore, by performing feature extraction on the palm skeleton information and the t×m second coordinates, the complexity of the model processing can be simplified, the model processing efficiency can be improved, and thus the gesture recognition efficiency can be improved. In addition, by subtracting the t×m first coordinates from the first coordinates of the target hand key points in the target image, t×m second coordinates are obtained, which can strengthen the connection between frames and highlight the dynamic change process between sequence frames.
[0081] In some embodiments, in S130, the target gesture type may be any one of a plurality of preset gesture types. The plurality of preset gesture types may include waving up, waving down, waving left, waving right, clenching a fist, and opening a fist. If the target gesture type does not belong to any one of the plurality of preset gesture types, the target gesture type may be other.
[0082] In addition, the dynamic gesture recognition model may also include a convolutional neural network layer (Convolutional Neural Networks, CNN). Based on this, the above S130 may specifically include:
[0083] The convolutional neural network layer is used to perform dynamic gesture recognition on the gesture feature information to obtain the target gesture type.
[0084] Based on this, in some embodiments, the above-mentioned dynamic gesture recognition is performed on the gesture feature information to obtain the target gesture type, which may specifically include:
[0085] Performing dynamic gesture recognition on the gesture feature information to obtain a plurality of preset gesture types and probabilities corresponding to the plurality of preset gesture types;
[0086] The probability with the largest value among multiple probabilities is determined as the target probability;
[0087] The preset gesture type corresponding to the target probability is determined as the target gesture type.
[0088] Here, the convolutional neural network layer may include four convolutional layers and one fully connected layer. The four convolutional layers may extract sequence information between t hand key point information according to gesture feature information. The fully connected layer may be used as a classifier.
[0089] As an example, each convolution layer can act on the channel dimension t, and extract the change features of the entire palm in t frames by combining the information of the same key point between different frames. Each convolution layer can continuously reduce the feature dimension of a single frame and increase the channel dimension until a feature matrix of (1024,1) dimension is obtained. After obtaining the feature matrix, the fully connected layer can classify the feature matrix to obtain multiple classification results and scores corresponding to the multiple classification results. Among them, the multiple classification results can represent multiple preset gesture types. The multiple classification results can be recorded as up, down, gist, plam, left, right and negative samples. Among them, negative samples can represent others. After obtaining multiple scores, the softmax() function can be used to normalize the multiple scores to obtain the probabilities corresponding to the multiple preset gesture types.
[0090] In this way, by determining the preset gesture type corresponding to the probability with the largest value among multiple probabilities as the target gesture type, the accuracy of the target gesture type can be improved.
[0091] Based on this, as an example, a schematic diagram of gesture recognition using a dynamic gesture recognition model can be as follows: Figure 3 shown.
[0092] Based on this, in order to obtain multiple preset gesture types and their corresponding probabilities, in some embodiments, before the above S120, the following steps may also be included:
[0093] Obtain palm skeleton information, hand key point sample information corresponding to each of the consecutive multiple frames of gesture sample images, and a first gesture type corresponding to the consecutive multiple hand key point sample information, where the first gesture type is any one of multiple preset gesture types, and the first gesture type is obtained by manually annotating the consecutive multiple hand key point sample information;
[0094] The initial dynamic gesture recognition model is used to extract features from the palm skeleton information and multiple hand key point sample information to obtain predicted gesture feature information;
[0095] Performing dynamic gesture recognition on the predicted gesture feature information to obtain a second gesture type;
[0096] determining a loss function value according to a similarity between the first gesture type and the second gesture type;
[0097] The model parameters of the initial dynamic gesture recognition model are adjusted according to the loss function value, and the dynamic gesture recognition model is trained.
[0098] Here, since the first gesture type is obtained by annotation, it can be ensured that the first gesture type is any one of multiple preset gesture types. In addition, manual annotation can ensure the accuracy of the first gesture type. It should be noted that each frame of gesture sample image can correspond to a hand key point sample information, and multiple consecutive frames of gesture sample images can correspond to a first gesture type. Therefore, multiple consecutive hand key point sample information can correspond to a first gesture type.
[0099] In this way, by using the preset gesture types and their corresponding multiple hand key point sample information to train the initial dynamic gesture recognition model, a dynamic gesture recognition model is obtained, which can ensure that the dynamic gesture recognition model can obtain multiple preset gesture types and their corresponding probabilities during the gesture recognition process.
[0100] In some embodiments, in S140, different gesture types may correspond to different control instructions. For example, the control instruction corresponding to waving down may be to turn on the screen; the control instruction corresponding to waving up may be to turn off the screen; the control instruction corresponding to clenching a fist may be to click the screen; the control instruction corresponding to opening the fist may be to release the click; the control instruction corresponding to waving left and right may be to return to the previous level; and the control instruction corresponding to other (negative samples) may be to do nothing.
[0101] In this way, by controlling the screen according to the control instructions corresponding to the target gesture type, the effect of air touch can be achieved.
[0102] In order to better describe the entire solution, some specific examples are given based on the above embodiments.
[0103] For example, Figure 4 As shown, a screen control method provided by an embodiment of the present application may include the following steps:
[0104] S41, obtaining palm skeleton information and t consecutive frames of gesture images;
[0105] S42, for each frame of the gesture image, performing hand key point detection on the gesture image to obtain first coordinates of m hand key points;
[0106] S43, subtracting the t×m first coordinates from the first coordinate of the wrist key point in the first frame gesture image, respectively, to obtain t×m second coordinates;
[0107] S44, using the graph convolutional neural network layer in the dynamic gesture recognition model to extract features of the palm skeleton information and t×m second coordinates to obtain gesture feature information;
[0108] S45, using the convolutional neural network layer in the dynamic gesture recognition model to perform dynamic gesture recognition on the gesture feature information to obtain a target gesture type;
[0109] S46: Control the screen according to the control instruction corresponding to the target gesture type.
[0110] In an embodiment of the present application, the hand key point information corresponding to each of the consecutive t-frame gesture images provides the association information between multiple key points between frames, and the palm skeleton information provides the association information between multiple key points in a single frame. Based on this, by using the dynamic gesture recognition model to extract features from the palm skeleton information and t×m second coordinates, gesture feature information is obtained, and gesture feature information including palm skeleton information can be obtained. By performing dynamic gesture recognition on the gesture feature information including the palm skeleton information, the model recognizes the dynamic change process of multiple key points between frames on the basis of paying attention to the association information between multiple key points in a single frame, which can improve the accuracy of the target gesture type recognized. Then, by controlling the screen according to the control instructions corresponding to the target gesture type, the accuracy of screen control can be improved, and the user's interactive experience can be enhanced.
[0111] Based on the screen control method provided in the above embodiment, the present application also provides a specific implementation of the screen control device. Please refer to the following embodiment.
[0112] like Figure 5 As shown, the screen control device 500 provided in the embodiment of the present application includes the following modules:
[0113] The first acquisition module 510 is used to acquire palm skeleton information and hand key point information corresponding to each of t consecutive frames of gesture images, where t is a positive integer;
[0114] A first extraction module 520 is used to extract features from the palm skeleton information and t hand key point information using a dynamic gesture recognition model to obtain gesture feature information;
[0115] The first recognition module 530 is used to perform dynamic gesture recognition on the gesture feature information to obtain a target gesture type;
[0116] The control module 540 is used to control the screen according to the control instruction corresponding to the target gesture type.
[0117] The screen control device 500 is described in detail below, as shown below:
[0118] In some embodiments, the dynamic gesture recognition model includes a graph convolutional neural network layer. Based on this, the first extraction module 520 may specifically include:
[0119] The first extraction submodule is used to extract features of the palm skeleton information and t hand key point information using a graph convolutional neural network layer to obtain gesture feature information.
[0120] In some embodiments, the first identification module 530 may specifically include:
[0121] The recognition submodule is used to perform dynamic gesture recognition on the gesture feature information to obtain a plurality of preset gesture types and the probabilities corresponding to the plurality of preset gesture types;
[0122] A first determination submodule is used to determine the probability with the largest value among the multiple probabilities as the target probability;
[0123] The second determination submodule is configured to determine the preset gesture type corresponding to the target probability as the target gesture type.
[0124] In some embodiments, the hand key point information includes first coordinates of m hand key points, where m is a positive integer.
[0125] Based on this, the first extraction module 520 may specifically include:
[0126] A calculation submodule, used for subtracting the t×m first coordinates from the first coordinates of the target hand key point in the target image respectively to obtain t×m second coordinates, where the target image is any frame in the t-frame gesture image, and the target hand key point is any one of the m hand key points;
[0127] The second extraction submodule is used to perform feature extraction on the palm skeleton information and t×m second coordinates to obtain gesture feature information.
[0128] In some embodiments, the screen control device 500 may further include:
[0129] The second acquisition module is used to obtain palm skeleton information, hand key point sample information corresponding to each of the consecutive multiple frames of gesture sample images, and a first gesture type corresponding to the consecutive multiple hand key point sample information before performing feature extraction on the palm skeleton information and t hand key point information using the dynamic gesture recognition model, wherein the first gesture type is any one of multiple preset gesture types, and the first gesture type is obtained by manually annotating the multiple hand key point sample information;
[0130] The second extraction module is used to extract features from palm skeleton information and multiple hand key point sample information using the initial dynamic gesture recognition model to obtain predicted gesture feature information;
[0131] A second recognition module, used to perform dynamic gesture recognition on the predicted gesture feature information to obtain a second gesture type;
[0132] a determination module, configured to determine a loss function value according to a similarity between the first gesture type and the second gesture type;
[0133] The adjustment module is used to adjust the model parameters of the initial dynamic gesture recognition model according to the loss function value, and train the dynamic gesture recognition model.
[0134] In some embodiments, the first acquisition module 510 may specifically include:
[0135] The acquisition submodule is used to acquire continuous t-frame gesture images;
[0136] The detection submodule is used to perform hand detection on the gesture image for each frame of the gesture image to obtain the hand position information;
[0137] The positioning submodule is used to locate the hand key points of the hand corresponding to the hand position information to obtain the hand key point information.
[0138] In the screen control device of the embodiment of the present application, the hand key point information corresponding to each of the continuous t-frame gesture images provides the association information between multiple key points between frames, and the palm skeleton information provides the association information between multiple key points in a single frame. Based on this, by using the dynamic gesture recognition model to extract features from the palm skeleton information and t hand key point information, gesture feature information is obtained, and gesture feature information including palm skeleton information can be obtained. By performing dynamic gesture recognition on the gesture feature information including the palm skeleton information, the model recognizes the dynamic change process of multiple key points between frames on the basis of paying attention to the association information between multiple key points in a single frame, which can improve the accuracy of the target gesture type recognized. Then, by controlling the screen according to the control instructions corresponding to the target gesture type, the accuracy of screen control can be improved, and the user's interactive experience can be enhanced.
[0139] Based on the screen control method provided in the above embodiment, the embodiment of the present application also provides a specific implementation of the electronic device. Figure 6 A schematic diagram of an electronic device 600 provided in an embodiment of the present application is shown.
[0140] The electronic device 600 may include a processor 610 and a memory 620 storing computer program instructions.
[0141] Specifically, the processor 610 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0142] The memory 620 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 620 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 620 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 620 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 620 is a non-volatile solid-state memory.
[0143] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Thus, typically, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect of the present application.
[0144] The processor 610 implements any one of the screen control methods in the above embodiments by reading and executing computer program instructions stored in the memory 620 .
[0145] In one example, the electronic device 600 may further include a communication interface 630 and a bus 640. Figure 6 As shown, the processor 610, the memory 620, and the communication interface 630 are connected via a bus 640 and communicate with each other.
[0146] The communication interface 630 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0147] Bus 640 includes hardware, software or both, and the parts of electronic equipment are coupled to each other. For example, but not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industrial standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industrial standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 640 may include one or more buses. Although the present application embodiment describes and shows a specific bus, the application considers any suitable bus or interconnection.
[0148] Exemplarily, the electronic device 600 may be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA).
[0149] The electronic device can execute the screen control method in the embodiment of the present application, thereby realizing the combination Figures 1 to 5 Described screen control method and device.
[0150] In addition, in combination with the screen control method in the above embodiment, the embodiment of the present application can provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any one of the screen control methods in the above embodiment is implemented.
[0151] In addition, the embodiment of the present application further provides a vehicle, which may include at least one of the following:
[0152] A screen control device as in any one of the embodiments of the second aspect;
[0153] An electronic device as in any one of the embodiments of the third aspect;
[0154] The computer-readable storage medium in any one of the embodiments of the fourth aspect will not be described in detail here.
[0155] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.
[0156] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0157] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be performed simultaneously.
[0158] The above reference is according to the method of the embodiment of the present application, the flow chart of the device (system) and the computer program product and / or the block diagram described various aspects of the present application.It should be understood that each square box in the flow chart and / or the block diagram and the combination of each square box in the flow chart and / or the block diagram can be realized by computer program instructions.These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the realization of the function / action specified in one or more square boxes of the flow chart and / or the block diagram.Such a processor can be but is not limited to a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit.It can also be understood that each square box in the block diagram and / or the flow chart and the combination of the square boxes in the block diagram and / or the flow chart can also be realized by the dedicated hardware that performs the specified function or action, or can be realized by the combination of dedicated hardware and computer instructions.
[0159] The above is only a specific implementation of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the protection scope of this application.
Claims
1. A screen control method, characterized in that: include: Obtain the palm skeleton information and the hand key point information corresponding to each of the consecutive t-frame gesture images, where t is a positive integer; Using a dynamic gesture recognition model to extract features from the palm skeleton information and the t key point information of the hand to obtain gesture feature information; Performing dynamic gesture recognition on the gesture feature information to obtain a target gesture type; The screen is controlled according to the control instruction corresponding to the target gesture type.
2. The method according to claim 1, characterized in that The dynamic gesture recognition model includes a graph convolutional neural network layer, and the dynamic gesture recognition model is used to extract features from the palm skeleton information and t key point information of the hand to obtain gesture feature information, including: The graph convolutional neural network layer is used to extract features of the palm skeleton information and t hand key point information to obtain gesture feature information.
3. The method according to claim 1, characterized in that The performing dynamic gesture recognition on the gesture feature information to obtain a target gesture type includes: Performing dynamic gesture recognition on the gesture feature information to obtain a plurality of preset gesture types and probabilities corresponding to the plurality of preset gesture types; Determine the probability with the largest value among the multiple probabilities as the target probability; The preset gesture type corresponding to the target probability is determined as the target gesture type.
4. The method according to claim 1 or 2, characterized in that: The hand key point information includes the first coordinates of m hand key points, where m is a positive integer; the feature extraction of the palm skeleton information and the t hand key point information is performed to obtain gesture feature information, including: Subtracting the t×m first coordinates from the first coordinates of the target hand key point in the target image respectively to obtain t×m second coordinates, wherein the target image is any one of the t frames of gesture images, and the target hand key point is any one of the m hand key points; Feature extraction is performed on the palm skeleton information and t×m second coordinates to obtain gesture feature information.
5. The method according to claim 1, characterized in that Before extracting features from the palm skeleton information and the t hand key point information using the dynamic gesture recognition model to obtain gesture feature information, the method further includes: Acquire the palm skeleton information, the hand key point sample information corresponding to each of the consecutive multiple frames of gesture sample images, and the first gesture type corresponding to the consecutive multiple hand key point sample information, wherein the first gesture type is any one of multiple preset gesture types, and the first gesture type is obtained by manually labeling the multiple hand key point sample information; Using an initial dynamic gesture recognition model, feature extraction is performed on the palm skeleton information and a plurality of hand key point sample information to obtain predicted gesture feature information; Performing dynamic gesture recognition on the predicted gesture feature information to obtain a second gesture type; determining a loss function value according to a similarity between the first gesture type and the second gesture type; The model parameters of the initial dynamic gesture recognition model are adjusted according to the loss function value, and the dynamic gesture recognition model is obtained by training.
6. The method according to claim 1, characterized in that Obtaining the hand key point information corresponding to each of the consecutive t-frame gesture images, including: Get t consecutive frames of gesture images; For each frame of the gesture image, performing hand detection on the gesture image to obtain hand position information; Perform hand key point positioning on the hand corresponding to the hand position information to obtain the hand key point information.
7. A screen control device, characterized in that: The device comprises: The first acquisition module is used to obtain palm skeleton information and hand key point information corresponding to each of t consecutive frames of gesture images, where t is a positive integer; A first extraction module is used to extract features of the palm skeleton information and the t hand key point information using a dynamic gesture recognition model to obtain gesture feature information; A first recognition module, used to perform dynamic gesture recognition on the gesture feature information to obtain a target gesture type; The control module is used to control the screen according to the control instruction corresponding to the target gesture type.
8. An electronic device, characterized in that: The electronic device comprises: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the screen control method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the screen control method according to any one of claims 1 to 6 is implemented.
10. A vehicle, characterized in that: Include at least one of the following: The screen control device as claimed in claim 7; The electronic device as claimed in claim 8; The computer readable storage medium of claim 9.
Citation Information
Cited By
Biomechanical modeling and normalized training data generation method and device for human hand posture estimation, equipment and medium
CN120705591A