A gesture monitoring and interaction method, device and equipment based on an intelligent cockpit and a storage medium
By improving the YOLOv7-ResCBAM network and data preprocessing technology, combined with model optimization and encryption processing, the problems of low accuracy, high latency and security in vehicle gesture recognition have been solved, achieving efficient and secure gesture recognition and interaction, and improving user experience and system performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DONGFENG LIUZHOU MOTOR
- Filing Date
- 2025-12-15
- Publication Date
- 2026-05-29
AI Technical Summary
Existing gesture recognition solutions face problems such as low recognition accuracy, high latency, and inability to interact securely in automotive scenarios. In particular, recognition accuracy decreases under complex lighting, occlusion, and changing viewing angles. Furthermore, high-precision models are difficult to run in real time on embedded automotive-grade chips, and fixed user habits pose a high risk of privacy leaks.
It adopts an improved YOLOv7-ResCBAM network structure and data preprocessing technology, combined with sparse pruning and quantization compression optimization model, to acquire image data in real time for preprocessing and gesture recognition, supports user-defined gesture commands, and ensures data security through encryption.
It achieves high-precision gesture recognition, and the optimized model can run at up to 243fps on the domain controller, meeting the requirements of real-time interaction, reducing system power consumption, extending vehicle range, improving ease of use and personalization, and enhancing user experience.
Smart Images

Figure CN122116459A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human-computer interaction technology in smart cockpits, and in particular to a gesture monitoring and interaction method, device, equipment and storage medium based on smart cockpits. Background Technology
[0002] With the rapid development of autonomous driving and smart cockpit technologies, traditional human-computer interaction methods such as buttons and touch controls are no longer sufficient to meet the demands for safety, convenience, and personalization. Gesture recognition, due to its advantages such as intuitiveness, non-contact operation, and natural interaction, has become a research hotspot for next-generation in-vehicle interaction.
[0003] However, existing gesture recognition solutions face several challenges when deployed in automotive scenarios, including decreased recognition accuracy due to complex lighting, occlusion, and changing viewing angles; large parameter sets in high-precision models that are difficult to run in real time on embedded automotive-grade chips; fixed instruction sets that cannot be flexibly adjusted according to user habits; and a lack of effective encryption during the collection, transmission, and storage of gesture data, posing a risk of privacy leaks.
[0004] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The main objective of this invention is to provide a gesture monitoring and interaction method, device, equipment, and storage medium based on a smart cockpit, aiming to solve the technical problems of low gesture recognition accuracy, high latency, and inability to interact securely in an in-vehicle embedded environment.
[0006] To achieve the above objectives, the present invention provides a gesture monitoring and interaction method based on a smart cockpit, the gesture monitoring and interaction method based on a smart cockpit comprising the following steps: Real-time acquisition of image data inside the vehicle cabin; The image data is preprocessed to obtain a preprocessed image; The preprocessed image is input into the gesture recognition model to obtain the gesture recognition result; The gesture recognition model is optimized to obtain an optimized model; Gesture recognition is performed based on the optimized model to obtain the final gesture command, and the target interactive action is executed according to the final gesture command.
[0007] In one embodiment, the step of preprocessing the image data to obtain a preprocessed image includes: The original RGB three-channel pixel values are parsed from the image data; Convert the original pixel values of the RGB three channels into a single-channel grayscale image; The grayscale image is subjected to local illumination compensation to obtain the compensated image; The pixel values of the compensated image are normalized to obtain a normalized image; The normalized image is superimposed with a random perturbation to obtain a perturbed image, which is then used as a preprocessed image.
[0008] In one embodiment, the step of inputting the preprocessed image into a gesture recognition model to obtain a gesture recognition result includes: Feature extraction is performed on the preprocessed image to obtain a gesture feature map; Attention weighting is applied to the gesture feature map to obtain a weighted feature map; Perform residual mapping on the weighted feature map to obtain the residual feature map; The residual feature map is subjected to target detection decoding to obtain gesture bounding box and category information, and the gesture bounding box and category information are used as the gesture recognition result.
[0009] In one embodiment, the step of optimizing the gesture recognition model to obtain an optimized model includes: The weights of each layer of the gesture recognition model are pruned layer by layer according to the sparsity threshold to obtain the pruned weights. The pruned weights are retrained and fine-tuned to obtain the fine-tuned weights. The fine-tuned weights are then quantized and compressed to obtain the optimized model.
[0010] In one embodiment, the method further includes: Receive custom gesture samples entered by the user through the interface; The custom gesture sample is augmented online to obtain the augmented sample; The enhanced sample is bound one-to-one with a preset function to obtain user-specific instructions; During the identification process, the user-specific command is output first to execute the corresponding target interaction action.
[0011] In one embodiment, the method further includes: The final gesture command and its associated information are subjected to symmetric or asymmetric encryption to obtain encrypted data. The encrypted data is sent to the vehicle communication bus, so that the vehicle communication bus forwards the encrypted data to the target node for decryption and verification, and obtains the decrypted instruction.
[0012] In one embodiment, the step of performing gesture recognition based on the optimized model to obtain the final gesture command includes: Non-maximum suppression is applied to the candidate boxes output by the optimized model to obtain the filtered boxes; The filtered boxes are compared with preset category thresholds to obtain qualified boxes; Perform coordinate mapping on the qualified frame to obtain the mapped coordinates; The category label corresponding to the mapped coordinates is used as the final gesture instruction.
[0013] Furthermore, to achieve the above objectives, the present invention also proposes a gesture monitoring and interaction device based on a smart cockpit, the device comprising: The acquisition module is used to acquire image data inside the vehicle cabin in real time; The preprocessing module is used to preprocess the image data to obtain a preprocessed image; The gesture recognition module is used to input the preprocessed image into the gesture recognition model to obtain the gesture recognition result; The model optimization module is used to optimize the gesture recognition model to obtain an optimized model; The interaction module is used to perform gesture recognition based on the optimized model, obtain the final gesture command, and execute the target interaction action according to the final gesture command.
[0014] Furthermore, to achieve the above objectives, the present invention also proposes a gesture monitoring and interaction device based on a smart cockpit, the device comprising: a memory, a processor, and a gesture monitoring and interaction program based on a smart cockpit stored in the memory and executable on the processor, the gesture monitoring and interaction program based on a smart cockpit being configured to implement the steps of the gesture monitoring and interaction method based on a smart cockpit as described above.
[0015] Furthermore, to achieve the above objectives, the present invention also proposes a storage medium storing a gesture monitoring and interaction program based on a smart cockpit, wherein when the gesture monitoring and interaction program based on a smart cockpit is executed by a processor, it implements the steps of the gesture monitoring and interaction method based on a smart cockpit as described above.
[0016] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the gesture monitoring and interaction method based on the smart cockpit described above.
[0017] One or more technical solutions proposed in this application have at least the following technical effects: By improving the YOLOv7-ResCBAM network structure and data preprocessing technology, high-precision gesture recognition is achieved. The optimized model can run at up to 243fps on the domain controller, meeting the requirements of real-time interaction and reducing driver operation latency. The use of lightweight models and efficient algorithm optimization technology reduces system power consumption and extends vehicle range. User-defined gesture commands are supported, improving system usability and personalization, and enhancing user experience. Based on the domain controller architecture design, a high degree of integration of gesture monitoring and interaction systems is achieved, simplifying the vehicle's electronic architecture and reducing system complexity and cost. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating an embodiment of the gesture monitoring and interaction method based on a smart cockpit in this application. Figure 2 This is a flowchart illustrating Embodiment 2 of the gesture monitoring and interaction method based on a smart cockpit in this application. Figure 3 This is a schematic diagram of the module structure of the gesture monitoring and interaction device based on the smart cockpit in an embodiment of this application; Figure 4 This is a schematic diagram of the device structure of the hardware operating environment involved in the gesture monitoring and interaction method based on the smart cockpit in the embodiments of this application.
[0021] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0022] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0023] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0024] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions, such as a gesture monitoring and interaction device based on a smart cockpit. The following description uses a gesture monitoring and interaction device based on a smart cockpit as an example to illustrate this embodiment and the subsequent embodiments.
[0025] Based on this, embodiments of this application provide a gesture monitoring and interaction method based on a smart cockpit, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the gesture monitoring and interaction method based on a smart cockpit in this application.
[0026] In this embodiment, the gesture monitoring and interaction method based on the smart cockpit includes steps S10 to S50: Step S10: Acquire image data of the vehicle cabin in real time; It should be noted that the purpose of this step is to register the original image signals inside the vehicle cabin into the controller cache, forming a unified data record for subsequent recognition and interaction.
[0027] Image data refers to the raw pixel set acquired in real time by an RGB camera without any format conversion or compression; vehicle cabin refers to the space in front of the driver's seat and passenger seat.
[0028] Understandably, the original image registration and cache writing are completed by reading the pixel stream from the camera interface and matching it with the preset cache address, followed by sampling period latching and timing alignment.
[0029] In its implementation, during the data acquisition phase, this invention uses a high-definition RGB camera to capture the driver's gestures in the cockpit in real time, ensuring that the acquired image data accurately reflects the driver's hand gestures. The acquired videos are preprocessed using the Mediapipe algorithm to remove images without gestures, ultimately constructing a dataset containing 20,000 different gesture images, covering 12 interactive gestures and 6 dangerous behavior gestures, providing rich data support for subsequent model training.
[0030] Step S20: Preprocess the image data to obtain a preprocessed image; It should be noted that the purpose of this step is to convert the raw image data into a preprocessed image, remove the effects of sudden changes in illumination, scale differences, and insufficient samples, and provide a unified format for network input.
[0031] A preprocessed image refers to a numerical image that has undergone grayscale conversion, local compensation, normalization, and perturbation; preprocessing refers to the action of transforming the original pixels according to a preset algorithm.
[0032] Understandably, format unification and cache writing are accomplished by reading the original image data and matching it with preset algorithm parameters.
[0033] In the specific implementation, the RGB three-channel image is converted into a single-channel grayscale image, and a contrast-limited adaptive histogram equalization algorithm is used for local illumination compensation; the pixel values are normalized to zero mean and unit variance, and the principle formula is as follows:
[0034] in, , The mean and standard deviation of the ROI pixels in the current frame are used to prevent gradient explosion caused by sudden changes in illumination; Gaussian noise, motion blur, downsampling and upsampling are randomly superimposed online to expand the training samples and improve the robustness of the model to low-resolution and shaky images.
[0035] In one feasible implementation, step S20 includes steps A11 to A15: Step A11: Extract the original RGB three-channel pixel values from the image data; It should be noted that the purpose of this step is to parse the original pixel values into calculable RGB components to avoid feature loss due to unclear format in subsequent processing.
[0036] RGB three-channel raw pixel values refer to the set of red, green, and blue components before splitting; parsing refers to the action of splitting the three components from the interface stream according to a preset bit width.
[0037] It is understandable that by reading the interface data packet and matching it with the preset bit width, and then splitting and writing it back through the channel, the RGB component parsing and cache writing are completed.
[0038] Step A12: Convert the original RGB three-channel pixel values into a single-channel grayscale image; It should be noted that the purpose of this step is to convert the RGB components into single-channel grayscale, thereby reducing the input dimensions of the subsequent network and reducing the amount of computation.
[0039] A single-channel grayscale image refers to a brightness image obtained by weighting and summing the RGB components according to preset weights; conversion refers to the action of weighting and summing according to preset coefficients.
[0040] It is understandable that grayscale conversion and cache writing are completed by reading RGB components, matching them with preset weights, and then performing weighted summation and writing back.
[0041] Step A13: Perform local illumination compensation on the grayscale image to obtain the compensated image; It should be noted that the purpose of this step is to balance the local brightness differences in the grayscale image to avoid feature distortion caused by sudden changes in illumination.
[0042] Local illumination compensation refers to the action of equalizing the brightness of a local area according to a preset window; the compensated image refers to the brightness map after equalization.
[0043] Understandably, by reading the grayscale image and matching it with a preset window, and then writing it back using a balancing algorithm, local compensation and cache writing are completed.
[0044] Step A14: Normalize the pixel values of the compensated image to obtain a normalized image; It should be noted that the purpose of this step is to map the pixel values of the compensated image to a uniform scale, so as to avoid gradient explosion caused by excessive numerical differences.
[0045] Normalization refers to the linear scaling of pixel values according to a preset mean and standard deviation; a normalized image refers to a numerical image after scaling.
[0046] Understandably, normalization and cache writing are completed by reading the compensated image and matching it with the preset mean and standard deviation, and then writing it back after linear scaling.
[0047] Step A15: Add random perturbation to the normalized image to obtain the perturbed image, and use the perturbed image as the preprocessed image.
[0048] It should be noted that the purpose of this step is to overlay random perturbations onto the normalized image, thereby increasing sample diversity and improving model robustness.
[0049] Random perturbation refers to the action of adding noise, blurring, or scaling to pixel values according to a preset probability; perturbed image refers to the numerical image after perturbation.
[0050] It is understandable that by reading the normalized image and matching it with preset perturbation parameters, and then randomly superimposing and writing it back, the perturbation generation and cache writing are completed.
[0051] Step S30: Input the preprocessed image into the gesture recognition model to obtain the gesture recognition result; It should be noted that the purpose of this step is to convert the preprocessed image into gesture bounding boxes and category information to form a recognizable quantification result.
[0052] Gesture recognition results refer to a structure containing bounding box coordinates and category labels; gesture recognition models refer to the unoptimized set of raw network weights.
[0053] Understandably, the recognition results are generated and cached by reading the preprocessed image and matching it with preset network weights.
[0054] It should be understood that during the model training phase, this invention proposes an improved YOLOv7-ResCBAM network structure. By introducing the CBAM attention mechanism and residual structure, the model's ability to extract deep features is enhanced, improving the accuracy and robustness of gesture recognition. Experimental results show that the improved model achieves 98.6% on the mAP50 metric, a 0.3% improvement compared to the original YOLOv7 model, and also shows improvement across various evaluation metrics, demonstrating the effectiveness of the proposed method.
[0055] In one feasible implementation, step S30 includes steps A21 to A24: Step A21: Extract features from the preprocessed image to obtain a gesture feature map; It should be noted that the purpose of this step is to map the perturbed image into high-dimensional gesture features, providing input for subsequent attention weighting.
[0056] Gesture feature map refers to a high-dimensional tensor after convolution operation; feature extraction refers to the mapping of actions according to a preset convolution kernel.
[0057] Understandably, feature extraction and cache writing are completed by reading the perturbed image, matching it with a preset convolutional kernel, and then writing it back after mapping operations.
[0058] Step A22: Apply attention weighting to the gesture feature map to obtain a weighted feature map; It should be noted that the purpose of this step is to reweight the gesture feature map according to attention weights to highlight key area information.
[0059] Weighted feature map refers to the tensor after being multiplied and added by attention coefficients; attention weighting refers to the action of multiplying and adding channels and space according to preset coefficients.
[0060] Understandably, by reading the gesture feature map and matching it with the preset attention coefficient, and then writing it back through multiplication and addition operations, the weighted generation and cache writing are completed.
[0061] Step A23: Perform residual mapping on the weighted feature map to obtain the residual feature map; It should be noted that the purpose of this step is to perform residual mapping on the weighted feature map, alleviate deep gradient vanishing, and enhance the feature representation capability.
[0062] The residual feature map refers to the tensor after being summed by skipping layers; the residual mapping refers to the summing action performed according to a preset skipping layer.
[0063] It is understandable that by reading the weighted feature map and matching it with the preset skip coefficients, and then performing a summation operation to write it back, the residual generation and cache writing are completed.
[0064] In the specific implementation, a gesture recognition algorithm based on the YOLOv7-ResCBAM network structure is adopted. By introducing the CBAM attention mechanism and residual structure, the ability to extract gesture features is enhanced, thereby improving the accuracy and robustness of gesture recognition.
[0065] Step A24: Perform target detection decoding on the residual feature map to obtain the gesture bounding box and category information, and use the gesture bounding box and category information as the gesture recognition result.
[0066] It should be noted that the purpose of this step is to decode the residual feature map into bounding boxes and categories, forming a recognizable quantization result.
[0067] Gesture bounding box and category information refers to a structure containing coordinates and labels; object detection decoding refers to the action of regression and classification based on preset anchor boxes.
[0068] Understandably, the decoding generation and cache writing are completed by reading the residual feature map and matching it with the preset anchor box, and then writing it back after regression and classification operations.
[0069] Step S40: Optimize the gesture recognition model to obtain the optimized model; It should be noted that the purpose of this step is to prune, fine-tune, and quantize the original network weights to obtain a smaller, faster optimized model, ensuring real-time embedded operation.
[0070] The optimized model refers to the set of weights after pruning, fine-tuning, and quantization; optimization refers to the action of transforming the weights according to the preset sparsity and quantization bit width.
[0071] Understandably, by reading the original network weights and matching them with a preset sparsity threshold, and then rewriting them back through a refinement process, the optimization generation and cache writing are completed.
[0072] It should be understood that, to further optimize model performance, this invention also investigated model pruning techniques and proposed an unstructured pruning algorithm based on layer adaptive sparsity. Using the LAMP pruning method, the number of model parameters was reduced from 35.7M to 2.42M, the computational cost from 105.6 GFLOPs to 26.0 GFLOPs, and the model file size from 71.52MB to 5.18MB, resulting in an overall size reduction of 83%. Accuracy reduction remained within acceptable limits, and the response speed reached 243fps, meeting the requirements for embedded device deployment.
[0073] Step S50: Perform gesture recognition based on the optimized model to obtain the final gesture command, and execute the target interactive action according to the final gesture command.
[0074] It should be noted that the purpose of this step is to put the optimized model back into the recognition process to complete the final gesture command generation and trigger the interaction, so as to achieve gesture control without delay or interruption.
[0075] The final gesture command refers to the executable code after post-processing, encryption, and customization; the target interaction action refers to the specific function call of the instrument panel, central control, or vehicle system.
[0076] Understandably, by reading the optimized model and matching it with preset post-processing parameters, and then through encryption and custom priority refinement steps, the instruction generation and cache writing are completed for forwarding and execution by the vehicle communication bus.
[0077] In one feasible implementation, step S50 includes steps A31 to A34: Step A31: Perform non-maximum suppression on the candidate boxes output by the optimized model to obtain the filtered boxes; It should be noted that the purpose of this step is to filter the candidate boxes output by the optimized model, remove overlapping boxes, and reduce redundant information.
[0078] Candidate boxes refer to the set of unfiltered bounding boxes; non-maximum suppression refers to the action of deduplication according to a preset intersection-union ratio threshold.
[0079] Understandably, by reading candidate boxes and matching them with preset intersection and comparison thresholds, and then performing deduplication and writing back, the filtering generation and cache writing are completed.
[0080] Step A32: Compare the filtered boxes with the preset category thresholds to obtain qualified boxes; It should be noted that the purpose of this step is to compare the filtered boxes with the preset category thresholds, remove low-confidence boxes, and retain high-confidence results.
[0081] The filtered bounding boxes refer to the set of bounding boxes after deduplication; the preset category threshold refers to the confidence lower limit value obtained from the calibration; the qualified boxes refer to the high-confidence boxes after comparison.
[0082] Understandably, by reading the filtered box and matching it with the preset category threshold, and then performing a comparison operation to write it back, the threshold filtering and cache writing are completed.
[0083] Step A33: Perform coordinate mapping on the qualified frame to obtain the mapped coordinates; It should be noted that the purpose of this step is to map the coordinates of the qualified bounding box to the scale of the input image, ensuring that the label corresponds one-to-one with the pixel.
[0084] A qualified bounding box refers to a high-confidence bounding box after threshold filtering; coordinate mapping refers to the action of transforming coordinates according to a preset scaling factor; mapped coordinates refer to the pixel coordinates after transformation.
[0085] Understandably, by reading the qualified frame and matching it with the preset scaling factor, and then writing it back after coordinate transformation, the mapping generation and cache writing are completed.
[0086] Step A34: Use the category label corresponding to the mapped coordinates as the final gesture command.
[0087] It should be noted that the purpose of this step is to assign values to the category labels corresponding to the mapped coordinates, forming a final instruction that can be issued.
[0088] Mapped coordinates refer to the transformed pixel coordinates; category labels refer to the text or code output by the model; final gesture commands refer to the text or code commands that can be issued.
[0089] Understandably, by reading the mapped coordinates and matching them with the preset label table, and then performing assignment operations to write back the values, the instruction generation and cache writing are completed for subsequent interactive execution.
[0090] Furthermore, symmetric or asymmetric encryption is performed on the final gesture command and its associated information to obtain encrypted data; The encrypted data is sent to the vehicle communication bus, which then forwards the encrypted data to the target node for decryption and verification, and obtains the decrypted instructions.
[0091] It should be noted that the purpose of this paragraph is to encrypt the final gesture command and its associated information before sending it to prevent eavesdropping or tampering on the bus.
[0092] The final gesture command's ancillary information refers to the set of coordinates, confidence level, and timestamp; symmetric or asymmetric encryption processing refers to the action of transforming data according to a preset key; encrypted data refers to the unreadable data after transformation; the target node refers to the receiving module of the instrument or central control system; decryption verification refers to the action of inversely transforming data according to a preset key.
[0093] Understandably, by reading the final gesture command and its associated information and matching it with a preset key, and then performing encryption operations to write it back, the encryption generation and cache writing are completed for subsequent bus forwarding.
[0094] In the specific implementation, identity authentication is performed through a camera; a key generated by ECDH-secp256r1 is used for session encryption; the key and integer are stored in a secure storage unit.
[0095] Furthermore, this plan also includes: Receive custom gesture samples entered by the user through the interface; Online data augmentation is performed on custom gesture samples to obtain augmented samples; The enhanced sample is bound one-to-one with the preset function to obtain user-specific instructions; During recognition, user-specific commands are output first to execute the corresponding target interaction action.
[0096] It should be noted that the purpose of this paragraph is to bind user-inputted custom gesture samples to preset functions to achieve personalized interaction.
[0097] Custom gesture samples refer to the set of images or coordinates entered by the user through the interface; online data augmentation refers to the actions that transform the samples according to preset perturbations; augmented samples refer to the set of samples after transformation; user-specific commands refer to text or codes that are bound one-to-one with the function; priority output refers to actions that are called first in the recognition queue.
[0098] Understandably, by reading custom gesture samples and matching them with preset perturbation parameters, and then enhancing and binding them back, exclusive instructions are generated and cached for subsequent priority recognition.
[0099] This solution also proposes a smart cockpit gesture monitoring and interaction system based on YOLOv7. The system comprises: a data acquisition module for real-time acquisition of image data within the vehicle cockpit, the data acquisition module including an RGB camera, transmitting image data to a processing module via MIPI and CSI interfaces; a storage module for storing gesture recognition models, optimized model parameters, pruned model parameters, and gesture recognition results, the storage module further including a sub-module for storing gesture datasets and model training parameters; and a processing module comprising: the data acquisition module for real-time acquisition of image data within the vehicle cockpit, the data acquisition module including an RGB camera, transmitting image data to a processing module via MIPI and CSI interfaces; and a gesture recognition sub-module based on an improved YOLOv7... The OLOv7-ResCBAM network structure performs gesture recognition on acquired image data using deep learning algorithms. This network structure enhances deep feature extraction and improves the accuracy and robustness of gesture recognition by incorporating the CBAM attention mechanism and residual structure. The model optimization module employs a layer-adaptive sparsity-based unstructured pruning algorithm to optimize the gesture recognition model, reducing model size and parameter redundancy and improving operational efficiency. The model conversion module converts the trained gesture recognition model to ONNX format and further to an NCNN model for efficient inference on embedded devices. The interaction module includes a customizable instruction unit to provide a personalized interactive experience. The communication module includes a data encryption unit to ensure the security and privacy of data transmission and storage.
[0100] The data acquisition module also includes a preprocessing unit for performing grayscale processing, normalization, and data augmentation on the acquired image data to improve the accuracy and robustness of gesture recognition. The preprocessing unit can automatically adjust the brightness and contrast of the image to adapt to different lighting conditions and generate more training samples through data augmentation techniques, thereby improving the model's generalization ability. The normalization processing can be expressed as: Where I is the original image pixel value, μ is the image pixel mean, and σ is the image pixel standard deviation. These are the normalized image pixel values.
[0101] The gesture recognition submodule also includes a model training unit for training the improved YOLOv7-ResCBAM network structure using a self-collected cockpit domain gesture dataset. This dataset includes various gesture categories, such as "like," "dislike," and "gesture 1" through "gesture 9," as well as dangerous driving behavior gestures including "using a mobile phone," "eating," "drinking water," and "smoking." The model training unit can automatically adjust the learning rate and optimize training parameters to ensure high accuracy and low false recognition rate across different gesture categories.
[0102] The model optimization module also includes a model fine-tuning unit, used to fine-tune the model after pruning to restore or improve its performance. The fine-tuning process includes adjusting the learning rate, optimizing training parameters, and retraining the model to ensure that the pruned model maintains high recognition accuracy while operating efficiently. The model optimization submodule can also use the Grad-CAM visualization analysis method to compare and analyze the models before and after pruning to verify the pruning effect and the improvement in model performance.
[0103] The sparsity of the pruning algorithm can be expressed as: ,in, This represents the total number of parameters in the original model; This represents the number of parameters retained after pruning; S represents the sparsity level, used to measure the degree of simplification of model parameters.
[0104] The interaction module also includes a command customization unit, allowing users to customize gesture commands according to their personal preferences. These customized gesture commands are set through the user interface and stored in the storage module. The interaction module can also automatically adjust the response speed and sensitivity of gesture commands based on the user's operating habits and preferences to provide a more personalized interactive experience.
[0105] The communication module also includes a data encryption unit for encrypting transmitted data to ensure data security and privacy. The encryption process employs symmetric or asymmetric encryption algorithms, effectively preventing data theft or tampering during transmission. The communication module also verifies the identity of the vehicle control system or driver's mobile device through a security authentication mechanism, ensuring the reliability and security of data transmission.
[0106] The model optimization module includes: a layer adaptive sparsity pruner, which calculates the global pruning threshold based on the target sparsity S and performs model sparsity pruning; a partial freeze strategy is used to fine-tune and restore the pruned model; and KL three-dimensional calibration is used to perform model quantization compression.
[0107] The custom instruction design of the interaction module includes: sampling and recording gesture instructions through a camera; enhancing and expanding key samples; and binding gesture instructions from the function list according to user needs.
[0108] The data acquisition module also includes a preprocessing unit for performing grayscale processing, normalization, and data augmentation on the acquired image data to improve the accuracy and robustness of gesture recognition. The preprocessing unit can automatically adjust the brightness and contrast of the image to adapt to different lighting conditions and generate more training samples through data augmentation techniques, thereby improving the model's generalization ability. The normalization processing can be expressed as: Where I is the original image pixel value, μ is the image pixel mean, and σ is the image pixel standard deviation. These are the normalized image pixel values.
[0109] The storage module is used to store the gesture recognition model, optimized model parameters, pruned model parameters, and gesture recognition results. The storage module also includes a sub-module for storing the gesture dataset and model training parameters.
[0110] The gesture recognition submodule also includes a model training unit for training the improved YOLOv7-ResCBAM network structure using a self-collected cockpit domain gesture dataset. This dataset includes various gesture categories, such as "like," "dislike," and "gesture 1" through "gesture 9," as well as dangerous driving behavior gestures including "using a mobile phone," "eating," "drinking water," and "smoking." The model training unit can automatically adjust the learning rate and optimize training parameters to ensure high accuracy and low false recognition rate across different gesture categories.
[0111] The processing module includes: a data acquisition module for real-time acquisition of image data inside the vehicle cabin, the data acquisition module including an RGB camera, transmitting image data to the processing module via MIPI and CSI interfaces; a gesture recognition submodule, based on an improved YOLOv7-ResCBAM network structure, performing gesture recognition on the acquired image data using a deep learning algorithm, the network structure enhancing deep feature extraction and improving the accuracy and robustness of gesture recognition by incorporating CBAM attention mechanism and residual structure; a model optimization module, employing a layer-adaptive sparsity-based unstructured pruning algorithm to optimize the gesture recognition model, reducing model size and parameter redundancy, and improving model operating efficiency; and a model conversion module, used to convert the trained gesture recognition model into ONNX format and further into an NCNN model for efficient inference on embedded devices.
[0112] The model optimization module also includes a model fine-tuning unit, used to fine-tune the model after pruning to restore or improve its performance. The fine-tuning process includes adjusting the learning rate, optimizing training parameters, and retraining the model to ensure that the pruned model maintains high recognition accuracy while operating efficiently. The model optimization submodule can also use the Grad-CAM visualization analysis method to compare and analyze the model before and after pruning to verify the pruning effect and the improvement in model performance.
[0113] The sparsity of the pruning algorithm can be expressed as: ,in, This represents the total number of parameters in the original model; This represents the number of parameters retained after pruning; S represents the sparsity level, used to measure the degree of simplification of model parameters.
[0114] The interaction module, including a command customization unit, provides a personalized interactive experience; the communication module, including a data encryption unit, ensures the security and privacy of data transmission and storage.
[0115] The interaction module also includes a command customization unit, allowing users to customize gesture commands according to their personal preferences. These customized gesture commands are set through the user interface and stored in the storage module. The interaction module can also automatically adjust the response speed and sensitivity of gesture commands based on the user's operating habits and preferences to provide a more personalized interactive experience.
[0116] The communication module also includes a data encryption unit for encrypting transmitted data to ensure data security and privacy. The encryption process employs symmetric or asymmetric encryption algorithms, effectively preventing data theft or tampering during transmission. The communication module also verifies the identity of the vehicle control system or driver's mobile device through a security authentication mechanism, ensuring the reliability and security of data transmission.
[0117] This embodiment provides a gesture monitoring and interaction method based on a smart cockpit. It not only improves the accuracy and real-time performance of gesture recognition but also optimizes model performance through model pruning techniques, making it more suitable for running on embedded devices. This system is of great significance for enhancing the human-computer interaction experience and driving safety in smart cockpits, providing strong technical support for the development of intelligent vehicles.
[0118] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 Step S40 includes steps S401 to S404: Step S401: Prune the weights of each layer of the gesture recognition model layer by layer according to the sparsity threshold to obtain the pruned weights. It should be noted that the purpose of this step is to prune the weights of each layer of the gesture recognition model layer by layer according to the preset sparsity threshold, remove redundant parameters, and reduce the model size.
[0119] The weights of each layer refer to the specific values in the convolution kernels and fully connected matrices of the model; the sparsity threshold refers to the maximum allowable non-zero proportion of each layer obtained from the calibration; the weights after pruning refer to the set of values after being zeroed out.
[0120] It is understandable that by reading the weights of each layer and comparing them with the preset sparsity threshold, and then writing them back after zeroing, the layer-by-layer pruning and cache writing are completed.
[0121] Step S402: Retrain and fine-tune the pruned weights to obtain the fine-tuned weights; It should be noted that the purpose of this step is to reinvest the pruned weights into the training, restore the accuracy lost due to setting them to zero, and ensure that the recognition performance is not reduced.
[0122] Post-pruning weights refer to the set of values after being set to zero; retraining fine-tuning refers to the action of updating non-zero weights according to a preset learning rate; fine-tuned weights refer to the set of updated values.
[0123] Understandably, by reading the pruned weights and matching them with the preset learning rate, and then iteratively updating and writing them back, the fine-tuning recovery and cache writing are completed.
[0124] Step S403: Quantize and compress the fine-tuned weights to obtain the optimized model.
[0125] It should be noted that the purpose of this step is to quantize and compress the fine-tuned weights to further reduce the size and improve the embedded running speed.
[0126] Fine-tuned weights refer to the set of values after fine-tuning; quantization compression refers to the action of discretizing the weights according to a preset bit width; optimized model refers to the complete set of weights after discretization.
[0127] Understandably, by reading the fine-tuned weights and matching them with the preset bit width, and then performing discretization operations to write them back, quantization compression and cache writing are completed.
[0128] It should be understood that the model storage area contains fixed files; the dataset sub-library is used to store cockpit domain gesture images and corresponding labels; the training parameter log records training process parameters; the user-defined table supports online reading and writing by the vehicle's UI; and the recognition result cache is used for debugging and functional safety backtracking.
[0129] In the implementation, an unstructured pruning algorithm based on layer adaptive sparsity is used to prune and optimize the gesture recognition model, significantly reducing the number of parameters and computational cost, and improving the model's running efficiency on the domain controller. The pruning formula is as follows: ,
[0130] Where MOP stands for Mathematical Optimization Probability, and MOA stands for Accelerated Mathematical Optimizer. This represents the current iteration number. α represents the maximum number of iterations, and α is a parameter that controls the descent curvature of the MOP.
[0131] This embodiment provides a gesture monitoring and interaction method based on a smart cockpit. Through layer-by-layer pruning, fine-tuning recovery and quantization compression processes, the system significantly reduces the model size and improves the running speed while ensuring recognition accuracy. This enables the gesture recognition model to run stably in the vehicle-mounted embedded chip, achieving high-precision, low-latency gesture recognition and interaction.
[0132] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the gesture monitoring and interaction method based on the smart cockpit in this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0133] This application also provides a gesture monitoring and interaction device based on a smart cockpit, please refer to... Figure 3 The gesture monitoring and interaction devices based on the smart cockpit include: Acquisition module 10 is used to acquire image data inside the vehicle cabin in real time; Preprocessing module 20 is used to preprocess image data to obtain a preprocessed image; The gesture recognition module 30 is used to input the preprocessed image into the gesture recognition model to obtain the gesture recognition result; The model optimization module 40 is used to optimize the gesture recognition model to obtain the optimized model; The interaction module 50 is used to perform gesture recognition based on the optimized model, obtain the final gesture command, and execute the target interaction action according to the final gesture command.
[0134] The gesture monitoring and interaction device based on a smart cockpit provided in this application, employing the gesture monitoring and interaction method based on a smart cockpit in the above embodiments, can solve the technical problems of low gesture recognition accuracy, high latency, and inability to interact securely in an in-vehicle embedded environment. Compared with the prior art, the beneficial effects of the gesture monitoring and interaction device based on a smart cockpit provided in this application are the same as those of the gesture monitoring and interaction method based on a smart cockpit provided in the above embodiments, and other technical features in the gesture monitoring and interaction device based on a smart cockpit are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0135] In one embodiment, the preprocessing module 20 is further configured to parse the original RGB three-channel pixel values from the image data; Convert the original pixel values of the RGB three channels into a single-channel grayscale image; Local illumination compensation is performed on the grayscale image to obtain the compensated image; The pixel values of the compensated image are normalized to obtain a normalized image; A random perturbation is superimposed on the normalized image to obtain a perturbed image, which is then used as the preprocessed image.
[0136] In one embodiment, the gesture recognition module 30 is further configured to extract features from the preprocessed image to obtain a gesture feature map; Attention weights are applied to the gesture feature map to obtain a weighted feature map; Perform residual mapping on the weighted feature map to obtain the residual feature map; The residual feature map is subjected to target detection decoding to obtain the gesture bounding box and category information, which are then used as the gesture recognition result.
[0137] In one embodiment, the model optimization module 40 is further used to prune the weights of each layer of the gesture recognition model layer by layer according to the sparsity threshold to obtain the pruned weights. The pruned weights are retrained and fine-tuned to obtain the fine-tuned weights. The fine-tuned weights are then quantized and compressed to obtain the optimized model.
[0138] In one embodiment, the interaction module 50 is further configured to receive custom gesture samples entered by the user through the interface; Online data augmentation is performed on custom gesture samples to obtain augmented samples; The enhanced sample is bound one-to-one with the preset function to obtain user-specific instructions; During recognition, user-specific commands are output first to execute the corresponding target interaction action.
[0139] In one embodiment, the interaction module 50 is further configured to perform symmetric encryption or asymmetric encryption on the final gesture instruction and its associated information to obtain encrypted data. The encrypted data is sent to the vehicle communication bus, which then forwards the encrypted data to the target node for decryption and verification, and obtains the decrypted instructions.
[0140] In one embodiment, the interaction module 50 is further configured to perform non-maximum suppression on the candidate boxes output by the optimized model to obtain the filtered boxes. The filtered boxes are compared with the preset category thresholds to obtain qualified boxes; Perform coordinate mapping on the qualified frame to obtain the mapped coordinates; The category label corresponding to the mapped coordinates is used as the final gesture command.
[0141] This application provides a gesture monitoring and interaction device based on a smart cockpit. The gesture monitoring and interaction device based on a smart cockpit includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the gesture monitoring and interaction method based on a smart cockpit in the first embodiment described above.
[0142] The following is for reference. Figure 4 This document illustrates a structural schematic diagram of a gesture monitoring and interaction device based on a smart cockpit, suitable for implementing embodiments of this application. The gesture monitoring and interaction device based on a smart cockpit in these embodiments may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 4 The gesture monitoring and interaction device based on the smart cockpit shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0143] like Figure 4 As shown, the gesture monitoring and interaction device based on the smart cockpit may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to programs stored in ROM (Read Only Memory) 1002 or programs loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the gesture monitoring and interaction device based on the smart cockpit. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the smart cockpit-based gesture monitoring and interaction device to exchange data wirelessly or via wired communication with other devices. Although the figure shows a smart cockpit-based gesture monitoring and interaction device with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.
[0144] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0145] The gesture monitoring and interaction device based on a smart cockpit provided in this application, employing the gesture monitoring and interaction method based on a smart cockpit in the above embodiments, can solve the technical problems of low gesture recognition accuracy, high latency, and inability to interact securely in an in-vehicle embedded environment. Compared with the prior art, the beneficial effects of the gesture monitoring and interaction device based on a smart cockpit provided in this application are the same as those of the gesture monitoring and interaction method based on a smart cockpit provided in the above embodiments, and other technical features in this gesture monitoring and interaction device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.
[0146] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0147] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0148] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the gesture monitoring and interaction method based on the smart cockpit in the above embodiments.
[0149] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory or Flash Memory), optical fibers, CD-ROM (CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0150] The aforementioned computer-readable storage medium may be included in a gesture monitoring and interaction device based on a smart cockpit; or it may exist independently and not be installed in a gesture monitoring and interaction device based on a smart cockpit.
[0151] The aforementioned computer-readable storage medium carries one or more programs. When the aforementioned one or more programs are executed by the gesture monitoring and interaction device based on the smart cockpit, the gesture monitoring and interaction device based on the smart cockpit: acquires image data inside the vehicle cockpit in real time; preprocesses the image data to obtain a preprocessed image; inputs the preprocessed image into a gesture recognition model to obtain a gesture recognition result; optimizes the gesture recognition model to obtain an optimized model; performs gesture recognition based on the optimized model to obtain a final gesture command; and executes the target interactive action according to the final gesture command.
[0152] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including LAN (Local Area Network) or WAN (Wide Area Network)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0153] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0154] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0155] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for executing the above-described gesture monitoring and interaction method based on a smart cockpit. This solves the technical problems of low gesture recognition accuracy, high latency, and inability to perform secure interaction in an in-vehicle embedded environment. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the gesture monitoring and interaction method based on a smart cockpit provided in the above embodiments, and will not be repeated here.
[0156] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the gesture monitoring and interaction method based on a smart cockpit as described above.
[0157] The computer program product provided in this application can solve the technical problems of low gesture recognition accuracy, high latency, and inability to interact securely in vehicle-mounted embedded environments. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the gesture monitoring and interaction method based on smart cockpit provided in the above embodiments, and will not be repeated here.
[0158] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A gesture monitoring and interaction method based on a smart cockpit, characterized in that, The method includes: Real-time acquisition of image data inside the vehicle cabin; The image data is preprocessed to obtain a preprocessed image; The preprocessed image is input into the gesture recognition model to obtain the gesture recognition result; The gesture recognition model is optimized to obtain an optimized model; Gesture recognition is performed based on the optimized model to obtain the final gesture command, and the target interactive action is executed according to the final gesture command.
2. The method as described in claim 1, characterized in that, The step of preprocessing the image data to obtain a preprocessed image includes: The original RGB three-channel pixel values are parsed from the image data; Convert the original pixel values of the RGB three channels into a single-channel grayscale image; The grayscale image is subjected to local illumination compensation to obtain the compensated image; The pixel values of the compensated image are normalized to obtain a normalized image; The normalized image is superimposed with a random perturbation to obtain a perturbed image, which is then used as a preprocessed image.
3. The method as described in claim 1, characterized in that, The step of inputting the preprocessed image into the gesture recognition model to obtain the gesture recognition result includes: Feature extraction is performed on the preprocessed image to obtain a gesture feature map; Attention weighting is applied to the gesture feature map to obtain a weighted feature map; Perform residual mapping on the weighted feature map to obtain the residual feature map; The residual feature map is subjected to target detection decoding to obtain gesture bounding box and category information, and the gesture bounding box and category information are used as the gesture recognition result.
4. The method as described in claim 1, characterized in that, The step of optimizing the gesture recognition model to obtain the optimized model includes: The weights of each layer of the gesture recognition model are pruned layer by layer according to the sparsity threshold to obtain the pruned weights. The pruned weights are retrained and fine-tuned to obtain the fine-tuned weights. The fine-tuned weights are then quantized and compressed to obtain the optimized model.
5. The method as described in claim 1, characterized in that, The method further includes: Receive custom gesture samples entered by the user through the interface; The custom gesture sample is augmented online to obtain the augmented sample; The enhanced sample is bound one-to-one with a preset function to obtain user-specific instructions; During the identification process, the user-specific command is output first to execute the corresponding target interaction action.
6. The method as described in claim 1, characterized in that, The method further includes: The final gesture command and its associated information are subjected to symmetric or asymmetric encryption to obtain encrypted data. The encrypted data is sent to the vehicle communication bus, so that the vehicle communication bus forwards the encrypted data to the target node for decryption and verification, and obtains the decrypted instruction.
7. The method as described in claim 1, characterized in that, The step of performing gesture recognition based on the optimized model to obtain the final gesture command includes: Non-maximum suppression is applied to the candidate boxes output by the optimized model to obtain the filtered boxes; The filtered boxes are compared with preset category thresholds to obtain qualified boxes; Perform coordinate mapping on the qualified frame to obtain the mapped coordinates; The category label corresponding to the mapped coordinates is used as the final gesture instruction.
8. A gesture monitoring and interaction device based on a smart cockpit, characterized in that, The device includes: The acquisition module is used to acquire image data inside the vehicle cabin in real time; The preprocessing module is used to preprocess the image data to obtain a preprocessed image; The gesture recognition module is used to input the preprocessed image into the gesture recognition model to obtain the gesture recognition result; The model optimization module is used to optimize the gesture recognition model to obtain an optimized model; The interaction module is used to perform gesture recognition based on the optimized model, obtain the final gesture command, and execute the target interaction action according to the final gesture command.
9. A gesture monitoring and interaction device based on a smart cockpit, characterized in that, The device includes: a memory, a processor, and a smart cockpit-based gesture monitoring and interaction program stored in the memory and executable on the processor, the smart cockpit-based gesture monitoring and interaction program being configured to implement the steps of the smart cockpit-based gesture monitoring and interaction method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium stores a gesture monitoring and interaction program based on a smart cockpit. When the gesture monitoring and interaction program based on a smart cockpit is executed by the processor, it implements the steps of the gesture monitoring and interaction method based on a smart cockpit as described in any one of claims 1 to 7.