A monocular camera-based vehicle-mounted gesture interaction method and system

CN116661594BActive Publication Date: 2026-09-15HEFEI UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202310547999.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-15
Publication Date
2026-09-15
Estimated Expiration
2043-05-15

AI Technical Summary

Benefits of technology

[0026] Compared with the prior art, the beneficial effects of the present invention include:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116661594B_ABST
    Figure CN116661594B_ABST
Patent Text Reader

Abstract

The application provides a kind of vehicle-mounted gesture interaction method and system based on monocular camera, belong to vehicle-mounted man-machine interaction and gesture recognition technical field.The vehicle-mounted gesture interaction method includes: monocular camera acquires image, in turn through hand recognition, segmentation, feature extraction, 2D detection, 3D detection and model fitting obtains gesture three-dimensional reconstruction model, then gesture three-dimensional reconstruction model is rendered as gesture depth map and gesture two-dimensional diagram, utilize interframe difference method to detect the three-dimensional space motion trajectory of hand, utilize pre-stored gesture dataset to match gesture, finally generate control instruction to vehicle-mounted multimedia system and man-machine interaction system.The method utilizes monocular camera to carry out gesture three-dimensional estimation, and fits out the image of multiple hands, compared with traditional gesture interaction method has lower equipment cost, better recognition accuracy and robustness, greater upgrade cost advantage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle-mounted human-computer interaction and gesture recognition technology, specifically to a vehicle-mounted gesture interaction method and system based on a monocular camera. Background Technology

[0002] Human-computer interaction (HCI) is a technology that studies the exchange of information between humans and computers. This technology involves multiple fields, including information science and intelligent science, and is guiding a hot research direction in 21st-century information and computer science. In recent years, with the development of artificial intelligence and related fields, HCI methods are moving towards a more natural and universal approach. There is an urgent need to use natural movements rather than traditional dedicated input devices and control systems to interact with digital content in virtual environments. HCI is shifting from a computer-centric to a user-centric approach. The human hand, as the most flexible organ in the human body, plays a crucial role in the field of HCI.

[0003] While gesture-based human-computer interaction technology has significant research and application value, as a new technological field, many problems remain to be solved, such as accurate hand position recognition and segmentation, and accurate finger recognition in self-occlusion situations. Currently, common solutions in the field of gesture interaction include: obtaining hand posture and spatial position information using wearable devices, obtaining hand depth information through depth cameras or binocular cameras for gesture interaction, and using gesture recognition models to recognize certain specific gestures.

[0004] However, the above solutions still have some problems. For example, the invention patent application publication "Data Interaction Method, Apparatus, Device and Medium Based on Wearable Devices" (CN 115840506 A), which uses wearable devices to detect and identify terminal devices and palms within a set spatial range, and generates a virtual hand corresponding to the user's palm in the wearable device. Data interaction between the wearable device and the terminal device is carried out when the target gesture of the virtual hand matches the preset interactive gesture. This method requires users to wear specific devices to perform gesture interaction, which greatly limits its use in daily life and application scenarios. The cumbersome wearing process also brings many inconveniences to gesture interaction and does not reflect the convenience of gesture interaction. In addition, wearable devices are expensive and difficult to popularize on a large scale.

[0005] The invention patent application published in CN 115620397 A, which employs a binocular camera solution, entitled "An In-Vehicle Gesture Recognition System Based on Leapmotion Sensor," utilizes binocular cameras to capture user gesture images, generate corresponding 3D gesture models, and acquire gesture data information of the captured object. The extracted gesture information is then matched with a gesture training library to output matched control signals for gesture interaction. This solution uses binocular cameras to acquire 3D gesture information, a process requiring significant computing resources. Furthermore, current mainstream depth cameras have a recognition range of 0.5-5m, resulting in a large blind spot. Additionally, the Leapmotion depth camera is costly and unsuitable for large-scale applications.

[0006] Patent applications using gesture recognition models, such as "A Method and Device for In-Cockpit Gesture Interaction" (CN 115424356 A), input real-time images from inside the cockpit into a gesture recognition model to obtain the first gesture category detection result and the first position category detection result output by the model. The system then controls the equipment inside the cockpit based on the control command corresponding to the first gesture at the first cockpit position. This approach uses a gesture recognition model to obtain gesture categories and their corresponding control commands, simply matching gestures with control commands. However, it does not acquire the three-dimensional spatial information of the gestures. The gesture interaction functionality relies heavily on the database within the gesture recognition model, which significantly limits the scalability and adaptability of gesture interaction, requiring substantial modification costs when adding new gestures.

[0007] With the support of deep learning algorithms, gesture tracking and 3D pose estimation can now be achieved through a single camera, freeing us from the limitations of wearable devices and expensive depth camera equipment. Gesture interaction is no longer limited to a specific scenario or a single function. In-vehicle gesture control, vehicle-machine gesture interaction, sign language recognition, and other functions can all be implemented on a single device. Furthermore, not only in 2D space, but also with the support of gesture pose estimation, 3D AR games and interactions in the car will become possible. Summary of the Invention

[0008] The technical problem to be solved by this invention is to address the limitations of existing gesture interaction application scenarios, insufficient scalability of interaction systems, and high cost of image acquisition equipment in the in-vehicle environment. This invention provides a method and device for gesture estimation and gesture interaction using a common monocular camera. This method achieves more accurate and efficient gesture interaction and reduces the cost of gesture interaction devices. In addition, rendering different gesture images increases the scalability of gesture interaction and reduces the cost of updating gesture interaction functions.

[0009] The objective of this invention is achieved as follows: This invention provides an in-vehicle gesture interaction system based on a monocular camera, comprising a monocular camera module, a data processing module, a control module, and a response module;

[0010] The data processing module includes a gesture segmentation unit, a neural network unit, and an image rendering unit connected in a unidirectional manner. The gesture segmentation unit includes a difference frame extractor, a skin color detector, and a gesture segmentation and background normalization algorithm module, wherein the difference frame extractor and the skin color detector are unidirectionally connected to the gesture segmentation and background normalization algorithm module, respectively. The neural network unit includes a feature extraction network, a 2D detector, and a 3D detector connected in a unidirectional manner, wherein the feature extraction network is unidirectionally connected to the 3D detector. The image rendering unit includes a model fitter and an image renderer connected in a unidirectional manner, wherein the model fitter pre-stores a parameterized hand model, and the 3D detector pre-stores a convolutional neural network.

[0011] The control module includes a trajectory detection unit, a gesture database, and a control unit. The trajectory detection unit and the gesture database are unidirectionally connected to the control unit. The trajectory detection unit pre-stores inter-frame difference method and gesture three-dimensional spatial motion trajectory detection algorithm. The gesture database pre-stores multiple gesture datasets corresponding to multiple application scenarios.

[0012] The monocular camera module is used to acquire images, and the monocular camera module is unidirectionally connected to the difference frame extractor and the skin color detector respectively;

[0013] The image renderer is unidirectionally connected to the trajectory detection unit and the gesture database, and the control unit is unidirectionally connected to the response module.

[0014] Preferably, the monocular camera module is a 1920×1080 resolution wide-angle RGB camera, installed on the vehicle's central control screen near the driver's seat.

[0015] Preferably, the response module includes a multimedia system and a human-computer interaction system.

[0016] Preferably, the data in the gesture database consists of pre-stored gesture images, including the following three types: pre-stored gesture depth maps, pre-stored gesture 2D maps, and pre-stored gesture 3D reconstruction models. The gesture dataset pre-stores any one of the above three types of gesture images or any two or more gesture images according to its application scenario.

[0017] This invention also provides an in-vehicle gesture interaction method based on a monocular camera. This interaction method is applied to gesture interaction and control of vehicles and includes the following steps:

[0018] Step 1: The monocular camera module acquires RGB images and transmits them to the difference frame extractor and skin color detector in the gesture segmentation unit, respectively.

[0019] Step 2: The difference frame extractor extracts the motion regions in the RGB image and transmits them to the gesture segmentation and background normalization algorithm module. The skin color detector simultaneously extracts the skin color regions in the RGB image and transmits them to the gesture segmentation and background normalization algorithm module.

[0020] The gesture segmentation and background normalization algorithm module first combines the motion area and the skin color area into a gesture area according to a predetermined program. Then, it segments the gesture area to obtain a gesture segmentation map. Next, it normalizes the background of the gesture segmentation map to obtain a background-normalized gesture segmentation map. Finally, it transmits the background-normalized gesture segmentation map to the feature extraction network in the neural network unit.

[0021] Step 3: The feature extraction network first extracts hand features from the background-normalized gesture segmentation map and records the extracted results as a feature map; then, the feature map is fed into a 2D detector to obtain a key point heatmap of the hand; then, the key point heatmap and the feature map are combined... Figure 1 The data is fed into a 3D detector, and a 3D key point location map of the hand is regressed through a pre-stored convolutional neural network. The 3D key point location map is then transmitted to the model fitter of the image rendering unit. The 3D key point location map contains the three-dimensional spatial location information of 21 key points of the hand.

[0022] Step 4: The model fitter fits the pre-stored parametric hand model with the 3D key point location map obtained in Step 3 to obtain the gesture 3D reconstruction model. The image renderer renders the gesture depth map and gesture 2D map from the camera perspective based on the gesture 3D reconstruction model. Then, the gesture depth map is transmitted to the trajectory monitoring unit in the control module. The gesture depth map, gesture 2D map, and gesture 3D reconstruction model are transmitted together to the gesture database in the control module.

[0023] Step 5: The trajectory monitoring unit uses the pre-stored inter-frame difference method to obtain the three-dimensional spatial displacement information of the root key points of the hand on the multiple continuous gesture depth maps obtained in Step 4, and obtains the three-dimensional spatial motion trajectory of the gesture; the gesture database matches the gesture depth map, gesture two-dimensional map and gesture three-dimensional reconstruction model obtained in Step 4 with the images in its corresponding gesture dataset to obtain the label information represented by the gesture features.

[0024] The control unit combines the tag information represented by the gesture features and the three-dimensional spatial motion trajectory of the gesture to generate control commands and transmit them to the response module.

[0025] Step 6: The response module responds to the control commands to realize human-machine interaction.

[0026] Compared with the prior art, the beneficial effects of the present invention include:

[0027] 1. A multi-mode gesture segmentation algorithm is used to crop and segment the palm portion of the image and normalize the palm background to avoid the influence of cluttered background on gesture estimation and improve the accuracy of hand key point estimation.

[0028] 2. By using neural networks to accurately obtain the three-dimensional spatial position information of key hand points from a monocular camera, the cost of gesture interaction devices is reduced without relying on sophisticated and expensive depth cameras and wearable devices.

[0029] 3. Reconstructing a 3D gesture model using a parametric hand model can effectively avoid recognition failures caused by hand self-occlusion or missing depth images in previous gesture recognition algorithms. The gesture depth map and gesture 2D map rendered in the 3D gesture reconstruction model have higher quality than traditional methods and avoid the influence of gesture background images, thus improving the accuracy of gesture recognition.

[0030] 4. Compared to traditional gesture recognition and gesture interaction methods where a specific gesture can only represent a specific meaning, the estimated three-dimensional spatial position information of hand key points has a higher dimension of gesture information. Different gestures can also contain more information during gesture interaction, such as the different angles of finger bending and the different distances of palm translation that are difficult to distinguish in gesture recognition.

[0031] 5. By matching three types of gesture image information with corresponding gesture databases, different human-computer interaction functions can be realized, such as sign language translation, AR games, and in-vehicle device control. When adding new gestures, only the corresponding gesture dataset needs to be added, which has a significant upgrade cost advantage. Attached Figure Description

[0032] Figure 1 This is a block diagram of the overall structure of the interactive system of the present invention;

[0033] Figure 2 This is a block diagram of the data processing module structure in an embodiment of the present invention;

[0034] Figure 3 This is a block diagram of the control module structure in an embodiment of the present invention. Detailed Implementation

[0035] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings.

[0036] Figure 1 This is a block diagram of the overall structure of the interactive system of the present invention. Figure 2 This is a block diagram of the data processing module structure in an embodiment of the present invention. Figure 3This is a block diagram of the control module structure in an embodiment of the present invention. Figures 1-3 As can be seen, the present invention provides an in-vehicle gesture interaction system based on a monocular camera, including a monocular camera module 10, a data processing module 20, a control module 30, and a response module 40.

[0037] The data processing module 20 includes a gesture segmentation unit 21, a neural network unit 22, and an image rendering unit 23 connected in a unidirectional manner. The gesture segmentation unit 21 includes a difference frame extractor 211, a skin color detector 212, and a gesture segmentation and background normalization algorithm module 213, wherein the difference frame extractor 211 and the skin color detector 212 are unidirectionally connected to the gesture segmentation and background normalization algorithm module 213, respectively. The neural network unit 22 includes a feature extraction network 221, a 2D detector 222, and a 3D detector 223 connected in a unidirectional manner, wherein the feature extraction network 221 is unidirectionally connected to the 3D detector 223. The image rendering unit 23 includes a model fitter 231 and an image renderer 232 connected in a unidirectional manner, wherein the model fitter 231 pre-stores a parameterized hand model, and the 3D detector 223 pre-stores a convolutional neural network.

[0038] The control module 30 includes a trajectory detection unit 31, a gesture database 32, and a control unit 33. The trajectory detection unit 31 and the gesture database 32 are unidirectionally connected to the control unit 33. The trajectory detection unit 31 is pre-stored with inter-frame difference method and gesture three-dimensional spatial motion trajectory detection algorithm. The gesture database 32 is pre-stored with various gesture datasets corresponding to various application scenarios.

[0039] The monocular camera module 10 is used to acquire images, and the monocular camera module 10 is unidirectionally connected to the difference frame extractor 211 and the skin color detector 212 respectively.

[0040] The image renderer 232 is unidirectionally connected to the trajectory detection unit 31 and the gesture database 32, respectively, and the control unit 33 is unidirectionally connected to the response module 40.

[0041] In this embodiment, the monocular camera module 10 is a 1920×1080 resolution wide-angle RGB camera, which is installed on the vehicle's central control screen near the driver's seat.

[0042] In this embodiment, the response module 40 includes a multimedia system 41 and a human-computer interaction system 42.

[0043] In this embodiment, the data in the gesture database 32 consists of pre-stored gesture images, including the following three types: pre-stored gesture depth maps, pre-stored gesture 2D maps, and pre-stored gesture 3D reconstruction models. The gesture dataset pre-stores any one of the above three types of gesture images or any two or more gesture images according to its application scenario.

[0044] In this embodiment, the application scenarios of the gesture interaction include: sign language translation, in-vehicle multimedia system control, virtual reality interaction, and in-vehicle device control. The in-vehicle device control includes window control, in-vehicle air conditioning control, seat adjustment, etc. For example, the gesture dataset corresponding to the in-vehicle multimedia system control scenario includes pre-stored gesture depth maps and pre-stored gesture 2D maps, while the gesture dataset corresponding to the virtual reality interaction scenario only includes pre-stored gesture 3D reconstruction models.

[0045] This invention also provides an in-vehicle gesture interaction method based on a monocular camera, comprising the following steps:

[0046] Step 1: The monocular camera module 10 acquires RGB images and transmits them to the difference frame extractor 211 and skin color detector 212 in the gesture segmentation unit 21, respectively.

[0047] Step 2: The difference frame extractor 211 extracts the regions of motion in the RGB image and transmits them to the gesture segmentation and background normalization algorithm module 213. The skin color detector 212 simultaneously extracts the skin color regions in the RGB image and transmits them to the gesture segmentation and background normalization algorithm module 213.

[0048] The gesture segmentation and background normalization algorithm module 213 first combines the motion area and the skin color area into a gesture area according to a predetermined program, then segments the gesture area to obtain a gesture segmentation map, and then normalizes the background of the gesture segmentation map to obtain a background-normalized gesture segmentation map, and then transmits the background-normalized gesture segmentation map to the feature extraction network 221 in the neural network unit 22.

[0049] Step 3: The feature extraction network 221 first extracts hand features from the background-normalized gesture segmentation map and records the extraction result as a feature map; then, the feature map is fed into the 2D detector 222 to obtain a key point heatmap of the hand; then, the key point heatmap and the feature map are combined... Figure 1 The data is fed into the 3D detector 223, which regresses the 3D key point location map of the hand through a pre-stored convolutional neural network, and transmits the 3D key point location map to the model fitter 231 of the image rendering unit 23; the 3D key point location map contains the three-dimensional spatial location information of 21 key points of the hand.

[0050] Step 4: The model fitter 231 fits the pre-stored parametric hand model with the 3D key point location map obtained in step 3 to obtain the gesture 3D reconstruction model. The image renderer 232 renders the gesture depth map and gesture 2D map from the camera perspective based on the gesture 3D reconstruction model. Then, the gesture depth map is transmitted to the trajectory monitoring unit 31 in the control module 30, and the gesture depth map, gesture 2D map, and gesture 3D reconstruction model are transmitted together to the gesture database 32 in the control module 30.

[0051] Step 5: The trajectory monitoring unit 31 uses the pre-stored inter-frame difference method to obtain the three-dimensional spatial displacement information of the root key points of the hand on the multiple continuous gesture depth maps obtained in step 4, and obtains the three-dimensional spatial motion trajectory of the gesture; the gesture database 32 matches the gesture depth map, gesture two-dimensional map and gesture three-dimensional reconstruction model obtained in step 4 with the images in its corresponding gesture dataset to obtain the label information represented by the gesture features.

[0052] The control unit 33 combines the tag information represented by the gesture features and the three-dimensional spatial motion trajectory of the gesture to generate control commands and transmit them to the response module 40;

[0053] Step 6: The response module 40 responds to the control commands to realize human-computer interaction.

[0054] In summary, this invention realizes an in-vehicle gesture interaction method and device based on a monocular camera. With the support of neural networks, a regular monocular camera is used to estimate the 3D pose of the hand, reducing the configuration cost of the gesture interaction device. The multi-mode gesture segmentation algorithm reduces external noise interference and improves the robustness of the hand keypoint estimation process. The 3D gesture reconstruction model obtained through gesture estimation can solve the gesture occlusion problem. Simultaneously, the rendered gesture depth map, compared to the depth image directly acquired by a depth camera, reduces the influence of the background, avoiding the impact of depth loss or depth image holes on gesture recognition. By matching three types of gesture image information with different gesture datasets, corresponding human-computer interaction functions are realized, broadening the application scenarios of gesture interaction. Furthermore, when a new gesture is added, only the corresponding gesture dataset needs to be added, resulting in a significant upgrade cost advantage and supporting more diverse gesture interaction functions.

Claims

1. A vehicle-mounted gesture interaction method based on a monocular camera, wherein the interaction method is applied to gesture interaction and control of a vehicle, characterized in that, Includes the following steps: Step 1: The monocular camera module (10) acquires RGB images and transmits them to the difference frame extractor (211) and skin color detector (212) in the gesture segmentation unit (21), respectively. Step 2, the difference frame extractor (211) extracts the region of motion in the RGB image and transmits it to the gesture segmentation and background normalization algorithm module (213). The skin color detector (212) synchronously extracts the skin color region in the RGB image and transmits it to the gesture segmentation and background normalization algorithm module (213). The gesture segmentation and background normalization algorithm module (213) first combines the motion area and the skin color area into a gesture area according to a predetermined program, then segments the gesture area to obtain a gesture segmentation map, and then normalizes the background of the gesture segmentation map to obtain a background normalized gesture segmentation map, and then transmits the background normalized gesture segmentation map to the feature extraction network (221) in the neural network unit (22). Step 3, Feature Extraction Network (221) First, it extracts hand features from the normalized hand segmentation map and records the extraction result as a feature map; then, it sends the feature map to the 2D detector (222) to obtain a key point heatmap of the hand; then, it sends the key point heatmap and the feature map together to the 3D detector (223), and regresses the 3D key point location map of the hand through the pre-stored convolutional neural network, and transmits the 3D key point location map to the model fitter (231) of the image rendering unit (23); the 3D key point location map contains the three-dimensional spatial location information of 21 key points of the hand; Step 4, the model fitter (231) fits the pre-stored parametric hand model with the 3D key point location map obtained in step 3 to obtain the hand gesture 3D reconstruction model. The image renderer renders the hand gesture depth map and hand gesture 2D map from the camera perspective based on the hand gesture 3D reconstruction model. Then, the hand gesture depth map is transmitted to the trajectory monitoring unit (31) in the control module (30). The hand gesture depth map, hand gesture 2D map, and hand gesture 3D reconstruction model are transmitted together to the hand gesture database (32) in the control module (30). Step 5, the trajectory monitoring unit (31) uses the pre-stored inter-frame difference method to obtain the three-dimensional spatial displacement information of the root key point of the hand on the multiple continuous gesture depth maps obtained in step 4, and obtains the three-dimensional spatial motion trajectory of the gesture; the gesture database (32) matches the gesture depth map, gesture two-dimensional map and gesture three-dimensional reconstruction model obtained in step 4 with the images in its corresponding gesture dataset to obtain the label information represented by the gesture features. The control unit (33) combines the tag information represented by the gesture features and the three-dimensional spatial motion trajectory of the gesture to generate control commands and transmit them to the response module (40); Step 6: The response module (40) responds to the control commands to realize human-machine interaction.

2. A vehicle-mounted gesture interaction system based on a monocular camera, used in the vehicle-mounted gesture interaction method based on a monocular camera as described in claim 1, characterized in that, It includes a monocular camera module (10), a data processing module (20), a control module (30), and a response module (40); The data processing module (20) includes a gesture segmentation unit (21), a neural network unit (22), and an image rendering unit (23) connected in a unidirectional manner. The gesture segmentation unit (21) includes a difference frame extractor (211), a skin color detector (212), and a gesture segmentation and background normalization algorithm module (213), wherein the difference frame extractor (211) and the skin color detector (212) are unidirectionally connected to the gesture segmentation and background normalization algorithm module (213). The neural network unit (22) includes a feature extraction network (221), a 2D detector (222), and a 3D detector (223) connected in a unidirectional manner, wherein the feature extraction network (221) is unidirectionally connected to the 3D detector (223). The image rendering unit (23) includes a model fitter (231) and an image renderer (232) connected in a unidirectional manner, wherein the model fitter (231) pre-stores a parameterized hand model, and the 3D detector (223) pre-stores a convolutional neural network. The control module (30) includes a trajectory detection unit (31), a gesture database (32), and a control unit (33). The trajectory detection unit (31) and the gesture database (32) are unidirectionally connected to the control unit (33). The trajectory detection unit (31) is pre-stored with inter-frame difference method and gesture three-dimensional space motion trajectory detection algorithm. The gesture database (32) is pre-stored with various gesture datasets corresponding to various application scenarios. The monocular camera module (10) is used to acquire images, and the monocular camera module (10) is unidirectionally connected to the difference frame extractor (211) and the skin color detector (212); The image renderer (232) is unidirectionally connected to the trajectory detection unit (31) and the gesture database (32), respectively, and the control unit (33) is unidirectionally connected to the response module (40).

3. The in-vehicle gesture interaction system based on a monocular camera according to claim 2, characterized in that, The monocular camera module (10) is a 1920×1080 resolution wide-angle RGB camera, which is installed on the vehicle's central control screen near the driver's seat.

4. The in-vehicle gesture interaction system based on a monocular camera according to claim 2, characterized in that, The response module (40) includes a multimedia system (41) and a human-computer interaction system (42).

5. The in-vehicle gesture interaction system based on a monocular camera according to claim 2, characterized in that, The data in the gesture database (32) consists of pre-stored gesture images, including the following three types: pre-stored gesture depth map, pre-stored gesture 2D map and pre-stored gesture 3D reconstruction model. The gesture dataset pre-stores any one of the above three types of gesture images or any two or more gesture images according to its application scenario.

Citation Information

Patent Citations

  • Gesture interaction method and device in cabin

    CN115424356A

  • Vehicle-mounted gesture recognition system based on Leapmotion sensor

    CN115620397A

  • Data interaction method and device based on wearable device, equipment and medium

    CN115840506A

  • Hand gesture segmentation method based on monocular vision complicated background

    CN104679242A

  • Embedded gesture control method and system based on monocular camera

    CN108629272A