Gesture interaction method and system of vehicle-mounted screen, readable storage medium and computer

By performing frame-by-frame processing and image preprocessing on the front-screen video data on the vehicle, combined with the optimization of the image detection model, the problems of large deviations and high costs of the existing gesture interaction methods are solved, and more accurate and efficient gesture recognition is achieved.

CN120066254APending Publication Date: 2025-05-30CHONGQING LIANGJIANG LIANCHUANG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510081476.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing gesture interaction method based on vehicle screens relies on manual feature extraction, with large deviations in the results and high costs, and gesture background and external interference affect the recognition accuracy.

Method used

By collecting video data on the front screen in real time, processing frame by frame to screen out background, image preprocessing is performed to reduce lighting effects, image detection models are constructed to identify edge information and coordinate outlines of moving targets, and key point features are used to optimize the model for gesture recognition.

Benefits of technology

It improves the accuracy and efficiency of gesture interaction, reduces manual intervention, reduces R&D costs, and enhances the robustness of gesture recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066254A_ABST
    Figure CN120066254A_ABST
Patent Text Reader

Abstract

The invention provides a gesture interaction method and system of a vehicle-mounted screen, a readable storage medium and a computer. The method comprises the following steps: performing frame-by-frame processing on collected video data of a preset area in front of the vehicle-mounted screen to obtain a target image; performing image preprocessing on the target image to obtain preprocessed image data, obtaining the distance between the moving target and the vehicle-mounted screen, constructing a coordinate system by taking the central point of the vehicle-mounted screen as an original point, and determining coordinate information of the moving target based on the coordinate system and the distance; determining a coordinate contour of the moving target according to the coordinate information and the edge information of the moving target; performing image mapping processing on the preprocessed image data to obtain key point features of a moving target in the preprocessed image data; and inputting the key point features and the coordinate contour into an image detection model to enable the image detection model to perform model optimization to obtain a gesture recognition model, and performing gesture recognition on a preset area in front of the vehicle-mounted screen by using the gesture recognition model to realize gesture interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a gesture interaction method, system, readable storage medium and computer for an in-vehicle screen. Background Art

[0002] With the rapid development of technology and the improvement of people's living standards, automobiles have become an indispensable part of people's lives.

[0003] As one of the commonly used components in vehicles, along with the intelligent development of the automotive industry, gesture interaction with the hand as the medium on the in-vehicle screen also has important research significance and value in human-computer interaction.

[0004] Currently, for the gesture interaction method based on the in-vehicle screen, it is usually based on traditional gesture recognition algorithms. After using gesture area segmentation, artificial features are extracted by means of histograms, etc., and a classifier is used to classify the extracted features to obtain gesture information. However, this method requires manual extraction of feature information specifically, which is relatively dependent on the experience of researchers, resulting in large result deviations and high R & D costs. Moreover, factors such as the gesture background and external interference also affect the recognition accuracy. Summary of the Invention

[0005] Based on this, the purpose of the present invention is to provide a gesture interaction method, system, readable storage medium and computer for an in-vehicle screen to at least solve the above-mentioned technical deficiencies.

[0006] The present invention provides a gesture interaction method for an in-vehicle screen, including: Real-time collecting video data of a preset area in front of the in-vehicle screen, and performing frame-by-frame processing on the video data to obtain corresponding target images, wherein at least one moving target is included in the target images; Performing image preprocessing on the target images to obtain preprocessed image data, and obtaining the distance between the moving target and the in-vehicle screen, constructing a coordinate system with the center point of the in-vehicle screen as the origin, and determining the coordinate information of the moving target based on the coordinate system and the distance; Constructing an image detection model, identifying the edge information of the moving target in the preprocessed image data, and determining the coordinate contour of the moving target according to the coordinate information and the edge information; Performing image mapping processing on the preprocessed image data to obtain the key point features of the moving target in the preprocessed image data; Input the key point features and the coordinate contours into the image detection model, so that the image detection model performs model optimization to obtain a gesture recognition model, and use the gesture recognition model to perform gesture recognition on a preset area in front of the vehicle-mounted screen to achieve gesture interaction.

[0007] Further, the steps of collecting video data of a preset area in front of the vehicle-mounted screen in real time and processing the video data frame by frame to obtain corresponding target images include: Decompose the video data frame by frame to obtain a number of frame-by-frame images, and perform a difference operation on each frame-by-frame image and a preset background image function to obtain corresponding first difference images; Perform gray-scale processing on the current frame and its previous frame in each frame-by-frame image, and perform binarization processing on the gray-scale processing result to obtain corresponding second difference images; Fuse the first difference image and the second difference image to obtain a corresponding target image.

[0008] Further, the steps of performing image preprocessing on the target image to obtain preprocessed image data, obtaining the distance between the moving target and the vehicle-mounted screen, constructing a coordinate system with the center point of the vehicle-mounted screen as the origin, and determining the coordinate information of the moving target based on the coordinate system and the distance include: Perform color space conversion on the target image, and perform denoising processing on the target image after color space conversion to obtain a corresponding denoised image; Input the denoised image into a convolutional neural network model for image processing to output a binary image of the denoised image; Obtain the distance between the moving target and the vehicle-mounted screen, construct a coordinate system with the center point of the vehicle-mounted screen as the origin, and determine the coordinate information of the moving target based on the image size of the target image, the coordinate system, and the distance.

[0009] Further, the steps of constructing an image detection model include: Construct an initial detection model, introduce a graph convolution algorithm into the initial detection model, and use the graph convolution algorithm to perform convolution processing on the input features of the initial detection model to obtain graph convolution output features; Fuse the output features of the initial detection model and the graph convolution output features to obtain final output features, and optimize the initial detection model based on the final output features to construct an image detection model.

[0010] Further, the step of identifying the edge information of the moving target in the preprocessed image data and determining the coordinate contour of the moving target according to the coordinate information and the edge information includes: Performing edge extraction on the preprocessed image to identify the edge information of the moving target in the preprocessed image and its corresponding heat map; Determining the coordinates of the maximum activation point and the second maximum activation point in the heat map according to the coordinate information; Predicting the coordinate contour of the moving target through the coordinates of the maximum activation point and the second maximum activation point in the heat map and the edge information of the moving target in the preprocessed image.

[0011] The present invention also proposes a gesture interaction system for an in-vehicle screen, including: A data acquisition module, configured to collect video data of a preset area in front of the in-vehicle screen in real time, and perform frame-by-frame processing on the video data to obtain corresponding target images, where at least one moving target is included in the target images; An information determination module, configured to perform image preprocessing on the target images to obtain preprocessed image data, and acquire the distance between the moving target and the in-vehicle screen, construct a coordinate system with the center point of the in-vehicle screen as the origin, and determine the coordinate information of the moving target based on the coordinate system and the distance; A data processing module, configured to construct an image detection model, identify the edge information of the moving target in the preprocessed image data, and determine the coordinate contour of the moving target according to the coordinate information and the edge information; A mapping processing module, configured to perform image mapping processing on the preprocessed image data to obtain key point features of the moving target in the preprocessed image data; A gesture interaction module, configured to input the key point features and the coordinate contour into the image detection model, so that the image detection model is optimized to obtain a gesture recognition model, and use the gesture recognition model to perform gesture recognition on a preset area in front of the in-vehicle screen to achieve gesture interaction.

[0012] Further, the data acquisition module includes: A frame-by-frame decomposition unit, configured to decompose the video data frame by frame to obtain a plurality of frame-by-frame images, and perform a difference operation on each frame-by-frame image and a preset background image function to obtain corresponding first difference images; A grayscale processing unit, configured to perform grayscale processing on the current frame and the previous frame of each frame-by-frame image, and perform binary processing on the grayscale processing result to obtain corresponding second difference images; An image fusion unit for fusing the first difference image and the second difference image to obtain a corresponding target image.

[0013] Further, the information determination module includes: A space conversion unit for performing color space conversion on the target image and denoising the target image after color space conversion to obtain a corresponding denoised image; An image processing unit for inputting the denoised image into a convolutional neural network model for image processing to output a binary image of the denoised image; An information determination unit for obtaining the distance between the moving target and the vehicle-mounted screen, constructing a coordinate system with the center point of the vehicle-mounted screen as the origin, and determining the coordinate information of the moving target based on the image size of the target image, the coordinate system, and the distance.

[0014] Further, the data processing module includes: A model construction unit for constructing an initial detection model, introducing a graph convolution algorithm into the initial detection model, and performing convolution processing on the input features of the initial detection model using the graph convolution algorithm to obtain graph convolution output features; A feature fusion unit for fusing the output features of the initial detection model and the graph convolution output features to obtain final output features, and optimizing the initial detection model based on the final output features to construct an image detection model.

[0015] Further, the data processing module further includes: An edge extraction unit for extracting the edges of the preprocessed image to identify the edge information of the moving target in the preprocessed image and its corresponding heat map; A coordinate determination unit for determining the coordinates of the maximum activation point and the second maximum activation point in the heat map according to the coordinate information; A contour prediction unit for predicting the coordinate contour of the moving target through the coordinates of the maximum activation point and the second maximum activation point in the heat map and the edge information of the moving target in the preprocessed image.

[0016] The present invention also proposes a readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the gesture interaction method of the vehicle-mounted screen described above is implemented.

[0017] The present invention also proposes a computer, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the gesture interaction method of the vehicle-mounted screen described above is implemented.

[0018] The gesture interaction method, system, readable storage medium and computer of the in-vehicle screen in the present invention collect video data in a preset area in front of the in-vehicle screen, and perform frame-by-frame processing on the video data. Through frame-by-frame processing, the background image is effectively screened out, thereby improving the accuracy of image processing. Image preprocessing is performed on the target image to reduce the influence of the illumination intensity on the target image. Moreover, by calculating the coordinate information of the moving target and constructing an image detection model, the coordinate contour of the moving target is determined by using the edge information and coordinate information of the moving target in the preprocessed image data. Further, local enhancement is performed on the hand features, and the obtained key point features and coordinate contour are used to optimize the image detection model, thereby providing a more effective and accurate interaction method for gesture recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a flowchart of the gesture interaction method of the in-vehicle screen in the first embodiment of the present invention; Figure 2 is Figure 1 a detailed flowchart of step S101 in Figure 3 is Figure 1 a detailed flowchart of step S102 in Figure 4 is Figure 1 a detailed flowchart of step S103 in Figure 5 It is a structural block diagram of the gesture interaction system of the in-vehicle screen in the second embodiment of the present invention; Figure 6 It is a structural block diagram of the computer in the third embodiment of the present invention.

[0020] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0023] Embodiment 1 Please refer to Figure 1 , which shows the gesture interaction method of the in-vehicle screen in the first embodiment of the present invention. The method specifically includes steps S101 to S105: S101, Real-time collect the video data of a preset area in front of the in-vehicle screen, and perform frame-by-frame processing on the video data to obtain corresponding target images, where at least one moving target is included in the target images; Further, please refer to Figure 2 , The step S101 specifically includes steps S1011~S1013: S1011, Decompose the video data frame by frame to obtain a number of frame-by-frame images, and perform a difference operation on each frame-by-frame image and a preset background image function to obtain corresponding first difference images; S1012, Perform gray-scale processing on the current frame and the previous frame of each frame-by-frame image, and perform binarization processing on the gray-scale processing result to obtain corresponding second difference images; S1013, Perform image fusion on the first difference image and the second difference image to obtain corresponding target images.

[0024] In specific implementation, real-time collect the video data of a preset area in front of the in-vehicle screen. Among them, the preset area is the collection range of a video collection device installed in the in-vehicle screen. The video collection device includes but is not limited to devices with collection functions such as cameras. The video data collected by the video collection device usually includes relevant action information such as hand movements of a person. Decompose the obtained video data frame by frame according to the frame rate to obtain a number of frame-by-frame images. Perform a difference operation on the obtained frame-by-frame images and the background image function in a preset background image library to obtain the first difference images corresponding to the frame-by-frame images. By calculating the gray-scale difference between two images, effectively filter out the background images from the frame-by-frame images, and perform smoothing filtering on the obtained first difference images to improve the accuracy of image processing.

[0025] Specifically, extract the current frame image and the previous frame of the current frame from the obtained frame-by-frame images, perform gray-scale processing on the two images, and perform binarization operation on the gray-scale processing result. Set the binarization threshold, and compare the data after the binarization operation with the binarization threshold to obtain a second difference image. Compare the difference in pixel gray-scale values after the binarization operation through the binarization threshold, and mark the pixel points with the absolute value of the difference greater than the threshold as target pixel points, otherwise mark the pixel points as background pixel points.

[0026] Further, the first difference image and the second difference image obtained above are subjected to image fusion to obtain a target image containing a moving target. It can be understood that both the first difference image and the second difference image contain target pixel points. By fusing the two images, the accuracy of image prediction is improved, and thus the gesture can be effectively segmented from the image background.

[0027] S102. Perform image preprocessing on the target image to obtain preprocessed image data, and obtain the distance between the moving target and the vehicle-mounted screen. Construct a coordinate system with the center point of the vehicle-mounted screen as the origin, and determine the coordinate information of the moving target based on the coordinate system and the distance. Further, please refer to Figure 3 , and the step S102 specifically includes steps S1021 to S1023: S1021. Perform color space conversion on the target image, and perform denoising processing on the target image after color space conversion to obtain a corresponding denoised image. S1022. Input the denoised image into a convolutional neural network model for image processing to output a binary image of the denoised image. S1023. Obtain the distance between the moving target and the vehicle-mounted screen, construct a coordinate system with the center point of the vehicle-mounted screen as the origin, and determine the coordinate information of the moving target based on the image size of the target image, the coordinate system, and the distance.

[0028] In specific implementation, color space conversion is performed on the obtained target image. Among them, since the images collected by video acquisition devices are usually in RGB encoding format, this format is easily affected by the illumination intensity. Therefore, the above target image is converted into the corresponding YCbCr color space. Through the space conversion, the color space of the image becomes linear, thereby reducing the calculation amount and improving the segmentation effect. Among them, Y represents the luminance component, and Cb and Cr respectively represent the chrominance offsets of blue and red. The target image after the above color space conversion is subjected to denoising processing to eliminate the noise in the target image.

[0029] Specifically, the OTSU algorithm in the convolutional neural network model is used to perform image processing on the denoised image to output a binary image of the denoised image, which resists the interference of brightness and contrast, thereby improving the calculation speed, getting rid of the trouble of manually setting thresholds, and reducing the possible error.

[0030] Further, a distance sensor is used to obtain the distance between the moving target and the vehicle-mounted screen in real time. A coordinate system is constructed with the center point of the vehicle-mounted screen as the origin. Coordinate transformation is performed based on the image size of the target image, this coordinate system, and the distance between the moving target and the vehicle-mounted screen, so as to determine the coordinate information of the moving target. Specifically, equations are constructed using the edge point coordinates of the moving target captured by the video acquisition device and the coordinates where the edge point coordinates are projected onto the vehicle-mounted screen in the vertical direction. The equations are solved, and the coordinate information of all contour points of the moving target is obtained by combining the corresponding compensation coefficients, and then the entire coordinate information of the moving target is determined.

[0031] S103. Construct an image detection model, identify the edge information of the moving target in the preprocessed image data, and determine the coordinate contour of the moving target according to the coordinate information and the edge information; Further, please refer to Figure 4 , the step S103 specifically includes steps S1031 to S1032: S1031. Construct an initial detection model, introduce a graph convolution algorithm into the initial detection model, and use the graph convolution algorithm to perform convolution processing on the input features of the initial detection model to obtain graph convolution output features; S1032. Perform feature fusion on the output features of the initial detection model and the graph convolution output features to obtain final output features, and optimize the initial detection model based on the final output features to construct an image detection model.

[0032] In specific implementation, an initial detection model is constructed. The initial detection model consists of two stacked fourth-order hourglasses, a joint graph inference module, and a multi-layer perception module. Among them, the hourglass model fuses features at different levels through an upsampling module and a downsampling module to obtain local information and global information of the hand, which helps to identify different joint points in the gesture. A graph convolution algorithm is introduced into the above initial detection model to enhance the features of the hand joint points in the gesture, and the graph convolution algorithm is used to perform convolution processing on the input features of the initial detection model; During the convolution processing, each node will interact with its neighbor nodes and calculate a new feature vector to represent the performance of the node in each feature dimension. Specifically, an undirected graph between joint points is defined, the features of each joint are input, the GCN algorithm is used to model the dependency relationship between joints, and matrix multiplication is used to perform operations on the features of all joint points to obtain graph convolution output features; Specifically, the output features of the initial detection model and the graph convolution output features obtained above are fused to obtain the final output features, which are the local enhanced features of the hand in the gesture. The local enhanced features are input into the initial detection model for model optimization to obtain the image detection model.

[0033] Further, step S103 further includes steps S1033 to S1035: S1033, perform edge extraction on the preprocessed image to identify the edge information of the moving target in the preprocessed image and its corresponding heat map; S1034, determine the coordinates of the maximum activation point and the second maximum activation point in the heat map according to the coordinate information; S1035, predict the coordinate contour of the moving target through the coordinates of the maximum activation point and the second maximum activation point in the heat map and the edge information of the moving target in the preprocessed image.

[0034] In specific implementation, perform edge extraction on the obtained preprocessed image to identify the edge information of the moving target in the preprocessed image and its corresponding heat map, determine the coordinates of the maximum activation point and the second maximum activation point in the heat map according to the above coordinate information, and calculate and predict the coordinate contour of the moving target through the following formula: ; In the formula, represents the position coordinates of the maximum activation point in the heat map, represents the position coordinates of the second maximum activation point in the heat map, represents the coordinate contour of the moving target, where the unit is pixel.

[0035] S104, perform image mapping processing on the preprocessed image data to obtain the key point features of the moving target in the preprocessed image data; S105, input the key point features and the coordinate contour into the image detection model, so that the image detection model performs model optimization to obtain a gesture recognition model, and use the gesture recognition model to perform gesture recognition on a preset area in front of the vehicle-mounted screen to achieve gesture interaction.

[0036] In specific implementation, input the key point features and the obtained coordinate contour into the above image detection model, so that the image detection model performs model optimization, use the distance between each joint in the gesture coordinates and its adjacent finger joint as the length standard for coordinate representation, thereby obtaining a gesture recognition model, and use the gesture recognition model to perform gesture recognition on a preset area in front of the vehicle-mounted screen, and realize gesture interaction according to the predefined meaning expression of the gesture dynamics, the gesture recognition result and the corresponding meaning expression.

[0037] In summary, for the gesture interaction method of the in-vehicle screen in the above embodiments of the present invention, video data is collected from a preset area in front of the in-vehicle screen, and the video data is processed frame by frame. By processing frame by frame, the background image is effectively screened out, thereby improving the accuracy of image processing. Image preprocessing is performed on the target image to reduce the influence of the light intensity on the target image. Moreover, by calculating the coordinate information of the moving target and constructing an image detection model, the coordinate contour of the moving target is determined using the edge information and coordinate information of the moving target in the preprocessed image data. Further, local enhancement is performed on the hand features, and the obtained key point features and coordinate contour are used to optimize the image detection model, thereby providing a more effective and accurate interaction method for gesture recognition.

[0038] Embodiment 2 On the other hand, the present invention also proposes a gesture interaction system for an in-vehicle screen. Please refer to Figure 5 , which shows the gesture interaction system for the in-vehicle screen in the second embodiment of the present invention. The system includes: A data acquisition module 11, configured to collect video data of a preset area in front of the in-vehicle screen in real time, and perform frame-by-frame processing on the video data to obtain corresponding target images, where at least one moving target is included in the target images; Furthermore, the data acquisition module 11 includes: A frame-by-frame decomposition unit, configured to decompose the video data frame by frame to obtain a plurality of frame-by-frame images, and perform a difference operation on each frame-by-frame image and a preset background image function to obtain corresponding first difference images; A grayscale processing unit, configured to perform grayscale processing on the current frame and the previous frame of each frame-by-frame image, and perform binarization processing on the grayscale processing result to obtain corresponding second difference images; An image fusion unit, configured to perform image fusion on the first difference images and the second difference images to obtain corresponding target images.

[0039] An information determination module 12, configured to perform image preprocessing on the target images to obtain preprocessed image data, and obtain the distance between the moving target and the in-vehicle screen, construct a coordinate system with the center point of the in-vehicle screen as the origin, and determine the coordinate information of the moving target based on the coordinate system and the distance; Furthermore, the information determination module 12 includes: A space conversion unit, configured to perform color space conversion on the target images, and perform denoising processing on the target images after color space conversion to obtain corresponding denoised images; An image processing unit for inputting the denoised image into a convolutional neural network model for image processing to output a binary image of the denoised image; An information determination unit for obtaining the distance between the moving target and the vehicle-mounted screen, constructing a coordinate system with the center point of the vehicle-mounted screen as the origin, and determining the coordinate information of the moving target based on the image size of the target image, the coordinate system, and the distance.

[0040] A data processing module 13 for constructing an image detection model, identifying the edge information of the moving target in the preprocessed image data, and determining the coordinate contour of the moving target according to the coordinate information and the edge information; Further, the data processing module 13 includes: A model construction unit for constructing an initial detection model, introducing a graph convolution algorithm into the initial detection model, and performing convolution processing on the input features of the initial detection model using the graph convolution algorithm to obtain graph convolution output features; A feature fusion unit for fusing the output features of the initial detection model and the graph convolution output features to obtain final output features, and optimizing the initial detection model based on the final output features to construct an image detection model.

[0041] Further, the data processing module 13 further includes: An edge extraction unit for extracting edges from the preprocessed image to identify the edge information of the moving target in the preprocessed image and its corresponding heat map; A coordinate determination unit for determining the coordinates of the maximum activation point and the second maximum activation point in the heat map according to the coordinate information; A contour prediction unit for predicting the coordinate contour of the moving target through the coordinates of the maximum activation point and the second maximum activation point in the heat map and the edge information of the moving target in the preprocessed image.

[0042] A mapping processing module 14 for performing image mapping processing on the preprocessed image data to obtain key point features of the moving target in the preprocessed image data; A gesture interaction module 15 for inputting the key point features and the coordinate contour into the image detection model, enabling the image detection model to optimize the model to obtain a gesture recognition model, and using the gesture recognition model to perform gesture recognition on a preset area in front of the vehicle-mounted screen to achieve gesture interaction.

[0043] The functions or operation steps implemented when the above modules and units are executed are substantially the same as those in the above method embodiments, and will not be elaborated here.

[0044] The gesture interaction system of the in-vehicle screen provided by the embodiments of the present invention has the same implementation principle and technical effects as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the system embodiments, reference may be made to the corresponding contents in the foregoing method embodiments.

[0045] Embodiment III The present invention also provides a computer. Please refer to Figure 6 , which shows the computer in the third embodiment of the present invention, including a memory 10, a processor 20, and a computer program 30 stored on the memory 10 and executable on the processor 20. When the processor 20 executes the computer program 30, the above-mentioned gesture interaction method for the in-vehicle screen is implemented.

[0046] Among them, the memory 10 includes at least one type of readable storage medium. The readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 10 can be an internal storage unit of the computer in some embodiments, such as the hard disk of the computer. The memory 10 can also be an external storage device in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 10 can also include both an internal storage unit and an external storage device of the computer. The memory 10 can be used not only to store application software and various types of data installed on the computer, but also to temporarily store data that has been output or will be output.

[0047] Among them, the processor 20 can be an Electronic Control Unit (ECU, also known as a vehicle computer), a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments, and is used to run the program code stored in the memory 10 or process data, such as executing an access restriction program, etc.

[0048] It should be noted that Figure 6 the structure shown does not constitute a limitation on the computer. In other embodiments, the computer may include fewer or more components than shown in the figure, or combine certain components, or have different component arrangements.

[0049] The embodiments of the present invention also provide a readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned gesture interaction method for the in-vehicle screen is implemented.

[0050] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.

[0051] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, a computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0052] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0053] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0054] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A gesture interaction method for a vehicle-mounted screen, characterized in that: include: Collecting video data of a preset area in front of the vehicle-mounted screen in real time, and processing the video data frame by frame to obtain a corresponding target image, wherein the target image contains at least one moving target; Performing image preprocessing on the target image to obtain preprocessed image data, and acquiring the distance between the moving target and the vehicle-mounted screen, constructing a coordinate system with the center point of the vehicle-mounted screen as the origin, and determining the coordinate information of the moving target based on the coordinate system and the distance; Constructing an image detection model, identifying edge information of the moving target in the preprocessed image data, and determining a coordinate contour of the moving target based on the coordinate information and the edge information; Performing image mapping processing on the preprocessed image data to obtain key point features of the moving target in the preprocessed image data; The key point features and the coordinate profile are input into the image detection model so that the image detection model is optimized to obtain a gesture recognition model, and the gesture recognition model is used to perform gesture recognition on a preset area in front of the vehicle-mounted screen to achieve gesture interaction.

2. The gesture interaction method of the vehicle-mounted screen according to claim 1, characterized in that: The steps of collecting video data of a preset area in front of the vehicle-mounted screen in real time and processing the video data frame by frame to obtain a corresponding target image include: Decomposing the video data frame by frame to obtain a plurality of frame-by-frame images, performing a difference process between each of the frame-by-frame images and a preset background image function to obtain a corresponding first difference image; Performing grayscale processing on the current frame and the previous frame of each of the frame-by-frame images, and binarizing the grayscale processing results to obtain a corresponding second differential image; The first differential image is fused with the second differential image to obtain a corresponding target image.

3. The gesture interaction method for a vehicle-mounted screen according to claim 1, characterized in that: The steps of performing image preprocessing on the target image to obtain preprocessed image data, acquiring the distance between the moving target and the vehicle-mounted screen, constructing a coordinate system with the center point of the vehicle-mounted screen as the origin, and determining the coordinate information of the moving target based on the coordinate system and the distance include: Performing color space conversion on the target image, and performing denoising on the target image after color space conversion to obtain a corresponding denoised image; Inputting the denoised image into a convolutional neural network model for image processing to output a binary image of the denoised image; The distance between the moving target and the vehicle-mounted screen is obtained, and a coordinate system is constructed with the center point of the vehicle-mounted screen as the origin, and the coordinate information of the moving target is determined based on the image size of the target image, the coordinate system and the distance.

4. The gesture interaction method for a vehicle-mounted screen according to claim 1, characterized in that: The steps to build an image detection model include: Constructing an initial detection model, and introducing a graph convolution algorithm into the initial detection model, and using the graph convolution algorithm to perform convolution processing on input features of the initial detection model to obtain graph convolution output features; The output features of the initial detection model and the graph convolution output features are feature fused to obtain final output features, and the initial detection model is optimized based on the final output features to construct an image detection model.

5. The gesture interaction method for a vehicle-mounted screen according to claim 1, characterized in that: The step of identifying edge information of the moving target in the preprocessed image data and determining the coordinate contour of the moving target according to the coordinate information and the edge information comprises: Performing edge extraction on the preprocessed image to identify edge information of a moving target in the preprocessed image and its corresponding heat map; Determine the coordinates of the maximum activation point and the second maximum activation point in the heat map according to the coordinate information; The coordinate contour of the moving object is predicted by the coordinates of the maximum activation point and the second maximum activation point in the heat map and the edge information of the moving object in the preprocessed image.

6. A gesture interaction system for a vehicle-mounted screen, characterized in that: include: A data acquisition module, used for real-time acquisition of video data of a preset area in front of the vehicle-mounted screen, and processing the video data frame by frame to obtain a corresponding target image, wherein the target image contains at least one moving target; An information determination module, configured to perform image preprocessing on the target image to obtain preprocessed image data, and obtain a distance between the moving target and the vehicle-mounted screen, construct a coordinate system with a center point of the vehicle-mounted screen as an origin, and determine coordinate information of the moving target based on the coordinate system and the distance; A data processing module, used to construct an image detection model, identify edge information of the moving target in the pre-processed image data, and determine the coordinate contour of the moving target according to the coordinate information and the edge information; A mapping processing module, used for performing image mapping processing on the pre-processed image data to obtain key point features of the moving target in the pre-processed image data; A gesture interaction module is used to input the key point features and the coordinate profile into the image detection model so that the image detection model is optimized to obtain a gesture recognition model, and the gesture recognition model is used to perform gesture recognition on a preset area in front of the vehicle-mounted screen to achieve gesture interaction.

7. The gesture interaction system for a vehicle-mounted screen according to claim 6, characterized in that: The data acquisition module comprises: A frame-by-frame decomposition unit, used for decomposing the video data frame by frame to obtain a plurality of frame-by-frame images, and performing a difference process between each of the frame-by-frame images and a preset background image function to obtain a corresponding first difference image; A grayscale processing unit, used for performing grayscale processing on the current frame and the previous frame of each of the frame-by-frame images, and performing binarization processing on the grayscale processing result to obtain a corresponding second differential image; An image fusion unit is used to fuse the first differential image with the second differential image to obtain a corresponding target image.

8. The gesture interaction system for a vehicle-mounted screen according to claim 6, characterized in that: The information determination module comprises: A space conversion unit, used to perform color space conversion on the target image, and perform denoising on the target image after color space conversion to obtain a corresponding denoised image; An image processing unit, used for inputting the denoised image into a convolutional neural network model for image processing, so as to output a binary image of the denoised image; An information determination unit is used to obtain the distance between the moving target and the vehicle-mounted screen, and to construct a coordinate system with the center point of the vehicle-mounted screen as the origin, and to determine the coordinate information of the moving target based on the image size of the target image, the coordinate system and the distance.

9. A readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the gesture interaction method for the vehicle-mounted screen as described in any one of claims 1 to 5 is implemented.

10. A computer comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the gesture interaction method for the vehicle-mounted screen as described in any one of claims 1 to 5 is implemented.