Anti-interference optimized gesture recognition method

By constructing a dynamic background model, optimizing image quality, real-time positioning of gesture areas and dynamic update of gesture trajectory, the problem of traditional gesture recognition methods degradation in complex dynamic backgrounds is solved, and high-precision gesture recognition and classification effects are achieved.

CN120183030APending Publication Date: 2025-06-20GUANGZHOU LANGO ELECTRONICS TECH CO LTD
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510105365.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In scenarios where complex dynamic backgrounds or multiple interference factors exist, traditional gesture recognition methods are difficult to effectively distinguish target areas from interference factors, resulting in a decrease in recognition accuracy and stability.

Method used

Dynamic background model is constructed through frame difference method and Gaussian mixed model, and static background and non-target action interference are eliminated; local brightness histogram analysis and area adaptive compensation algorithm are used, combined with multi-scale edge enhancement technology to optimize image quality; color histogram and feature matching algorithm are used to locate user gesture areas in real time, and dynamically update gesture trajectory through Kalman filtering; key features of gestures are extracted through deep learning and accurately compared with the standard model library.

Benefits of technology

Effectively eliminate non-target action interference, improve the accuracy and stability of gesture recognition; optimize image quality, improve feature extraction and recognition effects; realize high-precision classification of gesture categories, and meet the needs of diverse application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183030A_ABST
    Figure CN120183030A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of gesture recognition, in particular to an anti-interference optimized gesture recognition method. Comprising the following steps: acquiring a gesture video stream through a camera, constructing a dynamic background model by using a frame difference method and a Gaussian mixture model, eliminating a static background and interference, and extracting a target area image; performing local brightness histogram analysis on the target region image, and optimizing the image quality by adopting a region adaptive compensation algorithm and a multi-scale edge enhancement technology; positioning a gesture area in real time by using a color histogram and a feature matching algorithm, and dynamically updating a gesture track in combination with Kalman filtering; and extracting gesture shapes, tracks and dynamic mode features through deep learning, comparing the features with a standard model library, and outputting gesture categories and corresponding function instructions. According to the method, a multi-level optimization strategy is adopted for a complex background, a dynamic target and a changeable illumination environment, so that the anti-interference capability and the recognition precision of gesture recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gesture recognition, and particularly to a gesture recognition method with anti-interference optimization. Background Art

[0002] In the field of gesture recognition, complex dynamic environments often introduce a large number of interference factors, such as noise from non-target actions in dynamic backgrounds, confusion of gestures in multi-user scenarios, and image quality problems caused by changes in lighting conditions. These factors significantly affect the recognition performance of the system and limit the universality of gesture recognition technology. There are also the following problems currently: in scenarios with complex dynamic backgrounds or multiple interference factors (such as non-target action interference), traditional gesture recognition methods are difficult to effectively distinguish the target area from interference factors, resulting in a decrease in recognition accuracy and stability; limited by existing technologies, gesture images are prone to problems such as uneven brightness, insufficient contrast, or blurred edges under complex lighting conditions, affecting the subsequent feature extraction and recognition effects; traditional gesture recognition methods have limited ability to extract multi-dimensional gesture features (such as shape, trajectory, and dynamic pattern), and at the same time lack a high-precision comparison mechanism with standard models, resulting in the recognition accuracy of gesture categories being difficult to meet the actual application requirements. Summary of the Invention

[0003] To solve the above problems, the present invention provides a gesture recognition method with anti-interference optimization, which solves the problems of how to effectively eliminate non-target action interference in complex dynamic environments, cope with the impact of lighting condition changes on image quality, and improve the accuracy of gesture feature extraction and classification, thereby enhancing the anti-interference ability and recognition accuracy of gesture recognition.

[0004] To achieve the above object, the technical solution adopted by the present invention is:

[0005] A gesture recognition method with anti-interference optimization, comprising the following steps:

[0006] S1: Obtain a gesture video stream through a camera, and use the frame difference method and Gaussian mixture model to construct a dynamic background model to exclude static backgrounds and non-target action interference, and generate a target area image;

[0007] S2: Based on the target area image, perform local brightness histogram analysis on the dynamic target area, adopt a regional adaptive compensation algorithm, and at the same time combine multi-scale edge enhancement technology to optimize the image quality;

[0008] S3: Based on the optimized image, use color histogram and feature matching algorithms to real-time locate the user's gesture area, and adopt Kalman filtering to dynamically update the gesture trajectory;

[0009] S4: Extract the key features of the gesture, including shape, trajectory, and dynamic pattern, precisely compare them with the standard model library, and output the gesture category and the corresponding function instruction.

[0010] Further, the step S1 includes the following steps:

[0011] Collect the dynamic video stream containing the user's gesture in real time through the camera, perform frame segmentation on the video stream, and extract consecutive frame images;

[0012] Calculate the difference image between adjacent frames using the frame difference method to initially remove the static background area;

[0013] Construct a dynamic background model for the video stream based on the Gaussian mixture model, and perform background fitting and separation on the non-target area;

[0014] Perform post-processing on the difference image in combination with morphological processing, including dilation, erosion, and noise removal operations, to generate a clear target area image.

[0015] Even further, the formula of the frame difference method is as follows:

[0016] D(x, y) = |I t (x, y) - I t-1 (x, y)|

[0017] where D(x, y) represents the frame difference value of the current pixel point (x, y); I t (x, y) represents the gray value of the corresponding pixel point (x, y) at time t; I t-1 (x, y)) represents the gray value of the corresponding pixel point (x, y) at time t - 1.

[0018] Even further, the formula of the dynamic background model is as follows:

[0019]

[0020] where P(I t (x, y)) represents the probability that the pixel point (x, y) belongs to the background at time t; K represents the number of Gaussian components; ω k,t (x, y) represents the weight of the k-th Gaussian component; μ k,t (x, y) represents the mean of the k-th Gaussian component; represents the variance of the k-th Gaussian component; represents the Gaussian distribution function; I t (x, y) represents the gray value of the corresponding pixel point (x, y) at time t.

[0021] Further, the step S2 includes the following steps:

[0022] Based on the target region image, the local region sliding window method is used to calculate the partition of the luminance histogram, and the luminance distribution characteristics of each window region are extracted;

[0023] According to the luminance distribution characteristics, an adaptive compensation algorithm based on pixel weights is used to generate different compensation weights for low-luminance and high-luminance regions, and dynamically adjust the luminance curve of the region;

[0024] Using multi-scale edge enhancement technology, by detecting and enhancing the edge features in the image, the boundary clarity and contrast of the dynamic target region are optimized;

[0025] Combined with morphological processing technology, fine-grained operations are performed on the optimized image, including smoothing, denoising, and connected region segmentation, to generate the final optimized image.

[0026] Further, the step S3 includes the following steps:

[0027] Based on the optimized target image, calculate the color histogram, extract the color features of the gesture region, and lock the preliminary position of the gesture region through dynamic matching;

[0028] Within the preliminarily located gesture region, use the feature extraction algorithm to extract the local feature points of the gesture, and combine the matching algorithm to screen the candidate regions;

[0029] Adopt the Kalman filter algorithm, predict the region position of the next frame according to the position and motion state of the gesture region in the current frame, and dynamically adjust the Kalman filter parameters in combination with the actual observation values;

[0030] Perform time series analysis on the tracking trajectory of the gesture, and use the trajectory smoothing algorithm to optimize the gesture trajectory.

[0031] Further, the step S4 includes the following steps:

[0032] Based on the convolutional neural network algorithm, extract spatial features, trajectory features, and dynamic pattern features from the optimized gesture image to form a feature vector;

[0033] Use a large amount of standard gesture data to train the convolutional neural network algorithm to build a standard model library containing multiple gesture categories;

[0034] Match the extracted feature vector with the standard model library, calculate the similarity using cosine similarity, and output the category closest to the current gesture feature;

[0035] Based on the classification result, generate the corresponding function instruction through the preset mapping rule and output it.

[0036] Even further, the formula for the cosine similarity is as follows:

[0037]

[0038] Among them, Similarity(F, S) represents the similarity between the current gesture feature F and the standard gesture model S; F i represents the key feature vector of the current gesture; W i represents the importance weight of different feature modalities; S i represents the feature vector of a certain category of gesture in the standard gesture model library; ΔT represents the time difference between the current gesture trajectory and the standard gesture model trajectory; K represents the adjustment factor of the time characteristic; C represents the time decay coefficient; e -C·ΔT represents the time decay function; ∈ represents a small positive number; n represents the number of dimensions of the feature vector.

[0039] Furthermore, the adaptive compensation algorithm constructs a local illumination model for the dynamic target area and dynamically adjusts the image brightness by using the gradient direction weighting strategy.

[0040] Further, the gesture area localization based on the color histogram in the S3 step specifically performs hierarchical localization of the gesture area by fusing the global histogram and the local area histogram.

[0041] The beneficial effects of the present invention are as follows:

[0042] The present invention constructs a dynamic background model through the frame difference method and the Gaussian mixture model, which can effectively exclude the interference of the static background and non-target actions, improve the accuracy and stability of gesture recognition, and is particularly suitable for complex dynamic scenes. Through local brightness histogram analysis and regional adaptive compensation algorithm, combined with multi-scale edge enhancement technology, the quality of the target area image is optimized to ensure the clarity of the image under different illumination and background conditions, and further improve the subsequent feature extraction and recognition effects. The real-time localization of the user's gesture area is realized based on the color histogram and the feature matching algorithm, and at the same time, the gesture trajectory is dynamically updated through Kalman filtering, which can better adapt to the dynamic changes of the gesture and ensure the accuracy and continuity of the localization result. By extracting multi-dimensional key features (shape, trajectory, and dynamic pattern) of the gesture through deep learning and accurately comparing them with the standard model library, high-precision classification of gesture categories can be achieved to meet the needs of diverse application scenarios. Outputting the gesture category and the corresponding function instruction provides strong support for the intelligent interaction of the system, and at the same time has good scalability, and new gesture categories and functions can be added according to requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a schematic flowchart of a gesture recognition method with anti-interference optimization of the present invention.

[0044] Figure 2It is a schematic flowchart of step S1 provided by an embodiment of the present invention.

[0045] Figure 3 It is a schematic flowchart of step S4 provided by an embodiment of the present invention. Detailed implementation manners

[0046] Please refer to Figures 1-3 As shown, the present invention relates to a gesture recognition method with anti-interference optimization.

[0047] Embodiment

[0048] A gesture recognition method with anti-interference optimization includes the following steps:

[0049] S1: Obtain a gesture video stream through a camera, use the frame difference method and the Gaussian mixture model to construct a dynamic background model, exclude static background and non-target action interference, and generate a target region image;

[0050] Among them, the step S1 includes the following steps:

[0051] Real-time collect a dynamic video stream containing user gestures through a camera, perform frame segmentation on the video stream, and extract consecutive frame images;

[0052] Use the frame difference method to calculate the difference image between adjacent frames and preliminarily remove the static background area;

[0053] Specifically, select two consecutive frame images, compare their pixel changes, and extract the dynamically changing area. Filter the areas with small changes (such as noise or light fluctuations), and only retain the parts with significant actions. Use the binarization method to mark the significantly changed areas in the difference image as possible gesture areas, while the areas with small changes are retained as the background. Directly exclude the static background through the difference image, such as parts that do not participate in the action like the desktop and walls, providing a streamlined input for subsequent processing.

[0054] Based on the Gaussian mixture model, construct a dynamic background model for the video stream, and perform background fitting and separation on the non-target areas;

[0055] Specifically, initialize the first few seconds of frame data collected by the camera for constructing the background model. Take the change pattern of each pixel point as the background characteristic and record its historical distribution. As time goes by, update the background model to adapt to environmental changes (such as light fluctuations or other dynamic interfering objects in the background). By separating the foreground and the background, replace the background area with a fixed value to ensure that the dynamically extracted area only contains the gesture target. In each frame, extract the dynamic foreground area, that is, the area where the user gesture is located, by comparing the current frame pixels with the background model.

[0056] The differential image is post-processed in combination with morphological processing, including dilation, erosion, and noise removal operations, to generate a clear target region image.

[0057] Furthermore, the formula of the frame difference method is as follows:

[0058] D(x, y) = |I t (x,y) - I t-1 (x,y)|

[0059] where D(x, y) represents the frame difference value of the current pixel point (x, y); I t (x, y) represents the gray value of the corresponding pixel point (x, y) at time t; I t-1 (x, y) represents the gray value of the corresponding pixel point (x, y) at time t - 1.

[0060] The formula of the dynamic background model is as follows:

[0061]

[0062] where P(I t (x, y)) represents the probability that the pixel point (x, y) belongs to the background at time t; K represents the number of Gaussian components; ω k,t (x, y) represents the weight of the k-th Gaussian component, that is, the importance of this component in the current background model, and the weight is updated dynamically according to the video stream: for example, when the static background dominates, the weight of the main background component is higher; if there are gestures or other dynamic targets, the components with lower weights may be updated to new background patterns; μ k,t (x, y) represents the mean value of the k-th Gaussian component, that is, the background color or gray value of the pixel point (x, y). In a dynamic scene, such as when a gesture occludes the background, the mean value will be adjusted dynamically; represents the variance of the k-th Gaussian component, that is, the degree of fluctuation of the pixel point (x, y) in this component. For example: the variance of the static background is small, indicating that the change range of the color or gray value is limited; the variance in areas with large illumination changes (such as the window edge) is large; represents the Gaussian distribution function; I t (x, y) represents the gray value of the corresponding pixel point (x, y) at time t.

[0063] S2: Based on the target region image, perform local brightness histogram analysis on the dynamic target region, adopt the region adaptive compensation algorithm, and at the same time combine the multi-scale edge enhancement technology to optimize the image quality;

[0064] Among them, the step S2 includes the following steps:

[0065] Based on the target region image, the local region sliding window method is used to partition and calculate the luminance histogram, and the luminance distribution characteristics of each window region are extracted;

[0066] Specifically, the target region image is divided into multiple small regions (such as windows of 8×8 or 16×16 pixels). The step size of the sliding window can be set to half of the window width (for example, the step size of an 8-pixel window is 4 pixels) to ensure coverage of the entire target region and reduce boundary omissions. This partitioning method generates a grid of sliding windows, and the image is segmented into overlapping small regions.

[0067] Calculate the luminance histogram for the pixel values within each sliding window, and count the frequency distribution of the gray values (for example, the gray range of 0 to 255 is divided into 16 intervals, and each interval represents the number of pixels in that gray segment).

[0068] Extract the following luminance distribution characteristics based on the histogram results:

[0069] Average luminance: The average luminance of the pixels within the window region, used to determine whether the window is a bright region or a dark region.

[0070] Luminance variance: The range of variation of the luminance values within the window region, used to evaluate the luminance uniformity of the local region.

[0071] Main gray interval: The gray interval with the highest histogram frequency, used to determine whether the region is too dark or overexposed.

[0072] Record the luminance characteristics of each window (such as mean, variance, and main gray interval), and display the results in the form of a pseudo-color heat map: Dark regions are marked with cold colors (blue), and bright regions are marked with warm colors (red). Output the luminance characteristic map, providing a basis for subsequent adaptive compensation.

[0073] Based on the luminance distribution characteristics, use an adaptive compensation algorithm based on pixel weights to generate different compensation weights for low-luminance and high-luminance regions, and dynamically adjust the luminance curve of the region; the adaptive compensation algorithm constructs a local illumination model for the dynamic target region and uses the gradient direction weighting strategy to dynamically adjust the image luminance;

[0074] Specifically, based on the luminance characteristic map, construct a local illumination model for the entire target region image:

[0075] Dark region illumination: Marked as low-luminance regions, which may require significant luminance enhancement.

[0076] Bright region illumination: Marked as high-luminance regions, which may require moderate luminance suppression.

[0077] Uniform region illumination: Marked as normal-luminance regions, which usually do not require adjustment.

[0078] The constructed illumination model represents the brightness distribution of each region in the image and provides guidance for the compensation strategy.

[0079] Compensation methods are adopted for different brightness regions respectively:

[0080] Low brightness region: Non-linearly enhance the brightness of dark pixels. For example, enhance the details of the low brightness region through gamma correction (gamma value less than 1). Consider the brightness gradient direction of the surrounding region and weight the compensation amount to avoid increased noise caused by excessive enhancement.

[0081] High brightness region: Compress the overexposed region. For example, apply brightness stretching to the region with gray value higher than 200 and compress the high brightness value towards the middle value (such as 150 - 200). Use logarithmic mapping or histogram equalization technology for smooth transition and enhance the details of the overbright region.

[0082] Uniform region: Do not perform brightness compensation and retain the original brightness to avoid destroying the original uniformity of the image.

[0083] Calculate the gradient direction and magnitude of each pixel in the image (such as using the Sobel operator) to judge the directionality of illumination change. Apply higher weight to the brightness adjustment along the gradient direction to ensure continuous and natural brightness change in the edge region and avoid introducing brightness discontinuities. Apply the brightness adjustment result and renormalize the brightness-compensated image to the standard gray scale range (0 - 255) to output the optimized image.

[0084] Utilize multi-scale edge enhancement technology to optimize the boundary sharpness and contrast of the dynamic target region by detecting and enhancing edge features in the image;

[0085] It should be noted that use an edge detection algorithm (such as the Canny operator) to extract edge features from the optimized image: Set low and high thresholds (such as 50 and 150) to ensure that the main boundaries are detected while suppressing noise. Output the edge detection result image and mark the main edges of the target region. Decompose the target region image into different scales, for example, generate sub-images with different resolutions through Gaussian pyramid decomposition. Enhance the edge features separately at each scale, for example, enhance the boundary strength through the Laplacian filter. Fuse the multi-scale enhanced edge features to form complete boundary information. Superimpose the enhanced edge region on the original image to improve the boundary sharpness and overall contrast of the target region. Make the contour of the dynamic target region more prominent through the edge weighting strategy.

[0086] Combine morphological processing technology to perform fine operations on the optimized image, including smoothing, denoising, and connected region segmentation, to generate the final optimized image.

[0087] S3: Based on the optimized image, use color histogram and feature matching algorithm to locate the user gesture area in real time, and adopt Kalman filter to dynamically update the gesture trajectory; in the S3 step, the gesture area location based on color histogram is specifically carried out by fusing the global histogram and the local area histogram to hierarchically locate the gesture area.

[0088] Among them, the step S3 includes the following steps:

[0089] Based on the optimized target image, calculate the color histogram, extract the color features of the gesture area, and lock the preliminary position of the gesture area through dynamic matching;

[0090] Specifically, convert the optimized target image into a color space (such as HSV or RGB) to meet the requirements of color feature extraction for the gesture area:

[0091] HSV space: Separate statistics for hue, saturation, and value, suitable for complex lighting scenarios.

[0092] RGB space: Directly statistics the distribution of red, green, and blue channels, used for rapid extraction of color significant areas.

[0093] Statistically analyze the color histogram distribution of the entire image, record the color distribution characteristics (such as dividing each channel into 32 or 64 intervals and counting the number of pixels in each interval). Take the global color histogram as the overall reference template. Divide the target image into multiple local areas (such as 16×16 or 32×32 windows). Calculate the color histogram of each local window one by one, and record the color distribution characteristics of each area. By comparing the similarity between the global color histogram and each local histogram (such as using Bhattacharyya distance or cosine similarity), screen the areas that best match the global color features. Aggregate the local areas with high similarity to initially locate the possible positions of the gesture area. Dynamically update the global histogram in multiple frames of images, and further lock the gesture area through continuous frame matching: in each frame, use the gesture area of the previous frame as a reference and rematch the color histogram in the nearby area to ensure the continuity and stability of the positioning.

[0094] Within the initially located gesture area, use the feature extraction algorithm to extract the local feature points of the gesture, and combine the matching algorithm to screen the candidate areas;

[0095] Specifically, within the initially located gesture area, adopt a key point detection algorithm (such as ORB, SIFT or SURF) to extract local feature points:

[0096] ORB: Suitable for real-time requirements, capable of quickly extracting corner and edge features.

[0097] SIFT / SURF: Suitable for high-precision requirements, extracting more robust feature points, especially under rotation and scale changes.

[0098] Generate descriptors for the extracted feature points and record the feature vectors of each point. Match the extracted local feature points with the gesture standard template or the feature points of the previous frame. Use feature matching algorithms (such as KNN or BFMatcher) to compare the similarity between feature points:

[0099] KNN: Find the two nearest neighbor features for each feature point and use the distance ratio to screen out reliable matching points.

[0100] BFMatcher: Directly calculate the Euclidean distance and select the best match.

[0101] Filter out the most likely gesture area based on the distribution density of the matching points and the number of feature points. If multiple regions are similar, calculate the region confidence by combining the color histogram and the feature point distribution, and retain the candidate region with the highest confidence as the final gesture area.

[0102] Adopt the Kalman filter algorithm to predict the region position of the next frame according to the position and motion state of the gesture area in the current frame, and dynamically adjust the Kalman filter parameters in combination with the actual observation values;

[0103] Specifically, initialize the state variables of the Kalman filter, including: the gesture position in the current frame (center point coordinates and speed). The prediction covariance matrix, set the initial observation error and system noise parameters. Based on the gesture position and speed in the current frame, predict the possible position of the gesture area in the next frame. The prediction results include the center point and bounding box of the gesture area. In the next frame, update the state variables of the Kalman filter according to the deviation between the actually observed gesture area and the predicted position: adjust the weights of the system noise and observation noise to ensure accurate tracking in a dynamic environment. Correct the prediction error to make the filtering result more consistent with the actual situation. Record the final gesture position of each frame in the trajectory list to generate a sequence of trajectory points of the gesture movement.

[0104] Conduct time series analysis on the tracking trajectory of the gesture and optimize the gesture trajectory using the trajectory smoothing algorithm.

[0105] Specifically, analyze the time series data of the gesture trajectory and extract the following features:

[0106] Motion pattern: such as linear motion, arc motion or complex path.

[0107] Speed change: Calculate the distance change between trajectory points and analyze the speed change of the gesture movement.

[0108] Optimize the trajectory points using a smoothing algorithm (such as Kalman smoothing or exponential moving average): Remove discrete points that may be caused by jitter to ensure the continuity of the trajectory curve. Optimize the inflection points of the trajectory to make the motion pattern clearer. Output the smoothed and optimized trajectory, which can be used for dynamic feature analysis of gesture recognition or drawing the gesture path.

[0109] S4: Extract the key features of the gesture through deep learning, including shape, trajectory, and dynamic pattern, and make an accurate comparison with the standard model library, and output the gesture category and the corresponding function instruction.

[0110] Among them, the step S4 includes the following steps:

[0111] Extract spatial features, trajectory features, and dynamic pattern features from the optimized gesture image based on the convolutional neural network algorithm to form a feature vector;

[0112] Specifically, perform basic processing on the collected gesture image, such as adjusting brightness, contrast, and cropping the image to focus on the gesture area. Adjust the image to a unified size for subsequent processing (for example, 224x224 pixels). Denoise the image to reduce the interference of background noise. Use a convolutional neural network (CNN) to extract key spatial features in the image, such as the contour, edge, and texture information of the gesture. In a set of continuous gesture actions, use a deep model to capture the action trajectory and extract features related to the gesture movement direction and speed. Identify the temporal variation law of the gesture through a dynamic pattern analysis model and extract its dynamic features. Integrate the extracted spatial features, trajectory features, and dynamic pattern features into a unified feature vector as the representation of the current gesture for subsequent comparison.

[0113] Use a large amount of standard gesture data to train the convolutional neural network algorithm to build a standard model library containing multiple gesture categories;

[0114] Specifically, collect image and video data of various gesture actions to ensure that common gesture categories are included, such as waving, pointing, and digital gestures. The data collection should cover different lighting, backgrounds, and gesture execution speeds to enhance the generalization ability of the system. Classify and label the collected data, and assign a corresponding category label to each gesture. Clean up incomplete and unclear data to ensure the quality of the training data. Use a standard data set to train the convolutional neural network to let the model learn the features of each gesture category. During the training process, continuously adjust the model parameters to improve the recognition accuracy. Use a multi-round iteration method to train the model, and evaluate the model performance with a validation data set after each round of training. Save the trained model as a standard model library, which contains each gesture category and its corresponding feature representation as the benchmark for subsequent comparison.

[0115] Match the extracted feature vectors with the standard model library, calculate the similarity using cosine similarity, and output the category that is closest to the current gesture feature;

[0116] Based on the classification result, generate and output the corresponding function instruction through a preset mapping rule.

[0117] Furthermore, the formula for the cosine similarity is as follows:

[0118]

[0119] Among them, Similarity(F, S) represents the similarity between the current gesture feature F and the standard gesture model S, and the output range is from 0 to 1. The closer the value is to 1, the higher the matching degree between the current gesture and the standard model; F i represents the key feature vector of the current gesture; W i represents the importance weight of different feature modalities, which is dynamically adjusted according to the gesture category and task requirements; S i represents the feature vector of a certain category of gesture in the standard gesture model library; ΔT represents the time difference between the current gesture trajectory and the standard gesture model trajectory; K represents the adjustment factor of the time characteristic, which is used to balance the influence of the time characteristic on the overall similarity; C represents the time decay coefficient, which controls the influence intensity of the time difference on the similarity and is applicable to the precise matching of dynamic gestures; e -C·ΔT represents the time decay function; ∈ represents a small positive number; n represents the dimension number of the feature vector.

[0120] In summary, the present invention dynamically constructs a background model through the frame difference method and the Gaussian mixture model, efficiently excluding the interference of static backgrounds and non-target actions. Using morphological processing to further clean up noise and interference regions, generating a clear target region image, providing high-quality input for subsequent recognition. Based on local luminance histogram analysis and regional adaptive compensation algorithms, effectively enhancing the detail performance of low-luminance regions in the image and smoothing the overexposure problem in high-luminance regions. Combining multi-scale edge enhancement techniques, significantly enhancing the boundary clarity and contrast of the dynamic target region, improving the accuracy and robustness of gesture recognition.

[0121] The present invention is based on a method that combines color histogram and local feature matching to achieve precise positioning of the gesture region, adapting to different illuminations and complex backgrounds. Using the Kalman filter algorithm to dynamically update the gesture position and trajectory, effectively smoothing the trajectory noise and optimizing the continuity, providing reliable input data for dynamic gesture recognition. Extracting the spatial features, trajectory features, and dynamic pattern features of the gesture through a deep learning model, realizing multi-modal representation of gesture actions, and adapting to complex gesture changes. Using the trained standard model library combined with the cosine similarity algorithm to accurately compare gesture categories, improving the accuracy and efficiency of recognition.

[0122] The present invention can adapt to a variety of complex scenarios, including different lighting conditions, dynamic background changes, and rapid movements of gestures. The introduction of dynamic characteristics and temporal decay further enhances the robust recognition ability for dynamic gestures, especially maintaining high performance in long-duration and multi-action sequences.

[0123] The above embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary engineering and technical personnel in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A gesture recognition method with anti-interference optimization, characterized in that: The following steps are involved: S1: Obtain gesture video stream through the camera, use frame difference method and Gaussian mixture model to build dynamic background model, eliminate static background and non-target motion interference, and generate target area image; S2: Based on the target area image, a local brightness histogram analysis is performed on the dynamic target area, a regional adaptive compensation algorithm is adopted, and a multi-scale edge enhancement technology is combined to optimize the image quality; S3: Based on the optimized image, the color histogram and feature matching algorithm are used to locate the user gesture area in real time, and the Kalman filter is used to dynamically update the gesture trajectory; S4: Extract key features of gestures, including shape, trajectory and dynamic pattern, through deep learning, accurately compare them with the standard model library, and output gesture categories and corresponding functional instructions.

2. The anti-interference optimized gesture recognition method according to claim 1, characterized in that: The step S1 comprises the following steps: The dynamic video stream containing the user's gestures is collected in real time through the camera, and the video stream is frame-segmented to extract continuous frame images; The frame difference method is used to calculate the difference image between adjacent frames and preliminarily remove the static background area; Based on the Gaussian mixture model, a dynamic background model is constructed for the video stream, and background fitting and separation are performed on the non-target area. The difference image is post-processed in combination with morphological processing, including dilation, erosion and noise removal operations, to generate a clear image of the target area.

3. The anti-interference optimized gesture recognition method according to claim 2, characterized in that: The formula of the frame difference method is as follows: D(x,y)=|I t (x,y)-I t-1 (x,y)| Where D(x, y) represents the frame difference of the current pixel point (x, y); I t (x, y) represents the gray value of the corresponding pixel point (x, y) at time t; t-1 (x, y) represents the grayscale value of the corresponding pixel point (x, y) at time t-1.

4. The anti-interference optimized gesture recognition method according to claim 2, characterized in that: The formula of the dynamic background model is as follows: Among them, P(I t (x, y)) represents the probability that the pixel (x, y) belongs to the background at time t; K represents the number of Gaussian components; ω k,t (x, y) represents the weight of the kth Gaussian component; μ k,t (x, y) represents the mean of the kth Gaussian component; represents the variance of the kth Gaussian component; represents the Gaussian distribution function; I t (x, y) represents the grayscale value of the corresponding pixel point (x, y) at time t.

5. The anti-interference optimized gesture recognition method according to claim 1, characterized in that: The step S2 comprises the following steps: Based on the target area image, the local area sliding window method is used to partition the brightness histogram and extract the brightness distribution characteristics of each window area. According to the brightness distribution characteristics, an adaptive compensation algorithm based on pixel weights is used to generate different compensation weights for low-brightness and high-brightness areas, and dynamically adjust the brightness curve of the area; Using multi-scale edge enhancement technology, the edge features in the image are detected and enhanced to optimize the boundary clarity and contrast of the dynamic target area; The optimized image is refined by combining morphological processing techniques, including smoothing, denoising and connected region segmentation, to generate the final optimized image.

6. The anti-interference optimized gesture recognition method according to claim 1, characterized in that: The step S3 comprises the following steps: Based on the optimized target image, the color histogram is calculated, the color features of the gesture area are extracted, and the preliminary position of the gesture area is locked through dynamic matching; In the gesture area initially located, the local feature points of the gesture are extracted using the feature extraction algorithm, and the candidate areas are screened in combination with the matching algorithm; The Kalman filter algorithm is used to predict the position of the gesture area in the next frame based on the position and motion state of the gesture area in the current frame, and the Kalman filter parameters are dynamically adjusted based on the actual observation values; The tracking trajectory of the gesture is analyzed in time series, and the trajectory smoothing algorithm is used to optimize the gesture trajectory.

7. The anti-interference optimized gesture recognition method according to claim 1, characterized in that: The step S4 comprises the following steps: Based on the convolutional neural network algorithm, spatial features, trajectory features and dynamic pattern features are extracted from the optimized gesture image to form a feature vector; Use a large amount of standard gesture data to train the convolutional neural network algorithm and build a standard model library containing multiple gesture categories; The extracted feature vector is matched with the standard model library, and the cosine similarity is used to calculate the similarity, and the category closest to the current gesture feature is output; Based on the classification results, the corresponding functional instructions are generated and output through the preset mapping rules.

8. The anti-interference optimized gesture recognition method according to claim 7, characterized in that: The formula for the cosine similarity is as follows: Where Similarity(F, S) represents the similarity between the current gesture feature F and the standard gesture model S; F i Represents the key feature vector of the current gesture; W i Represents the importance weights of different feature modes; S i represents the feature vector of a certain category of gestures in the standard gesture model library; ΔT represents the time difference between the current gesture trajectory and the standard gesture model trajectory; K represents the adjustment factor of the time characteristic; C represents the time attenuation coefficient; e -C·ΔT represents the time decay function; ∈ represents a small positive number; n represents the dimension of the feature vector.

9. The anti-interference optimized gesture recognition method according to claim 5, characterized in that: The adaptive compensation algorithm constructs a local illumination model of the dynamic target area and dynamically adjusts the image brightness using a gradient direction weighting strategy.

10. The anti-interference optimized gesture recognition method according to claim 1, characterized in that: The gesture area positioning based on the color histogram in step S3 is specifically performed by fusing the global histogram and the local area histogram to hierarchically position the gesture area.

Citation Information

Cited By

  • Intelligent light-operated product power supply control system based on gesture recognition

    CN120499900A

  • A power control system for intelligent light-controlled products based on gesture recognition

    CN120499900B

  • Fiber endoscope image focus detection method and system

    CN120543553A

  • Adaptive dimming method for dynamic interference in infrared view field range

    CN120730149A

  • Natural man-machine interaction sign language recognition method

    CN120821369A