Gesture optical motion capture recognition method based on thermal imaging
Through infrared thermal imaging and bone point patch combined with YOLOv5 and HRNet models, high-precision gesture recognition in low-visibility environments is achieved, solving the recognition problem of traditional technology under low light conditions, reducing cost and complexity.
Patent Information
- Application Number
- CN202510115701.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional visible light motion capture technology is difficult to achieve high-precision hand motion recognition in low visibility and weak light environments, and is costly and inconvenient to arrange the scene.
Infrared thermal imaging technology is adopted to attach bone point patches to finger joints, infrared images are collected using thermal imaging equipment, grayscale expansion, edge detection, image binarization and bone point tracking, gesture recognition is combined with YOLOv5 and HRNet models, and gesture classification is used using graph convolution network.
Achieve high-precision gesture recognition in low-visibility environments solves the accuracy problems of traditional technologies under low-light conditions and reduces cost and complexity.
Smart Images

Figure CN120108033A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gesture recognition, and in particular to a method for gesture optical motion capture and recognition based on thermal imaging. Background Art
[0002] With the continuous advancement of science and technology, human-computer interaction (HCI) technology is also experiencing rapid development. Traditional human-computer interaction methods mainly rely on input devices such as keyboards, mice and touch screens, which can meet the needs of users in many cases. However, they usually require users to make physical contact, and may not be intuitive or efficient in complex or dynamic application environments, and may not even guarantee recognition accuracy and sensitivity in extreme environments. In recent years, gesture recognition methods based on infrared thermal imaging technology have become an emerging research direction in the field of HCI because of their advantages such as no physical contact, ability to work in dark environments, and ability to capture subtle movements.
[0003] Gesture recognition technology has been widely used in C-end scenarios such as VR games. Ultraleap's gesture recognition module has a built-in infrared LED light source. After the infrared LED light shines on the hands, the light is reflected back to the infrared camera of the gesture recognition module, realizing gesture recognition based on optical data. Based on this, gesture recognition does not have to rely solely on the surrounding visible light, thereby ensuring relatively stable optical tracking. At the same time, the mate60 series mobile terminals recently launched by Huawei use gesture sensors and micro-core NPUs to recognize people's waving gestures in real time. Because it uses infrared sensors, it can recognize simple gestures in darker environments to a certain extent, but the sensitivity and accuracy will also be greatly reduced.
[0004] Considering that the current optical motion capture technology almost all uses the principle of computer vision, multiple high-speed cameras track the feature points of the object from different angles to complete the capture of the whole body movement, which has high requirements for ambient light. In animation production, due to the need to capture high-precision character movements, the characters are often required to wear many reflective marking points (marker points), and in a specific environment, multiple high-speed cameras at different angles are required to simultaneously collect the motion trajectory. The cost of use is high and the scene arrangement is inconvenient. In the real environment with low visibility, it is impossible to capture the character movements with high precision. A method of applying infrared thermal imaging technology to motion capture and recognition is needed to effectively solve this problem. Summary of the invention
[0005] In order to solve the technical problems existing in the prior art, the present invention provides a method for gesture optical motion capture and recognition based on thermal imaging, which can capture human gestures with high precision in a low-visibility environment, perform gesture recognition on the gestures and obtain the categories of the gestures, and can solve the problem of high-precision hand motion recognition faced by traditional visible light motion capture technology in low-visibility and weak light environments.
[0006] The present invention aims to provide a method for capturing and recognizing hand gesture optical motions based on thermal imaging.
[0007] The purpose of the present invention can be achieved by adopting the following technical solutions:
[0008] A method for capturing and recognizing gesture optical motion based on thermal imaging comprises the following steps:
[0009] S1. Collect infrared images of gestures using an infrared thermal imaging gesture recognition device, where the infrared thermal imaging gesture recognition device includes a thermal imaging device and a plurality of bone point patches; the plurality of bone point patches are attached to a plurality of finger joints, and the thermal imaging device locates the bone points of the fingers through the plurality of bone point patches;
[0010] S2, preprocessing the collected infrared image of the gesture to obtain a preprocessed infrared image;
[0011] S3, performing grayscale expansion, gesture edge detection, edge image filling and image binarization segmentation on the pre-processed infrared image to obtain a binary gesture area image;
[0012] S4, extracting hand information from the binary gesture area image, identifying and marking the hand skeleton points based on the hand information, and obtaining a connection diagram of the hand skeleton point motion trajectories;
[0013] S41, training a YOLOv5 target detection model, extracting hand information from the gesture area image through the trained YOLOv5 target detection model, and cropping a hand area from the gesture area image according to the hand information;
[0014] S42, building a skeleton point running tracking model based on the HRNet model, training the skeleton point running tracking model, and obtaining a trained skeleton point running tracking model;
[0015] S43, using the trained skeleton point running tracking model to identify and mark the skeleton points of the hand area, tracking and analyzing the changes of the skeleton points of the hand and predicting the motion trajectory, and drawing a motion trajectory diagram of the skeleton points of the hand.
[0016] S5. Perform gesture recognition on the motion trajectory of the hand skeleton points, and output a sequence of gesture categories and the corresponding Chinese meanings.
[0017] Specifically, step S3 includes:
[0018] S31, performing grayscale expansion on the infrared image to improve the contrast of the infrared image, and obtaining an infrared image after grayscale expansion;
[0019] S32, performing gesture edge detection on the infrared image, and extracting the contour edge of the gesture area in the infrared image after grayscale expansion based on the Canny edge detection algorithm;
[0020] S33, based on the obtained contour edge of the gesture area, performing area filling on the contour edge of the gesture area to obtain a complete gesture area;
[0021] S34, performing image binarization segmentation on the gesture area to obtain a binary gesture area image.
[0022] Specifically, the grayscale expansion of the infrared image includes: converting the infrared image into a grayscale image and removing the noise from the grayscale image; calculating the original grayscale value range of the grayscale image of the infrared image, setting the target grayscale range, transforming the grayscale value using a linear formula, and stretching the original grayscale value distribution of the grayscale image.
[0023] Specifically, the gesture edge detection is performed on the infrared image, and the contour of the gesture in the infrared image after grayscale expansion is extracted based on the Canny edge detection algorithm, including:
[0024] S321, performing Gaussian filtering on the infrared image after grayscale expansion through a Gaussian filter;
[0025] S322, calculating the gradient amplitude image and angle image of the smoothed image by using the Sobel operator, applying non-maximum suppression to the gradient amplitude image, retaining the image point with the maximum local gradient amplitude image, removing non-edge pixels, and obtaining a gesture edge image;
[0026] S323, setting a low threshold and a high threshold, extracting strong edges and potential edges from the gesture edge image based on the low threshold and the high threshold, connecting the potential edges with the strong edges to form a complete edge, and obtaining the edge of the gesture area.
[0027] Specifically, the step S41 includes:
[0028] Collect image data sets and annotate the image data sets. The annotated information includes the bounding box of the hand area in each image, the size of the hand, and the category of the hand, so as to obtain an annotated data set.
[0029] Use the labeled data set to train the YOLOv5 target detection model to obtain a trained YOLOv5 target detection model;
[0030] The hand information is extracted from the binarized gesture area image of the trained YOLOv5 target detection model to separate the hand area.
[0031] Specifically, the step S42 includes:
[0032] A skeleton point running tracking model is built based on the HRNet model. The network structure of the HRNet model is improved. A deep convolutional attention mechanism is introduced when fusing features at each layer of the HRNet model. Sequential modeling is performed when the Transformer module is added after the output of the HRNet model to obtain a skeleton point running tracking model.
[0033] Collect and annotate the hand or human skeleton point data as training data, combine the key point detection loss and sequence prediction loss to get the joint loss, and use the training data to train the skeleton point tracking model based on the joint loss.
[0034] Specifically, the step S43 extracts human skeleton points from the dynamic hand contour features of continuous frames and predicts the motion trajectory through the skeleton point running tracking model to obtain a skeleton point motion trajectory connection diagram, including:
[0035] The trained skeleton point tracking model is used to detect skeleton points in the hand area, and the positions of skeleton points in the gesture area are predicted and marked in the hand area in each frame to generate skeleton point coordinates.
[0036] The marked hand skeleton points are connected according to the skeleton point coordinates to form a complete hand skeleton image, and the changes of the hand skeleton points in multiple hand skeleton image frames are tracked and analyzed to obtain the movement trajectory of the hand skeleton points in continuous frames.
[0037] Specifically, the gesture recognition is performed on the motion trajectory diagram of the hand skeleton points, and the gesture category sequence and the corresponding Chinese meaning are output, including:
[0038] The graph convolutional network (GCN) is used to perform gesture recognition on the motion trajectory of the hand skeleton points and output a sequence of gesture categories.
[0039] The gesture category sequence output by the graph convolutional network GCN is input into the pre-trained language model, and the gesture category sequence is converted into Chinese description through the pre-trained language model to obtain the Chinese meaning corresponding to the gesture category.
[0040] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0041] The present invention provides a method for capturing and identifying gesture optical motion based on thermal imaging, wherein multiple bone point patches are attached to multiple finger joints, a thermal imaging device locates the finger bone points through the multiple bone point patches, an infrared thermal imaging gesture recognition device is used to collect infrared images of gestures, grayscale expansion, gesture edge detection, edge image filling and image binarization segmentation are performed on the infrared images, a bone point operation tracking model is constructed, and bone point recognition and annotation of the hand area are performed through the bone point operation tracking model, a motion trajectory diagram of the hand bone points is drawn, and finally, gesture recognition is performed on the bone point motion trajectory connection diagram to obtain the Chinese meaning of the gesture, and the gesture actions of people can be captured with high accuracy in a low-visibility environment, and gesture recognition is performed on the gesture actions and the gesture categories are obtained. The method can solve the high-precision hand motion recognition problem faced by traditional visible light motion capture technology in low-visibility and weak-light environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying creative work.
[0043] Figure 1 is a flow chart of a method for optical motion capture and recognition of gestures based on thermal imaging in an embodiment of the present invention;
[0044] Figure 2 is a comparison diagram of edge detection results of infrared gesture images in an embodiment of the present invention;
[0045] Figure 3 is an edge image filling effect diagram in an embodiment of the present invention;
[0046] Figure 4 This is a skeleton point marking diagram effect diagram in an embodiment of the present invention. DETAILED DESCRIPTION
[0047] The technical solution of the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is obvious that the described embodiments are part of the embodiments of the present invention, not all of the embodiments, and the implementation of the present invention is not limited to this. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0048] Embodiment 1:
[0049] like Figure 1As shown, the method for capturing and recognizing gesture optical motion based on thermal imaging of the present invention comprises the following steps:
[0050] S1. Collect infrared images of gesture actions using an infrared thermal imaging gesture recognition device. The infrared thermal imaging gesture recognition device includes a thermal imaging device and a plurality of bone point patches. The plurality of bone point patches are attached to a plurality of bone points on the hand.
[0051] Thermal imaging equipment uses infrared radiation to detect and measure the temperature distribution of objects, thereby generating a visual heat map. In motion capture, especially gesture recognition, human-computer interaction, infrared thermal imaging technology can not only capture the temperature changes caused by human muscle activity like traditional computer vision, thereby identifying different gestures and actions, but also solve the problem that traditional visible light optical capture technology cannot be used normally under low visibility and low light conditions. Even in a low-visibility environment, infrared images can still characterize the temperature difference between the hand and the environment to enable the system to recognize the gesture target.
[0052] In this embodiment, the thermal imaging device can use a thermal imager from Hikvision, with an infrared resolution of 256×192, a pixel size of 12μm, and a frame rate of 25Hz. It has good parameters such as resolution, temperature sensitivity (thermal sensitivity), and frame rate to ensure sufficient detail capture and fast response. The thermal imager collects infrared images of gestures, saves the collected infrared images in a standard format (such as TIFF or PNG), and feeds the collected information into a convolutional neural network model to achieve fast and accurate recognition of multiple gestures.
[0053] In order to improve the recognition accuracy of different finger joints, distinguish and filter out the user's own movements, and track the special radiation energy points in the acquired thermal imaging images, the present invention attaches multiple bone point patches with different heat insulation capabilities to multiple key finger joints, and the thermal imaging device locates the finger bone points through the multiple bone point patches. Preferably, the bone point patch is a mixture of thermoplastic polyurethane (TPU) and boron nitride (BN), and the boron nitride content is 10%-30%. The mixture of thermoplastic polyurethane (TPU) and boron nitride (BN) is pressed into a plate and then cut into a patch of approximately a preset size to obtain a bone point patch of a preset size. Multiple bone point patches are attached to multiple bone points on the hand, and the thermal imager can accurately locate the finger bone points through the multiple bone point patches. Thermoplastic polyurethane (TPU) is a polymer material with elastic and thermoplastic properties. It has high elasticity and excellent mechanical properties. It has rubber-like elasticity and can recover to its original state within a large strain range. It also has excellent tensile strength and can withstand large mechanical stress. It is extremely wear-resistant and suitable for high-friction environments. Good flexibility and low-temperature performance allow TPU to maintain good flexibility over a wide temperature range and is suitable for low-temperature environments. TPU is also an environmentally friendly material with recyclability. However, the thermal conductivity of TPU is not high and cannot meet the performance requirements of bone point patches. When TPU (thermoplastic polyurethane) is mixed with boron nitride (BN), the properties of the two materials are combined to form a composite material with special properties. The performance of this mixture mainly depends on the morphology of boron nitride (such as nanosheets, micron particles) and the filling ratio.
[0054] In this embodiment, 80% thermoplastic polyurethane (TPU) and 20% boron nitride are mixed to obtain a mixture, and the mixture is pressed into a plate by a vulcanization molding machine and then cut into a square patch of about 1 cm long to obtain a bone point patch. The bone point patch is a material with high thermal conductivity and good flexibility. When TPU is mixed with boron nitride in different proportions, the properties of the obtained mixture will also be different, and the specific changes are shown in Table 1.
[0055]
[0056]
[0057] The addition of bone point patches significantly improves the thermal conductivity of TPU and has excellent thermal conductivity. The filling network formed by boron nitride particles in the TPU matrix can significantly improve the rigidity and tensile strength of the composite material. Despite the increased rigidity, the flexibility and elasticity of TPU are retained to a certain extent, depending on the filling ratio of boron nitride. Improved friction performance, lubrication effect: Boron nitride has a layered structure similar to graphite, which gives the material good lubrication properties and reduces the friction coefficient. Enhanced wear resistance: The composite material shows a lower wear rate in dynamic applications and is suitable for high friction environments.
[0058] S2, preprocessing the infrared image of the gesture to obtain a preprocessed infrared image;
[0059] The preprocessing of the infrared image of the gesture includes image calibration, digital noise reduction, data enhancement, target temperature range highlighting and other operations on the captured infrared image to obtain a preprocessed infrared image. The image noise of the preprocessed infrared image is significantly reduced, the image contrast is enhanced, and the hand features are prominent.
[0060] Among them, image calibration includes operations such as rotation and flipping. 3D DNR digital noise reduction is used to digitally reduce the noise of infrared images. When infrared thermal imaging technology is used, image noise will appear randomly, and the noise in each frame is different. 3DDNR automatically filters non-overlapping information (noise) by comparing several adjacent frames of images, so that the image noise is significantly reduced, and the output infrared image will be purer and more delicate. Digital detail enhancement of infrared images can be performed through the DDE (Digital Detail Enhancement) nonlinear image processing algorithm, which can retain the details in high dynamic range images, thereby matching the total dynamic range of the original image. Even in extreme temperature dynamic range scenes, the details of the target object can be seen clearly. DDE is mainly used to enhance the details in infrared images so that they can show the details of high dynamic range images within a limited dynamic range (such as 8-bit images). The target temperature range is highlighted by the isothermal mode, and the target temperature range (human skin surface temperature range 36~38℃) or Y16 range that needs to be paid attention to in the infrared image is highlighted by a bright pseudo-color band for user analysis and processing.
[0061] S3. Perform grayscale expansion, gesture edge detection, edge image filling and image binarization segmentation processing on the pre-processed infrared image to obtain a binary gesture area image.
[0062] S31, performing grayscale expansion on the infrared image to improve the contrast of the infrared image and obtain the infrared image after grayscale expansion.
[0063] In an environment with a large temperature difference, the contrast of the infrared thermal image of the human body is low and the details are blurred. The infrared image needs to be grayscale expanded. Grayscale expansion can improve the contrast of the infrared image, making the bright and dark areas more distinct, and making the target area in the infrared image clearer.
[0064] Specifically, the grayscale expansion of infrared images includes: converting the infrared image into a grayscale image, denoising the grayscale image, such as using median filtering, mean filter, etc. to denoise. Calculate the original grayscale value range of the grayscale image of the infrared image, the original grayscale value range includes the original grayscale value minimum value and the original grayscale value maximum value, set the target grayscale range (such as 10,255), use a linear transformation formula or a nonlinear formula to transform the grayscale value, and stretch the original grayscale value distribution to a wider range.
[0065] Among them, the linear transformation formula is:
[0066]
[0067] Among them, g(x, y) is the grayscale value after expansion, f(x, y) is the original grayscale value, a and b are the target grayscale ranges, and c and d are the minimum and maximum grayscale values of the original image, respectively.
[0068] S32. Perform gesture edge detection on the infrared image, extract the contour of the gesture in the infrared image after grayscale expansion based on the Canny edge detection algorithm, enhance the boundary information of the gesture area, and provide a basis for subsequent gesture recognition or segmentation.
[0069] Edge detection mainly identifies significant change areas (i.e. edges) by calculating the gradient of image grayscale value changes. The Canny edge detection algorithm is selected to perform gesture edge detection on infrared images. Canny edge detection is a multi-level edge detection algorithm. The threshold range (low threshold and high threshold) of the algorithm is set to extract the contour of the gesture in the infrared image after grayscale expansion. The identified edges are detailed and continuous, which is suitable for accurate detection. Canny edge detection is a multi-level edge detection algorithm and a filter-based detection algorithm. However, non-maximum suppression and double threshold processing are added on the basis of filtering. The input image is smoothed using a Gaussian filter. The gradient amplitude image and angle image are calculated for the smoothed image. Non-maximum suppression is applied to the gradient amplitude image, and double threshold processing is performed on the image after non-maximum suppression.
[0070] Specifically, step S32 includes the following steps:
[0071] S321, smoothing the infrared image after grayscale expansion by using a Gaussian filter to obtain a smoothed infrared image. Using a Gaussian filter can smooth the image, reduce background interference, and reduce the interference of noise on edge detection.
[0072] Gaussian smoothing filter belongs to the category of linear filters. Its working principle is very similar to that of mean filter, and it can also reduce noise and smooth images. The biggest difference between Gaussian filter and mean filter lies in the filter window. Mean filter determines the pixel of the point by the average value of all pixels in the adjacent area, while Gaussian filter selects a filter window of a certain size, and the weight of each pixel in the window is not the same. The weighted average of all pixels is used as the final output value of the pixel of the point. It is precisely because of the characteristic that the weight coefficient of the pixel of the Gaussian filter is smaller as the distance from the center pixel is larger, that the loss of details in the Gaussian filter is much smaller than that of the mean filter during the filtering process. Gaussian smoothing filter retains more image information while filtering out noise.
[0073] S322, calculating the gradient amplitude image and angle image of the smoothed image by using the Sobel operator, applying non-maximum suppression to the gradient amplitude image, retaining the image points with the largest local gradient amplitude image, removing non-edge pixels, and obtaining a gesture edge image.
[0074] The gradient amplitude image and angle image of the smoothed image are calculated by the Sobel operator, and then non-maximum suppression is applied to the gradient amplitude image, which mainly retains the significant edge information in the image and removes the interference of non-edge pixels, so that a fine gesture edge image can be obtained. The obtained gesture edge image has a clear edge contour. After non-maximum suppression, the edge becomes clearer, and only the pixels at the position of strong gradient amplitude are retained. Non-edge pixels are removed, and pixels that do not meet the edge characteristics (such as small gradient amplitude or non-local maximum) are set to zero, reducing noise interference. The edges are refined, and the retained edges are close to the width of a single pixel, which is convenient for subsequent edge detection or image analysis.
[0075] Specifically, the Sobel operator calculates the image gradient magnitude and direction including the following steps:
[0076] The Sobel operator is used to calculate the gradient. The Sobel operator is an edge detection operator based on convolution operation. It estimates the gradient by detecting the speed of the change of the gray value of the image. The Sobel convolution kernel is defined, including the horizontal convolution kernel and the vertical convolution kernel. Among them, the horizontal convolution kernel is defined as G x :
[0077]
[0078] The vertical convolution kernel is defined as G y :
[0079]
[0080] Convolution calculation, G xAnd the vertical convolution kernel G y Convolve with the input image I to obtain the horizontal gradient image I x (x, y) and vertical gradient images I y (x, y):
[0081] I x (x, y) = I*G x , I y (x, y) = I*G y
[0082] Among them, * represents the convolution operation.
[0083] Calculate the gradient magnitude and gradient direction of each pixel to obtain the gradient magnitude image and gradient direction image respectively. The calculation formula of the gradient magnitude is:
[0084]
[0085] In the formula, the gradient amplitude M(x, y) represents the intensity of grayscale change of each pixel.
[0086] The gradient amplitude reflects the intensity of the grayscale change of the image. In the image, the places where the grayscale changes dramatically are usually edges, and the gradient amplitude is large. For example, the grayscale value of the object boundary or the area with strong contrast in the image will change significantly, and the calculated gradient amplitude will be very large at this time. By thresholding the gradient amplitude, the significant edges in the image can be extracted. The gradient amplitude is an important basis for Canny edge detection. By processing the gradient amplitude image (such as applying thresholds, non-maximum suppression, etc.), the edge information in the image can be effectively extracted. Larger gradient amplitudes usually correspond to edges in the image, and smaller gradient amplitudes usually correspond to flat areas or backgrounds in the image. In computer vision tasks, the gradient amplitude can be used as a feature of the image for subsequent feature matching, object recognition, image segmentation and other tasks. For example, the gradient amplitude image can be used as an input feature for training a machine learning model to help the model understand the edge information of objects in the image. The gradient amplitude can also be used to detect changes in the contour of an object under different viewing angles, thereby inferring the depth information of the object.
[0087] The calculation formula for the gradient direction is:
[0088]
[0089] Among them, the gradient direction θ(x, y) represents the direction of the grayscale change of each pixel. Usually, the gradient direction is expressed as an angle, ranging from [0°, 180°] or [-180°, 180°].
[0090] The gradient direction indicates the direction in which the grayscale changes fastest in the image, that is, the direction of the edge. At each pixel in the image, the gradient direction points to the direction in which the grayscale value of the point changes the most. The calculation of the edge direction is crucial to understanding the structure of the image. For example, the gradient direction can help us understand the boundary morphology of an object and provide information about the direction of the object's contour. When performing edge detection, the non-maximum suppression (NMS) step relies on the gradient direction. In this process, the gradient direction helps determine which pixels should be retained as edges and which pixels should be suppressed. In NMS, the algorithm compares each pixel with its two neighboring pixels in that direction based on the gradient direction, and only retains it as an edge if the gradient magnitude of the current pixel is greater than that of the neighboring pixels. In some segmentation algorithms, such as edge-based image segmentation, the gradient direction helps guide the segmentation process. By analyzing the gradient direction and magnitude of different regions in the image, the segmentation algorithm can be helped to identify objects and backgrounds in the image, thereby segmenting the image into different regions.
[0091] The combined effect of gradient magnitude and direction, the combined use of gradient magnitude and gradient direction is the most important step in edge detection. After the Sobel operator is calculated, the gradient magnitude is used to find strong edges, and the gradient direction helps locate the exact position of the edge. The gradient magnitude is used to determine which positions are candidate points for the edge, and the gradient direction helps determine the specific direction of the edge and accurately locate the shape of the edge. The gradient magnitude and gradient direction provide local features of different regions in the image. By counting these features, texture analysis, object recognition and classification can be performed. For example, some texture patterns may have strong gradient magnitudes and specific directional laws, and calculating these features can help perform more accurate image analysis. In target detection and tracking, gradient magnitude and direction can be used as effective information to describe target features. By analyzing the edge strength (gradient magnitude) and edge direction of the target area, the accuracy of detection and the robustness of tracking can be improved.
[0092] Local Gradient Magnitude refers to the gradient magnitude of a pixel in an image within its local neighborhood. The gradient magnitude itself measures the intensity of the grayscale change of the image, while the local gradient magnitude is calculated in a small range around the pixel (such as a 3x3 or 5x5 neighborhood) to reflect the intensity of the gradient change in the area.
[0093] Calculate the local gradient magnitude, retain the image point with the largest local gradient magnitude image, and remove non-edge pixels, including the following steps:
[0094] Calculate the gradient of the image and use the Sobel operator to calculate the horizontal gradient (G) and vertical gradient (G) of each pixel. The formula for calculating the gradient amplitude is:
[0095]
[0096] Among them, M(x,y) is the gradient amplitude of the pixel.
[0097] Select a local neighborhood range and define a neighborhood window (usually a 3x3 or 5x5 area). For example, select several pixels around the current pixel in the image (including the current pixel itself) to calculate the local gradient amplitude.
[0098] To calculate the local gradient magnitude, you can calculate the mean, maximum, or other statistical values of the gradient magnitude within the neighborhood window. These statistical values represent the local gradient magnitude. For example, you can use the maximum gradient magnitude to represent a strong edge in the local area:
[0099] M local (x, y)=max{M(x′, y′)|(x′, y′)∈ domain (x, y)}
[0100] Among them, M local (x, y) is the local gradient magnitude of the current pixel (x, y), indicating the maximum gradient magnitude in its neighborhood.
[0101] S323, dual threshold detection and edge connection, set a low threshold (such as 50) and a high threshold (such as 150), extract strong edges and potential edges from the edge image of the gesture based on the low threshold and the high threshold, use the strong edge to connect the potential edge to form a complete edge, and obtain the contour edge of the gesture area. Among them, the edge above the high threshold is a strong edge and the edge between the high and low values is a potential edge.
[0102] Target objects of dual threshold detection: Contours and edges of gestures: Edge information of gestures is extracted through edge detection. Dual threshold detection filters the edges in the image through high and low thresholds, retaining the most significant edges (such as fingers, palm contours, etc.). Noise suppression: The low threshold part helps to remove noise and maintain edge continuity.
[0103] Target objects of edge connection: Gesture outline: For example, the overall shape of a gesture may include the edges of fingers and palms. In gesture recognition, the detected edges need to be connected to obtain the complete gesture outline. Gesture details: such as the gaps between fingers, the boundaries of the palm, etc., can be expressed more finely through edge connection.
[0104] like Figure 2 As shown in the figure, the comparison of edge detection results of infrared gesture images is shown in the figure. Figure 2 The left image is the original infrared image, showing the simulated low-contrast gesture area with blurred boundaries and poor distinction from the background. Figure 2The right picture shows the Canny edge detection result, which successfully extracts the contour edge of the gesture area. The gesture area has clear boundaries and little voice interference, making it suitable for subsequent gesture recognition or area segmentation tasks.
[0105] S33, edge image filling, based on the edge of the obtained gesture area, the contour edge of the gesture area is filled to obtain a complete gesture area. The complete gesture area is restored to provide a complete target contour for subsequent segmentation and recognition tasks.
[0106] Based on the edge of the gesture area obtained through edge detection of the image, the drawcontours method of the cross-platform computer vision and machine learning software library OpenCV can be used to fill the contour edge of the gesture area. The detected contour is drawn on the image according to the edge of the gesture area through the drawcontours method of OpenCV, and then the contour of the gesture area is filled with a specific color. The various parts of the gesture are identified and distinguished by filling the contour of the gesture to obtain a complete gesture area.
[0107] like Figure 3 As shown in the figure, the edge image filling effect diagram, Figure 3 The left image is a Canny edge detection image showing the preliminary edge detection results. The edges are clear but there are gaps. Figure 3 The right picture shows the contour filling result. Based on the edge detection contour, the target area is filled to generate a clear gesture area.
[0108] S34, by image binarization segmentation, the gesture area (target area) in the infrared image is separated from the background to obtain a binary gesture area image. The gesture area image is separated from the background to provide input for subsequent recognition and analysis.
[0109] The gesture area image is binarized by thresholding, which separates the background from the foreground (i.e., the gesture part) in the image. The gesture area in the image is separated, and the background is usually black (0) and the foreground (gesture) is white (255). Through binarization, the outline of the gesture becomes clear, providing a clear foreground area for subsequent feature extraction and skeleton point recognition. It helps to reduce noise interference and make subsequent processing more accurate.
[0110] The image is segmented using an adaptive threshold segmentation method, which dynamically determines the threshold based on the local area of the image, and is suitable for scenes with uneven lighting. The adaptive threshold segmentation method defines the edge of the image containing solid matter in contrast to the background, providing a binary output from a grayscale image. This method involves the selection of an appropriate threshold T to convert a grayscale value image into a binary image. The advantage of obtaining a binary image is that it simplifies the complexity of the data and the process of recognition and classification. To describe its mathematics, a threshold T is defined using image pixel labels, with label 1 corresponding to the object and 0 corresponding to the background. The infrared image can be defined as a function f(x, y), and the threshold image g(x, y) can be defined as follows:
[0111] 1, f(x, y)>T;
[0112] 0, f(x, y) <T;
[0113] Calculate the threshold T in the local area. The formula of threshold T is:
[0114]
[0115] Where N is the number of pixels in the local window, R is the local window, and C is a constant, which is usually used to fine-tune the threshold to help distinguish the background from the foreground. A common choice is a small positive number (such as 2 or 3) to reduce the impact of the background when calculating the threshold. (x, y) is the center coordinate of the local window, which represents the reference point when calculating the threshold in the image. T(x, y) is the threshold calculated at the point (x, y) and is used for image binarization. This threshold can change dynamically based on the average value of the pixel values in the window. R(x, y) represents the local window area centered on (x, y), usually a wxw area, and all pixels contained in the window are represented by (i, j). I(i, j) represents the grayscale value or intensity value of the pixel point (i, j) in the image, which represents the brightness value of each pixel in the image (usually a value between 0 and 255 in a grayscale image).
[0116] N represents the number of pixels in the local window. For example, if a 3x 3 window is selected, then N = 9. R(x, y) is the window area centered at (x, y), which includes pixels within a certain range around the center point. When calculating the threshold, select a window of appropriate size.
[0117] S4. Extract hand information from the binary gesture area image, identify and mark the hand skeleton points based on the hand information, and obtain a connection diagram of the hand skeleton point motion trajectories.
[0118] S41. Train a YOLOv5 target detection model, extract hand information from the gesture area image through the trained YOLOv5 target detection model, and crop the hand area from the gesture area image according to the hand information; wherein the hand information includes the hand's bounding box coordinates, hand size, and hand category.
[0119] The YOLOv5 target detection model is a lightweight and fast target detection model that is very suitable for hand detection tasks. The hand information is extracted from the gesture area image through the trained YOLOv5 target detection model, and the position of the palm and the distribution of the fingers are determined through the hand information to locate the skeleton points. The hand information extraction provides the basis for the subsequent skeleton point detection, which helps us find the shape of the gesture and further determine the position of the palm and the distribution of the fingers through the contour.
[0120] S411. Collect image data sets and annotate the image data sets. The annotated information includes: a bounding box of the hand area in each image, the size of the hand, and the category of the hand, to obtain an annotated data set.
[0121] Image datasets can use public hand datasets (such as the Hand-Tracking dataset and the 300W-LP dataset) or collect data yourself. Label the image dataset data and label the hand area in each image as a rectangular bounding box, that is, give the position of the hand. In order to improve the generalization ability of the model, the data can be enhanced, such as rotation, flipping, zooming, and cropping. Collect image datasets and high-quality labeled data to provide the model with sufficient training samples. The model is trained by labeling the position of the hand in the image so that the model can learn to recognize and locate the hand.
[0122] S412: Model training: Use the labeled data set to train the YOLOv5 target detection model to obtain a trained YOLOv5 target detection model.
[0123] The model training phase is based on the YOLOv5 framework. During the training process, YOLOv5 automatically optimizes the model parameters so that it can learn how to recognize the hands in the image. By continuously optimizing the loss function, the model can improve its accuracy and obtain a trained model. The purpose of training the model is to enable the YOLOv5 object detection model to accurately detect the hand area, including the bounding box and category information.
[0124] S413, extracting hand information from the binarized gesture area image of the trained YOLOv5 target detection model and separating the hand area.
[0125] The trained YOLOv5 object detection model is used to infer the input image to obtain a prediction result, where the prediction result includes the bounding box coordinates, hand size, and hand category of the hand detected in the image. The bounding box coordinates indicate the position of the hand in the image, and the hand size is represented by the width and height of the bounding box. The hand category is such as left or right hand, or gesture classification (such as fist, heart shape, etc. in gesture recognition). The hand region is separated from the original image and subsequently processed, such as gesture recognition and tracking. By extracting the hand region, the interference part in the image can be removed, and the focus can be placed on the detection and analysis of the hand. The trained model is used to perform hand detection on the new image to obtain a prediction result, and the hand region is cropped from the original image according to the bounding box coordinates, the hand region is separated from the original image, and the hand information such as hand size and hand category is extracted from the hand region. By extracting the hand region, the interference part in the image can be removed, and the focus can be placed on the detection and analysis of the hand.
[0126] The YOLOv5 target detection model can be optimized, and the detection results can be optimized through steps such as non-maximum suppression and confidence screening to ensure high-precision hand detection in the end. Although YOLOv5 itself has been optimized to a certain extent, a threshold can be set through non-maximum suppression (NMS) to only retain detection results with a confidence greater than a certain value, such as selecting results with a confidence greater than 0.5. Redundant detection boxes can be removed and the boxes with the highest confidence can be retained to reduce the problem of repeated detections. Duplicate boxes can be removed and the boxes that are most likely to represent hands can be retained, thereby improving the accuracy and efficiency of the overall detection and improving the quality of the detection results.
[0127] S42. Building a skeleton point motion tracking model based on the HRNet model, training the skeleton point motion tracking model, and obtaining a trained skeleton point motion tracking model.
[0128] S421. Build a skeleton point running tracking model based on the HRNet model, improve the network structure of the HRNet model, introduce a deep convolutional attention mechanism when fusing features at each layer of the HRNet model, and perform sequential modeling when adding a Transformer module after the output of the HRNet model to obtain the skeleton point running tracking model.
[0129] HRNet (High-Resolution Network) is an efficient skeleton point detection model. It retains high-resolution features through multi-scale feature fusion while enhancing low-resolution features. It is suitable for skeleton point detection and tracking tasks. The skeleton point sequence recognition algorithm of the HRNet model includes the following steps:
[0130] Skeleton point detection, extracting key points of the human body (such as joints) from a single frame image. Time series modeling, combining the information of skeleton points in multiple frames to identify motion patterns. Classification / prediction, performing action classification or behavior prediction based on the features of skeleton point trajectories.
[0131] Specifically, a skeleton point running tracking model is built based on the HRNet model, and the network structure of the HRNet model is improved, including:
[0132] When fusing features at each layer of the HRNet model, a deep convolutional attention mechanism is introduced to enhance the features of specific areas, so that the HRNet model can focus on the areas where joints or important bone points are located when detecting bone points, and extract human key points from a single frame image.
[0133] Introducing the attention mechanism, enhancing the response of feature extraction to skeleton points, enhancing skeleton point detection, and enhancing the feature fusion of the HRNet model. The original HRNet relies on multi-scale feature fusion, but in hand skeleton point detection, local details are very important, so it is necessary to further strengthen the focus on high-resolution features. In the original version, HRNet maintains fine-grained information through high-resolution feature maps, which is particularly effective for skeleton point detection tasks. The core idea of HRNet is to process multi-resolution (different scales) feature maps in parallel and maintain high-resolution information flow. Unlike traditional bottom-up networks, HRNet avoids information loss by continuously fusing high-resolution and low-resolution feature maps. High-resolution feature maps, high-resolution features (fine information) extracted by the model from the input image, are used to retain fine-grained spatial information. Low-resolution feature maps, low-resolution features (global information) obtained by downsampling, are mainly used to capture global context. Feature fusion, HRNet enhances the interaction of information at different scales by continuously fusing high-resolution and low-resolution features. The attention mechanism is introduced, including the use of SE (Squeeze-and-Excitation), CBAM (Convolutional Block Attention Module) and other attention mechanisms to enhance the features of specific areas, so that the model can focus more on the details of important bone points such as joints and fingertips.
[0134] A Transformer module is added after the output of the HRNet model for timing modeling. The Transformer module is used to process the timing information, track the timing of the skeleton points, and capture the motion trajectory of the skeleton points between consecutive frames.
[0135] Skeleton point detection is not only the recognition of a single frame, but also involves the temporal tracking of skeleton points. Therefore, the recognition of skeleton points in a single frame needs to be combined with temporal information to better capture the motion trajectory of skeleton points. The Transformer module can process temporal data and capture the motion pattern of skeleton points between consecutive frames. By introducing the temporal modeling module, the network not only considers the features of the current frame, but also improves the accuracy of skeleton point recognition through contextual information.
[0136] The Transformer module is a deep learning model architecture in the fields of natural language processing (NLP) and computer vision (CV). By adding the Transformer module for time series modeling, the temporal dependency of skeleton points can be captured. In the task of skeleton point detection, we usually not only focus on the position of skeleton points in a single frame, but also hope to capture the motion changes of skeleton points in the time series. For example, the motion trajectory of finger joints, the dynamic changes of body joints, etc. Through the Transformer module, the model can use the information between the previous and next frames to predict the position of skeleton points in the current frame, thereby improving the accuracy of overall skeleton point positioning and motion coherence. In many application scenarios, skeleton point detection may face some challenges, such as occlusion, illumination changes, posture changes, etc. Traditional skeleton point detection methods can usually only make predictions based on single frame information and are easily affected by these factors. By introducing the Transformer module for time series modeling, the model can make smooth transitions between consecutive frames, overcome the limitations of single frame images, and improve the robustness of skeleton point detection in complex scenes. For some tasks involving action recognition and gesture recognition, it is necessary not only to identify the position of skeleton points in the current frame, but also to model and predict the motion trajectory of skeleton points. The Transformer module can effectively learn the temporal changes of skeleton points and predict the future positions of skeleton points. It can even predict reasonable positions of skeleton points when there is partial occlusion or information loss.
[0137] In time series tasks such as skeleton point sequence modeling, the Transformer module can extract global dependency information of longer sequences. Especially when processing large-scale data and long time series, the core idea of Transformer is the self-attention mechanism, which makes the output of each time step not only related to its own input, but also interact with the input of all other time steps. The self-attention mechanism is the core of Transformer. It determines which other parts each part should pay attention to by calculating the dependencies between different parts in the input sequence. For the input of each position, the attention mechanism calculates its attention to other positions in the sequence (i.e., weights) and adjusts the current representation based on these weights.
[0138] Self-attention mechanism formula:
[0139]
[0140] Where Q is the query, which indicates the current input position. K is the key, which indicates other information related to the current input position. V is the value, which indicates the actual information content. qi is the dimension of the key vector, which is usually used to scale the attention score. The attention mechanism calculates how each position is related to other positions in the sequence and generates a weighted sum
[0141] Through the self-attention mechanism, Transformer can directly calculate the dependency between any two positions, which is not affected by the length of the sequence. This allows Transformer to handle global dependencies in long sequences, and is particularly suitable for processing long-term skeletal point motion trajectory data. Each layer of Transformer can be calculated in parallel because the self-attention mechanism calculates the relationship between all positions in the sequence, regardless of the time order. This allows Transformer to be trained efficiently, and is particularly suitable for large-scale data and high-frame-rate video data.
[0142] Transformer can capture the long-distance dependencies between skeleton points and understand the changes of the entire skeleton structure in the time dimension, rather than being limited to the relationship between local skeleton points. By combining the Transformer module with the HRNetHRNet model, the HRNet model extracts the positions of the skeleton points from each frame of the image and performs local spatial feature extraction. The Transformer module can further process the temporal evolution of these skeleton points, perform time series modeling on the skeleton point sequence through the self-attention mechanism and the multi-head attention mechanism, capture the global dependency of the skeleton points, capture the time series information of the skeleton points between different frames, and improve the prediction ability of the skeleton point sequence.
[0143] S422. Collect and annotate the hand or human body bone point data as training data, combine the key point detection loss and the sequence prediction loss to obtain the joint loss, and use the training data to train the bone point running tracking model based on the joint loss.
[0144] Collect a large amount of image or video data with skeleton annotations as training data to train the skeleton point tracking model. These training data need to mark the positions of the skeleton points (such as hand joints, fingertips, etc.). Public datasets such as COCO and MPII can be used, or self-collected data can be used for annotation. The generalization ability of the model can be improved by enhancing the training data (for example, by rotating, flipping, adding noise, etc.) and solving the problem of class imbalance. Especially in the case of scarce data, poor lighting or severe occlusion, the optimization of the training process can effectively improve the robustness of the model. Adaptive learning rate adjustment, through the learning rate scheduler (such as cosine annealing, periodic learning rate adjustment, etc.) to improve the training efficiency of the model and prevent overfitting. Use the training data to train the skeleton point tracking model so that the skeleton point tracking model has better performance in skeleton point detection. The goal of the training model is to predict the position of the skeleton points in each frame of the image through the feedforward process (forward propagation). In addition, the temporal changes of the skeleton points need to be processed by the Transformer module to generate the motion trajectory of consecutive frames. The trained skeleton point motion tracking model is used for skeleton point detection and time series recognition, combined with the Transformer module to capture the skeleton point motion trajectory generation: according to the skeleton point position of continuous frames, skeleton point recognition and tracking can be performed to generate the motion trajectory of the hand or human body, providing support for gesture recognition, action recognition and other tasks.
[0145] The key point detection loss and the sequence prediction loss are combined to obtain the joint loss, and the skeleton point tracking model is trained using the training data based on the joint loss.
[0146] The application of joint loss in deep learning refers to combining multiple loss functions for optimization, which is usually used in multi-task learning scenarios. In skeleton point detection and skeleton point sequence prediction, joint loss combines key point detection loss (such as regression error of key point position) and sequence prediction loss (such as action recognition, prediction error of skeleton point motion trajectory) for optimization, so that the model can optimize multiple tasks at the same time and improve its performance.
[0147] Specifically, the key point detection loss and sequence prediction loss are weighted and summed to obtain the joint loss. Among them, the key point detection loss is used to optimize the position prediction of each bone point (key point) in the model, that is, to predict the (x, y) coordinates of the bone point (such as the finger joint fingertip, etc.) by regression. The formula of the key point detection loss is as follows:
[0148]
[0149] Among them, N is the total number of skeleton points, which include important points such as joints and fingertips of the human body or hand. is the coordinate of the bone point predicted by the model, x i ,y i is the real bone point coordinate.
[0150] Sequence prediction loss is used to optimize the timing tracking and motion prediction of skeleton points. This usually involves modeling the motion of skeleton points in consecutive frames. It can be based on the prediction of future actions or postures of timing models such as LSTM, GRU or Transformer. The formula for sequence prediction loss is:
[0151]
[0152] Where T is the number of frames in the sequence, indicating the length of the time sequence, usually the number of video frames or the length of an action sequence; is the predicted bone point position of the model at time t, x t ,y t is the actual bone point position.
[0153] The joint loss function is the weighted sum of the key point detection loss and the sequence prediction loss. The joint loss function is expressed as:
[0154]
[0155] in, is the keypoint detection loss, is the sequence prediction loss, and λ is a hyperparameter used to adjust the weight between the key point detection loss and the sequence prediction loss. By adjusting the relative importance of the two losses, a balance can be found between the key point detection accuracy and the time series prediction accuracy. By introducing the adjustment of hyperparameters, the optimization direction of key point detection and sequence prediction can be controlled. If λ is large, the model will pay more attention to sequence prediction, which may perform better in time coherence, but may affect the accuracy of single-frame skeleton points. If the hyperparameter is small, the model will focus more on the accuracy of key point detection. This flexible adjustment can adapt to different application scenarios, such as real-time gesture recognition or dynamic human posture estimation.
[0156] Through the joint loss, the model not only optimizes the regression of the skeleton point position and the prediction of the skeleton point sequence during training, but also enhances information sharing between tasks. The key point detection loss ensures that the model can accurately locate the skeleton points, and the sequence prediction loss ensures that the model can accurately predict the position changes of the skeleton points in consecutive frames and maintain the consistency of the motion trajectory. Especially when the skeleton points are occluded or move quickly, the timing loss can help the model predict and maintain the correct skeleton point position. In this way, the network can learn how to extract the most critical features in key point detection and learn how to maintain the consistency of the skeleton points in timing prediction, thereby improving the prediction of the motion trajectory of the skeleton points.
[0157] Based on each frame of image obtained from the real-time video stream of the thermal imaging device, the position of the skeleton points of each frame is inferred in real time through the trained model, and the skeleton points are tracked in combination with the information of consecutive frames to maintain the identity consistency of each skeleton point and generate a smooth motion trajectory. According to the position of the skeleton points in consecutive frames, the motion trajectory of the skeleton points is drawn, and key information points such as finger joints and wrists are connected to form a visual motion trajectory. The final output result is a path containing the position of the hand or human skeleton points in each frame and the connection path between these positions. This trajectory diagram can be used for tasks such as action recognition and gesture analysis.
[0158] Based on HRNet, we built a skeleton point operation tracking model and adapted it to thermal imagers for gesture recognition. We combined HRNet's high-precision ability in skeleton point detection with the characteristics of thermal imaging data, and innovatively added attention mechanisms and time series modeling enhancement modules. We combined HRNet with the time series model to achieve thermal imaging gesture recognition. The thermal imaging gesture recognition based on HRNet has higher-precision skeleton point detection (optimized for thermal imaging data) and is suitable for tracking and recognizing complex gestures through dynamic time series modeling. Multimodal data fusion significantly improves robustness in complex scenarios.
[0159] S43, using the trained skeleton point running tracking model to identify and mark the skeleton points of the hand area, tracking and analyzing the changes of the skeleton points of the hand and predicting the motion trajectory, and drawing a motion trajectory diagram of the skeleton points of the hand.
[0160] S431, detecting skeleton points in the hand area through the trained skeleton point running tracking model, predicting and marking the positions of skeleton points in the gesture area in each frame of the hand area, and generating skeleton point coordinates.
[0161] Specifically, the skeleton points of the hand refer to the key points in the hand structure, such as the wrist, various joints, fingertips, and the outline of the palm, forming the skeleton structure of the gesture. The wrist is usually the base point of the hand, and each joint and fingertip of the finger is a key point. The skeleton points of the finger are usually composed of multiple joint points. Each finger can have multiple joints (such as the bottom, middle, and fingertips). The outline of the palm and the wrist also need to be marked as key points. The tracking model running through the skeleton points automatically extracts and identifies these skeleton points according to the contour shape of the hand area map and the position of the fingertips, identifies and marks these skeleton points, and generates the coordinates of the skeleton points. The coordinates of the skeleton points refer to the position of each skeleton point (for example, (x, y) coordinates). The recognition and annotation of skeleton points help describe the specific form of the gesture and can provide accurate gesture action data.
[0162] S432, connecting the marked hand skeleton points according to the skeleton point coordinates to form a complete hand skeleton image, tracking and analyzing the changes of the hand skeleton points in multiple hand skeleton image frames, and obtaining a motion trajectory diagram of the hand skeleton points in continuous frames.
[0163] According to the bone structure of the fingers and palms, each joint point and fingertip of the fingers are connected in sequence, and the wrist and finger base are also connected to form a complete hand skeleton image. Connecting the bone points to form a complete hand skeleton can more clearly depict the overall structure of the gesture and provide a basis for subsequent analysis (such as gesture recognition and motion trajectory analysis). According to the changes of the hand bone points in different time frames, the same bone points in adjacent frames are connected with lines, and the motion trajectories of all bone points are connected together to form a motion trajectory diagram of the gesture bone points. The hand posture is inferred based on the positional relationship of the bone points. Hand postures include clenching a fist, opening the hand, bending fingers, etc. The bone point motion trajectory connection diagram is a graph structure generated by modeling the spatial and temporal characteristics of the bone point sequence, which provides a dynamic feature basis for further gesture classification and recognition. Such as Figure 4 The figure shows the skeleton point marking diagram effect diagram in the embodiment of the present invention.
[0164] The final output of the skeletal point tracking model is the motion trajectory of the hand skeletal points in the time dimension, which can show the changes and motion patterns of the skeletal points between consecutive frames. The skeletal point sequence means that in each frame, the model predicts the continuous positions of multiple skeletal points of the hand in the time dimension (such as joints, wrists, fingertips, etc.). For each frame, the model outputs an array containing the coordinates of the skeletal points (usually (x, y) coordinates, and may also include z coordinates and other key information). These coordinates change over time (i.e., the output of different frames) to form a complete skeletal point motion trajectory. The positions of the same skeletal points in consecutive frames are connected by lines to obtain a skeletal point motion trajectory connection diagram, forming a visual motion trajectory. The skeletal point motion trajectory connection diagram helps to show the changes of skeletal points in space and time. The motion trajectory of the hand skeletal points can be drawn by graphical tools or image processing algorithms, connecting the key points of the finger joints to form a dynamic trajectory diagram of the skeleton.
[0165] S5. Perform gesture recognition on the motion trajectory of the hand skeleton points, and output a sequence of gesture categories and the corresponding Chinese meanings.
[0166] S51. Perform gesture recognition on the motion trajectory graph of the hand skeleton points through the graph convolutional network GCN, and output a sequence of gesture categories.
[0167] According to the connection graph of the skeletal point motion trajectory, the skeletal points are modeled as a graph through the graph convolutional network GCN (Graph Convolutional Network), and the human skeletal points are regarded as nodes of the graph. The joint connections between the skeletal points are defined as edges, and the coordinates (x, yz) of each skeletal point or its dynamic characteristics in the time series (speed, acceleration, etc.) are used as node features. The spatial relationship features of the skeletal points are extracted through graph convolution operations, and the dynamic changes of gestures are captured by combining the information of the time dimension. The features in the spatial and time dimensions are extracted, and finally the gesture category is output. The Chinese meaning corresponding to the gesture is obtained based on the gesture category.
[0168] In this embodiment of the graph convolutional network (GCN), the final processed object can be understood as a sequence of bone points, but the core input representation is a connection graph of the motion trajectory of the bone points. Specifically, the connection graph of the motion trajectory of the bone points is a graph structure generated by modeling the spatial and temporal characteristics of the bone point sequence for use by GCN. Input the connection graph of the motion trajectory of the bone point to obtain a sequence of bone points, expressed as (T, N, C). T is the number of time frames, indicating how many frames there are in the sequence. N is the number of bone points (such as joint points, fingertip points, etc.). C is the characteristic dimension of the bone point (such as coordinates (,y,z), speed, acceleration, etc.). Each frame in the bone point sequence contains the coordinates and features of each bone point, and the entire sequence forms a trajectory in the time dimension. The bone points (such as joint points, fingertip points, etc.) are used as nodes of the graph. The features of each node are composed of the (,y,z) coordinates of the bone point and other dynamic characteristics (such as speed, acceleration). Constructing edges: In the spatial dimension, edges represent the anatomical connection between bone points (such as the connection between the joint point of the finger and the fingertip). In the temporal dimension, edges represent the dynamic connection between the same bone point in consecutive frames. This construction method converts the sequence of bone point motion trajectories into a spatio-temporal graph, i.e., a "bone point motion trajectory connection graph."
[0169] The object processed by the graph convolutional network (GCN) is the connection graph of the motion trajectory of the skeleton points. GCN extracts features in the spatial dimension and uses the graph structure to extract the spatial relationship between the skeleton points, such as the relative position between the finger joints, the connection between the palm and the wrist, etc. The feature extraction in the time dimension captures the dynamic changes of the skeleton points in time (such as speed and acceleration), and extracts the temporal characteristics of the motion trajectory through the connection edges of the time series.
[0170] The graph convolutional network (GCN) processes the nodes and edges of the graph. The nodes of the graph represent the skeleton points. The features of each node are the coordinates and their dynamic characteristics. The edges of the graph connect the joint relationship between the skeleton points in space and connect the same skeleton points in consecutive frames in time. It realizes the recognition of gesture categories and finally outputs the gesture categories.
[0171] The basic formula of graph convolutional network GCN is expressed as:
[0172] H (l+1) =σ(AH (l) W (l) )
[0173] Among them, H (l) is the feature representation of the l-th layer node, A is the adjacency matrix of the graph, W (l) A learnable weight matrix. σ Non-linear activation function (such as ReLU).
[0174] S52. Input the gesture category sequence output by the graph convolutional network GCN into a pre-trained language model, convert the gesture category sequence into Chinese description through the pre-trained language model, and obtain the Chinese meaning corresponding to the gesture category.
[0175] After completing gesture recognition through the graph convolutional network GCN and obtaining the gesture category, the recognized gesture category (such as the classification result) is mapped to its corresponding Chinese meaning. The gesture category sequence output by the graph convolutional network GCN is represented as [2, 4, 1], where 2 represents "clenching a fist", 4 represents "opening the hand", and 1 represents "finger gesture number 1". Each number represents the gesture recognition result of the model for a certain time frame. The meanings corresponding to these numbers require a pre-defined mapping relationship from a category to a Chinese description to associate the gesture category number with its Chinese meaning. For example: the Chinese meaning of the gesture category number 1 is: gesture number 1, the Chinese meaning of 2 is: clenching a fist, the Chinese meaning of 3 is: pointing in a certain direction, the Chinese meaning of 4 is: opening the hand, and the Chinese meaning of 5 is: finger OK gesture.
[0176] Based on the output gesture category sequence [2, 4, 1], the gesture category sequence can be converted into Chinese descriptions through pre-trained language models (such as BERT, GPT, etc.) to generate corresponding Chinese descriptions. First, convert the category sequence [2, 4, 1] into Chinese meanings; then query the mapping table: 2-> clench, 4-> open hand, 1-> gesture number 1. Finally, convert it into Chinese description: clench fist, open hand, gesture number 1.
[0177] In summary, the present invention provides a method for capturing and recognizing gesture optical motion based on thermal imaging, wherein multiple bone point patches are attached to multiple finger joints, a thermal imaging device locates the finger bone points through multiple bone point patches, an infrared thermal imaging gesture recognition device is used to collect infrared images of gestures, and grayscale expansion, gesture edge detection, edge image filling and image binarization segmentation are performed on the infrared images. A YOLOv5 target detection model is trained, hand information is extracted from the gesture area image through the trained YOLOv5 target detection model, and the hand area is cropped from the gesture area image according to the hand information. A bone point operation tracking model is constructed based on the HRNet model, the bone point operation tracking model is trained, and the trained bone point operation tracking model is obtained. The bone point operation tracking model is used to identify and mark the hand area, the changes of the hand bone points are tracked and analyzed, and the motion trajectory is predicted, and a motion trajectory diagram of the hand bone points is drawn. Finally, the bone point motion trajectory connection diagram is used to perform gesture recognition to obtain the Chinese meaning of the gesture, and the gesture action of the person can be captured with high accuracy in a low-visibility environment, and the gesture action is gesture recognized and the Chinese meaning of the gesture action is obtained. It can solve the problem of high-precision hand motion recognition faced by traditional visible light motion capture technology in low visibility and weak lighting environments.
[0178] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. A method for capturing and recognizing gesture optical motion based on thermal imaging, characterized in that: The following steps are included S1. Collecting an infrared image of a gesture using an infrared thermal imaging gesture recognition device, where the infrared thermal imaging gesture recognition device includes a thermal imaging device and a plurality of bone point patches; Multiple bone point patches are attached to multiple finger joints, and the thermal imaging device locates the finger bone points through the multiple bone point patches; S2, preprocessing the collected infrared image of the gesture to obtain a preprocessed infrared image; S3, performing grayscale expansion, gesture edge detection, edge image filling and image binarization segmentation on the pre-processed infrared image to obtain a binary gesture area image; S4, extracting hand information from the binary gesture area image, identifying and marking the hand skeleton points based on the hand information, and obtaining a connection diagram of the hand skeleton point motion trajectories; S41, training a YOLOv5 target detection model, extracting hand information from the gesture area image through the trained YOLOv5 target detection model, and cropping a hand area from the gesture area image according to the hand information; S42, building a skeleton point running tracking model based on the HRNet model, training the skeleton point running tracking model, and obtaining a trained skeleton point running tracking model; S43, using the trained skeleton point running tracking model to identify and mark the skeleton points of the hand area, tracking and analyzing the changes of the skeleton points of the hand and predicting the motion trajectory, and drawing a motion trajectory diagram of the skeleton points of the hand; S5. Perform gesture recognition on the motion trajectory of the hand skeleton points, and output a sequence of gesture categories and the corresponding Chinese meanings.
2. The method for capturing and recognizing gesture optical motion based on thermal imaging according to claim 1, characterized in that: The bone point patch is a mixture of thermoplastic polyurethane and boron nitride, with the boron nitride content being 10%-30%.
3. The method for capturing and recognizing gesture optical motion based on thermal imaging according to claim 1, characterized in that: The step S3 comprises: S31, performing grayscale expansion on the infrared image to improve the contrast of the infrared image, and obtaining an infrared image after grayscale expansion; S32, performing gesture edge detection on the infrared image, and extracting the contour edge of the gesture area in the infrared image after grayscale expansion based on the Canny edge detection algorithm; S33, based on the obtained contour edge of the gesture area, performing area filling on the contour edge of the gesture area to obtain a complete gesture area; S34, performing image binarization segmentation on the gesture area to obtain a binary gesture area image.
4. The method for capturing and recognizing gesture optical motion based on thermal imaging according to claim 3, characterized in that: The grayscale expansion of the preprocessed infrared image includes: converting the infrared image into a grayscale image and removing the noise from the grayscale image; calculating the original grayscale value range of the grayscale image of the infrared image, setting a target grayscale range, transforming the grayscale value using a linear formula, and stretching the original grayscale value distribution of the grayscale image.
5. The method for capturing and recognizing gesture optical motion based on thermal imaging according to claim 3, characterized in that: The step of performing gesture edge detection on the infrared image and extracting the contour of the gesture in the infrared image after grayscale expansion based on the Canny edge detection algorithm includes: S321, performing Gaussian filtering on the infrared image after grayscale expansion through a Gaussian filter; S322, calculating the gradient amplitude image and angle image of the smoothed image by using the Sobel operator, applying non-maximum suppression to the gradient amplitude image, retaining the image point with the maximum local gradient amplitude image, removing non-edge pixels, and obtaining a gesture edge image; S323, setting a low threshold and a high threshold, extracting strong edges and potential edges from the gesture edge image based on the low threshold and the high threshold, connecting the potential edges with the strong edges to form a complete edge, and obtaining the edge of the gesture area.
6. The method for capturing and recognizing gesture optical motion based on thermal imaging according to claim 1, characterized in that: The step S41 comprises: Collect image data sets and annotate the image data sets. The annotated information includes the bounding box of the hand area in each image, the size of the hand, and the category of the hand, so as to obtain an annotated data set. Use the labeled data set to train the YOLOv5 target detection model to obtain a trained YOLOv5 target detection model; The hand information is extracted from the binarized gesture area image of the trained YOLOv5 target detection model to separate the hand area.
7. The method for capturing and recognizing gesture optical motion based on thermal imaging according to claim 6, characterized in that: The step S42 comprises: A skeleton point running tracking model is built based on the HRNet model. The network structure of the HRNet model is improved. A deep convolutional attention mechanism is introduced when fusing features at each layer of the HRNet model. Sequential modeling is performed when the Transformer module is added after the output of the HRNet model to obtain a skeleton point running tracking model. Collect and annotate the hand or human skeleton point data as training data, combine the key point detection loss and sequence prediction loss to get the joint loss, and use the training data to train the skeleton point tracking model based on the joint loss.
8. The method for capturing and recognizing gesture optical motion based on thermal imaging according to claim 7, characterized in that: The formula for the key point detection loss is as follows: Among them, N is the total number of skeleton points, which include important points such as joints and fingertips of the human body or hand. is the coordinate of the bone point predicted by the model, x i ,y i is the real bone point coordinate; The formula for the sequence prediction loss is: Where T is the number of frames in the sequence, indicating the length of the time sequence, usually the number of video frames or the length of an action sequence; is the predicted bone point position of the model at time t, x t ,y t is the actual bone point position; The combined loss is expressed as: in, is the keypoint detection loss, is the sequence prediction loss and λ is a hyperparameter.
9. The method for capturing and recognizing gesture optical motion based on thermal imaging according to claim 1, characterized in that: The step S43 comprises: The trained skeleton point tracking model is used to detect skeleton points in the hand area, and the positions of skeleton points in the gesture area are predicted and marked in the hand area in each frame to generate skeleton point coordinates. The marked hand skeleton points are connected according to the skeleton point coordinates to form a complete hand skeleton image, and the changes of the hand skeleton points in multiple hand skeleton image frames are tracked and analyzed to obtain the movement trajectory of the hand skeleton points in continuous frames.
10. The method for capturing and recognizing gesture optical motion based on thermal imaging according to claim 1, characterized in that: The step S5 comprises: The graph convolutional network (GCN) is used to perform gesture recognition on the motion trajectory of the hand skeleton points and output a sequence of gesture categories. The gesture category sequence output by the graph convolutional network GCN is input into the pre-trained language model, and the gesture category sequence is converted into Chinese description through the pre-trained language model to obtain the Chinese meaning corresponding to the gesture category.
Citation Information
Cited By
Industrial operation gesture recognition model based on deep learning
CN121053703A