Acupuncture manipulation recognition method based on global hand shape and local finger movement

By combining optical flow features and skeleton key point technology with interactive attention and random weight fusion, the accuracy problem of acupuncture technique recognition is solved, a scientific quantitative description of global and local motion features is achieved, and the recognition accuracy is improved.

CN117238033BActive Publication Date: 2025-09-19BEIJING UNIV OF CHEM TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311251798.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-26
Publication Date
2025-09-19
Estimated Expiration
2043-09-26

AI Technical Summary

Technical Problem

Existing technologies are unable to accurately and effectively identify complex and varied acupuncture techniques, and lack comprehensive and detailed information on the recognition and quantification of global hand shapes and local finger movements.

Method used

By combining the RGB image sequence with the color system to convert it into a visual optical flow sequence, global motion features are extracted, and local motion features are obtained using the coordinates of skeleton key points. Combined with interactive attention and random weight fusion technology, acupuncture manipulation recognition is achieved.

Benefits of technology

It improves the accuracy of acupuncture manipulation recognition, overcomes background and environmental interference, and scientifically and quantitatively describes the global and local movement characteristics of the manipulation, which has practical application significance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117238033B_ABST
    Figure CN117238033B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for identifying acupuncture manipulations based on global hand shapes and local finger movements, comprising: converting an image sequence of acupuncture manipulation videos into a visual optical flow sequence, and obtaining global motion feature information through feature extraction; obtaining a skeleton image and key point coordinates of the finger from the acupuncture manipulation videos, and extracting and fusing the motion representation parameters and skeleton image obtained based on the key point coordinates to obtain local motion feature information; performing multi-scale feature interactive fusion on the global motion feature information and the local motion feature information, and then performing random weight fusion on the two sets of interactively fused feature information to output a recognition result. The present invention extracts features and quantitatively describes the acupuncture manipulation process from both global and local perspectives, and then performs interactive fusion analysis on the global features and local features, so that the recognition result fully utilizes the global advantages and detailed descriptions of the two parts of data, and has a high accuracy rate in the recognition of acupuncture manipulations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of gesture recognition technology, and in particular to an acupuncture manipulation recognition method oriented to global hand shape and local finger movement. Background Art

[0002] Acupuncture is a key treatment approach in Traditional Chinese Medicine. It involves the physician inserting a needle into a patient's specific acupuncture points at a certain angle, depth, and frequency, stimulating the acupuncture points through lifting, inserting, and twisting to achieve the purpose of treatment. Different operating techniques and hand shapes also determine the final treatment effect, so the inheritance and learning of the techniques are very important.

[0003] Based on the research on the identification of acupuncture manipulation, many scholars have also conducted a series of research works, as shown below:

[0004] (1) Statistical analysis method: In traditional manipulation recognition research, the manipulation motion parameters and waveforms are usually obtained by the acquisition equipment, and statistical analysis is performed on different manipulation parameters to obtain the overall change pattern. In the existing technology, an acupuncture scale is designed, and a scoring link is set during the acupuncture process. The operation process is quantitatively scored step by step, and finally a descriptive statistical analysis is performed based on the scale. This can realize the evaluation of different manipulations to determine whether the manipulation operation is standardized and whether the actual treatment effect meets the requirements. However, the manipulation recognition method based on statistical laws lacks scientificity and cannot objectively evaluate the complex and changeable manipulation process.

[0005] (2) Shallow data-driven modeling method: With the development of technology, some scholars have used data mining methods to explore the changes in parameter information during acupuncture and summarize the regular information of different techniques. The existing technology uses data mining technology to mine regular information from a large number of acupuncture technique parameters collected, explore the association rules between technique parameters, and evaluate and analyze different techniques. However, data mining methods can only summarize general regular characteristics based on raw data and lack refined quantitative descriptions.

[0006] (3) Shallow machine learning methods: With the development of random computer technology, some scholars have begun to use advanced data processing technology to perform a series of preprocessing operations on the original data, and then use machine learning methods to study technique recognition. One existing solution starts from a clustering method based on machine learning, explores the internal laws of acupuncture techniques, and conducts fuzzy clustering analysis based on the acupuncture efficacy under different parameters; another existing solution summarizes the current research on acupuncture technique recognition based on data mining methods, and realizes the classification of four common techniques. However, most of the current recognition methods based on machine learning attempt to capture some manual features, and lack effective feature extraction methods for dealing with some complex action recognition problems.

[0007] (4) Deep learning methods: With the rise of deep learning, motion recognition methods based on deep learning have been widely used. Some scholars often use the sliding window method to extract the changing characteristics of the piezoelectric waveform of the hand, obtain the feature information defined by "effective peak", "independent window", "motion window", etc., model the acupuncture movement change process based on the feature information, and then input the features into the proposed ensemble learning classifier to achieve accurate recognition of the four techniques. In addition, some scholars use deep networks such as ResNet, MobileNet, InceptionNet, DenseNet and other networks to detect and recognize gestures. However, the current deep learning-based methods often describe image features at a single scale. When dealing with the complex motion change process of acupuncture manipulation, it is difficult to accurately describe the global contour information and local subtle changes of acupuncture manipulation, which are key parameters for manipulation recognition.

[0008] Therefore, it is necessary to explore a more comprehensive and scientific method for extracting and identifying acupuncture technique features. Summary of the Invention

[0009] (1) Technical issues to be resolved

[0010] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides an acupuncture manipulation recognition method for global hand shape and local finger movement, which solves the technical problem that there is no accurate, effective and comprehensive recognition and quantification solution for complex and changeable acupuncture manipulations.

[0011] (2) Technical solution

[0012] In order to achieve the above objectives, the main technical solutions adopted by the present invention include:

[0013] In a first aspect, an embodiment of the present invention provides a method for acupuncture manipulation recognition based on global hand shape and local finger movement, comprising:

[0014] By utilizing the pixel changes between different frames, the RGB image sequence of the acupuncture manipulation video is converted into a visual optical flow sequence in combination with the color system. The global motion feature information of the hand is then obtained by feature extraction of the visual optical flow sequence.

[0015] The skeleton image of the finger and the corresponding key point coordinates are obtained from the acupuncture manipulation video. The motion representation parameters including characteristic angle, characteristic distance and spatial feature information are calculated based on the key point coordinates. The local motion feature information of the hand is then obtained by feature extraction and feature fusion of the motion representation parameters and the skeleton image.

[0016] The global motion feature information and the local motion feature information are linearized separately. Based on the two sets of query vectors, key vectors, and value vectors, interactive attention operations, normalization processing, and weight calculation are performed to obtain two sets of interactively fused feature information.

[0017] Random parameters are introduced to perform random weight fusion on the two sets of interactively fused feature information, and finally the acupuncture manipulation recognition results are obtained.

[0018] Optionally, the RGB image sequence of the acupuncture manipulation video is converted into a visual optical flow sequence by combining the color system using pixel changes between different frames. The global motion feature information of the hand is then obtained by feature extraction of the visual optical flow sequence, including:

[0019] By using the pixel changes between different frames, the RGB image sequence of the acupuncture manipulation video is obtained through optical flow calculation to obtain an optical flow sequence;

[0020] The Munsell color system is used to convert the velocity vectors of the optical flow sequence in the u and v directions into the color changes in the horizontal and vertical directions, thereby obtaining a visual optical flow sequence.

[0021] The visualized optical flow sequence is input into a pre-set feature extraction network for feature extraction to obtain the global motion feature information of the hand.

[0022] Optionally, a skeleton image of the finger and the corresponding key point coordinates are obtained from the acquired acupuncture manipulation video, and motion representation parameters including characteristic angles, characteristic distances, and spatial feature information are calculated based on the key point coordinates. Subsequently, feature extraction and feature fusion of the motion representation parameters and the skeleton image are performed to obtain local motion feature information of the hand, including:

[0023] Using a pre-used human skeleton posture detection library, key points of the hand in the acupuncture manipulation video are marked to obtain skeleton images and key point coordinate information; the key points include the joints of the index finger and thumb;

[0024] The characteristic angle and characteristic distance of each frame image are calculated based on the key point coordinates, and then the difference operation is performed between each frame based on the key point coordinates to obtain the spatial change value between adjacent frame images. Finally, the motion change between all frames is accumulated to obtain the spatial feature information;

[0025] The obtained motion feature parameters including characteristic angle, characteristic distance and spatial feature information are input into a pre-set LSTM network for feature extraction to obtain the kinematic parameter feature vector of the acupuncture manipulation;

[0026] The obtained skeleton image data is input into a pre-set feature extraction network for feature extraction to obtain an image feature vector of the hand skeleton;

[0027] The kinematic parameter feature vector of the acupuncture manipulation and the image feature vector of the hand skeleton are fused to obtain the local motion feature information of the hand.

[0028] Optionally, the characteristic angle and characteristic distance of each frame image are calculated based on the key point coordinates, and then the difference operation is performed between each frame based on the key point coordinates to obtain the spatial change value between adjacent frame images. Finally, the motion change between all frames is accumulated to obtain the spatial feature information including:

[0029] The characteristic angle of each frame is obtained by solving the key point coordinates and geometric relationships. Then, the characteristic angles between adjacent frames are calculated to obtain the change angle between frames. All the change angles are accumulated to obtain the total characteristic angle.

[0030] The characteristic distance of each frame image is obtained based on the key point coordinates and the spatial position relationship. Then, the characteristic distances between adjacent frames are calculated to obtain the change distance between frames. All the change distances are accumulated to obtain the total characteristic distance.

[0031] The change differences of the horizontal and vertical coordinates of adjacent frames are calculated based on the key point coordinates, and all the change differences of the horizontal and vertical coordinates are accumulated to obtain the total spatial feature information.

[0032] Optionally,

[0033] Based on the key point coordinates and geometric relationships, the characteristic angle of each frame is calculated. Then, the characteristic angles between adjacent frames are calculated to obtain the change angle between frames. All the change angles are accumulated to obtain the total characteristic angle. The following angle formula and angle accumulation formula are used for implementation:

[0034]

[0035] Where n is the number of image frames, θ i is the characteristic angle of the i-th frame image, i=1,2,……n,x j and y j are the coordinates of the jth key point in the i-th frame image, j is a positive integer, and θ is the total characteristic angle of a set of action sequences;

[0036] The characteristic distance of each frame image is obtained based on the key point coordinates combined with the spatial position relationship. Then, the characteristic distances between adjacent frame images are calculated to obtain the change distance between frames. All the change distances are accumulated to obtain the total characteristic distance through the following distance formula and distance accumulation implementation:

[0037]

[0038] d=Σ n |(d i+1 -d i )|

[0039] Where, d i is the feature distance of the current frame image, and d is the total feature distance;

[0040] Based on the key point coordinates, the change difference of the horizontal and vertical coordinates of adjacent frames is calculated, and all the change differences of the horizontal and vertical coordinates are accumulated to obtain the total spatial feature information. The following coordinate difference accumulation formula is used for implementation:

[0041]

[0042] Where X is the total change difference of the horizontal coordinate, and Y is the total change difference of the vertical coordinate. Represents the coordinate information of the current frame, Represents the coordinate information of the next frame.

[0043] Optionally, linearization operations are performed on the global motion feature information and the local motion feature information respectively, and interactive attention operations, normalization processing, and weight calculation are performed based on the two sets of query vectors, key vectors, and value vectors obtained to obtain the two sets of interactively fused feature information including:

[0044] Perform linearization operations on the global motion feature information and the local motion feature information respectively to obtain two sets of corresponding query vectors, key vectors, and value vectors;

[0045] Performing a product operation on the query vector corresponding to the global motion feature information and the key vector corresponding to the local motion feature information to obtain a first degree of association;

[0046] Performing a product operation on the query vector corresponding to the local motion feature information and the key vector corresponding to the global motion feature information to obtain a second degree of association;

[0047] The obtained first correlation degree and second correlation degree are respectively normalized by the Softmax function so that the first correlation degree and the second correlation degree are always kept between 0 and 1;

[0048] A first weight value is obtained by performing a matrix product operation on the value vector corresponding to the global motion feature information and the first correlation degree, and a second weight value is obtained by performing a matrix product operation on the value vector corresponding to the local motion feature information and the second correlation degree.

[0049] Optionally, random parameters are introduced to perform random weight fusion on the two sets of interactively fused feature information, and the final acupuncture manipulation recognition results include:

[0050] Introduce a random parameter λ;

[0051] For the first weight value and the second weight value at the two scales, a first weight coefficient λ and a second weight coefficient 1-λ are assigned respectively;

[0052] Perform random weight processing on the first weight value and the second weight value according to the first weight coefficient λ and the second weight coefficient 1-λ respectively to obtain a total output result;

[0053] Among them, the total output result is obtained through the following weight formula:

[0054] Output=λ*W G +(1-λ)*W L

[0055] λ∈Beta(α,α)

[0056] Where W G is the first weight value, W L is the second weight value, λ∈Beta(α,α) means that λ is the number of random samples obeying the beta distribution, and the parameter α is an adjustable parameter.

[0057] In a second aspect, the present invention provides an acupuncture manipulation recognition system for global hand shape and local finger movement, comprising:

[0058] The global information acquisition module is used to convert the RGB image sequence of the acupuncture manipulation video into a visual optical flow sequence by combining the color system and using the pixel changes between different frames. The global motion feature information of the hand is then obtained by feature extraction from the visual optical flow sequence.

[0059] The local information acquisition module is used to obtain the skeleton image of the finger and the corresponding key point coordinates from the acquired acupuncture manipulation video. Based on the key point coordinates, motion representation parameters including characteristic angles, characteristic distances, and spatial feature information are calculated. The local motion feature information of the hand is then obtained by feature extraction and feature fusion of the motion representation parameters and the skeleton image.

[0060] The multi-scale interactive fusion module is used to linearize the global motion feature information and the local motion feature information respectively. Based on the two sets of query vectors, key vectors, and value vectors, interactive attention operations, normalization processing, and weight calculation are performed to obtain two sets of interactive fused feature information.

[0061] The random weight fusion module is used to introduce random parameters to perform random weight fusion on the two sets of interactively fused feature information, and finally obtain the acupuncture manipulation recognition result.

[0062] In a third aspect, the present invention provides an acupuncture manipulation recognition device for global hand shape and local finger movement, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the acupuncture manipulation recognition method for global hand shape and local finger movement as described above.

[0063] In a fourth aspect, the present invention provides a computer-readable medium having computer-executable instructions stored thereon, which, when executed by a processor, implement the acupuncture manipulation recognition method for global hand shape and local finger movement as described above.

[0064] (3) Beneficial effects

[0065] The beneficial effects of the present invention are as follows: considering that the acupuncture process has a small amplitude and a fast frequency of change, and that different doctors have individual differences in their techniques, the identification and quantification of techniques are difficult. The present invention starts from global features and local features. In terms of global features, in order to overcome the interference of background, environmental factors and other factors in traditional RGB image data, the global contour motion changes of the acupuncture hand shape are described by calculating optical flow features; in terms of local features, in order to scientifically and quantitatively describe the motion changes of the acupuncture manipulation, based on the coordinates of the skeleton key points, the kinematic feature parameters describing the manipulation operation process are obtained. The individual kinematic feature parameters are calculated through the coordinate information, and the kinematic feature parameters are combined with the skeleton image to jointly describe the local motion feature changes of the manipulation. Furthermore, in order to effectively perform multi-scale information interaction and fusion of global feature information and local feature information, the two parts of data features are randomly fused through operations such as attention calculation and random weight fusion.

[0066] The present invention extracts features and quantitatively describes the acupuncture manipulation process from both global and local perspectives, and then interactively fuses and analyzes the global and local features, so that the recognition results fully utilize the global advantages and detailed descriptions of the two parts of data, have a high accuracy in the recognition of acupuncture manipulation, and have practical application significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 A flowchart of a method for acupuncture manipulation recognition based on global hand shape and local finger movement, proposed in an embodiment of the present invention;

[0068] Figure 2 A schematic diagram of the optical flow graph calculation process of an acupuncture manipulation recognition method for global hand shape and local finger movement proposed in an embodiment of the present invention;

[0069] Figure 3This is a schematic diagram of a specific flow of step S1 of a method for acupuncture manipulation recognition based on global hand shape and local finger movement, proposed in an embodiment of the present invention;

[0070] Figure 4 (a), (b), and (c) are respectively the original RGB manipulation sequence process, the manipulation sequence process after Gaussian filtering, and the manipulation sequence process after optical flow processing of an acupuncture manipulation recognition method for global hand shape and local finger movement proposed in an embodiment of the present invention;

[0071] Figure 5 This is a schematic diagram of a specific flow of step S2 of a method for acupuncture manipulation recognition based on global hand shape and local finger movement, proposed in an embodiment of the present invention;

[0072] Figure 6 A schematic diagram of the coordinates of 21 hand joints in an acupuncture manipulation recognition method for global hand shape and local finger movement proposed in an embodiment of the present invention;

[0073] Figure 7 Schematic diagram of the index finger and thumb joints of an acupuncture manipulation recognition method based on global hand shape and local finger movement proposed in an embodiment of the present invention;

[0074] Figure 8 This is a schematic diagram of a specific flow of step S22 of a method for acupuncture manipulation recognition based on global hand shape and local finger movement, as proposed in an embodiment of the present invention;

[0075] Figure 9 A schematic diagram of the characteristic angle change process of an acupuncture manipulation recognition method for global hand shape and local finger movement proposed in an embodiment of the present invention;

[0076] Figure 10 This is a schematic diagram of a specific flow of step S3 of a method for acupuncture manipulation recognition based on global hand shape and local finger movement, proposed in an embodiment of the present invention;

[0077] Figure 11 This is a schematic diagram of a specific flow of step S41 of a method for acupuncture manipulation recognition based on global hand shape and local finger movement, as proposed in an embodiment of the present invention;

[0078] Figure 12 A schematic diagram of the multi-scale interactive fusion and random weight fusion process of an acupuncture manipulation recognition method for global hand shape and local finger movement proposed in an embodiment of the present invention;

[0079] Figure 13 Beta distribution probability curves for different α values ​​of an acupuncture manipulation recognition method for global hand shape and local finger movement proposed in an embodiment of the present invention;

[0080] Figure 14 A schematic structural diagram of a model of an acupuncture manipulation recognition method based on global hand shape and local finger movement proposed in an embodiment of the present invention;

[0081] Figure 15 This is a comparison diagram of confusion matrices under different models of an acupuncture manipulation recognition method for global hand shape and local finger movement proposed in an embodiment of the present invention. DETAILED DESCRIPTION

[0082] In order to better explain the present invention and facilitate understanding, the present invention is described in detail below through specific implementation methods in conjunction with the accompanying drawings.

[0083] like Figure 1 As shown, an embodiment of the present invention proposes an acupuncture manipulation recognition method for global hand shape and local finger movement, including: first, utilizing the pixel changes between different frames, the RGB image sequence of the acquired acupuncture manipulation video is combined with the color system to be converted into a visual optical flow sequence, and then the global motion feature information of the hand is obtained by feature extraction of the visual optical flow sequence; secondly, the skeleton image of the finger and the corresponding key point coordinates are obtained from the acquired acupuncture manipulation video, and the motion representation parameters including feature angle, feature distance and spatial feature information are solved based on the key point coordinates, and then the local motion feature information of the hand is obtained by feature extraction and feature fusion of the motion representation parameters and the skeleton image; then, linearization operations are performed on the global motion feature information and the local motion feature information respectively, and interactive attention operations, normalization processing and weight calculation are performed based on the two sets of query vectors, key vectors and value vectors obtained to obtain two sets of interactively fused feature information; finally, random parameters are introduced to perform random weight fusion on the two sets of interactively fused feature information, and finally the acupuncture manipulation recognition result is obtained.

[0084] Considering that the acupuncture process has a small amplitude and a fast frequency of change, and that different doctors have personalized differences in their techniques, the identification and quantification of techniques are difficult. This paper starts from global features and local features. In terms of global features, in order to overcome the interference of background, environmental factors and other factors in traditional RGB image data, the global contour motion changes of the acupuncture hand shape are described by calculating optical flow features; in terms of local features, in order to scientifically and quantitatively describe the motion changes of acupuncture manipulation, based on the coordinates of the skeleton key points, the kinematic feature parameters describing the manipulation operation process are obtained. The various kinematic feature parameters are calculated through the coordinate information, and the kinematic feature parameters are combined with the skeleton image to jointly describe the local motion feature changes of the manipulation. Furthermore, in order to effectively perform multi-scale information interaction and fusion of global feature information and local feature information, the two parts of data features are randomly fused through operations such as attention calculation and random weight fusion.

[0085] The present invention extracts features and quantitatively describes the acupuncture manipulation process from both global and local perspectives, and then interactively fuses and analyzes the global and local features, so that the recognition results fully utilize the global advantages and detailed descriptions of the two parts of data, have a high accuracy in the recognition of acupuncture manipulation, and have practical application significance.

[0086] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to enable a clearer and more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0087] Specifically, the present invention provides a method for acupuncture manipulation recognition based on global hand shape and local finger movement, which includes:

[0088] S1. Using the pixel changes between different frames, the RGB image sequence of the acupuncture manipulation video is converted into a visual optical flow sequence by combining the color system. Then, the global motion feature information of the hand is obtained by feature extraction of the visual optical flow sequence. In order to better extract the global motion features and subtle motion change patterns of the hand, reference Figure 2 , the RGB video image sequence is converted into an optical flow image after preprocessing, and the optical flow sequence is used as the input of the network for feature extraction.

[0089] Further, if Figure 3 As shown, step S1 includes:

[0090] S11. Using pixel changes between different frames, the RGB image sequence of the acquired acupuncture manipulation video is subjected to optical flow calculation to obtain an optical flow sequence.

[0091] S12. The Munsell color system is used to convert the velocity vectors of the optical flow sequence in the u and v directions into color changes in the horizontal and vertical directions, thereby obtaining a visualized optical flow sequence.

[0092] S13. Input the visualized optical flow sequence into a pre-set feature extraction network for feature extraction to obtain global motion feature information of the hand.

[0093] In step S1, the specific optical flow calculation process is:

[0094] First, in a video clip, the change of the set time will not cause a drastic change in the target position, that is, the displacement between adjacent frames is small, so the constraint equation can be obtained:

[0095] I(x,y,t)=I(x+dx,y+dy,t+dt) (1)

[0096] In formula (1), I(x, y, t) represents the light intensity of a pixel in the first frame. It takes dt to move the distance (dx, dy) to the next frame. Taylor expansion of the right side of the equation yields:

[0097]

[0098] In formula (2), ε represents a second-order infinitesimal. Dividing both sides of (1) by dt, we can obtain:

[0099]

[0100] In formula (3), Represent the partial derivatives of the grayscale of the pixel in the image along the X, Y, and T directions respectively, and let u and v be the velocity vectors of the optical flow along the X and Y axes respectively, recorded as:

[0101]

[0102] Then formula (3) can be written as:

[0103] I x u+I y v+I t =0 (5)

[0104] Different optical flow calculation methods are obtained from formula (5). For small changes in movements such as acupuncture manipulation, in order to more effectively extract global motion contour features, the present invention uses the proposed Gunnar-Farneback algorithm to estimate the optical flow of the hand. For subtle changes in hand movements, by calculating the optical flow information, small motion changes can be made more obvious. At the same time, in the motion processing process, in order to better overcome background interference and improve the visualization effect, threshold parameters are set to overcome the interference effect, so that the algorithm can improve the detection effect of fast-moving targets. In addition, in order to better visualize the optical flow map, the Munsell color system is cited to convert the velocity vectors in the u and v directions into color changes in the horizontal and vertical directions. Different colors can represent different motion directions, and the depth can represent the speed of the movement. Compared with the original optical flow image, the hand contour after visualization can clearly reflect the changes in direction and amplitude of twisting and lifting. Figure 4 Taking twisting and complementing as an example, a, b, and c respectively compare the original manipulation movement process, the manipulation movement process after Gaussian filtering, and the optical flow processing process of the present invention. It can be clearly seen that after processing by the optical flow method and the color system, while overcoming background interference, the global contour movement changes of the hand are clearer and more obvious, and the movement change characteristics are easier to capture.

[0105] The processed optical flow action image sequence is then represented as a set of hand movements every 5 frames. The complete optical flow manipulation dataset is input into the feature extraction network for feature extraction, and a feature vector of the global scale of the hand movement is obtained.

[0106] S2. Obtain the skeleton image of the finger and the corresponding key point coordinates from the acquired acupuncture manipulation video, and calculate the motion characterization parameters including characteristic angle, characteristic distance and spatial feature information based on the key point coordinates, and then obtain the local motion feature information of the hand by performing feature extraction and feature fusion on the motion characterization parameters and the skeleton image. In literature research and communication with a large number of doctors, it was found that for these four common acupuncture manipulations (twisting to tonify, twisting to drain, lifting and inserting to tonify, lifting and inserting to drain), the key to identifying the manipulation is to quantify the kinematic parameters of the manipulation. Therefore, in order to quantitatively describe the movement change law of the hand, the present invention starts from the local characteristics of the hand movement and explores the motion parameters that can quantitatively describe different manipulations.

[0107] Further, if Figure 5 As shown, step S2 includes:

[0108] S21. Mark key points of the hand in the acquired acupuncture manipulation video using a pre-called human skeleton posture detection library to obtain a skeleton image and key point coordinate information; wherein the key points include the joints of the index finger and the thumb.

[0109] In one specific implementation, Mediapipe is used to mark key points on the hand and obtain skeleton images and key point coordinate information. Mediapipe is a human skeletal pose detection library developed by Google based on deep learning and computer vision. Compared to human skeletal pose detection libraries such as OpenPose and BlazePose, Mediapipe offers advantages in real-time performance and robustness.

[0110] When using Mediapipe to mark the hand joints, such as Figure 6 As shown in the figure, there are 21 joints in the hand. In actual acupuncture, the changes of the index finger and thumb can represent the complete acupuncture manipulation. Figure 7 In this example, the joints of the index finger and thumb are individually marked and defined as key points. The changes in the movement of these key points represent the entire movement process. The coordinates of key points 8, 6, 5, and 4 are recorded as P8(x8, y8), P6(x6, y6), P5(x5, y5), and P4(x4, y4), respectively.

[0111] S22. Calculate the characteristic angle and characteristic distance of each frame image based on the key point coordinates, then perform difference operation between each frame based on the key point coordinates to obtain the spatial change value between adjacent frame images, and finally accumulate the motion change between all frames to obtain spatial feature information.

[0112] Because the keypoint coordinates are based on the pixel positions of the image, the coordinate information between different movements is not meaningful. Therefore, for a complete manipulative movement, the coordinate information changes between the previous and next frames are considered. By calculating the difference in the motion parameters of the previous and next frames, the motion change pattern is derived.

[0113] Taking a sleight of hand movement composed of n frames of images as an example, we first calculate the coordinate information of the required key points separately, and then use the coordinates to calculate the characteristic angle and characteristic distance of each frame of image. Then, we perform difference operation between each frame to obtain the motion change between the previous and next frame images. Finally, we accumulate the motion change between all frames. The accumulated result is the motion representation parameter of a sleight of hand movement sequence.

[0114] Furthermore, if Figure 8 As shown, step S22 includes:

[0115] S221. Calculate the characteristic angle of each frame based on the key point coordinates and geometric relationships, then perform difference calculation on the characteristic angles between adjacent frames to obtain the change angle between frames, and accumulate all the change angles to obtain the total characteristic angle.

[0116] The above step S221 is implemented by the following angle formula and angle accumulation formula:

[0117]

[0118] In equations (6) and (7), n is the number of image frames, θ i is the characteristic angle of the i-th frame image, i=1,2,……n,x j and y j are the coordinates of the jth key point in the i-th frame image, j is a positive integer, and θ is the total characteristic angle of a set of action sequences.

[0119] The size and speed of the angle change during acupuncture manipulation are the key to distinguishing the four different manipulations. Therefore, in a specific implementation, for a set of action sequences consisting of n frames of images, after obtaining the coordinate information P8 (x8, y8), P6 (x6, y6), and P5 (x5, y5) of each frame, the characteristic angle θ of each frame can be calculated based on the geometric relationship and angle calculation formula. i, then the difference between the characteristic angles of the previous and next frames is calculated to obtain an angle change between frames, and finally all the changed angles are accumulated to obtain a set of total characteristic angles θ.

[0120] Taking the twisting technique as an example, 5 frames are selected as a set of action sequences. Figure 9 The characteristic angle change during the twisting process within a movement cycle is measured by the change in the characteristic angle during continuous movements to measure the amplitude and speed of the movement frequency of the manipulation. Therefore, the angles θ1, θ2, θ3, θ4, and θ5 between the joints are calculated frame by frame. The total characteristic angle change during the entire twisting process can be calculated as follows:

[0121] θ=|(θ2-θ1)|+|θ3-θ2|+(θ4-θ3)|+(θ5-θ4)| (8)

[0122] S222. The characteristic distance of each frame image is obtained based on the key point coordinates combined with the spatial position relationship, and then the difference between the characteristic distances of adjacent frame images is calculated to obtain the change distance between each frame, and all the change distances are accumulated to obtain the total characteristic distance.

[0123] The above step S222 is implemented by the following distance formula and distance accumulation formula:

[0124]

[0125] d=Σ n |(d i+1 -d i )| (10)

[0126] In formulas (9) and (10), d i is the feature distance of the current frame image, and d is the total feature distance.

[0127] In practice, the relative displacement between different joints during acupuncture manipulation is also key information for distinguishing manipulations. In actual data collection and communication with physicians, the distance between joint point P4 (x4, y4) and joint point P5 (x5, y5) is used to represent the amplitude and speed of the movement. Therefore, the spatial position relationship and the characteristic distance d of each frame image can be calculated according to the Euclidean distance formula. i , and then the feature distance differences between the previous and next frame images are accumulated to obtain the total feature distance d.

[0128] S223 . Calculate the change differences of the horizontal and vertical coordinates of adjacent frames based on the key point coordinates, and accumulate all the change differences of the horizontal and vertical coordinates to obtain the total spatial feature information.

[0129] The above step S223 is implemented by the following coordinate difference accumulation formula:

[0130]

[0131] Where, Represents the coordinate information of the current frame, Represents the coordinate information of the next frame.

[0132] In a specific embodiment, because the hand undergoes a spatial change during acupuncture, the spatial change information of the hand in the horizontal and vertical directions is represented by the coordinate change of the thumb joint point P4 (x4, y4). The change differences between the horizontal and vertical coordinates between each frame are calculated respectively, and then accumulated to obtain the total spatial change.

[0133] S23. Input the obtained motion feature parameters including the characteristic angle, characteristic distance and spatial feature information into a pre-set LSTM network for feature extraction to obtain a kinematic parameter feature vector of the acupuncture manipulation.

[0134] The original coordinate information for each set of movements is represented by a set of one-dimensional vectors (X, Y, θ, d) after the motion feature parameters are calculated. To enable the network to better learn the subtle local motion characteristics of the hand, all processed motion feature data is first input into the LSTM network for feature extraction, resulting in a feature vector for the manipulation kinematic parameters. Simultaneously, the skeleton image data is extracted using a feature extraction network to obtain a skeleton image feature vector. Finally, the kinematic parameter feature vector and the skeleton image feature vector are fused to obtain the local motion feature vector of the manipulation.

[0135] S24: Input the obtained skeleton image data into a pre-set feature extraction network for feature extraction to obtain an image feature vector of the hand skeleton.

[0136] S25. Fusing the kinematic parameter feature vector of the acupuncture manipulation with the image feature vector of the hand skeleton to obtain local motion feature information of the hand.

[0137] S3. Perform linearization operations on the global motion feature information and the local motion feature information respectively, and perform interactive attention operations, normalization processing, and weight calculation based on the two sets of query vectors, key vectors, and value vectors to obtain two sets of interactively fused feature information.

[0138] Furthermore, if Figure 10 As shown, step S3 includes:

[0139] S31. Perform linearization operations on the global motion feature information and the local motion feature information respectively to obtain two sets of corresponding query vectors, key vectors, and value vectors.

[0140] like Figure 12 As shown in the figure, for global motion feature information and local motion feature information, we first need to linearize these two types of feature vectors and use a linear layer function to obtain their corresponding query vector Q (query), key vector K (key), and value vector V (value). The formula is as follows:

[0141] Q G ,K G ,V G =Linear(F G ) (13)

[0142] Q L ,K L ,V L =Linear(F L ) (14)

[0143] Among them, Linear is the linear layer function.

[0144] S32: Perform a product operation on the query vector corresponding to the global motion feature information and the key vector corresponding to the local motion feature information to obtain a first degree of association.

[0145] S33 . Perform a product operation on the query vector corresponding to the local motion feature information and the key vector corresponding to the global motion feature information to obtain a second degree of association.

[0146] Unlike the traditional attention mechanism, the present invention no longer calculates the similarity of Q and K between each independent attention module. What needs to be paid more attention to is the connection between features of different scales. Therefore, in order to obtain the mutual information between the two scales, that is, the correlation between the global feature and the local feature, first calculate the Q of the global feature. G With local features K L Perform product operation to obtain the first correlation degree. Similarly, the Q of the local feature is L K with global features G Multiplying ∠ with ∠ to obtain the second degree of association, the formula is as follows:

[0147] a GL =Q G ·K L (15)

[0148] a LG =Q L ·K G (16)

[0149] S34 , respectively normalize the obtained first correlation degree and second correlation degree through a Softmax function, so that the first correlation degree and the second correlation degree always remain between 0 and 1.

[0150] S35. Perform a matrix product operation on the value vector corresponding to the global motion feature information and the first correlation to obtain a first weight value, and perform a matrix product operation on the value vector corresponding to the local motion feature information and the second correlation to obtain a second weight value. The calculation formula is as follows:

[0151] W G =V G Softmax(a GL ) (17)

[0152] W L =V L Softmax(a LG ) (18) S4. Introduce random parameters to perform random weight fusion on the two sets of interactive fusion feature information, and finally obtain the acupuncture manipulation recognition result.

[0153] Furthermore, if Figure 11 As shown, step S4 includes:

[0154] S41. Introduce a random parameter λ.

[0155] S42 : Assign a first weight coefficient λ and a second weight coefficient 1-λ to the first weight value and the second weight value at the two scales, respectively.

[0156] S43 , performing random weight processing on the first weight value and the second weight value according to the first weight coefficient λ and the second weight coefficient 1-λ respectively, to obtain a total output result.

[0157] In order to more effectively fuse the interactive information of the two scale feature data and improve the generalization ability of the overall network, the output weights W of the two scales are G and W L , weighted by random parameters, and then added based on the weighted results.

[0158] First, we introduce a parameter λ, which is the number of random samples that obey the beta distribution, and then we calculate the output weights W for the two scales. G and W L , assign weight coefficients λ and 1-λ to them respectively, and finally add the output results of the two scale features after random weight processing to obtain the total output result.

[0159] Among them, the total output result is obtained through the following weight formula:

[0160] Output=λ*W G +(1-λ)*W L

[0161] λ∈Beta(α,α)

[0162] Where W G is the first weight value, W L is the second weight value, λ∈Beta(α,α) means that λ is the number of random samples that obey the beta distribution, and the parameter α is an adjustable parameter. Its value can be adjusted through experiments to obtain the best parameter results. The results of different parameters are shown in Table 1 and Figure 13 As shown,

[0163] Table 1 Model recognition accuracy under different α values

[0164]

[0165] From Table 1 and Figure 13 It can be found that when the α value is small, that is, when α < 1, the probability of the values ​​at both ends of the Beta distribution is greater, and the probability of the value in the middle is smaller. When α → ∞, the Beta distribution is closer to the binomial distribution, indicating that the output results of the two attention modules are more independent, resulting in the final model output result. One set of feature data has a larger weight, and the other set of feature data contributes almost no feature information. At this time, the two parts of the data cannot be effectively fused after random weight processing, resulting in a low final recognition accuracy. When the α value is large, that is, when α > 1, as α increases, the probability curve shows a larger probability in the middle and a smaller probability at both ends. When α → ∞, the probability value will approach 0.5, indicating that after random weight fusion processing, the output results of the two parts of the feature data are assigned according to a weight ratio of 0.5. At this time, the final output result of the network is the linear addition of the two parts of the feature data. Linear fusion cannot effectively fuse the two parts of the feature data, resulting in a general final recognition effect.

[0166] pass Figure 13 The Beta distribution curve shown shows that when 0.7<α<1, the probability density curve is close to the uniform distribution between (0, 1). When α=1, the probability distribution result is equal to the uniform distribution between (0, 1), indicating that for the feature information of the two parts of the data, the final output result can be uniformly and effectively fused with the feature information processed by the attention module, so that the final output result contains both the global motion feature information and the local subtle motion feature information of the finger. Compared with the linear fusion method, this fusion method can improve the fusion efficiency between different feature data, effectively combine the feature information of the two, and improve the recognition accuracy.

[0167] In addition, the present invention also provides an acupuncture manipulation recognition system for global hand shape and local finger motion, comprising: a global information acquisition module for converting the RGB image sequence of the acquired acupuncture manipulation video into a visual optical flow sequence by combining the color system with the pixel changes between different frames, and then extracting the global motion feature information of the hand by performing feature extraction on the visual optical flow sequence; a local information acquisition module for obtaining the skeleton image of the finger and the corresponding key point coordinates from the acquired acupuncture manipulation video, and calculating the motion representation parameters including feature angle, feature distance and spatial feature information based on the key point coordinates, and then extracting and fusing the motion representation parameters and the skeleton image to obtain the local motion feature information of the hand; a multi-scale interactive fusion module for linearizing the global motion feature information and the local motion feature information, and performing interactive attention operation, normalization processing and weight calculation based on the two sets of query vectors, key vectors and value vectors obtained to obtain the two sets of interactive fused feature information; and a random weight fusion module for introducing random parameters to perform random weight fusion on the two sets of interactive fused feature information to ultimately obtain the acupuncture manipulation recognition result.

[0168] refer to Figure 14 It can be seen that the network structure of the acupuncture manipulation recognition model is mainly composed of three parts:

[0169] The first part is global feature extraction, which addresses the issues of background and lighting interference in the original RGB video. To overcome these interferences and better extract the global motion features of the hand movements, the RGB image sequence is converted into an optical flow image sequence by utilizing pixel variations between frames. The optical flow features are then used to describe the global contour motion changes of the hand. The optical flow sequence is then used as a channel data for the hand gesture and input into the residual backbone network for feature extraction, obtaining the global motion feature information of the hand gesture changes.

[0170] The second part is local feature extraction, which focuses on the recognition difficulties caused by the small amplitude and high frequency of acupuncture manipulation. In order to quantitatively study and analyze the movement patterns of the manipulation, starting from a specific local finger, Mediapipe is used to obtain the skeleton image of the finger and the corresponding key point coordinates. The key point coordinates are used to calculate the kinematic parameters, and the changes in characteristic angles, characteristic distances, and spatial feature information during the acupuncture manipulation movement are obtained. The changes in parameter values ​​are used to quantitatively describe the local movement changes of the manipulation. The parameter information is then passed through the LSTM network for feature extraction, and the skeleton image is passed through the residual network for feature extraction. The outputs of the two are then fused on the same dimension to obtain local motion feature information.

[0171] The third part is the attention feature fusion part. In order to more effectively fuse the global feature information and local feature information extracted by the first two parts, and at the same time to explore the intrinsic connection between the two parts of feature information, the two parts of feature information are effectively combined and utilized. An "interactive attention module" is proposed to effectively fuse the two parts of feature vectors through a series of operations such as interactive attention calculation and random weight fusion, thereby improving the overall performance of the model.

[0172] Finally, in order to verify the scientificity and effectiveness of the acupuncture manipulation recognition network proposed in this invention in solving the acupuncture manipulation recognition problem, the recognition network of this invention was compared with five classic action recognition networks. The original RGB manipulation image or video dataset was used as the input of the classic action recognition network, and 100 iterations of training were performed to obtain the classification results. The recognition accuracy of each action recognition network is shown in Table 2, where the confusion matrix of each model is shown in Table 2. Figure 15 As shown:

[0173] Table 2 Comparison of accuracy of different network models

[0174] Num Methods Accuracy a Tran et al (2015)

[62] 70.3% b Qiu et al (2017)

[63] 75.2% c Carreira et al (2017)

[64] 77.5% d Hera at al(2017

[65] 80.7% f Du et al (2019)

[66] 85.1% Our method 95.3%

[0175] Through Table 2 and Figure 15The experimental results show that when the classic three-dimensional convolutional action network, that is, the action recognition network represented by C3D

[62] , P3D

[63] , I3D

[64] and ResNet3D

[65] , is used to process acupuncture manipulation data, the confusion matrix results show that the three-dimensional convolutional network can effectively learn the specificity between twisting and lifting and inserting manipulations when processing manipulation movements, and has a high recognition accuracy between twisting and lifting and inserting. However, it is found that the network cannot effectively distinguish between tonification and drainage, that is, twisting tonification and twisting drainage, and lifting and inserting tonification and lifting and inserting drainage. This is mainly because the three-dimensional convolutional network has an advantage in processing global motion features. It can extract spatial motion features for the changes in twisting and lifting and inserting manipulations in the overall range, but cannot extract effective motion features for the details of local motion changes, such as the microscopic motion changes in angle and distance of tonification and drainage, resulting in a low final recognition rate. When the CSN

[66] network is used to process the technique data, the result is shown in Figure f. Compared with the previous three-dimensional convolutional network, the CSN network has advantages in processing temporal information and spatial information during the movement process. Therefore, it can more effectively extract the temporal and spatial features of the movement process, which to a certain extent makes up for the shortcomings of the three-dimensional convolutional network. Therefore, the recognition accuracy of tonification and purgation methods has been improved, and the final recognition accuracy is 85.1%. However, the network structure still does not propose a specific and effective feature extraction method for the difference between tonification and purgation techniques, resulting in the recognition accuracy not being able to achieve the ideal effect.

[0176] In the network model designed by the present invention, in view of the difficulties that the previous action recognition network has in the tonifying and purging methods, the present invention starts from the global features and local features respectively. In order to overcome the interference of the background, environment and other factors in the RGB data on the recognition process, the RGB data is converted into optical flow image data in the global features. While overcoming the background interference, the contour information of the hand movement is further visualized. Through the movement changes of the hand contour, the change amplitude and frequency of the manipulation movement process can be better reflected. At the same time, for the subtle changes in the manipulation operation process, the manipulation movement process is quantitatively described by exploring its kinematic change law. The spatial coordinate information of the corresponding key points is obtained based on the skeleton image, and the kinematic feature parameters of the complete movement process are calculated according to the change of the coordinates. The kinematic parameter information obtained by calculation can effectively reflect the subtle changes in the movement amplitude and movement frequency between the tonifying and purging methods, and the subtle local movement change law is difficult for the deep network to learn. Finally, through the confusion matrix, it can be found that when the proposed feature processing method and network model are used, the recognition accuracy of twisting to tonify and twisting to drain, lifting and inserting to tonify and lifting and inserting to drain is significantly improved. Therefore, the network can effectively solve the difficulties brought by the tonifying and draining methods in the manipulation recognition process, and the recognition accuracy of the four manipulations can reach 95.3%. Therefore, in dealing with the problem of acupuncture manipulation action recognition, the solution provided by the present invention has significant advantages. In addition, after comparison with the mainstream action recognition network, it was found that the network model of the present invention can effectively solve the recognition difficulty problem between the tonifying and draining methods, greatly improving the recognition accuracy of the four manipulations.

[0177] Furthermore, the present invention provides an acupuncture manipulation recognition device for global hand shape and local finger movement, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the acupuncture manipulation recognition method for global hand shape and local finger movement as described above.

[0178] Furthermore, the present invention provides a computer-readable medium having computer-executable instructions stored thereon, which, when executed by a processor, implement the acupuncture manipulation recognition method for global hand shape and local finger movement as described above.

[0179] In summary, the present invention provides a solution, model, device and medium for acupuncture manipulation recognition for global hand shape and local finger movement. The present invention extracts features and quantitatively describes the acupuncture manipulation process from the perspectives of global features and local features. In terms of global features, the original RGB data is converted into optical flow image data to overcome the influence of background, lighting and other factors on the manipulation recognition in the RGB data. After processing, the contour information of the hand is more obvious and the changes in the movement process are clearer. In terms of local features, in order to more specifically quantify the local movement change rules of the manipulation, Mediapipe is used to extract the key point coordinate information of the hand, record the coordinate changes during the manipulation movement, and calculate the defined manipulation kinematic feature parameters based on the changes in the coordinate information. Subsequently, the kinematic parameter feature vector extracted by the feature network is combined with the skeleton image feature vector to obtain a feature vector that can represent the local motion information. Finally, in order to effectively combine the global feature information with the local feature information, an "interactive attention module" is proposed to mine the correlation between different data so that the final output result contains the feature information of the two parts of the data.

[0180] Since the systems / devices described in the above embodiments of the present invention are systems / devices used to implement the methods of the above embodiments of the present invention, those skilled in the art will be able to understand the specific structures and variations of these systems / devices based on the methods described in the above embodiments of the present invention, and thus will not be described in detail here. All systems / devices used in the methods of the above embodiments of the present invention are within the scope of protection of the present invention.

[0181] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0182] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions.

[0183] It should be noted that the word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The present invention may be implemented by means of hardware comprising several distinct components and by means of a suitably programmed computer. In a claim enumerating several means, several of these means may be embodied by the same piece of hardware. The use of the words first, second, third, etc., is merely for convenience and does not imply any order. These words should be understood as part of the component name.

[0184] In addition, it should be noted that, in the description of this specification, the description of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.

[0185] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments after learning the basic creative concept. Therefore, the claims should be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0186] Obviously, those skilled in the art may make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if such modifications and variations fall within the scope of the claims and their equivalents, the present invention shall also include such modifications and variations.

Claims

1. A method for acupuncture manipulation recognition based on global hand shape and local finger movement, characterized by: include: By utilizing the pixel changes between different frames, the RGB image sequence of the acupuncture manipulation video is converted into a visual optical flow sequence in combination with the color system. The global motion feature information of the hand is then obtained by feature extraction of the visual optical flow sequence. The skeleton image of the finger and the corresponding key point coordinates are obtained from the acupuncture manipulation video. The motion representation parameters including characteristic angle, characteristic distance and spatial feature information are calculated based on the key point coordinates. The local motion feature information of the hand is then obtained by feature extraction and feature fusion of the motion representation parameters and the skeleton image. The global motion feature information and the local motion feature information are linearized separately. Based on the two sets of query vectors, key vectors, and value vectors, interactive attention operations, normalization processing, and weight calculation are performed to obtain two sets of interactively fused feature information. Random parameters are introduced to perform random weight fusion on the two sets of interactively fused feature information, and finally the acupuncture manipulation recognition results are obtained.

2. The acupuncture manipulation recognition method for global hand shape and local finger movement according to claim 1, characterized in that: By utilizing the pixel changes between different frames, the RGB image sequence of the acupuncture manipulation video is converted into a visual optical flow sequence in combination with the color system. The global motion feature information of the hand is then obtained by feature extraction of the visual optical flow sequence, including: By using the pixel changes between different frames, the RGB image sequence of the acupuncture manipulation video is obtained through optical flow calculation to obtain an optical flow sequence; The Munsell color system is used to convert the velocity vectors of the optical flow sequence in the u and v directions into the color changes in the horizontal and vertical directions, thereby obtaining a visual optical flow sequence. The visualized optical flow sequence is input into a pre-set feature extraction network for feature extraction to obtain the global motion feature information of the hand.

3. The acupuncture manipulation recognition method for global hand shape and local finger movement according to claim 1, characterized in that: The skeleton image of the finger and the corresponding key point coordinates are obtained from the acupuncture manipulation video. Based on the key point coordinates, motion representation parameters including characteristic angles, characteristic distances, and spatial feature information are obtained. Then, feature extraction and feature fusion of the motion representation parameters and the skeleton image are performed to obtain the local motion feature information of the hand, including: Using a pre-used human skeleton posture detection library, key points of the hand in the acupuncture manipulation video are marked to obtain skeleton images and key point coordinate information; the key points include the joints of the index finger and thumb; The characteristic angle and characteristic distance of each frame image are calculated based on the key point coordinates, and then the difference operation is performed between each frame based on the key point coordinates to obtain the spatial change value between adjacent frame images. Finally, the motion change between all frames is accumulated to obtain the spatial feature information; The obtained motion feature parameters including characteristic angle, characteristic distance and spatial feature information are input into a pre-set LSTM network for feature extraction to obtain the kinematic parameter feature vector of the acupuncture manipulation; The obtained skeleton image data is input into a pre-set feature extraction network for feature extraction to obtain an image feature vector of the hand skeleton; The kinematic parameter feature vector of the acupuncture manipulation and the image feature vector of the hand skeleton are fused to obtain the local motion feature information of the hand.

4. The acupuncture manipulation recognition method for global hand shape and local finger movement according to claim 3, characterized in that: Based on the key point coordinates, the characteristic angle and characteristic distance of each frame image are calculated. Then, based on the key point coordinates, the difference operation is performed between each frame to obtain the spatial change value between adjacent frame images. Finally, the motion change between all frames is accumulated to obtain the spatial feature information including: The characteristic angle of each frame is obtained by solving the key point coordinates and geometric relationships. Then, the characteristic angles between adjacent frames are calculated to obtain the change angle between frames. All the change angles are accumulated to obtain the total characteristic angle. The characteristic distance of each frame image is obtained based on the key point coordinates and the spatial position relationship. Then, the characteristic distances between adjacent frames are calculated to obtain the change distance between frames. All the change distances are accumulated to obtain the total characteristic distance. The change differences of the horizontal and vertical coordinates of adjacent frames are calculated based on the key point coordinates, and all the change differences of the horizontal and vertical coordinates are accumulated to obtain the total spatial feature information.

5. The acupuncture manipulation recognition method for global hand shape and local finger movement according to claim 4, characterized in that: Based on the key point coordinates and geometric relationships, the characteristic angle of each frame is calculated. Then, the characteristic angles between adjacent frames are calculated to obtain the change angle between frames. All the change angles are accumulated to obtain the total characteristic angle. The following angle formula and angle accumulation formula are used for implementation: θ=Σ n |(θ i+1 -θ i )| Where n is the number of image frames, θ i is the characteristic angle of the i-th frame image, i=1,2,……n,x j and y j are the coordinates of the jth key point in the i-th frame image, j is a positive integer, and θ is the total characteristic angle of a set of action sequences; The characteristic distance of each frame image is obtained based on the key point coordinates combined with the spatial position relationship. Then, the characteristic distances between adjacent frame images are calculated to obtain the change distance between frames. All the change distances are accumulated to obtain the total characteristic distance through the following distance formula and distance accumulation implementation: d=Σ n |(d i+1 -d i )| Where, d i is the feature distance of the current frame image, and d is the total feature distance; Based on the key point coordinates, the change difference of the horizontal and vertical coordinates of adjacent frames is calculated, and all the change differences of the horizontal and vertical coordinates are accumulated to obtain the total spatial feature information. The following coordinate difference accumulation formula is used for implementation: Where X is the total change difference of the horizontal coordinate, and Y is the total change difference of the vertical coordinate. Represents the coordinate information of the current frame, Represents the coordinate information of the next frame.

6. The acupuncture manipulation recognition method for global hand shape and local finger movement according to claim 1, characterized in that: The global motion feature information and the local motion feature information are linearized separately. Based on the two sets of query vectors, key vectors, and value vectors, interactive attention operations, normalization processing, and weight calculation are performed. The feature information obtained after the two sets of interactive fusion includes: Perform linearization operations on the global motion feature information and the local motion feature information respectively to obtain two sets of corresponding query vectors, key vectors, and value vectors; Performing a product operation on the query vector corresponding to the global motion feature information and the key vector corresponding to the local motion feature information to obtain a first degree of association; Performing a product operation on the query vector corresponding to the local motion feature information and the key vector corresponding to the global motion feature information to obtain a second degree of association; The obtained first correlation degree and second correlation degree are respectively normalized by the Softmax function so that the first correlation degree and the second correlation degree are always kept between 0 and 1; A first weight value is obtained by performing a matrix product operation on the value vector corresponding to the global motion feature information and the first correlation degree, and a second weight value is obtained by performing a matrix product operation on the value vector corresponding to the local motion feature information and the second correlation degree.

7. The acupuncture manipulation recognition method for global hand shape and local finger movement according to claim 6, characterized in that: The random parameters are introduced to perform random weight fusion on the two sets of interactive fusion feature information, and the final acupuncture manipulation recognition results include: Introduce a random parameter λ; For the first weight value and the second weight value at the two scales, a first weight coefficient λ and a second weight coefficient 1-λ are assigned respectively; Perform random weight processing on the first weight value and the second weight value according to the first weight coefficient λ and the second weight coefficient 1-λ respectively to obtain a total output result; Among them, the total output result is obtained through the following weight formula: Output=λ*W G +(1-λ)*W L λ∈Beta(α,α) Where W G is the first weight value, W L is the second weight value, λ∈Beta(α,α) means that λ is the number of random samples obeying the beta distribution, and the parameter α is an adjustable parameter.

8. An acupuncture manipulation recognition system for global hand shape and local finger movement, characterized by: include: The global information acquisition module is used to convert the RGB image sequence of the acupuncture manipulation video into a visual optical flow sequence by combining the color system and using the pixel changes between different frames. The global motion feature information of the hand is then obtained by feature extraction from the visual optical flow sequence. The local information acquisition module is used to obtain the skeleton image of the finger and the corresponding key point coordinates from the acquired acupuncture manipulation video. Based on the key point coordinates, motion representation parameters including characteristic angles, characteristic distances, and spatial feature information are calculated. The local motion feature information of the hand is then obtained by feature extraction and feature fusion of the motion representation parameters and the skeleton image. The multi-scale interactive fusion module is used to linearize the global motion feature information and the local motion feature information respectively. Based on the two sets of query vectors, key vectors, and value vectors, interactive attention operations, normalization processing, and weight calculation are performed to obtain two sets of interactive fused feature information. The random weight fusion module is used to introduce random parameters to perform random weight fusion on the two sets of interactively fused feature information, and finally obtain the acupuncture manipulation recognition result.

9. An acupuncture manipulation recognition device for global hand shape and local finger movement, characterized by: include: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the acupuncture manipulation recognition method for global hand shape and local finger movement as described in any one of claims 1 to 7.

10. A computer-readable medium having computer-executable instructions stored thereon, characterized in that: When the executable instructions are executed by the processor, the acupuncture manipulation recognition method for global hand shape and local finger movement according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Video-based human body interaction action recognition method

    CN108241849A

  • Global-local RGB-D multimode-based gesture recognition method

    CN108388882A