A Needling Manipulation Recognition Method, System, Device and Medium Based on Ultrasonic Dynamic Optical Flow and Self-Supervised Mask

Through ultrasonic dynamic optical flow and self-supervised mask technology, the problem of insufficient dynamic biomechanical characteristics in acupuncture technique recognition is solved, efficient end-to-end recognition is achieved, and the recognition accuracy and robustness of acupuncture technique is improved.

CN120107319BActive Publication Date: 2025-07-18BEIJING UNIV OF CHEM TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510592016.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-07-18
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively characterize the dynamic biomechanical characteristics of subcutaneous tissue under the action of acupuncture techniques, and there is a lack of an end-to-end recognition model that integrates space-time dynamic effects, resulting in strong subjectivity and inefficiency in the recognition process.

Method used

Using a method based on ultrasonic dynamic optical flow and self-supervised mask, keyframe fragments are screened through optical flow analysis, input sequences are constructed, and direction-specific multi-head attention calculation and cross-layer mask conduction modulation are used, combined with a cascaded spatiotemporal pyramid pooling module and multi-expert classifiers, the end-to-end identification of acupuncture techniques is achieved.

Benefits of technology

It significantly improves the robustness and accuracy of acupuncture techniques, reduces the degree of manual intervention, and establishes an end-to-end identification system that integrates photofluidic dynamic characteristics, tissue deformation response and space-time coupling effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107319B_ABST
    Figure CN120107319B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image recognition technology, and in particular to a method, system, device and medium for acupuncture manipulation recognition based on ultrasonic dynamic optical flow and self-supervised mask. The method includes: screening out key frame segments based on the inter-frame optical flow intensity difference of ultrasonic image sequences, and constructing an input sequence by expanding buffer frames; obtaining the forward and backward optical flow of adjacent frames and performing bidirectional alignment, performing direction-specific multi-head attention calculation on the aligned frame group to output a time-varying probability mask, and injecting it into the first layer of the spatio-temporal separation attention network for temporal attention calculation, and performing cross-layer mask conduction modulation in subsequent predetermined levels; inputting the gated modulation features output by each level into a multi-expert classifier after granularity feature extraction and cross-scale channel attention fusion to output the recognition result. The present invention reduces the degree of manual intervention while establishing the first end-to-end recognition system that integrates the dynamic characteristics of optical flow, tissue deformation response and spatio-temporal coupling effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and particularly to a method, system, device and medium for identifying acupuncture manipulation based on ultrasonic dynamic optical flow and self-supervised mask. Background Art

[0002] In the research of traditional Chinese medicine acupuncture manipulation, different schools and master-disciple relationships lead to significant differences in hand type parameters when physicians perform the same manipulation. Taking the twirling reinforcing manipulation as an example, the differences in the needle-holding postures of different physicians directly affect the unified quantification of manipulation parameters and the classification accuracy. Traditional research usually realizes the characterization of manipulation by collecting external physical parameters such as the rotation angle of the needle body and the lifting and thrusting amplitude. However, such parameters are difficult to reveal the dynamic stimulation effect of the manipulation on subcutaneous tissues. Therefore, the academic community has gradually turned to combining medical imaging technologies to explore the response of subcutaneous tissues, such as observing the displacement of connective tissues through ultrasonic waves or recording the changes in functional connections of brain regions using magnetic resonance imaging.

[0003] The existing technologies mainly analyze the subcutaneous stimulation effect through two paths: one is to indirectly obtain parameters such as blood flow and temperature by means of biosignal sensing technology, and the other is to capture the dynamic characteristics of tissues relying on medical imaging technologies such as ultrasonic, elastography or functional magnetic resonance. However, the existing methods face multiple challenges: First, the movement of subcutaneous tissues has highly complex three-dimensional dynamic characteristics, but there is a lack of a standardized image dataset covering various manipulations, resulting in research being limited to small-sample manually annotated data; Second, the feature extraction process depends on manually annotating anatomical structures or specific algorithm processing, which has problems of strong subjectivity and low repeatability, and it is difficult to achieve cross-case generalization; Finally, existing research mostly focuses on evidence-based analysis of curative effects and fails to construct an end-to-end recognition model from the dynamic response of subcutaneous tissues to manipulation categories, restricting its clinical application value. The breakthrough of these bottlenecks requires solving key technical problems such as the biomechanical characterization of subcutaneous tissues, adaptive feature extraction and spatio-temporal feature fusion. Summary of the Invention

[0004] (I) Technical Problems to be Solved

[0005] In view of the above-mentioned disadvantages and deficiencies of the existing technology, the present invention provides a method, system, device and medium for identifying acupuncture manipulation based on ultrasonic dynamic optical flow and self-supervised mask, which solves the technical problems that the existing technology is difficult to effectively characterize the dynamic biomechanical characteristics of subcutaneous tissues under acupuncture manipulation, the recognition process is highly subjective and inefficient due to its reliance on manually annotated image feature extraction methods, and there is a lack of an end-to-end recognition model integrating spatio-temporal dynamic effects.

[0006] (II) Technical Solutions

[0007] In order to achieve the above object, the main technical solutions adopted by the present invention include:

[0008] In a first aspect, an embodiment of the present invention provides a method for identifying acupuncture manipulation based on ultrasonic dynamic optical flow and self-supervised mask, including:

[0009] Perform optical flow analysis on the obtained ultrasonic image sequence, screen out key frame segments based on the difference in inter-frame optical flow intensity, and construct an input sequence by extending buffer frames forward and backward.

[0010] Adopt an overlapping sliding window on the input sequence to intercept local segments, calculate the forward and backward optical flow of adjacent frames within the local segments and perform bidirectional alignment, and perform direction-specific multi-head attention calculation on the aligned frame group to output a time-varying probability mask.

[0011] Inject the time-varying probability mask into the first layer of the spatio-temporal separation attention network for temporal attention calculation, and use a gated unit in subsequent predetermined levels to perform cross-layer mask conduction modulation on the mask feature of the first layer and the current layer feature.

[0012] Concatenate the gated modulation features output by each level and input them into a cascaded spatio-temporal pyramid pooling module to extract time-sharing granularity motion features. After cross-scale channel attention fusion, input them into a multi-expert classifier to output the acupuncture manipulation identification result.

[0013] Optionally, performing optical flow analysis on the obtained ultrasonic image sequence, screening out key frame segments based on the difference in inter-frame optical flow intensity, and constructing an input sequence by extending buffer frames forward and backward includes:

[0014] Obtain and standardize the original ultrasonic video data, calculate the optical flow field of each frame of the obtained ultrasonic image sequence, and generate a visual inter-frame optical flow image sequence encoded in HSV color.

[0015] Convert the visual inter-frame optical flow image sequence to the HSV color space and perform independent normalization processing on each channel, calculate the global optical flow intensity of each frame, and input it into a Gaussian difference filter with temporal awareness to screen out significant time points where the difference value of the optical flow intensity exceeds three times the average fluctuation level of adjacent frames.

[0016] On the ultrasonic image sequence, use the significant time point as the center and extend to both sides of the time series until the continuous region where the optical flow intensity change rate drops to the baseline level is used as the key frame segment.

[0017] Based on the key frame segment, adaptively extend buffer frames bidirectionally forward and backward to obtain an initial input sequence.

[0018] Perform optical flow gradient consistency detection on the initial input sequence. If the motion direction mutation between adjacent frames exceeds a preset mutation angle, insert an interpolation transition frame, and finally output an input sequence resistant to motion breakage.

[0019] Optionally, the visual inter-frame optical flow image sequence is converted to the HSV color space and each channel is independently normalized, the global optical flow intensity of each frame is calculated, and a Gaussian difference filter with time-series perception is input to screen the significant time points where the optical flow intensity difference value exceeds three times the average fluctuation level of adjacent frames, including:

[0020] Convert the visual inter-frame optical flow image sequence to the HSV color space, and independently normalize the hue channel, saturation channel, and brightness channel of the HSV color space;

[0021] Perform pixel-by-pixel multiplication and accumulation on the normalized values of the hue channel, saturation channel, and brightness channel to obtain the global optical flow intensity of each optical flow image that represents the dynamic change intensity of the muscle tissue, and then obtain a global optical flow intensity sequence;

[0022] The first sliding time window is used to perform sliding average filtering on the global optical flow intensity sequence to generate a baseline trend component, and the global optical flow intensity sequence is subtracted from the baseline component to obtain a residual sequence;

[0023] Applying a second sliding time window to the residual sequence, obtaining a second-order derivative response of the residual signal within the second sliding time window through a Gaussian difference filter to generate an enhanced residual sequence;

[0024] In the enhanced residual sequence, the standard deviation of the optical flow intensity fluctuation of the N frames before and after the current frame is used as the benchmark. If the instantaneous difference value exceeds three times the standard deviation, it is marked as a significant time point.

[0025] Optionally, the adaptive extended buffer frame includes:

[0026] Forward expansion intercepts an optical flow compensation frame group containing at least two consecutive frames before the key frame segment, and when the time interval between the starting point of the key frame segment and the ending point of the preceding key frame segment is less than a preset merging threshold, merges the optical flow compensation frame group into the backward expansion sequence of the preceding key frame segment;

[0027] Backward expansion intercepts a motion attenuation observation frame group containing at least three consecutive frames after the key frame segment, and the interception termination condition of the motion attenuation observation frame group is that the amplitude of the optical flow between frames decays to below a preset amplitude threshold of the central frame amplitude of the key frame segment;

[0028] When the starting position of the key frame segment is located at the starting point of the ultrasound image sequence, the optical flow field of the first frame is spatially mirrored along the main axis of the muscle fiber of the corresponding ultrasound image, and a random perturbation vector with an amplitude of 15%-25% of the standard deviation of the optical flow of the current frame is superimposed to generate a virtual forward compensation frame;

[0029] When the key frame segment ends at the end of the ultrasound image sequence, the motion trajectory is extrapolated according to 1.2-1.8 times the duration of the optical flow vector of the last frame to generate a virtual backward observation frame.

[0030] Optionally, an overlapping sliding window is adopted on the input sequence to intercept local segments, the forward and backward optical flows of adjacent frames within the local segments are obtained and two-way alignment is performed, and direction-specific multi-head attention calculation is performed on the aligned frame groups. The output time-varying probability mask includes:

[0031] Intercept a local segment containing at least one key frame segment, at least two forward buffer frames, and at least three backward buffer frames on the input sequence with a sliding window at an adaptive overlapping rate;

[0032] Use a lightweight optical flow network to perform forward and backward optical flow estimation on each adjacent frame of each local segment. Based on the forward optical flow, the previous frame is spatially deformed and mapped to align with the current frame coordinate space. Based on the backward optical flow, the subsequent frame is reversely deformed and aligned to the current frame coordinate space. The three aligned frames are concatenated along the channel dimension to generate a spatio-temporally aligned superimposed frame group;

[0033] Parallelly input the horizontally moving attention head group, vertically moving attention head group, and composite motion analysis head group in the multi-head attention module to process the features of the spatio-temporally aligned superimposed frame group;

[0034] Orthogonally project and fuse the output features of the horizontally moving attention head group and the vertically moving attention head group to generate a feature tensor representing the joint distribution of motion vectors;

[0035] Concatenate the output features of the composite analysis head group and the feature tensor representing the joint distribution of motion vectors along the channel axis to generate a cross-motion mode fusion feature;

[0036] Perform spatial pyramid pooling operation on the cross-motion mode fusion feature, extract multi-granularity motion features at different scales respectively, and use channel attention gating to weight and fuse the features at each scale, retain the cross-dimensional coupling features representing horizontal muscle fiber contraction, vertical subcutaneous deformation, and non-directional motion, perform convolution and non-linear activation on the coupling features, and output an initial mask;

[0037] Based on the optical flow vectors within adjacent sliding windows, perform cross-window fusion on the overlapping regions of adjacent initial masks, perform morphological closing operation on the fused mask to eliminate holes and isolated noises, and refine the mask boundary through sub-pixel optical flow-guided interpolation, and output a time-varying probability mask with continuous boundaries for representing the motion probability distribution of subcutaneous tissues;

[0038] Among them,

[0039] The lateral motion focus head group is configured as a motion parsing module based on horizontal optical flow gradient constraints. By enhancing the attention weight of the horizontal motion component, the motion features related to the contraction of horizontal muscle fibers are specifically extracted. The horizontal optical flow gradient constraint is: when calculating the query-key similarity matrix, a dynamic weight gain is applied to the pixel points whose horizontal optical flow gradient component exceeds the preset ratio threshold of the vertical component.

[0040] The vertical motion focus head group is configured as a motion analysis module based on the vertical optical flow gradient constraint. By enhancing the characteristic response intensity of the vertical optical flow gradient component, the motion features related to the vertical deformation of the subcutaneous tissue caused by the acupuncture insertion and lifting operation are specifically extracted; the vertical optical flow gradient constraint is: when calculating the attention weight, the exponential feature enhancement of the channel dimension is implemented for the area where the vertical optical flow gradient component exceeds the preset ratio threshold of the horizontal optical flow gradient component;

[0041] The compound motion analysis head group is configured as a full-degree-of-freedom motion analysis module without directional constraints, which is used to capture nonlinear motion characteristics related to rotational motion, oblique shear motion and multi-directional compound motion.

[0042] Optionally, injecting the time-varying probability mask into the first layer of the spatiotemporal separation attention network to perform temporal attention calculation, and using a gating unit in a subsequent predetermined layer to perform cross-layer mask conduction modulation on the first layer mask feature and the current layer feature, including:

[0043] The time-varying probability mask is input into the Patch embedding module synchronized with the spatiotemporal separation attention network. Through the two-dimensional convolution operation with preset block size and step length, the spatial area of the mask is processed into non-overlapping blocks to generate mask embedding blocks.

[0044] Multiple mask embedding blocks are stacked sequentially along the time dimension, and the channel dimension is expanded and transformed through a learnable linear projection layer to generate a mask feature tensor consistent with the dimension of the spatiotemporal attention network input tensor;

[0045] In the temporal attention calculation stage of the first layer of the spatiotemporal separation attention network, the mask feature tensor is added to the original query tensor element by element, and then a layer normalization operation is performed to tilt the temporal attention weight toward the motion probability area marked by the mask that exceeds the preset probability threshold;

[0046] Through the gated recalibration units set at the 2nd, 4th and 6th layers, the mask feature tensor is reduced in dimension by depthwise separable convolution and then transferred across layers, and the query tensor of the current layer is input into the two-level fully connected network to generate the gated weights;

[0047] The gated weights are used to dynamically scale the reduced-dimensional mask features and superimposed on the current layer features through residual connections to obtain gated modulation features containing mask-guided information.

[0048] Through a pre-deployed auxiliary supervision branch, calculate the mutual information entropy between the deep features and the first-layer mask feature tensor. If the entropy value is lower than the set entropy threshold, trigger the backtracking adjustment of the gating weights.

[0049] Optionally, splice the gated modulation features output by each layer and input them into a cascaded spatio-temporal pyramid pooling module to extract the time-sharing granularity motion features. After cross-scale channel attention fusion, input them into a multi-expert classifier. The output acupuncture manipulation recognition results include:

[0050] Splice the gated modulation features output by each layer along the channel dimension to obtain a fused feature tensor containing shallow motion details and deep semantic associations, and perform batch normalization and spatio-temporal dimension reorganization to match the input structure of the cascaded spatio-temporal pyramid.

[0051] Input the fused feature tensor after batch normalization and spatio-temporal dimension reorganization into the spatio-temporal pyramid module, perform dense sliding max pooling within a 4-8 frame window respectively, obtain short-term granularity features by extracting the myofibril fibrillation features caused by the instantaneous action of acupuncture, perform sparse self-attention pooling within a 16-24 frame window, obtain medium-term granularity features by acquiring the motion rhythm features of a single lifting-thrusting or twirling operation, and perform adaptive average pooling within a 32-48 frame window to extract the cumulative biomechanical features under the continuous action of multiple manipulations to obtain long-term granularity features.

[0052] Perform multi-scale alignment and channel attention weighted fusion on the short-term, medium-term, and long-term granularity features to obtain the final fused feature.

[0053] Input the final fused feature into parallel optical flow motion experts, tissue deformation experts, and spatio-temporal coupling experts, and perform weighted voting based on the confidence scores output by each expert to output the probability distribution of acupuncture manipulation categories.

[0054] In a second aspect, an acupuncture manipulation recognition system based on ultrasonic dynamic optical flow and self-supervised mask provided by an embodiment of the present invention includes:

[0055] An input sequence construction module, configured to perform optical flow analysis on the obtained ultrasonic image sequence, screen out key frame segments based on the frame-interval optical flow intensity difference, and construct an input sequence by extending buffer frames forward and backward.

[0056] A mask output module, configured to intercept local segments on the input sequence using an overlapping sliding window, calculate the forward and backward optical flow of adjacent frames within the local segments and perform bidirectional alignment, and perform direction-specific multi-head attention calculation on the aligned frame group to output a time-varying probability mask.

[0057] The cross-layer mask conduction modulation module is used to inject a time-varying probability mask into the first layer of the spatio-temporal separated attention network for temporal attention calculation, and in subsequent predetermined layers, a gating unit is used to perform cross-layer mask conduction modulation on the mask features of the first layer and the current layer features;

[0058] The manipulation recognition module splices the gated modulation features output from each layer and inputs them into a cascaded spatio-temporal pyramid pooling module to extract time-sharing granularity motion features. After cross-scale channel attention fusion, the features are input into a multi-expert classifier to output the acupuncture manipulation recognition result.

[0059] In a third aspect, an embodiment of the present invention provides an acupuncture manipulation recognition device based on ultrasonic dynamic optical flow and self-supervised mask, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the acupuncture manipulation recognition method based on ultrasonic dynamic optical flow and self-supervised mask as described above.

[0060] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which computer-executable instructions are stored, and when the executable instructions are executed by a processor, the acupuncture manipulation recognition method based on ultrasonic dynamic optical flow and self-supervised mask as described above is implemented.

[0061] (III) Beneficial effects

[0062] The beneficial effects of the present invention are as follows: By constructing a multi-level cascaded intelligent analysis framework, the present invention systematically solves the technical bottlenecks of insufficient characterization of subcutaneous tissue dynamic features, high artificial dependence, and spatio-temporal model fragmentation in the prior art. First, based on the adaptive screening mechanism of optical flow intensity difference and the buffer frame extension strategy, an input sequence with temporal coherence is constructed, effectively overcoming the problem of feature distortion caused by probe displacement and tissue movement breakage; second, through sliding window interception and bidirectional optical flow alignment operations, high-precision motion compensation is performed within local spatio-temporal segments, significantly improving the biological rationality of spatio-temporal feature alignment; at the same time, the direction-specific multi-head attention mechanism is innovatively introduced to achieve decoupled analysis of motion directions during the generation of time-varying probability masks, breaking through the limitations of traditional methods in parsing complex motion patterns; further, a cross-layer gating conduction modulation strategy is adopted to maintain the dynamic coupling of motion features and semantic features in the deep network, solving the problem of feature decoupling caused by deepening of the existing model; finally, through the synergistic effect of cascaded spatio-temporal pooling and multi-expert decision-making mechanism, the organic fusion of cross-scale features from micro myofibril fibrillation to macro biomechanical effects is realized, significantly improving the recognition robustness of complex manipulation action patterns. While reducing the degree of manual intervention, the present invention establishes the first end-to-end recognition system integrating the dynamic characteristics of optical flow, tissue deformation response, and spatio-temporal coupling effect. Brief Description of the Drawings

[0063] Figure 1 It is a schematic flowchart of the method provided by the embodiment of the present invention;

[0064] Figure 2 It is a specific schematic flowchart of step S1 of the method provided by the embodiment of the present invention;

[0065] Figure 3 It is a data acquisition flowchart of the method provided by the embodiment of the present invention;

[0066] Figure 4 It is a schematic diagram of the original data cropping of the method provided by the embodiment of the present invention;

[0067] Figure 5 It is a schematic diagram of the size reshaping of the method provided by the embodiment of the present invention;

[0068] Figure 6 It is a specific schematic flowchart of step S12 of the method provided by the embodiment of the present invention;

[0069] Figure 7 It is a specific schematic flowchart of step S2 of the method provided by the embodiment of the present invention;

[0070] Figure 8 It is a specific schematic flowchart of step S3 of the method provided by the embodiment of the present invention;

[0071] Figure 9 It is a comparison diagram of the effect of the mask-guided attention calculation of the method provided by the embodiment of the present invention;

[0072] Figure 10 It is a schematic diagram of the cross-layer gating mechanism of the method provided by the embodiment of the present invention;

[0073] Figure 11 It is a specific schematic flowchart of step S4 of the method provided by the embodiment of the present invention;

[0074] Figure 12 It is a schematic diagram of the overall process of the method provided by the embodiment of the present invention. Detailed Description of the Embodiment

[0075] In order to better explain the present invention for easy understanding, the present invention will be described in detail below with reference to the drawings through specific embodiments.

[0076] As Figure 1As shown, a method for identifying acupuncture manipulation based on ultrasonic dynamic optical flow and self-supervised mask proposed in an embodiment of the present invention includes: performing optical flow analysis on the obtained ultrasonic image sequence, screening out key frame segments based on the difference in inter-frame optical flow intensity, and constructing an input sequence by expanding buffer frames forward and backward; intercepting local segments on the input sequence using an overlapping sliding window, calculating the forward and backward optical flow of adjacent frames within the local segment and performing bidirectional alignment, and performing direction-specific multi-head attention calculation on the aligned frame group to output a time-varying probability mask; injecting the time-varying probability mask into the first layer of a spatio-temporal separated attention network for temporal attention calculation, and using a gated unit in subsequent predetermined levels to perform cross-layer mask conduction modulation on the mask feature of the first layer and the current layer feature; splicing the gated modulation features output by each level and inputting them into a cascaded spatio-temporal pyramid pooling module to extract time-sharing granularity motion features, and inputting them into a multi-expert classifier after cross-scale channel attention fusion to output the identification result of acupuncture manipulation.

[0077] The present invention systematically solves the technical bottlenecks of insufficient characterization of subcutaneous tissue dynamic features, high artificial dependence, and spatio-temporal model fragmentation in the prior art by constructing a multi-level cascaded intelligent analysis framework. First, an adaptive screening mechanism based on the difference in optical flow intensity and a buffer frame expansion strategy are used to construct an input sequence with temporal coherence, effectively overcoming the problem of feature distortion caused by probe displacement and tissue motion breakage; second, through sliding window interception and bidirectional optical flow alignment operations, high-precision motion compensation is performed within local spatio-temporal segments, significantly improving the biological rationality of spatio-temporal feature alignment; at the same time, a direction-specific multi-head attention mechanism is innovatively introduced to achieve decoupled analysis of motion directions during the generation of the time-varying probability mask, breaking through the limitation of the traditional method's ability to analyze complex motion patterns; further, a cross-layer gated conduction modulation strategy is adopted to maintain the dynamic coupling of motion features and semantic features in the deep network, solving the problem of feature decoupling caused by the deepening of the existing model; finally, through the synergistic effect of the cascaded spatio-temporal pooling and multi-expert decision-making mechanism, an organic fusion of cross-scale features from microscopic myofibril fibrillation to macroscopic biomechanical effects is achieved, significantly improving the recognition robustness of complex manipulation action patterns. The present invention establishes the first end-to-end recognition system that integrates the dynamic characteristics of optical flow, tissue deformation response, and spatio-temporal coupling effect while reducing the degree of manual intervention.

[0078] To better understand the above technical solution, the exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a clearer and more thorough understanding of the present invention and to be able to convey the scope of the present invention completely to those skilled in the art.

[0079] Specifically, the embodiment of the present invention provides an acupuncture manipulation recognition method based on ultrasonic dynamic optical flow and self-supervised mask, including:

[0080] S1. Perform optical flow analysis on the obtained ultrasonic image sequence, screen out key frame segments based on the inter-frame optical flow intensity difference, and construct an input sequence resistant to motion breaks by extending buffer frames forward and backward. Key frame extraction can effectively screen out data of a certain time period containing key information to reduce information redundancy. This embodiment proposes a method for extracting key frames of a motion sequence (Motion sequence keyframe extraction, MSKE). While not losing the key scene information of the changes in muscle tissue affected by acupuncture in the image, continuous frame segments are obtained to form an input sequence, ensuring that the subsequent spatio-temporal features extracted can effectively reflect the stimulation effect on the subcutaneous muscle tissue during continuous operation of each type of manipulation in the acupuncture process.

[0081] Further, as Figure 2 shown, step S1 includes:

[0082] S11. Obtain and standardize the original ultrasonic video data, calculate the optical flow field of each frame for the obtained ultrasonic image sequence, and generate a visual inter-frame optical flow image sequence encoded in HSV color.

[0083] To obtain the original ultrasonic video data of the acupuncture motion area, 6 senior acupuncture physicians, 2 ultrasonic acquisition physicians from 4 medical institutions, and 53 healthy subjects were invited to participate in the ultrasonic data acquisition work.

[0084] In the specific acquisition work, referring to Figure 3 in which (a) shows the actual operation scene of data acquisition, Figure 3 in which (b) is a schematic diagram of data acquisition. It is required that the acupuncture physician insert the needle at the Kongzui acupoint (LU6) of the subject and slightly lift and insert the needle body. At the same time, the ultrasonic physician moves the probe close to the needle insertion point to find the muscle movement center affected by acupuncture. After the ultrasonic physician determines the probe position, the acupuncture physician operates the four types of manipulations of twirling reinforcing, twirling reducing, lifting-thrusting reinforcing, and lifting-thrusting reducing in sequence according to the recorder's prompt. Each type of manipulation is performed for 4 - 7 s, and the ultrasonic physician records the images in real time. After each type of manipulation is completed, the needle is withdrawn and the subject's muscles are allowed to relax for 3 - 5 min. Before operating the next manipulation, use ultrasound to observe whether the subject's muscles have returned to the resting state, and the interval time can be appropriately extended according to the specific situation. Figure 3 in which (c) shows a set of finally acquired ultrasonic video frame sequence data. The actually screened and recorded images with a duration of 4 s or more are saved and exported in the formats of.avi and.dcm respectively.

[0085] Based on the above operation steps, a real-time ultrasound image dataset of four acupuncture manipulation methods, namely Reinforcing by twisting and rotating (RFTR), Reducing by twisting and rotating (RDTR), Reinforcing by lifting and thrusting (RFLT), and Reducing by lifting and thrusting (RDLT), acting on subcutaneous muscle tissue was collected and constructed. Immediately following, the following data standardization processes were carried out:

[0086] (1) Original data cropping: From the collected original ultrasound video dataset, randomly select one video data and intercept video frames as shown in (a) of Figure 4 . In the picture of the original ultrasound video, there are border and text information, which will interfere with subsequent image analysis. Therefore, use the ScreenToGif software to intercept the image part without border and text information from the original video data. The window size is 855×680 (pixels) to obtain new video data, as shown in (b) of Figure 4 .

[0087] (2) Dimension reshaping: In order to conveniently meet the input data dimension requirements of most deep networks and appropriately reduce the scale of data tensors during model training, based on the video data obtained in step 1, perform dimension reshaping operations on all video data, and uniformly convert the 855×680 (pixel) video data into 512×512 (pixel) video frame data frame by frame as shown in Figure 5 .

[0088] (3) Standardization processing: Restore the frame sequence processed in (2) to.avi format video data and unify the video data parameters. Since there are manual errors when ultrasound acquisition physicians manually record data videos, it is difficult to ensure that the durations are exactly the same. Therefore, uniformly intercept the middle 4s video data from each ultrasound video with a duration greater than or equal to 4s, and unify the frame rate of the video to 30 frames / s (restore to the frame rate of the original ultrasound image video). It should be emphasized that each video data undergoes a unified standard preprocessing process to obtain a unique independent sample, and there is no situation where multiple samples are intercepted repeatedly from a single original image data.

[0089] In a specific embodiment, for any collected ultrasound image sequence I wherein, it is denoted as:

[0090] ;

[0091] wherein, Irepresents the entire video sequence; I t represents t the video frame corresponding to the moment, while I T represents the sequence I of the last video frame; R represents the variable dimension, where I t is a three-dimensional matrix variable with dimensions of H × W × C (height × width × number of channels), I is a 4D tensor with dimensions of T × H × W × C (duration × height × width × number of channels). The ultrasonic video sequence I passes through the built-in flownet2 optical flow estimation network in the mmflow library to obtain the inter-frame optical flow map (Flow map) of each video. If the inter-frame optical flow sequence data of all data is saved in the.flo format file, the subsequent call of the optical flow data will occupy a large amount of memory, and the current extraction stage does not require providing very accurate optical flow information, only providing a screening basis for the subsequent steps. Therefore, it is selected to represent with the visual optical flow image (Optical flow visualization). The ultrasonic video inter-frame optical flow sequence is O wherein, denoted as:

[0092] ;

[0093] wherein, is obtained by calculating the optical flow between two consecutive frames I t and I t+1 ; O adopts the HSV color encoding, where the hue (Hue, H) represents the motion direction, and the saturation (Saturation, S) and value (Value, V) represent the motion amplitude; R represents the variable dimension, where O t,t+1 is a three-dimensional matrix variable with dimensions of H × W × C (height × width × number of channels), O is a 4D tensor with dimensions of T × H × W × C (duration × height × width × number of channels). Taking OTaking one frame as an example, the pixel-level encoding in the HSV color space is as follows:

[0094] ;

[0095] Among them, ( x , y ) represents the position of the pixel in the optical flow image. H t,t-1 ( x , y ), S t,t-1 (x, y ) and V t,t-1 (x, y ) respectively characterize the motion information of this pixel point.

[0096] S12. Convert the visualized inter-frame optical flow image sequence to the HSV color space, perform independent normalization processing on each channel, calculate the global optical flow intensity of each frame, and input it into a Gaussian difference filter with temporal awareness to extract significant time points where the difference value of the filtered optical flow intensity exceeds three times the average fluctuation level of adjacent frames.

[0097] Furthermore, as Figure 6 shown, step S11 includes:

[0098] S121. Convert the visualized inter-frame optical flow image sequence to the HSV color space, and perform independent normalization processing on the hue channel, saturation channel, and brightness channel of the HSV color space.

[0099] S122. Perform pixel-by-pixel multiplication and accumulation on the normalized values of the hue channel, saturation channel, and brightness channel to obtain the global optical flow intensity representing the dynamic change intensity of muscle tissue for each optical flow image, and further obtain the global optical flow intensity sequence.

[0100] S123. Use the first sliding time window to perform sliding average filtering on the global optical flow intensity sequence to generate a baseline trend component, subtract the global optical flow intensity sequence from the baseline component to obtain a residual sequence.

[0101] S124. Apply the second sliding time window to the residual sequence, and generate an enhanced residual sequence by obtaining the second derivative response of the residual signal within the second sliding time window through a Gaussian difference filter.

[0102] Specifically, this step adopts a first sliding time window with a time window length set to 0.5 - 1.5 seconds, corresponding to 20 - 38 frames, which mainly suppresses low-frequency physiological interferences such as respiration (0.2 - 0.3 Hz) and vascular pulsation (1 - 1.5 Hz), while retaining the medium and high-frequency motion components related to acupuncture operations. On the residual sequence after baseline removal, a second sliding time window is applied, and the time window length is matched with the minimum effective stimulation duration of clinically verified acupuncture manipulations (50 - 150 ms), corresponding to 3 - 9 frames, which is used to extract the transient tissue response characteristics triggered by operations such as lifting-thrusting and twirling.

[0103] S125. In the enhanced residual sequence, based on the standard deviation of the optical flow intensity fluctuations of N frames before and after the current frame as the center, if the instantaneous difference value exceeds three times the standard deviation, it is marked as a significant time point.

[0104] In a specific embodiment, for each group of optical flow sequences O optical flow intensity I is calculated. The given optical flow frame is converted to the HSV color space and the H, S, and V channels are normalized so that their values are all within the range of [0, 1]. O t,t+1 The calculation formula for the optical flow intensity I of

[0105] ;

[0106] where H norm (.) is the pixel hue normalization processing function; S norm (.) is the pixel saturation normalization processing function; V norm (.) is the pixel brightness normalization processing function. Thus, the optical flow intensity I of each optical flow image O t,t+1 can be obtained t,t+1 . The change in the optical flow intensity I reflects the intensity of muscle tissue movement under acupuncture, and reflects the effectiveness of the manipulation on the subcutaneous layer.

[0107] Next, taking the sampling moment as the sampling point, first custom parameters n 、 N and L are given. These custom sampling parameters are used to control the size of the model input sequence, and the specific definitions and calculation formulas are shown in Table 1 Parameter Definition Table below.

[0108] Table 1 Parameter Definition Table

[0109]

[0110] where Nis the total number of frames in the input sequence of the subsequent classification network. Common values are integers such as 8, 16, 32, etc.; n is the total number of sampling points for each segment of ultrasonic image data, which can be custom-adjusted according to the input sequence length N The required sampling frequency. The larger the value, the higher the sampling frequency; L is the number of local consecutive frames at each sampling point, used to ensure the continuity of data at the sampling point and guarantee the true and complete temporal information of the movement of muscle bundles in the acupuncture action area.

[0111] To select the sampling points, first calculate the change in optical flow intensity of the inter-frame optical flow sequence , and the formula is as follows:

[0112] ;

[0113] where I t+1,t+2 and I t,t+1 are the optical flow intensities corresponding to the inter-frame optical flow O t+1,t+2 and O t,t+1 respectively. reflects the moments when there are significant changes in the dynamic information in the image and can reflect important movement information. Therefore, according to the magnitude of , the first n largest are selected and arranged as follows:

[0114] ;

[0115] At this time, the moments corresponding to are the sampling points. Re-sort the sampling points according to the true time order as follows: n ;

[0116] ;

[0117] In this way, it can be ensured that both the sampling order and the composed key frame input sequence conform to the true time order later, and there will be no temporal chaos where the later video frames in the original video appear before the earlier video frames.

[0118] After determining the n sampling points, each sampling point corresponds to two frames of the original image sequence. For any sampling point, obtain L consecutive video frames to form a continuous key frame segment segment i :

[0119] ;

[0120] S13. On the ultrasound image sequence, a continuous region centered on the significant time point and extending to both sides of the time series until the change rate of the optical flow intensity drops to the baseline level is used as the key frame segment.

[0121] S14. Based on the key frame segment, adaptively expand the buffer frames bidirectionally forward and backward to obtain the initial input sequence.

[0122] It should be understood that the core role of the buffer frame is to prevent subsequent mask weights from jumping through spatio-temporal continuity constraints. Specifically, the bidirectional adaptive expansion of the buffer frame forward and backward includes the following steps:

[0123] (1) Forward expansion: Intercept an optical flow compensation frame group including at least two consecutive frames before the key frame segment. The optical flow compensation frame group is used to reconstruct the initial momentum accumulation process of the acupuncture force. When the time interval between the starting point of the key frame segment and the ending point of the previous key frame segment is less than the preset merging threshold (such as 50 milliseconds), the optical flow compensation frame group is incorporated into the backward expansion sequence of the previous key frame segment.

[0124] (2) Backward expansion: Intercept a motion attenuation observation frame group including at least three consecutive frames after the key frame segment. The motion attenuation observation frame group covers the complete attenuation cycle of the elastic rebound of muscle fibers. The termination condition for intercepting the motion attenuation observation frame group is that the amplitude of the inter-frame optical flow decays to below the preset amplitude threshold (such as 20%) of the amplitude of the central frame of the key frame segment.

[0125] (3) When the starting position of the key frame segment is at the starting point of the ultrasound image sequence, perform spatial mirror flipping on the optical flow field of the first frame along the main axis of the muscle fibers of the corresponding ultrasound image, and superimpose a random perturbation vector with an amplitude of 15%-25% of the standard deviation of the optical flow of the current frame to generate a virtual forward compensation frame.

[0126] (4) When the ending position of the key frame segment is at the ending point of the ultrasound image sequence, extrapolate the motion trajectory according to 1.2-1.8 times the duration of the optical flow vector of the last frame to generate a virtual backward observation frame; among them, the timestamps of the inserted virtual forward compensation frame or virtual backward observation frame are evenly distributed with the adjacent frames, and the difference in the motion direction between the adjacent virtual frames and the real frames is not greater than 15 degrees.

[0127] S15. Perform optical flow gradient consistency detection on the initial input sequence. If the sudden change in the motion direction between adjacent frames exceeds the preset mutation angle, insert an interpolation transition frame, and finally output an input sequence resistant to motion breakage.

[0128] S2. On the input sequence, use an overlapping sliding window to intercept local segments, calculate the forward and backward optical flow between adjacent frames in the local segments and perform bidirectional alignment, and perform direction-specific multi-head attention calculation on the aligned frame group to output a time-varying probability mask.

[0129] In this step, overlapping sliding windows are used to intercept local segments on the input sequence. The forward and backward optical flow fields of adjacent frames within the local segments are calculated through an optical flow estimation network, and pixel-level spatial deformation alignment is performed on the front and back frames based on the bidirectional optical flow to eliminate spatio-temporal misalignment caused by probe displacement or tissue elastic deformation. A direction-specific multi-head attention mechanism is deployed on the aligned frame group, and independent attention head groups are designed respectively for horizontal motion (such as transverse contraction of muscle fibers), vertical motion (such as longitudinal deformation of subcutaneous tissue), and composite motion (such as rotational shear). Through the motion direction threshold constraint and optical flow amplitude weighting mechanism, the feature responses of motion-sensitive regions in different directions are dynamically enhanced. Finally, a time-varying probability mask is output, and its pixel values represent the motion saliency of each region under the action of the manipulation method.

[0130] Further, as Figure 7 shown, step S2 includes:

[0131] S21. Intercept local segments containing at least one key frame segment, at least two forward buffer frames, and at least three backward buffer frames on the input sequence with a sliding window of adaptive overlap rate.

[0132] S22. Perform forward and backward optical flow estimation on each adjacent frame of each local segment using a lightweight optical flow network. Align the previous frame to the current frame coordinate space through spatial deformation mapping based on the forward optical flow, and align the subsequent frame to the current frame coordinate space through reverse deformation based on the backward optical flow. Concatenate the three aligned frames along the channel dimension to generate a spatio-temporally aligned superimposed frame group.

[0133] In a specific embodiment, for any continuous local segment with a frame sequence , the sliding window t at time I t uses the current frame as the query I t-1 , I t , I t+1 as the keys k t-1 , , k t+1 . In order to more accurately locate the motion area and explicitly model the temporal motion information, the optical flow alignment method is used to align the key k t-1 of the previous frame with the key k t+1 of the subsequent frame to , that is, align the previous frame and the subsequent frame to the current frame, and minimize the spatio-temporal feature misalignment caused by probe jitter, displacement, etc.

[0134] Specifically, the lightweight SPyNet model is introduced to calculate the inter-frame optical flow. SpyNet(.) represents calculating the inter-frame optical flow using the SPyNet model. The optical flow calculation formula is as follows:

[0135] ;

[0136] where flow forward : , represents the forward optical flow from frame I t-1 to frame I t . R represents the variable dimension. Here f t-1,t is a four-dimensional matrix variable with a dimension of B ×2× H × W (batch size × number of channels (2) × height × width); flow backward : , represents the backward optical flow from frame I t+1 to frame I t . R represents the variable dimension. Here f t+1 , t is a four-dimensional matrix with a variable dimension of B ×2× H × W (batch size × number of channels (2) × height × width).

[0137] For a continuous key-frame segment

[0138] segment i the following optical flow set is obtained:

[0139] ;

[0140] where F forward and F backward are the forward optical flow set and the backward optical flow set of the key-frame segment segment i respectively; for each continuous key-frame segment segment i there are L' frames ( segment i the previous frame and the next frame of the multi-retained segment, L' =L+(2), when calculating the forward and backward optical flow information segment i The first and last frames of L only provide a reference for the optical flow calculation of the middle frames of the segment i forward optical flow set F forward and the backward optical flow set F backward When constructing, the reserved continuous frame time range is [2, L' -1].

[0141] Subsequently, the forward and backward optical flow information is respectively used to align the keys k t-1 , k t+1 .

[0142] The key of the previous frame k t-1 is aligned to the key of the current frame k t to obtain the aligned key k ' t-1 :

[0143] ;

[0144] The key of the next frame k t+1 is aligned to the key of the current frame k t to obtain the aligned key k ' t+1 :

[0145] ;

[0146] After optical flow alignment, the keys are concatenated as follows to obtain the keys k and values v of the superimposed frame group:

[0147] ;

[0148] S23. Parallelly input the horizontally moving attention head group, vertically moving attention head group, and composite motion analysis head group in the multi-head attention module into the spatio-temporally aligned superimposed frame group for feature processing;

[0149] Specifically, the horizontal motion focusing head group is configured as a motion analysis module based on the horizontal optical flow gradient constraint. By enhancing the attention weight of the horizontal motion component, it specifically extracts the motion features related to the contraction motion of muscle fibers in the horizontal direction. The horizontal optical flow gradient constraint is as follows: when calculating the query-key similarity matrix, a dynamic weight gain of 3-5 times is applied to the pixel points where the horizontal optical flow gradient component exceeds a preset ratio threshold (more than 2 times) of the vertical optical flow gradient component.

[0150] The longitudinal motion focusing head group is configured as a motion analysis module based on the vertical optical flow gradient constraint. By enhancing the feature response intensity of the vertical optical flow gradient component, it specifically extracts the motion features related to the vertical deformation of the subcutaneous tissue caused by the lifting and inserting operations of acupuncture needles. The vertical optical flow gradient constraint is as follows: when calculating the attention weight, for the region where the vertical optical flow gradient component exceeds the preset ratio threshold (more than 2 times) of the horizontal optical flow gradient component, that is, the region where the vertical optical flow gradient component dominates, exponential feature enhancement in the channel dimension is implemented.

[0151] The composite motion analysis head group is configured as a full-degree-of-freedom motion analysis module without direction constraints, which is used to capture rotational motion, oblique shear motion, and multi-directional composite motion patterns, and model the non-linear motion features of the multi-dimensional motion vector field through the spatio-temporal joint attention mechanism of aligning frame groups.

[0152] S24. Orthogonally project and fuse the output features of the horizontal motion focusing head group and the longitudinal focusing head group to generate a feature tensor representing the joint distribution of motion vectors.

[0153] S25. Concatenate the output features of the composite analysis head group with the feature tensor representing the joint distribution of motion vectors along the channel axis to generate cross-motion mode fusion features.

[0154] S26. Perform spatial pyramid pooling operations on the cross-motion mode fusion features, extract multi-granularity motion features at different scales respectively, and use channel attention gating to weight and fuse the features at each scale, retain the cross-dimensional coupling features representing muscle fiber contraction in the horizontal direction, subcutaneous deformation in the vertical direction, and directionless motion, perform convolution and non-linear activation on the coupling features, and output an initial mask.

[0155] S27. Perform cross-window fusion on the overlapping regions of adjacent initial masks based on the optical flow vectors in adjacent sliding windows, perform morphological closing operations on the fused masks to eliminate holes and isolated noises, and refine the mask boundaries through sub-pixel optical flow-guided interpolation to output a time-varying probability mask with continuous boundaries for representing the motion probability distribution of the subcutaneous tissue. The continuously transitioning time-varying probability mask represents the motion probability distribution of the subcutaneous tissue, and its high-response region is consistent with the dynamic deformation region of muscle bundles under acupuncture.

[0156] S3. Inject the time-varying probability mask into the first layer of the spatio-temporal separation attention network for temporal attention calculation, and use a gated unit in subsequent predetermined layers to perform cross-layer mask conduction modulation on the mask features of the first layer and the features of the current layer.

[0157] Further, as Figure 8 shown, step S3 includes:

[0158] S31. Input the time-varying probability mask into the Patch embedding module synchronized with the spatio-temporal separation attention network, and perform non-overlapping block processing on the spatial region of the mask through a two-dimensional convolution operation with a preset block size and stride to generate mask embedding blocks.

[0159] S32. Stack multiple mask embedding blocks sequentially along the time dimension, and perform an expansion transformation on the channel dimension through a learnable linear projection layer to generate a mask feature tensor with the same dimension as the input tensor of the spatio-temporal attention network. Among them, the learnable linear projection layer is a parameterized linear transformation layer, whose role is to map the input features from the original dimension to the target dimension, and the transformation matrix (weight) in the mapping process is automatically optimized through the backpropagation algorithm.

[0160] S33. In the temporal attention calculation stage of the first layer of the spatio-temporal separation attention network, add the mask feature tensor and the original query tensor element by element, and then perform layer normalization operation to make the temporal attention weights tilt towards the motion probability region of the mask annotation that exceeds the preset probability threshold (0.7).

[0161] S34. Through the gated recalibration units set in the 2nd, 4th, and 6th layers, perform cross-layer transfer on the mask feature tensor after dimensionality reduction by depthwise separable convolution, and input the query tensor of the current layer into a two-level fully connected network to generate gated weights.

[0162] S35. Dynamically scale the dimensionality-reduced mask features using the gated weights, and stack them to the current layer features through residual connection to obtain gated modulation features containing mask guidance information.

[0163] S36. Through an auxiliary supervision branch deployed in front of the final classification layer, calculate the mutual information entropy between the deep features and the mask feature tensor of the first layer. If the entropy value is lower than the set entropy threshold, trigger the backtracking adjustment of the gated weights. Specifically, the backtracking adjustment includes: freezing the parameters of the high-level network, and backpropagating to adjust the weights of the fully connected layer of the gated unit; if the entropy threshold is not reached for 3 consecutive iterations, close the mask conduction path of the current channel.

[0164] In yet another specific embodiment, TimeSformer with Divided space-time attention (Div-ST) is selected as the basic structure of the backbone network. First, a Patch Embedding module consistent with the time and space encoding of TimeSformer is used to convert the SSAD ROI mask M into the tensor form applicable to the input of the Transformer module in the original network. The following specific operations are performed: The space of a single-frame mask is non-overlappingly partitioned using a two-dimensional convolution operation with a kernel size of P×P and a stride of P. Based on the original resolution H×W of the ultrasound image frame, the mask with a size of H×W is divided into NUM = (H / P) × (W / P) mask embedding blocks. The spatial dimension of each embedding block is P×P, and after being unfolded (Flattened) in the channel dimension by a learnable linear projection layer, it generates C × P ² feature vectors, where C is the number of mask channels.

[0165] Finally, the SSAD ROI mask M is transformed as follows:

[0166] ;

[0167] where PatchEmbed(.) represents Patch embedding processing, and the dimension of the processed SSAD ROI mask M is also adjusted to the tensor form applicable to the input of the Transformer module; R represents the variable dimension. Here, M ' is a three-dimensional matrix variable with a dimension of B × NUM ×( C × P ²). The main input feature after Patch embedding processing is defined as X , X and in the temporal attention calculation of the first-layer Transformer module, the corresponding query tensor q x is obtained. The SSAD ROI mask M' is added to q x and layer normalization is performed as follows:

[0168] ;

[0169] Thus, based on the temporal attention of the basic backbone network Div-ST TimeSformer, the implementation based on the SSAD ROI mask MGuide the spatio-temporal attention calculation in the first-layer Transformer module to enhance the model's self-supervised and adaptive attention to the muscle movement characteristics around the acupuncture point area affected by the acupuncture manipulation. Figure 9 Shows whether to add the SSADROI mask M Comparison chart of the effect of guiding attention calculation.

[0170] Figure 9 The wireframe in part a: shows the original spatio-temporal attention calculation effect of the backbone network Div-ST TimeSformer. It can be seen that the temporal attention part performs cross-attention calculation on the pixels at the same spatial position in different time steps. For t each pixel at each spatial position at a moment, it performs cross-attention calculation with the global spatial pixels at the current moment to extract temporal global features.

[0171] Figure 9 The wireframe in part b: shows the temporal attention calculation effect guided by the SSAD ROI mask M That is, the optimized temporal attention part not only performs cross-attention calculation on the pixels at the same spatial position in different time frames, but also assigns higher weights to the ROI areas containing image motion information according to the guidance of the SSAD ROI mask M and reduces the attention to relatively static background information, while also avoiding the neglect of fine dynamic features by the binary mask, thus realizing more accurate temporal local attention allocation. Combining with the original global spatial attention mechanism, it can better highlight the feature information of the manipulation area, enabling the model to have a more accurate understanding of the subcutaneous effects of different acupuncture manipulations.

[0172] It should be noted that the deep feature extraction part of Div-ST TimeSformer is stacked by multiple Transformer modules with spatio-temporal attention separation, and the information of the SSAD ROI mask is only introduced in the first layer (layer 0) M It is easy to cause the loss of the guiding information of the SSAD ROI mask for the acupuncture area gradually as the network deepens layer by layer M Therefore, based on the principle of the gating mechanism, a simple gating module is designed to introduce the information of the SSADROI mask across layers without introducing too much computational complexity and significantly increasing the network complexity. M The specific structure is as follows Figure 10 as shown.

[0173] As Figure 10 can be seen, the SSADROI mask is introduced in the temporal attention calculation of the first-layer (layer 0) Transformer moduleM' , as shown in the formula, the corresponding query tensor in the temporal attention calculation is obtained q' x , similarly, in the subsequent i layer of the Transformer module, the SSAD ROI mask is also introduced M' . A gating module is set up to perform weighted adjustment on the query tensor i of the temporal attention in the q xi layer and the SSAD ROI mask M' . The gating mechanism will adaptively adjust the retention degree of the SSAD ROI mask M' according to the features of this layer. First, the gating weight g t is set as a tensor with the same shape as the current query tensor q xi , where g t is obtained through the calculation of the gating layer, and the formula is as follows:

[0174] ;

[0175] where FC1(.) is the linear transformation process of the first fully connected layer, FC2(.) is the linear transformation process of the second fully connected layer. It should be noted that the RELU activation function is enabled between the two fully connected layers. And represents the mutual information entropy calculated between the output features of the i layer (where the SSAD ROI mask M ' is introduced by the gating mechanism) and the output features of the first layer (layer 0). If it is less than the threshold τ , the gating weight backtracking adjustment is triggered, and ▽ g t (.) is the gating weight gradient, which adjusts the gating weight during the optimization process, λ is a learnable parameter, g t ' (.) is the optimized gating weight.

[0176] Finally, for the q xi introduced SSAD ROI mask M' , the information is as follows:

[0177] ;

[0178] The gating weight g t will adaptively retain the SSAD ROI mask introduced by the underlying layer according to the calculation results of the i layerM Regarding the influence, choosing to introduce across layers instead of layer by layer is to appropriately control the coupling degree of the underlying motion features to the deep semantic features, so that the model will neither lose a large amount of attention to the acupuncture area, nor be overly interfered by the underlying features in the extraction of deep semantic features.

[0179] It is experimentally found that when stacking 12 layers of Transformer modules with spatio-temporal attention separation and introducing the gating mechanism across layers at the 2nd, 4th, and 6th layers (i.e., relatively shallow layers), the model can achieve the best recognition effect.

[0180] S4. Concatenate the gated modulation features output from each level and input them into the cascaded spatio-temporal pyramid pooling module to extract the motion features at different time granularities. After cross-scale channel attention fusion, input them into the multi-expert classifier to output the recognition results of acupuncture manipulation techniques.

[0181] Furthermore, as Figure 11 shown, step S4 includes:

[0182] S41. Concatenate the gated modulation features output from each layer along the channel dimension to obtain a fusion feature tensor containing shallow motion details and deep semantic associations, and perform batch normalization and spatio-temporal dimension reorganization to match the input structure of the cascaded spatio-temporal pyramid.

[0183] In this step, batch normalization is to perform mean-variance normalization in the batch dimension to eliminate the feature distribution shift caused by the difference in optical flow intensity of different video samples, and spatio-temporal dimension reorganization is to reorganize the dimension of the feature tensor from [batch, time, space, channel] to [batch, channel, time×space] to match the input requirements of the cascaded spatio-temporal pyramid module.

[0184] S42. Input the fusion feature tensor after batch normalization and spatio-temporal dimension reorganization into the spatio-temporal pyramid module, perform dense sliding max pooling within a 4 - 8 frame window to obtain short-term granularity features by extracting the myofibril fibrillation features caused by the instantaneous action of acupuncture, perform sparse self-attention pooling within a 16 - 24 frame window to obtain medium-term granularity features by acquiring the motion rhythm features of a single lifting-thrusting or twirling operation, and perform adaptive average pooling within a 32 - 48 frame window to extract the cumulative biomechanical features under the continuous action of multiple manipulations to obtain long-term granularity features.

[0185] S43. Perform multi-scale alignment and channel attention weighted fusion on the short-term, medium-term, and long-term granularity features to obtain the final fusion feature.

[0186] In this step, multi-scale alignment is to map short-term, medium-term, and long-term features to a unified spatio-temporal coordinate system through deformation convolution guided by optical flow, eliminating the misalignment caused by scale differences; while channel attention weighting is to calculate the channel importance weights for each granularity feature respectively, and a gating mechanism is used to strengthen the transient response channels of short-term features (gain coefficient 2.0 - 3.5), the rhythm marking channels of medium-term features (gain 1.5 - 2.0), and the cumulative effect channels of long-term features (gain 1.0 - 1.2); based on the dynamic routing algorithm to determine the cross-scale fusion priority, the short-term:medium-term:long-term weight ratio is initially set to 4:3:3.

[0187] S44. Input the finally fused features into the parallel optical flow motion expert, tissue deformation expert, and spatio-temporal coupling expert, and perform weighted voting based on the confidence scores output by each expert to output the probability distribution of acupuncture manipulation categories.

[0188] Specifically, the motion pattern expert is used to analyze the time-frequency characteristics (main frequency, harmonic energy ratio) of the optical flow vector field, the tissue deformation expert is used to analyze the texture strain energy distribution and the time-varying curve of the shear strain rate, and the spatio-temporal coupling expert is used to evaluate the motion-deformation phase difference and the energy transfer efficiency. Each expert independently outputs the probability distribution of the manipulation category, and performs weighted voting based on the real-time confidence scores (40% for the optical flow motion expert, 30% for the tissue deformation, and 30% for the spatio-temporal coupling); when the confidence difference of the highest-voted category < 15%, trigger sub-pixel level feature reprojection verification. Finally, a result including the manipulation category label (such as lifting-thrusting / rotating / composite) and the confidence score (0 - 1) is generated.

[0189] In addition, an acupuncture manipulation recognition system based on ultrasonic dynamic optical flow and self-supervised mask provided by an embodiment of the present invention includes:

[0190] An input sequence construction module, configured to perform optical flow analysis on the obtained ultrasonic image sequence, screen out key frame segments based on the inter-frame optical flow intensity difference, and construct an input sequence resistant to motion fracture by extending buffer frames forward and backward.

[0191] A mask output module, configured to intercept local segments on the input sequence by using an overlapping sliding window, calculate the forward and backward optical flow of adjacent frames within the local segments and perform bidirectional alignment, and perform direction-specific multi-head attention calculation on the aligned frame group to output a time-varying probability mask.

[0192] A cross-layer mask conduction modulation module, configured to inject the time-varying probability mask into the first layer of the spatio-temporal separation attention network for temporal attention calculation, and use a gating unit to perform cross-layer mask conduction modulation on the mask feature of the first layer and the current layer feature in subsequent predetermined layers.

[0193] The manipulation recognition module splices the gated modulation features output at each level and inputs them into a cascaded spatio-temporal pyramid pooling module to extract the motion features at different time granularities. After cross-scale channel attention fusion, the features are input into a multi-expert classifier to output the recognition result of acupuncture manipulations.

[0194] Moreover, an embodiment of the present invention provides an acupuncture manipulation recognition device based on ultrasonic dynamic optical flow and self-supervised masks, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned acupuncture manipulation recognition method based on ultrasonic dynamic optical flow and self-supervised masks.

[0195] Furthermore, an embodiment of the present invention provides a computer-readable storage medium, on which computer-executable instructions are stored, and when the executable instructions are executed by a processor, the above-mentioned acupuncture manipulation recognition method based on ultrasonic dynamic optical flow and self-supervised masks is implemented.

[0196] In summary, an embodiment of the present invention provides an acupuncture manipulation recognition method, system, device and medium based on ultrasonic dynamic optical flow and self-supervised masks. Referring to Figure 12 , its overall process is as follows:

[0197] In the first stage, an ultrasonic video dataset containing four acupuncture manipulations, namely reinforcing twirling, reducing twirling, reinforcing lifting-thrusting, and reducing lifting-thrusting, was first constructed to explore the differences in the stimulation effects of different acupuncture manipulations on subcutaneous tissues. In order to verify the effectiveness and scientificity of the dataset, a pre-experiment analysis based on radiomics was carried out, the texture features based on ultrasonic images were extracted, and the significance analysis of the features was performed. It was found that there were significant differences among the four types of manipulation data collected, and the data samples were representative, providing effective support for subsequent analysis. Then, optical flow analysis was performed on the obtained ultrasonic image sequences, key frame segments were selected based on the inter-frame optical flow intensity differences, and an input sequence resistant to motion breaks was constructed by extending the buffer frames forward and backward.

[0198] In the second stage, overlapping sliding windows are used to intercept local segments on the input sequence. The forward and backward optical flows of adjacent frames within the local segments are obtained and bidirectional alignment is performed. Direction-specific multi-head attention calculation is carried out on the aligned frame groups, and a time-varying probability mask is output. Based on the common Region of Interest (ROI) mask annotation idea in medical image analysis, the present invention designs a self-supervised-adaptive dynamic (SSAD) ROI mask extraction module for the muscle movement area under manipulation. On the one hand, it solves the problem of mask annotation when the target area has no clear contour or no clear boundary range threshold in the analysis of medical image videos. On the other hand, the ROI mask that adaptively adjusts the weight distribution avoids the disadvantage of binary masks missing global information and can better guide the feature extraction network to extract more accurate spatio-temporal motion features in muscle ultrasound images under different acupuncture manipulation techniques.

[0199] In the third stage, the time-varying probability mask is injected into the first layer of the spatio-temporal separation attention network for temporal attention calculation, and a gated unit is used in subsequent predetermined levels to perform cross-layer mask conduction modulation on the mask features of the first layer and the current layer features; the gated modulation features output by each level are concatenated and input into a cascaded spatio-temporal pyramid pooling module to extract time-sharing granularity motion features. After cross-scale channel attention fusion, they are input into a multi-expert classifier to output the acupuncture manipulation recognition result.

[0200] Thus, the present invention proposes an MSKE (Motion sequence keyframe extraction) module and an SSAD (Self-supervised-adaptive dynamic) ROI mask extraction module, which respectively solve the problems of extracting key information from long video sequences and self-supervised-adaptive extraction of dynamic ROI masks, greatly reducing the manual and time costs. And an optimization module and a cross-layer gating mechanism are added to the conventional spatio-temporal attention separation (Div-ST) timeSformer model to achieve spatio-temporal attention feature extraction guided by the SSAD ROI mask, and a manipulation recognition model based on four types of manipulations, namely twirling reinforcing, twirling reducing, lifting-thrusting reinforcing, and lifting-thrusting reducing, is constructed. The accuracy of the proposed optimized recognition model reaches 92.41%. Through experiments, it is verified that the proposed recognition model has good stability and robustness, and the proposed optimization modules all have a certain degree of generality, which also provides an idea and method for medical image analysis.

[0201] Since the system / apparatus described in the above embodiments of the present invention is the system / apparatus adopted for implementing the method of the above embodiments of the present invention, based on the method described in the above embodiments of the present invention, those skilled in the art can understand the specific structure and variations of the system / apparatus, and thus will not be elaborated herein. Any system / apparatus adopted for the method of the above embodiments of the present invention falls within the scope of protection of the present invention.

[0202] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0203] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can also be implemented.

[0204] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications after learning the basic creative concepts. Therefore, the solution of the present invention should be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.

[0205] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalent technologies, the present invention should also include these modifications and variations.

Claims

1. A method for identifying acupuncture manipulation based on ultrasonic dynamic optical flow and self-supervised mask, characterized in that, Comprising: Performing optical flow analysis on the obtained ultrasonic image sequence, screening out key frame segments based on the inter-frame optical flow intensity difference, and constructing an input sequence by expanding the buffer frames forward and backward; Adopting an overlapping sliding window on the input sequence to intercept local segments, calculating the forward and backward optical flows of adjacent frames within the local segments and performing bidirectional alignment, and performing direction-specific multi-head attention calculation on the aligned frame groups to output a time-varying probability mask; Injecting the time-varying probability mask into the first layer of the spatio-temporal separation attention network for temporal attention calculation, and using a gated unit in subsequent predetermined levels to perform cross-layer mask conduction modulation on the mask features of the first layer and the current layer features; Concatenating the gated modulation features output by each level and inputting them into a cascaded spatio-temporal pyramid pooling module to extract time-sharing granularity motion features, and inputting them into a multi-expert classifier after cross-scale channel attention fusion to output the recognition result of acupuncture manipulation; 2. The acupuncture manipulation recognition method based on ultrasonic dynamic optical flow and self-supervised mask according to claim 1, wherein, Performing optical flow analysis on the obtained ultrasonic image sequence, screening out key frame segments based on the inter-frame optical flow intensity difference, and constructing an input sequence by expanding the buffer frames forward and backward includes: Obtaining and standardizing the original ultrasonic video data, calculating the optical flow field of each frame of the obtained ultrasonic image sequence, and generating a visual inter-frame optical flow image sequence encoded in HSV color; Converting the visual inter-frame optical flow image sequence to the HSV color space and performing independent normalization processing on each channel, calculating the global optical flow intensity of each frame, and inputting it into a Gaussian difference filter with temporal awareness to screen out significant time points where the optical flow intensity difference value exceeds three times the average fluctuation level of adjacent frames; On the ultrasonic image sequence, using the significant time points as the center, extending to both sides of the time series until the continuous region where the optical flow intensity change rate drops to the baseline level is used as the key frame segment; Taking the key frame segment as a reference, adaptively expanding the buffer frames bidirectionally forward and backward to obtain an initial input sequence; Performing optical flow gradient consistency detection on the initial input sequence, and inserting interpolation transition frames if the motion direction mutation between adjacent frames exceeds a preset mutation angle, and finally outputting an input sequence resistant to motion breakage; 3. The acupuncture manipulation recognition method based on ultrasonic dynamic optical flow and self-supervised mask according to claim 2, wherein, Converting the visual inter-frame optical flow image sequence to the HSV color space and performing independent normalization processing on each channel, calculating the global optical flow intensity of each frame, and inputting it into a Gaussian difference filter with temporal awareness to screen out significant time points where the optical flow intensity difference value exceeds three times the average fluctuation level of adjacent frames includes: Converting the visual inter-frame optical flow image sequence to the HSV color space, and performing independent normalization processing on the hue channel, saturation channel, and brightness channel of the HSV color space; Performing pixel-by-pixel multiplication accumulation on the normalized values of the hue channel, saturation channel, and brightness channel to obtain the global optical flow intensity representing the dynamic change intensity of muscle tissue for each optical flow image, and further obtaining a global optical flow intensity sequence; Using a first sliding time window to perform sliding average filtering on the global optical flow intensity sequence to generate a baseline trend component, subtracting the global optical flow intensity sequence from the baseline component to obtain a residual sequence; Applying a second sliding time window to the residual sequence, and generating an enhanced residual sequence by obtaining the second derivative response of the residual signal within the second sliding time window through a Gaussian difference filter; In the enhanced residual sequence, based on the standard deviation of the optical flow intensity fluctuations of N frames before and after the current frame as the center, if the instantaneous difference value exceeds three times the standard deviation, it is marked as a significant time point.

4. The acupuncture manipulation recognition method based on ultrasonic dynamic optical flow and self-supervised mask according to claim 2, characterized in that, The adaptive extended buffer frames include: Forward extension intercepts an optical flow compensation frame group including at least two consecutive frames before the key frame segment. When the time interval between the starting point of the key frame segment and the ending point of the previous key frame segment is less than the preset merging threshold, the optical flow compensation frame group is incorporated into the backward extension sequence of the previous key frame segment; Backward extension intercepts a motion attenuation observation frame group including at least three consecutive frames after the key frame segment. The termination condition for intercepting the motion attenuation observation frame group is that the inter-frame optical flow amplitude decays below the preset amplitude threshold of the amplitude of the center frame of the key frame segment; When the starting position of the key frame segment is at the start of the ultrasonic image sequence, perform a spatial mirror flip on the optical flow field of the first frame along the main direction axis of the muscle fibers of the corresponding ultrasonic image, and superimpose a random perturbation vector with an amplitude of 15%-25% of the current frame's optical flow standard deviation to generate a virtual forward compensation frame; When the ending position of the key frame segment is at the end of the ultrasonic image sequence, extrapolate the motion trajectory according to 1.2-1.8 times the duration of the last frame's optical flow vector to generate a virtual backward observation frame.

5. The acupuncture manipulation recognition method based on ultrasonic dynamic optical flow and self-supervised mask according to claim 1, characterized in that, On the input sequence, use an overlapping sliding window to intercept local segments, calculate the forward and backward optical flow between adjacent frames in the local segments and perform bidirectional alignment, and perform direction-specific multi-head attention calculation on the aligned frame group. The output time-varying probability mask includes: On the input sequence, use a sliding window with an adaptive overlapping rate to intercept local segments including at least one key frame segment, at least two forward buffer frames, and at least three backward buffer frames; Use a lightweight optical flow network to perform forward and backward optical flow estimation on each adjacent frame of each local segment. Based on the forward optical flow, map the previous frame to the current frame coordinate space through spatial deformation alignment, and based on the backward optical flow, align the subsequent frame to the current frame coordinate space through reverse deformation. Concatenate the three aligned frames along the channel dimension to generate a spatio-temporally aligned superimposed frame group; Parallel input the spatio-temporally aligned superimposed frame group into the lateral motion attention head group, longitudinal motion attention head group, and composite motion analysis head group in the multi-head attention module for feature processing; Orthogonally project and fuse the output features of the lateral motion attention head group and the longitudinal attention head group to generate a feature tensor representing the joint distribution of motion vectors; Concatenate the output features of the composite analysis head group and the feature tensor representing the joint distribution of motion vectors along the channel axis to generate a cross-motion mode fusion feature; Perform a spatial pyramid pooling operation on the cross-motion mode fusion feature, extract multi-granularity motion features at different scales respectively, and use channel attention gating to weight and fuse the features of each scale, retain the cross-dimensional coupling features representing the contraction of muscle fibers in the horizontal direction, subcutaneous deformation in the vertical direction, and motion without direction, perform convolution and non-linear activation on the coupling features, and output the initial mask; Cross-window fusion is performed on the adjacent initial mask overlapping regions based on the optical flow vectors within adjacent sliding windows. Morphological closing operation is executed on the fused mask to eliminate holes and isolated noises, and the mask boundary is refined by sub-pixel optical flow-guided interpolation, outputting a time-varying probability mask with continuous boundary for characterizing the probability distribution of subcutaneous tissue movement; Wherein, The horizontal motion focused head group is configured as a motion analysis module based on the horizontal optical flow gradient constraint. By enhancing the attention weight of the horizontal motion component, it specifically extracts the motion features related to the contraction motion of muscle fibers in the horizontal direction. The horizontal optical flow gradient constraint is: when calculating the query-key similarity matrix, a dynamic weight gain is applied to the pixel points where the horizontal optical flow gradient component exceeds a preset ratio threshold of the vertical component; The vertical motion focused head group is configured as a motion analysis module based on the vertical optical flow gradient constraint. By enhancing the feature response intensity of the vertical optical flow gradient component, it specifically extracts the motion features related to the vertical deformation of subcutaneous tissue caused by the lifting and inserting operations of acupuncture. The vertical optical flow gradient constraint is: when calculating the attention weight, exponential feature enhancement in the channel dimension is implemented for the region where the vertical optical flow gradient component exceeds a preset ratio threshold of the horizontal optical flow gradient component; The composite motion analysis head group is configured as a full-degree-of-freedom motion analysis module without direction constraint, used to capture the non-linear motion features related to rotational motion, oblique shear motion and multi-directional composite motion.

6. The acupuncture manipulation recognition method based on ultrasonic dynamic optical flow and self-supervised mask according to claim 1, characterized in that, Inject the time-varying probability mask into the first layer of the spatio-temporal separation attention network for temporal attention calculation, and in subsequent predetermined levels, use a gated unit to perform cross-layer mask conduction modulation on the mask feature of the first layer and the current layer features, including: Input the time-varying probability mask into the Patch embedding module synchronized with the spatio-temporal separation attention network. Through the two-dimensional convolution operation with a preset block size and stride, the spatial region of the mask is processed into non-overlapping blocks to generate mask embedding blocks; Stack multiple mask embedding blocks sequentially along the time dimension, and perform an expansion transformation on the channel dimension through a learnable linear projection layer to generate a mask feature tensor with the same dimension as the input tensor of the spatio-temporal attention network; In the temporal attention calculation stage of the first layer of the spatio-temporal separation attention network, add the mask feature tensor and the original query tensor element by element, and then perform layer normalization operation to make the temporal attention weight tilt towards the motion probability region where the mask annotation exceeds the preset probability threshold; Through the gated recalibration units set in the 2nd, 4th and 6th levels, the mask feature tensor is dimension-reduced by depthwise separable convolution and then transmitted across layers. The query tensor of the current layer is input into a two-level fully connected network to generate gated weights; Dynamically scale the dimension-reduced mask feature using the gated weights, and stack it to the current layer feature through residual connection to obtain the gated modulation feature containing mask-guided information; Through the pre-deployed auxiliary supervision branch, calculate the mutual information entropy between the deep feature and the mask feature tensor of the first layer. If the entropy value is lower than the set entropy threshold, trigger the backtracking adjustment of the gated weights.

7. The acupuncture manipulation recognition method based on ultrasonic dynamic optical flow and self-supervised mask according to claim 1, characterized in that, The gated modulation features output at each level are concatenated and then input into a cascaded spatio-temporal pyramid pooling module to extract motion features at different time scales. After cross-scale channel attention fusion, they are input into a multi-expert classifier. The output of the acupuncture manipulation recognition results includes: The gated modulation features output at each layer are concatenated along the channel dimension to form a fused feature tensor that contains both shallow motion details and deep semantic associations, and then batch normalization and spatio-temporal dimension reorganization are performed to match the input structure of the cascaded spatio-temporal pyramid; The fused feature tensor after batch normalization and spatio-temporal dimension reorganization is input into the spatio-temporal pyramid module. Dense sliding maximum pooling is performed within a 4-8 frame window to obtain short-term granularity features by extracting the myofibril fibrillation features caused by the instantaneous action of acupuncture. Sparse self-attention pooling is performed within a 16-24 frame window to obtain medium-term granularity features by capturing the motion rhythm features of a single lifting-thrusting or twirling manipulation. Adaptive average pooling is performed within a 32-48 frame window to extract the cumulative biomechanical features under the continuous action of multiple manipulations to obtain long-term granularity features; The short-term, medium-term, and long-term granularity features are fused through multi-scale alignment and channel attention weighting to obtain the final fused features; The final fused features are input into parallel optical flow motion experts, tissue deformation experts, and spatio-temporal coupling experts. Weighted voting is performed based on the confidence scores output by each expert to output the probability distribution of acupuncture manipulation categories.

8. An acupuncture manipulation recognition system based on ultrasonic dynamic optical flow and self-supervised mask, characterized in that, Including: An input sequence construction module for performing optical flow analysis on the obtained ultrasound image sequence, screening out key frame segments based on the frame-interval optical flow intensity difference, and constructing an input sequence by extending buffer frames forward and backward; A mask output module for intercepting local segments on the input sequence using overlapping sliding windows, calculating the forward and backward optical flow of adjacent frames within the local segments and performing bidirectional alignment, and performing direction-specific multi-head attention calculation on the aligned frame groups to output a time-varying probability mask; A cross-layer mask conduction modulation module for injecting the time-varying probability mask into the first layer of the spatio-temporal separated attention network for temporal attention calculation, and using a gating unit to perform cross-layer mask conduction modulation of the mask features of the first layer and the current layer features at subsequent predetermined levels; A manipulation recognition module that concatenates the gated modulation features output at each level and inputs them into a cascaded spatio-temporal pyramid pooling module to extract motion features at different time scales. After cross-scale channel attention fusion, they are input into a multi-expert classifier to output the acupuncture manipulation recognition results.

9. An acupuncture manipulation recognition device based on ultrasonic dynamic optical flow and self-supervised mask, characterized in that, Including: At least one processor; And a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the acupuncture manipulation recognition method based on ultrasound dynamic optical flow and self-supervised mask as described in any one of claims 1-7.

10. A computer-readable storage medium having computer-executable instructions stored thereon, characterized in that, When the executable instructions are executed by the processor, the acupuncture manipulation recognition method based on ultrasound dynamic optical flow and self-supervised mask as described in any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Multi-mode acupuncture manipulation recognition method and system fusing vision and touch

    CN116597517A

  • Acupuncture manipulation recognition method, system and equipment based on subcutaneous needle feeling and medium

    CN118038076A