Acupuncture manipulation identification method, system and equipment based on ultrasonic dynamic optical flow and self-supervised mask, and medium
Through the method based on ultrasonic dynamic optical flow and self-supervised mask, the problems of strong subjectivity and inefficiency in the identification process of acupuncture techniques in the prior art are solved, and effective characterization of dynamic biomechanical characteristics of subcutaneous tissues and end-to-end identification of acupuncture techniques are realized.
Patent Information
- Application Number
- CN202510592016.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The prior art is difficult to effectively characterize the dynamic biomechanical characteristics of subcutaneous tissue under the action of acupuncture techniques, and the image feature extraction method that relies on artificial labels leads to strong subjectivity and low efficiency in the recognition process, and lacks an end-to-end recognition model that integrates space-time dynamic effects.
Using acupuncture technique identification method based on ultrasonic dynamic optical flow and self-supervised mask, keyframe fragments are screened through optical flow analysis, input sequence is constructed, and overlapping sliding window interception, bidirectional optical flow alignment and direction-specific multi-head attention calculation are performed on the input sequence, and time-varying probability mask is output. The time-varying probability mask is injected into the space-time separation attention network for timing attention calculation, and cross-layer mask conduction modulation is performed through the gated unit. Finally, the cascading space-time pyramid pooling module and multi-expert classifier are input for feature extraction and recognition.
The systematic characterization of dynamic biomechanical characteristics of subcutaneous tissue is realized, the degree of manual intervention is reduced, and an end-to-end acupuncture technique recognition model is established, which significantly improves the recognition robustness of complex technique action patterns.
Smart Images

Figure CN120107319A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to an acupuncture manipulation recognition method, system, device and medium based on ultrasonic dynamic optical flow and self-supervisory mask. Background Art
[0002] In the study of TCM acupuncture techniques, different schools and master-disciple relationships lead to significant differences in hand parameters when doctors perform the same technique. Taking the twisting and supplementing technique as an example, the differences in needle holding postures of different doctors directly affect the unified quantification and classification accuracy of the technique parameters. Traditional research usually achieves technique characterization by collecting in vitro physical parameters such as the needle body rotation angle and lifting and insertion amplitude, but such parameters are difficult to reveal the dynamic stimulation effect of the technique on the subcutaneous tissue. Therefore, the academic community has gradually turned to combining medical imaging technology to explore the response of subcutaneous tissue, such as observing connective tissue displacement through ultrasound or using magnetic resonance imaging to record changes in functional connectivity of brain regions.
[0003] Existing technologies mainly analyze the effects of subcutaneous stimulation through two paths: one is to indirectly obtain parameters such as blood flow and temperature with the help of biosignal sensing technology; the other is to capture the dynamic characteristics of tissues based on medical imaging technologies such as ultrasound, elastic imaging or functional magnetic resonance imaging. However, existing methods face multiple challenges: first, subcutaneous tissue movement has highly complex three-dimensional dynamic characteristics, but there is a lack of standardized imaging data sets covering a variety of techniques, which limits research to small samples of manually annotated data; second, the feature extraction process relies on manually annotated anatomical structures or specific algorithm processing, which has the problems of strong subjectivity and low repeatability, making it difficult to achieve cross-case generalization; finally, existing research focuses on evidence-based analysis of efficacy, and has failed to build an end-to-end recognition model from the dynamic response of subcutaneous tissue to the category of technique, which restricts its clinical application value. Breaking through these bottlenecks requires solving key technical problems such as subcutaneous tissue biomechanical characterization, adaptive feature extraction, and spatiotemporal feature fusion. Summary of the invention
[0004] 1. Technical issues to be resolved In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a method, system, device and medium for acupuncture manipulation recognition based on ultrasonic dynamic optical flow and self-supervised mask, which solves the technical problems that the prior art is difficult to effectively characterize the dynamic biomechanical characteristics of subcutaneous tissue under the action of acupuncture manipulation, and its reliance on manually labeled image feature extraction leads to strong subjectivity and low efficiency in the recognition process, and lacks an end-to-end recognition model that integrates spatiotemporal dynamic effects.
[0005] (II) Technical solution In order to achieve the above object, the main technical solutions adopted by the present invention include: In a first aspect, an embodiment of the present invention provides a method for acupuncture manipulation recognition based on ultrasonic dynamic optical flow and self-supervised mask, comprising: Perform optical flow analysis on the acquired ultrasound image sequence, select key frame segments based on the difference in optical flow intensity between frames, and construct the input sequence by extending the buffer frames forward and backward; An overlapping sliding window is used to capture local segments on the input sequence, the forward and backward optical flows of adjacent frames in the local segments are calculated and bidirectional alignment is performed, direction-specific multi-head attention calculation is performed on the aligned frame group, and a time-varying probability mask is output; The time-varying probability mask is injected into the first layer of the spatiotemporal separation attention network for temporal attention calculation, and the gate control unit is used in the subsequent predetermined layers to modulate the first layer mask features with the current layer features through cross-layer mask conduction; The gated modulation features output from each level are spliced and input into the cascaded spatiotemporal pyramid pooling module to extract the time-sharing granular motion features. After cross-scale channel attention fusion, they are input into the multi-expert classifier to output the acupuncture technique recognition results.
[0006] Optionally, performing optical flow analysis on the acquired ultrasound image sequence, selecting key frame segments based on the difference in optical flow intensity between frames, and constructing an input sequence by extending the buffer frames forward and backward includes: The original ultrasound video data is acquired and standardized, and the optical flow field is calculated frame by frame for the obtained ultrasound image sequence to generate a visual inter-frame optical flow image sequence with HSV color coding; The visual inter-frame optical flow image sequence is converted to the HSV color space and each channel is independently normalized. The global optical flow intensity of each frame is calculated and input into a Gaussian difference filter with time-series perception to filter out significant time points where the optical flow intensity difference value exceeds three times the average fluctuation level of adjacent frames. In the ultrasound image sequence, the continuous area with the significant time point as the center and extending to both sides of the time sequence until the rate of change of the optical flow intensity drops to the baseline level is taken as the key frame segment; Based on the key frame fragments, the buffer frames are adaptively expanded forward and backward bidirectionally to obtain the initial input sequence; The initial input sequence is tested for consistency of optical flow gradients. If the motion direction mutation between adjacent frames exceeds the preset mutation angle, an interpolation transition frame is inserted, and finally an input sequence that is resistant to motion breaks is output.
[0007] Optionally, the visual inter-frame optical flow image sequence is converted to the HSV color space and each channel is independently normalized, the global optical flow intensity of each frame is calculated, and a Gaussian difference filter with time-series perception is input to screen the significant time points where the optical flow intensity difference value exceeds three times the average fluctuation level of adjacent frames, including: Convert the visual inter-frame optical flow image sequence to the HSV color space, and independently normalize the hue channel, saturation channel, and brightness channel of the HSV color space; Perform pixel-by-pixel multiplication and accumulation on the normalized values of the hue channel, saturation channel, and brightness channel to obtain the global optical flow intensity of each optical flow image that represents the dynamic change intensity of the muscle tissue, and then obtain a global optical flow intensity sequence; The first sliding time window is used to perform sliding average filtering on the global optical flow intensity sequence to generate a baseline trend component, and the global optical flow intensity sequence is subtracted from the baseline component to obtain a residual sequence; Applying a second sliding time window to the residual sequence, obtaining a second-order derivative response of the residual signal within the second sliding time window through a Gaussian difference filter to generate an enhanced residual sequence; In the enhanced residual sequence, the standard deviation of the optical flow intensity fluctuation of the N frames before and after the current frame is used as the benchmark. If the instantaneous difference value exceeds three times the standard deviation, it is marked as a significant time point.
[0008] Optionally, the adaptive extended buffer frame includes: Forward expansion intercepts an optical flow compensation frame group containing at least two consecutive frames before the key frame segment, and when the time interval between the starting point of the key frame segment and the ending point of the preceding key frame segment is less than a preset merging threshold, merges the optical flow compensation frame group into the backward expansion sequence of the preceding key frame segment; Backward expansion intercepts a motion attenuation observation frame group containing at least three consecutive frames after the key frame segment, and the interception termination condition of the motion attenuation observation frame group is that the amplitude of the optical flow between frames decays to below a preset amplitude threshold of the central frame amplitude of the key frame segment; When the starting position of the key frame segment is located at the starting point of the ultrasound image sequence, the optical flow field of the first frame is spatially mirrored along the main axis of the muscle fiber of the corresponding ultrasound image, and a random perturbation vector with an amplitude of 15%-25% of the standard deviation of the optical flow of the current frame is superimposed to generate a virtual forward compensation frame; When the key frame segment ends at the end of the ultrasound image sequence, the motion trajectory is extrapolated according to 1.2-1.8 times the duration of the optical flow vector of the last frame to generate a virtual backward observation frame.
[0009] Optionally, overlapping sliding windows are used to capture local segments on the input sequence, forward and backward optical flows of adjacent frames in the local segments are calculated and bidirectional alignment is performed, direction-specific multi-head attention calculation is performed on the aligned frame group, and the output time-varying probability mask includes: Intercepting a local segment including at least one key frame segment, at least two forward buffer frames, and at least three backward buffer frames on the input sequence using a sliding window with an adaptive overlap rate; A lightweight optical flow network is used to perform forward and backward optical flow estimation on each adjacent frame of each local segment. Based on the forward optical flow, the previous frame is spatially deformed and mapped to the coordinate space of the current frame. Based on the backward optical flow, the next frame is reversely deformed and aligned to the coordinate space of the current frame. The aligned three frames are spliced along the channel dimension to generate a spatiotemporally aligned superimposed frame group. The spatiotemporally aligned superimposed frame groups are input in parallel into the lateral motion focus head group, the longitudinal motion focus head group and the compound motion parsing head group in the multi-head attention module for feature processing; The output features of the horizontal motion focused head group and the vertical motion focused head group are orthogonally projected and fused to generate a feature tensor representing the joint distribution of motion vectors; The output features of the composite parsing head group are concatenated with the feature tensor representing the joint distribution of motion vectors along the channel axis to generate cross-motion mode fusion features; Perform spatial pyramid pooling on the cross-motion mode fusion features, extract multi-granularity motion features at different scales, and use channel attention gating to weightedly fuse features of each scale, retaining cross-dimensional coupling features that represent horizontal muscle fiber contraction, vertical subcutaneous deformation, and non-directional motion, perform convolution and nonlinear activation on the coupling features, and output the initial mask; Based on the optical flow vectors in adjacent sliding windows, the overlapping areas of adjacent initial masks are fused across windows, and a morphological closing operation is performed on the fused masks to eliminate holes and isolated noise. The mask boundaries are refined through sub-pixel optical flow-guided interpolation, and a time-varying probability mask with continuous boundaries is output to characterize the probability distribution of subcutaneous tissue motion. in, The lateral motion focus head group is configured as a motion parsing module based on horizontal optical flow gradient constraints. By enhancing the attention weight of the horizontal motion component, the motion features related to the contraction of horizontal muscle fibers are specifically extracted. The horizontal optical flow gradient constraint is: when calculating the query-key similarity matrix, a dynamic weight gain is applied to the pixel points whose horizontal optical flow gradient component exceeds the preset ratio threshold of the vertical component. The vertical motion focus head group is configured as a motion analysis module based on the vertical optical flow gradient constraint. By enhancing the characteristic response intensity of the vertical optical flow gradient component, the motion features related to the vertical deformation of the subcutaneous tissue caused by the acupuncture insertion and lifting operation are specifically extracted; the vertical optical flow gradient constraint is: when calculating the attention weight, the exponential feature enhancement of the channel dimension is implemented for the area where the vertical optical flow gradient component exceeds the preset ratio threshold of the horizontal optical flow gradient component; The compound motion analysis head group is configured as a full-degree-of-freedom motion analysis module without directional constraints, which is used to capture nonlinear motion characteristics related to rotational motion, oblique shear motion and multi-directional compound motion.
[0010] Optionally, injecting the time-varying probability mask into the first layer of the spatiotemporal separation attention network to perform temporal attention calculation, and using a gating unit in a subsequent predetermined layer to perform cross-layer mask conduction modulation on the first layer mask feature and the current layer feature, including: The time-varying probability mask is input into the Patch embedding module synchronized with the spatiotemporal separation attention network. Through the two-dimensional convolution operation with preset block size and step length, the spatial area of the mask is processed into non-overlapping blocks to generate mask embedding blocks. Multiple mask embedding blocks are stacked sequentially along the time dimension, and the channel dimension is expanded and transformed through a learnable linear projection layer to generate a mask feature tensor consistent with the dimension of the spatiotemporal attention network input tensor; In the temporal attention calculation stage of the first layer of the spatiotemporal separation attention network, the mask feature tensor is added to the original query tensor element by element, and then a layer normalization operation is performed to tilt the temporal attention weight toward the motion probability area marked by the mask that exceeds the preset probability threshold; Through the gated recalibration units set at the 2nd, 4th and 6th layers, the mask feature tensor is reduced in dimension by depthwise separable convolution and then transferred across layers, and the query tensor of the current layer is input into the two-level fully connected network to generate the gated weights; The gated weights are used to dynamically scale the reduced-dimensional mask features and superimposed on the current layer features through residual connections to obtain gated modulation features containing mask-guided information. Through the pre-deployed auxiliary supervision branch, the mutual information entropy of the deep features and the first-layer mask feature tensor is calculated. If the entropy value is lower than the set entropy threshold, the gated weight backtracking adjustment is triggered.
[0011] Optionally, the gated modulation features output from each level are spliced and input into the cascaded spatiotemporal pyramid pooling module to extract the time-sharing granular motion features, which are then input into the multi-expert classifier after cross-scale channel attention fusion. The output acupuncture manipulation recognition results include: The gated modulation features output by each layer are concatenated along the channel dimension to form a fused feature tensor containing shallow motion details and deep semantic associations, and batch normalization and spatiotemporal dimension reorganization are performed to match the input structure of the cascaded spatiotemporal pyramid. The fused feature tensor after batch normalization and spatiotemporal dimension reorganization is input into the spatiotemporal pyramid module, and dense sliding maximum pooling is performed in the 4-8 frame window to extract the muscle fiber tremor characteristics caused by the instantaneous effect of acupuncture to obtain short-term granularity features, sparse self-attention pooling is performed in the 16-24 frame window to obtain the movement rhythm characteristics of a single operation of lifting, inserting or twisting to obtain medium-term granularity features, and adaptive average pooling is performed in the 32-48 frame window to extract the cumulative biomechanical characteristics under the continuous action of multiple manipulations to obtain long-term granularity features; The short-term, medium-term, and long-term granular features are multi-scale aligned and fused with channel attention weighted to obtain the final fused features; The final fusion features are input into parallel optical flow motion experts, tissue deformation experts and spatiotemporal coupling experts, and weighted voting is performed based on the confidence scores output by each expert to output the probability distribution of acupuncture technique categories.
[0012] In a second aspect, an embodiment of the present invention provides an acupuncture manipulation recognition system based on ultrasonic dynamic optical flow and self-supervised mask, comprising: An input sequence construction module is used to perform optical flow analysis on the acquired ultrasound image sequence, select key frame segments based on the difference in optical flow intensity between frames, and construct an input sequence by extending the buffer frames forward and backward; The mask output module is used to intercept local segments on the input sequence using overlapping sliding windows, obtain the forward and backward optical flows of adjacent frames in the local segments and perform bidirectional alignment, perform direction-specific multi-head attention calculation on the aligned frame group, and output a time-varying probability mask; A cross-layer mask conduction modulation module is used to inject the time-varying probability mask into the first layer of the spatiotemporal separation attention network for temporal attention calculation, and use a gating unit in the subsequent predetermined layers to perform cross-layer mask conduction modulation on the first layer mask features and the current layer features; The technique recognition module concatenates the gated modulation features output by each level and inputs them into the cascaded spatiotemporal pyramid pooling module to extract the time-sharing granular motion features. After cross-scale channel attention fusion, they are input into the multi-expert classifier to output the acupuncture technique recognition results.
[0013] In a third aspect, an embodiment of the present invention provides an acupuncture technique recognition device based on ultrasonic dynamic optical flow and self-supervised mask, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the acupuncture technique recognition method based on ultrasonic dynamic optical flow and self-supervised mask as described above.
[0014] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having computer executable instructions stored thereon, which, when executed by a processor, implements the acupuncture technique recognition method based on ultrasonic dynamic optical flow and self-supervised mask as described above.
[0015] (III) Beneficial effects The beneficial effects of the present invention are as follows: the present invention systematically solves the technical bottlenecks of insufficient representation of dynamic characteristics of subcutaneous tissue, high artificial dependence and fragmentation of time and space models in the prior art by constructing a multi-layer cascade intelligent analysis framework. Firstly, based on the adaptive screening mechanism and buffer frame expansion strategy of optical flow intensity difference, an input sequence with temporal coherence is constructed, which effectively overcomes the problem of feature distortion caused by probe displacement and tissue movement disruption; secondly, through sliding window interception and bidirectional optical flow alignment operations, high-precision motion compensation is performed in local spatiotemporal segments, which significantly improves the biological rationality of spatiotemporal feature alignment; at the same time, a direction-specific multi-head attention mechanism is innovatively introduced to realize motion direction decoupling analysis in the process of generating time-varying probability masks, breaking through the limitations of traditional methods in the ability to analyze complex motion patterns; further, a cross-layer gated conduction modulation strategy is adopted to maintain the dynamic coupling of motion features and semantic features in deep networks, solving the problem of feature decoupling caused by the deepening of the hierarchy of existing models; finally, through the synergy of cascaded spatiotemporal pooling and multi-expert decision-making mechanism, the organic fusion of cross-scale features from microscopic muscle fiber tremors to macroscopic biomechanical effects is achieved, significantly improving the recognition robustness of complex manipulation action patterns. While reducing the degree of human intervention, the present invention establishes the first end-to-end recognition system that integrates optical flow dynamics, tissue deformation response and spatiotemporal coupling effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A schematic diagram of a flow chart of a method provided by an embodiment of the present invention; Figure 2 A schematic diagram of a specific flow chart of step S1 of the method provided in an embodiment of the present invention; Figure 3 A data collection flow chart of the method provided by an embodiment of the present invention; Figure 4 A schematic diagram of raw data clipping of the method provided in an embodiment of the present invention; Figure 5 A schematic diagram of resizing the method provided by an embodiment of the present invention; Figure 6 A schematic diagram of a specific flow chart of step S12 of the method provided in an embodiment of the present invention; Figure 7 A schematic diagram of a specific flow chart of step S2 of the method provided in an embodiment of the present invention; Figure 8 A schematic diagram of a specific flow chart of step S3 of the method provided in an embodiment of the present invention; Fig. 9 A comparison diagram of the effects of mask-guided attention calculation of the method provided in an embodiment of the present invention; Fig.10 A schematic diagram of a cross-layer gating mechanism of a method provided in an embodiment of the present invention; Fig.11 A schematic diagram of a specific flow chart of step S4 of the method provided in an embodiment of the present invention; Fig.12 The present invention provides a schematic diagram of the overall process of the method. DETAILED DESCRIPTION
[0017] In order to better explain the present invention and facilitate understanding, the present invention is described in detail below through specific implementation modes in conjunction with the accompanying drawings.
[0018] like Figure 1 As shown, an acupuncture technique recognition method based on ultrasonic dynamic optical flow and self-supervised mask proposed in an embodiment of the present invention includes: performing optical flow analysis on the obtained ultrasonic image sequence, screening out key frame segments based on the difference in optical flow intensity between frames, and constructing an input sequence by extending the buffer frames forward and backward; using overlapping sliding windows to intercept local segments on the input sequence, calculating the forward and backward optical flows of adjacent frames in the local segments and performing bidirectional alignment, performing direction-specific multi-head attention calculation on the aligned frame group, and outputting a time-varying probability mask; injecting the time-varying probability mask into the first layer of the spatiotemporal separation attention network for temporal attention calculation, and using a gating unit in subsequent predetermined levels to perform cross-layer mask conduction modulation on the first layer mask features and the current layer features; splicing the gated modulation features output by each level and inputting them into a cascaded spatiotemporal pyramid pooling module, extracting time-sharing granular motion features, and inputting them into a multi-expert classifier after cross-scale channel attention fusion to output acupuncture technique recognition results.
[0019] The present invention systematically solves the technical bottlenecks of insufficient representation of dynamic characteristics of subcutaneous tissue, high artificial dependence and fragmentation of spatiotemporal models in the prior art by constructing a multi-layer cascade intelligent analysis framework. Firstly, based on the adaptive screening mechanism and buffer frame expansion strategy of optical flow intensity difference, an input sequence with temporal coherence is constructed, which effectively overcomes the problem of feature distortion caused by probe displacement and tissue movement disruption; secondly, through sliding window interception and bidirectional optical flow alignment operations, high-precision motion compensation is performed in local spatiotemporal segments, which significantly improves the biological rationality of spatiotemporal feature alignment; at the same time, a direction-specific multi-head attention mechanism is innovatively introduced to realize motion direction decoupling analysis in the process of generating time-varying probability masks, breaking through the limitations of traditional methods in the ability to analyze complex motion patterns; further, a cross-layer gated conduction modulation strategy is adopted to maintain the dynamic coupling of motion features and semantic features in deep networks, solving the problem of feature decoupling caused by the deepening of the hierarchy of existing models; finally, through the synergy of cascaded spatiotemporal pooling and multi-expert decision-making mechanism, the organic fusion of cross-scale features from microscopic muscle fiber tremors to macroscopic biomechanical effects is achieved, significantly improving the recognition robustness of complex manipulation action patterns. While reducing the degree of human intervention, the present invention establishes the first end-to-end recognition system that integrates optical flow dynamics, tissue deformation response and spatiotemporal coupling effects.
[0020] In order to better understand the above technical solution, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a clearer and more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0021] Specifically, an embodiment of the present invention provides a method for identifying acupuncture techniques based on ultrasonic dynamic optical flow and self-supervised mask, comprising: S1. Perform optical flow analysis on the acquired ultrasound image sequence, select key frame segments based on the difference in optical flow intensity between frames, and construct an input sequence that is resistant to motion breaks by extending the buffer frames forward and backward. Key frame extraction can effectively select data for a certain period of time containing key information to reduce information redundancy. This embodiment proposes a motion sequence keyframe extraction (MSKE) extraction method, which obtains continuous frame segments to form an input sequence without losing key scene information of changes in muscle tissue caused by acupuncture in the image, ensuring that the subsequent extracted spatiotemporal features can effectively reflect the stimulation effect of each type of technique on subcutaneous muscle tissue during continuous operation during acupuncture.
[0022] Furthermore, if Figure 2 As shown, step S1 includes: S11. Acquire and standardize the original ultrasound video data, perform frame-by-frame optical flow field calculation on the obtained ultrasound image sequence, and generate a HSV color-coded visual inter-frame optical flow image sequence.
[0023] In order to obtain the original ultrasound video data of the acupuncture movement area, six senior acupuncturists and two ultrasound acquisition physicians from four medical institutions and 53 healthy subjects were invited to participate in the ultrasound data collection.
[0024] In the specific collection work, refer to Figure 3 (a) shows the actual operation scenario of data collection. Figure 3(b) is a schematic diagram of data collection, requiring the acupuncturist to insert the needle at the Kongzui acupoint (LU6) of the subject and slightly lift the needle body. At the same time, the ultrasound physician moves the probe close to the needle insertion point to find the muscle movement center of the acupuncture effect. After the ultrasound physician determines the probe position, the acupuncturist follows the recorder's instructions to perform four types of manipulations in sequence: twisting to tonify, twisting to drain, lifting and inserting to tonify, and lifting and inserting to drain. Each type of manipulation is performed for 4-7 seconds, and the ultrasound physician records the image in real time. After each manipulation, the needle is removed and waited for 3-5 minutes to relax the subject's muscles. Before performing the next manipulation, ultrasound is used to observe whether the subject's muscles have returned to a resting state. The interval time can be appropriately extended according to the specific situation. Figure 3 (c) in the figure shows a set of ultrasound video frame sequence data finally acquired. The actual data screening images with a recording time of 4 seconds or more are saved and exported in .avi and .dcm formats respectively.
[0025] Based on the above operation steps, a real-time ultrasound image dataset of four types of acupuncture manipulations acting on subcutaneous muscle tissue, including Reinforcing by twisting and rotating (RFTR), Reducing by twisting and rotating (RDTR), Reinforcing by lifting and thrusting (RFLT), and Reducing by lifting and thrusting (RDLT), was collected and constructed. The following data standardization processing was then performed: (1) Raw data cropping: Select a video data from the acquired raw ultrasound video data set and extract the video frame. Figure 4 As shown in (a) in the figure, the original ultrasound video contains borders and text information, which will interfere with the subsequent image analysis. Therefore, ScreenToGif software is used to cut out the image part of the original video data without borders and text information. The window size is 855×680 (pixels) to obtain new video data, as shown in Figure 4 As shown in (b) in .
[0026] (2) Resizing: In order to meet the input data size requirements of most deep networks and appropriately reduce the size of the data tensor during model training, based on the video data obtained in step 1, all video data is resized to unify the 855×680 (pixel) video data into 512×512 (pixel) video frame data frame by frame as follows: Figure 5 shown.
[0027] (3) Standardization: The frame sequence processed in (2) is restored to video data in .avi format, and the video data parameters are unified. Due to the manual errors when ultrasound acquisition physicians manually record data videos, it is difficult to ensure that the duration is completely uniform. Therefore, the middle segment of 4 seconds of video data is uniformly cut from each ultrasound video segment greater than or equal to 4 seconds, and the frame rate of the video is unified to 30 frames / s (restored to the frame rate of the original ultrasound image video). It should be emphasized that each video data is processed through a unified standardization preprocessing process to obtain a unique independent sample, and there is no situation where multiple samples are repeatedly cut from a piece of original image data.
[0028] In a specific embodiment, for any acquired ultrasound image sequence I Among them, it is recorded as: ; in, I Represents the entire video sequence; I t express t The video frame corresponding to the moment, and I T Representation sequence I The last video frame of ; R represents the variable dimension, here I t is a three-dimensional matrix variable with dimensions H × W × C (height × width × number of channels), I is a 4D tensor with dimensions T × H × W × C (Duration × height × width × number of channels). The ultrasound video sequence I is passed through the built-in flownet2 optical flow estimation network under the mmflow library to obtain the inter-frame optical flow map of each video. If the inter-frame optical flow sequence data of all data is saved in .flo format, the subsequent call of the optical flow data will occupy a lot of memory. The current extraction stage does not need to provide very accurate optical flow information, but only provides a basis for screening in subsequent steps, so it is chosen to be represented by an optical flow visualization. The inter-frame optical flow sequence of the ultrasound video is O Among them, it is recorded as: ; in, is calculated by calculating two consecutive frames I t and I t+1 The optical flow between OHSV color coding is used, where hue (H) indicates the direction of movement, and saturation (S) and value (V) indicate the amplitude of movement; R represents the variable dimension, here O t,t+1 is a three-dimensional matrix variable with dimensions H × W × C (height × width × number of channels), O is a 4D tensor with dimensions T × H × W × C (Duration × Height × Width × Number of Channels). O Taking one frame as an example, the pixel-level encoding in HSV color space is: ; in,( x , y ) represents the position of the pixel in the optical flow image, H t,t-1 ( x , y ), S t,t-1 (x, y )and V t,t-1 (x, y ) respectively represent the motion information of the pixel.
[0029] S12. Convert the visualized inter-frame optical flow image sequence into the HSV color space and perform independent normalization processing on each channel, calculate the global optical flow intensity of each frame, and input it into a time-series-aware Gaussian difference filter to extract and filter significant time points where the optical flow intensity difference value exceeds three times the average fluctuation level of adjacent frames.
[0030] Furthermore, if Figure 6 As shown, step S11 includes: S121, converting the visualized inter-frame optical flow image sequence into the HSV color space, and independently normalizing the hue channel, saturation channel, and brightness channel of the HSV color space.
[0031] S122, performing pixel-by-pixel multiplication and accumulation on the normalized values of the hue channel, the saturation channel, and the brightness channel to obtain a global optical flow intensity of each optical flow image that represents the intensity of dynamic changes in muscle tissue, and further obtain a global optical flow intensity sequence.
[0032] S123, using the first sliding time window to perform sliding average filtering on the global optical flow intensity sequence to generate a baseline trend component, and subtracting the global optical flow intensity sequence from the baseline component to obtain a residual sequence.
[0033] S124: Apply a second sliding time window to the residual sequence, and obtain a second-order derivative response of the residual signal in the second sliding time window through a Gaussian difference filter to generate an enhanced residual sequence.
[0034] Specifically, this step uses the first sliding time window, with the time window length set to 0.5-1.5 seconds, corresponding to 20-38 frames, focusing on suppressing low-frequency physiological interference such as breathing (0.2-0.3Hz) and vascular pulsation (1-1.5Hz), while retaining the medium and high frequency motion components related to acupuncture operations. On the residual sequence after baseline removal, the second sliding time window is applied, with the time window length matching the minimum effective stimulation duration of clinically verified acupuncture techniques (50-150ms), corresponding to 3-9 frames, to extract transient tissue response characteristics caused by operations such as lifting, inserting, and twisting.
[0035] S125. In the enhanced residual sequence, the standard deviation of the optical flow intensity fluctuation of the N frames before and after the current frame is used as a benchmark. If the instantaneous difference value exceeds three times the standard deviation, it is marked as a significant time point.
[0036] In a specific embodiment, for each set of optical flow sequences O Calculate the optical flow intensity I. Convert to HSV color space and normalize the H, S, and V channels so that their values are in the range of [0, 1]. O t,t+1 The calculation formula of the optical flow intensity I is as follows: ; in, H norm (.) is the pixel tone normalization processing function; S norm (.) is the pixel saturation normalization function; V norm (.) is the pixel brightness normalization function. Thus, each optical flow image can be obtained O t,t+1 The optical flow intensity I t,t+1 The change of the optical flow intensity I reflects the intensity of the muscle tissue movement under the action of acupuncture, and reflects the effect of the manipulation on the subcutaneous tissue.
[0037] Next, take the sampling time as the sampling point and give the custom parameters first n , N and LThese custom sampling parameters are used to control the size of the model input sequence. The specific definitions and calculation formulas are shown in the parameter definition table in Table 1 below.
[0038] Table 1 Parameter definition table
[0039] in, N It is the total number of frames in the subsequent classification network input sequence. Common values are integers such as 8, 16, and 32. n is the total number of sampling points for each segment of ultrasound image data, which can be calculated based on the input sequence length. N The required sampling frequency can be adjusted by your own settings. The larger the value, the higher the sampling frequency. L It is the number of local continuous frames at each sampling point, which is used to ensure the continuity of data at the sampling point and to ensure that the timing information of the muscle bundle movement in the acupuncture area is true and complete.
[0040] In order to select the sampling points, we first calculate the optical flow intensity change of the inter-frame optical flow sequence , the formula is as follows: ; Among them, I t+1,t+2 and I t,t+1 They are inter-frame optical flow O t+1,t+2 and O t,t+1 The corresponding optical flow intensity. It reflects the moment when the dynamic information in the image changes significantly, and can reflect important motion information. The size of the filter n The largest , arranged as follows: ; at this time The corresponding moment is the sampling point, and Corresponding real time order n The sampling points are as follows: ; In this way, it can be ensured that the sampling order and the key frame input sequence are consistent with the real time order, and the timing confusion in which the later video frames in the original video appear before the earlier video frames will not occur.
[0041] Sure n After sampling points, each sampling point corresponds to two frames of the original image sequence, and continuous L Frames of video form a continuous keyframe segment segment i : ; S13. In the ultrasound image sequence, the continuous area with the significant time point as the center and extending to both sides of the time sequence until the rate of change of the optical flow intensity drops to the baseline level is taken as the key frame segment.
[0042] S14, taking the key frame segments as a reference, adaptively expanding the buffer frames in both forward and backward directions to obtain an initial input sequence.
[0043] It should be understood that the core function of the buffer frame is to prevent subsequent mask weight jumps through spatiotemporal continuity constraints. Specifically, the forward and backward bidirectional adaptive expansion of the buffer frame includes the following steps: (1) Forward expansion intercepts an optical flow compensation frame group containing at least two consecutive frames before the key frame segment. The optical flow compensation frame group is used to reconstruct the initial momentum accumulation process of the acupuncture force. When the time interval between the starting point of the key frame segment and the ending point of the previous key frame segment is less than a preset merging threshold (e.g., 50 milliseconds), the optical flow compensation frame group is merged into the backward extension sequence of the previous key frame segment.
[0044] (2) Backward expansion intercepts a motion attenuation observation frame group containing at least three consecutive frames after the key frame segment. The motion attenuation observation frame group covers the complete attenuation cycle of muscle fiber elastic rebound. The interception termination condition of the motion attenuation observation frame group is that the inter-frame optical flow amplitude decays to below a preset amplitude threshold (e.g., 20%) of the amplitude of the central frame of the key frame segment.
[0045] (3) When the starting position of the key frame segment is located at the starting point of the ultrasound image sequence, the optical flow field of the first frame is spatially mirrored along the main axis of the muscle fiber of the corresponding ultrasound image, and a random perturbation vector with an amplitude of 15%-25% of the standard deviation of the optical flow of the current frame is superimposed to generate a virtual forward compensation frame.
[0046] (4) When the end position of the key frame segment is at the end point of the ultrasound image sequence, the motion trajectory is extrapolated according to 1.2-1.8 times the duration of the optical flow vector of the last frame to generate a virtual backward observation frame; wherein, the timestamps of the inserted virtual forward compensation frames or virtual backward observation frames are evenly spaced with adjacent frames, and the difference in motion direction between adjacent virtual frames and real frames is no more than 15 degrees.
[0047] S15, performing optical flow gradient consistency detection on the initial input sequence, inserting an interpolation transition frame if the motion direction mutation between adjacent frames exceeds a preset mutation angle, and finally outputting an input sequence that is resistant to motion breaks.
[0048] S2. Use overlapping sliding windows to capture local segments on the input sequence, calculate the forward and backward optical flows of adjacent frames in the local segments and perform bidirectional alignment, perform direction-specific multi-head attention calculation on the aligned frame group, and output a time-varying probability mask.
[0049] In this step, overlapping sliding windows are used to capture local segments on the input sequence, and the forward and backward optical flow fields of adjacent frames in the local segments are calculated through the optical flow estimation network. Pixel-level spatial deformation alignment is performed on the front and rear frames based on the bidirectional optical flow to eliminate the spatiotemporal misalignment caused by probe displacement or tissue elastic deformation. A direction-specific multi-head attention mechanism is deployed on the aligned frame group, and independent attention head groups are designed for horizontal motion (such as lateral contraction of muscle fibers), vertical motion (such as longitudinal deformation of subcutaneous tissue) and compound motion (such as rotational shear). The feature responses of motion-sensitive areas in different directions are dynamically enhanced through motion direction threshold constraints and optical flow amplitude weighting mechanisms. Finally, a time-varying probability mask is output, and its pixel value characterization technique represents the motion significance of each area.
[0050] Furthermore, if Figure 7 As shown, step S2 includes: S21. Using a sliding window with an adaptive overlap rate on an input sequence, intercept a local segment including at least one key frame segment, at least two forward buffer frames, and at least three backward buffer frames.
[0051] S22. A lightweight optical flow network is used to perform forward and backward optical flow estimation on adjacent frames of each local segment. Based on the forward optical flow, the previous frame is spatially deformed and mapped to the current frame coordinate space. Based on the backward optical flow, the next frame is reversely deformed and aligned to the current frame coordinate space. The aligned three frames are spliced along the channel dimension to generate a spatiotemporally aligned superimposed frame group.
[0052] In one embodiment, for any continuous local segment, there is a frame sequence , sliding window t Current frame at time I t For query , take itself and two adjacent frames in time, that is I t-1 , I t , I t+1 Key k t-1 , , k t+1 In order to locate the moving area more accurately and explicitly model the temporal motion information, the key image of the previous frame is aligned using the optical flow method. k t-1 Key with next frame k t+1 Align to , that is, aligning the previous frame and the next frame to the current frame to minimize the misalignment of temporal and spatial features caused by probe jitter and displacement.
[0053] Specifically, a lightweight SPyNet model is introduced to calculate the inter-frame optical flow. SpyNet(.) indicates that the inter-frame optical flow is calculated using the SPyNet model. The optical flow calculation formula is as follows: ; Among them, flow forward : , indicating that from the frame I t-1 To Frame I t The forward optical flow of R represents the variable dimension, here f t-1,t is a four-dimensional matrix variable with dimensions B ×2× H × W (batch size × number of channels (2) × height × width); flow backward : , indicating that from the frame I t+1 To Frame I t The backward optical flow, R represents the variable dimension, here f t+1 , t is a four-dimensional matrix with variable dimensions B ×2× H × W (batch size × number of channels (2) × height × width).
[0054] For a continuous keyframe segment segment i The optical flow set is as follows: ; in, F forward and F backward Key frame segments segment i The forward optical flow set and the backward optical flow set of each consecutive key frame segment segment i have L' frame( segment i Keep the previous frame and the next frame of the clip. L' =L+ 2) When calculating the forward and backward optical flow information segment i The first and last two frames are only L The optical flow calculation of the middle frame provides a reference for constructing segment i The forward optical flow set F forward and the backward optical flow set F backward When , the time range of the retained continuous frames is [2, L' -1].
[0055] Then use the forward and backward optical flow information to key k t-1 , k t+1 to align.
[0056] Previous frame key k t-1 Align to current frame key k t Get the aligned key k ' t-1 : ; Next frame key k t+1 Align to current frame key k t Get the aligned key k ' t+1 : ; After optical flow alignment, the key The splicing is as follows to get the key of the superimposed frame group k Sum v : ; S23, inputting the spatiotemporally aligned superimposed frame group in parallel into the lateral motion focus head group, the longitudinal motion focus head group, and the compound motion parsing head group in the multi-head attention module for feature processing; Specifically, the lateral motion focus head group is configured as a motion analysis module based on horizontal optical flow gradient constraints. By enhancing the attention weight of the horizontal motion component, it specifically extracts motion features related to the horizontal muscle fiber contraction movement; the horizontal optical flow gradient constraint is: when calculating the query-key similarity matrix, a 3-5 times dynamic weight gain is applied to the pixel points whose horizontal optical flow gradient component exceeds the preset ratio threshold (more than 2 times) of the vertical optical flow gradient component.
[0057] The longitudinal motion focus head group is configured as a motion analysis module based on the vertical optical flow gradient constraint. By enhancing the characteristic response intensity of the vertical optical flow gradient component, it specifically extracts the motion features related to the vertical deformation of the subcutaneous tissue caused by the acupuncture insertion and lifting operation; the vertical optical flow gradient constraint is: when calculating the attention weight, the vertical optical flow gradient component exceeds the preset ratio threshold (more than 2 times) of the horizontal optical flow gradient component, that is, the area where the vertical optical flow gradient component dominates is subjected to exponential feature enhancement in the channel dimension.
[0058] The composite motion parsing head group is configured as a full-degree-of-freedom motion parsing module without directional constraints, which is used to capture rotational motion, oblique shear motion and multi-directional composite motion patterns, and models the nonlinear motion characteristics of the multi-dimensional motion vector field through the spatiotemporal joint attention mechanism of the aligned frame group.
[0059] S24, performing orthogonal projection fusion on the output features of the lateral motion focused head group and the longitudinal motion focused head group to generate a feature tensor representing the joint distribution of motion vectors.
[0060] S25. Splicing the output features of the composite parsing head group and the feature tensor representing the joint distribution of motion vectors along the channel axis to generate cross-motion mode fusion features.
[0061] S26. Perform spatial pyramid pooling on the cross-motion mode fusion features, extract multi-granularity motion features at different scales, and use channel attention gating to weightedly fuse features of each scale, retaining the cross-dimensional coupling features that characterize horizontal muscle fiber contraction, vertical subcutaneous deformation, and non-directional motion, convolve and nonlinearly activate the coupling features, and output the initial mask.
[0062] S27, based on the optical flow vectors in the adjacent sliding windows, the overlapping areas of the adjacent initial masks are fused across windows, a morphological closing operation is performed on the fused mask to eliminate holes and isolated noise, and the mask boundary is refined through sub-pixel optical flow guided interpolation, and a time-varying probability mask with a continuous boundary is output for representing the probability distribution of subcutaneous tissue movement. The continuously transitioned time-varying probability mask represents the probability distribution of subcutaneous tissue movement, and its high response area is consistent with the dynamic deformation area of the muscle bundle under the action of acupuncture.
[0063] S3. Inject the time-varying probability mask into the first layer of the spatiotemporal separation attention network for temporal attention calculation, and use the gating unit in the subsequent predetermined layers to perform cross-layer mask conduction modulation on the first-layer mask features and the current-layer features.
[0064] Furthermore, if Figure 8 As shown, step S3 includes: S31. The time-varying probability mask is input into the Patch embedding module synchronized with the spatiotemporal separation attention network. Through the two-dimensional convolution operation with preset block size and step size, the spatial area of the mask is processed into non-overlapping blocks to generate a mask embedding block.
[0065] S32. Multiple mask embedding blocks are stacked sequentially along the time dimension, and the channel dimension is expanded and transformed through a learnable linear projection layer to generate a mask feature tensor that is consistent with the dimension of the spatiotemporal attention network input tensor. Among them, the learnable linear projection layer is a parameterized linear transformation layer, which maps the input features from the original dimension to the target dimension. The transformation matrix (weight) in the mapping process is automatically optimized through the back-propagation algorithm.
[0066] S33. In the temporal attention calculation stage of the first layer of the spatiotemporal separation attention network, the mask feature tensor is added to the original query tensor element by element, and then a layer normalization operation is performed to tilt the temporal attention weight toward the motion probability area marked by the mask that exceeds the preset probability threshold (0.7).
[0067] S34, through the gated recalibration units set at the 2nd, 4th and 6th layers, the mask feature tensor is reduced in dimension by the depthwise separable convolution and then transferred across layers, and the query tensor of the current layer is input into the two-level fully connected network to generate the gated weights.
[0068] S35. Use the gating weight to dynamically scale the mask features after dimensionality reduction, and superimpose them to the current layer features through residual connection to obtain the gated modulation features containing the mask guidance information.
[0069] S36. By pre-deploying the auxiliary supervision branch before the final classification layer, the mutual information entropy of the deep features and the first-layer mask feature tensor is calculated. If the entropy value is lower than the set entropy threshold, the gating weight backtracking adjustment is triggered. Specifically, the backtracking adjustment includes: freezing the high-level network parameters, back-propagating the fully connected layer weights of the gating unit; if the entropy threshold is not reached for three consecutive iterations, the mask conduction path of the current channel is closed.
[0070] In another specific embodiment, the TimeSformer with divided space-time attention (Div-ST) is selected as the basic structure of the backbone network. First, the Patch Embedding module (Patch Embedding) consistent with the time and space encoding of TimeSformer is used to embed the SSAD ROI mask MConverted into a tensor form suitable for Transformer module input in the original network. The following operations are performed: Use a two-dimensional convolution operation with a kernel size of P×P and a stride of P to perform non-overlapping block processing on the space of a single frame mask. Based on the original resolution H×W of the ultrasound image frame, the mask of size H×W is divided into NUM = (H / P) × (W / P) mask embedding blocks. The spatial dimension of each embedding block is P×P, which is generated after the channel dimension is expanded (Flatten) by the learnable linear projection layer. C × P ², where C is the number of mask channels.
[0071] Final SSAD ROI mask M The conversion is as follows: ; Among them, PatchEmbed(.) indicates Patch embedding processing, and SSAD ROI mask M The processed dimensions are also adjusted to a tensor form suitable for Transformer module input; R represents the variable dimension, here M ' is a three-dimensional matrix variable with dimensions B × NUM ×( C × P ²). The main input feature after Patch embedding is defined as X , X The corresponding query tensor is obtained in the temporal attention calculation in the first-layer Transformer module q x , mask the SSAD ROI M' and q x Add and normalize the layers as follows: ; Therefore, we implement SSAD ROI mask based on the temporal attention of the basic backbone network Div-ST TimeSformer. M Guide the spatiotemporal attention calculation in the first-layer Transformer module to instruct the model's self-supervision-adaptive implementation of enhanced attention to the muscle movement characteristics around the acupoints affected by acupuncture manipulation. Fig. 9 Shows whether to add SSADROI mask M Comparison of the effects of guiding attention calculation.
[0072] Fig. 9The wireframe in part a shows the original spatiotemporal attention calculation effect of the backbone network Div-ST TimeSformer. It can be seen that the temporal attention part performs cross-attention calculations on pixels at the same spatial position at different time steps. t Each spatial position pixel at a moment performs cross-attention calculation on the global spatial pixel at the current moment to extract temporal global features.
[0073] Fig. 9 Wireframe in part b: showing the SSAD ROI mask M The optimized temporal attention part not only performs cross-attention calculations on pixels at the same spatial position in different time frames, but also performs cross-attention calculations based on the SSAD ROI mask. M Under the guidance of the algorithm, a higher weight is given to the ROI area containing image motion information and relatively static background information is paid less attention. At the same time, it also avoids the neglect of subtle dynamic features by the binary mask, thereby achieving more accurate temporal local attention allocation. Combined with the original global spatial attention mechanism, it can better highlight the characteristic information of the manipulation area, so that the model can have a more accurate understanding of the subcutaneous effects of different acupuncture effects.
[0074] It is important to understand that the deep feature extraction part of Div-ST TimeSformer is composed of multiple layers of spatiotemporal attention separation Transformer modules, and only the SSAD ROI mask is introduced in the first layer (layer 0). M The information can easily cause the SSAD ROI mask to be gradually lost as the network deepens layer by layer. M Aiming at the guidance information of the acupuncture area, the high-level semantic features are decoupled from the low-level motion features. For this reason, a simple gating module is designed based on the principle of the gating mechanism. It can realize the cross-layer introduction of SSADROI mask without introducing too much calculation and significantly increasing the network complexity. M The specific structure is as follows Fig.10 shown.
[0075] Depend on Fig.10 It can be seen that the SSADROI mask is introduced in the temporal attention calculation of the first layer (layer 0) Transformer module M' , as shown in the formula, the corresponding query tensor is obtained in the temporal attention calculation q' x Similarly, in the subsequent i The SSAD ROI mask is also introduced in the layer Transformer module M' . Set up a gating module to iQuery tensor for layer-time attention q xi and SSAD ROI mask M' The gating mechanism will adaptively adjust the SSAD ROI mask according to the characteristics of the layer. M' First, set the gating weight g t For a tensor with the current query q xi Tensors of the same shape, where g t It is calculated through the gating layer, and the formula is as follows: ; Among them, FC1(.) is the linear transformation process of fully connected layer 1, and FC2(.) is the linear transformation process of fully connected layer 2. It is worth noting that the RELU activation function is enabled between the two fully connected layers. Indicates i Layer (SSAD ROI mask M introduced by the gating mechanism ' The output features of the first layer (layer 0) are used to calculate the mutual information entropy. If the output features of the first layer (layer 0) are less than the threshold τ , then the gated weights are adjusted retrospectively, ▽ g t (.) is the gate weight gradient, which adjusts the gate weight during the optimization process. λ is a learnable parameter, g t ' (.) is the optimized gating weight.
[0076] Finally, for q xi Introducing SSAD ROI Mask M' The information is as follows: ; Gating weight g t Will be based on i The calculation results of the layer adaptively retain the SSAD ROI mask introduced by the bottom layer M The reason for choosing cross-layer introduction instead of layer-by-layer introduction is to properly control the degree of coupling of the underlying motion features to the deep semantic features, so that the model will not lose a lot of attention to the acupuncture area, and at the same time will not interfere with the extraction of deep semantic features due to too much underlying features.
[0077] Experiments have shown that stacking 12 layers of spatiotemporal attention-separated Transformer modules and introducing a gating mechanism across the 2nd, 4th, and 6th layers can achieve the best recognition effect.
[0078] S4. The gated modulation features output from each level are spliced and input into the cascaded spatiotemporal pyramid pooling module to extract the time-sharing granular motion features. After cross-scale channel attention fusion, they are input into the multi-expert classifier to output the acupuncture technique recognition results.
[0079] Furthermore, if Fig.11 As shown, step S4 includes: S41. The gated modulation features output by each layer are concatenated along the channel dimension to form a fused feature tensor containing shallow motion details and deep semantic associations, and batch normalization and spatiotemporal dimension reorganization are performed to match the input structure of the cascaded spatiotemporal pyramid.
[0080] In this step, batch normalization is performed as follows: mean-variance normalization is performed on the batch dimension to eliminate the feature distribution offset caused by the difference in optical flow intensity of different video samples, and the spatiotemporal dimension is reorganized as follows: the dimension of the feature tensor is reorganized from [batch, time, space, channel] to [batch, channel, time × space] to match the input requirements of the cascaded spatiotemporal pyramid module.
[0081] S42. Input the fused feature tensor after batch normalization and reorganization of the spatiotemporal dimension into the spatiotemporal pyramid module, perform dense sliding maximum pooling in 4-8 frame windows, extract the muscle fiber tremor characteristics caused by the instantaneous effect of acupuncture, and obtain short-term granularity characteristics; perform sparse self-attention pooling in 16-24 frame windows, obtain the medium-term granularity characteristics by acquiring the motion rhythm characteristics of a single operation of lifting, inserting or twisting, and perform adaptive average pooling in 32-48 frame windows to extract the cumulative biomechanical characteristics under the continuous action of multiple manipulations to obtain long-term granularity characteristics.
[0082] S43. Perform multi-scale alignment and channel attention weighted fusion on the short-term, medium-term and long-term granular features to obtain the final fused features.
[0083] In this step, multi-scale alignment is to map the short-time, medium-time and long-time features to a unified spatiotemporal coordinate system through deformable convolution guided by optical flow, eliminating the misalignment caused by scale differences; and channel attention weighting is to calculate the channel importance weights for each granular feature separately, and use a gating mechanism to strengthen the transient response channel of short-time features (gain coefficient 2.0-3.5), the rhythmic marking channel of medium-time features (gain 1.5-2.0), and the cumulative effect channel of long-time features (gain 1.0-1.2); the cross-scale fusion priority is determined based on the dynamic routing algorithm, and the short-time: medium-time: long-time weight ratio is initialized to 4:3:3.
[0084] S44, inputting the final fusion features into the parallel optical flow motion experts, tissue deformation experts and spatiotemporal coupling experts, performing weighted voting based on the confidence scores output by each expert, and outputting the probability distribution of acupuncture manipulation categories.
[0085] Specifically, the motion pattern expert is used to analyze the time-frequency characteristics (main frequency, harmonic energy ratio) of the optical flow vector field, the tissue deformation expert is used to analyze the texture strain energy distribution and the shear strain rate time-varying curve, and the spatiotemporal coupling expert is used to evaluate the motion-deformation phase difference and energy transfer efficiency. Each expert independently outputs the probability distribution of the manipulation category and performs weighted voting based on the real-time confidence score (optical flow motion expert 40%, tissue deformation 30%, spatiotemporal coupling 30%); when the confidence difference of the highest-voted category is <15%, sub-pixel feature reprojection verification is triggered. Finally, a result containing the manipulation category label (such as lifting / twisting / combination) and confidence score (0-1) is generated.
[0086] In addition, an embodiment of the present invention provides an acupuncture technique recognition system based on ultrasonic dynamic optical flow and self-supervised mask, comprising: The input sequence construction module is used to perform optical flow analysis on the acquired ultrasound image sequence, select key frame segments based on the difference in optical flow intensity between frames, and construct an input sequence that is resistant to motion breaks by extending the buffer frames forward and backward.
[0087] The mask output module is used to intercept local segments on the input sequence using overlapping sliding windows, calculate the forward and backward optical flows of adjacent frames in the local segments and perform bidirectional alignment, perform direction-specific multi-head attention calculation on the aligned frame group, and output a time-varying probability mask.
[0088] The cross-layer mask conduction modulation module is used to inject the time-varying probability mask into the first layer of the spatiotemporal separation attention network for temporal attention calculation, and use the gating unit in the subsequent predetermined levels to perform cross-layer mask conduction modulation on the first-layer mask features and the current layer features.
[0089] The technique recognition module concatenates the gated modulation features output by each level and inputs them into the cascaded spatiotemporal pyramid pooling module to extract the time-sharing granular motion features. After cross-scale channel attention fusion, they are input into the multi-expert classifier to output the acupuncture technique recognition results.
[0090] Furthermore, an embodiment of the present invention provides an acupuncture technique recognition device based on ultrasonic dynamic optical flow and self-supervised mask, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the acupuncture technique recognition method based on ultrasonic dynamic optical flow and self-supervised mask as described above.
[0091] Furthermore, an embodiment of the present invention provides a computer-readable storage medium having computer-executable instructions stored thereon. When the executable instructions are executed by a processor, the acupuncture technique recognition method based on ultrasonic dynamic optical flow and self-supervised mask as described above is implemented.
[0092] In summary, the embodiments of the present invention provide a method, system, device and medium for acupuncture technique recognition based on ultrasonic dynamic optical flow and self-supervised mask. Fig.12 , the overall process is as follows: In the first stage, we first constructed an ultrasound video dataset containing four acupuncture techniques: twisting and replenishing, twisting and draining, lifting and inserting to replenish, and lifting and inserting to drain, and explored the differences in the stimulation effects of different acupuncture techniques on subcutaneous tissue. In order to verify the validity and scientificity of the dataset, we conducted a preliminary experimental analysis based on imaging omics, extracted texture features based on ultrasound images, and performed a significance analysis of the features. It was found that there were significant differences between the four types of technique data collected, and the data samples were representative, providing effective support for subsequent analysis. Then, we performed optical flow analysis on the acquired ultrasound image sequence, screened out key frame segments based on the difference in optical flow intensity between frames, and constructed an input sequence that resisted motion fracture by extending the buffer frames forward and backward.
[0093] In the second stage, overlapping sliding windows are used to capture local segments on the input sequence, and the forward and backward optical flows of adjacent frames in the local segments are calculated and bidirectional alignment is performed. Direction-specific multi-head attention calculations are performed on the aligned frame group to output a time-varying probability mask. Based on the common Region of Interest (ROI) mask annotation idea in medical image analysis, the present invention designs a self-supervised-adaptive dynamic (SSAD) ROI mask extraction module for the muscle-affected area under the manipulation. On the one hand, it solves the mask annotation problem when the target area has no clear outline or no clear boundary range threshold when analyzing medical image videos. On the other hand, the ROI mask that adaptively adjusts the weight distribution avoids the disadvantage of binary mask missing global information, and can better guide the feature extraction network to extract more accurate spatiotemporal motion features in muscle ultrasound images under different acupuncture techniques.
[0094] In the third stage, the time-varying probability mask is injected into the first layer of the spatiotemporal separation attention network for temporal attention calculation, and the gating unit is used in the subsequent predetermined layers to perform cross-layer mask conduction modulation on the first-layer mask features and the current layer features; the gated modulation features output by each layer are spliced and input into the cascaded spatiotemporal pyramid pooling module to extract the time-sharing granular motion features, which are then input into the multi-expert classifier after cross-scale channel attention fusion to output the acupuncture technique recognition results.
[0095] Therefore, the present invention proposes a MSKE (Motion sequence keyframe extraction) module and a SSAD (Self-supervised-adaptive dynamic) ROI mask extraction module, which respectively solve the problems of key information extraction of long video sequences and self-supervised-adaptive extraction of dynamic ROI masks, greatly reducing labor and time costs. The optimization module and cross-layer gating mechanism are added to the conventional timeSformer model of spatiotemporal attention separation (Div-ST) to achieve spatiotemporal attention feature extraction guided by the SSAD ROI mask, and construct a technique recognition model based on four types of techniques: twisting and supplementing, twisting and draining, lifting and inserting, and lifting and inserting. The accuracy of the proposed optimized recognition model reaches 92.41%. The proposed recognition model has good stability and robustness through experiments, and the proposed optimization modules have certain versatility, which also provides a way of thinking and method for medical image analysis.
[0096] Since the system / device described in the above embodiments of the present invention is a system / device used to implement the method of the above embodiments of the present invention, a person skilled in the art can understand the specific structure and deformation of the system / device based on the method described in the above embodiments of the present invention, and thus will not be described in detail here. All systems / devices used in the method of the above embodiments of the present invention belong to the scope of protection of the present invention.
[0097] It should be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0098] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions.
[0099] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments after knowing the basic creative concept. Therefore, the present invention should be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.
[0100] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention belong to the scope of the present invention and its equivalent technology, the present invention should also include these modifications and variations.
Claims
1. A method for acupuncture technique recognition based on ultrasonic dynamic optical flow and self-supervised mask, characterized in that: include: Perform optical flow analysis on the acquired ultrasound image sequence, select key frame segments based on the difference in optical flow intensity between frames, and construct the input sequence by extending the buffer frames forward and backward; An overlapping sliding window is used to capture local segments on the input sequence, the forward and backward optical flows of adjacent frames in the local segments are calculated and bidirectional alignment is performed, direction-specific multi-head attention calculation is performed on the aligned frame group, and a time-varying probability mask is output; The time-varying probability mask is injected into the first layer of the spatiotemporal separation attention network for temporal attention calculation, and the gate control unit is used in the subsequent predetermined layers to modulate the first layer mask features with the current layer features through cross-layer mask conduction; The gated modulation features output from each level are spliced and input into the cascaded spatiotemporal pyramid pooling module to extract the time-sharing granular motion features. After cross-scale channel attention fusion, they are input into the multi-expert classifier to output the acupuncture technique recognition results.
2. The acupuncture technique recognition method based on ultrasonic dynamic optical flow and self-supervised mask according to claim 1, characterized in that: The obtained ultrasound image sequence is subjected to optical flow analysis, key frame segments are selected based on the difference in optical flow intensity between frames, and the input sequence is constructed by extending the buffer frames forward and backward, including: The original ultrasound video data is acquired and standardized, and the optical flow field is calculated frame by frame for the obtained ultrasound image sequence to generate a visual inter-frame optical flow image sequence with HSV color coding; The visual inter-frame optical flow image sequence is converted to the HSV color space and each channel is independently normalized. The global optical flow intensity of each frame is calculated and input into a Gaussian difference filter with time-series perception to filter out significant time points where the optical flow intensity difference value exceeds three times the average fluctuation level of adjacent frames. In the ultrasound image sequence, the continuous area with the significant time point as the center and extending to both sides of the time sequence until the rate of change of the optical flow intensity drops to the baseline level is taken as the key frame segment; Based on the key frame fragments, the buffer frames are adaptively expanded forward and backward bidirectionally to obtain the initial input sequence; The initial input sequence is tested for consistency of optical flow gradients. If the motion direction mutation between adjacent frames exceeds the preset mutation angle, an interpolation transition frame is inserted, and finally an input sequence that is resistant to motion breaks is output.
3. The acupuncture technique recognition method based on ultrasonic dynamic optical flow and self-supervised mask according to claim 2, characterized in that: The visual inter-frame optical flow image sequence is converted to the HSV color space and each channel is independently normalized. The global optical flow intensity of each frame is calculated and input into the time-series-aware Gaussian difference filter to screen the significant time points where the optical flow intensity difference value exceeds three times the average fluctuation level of adjacent frames, including: Convert the visual inter-frame optical flow image sequence to the HSV color space, and independently normalize the hue channel, saturation channel, and brightness channel of the HSV color space; Perform pixel-by-pixel multiplication and accumulation on the normalized values of the hue channel, saturation channel, and brightness channel to obtain the global optical flow intensity of each optical flow image that represents the dynamic change intensity of the muscle tissue, and then obtain a global optical flow intensity sequence; The first sliding time window is used to perform sliding average filtering on the global optical flow intensity sequence to generate a baseline trend component, and the global optical flow intensity sequence is subtracted from the baseline component to obtain a residual sequence; Applying a second sliding time window to the residual sequence, obtaining a second-order derivative response of the residual signal within the second sliding time window through a Gaussian difference filter to generate an enhanced residual sequence; In the enhanced residual sequence, the standard deviation of the optical flow intensity fluctuation of the N frames before and after the current frame is used as the benchmark. If the instantaneous difference value exceeds three times the standard deviation, it is marked as a significant time point.
4. The acupuncture technique recognition method based on ultrasonic dynamic optical flow and self-supervised mask according to claim 2, characterized in that: The adaptive extended buffer frame includes: Forward expansion intercepts an optical flow compensation frame group containing at least two consecutive frames before the key frame segment, and when the time interval between the starting point of the key frame segment and the ending point of the preceding key frame segment is less than a preset merging threshold, merges the optical flow compensation frame group into the backward expansion sequence of the preceding key frame segment; Backward expansion intercepts a motion attenuation observation frame group containing at least three consecutive frames after the key frame segment, and the interception termination condition of the motion attenuation observation frame group is that the amplitude of the optical flow between frames decays to below a preset amplitude threshold of the central frame amplitude of the key frame segment; When the starting position of the key frame segment is located at the starting point of the ultrasound image sequence, the optical flow field of the first frame is spatially mirrored along the main axis of the muscle fiber of the corresponding ultrasound image, and a random perturbation vector with an amplitude of 15%-25% of the standard deviation of the optical flow of the current frame is superimposed to generate a virtual forward compensation frame; When the key frame segment ends at the end of the ultrasound image sequence, the motion trajectory is extrapolated according to 1.2-1.8 times the duration of the optical flow vector of the last frame to generate a virtual backward observation frame.
5. The acupuncture technique recognition method based on ultrasonic dynamic optical flow and self-supervised mask according to claim 1, characterized in that: An overlapping sliding window is used to capture local segments on the input sequence. The forward and backward optical flows of adjacent frames in the local segments are calculated and bidirectional alignment is performed. Direction-specific multi-head attention calculation is performed on the aligned frame group. The output time-varying probability mask includes: Intercepting a local segment including at least one key frame segment, at least two forward buffer frames, and at least three backward buffer frames on the input sequence using a sliding window with an adaptive overlap rate; A lightweight optical flow network is used to perform forward and backward optical flow estimation on each adjacent frame of each local segment. Based on the forward optical flow, the previous frame is spatially deformed and mapped to the coordinate space of the current frame. Based on the backward optical flow, the next frame is reversely deformed and aligned to the coordinate space of the current frame. The aligned three frames are spliced along the channel dimension to generate a spatiotemporally aligned superimposed frame group. The spatiotemporally aligned superimposed frame groups are input in parallel into the lateral motion focus head group, the longitudinal motion focus head group and the compound motion parsing head group in the multi-head attention module for feature processing; The output features of the horizontal motion focused head group and the vertical motion focused head group are orthogonally projected and fused to generate a feature tensor representing the joint distribution of motion vectors; The output features of the composite parsing head group are concatenated with the feature tensor representing the joint distribution of motion vectors along the channel axis to generate cross-motion mode fusion features; Perform spatial pyramid pooling on the cross-motion mode fusion features, extract multi-granularity motion features at different scales, and use channel attention gating to weightedly fuse features of each scale, retaining cross-dimensional coupling features that represent horizontal muscle fiber contraction, vertical subcutaneous deformation, and non-directional motion, perform convolution and nonlinear activation on the coupling features, and output the initial mask; Based on the optical flow vectors in adjacent sliding windows, the overlapping areas of adjacent initial masks are fused across windows, and a morphological closing operation is performed on the fused masks to eliminate holes and isolated noise. The mask boundaries are refined through sub-pixel optical flow-guided interpolation, and a time-varying probability mask with continuous boundaries is output to characterize the probability distribution of subcutaneous tissue motion. in, The lateral motion focus head group is configured as a motion parsing module based on horizontal optical flow gradient constraints. By enhancing the attention weight of the horizontal motion component, the motion features related to the contraction of horizontal muscle fibers are specifically extracted. The horizontal optical flow gradient constraint is: when calculating the query-key similarity matrix, a dynamic weight gain is applied to the pixel points whose horizontal optical flow gradient component exceeds the preset ratio threshold of the vertical component. The vertical motion focus head group is configured as a motion analysis module based on the vertical optical flow gradient constraint. By enhancing the characteristic response intensity of the vertical optical flow gradient component, the motion features related to the vertical deformation of the subcutaneous tissue caused by the acupuncture insertion and lifting operation are specifically extracted; the vertical optical flow gradient constraint is: when calculating the attention weight, the exponential feature enhancement of the channel dimension is implemented for the area where the vertical optical flow gradient component exceeds the preset ratio threshold of the horizontal optical flow gradient component; The compound motion analysis head group is configured as a full-degree-of-freedom motion analysis module without directional constraints, which is used to capture nonlinear motion characteristics related to rotational motion, oblique shear motion and multi-directional compound motion.
6. The acupuncture technique recognition method based on ultrasonic dynamic optical flow and self-supervised mask according to claim 1, characterized in that: The time-varying probability mask is injected into the first layer of the spatiotemporal separation attention network for temporal attention calculation, and the gate control unit is used in the subsequent predetermined layers to modulate the first layer mask features with the current layer features for cross-layer mask conduction, including: The time-varying probability mask is input into the Patch embedding module synchronized with the spatiotemporal separation attention network. Through the two-dimensional convolution operation with preset block size and step length, the spatial area of the mask is processed into non-overlapping blocks to generate mask embedding blocks. Multiple mask embedding blocks are stacked sequentially along the time dimension, and the channel dimension is expanded and transformed through a learnable linear projection layer to generate a mask feature tensor consistent with the dimension of the spatiotemporal attention network input tensor; In the temporal attention calculation stage of the first layer of the spatiotemporal separation attention network, the mask feature tensor is added to the original query tensor element by element, and then a layer normalization operation is performed to tilt the temporal attention weight toward the motion probability area marked by the mask that exceeds the preset probability threshold; Through the gated recalibration units set at the 2nd, 4th and 6th layers, the mask feature tensor is reduced in dimension by depthwise separable convolution and then transferred across layers, and the query tensor of the current layer is input into the two-level fully connected network to generate the gated weights; The gated weights are used to dynamically scale the reduced-dimensional mask features and superimposed on the current layer features through residual connections to obtain gated modulation features containing mask-guided information. Through the pre-deployed auxiliary supervision branch, the mutual information entropy of the deep features and the first-layer mask feature tensor is calculated. If the entropy value is lower than the set entropy threshold, the gated weight backtracking adjustment is triggered.
7. The acupuncture technique recognition method based on ultrasonic dynamic optical flow and self-supervised mask according to claim 1, characterized in that: The gated modulation features output from each level are spliced and input into the cascaded spatiotemporal pyramid pooling module to extract the time-sharing granular motion features. After cross-scale channel attention fusion, they are input into the multi-expert classifier. The output acupuncture manipulation recognition results include: The gated modulation features output by each layer are concatenated along the channel dimension to form a fused feature tensor containing shallow motion details and deep semantic associations, and batch normalization and spatiotemporal dimension reorganization are performed to match the input structure of the cascaded spatiotemporal pyramid. The fused feature tensor after batch normalization and spatiotemporal dimension reorganization is input into the spatiotemporal pyramid module, and dense sliding maximum pooling is performed in the 4-8 frame window to extract the muscle fiber tremor characteristics caused by the instantaneous effect of acupuncture to obtain short-term granularity features, sparse self-attention pooling is performed in the 16-24 frame window to obtain the movement rhythm characteristics of a single operation of lifting, inserting or twisting to obtain medium-term granularity features, and adaptive average pooling is performed in the 32-48 frame window to extract the cumulative biomechanical characteristics under the continuous action of multiple manipulations to obtain long-term granularity features; The short-term, medium-term, and long-term granular features are multi-scale aligned and fused with channel attention weighted to obtain the final fused features; The final fusion features are input into parallel optical flow motion experts, tissue deformation experts and spatiotemporal coupling experts, and weighted voting is performed based on the confidence scores output by each expert to output the probability distribution of acupuncture technique categories.
8. An acupuncture technique recognition system based on ultrasonic dynamic optical flow and self-supervised mask, characterized in that: include: An input sequence construction module is used to perform optical flow analysis on the acquired ultrasound image sequence, select key frame segments based on the difference in optical flow intensity between frames, and construct an input sequence by extending the buffer frames forward and backward; The mask output module is used to intercept local segments on the input sequence using overlapping sliding windows, obtain the forward and backward optical flows of adjacent frames in the local segments and perform bidirectional alignment, perform direction-specific multi-head attention calculation on the aligned frame group, and output a time-varying probability mask; A cross-layer mask conduction modulation module is used to inject the time-varying probability mask into the first layer of the spatiotemporal separation attention network for temporal attention calculation, and use a gating unit in the subsequent predetermined layers to perform cross-layer mask conduction modulation on the first layer mask features and the current layer features; The technique recognition module concatenates the gated modulation features output by each level and inputs them into the cascaded spatiotemporal pyramid pooling module to extract the time-sharing granular motion features. After cross-scale channel attention fusion, they are input into the multi-expert classifier to output the acupuncture technique recognition results.
9. An acupuncture technique recognition device based on ultrasonic dynamic optical flow and self-supervised mask, characterized in that: include: at least one processor; And a memory communicatively connected to at least one processor; wherein the memory stores instructions executable by at least one processor, and the instructions are executed by at least one processor so that at least one processor can execute the acupuncture technique recognition method based on ultrasonic dynamic optical flow and self-supervised mask as described in any one of claims 1-7.
10. A computer-readable storage medium having computer-executable instructions stored thereon, characterized in that: When the executable instructions are executed by the processor, the acupuncture manipulation recognition method based on ultrasonic dynamic optical flow and self-supervised mask as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-mode acupuncture manipulation recognition method and system fusing vision and touch
CN116597517A
Acupuncture manipulation recognition method, system and equipment based on subcutaneous needle feeling and medium
CN118038076A
System and method for self-supervised video transformer
US20240169692A1
Cited By
Semi-autonomous remote sensing water area generalization segmentation method fused with multispectral inversion
CN120808175A