AI-Based Intelligent Lesion Detection Method in Virtual Hysteroscopic Surgery Training

By adopting AI-based lesion intelligent detection method in virtual hysteroscopic surgery training, using the LSTM layer and Transformer encoder to process surgical video features, the limitations of lesion recognition in the virtual environment are solved, achieving more accurate lesion detection and reducing the dependence of manual guidance.

CN120125921BActive Publication Date: 2025-07-18JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510620037.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-07-18
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The existing virtual hysteroscopic training methods have limitations in lesion identification and diagnosis. Students lack real-time and intuitive lesion recognition feedback in the virtual environment, making it difficult for them to understand immediately whether their lesion identification is accurate during the operation.

Method used

Using AI-based intelligent lesion detection method, the surgical process context recognition module is constructed through the LSTM layer, the fully connected layer and the Softmax activation function, combined with the discrete cosine change mechanism and the interactive perceptual feature approximation module, the surgical video features are processed using the Transformer encoder to generate lesion probability and uncertainty scores.

Benefits of technology

It improves the accuracy of lesion detection, reduces the dependence on manual guidance, provides richer temporal and spatial dimension information, and enhances students' lesion recognition ability in virtual environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125921B_ABST
    Figure CN120125921B_ABST
Patent Text Reader

Abstract

The present invention proposes an AI-based intelligent lesion detection method in virtual hysteroscopic surgery training. The method includes: obtaining an input feature sequence, constructing a surgical procedure context recognition module based on an LSTM layer, a fully connected layer, and a Softmax activation function, using the surgical procedure context recognition module to process the input feature sequence to obtain surgical stage probabilities, constructing a discrete cosine transform channel attention module based on a discrete cosine transform mechanism, a fully connected layer, and a layer normalization mechanism, using the discrete cosine transform channel attention module to process the input feature sequence to obtain features weighted by channel attention. The present invention designs an interactive perception feature approximation module. By analyzing the feature changes between consecutive frames, the dynamic information of changes in surgical operations or lesion states is captured. The interactive perception feature approximation module can reflect the operation behaviors and lesion change trends during the surgical process, providing richer spatio-temporal dimensional information for lesion detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical training, and particularly relates to an AI-based intelligent lesion detection method in virtual hysteroscopic surgery training. Background Art

[0002] Hysteroscopic surgery is a commonly used minimally invasive gynecological surgery for diagnosing and treating various intrauterine diseases, such as endometrial polyps, uterine fibroids, intrauterine adhesions, and early endometrial cancer. For obstetricians and gynecologists, it is crucial to master the operation skills of hysteroscopy and accurately identify various lesions in the uterine cavity.

[0003] Traditional hysteroscopic surgery training usually relies on the direct guidance of a tutor and the practice of trainees in real surgeries. However, the opportunities for real surgeries are limited and there are certain risks. As an emerging educational method, virtual hysteroscopic surgery training provides a safe and repeatable practice environment for trainees by simulating real surgical scenarios. Trainees can familiarize themselves with the surgical procedures and practice operation skills in the virtual environment, thus making up for the deficiencies of traditional training to a certain extent.

[0004] However, the existing virtual hysteroscopic surgery training methods still have some limitations in lesion identification and diagnosis. When trainees operate in a virtual environment, they often lack real-time and intuitive feedback on lesion identification. Tutors usually need to conduct evaluations afterwards or provide manual guidance during the operation, which reduces the training efficiency and realism to a certain extent. It is difficult for trainees to immediately understand whether their identification of lesions is accurate during the operation, and it is also difficult for them to effectively learn the characteristics and diagnostic key points of different lesions. Summary of the Invention

[0005] In view of the above situation, the main purpose of the present invention is to propose an AI-based intelligent lesion detection method in virtual hysteroscopic surgery training to solve the above technical problems.

[0006] The present invention proposes an AI-based intelligent lesion detection method in virtual hysteroscopic surgery training, and the method includes the following steps:

[0007] Step 1, obtain an input feature sequence, construct a surgical procedure context recognition module based on an LSTM layer, a fully connected layer, and a Softmax activation function, and use the surgical procedure context recognition module to process the input feature sequence to obtain the surgical stage probability;

[0008] Step 2, construct a discrete cosine transform channel attention module based on a discrete cosine transform mechanism, a fully connected layer, and a layer normalization mechanism, and use the discrete cosine transform channel attention module to process the input feature sequence to obtain the feature weighted by channel attention;

[0009] Step 3: Construct an interactive perception feature approximation module based on a feature difference calculation mechanism, a fully connected layer, and a ReLU activation function, and use the interactive perception feature approximation module to process the features weighted by channel attention to obtain interactive perception features;

[0010] Step 4: Construct an interactive-guided attention mechanism module based on an attention mechanism and a linear layer, and use the interactive-guided attention mechanism module to process the surgical stage probability, the features weighted by channel attention, and the interactive perception features to obtain attention-weighted features;

[0011] Step 5: Sequentially process the attention-weighted features through a Transformer encoder and average pooling to obtain the final feature representation:

[0012] Step 6: Sequentially process the final feature representation through a ReLU activation function, a fully connected layer, and a Sigmoid activation function to obtain the lesion probability and the uncertainty score respectively.

[0013] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0014] 1. The present invention uses a surgical process context recognition module to understand the current surgical stage, incorporates the context information of the surgical process into the lesion detection process, and combines this context information with visual features, thereby improving the accuracy of lesion detection and reducing the dependence on manual guidance;

[0015] 2. The present invention analyzes the feature changes between consecutive frames through an interactive perception feature approximation module, captures the dynamic information of surgical operations or lesion states. The interactive perception feature approximation module can reflect the operation behaviors and the change trends of lesions during the surgical process, providing richer spatio-temporal dimension information for lesion detection.

[0016] 3. The present invention uses a Transformer encoder to process the attention-weighted features, which can effectively capture the long-range dependencies within the attention-weighted features, thereby better understanding the temporal information in the surgical video or virtual environment and providing a comprehensive feature representation for the final lesion detection.

[0017] The additional aspects and advantages of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flowchart of the steps of the AI-based intelligent lesion detection method in virtual hysteroscopic surgery training proposed by the present invention;

[0019] Figure 2This is the method framework diagram of the AI-based intelligent lesion detection method in virtual hysteroscopic surgery training proposed by the present invention;

[0020] Figure 3 This is the module architecture diagram of the AI-based intelligent lesion detection method in virtual hysteroscopic surgery training proposed by the present invention. Specific embodiments

[0021] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.

[0022] These and other aspects of the embodiments of the present invention will be clear from the following description and drawings. In these descriptions and drawings, some specific embodiments of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0023] Please refer to Figure 1 , this embodiment provides an AI-based intelligent lesion detection method in virtual hysteroscopic surgery training. The method includes the following steps:

[0024] Step 1: Obtain an input feature sequence, construct a surgical procedure context recognition module based on an LSTM layer, a fully connected layer, and a Softmax activation function, and use the surgical procedure context recognition module to process the input feature sequence to obtain the surgical stage probability.

[0025] Please refer to Figure 2 and Figure 3 , in Step 1, obtain an input feature sequence, construct a surgical procedure context recognition module based on an LSTM layer, a fully connected layer, and a Softmax activation function, and use the surgical procedure context recognition module to process the input feature sequence to obtain the surgical stage probability, which specifically includes the following sub-steps:

[0026] Obtain an input feature sequence, input the input feature sequence into the LSTM layer for processing to obtain the output sequence of the LSTM. The following relationship exists in the corresponding process:

[0027] ;

[0028] Among them, represents the output sequence of the LSTM, represents the hidden state of the last time step, represents the cell state of the last time step, Indicates being processed by the LSTM layer, Indicates the input feature sequence and , Indicates the batch size, Indicates the sequence length, Indicates the video feature dimension;

[0029] Based on the output sequence of the LSTM, the output of the last time step is successively passed through the first fully connected layer and the Softmax activation process to obtain the surgical stage probability. There are the following relational expressions in the corresponding process:

[0030] ;

[0031] Among them, Indicates the log-odds of the surgical stage and , Indicates 5 different process stages included in the virtual hysteroscopic surgery training (introducing basic knowledge, instrument operation and navigation, diagnostic hysteroscopic examination procedures, therapeutic hysteroscopic examination procedures, advanced scenarios and complication management), Indicates the output of the last time step in the output sequence of the LSTM, Indicates the weight matrix of the first fully connected layer, Indicates the bias vector of the first fully connected layer, Indicates the surgical stage probability and , Indicates being processed by the Softmax activation.

[0032] It should be noted that the input feature sequence is the hysteroscopic surgery training video stream. The surgical process context recognition module is used to receive the input feature sequence and extract the context information of the surgical process based on the time series, and output the probability distribution representing the current surgical stage.

[0033] Step 2: Construct a discrete cosine transform channel attention module based on the discrete cosine transform mechanism, the fully connected layer, and the layer normalization mechanism, and use the discrete cosine transform channel attention module to process the input feature sequence to obtain the feature weighted by channel attention;

[0034] In step 2, a discrete cosine transform channel attention module is constructed based on the discrete cosine transform mechanism, the fully connected layer, and the layer normalization mechanism, and the discrete cosine transform channel attention module is used to process the input feature sequence to obtain the feature weighted by channel attention, which specifically includes the following sub-steps:

[0035] Perform discrete cosine transform processing on the input feature sequence to obtain discrete cosine transform coefficients. There are the following relational expressions in the corresponding process:

[0036] ;

[0037] Among them, represents the th coefficient in the th channel after the input feature sequence undergoes discrete cosine transform, represents the th element in the th channel of the input feature, represents the sequence length, represents the cosine function, represents pi, represents the index of the channel, represents the index of the coefficient, represents the index of the element;

[0038] The discrete cosine transform coefficients are successively passed through the second fully connected layer, Dropout layer, and ReLU activation process to obtain the output features after passing through the ReLU activation function. The following relationship exists in the corresponding process:

[0039] ;

[0040] Among them, represents the output features of the second fully connected layer, represents the discrete cosine transform coefficients, represents the weight matrix of the second fully connected layer, represents the bias vector of the second fully connected layer, represents the output features after passing through the Dropout layer, represents performing the Dropout operation, represents the dropout rate of Dropout, represents the output features after passing through the ReLU activation function, represents passing through the ReLU activation function;

[0041] The output features after passing through the ReLU activation function are successively passed through the third fully connected layer, Sigmoid, and layer normalization process to obtain the channel weights after layer normalization. The following relationship exists in the corresponding process:

[0042] ;

[0043] Among them, represents the output features of the third fully connected layer, represents the weight matrix of the third fully connected layer, represents the bias vector of the third fully connected layer, represents the channel weights, Indicates after being processed by the Sigmoid function, represents the channel weights after layer normalization, represents the layer normalization operation;

[0044] The input feature sequence is weighted and fused with the channel weights after layer normalization to obtain the features weighted by channel attention. The following relational expressions exist in the corresponding process:

[0045] ;

[0046] Among them, represents the features weighted by channel attention, represents element-wise multiplication.

[0047] It should be noted that the discrete cosine transform channel attention module is used to perform discrete cosine transform on each channel of the input features and learn the attention weights between channels to enhance the feature representation of key channels.

[0048] Step 3: Construct an interaction-aware feature approximation module based on the feature difference calculation mechanism, fully connected layer, and ReLU activation function, and use the interaction-aware feature approximation module to process the features weighted by channel attention to obtain interaction-aware features;

[0049] In Step 3, an interaction-aware feature approximation module is constructed based on the feature difference calculation mechanism, fully connected layer, and ReLU activation function, and the interaction-aware feature approximation module is used to process the features weighted by channel attention to obtain interaction-aware features, which specifically includes the following sub-steps:

[0050] Based on the features weighted by channel attention, calculate the feature difference between adjacent frames to obtain the difference features. The following relational expressions exist in the corresponding process:

[0051] ;

[0052] Among them, represents the difference features, represents all the features from the second frame to the last frame, represents all the features from the first frame to the penultimate frame;

[0053] Concatenate all the features from the first frame to the penultimate frame with the difference features to obtain the concatenated feature tensor. The following relational expressions exist in the corresponding process:

[0054] ;

[0055] Among them, represents the concatenated feature tensor, represents the feature concatenation operation, Indicates the dimension of splicing;

[0056] The spliced feature tensor is successively processed by the fourth fully connected layer and the ReLU activation function to obtain interaction-aware features. The following relational expressions exist in the corresponding process:

[0057] ;

[0058] Among them, represents the interaction-aware feature and , represents the weight matrix of the fourth fully connected layer, represents the transpose, represents the bias vector of the fourth fully connected layer.

[0059] It should be noted that the interaction-aware feature approximation module is used to analyze the feature changes between consecutive frames, capture the dynamic information of surgical operations or lesion states, and generate interaction-aware features.

[0060] Step 4: Construct an interaction-guided attention mechanism module based on the attention mechanism and the linear layer, and use the interaction-guided attention mechanism module to process the surgical stage probability, the feature weighted by channel attention, and the interaction-aware feature to obtain the attention-weighted feature;

[0061] In Step 4, an interaction-guided attention mechanism module is constructed based on the attention mechanism and the linear layer, and the interaction-guided attention mechanism module is used to process the surgical stage probability, the feature weighted by channel attention, and the interaction-aware feature to obtain the attention-weighted feature, which specifically includes the following sub-steps:

[0062] Repeat the surgical stage probability a certain number of times in the time dimension to obtain the repeated surgical stage probability in the time dimension. Perform context fusion processing on the repeated surgical stage probability in the time dimension and the feature weighted by channel attention to obtain the context-fused visual feature. The following relational expressions exist in the corresponding process:

[0063] ;

[0064] Among them, represents the repeated surgical stage probability in the time dimension, represents the index of the number of repetitions, represents the total number of repetitions, represents the context-fused visual feature;

[0065] It should be noted that repeating the surgical stage probability several times in the time dimension is to match the shape of the features weighted by channel attention. The number of times the surgical stage probability is repeated in the time dimension is equal to the time dimension of the features weighted by channel attention.

[0066] The visually fused features are processed through the fifth fully connected layer and the ReLU activation function in sequence to obtain the features after dimensionality reduction mapping. The following relational expression exists in the corresponding process:

[0067] ;

[0068] Wherein, represents the features after dimensionality reduction mapping, represents the weight matrix of the fifth fully connected layer, represents the bias vector of the fifth fully connected layer;

[0069] The interaction-aware features are processed through a linear layer and the Sigmoid activation function in sequence to obtain the attention weights. The following relational expression exists in the corresponding process:

[0070] ;

[0071] Wherein, represents the output of the linear layer, represents the weight matrix of the linear layer, represents the bias vector of the linear layer, represents the attention weights and ;

[0072] The attention weights and the features after dimensionality reduction mapping are weighted and fused to obtain the features after attention weighting. The following relational expression exists in the corresponding process:

[0073] ;

[0074] Wherein, represents the features after attention weighting.

[0075] It should be noted that in order to match the time dimensions of and , can be padded or cropped. The interaction-guided attention mechanism module is used to utilize the output of the interaction-aware feature approximation module to guide the attention to the visually fused features and highlight the feature regions related to the interaction behavior.

[0076] Step 5: The features after attention weighting are processed through the Transformer encoder and average pooling in sequence to obtain the final feature representation;

[0077] In step 5, the attention-weighted features are successively processed by a Transformer encoder and average pooling to obtain the final feature representation, which specifically includes the following sub-steps:

[0078] The attention-weighted features are encoded by the Transformer encoder to obtain the encoded features. The following relationship exists during the corresponding process:

[0079] ;

[0080] where represents the encoded features and , represents being encoded by the Transformer encoder;

[0081] The encoded features are averaged and pooled in the time series dimension to obtain the final feature representation. The following relationship exists during the corresponding process:

[0082] ;

[0083] where represents the value of the -th element of the final feature representation of the -th sample in the batch, represents the index in the batch dimension and , represents the index in the time series dimension, represents the index in the feature dimension and .

[0084] It should be noted that the Transformer encoder is used to encode the attention-weighted feature sequence, capture the long-range dependencies within the sequence, and generate the final feature representation for lesion detection.

[0085] Step 6: The final feature representation is successively processed by the ReLU activation function, a fully connected layer, and the Sigmoid activation function to obtain the lesion probability and uncertainty score respectively.

[0086] In step 6, the final feature representation is successively processed by the ReLU activation function, a fully connected layer, and the Sigmoid activation function to obtain the lesion probability and uncertainty score respectively, which specifically includes the following sub-steps:

[0087] The final feature representation is successively processed by the ReLU activation function, the sixth fully connected layer, and the Sigmoid activation function to obtain the lesion probability. The following relationship exists during the corresponding process:

[0088] ;

[0089] Among them, represents the output after being processed by the ReLU activation function, represents the final feature representation and , represents the output of the sixth fully connected layer, represents the weight matrix of the sixth fully connected layer, represents the bias vector of the sixth fully connected layer, represents the lesion probability and ;

[0090] The final feature representation is successively processed by the seventh fully connected layer and the Sigmoid activation function to obtain an uncertainty score. There are the following relational expressions in the corresponding process:

[0091] ;

[0092] Among them, represents the output of the seventh fully connected layer, represents the weight matrix of the seventh fully connected layer, represents the bias vector of the seventh fully connected layer, represents the uncertainty score and .

[0093] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are successively shown according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps does not have a strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0094] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following well-known technologies in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0095] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0096] The above-described embodiments only represent several implementation manners of the present invention. The descriptions are relatively specific and detailed, but should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention should be subject to the appended claims.

Claims

1. An AI-based intelligent lesion detection method in virtual hysteroscopic surgery training, characterized in that, The method includes the following steps: Step 1: Obtain an input feature sequence, construct a surgical procedure context recognition module based on an LSTM layer, a fully connected layer, and a Softmax activation function, and use the surgical procedure context recognition module to process the input feature sequence to obtain surgical stage probabilities; Step 2: Construct a discrete cosine transform channel attention module based on a discrete cosine transform mechanism, a fully connected layer, and a layer normalization mechanism, and use the discrete cosine transform channel attention module to process the input feature sequence to obtain features weighted by channel attention; Step 3: Construct an interactive perception feature approximation module based on a feature difference calculation mechanism, a fully connected layer, and a ReLU activation function, and use the interactive perception feature approximation module to process the features weighted by channel attention to obtain interactive perception features; Step 4: Construct an interaction-guided attention mechanism module based on an attention mechanism and a linear layer, and use the interaction-guided attention mechanism module to process the surgical stage probabilities, the features weighted by channel attention, and the interactive perception features to obtain features weighted by attention; Step 5: Sequentially pass the features weighted by attention through a Transformer encoder and average pooling to obtain a final feature representation; Step 6: Sequentially pass the final feature representation through a ReLU activation function, a fully connected layer, and a Sigmoid activation function to obtain lesion probabilities and uncertainty scores respectively.

2. The AI-based intelligent lesion detection method in virtual hysteroscopy surgery training according to claim 1, wherein, In Step 1, to obtain an input feature sequence, construct a surgical procedure context recognition module based on an LSTM layer, a fully connected layer, and a Softmax activation function, and use the surgical procedure context recognition module to process the input feature sequence to obtain surgical stage probabilities, the following specific sub-steps are included: Obtain an input feature sequence, input the input feature sequence into the LSTM layer for processing to obtain an output sequence of the LSTM. There are the following relational expressions during the corresponding process: ; Among them, represents the output sequence of the LSTM, represents the hidden state at the last time step, represents the cell state at the last time step, represents the input feature sequence, represents being processed by the LSTM layer; Based on the output sequence of the LSTM, sequentially pass the output of the last time step through a first fully connected layer and Softmax activation processing to obtain surgical stage probabilities. There are the following relational expressions during the corresponding process: ; Among them, represents the log odds of the surgical stage, represents the output of the last time step in the output sequence of the LSTM, represents the weight matrix of the first fully connected layer, represents the bias vector of the first fully connected layer, represents the surgical stage probability, represents being processed by Softmax activation.

3. The AI-based intelligent lesion detection method in virtual hysteroscopy surgery training according to claim 2, characterized in that, In Step 2, to construct a discrete cosine transform channel attention module based on a discrete cosine transform mechanism, a fully connected layer, and a layer normalization mechanism, and use the discrete cosine transform channel attention module to process the input feature sequence to obtain features weighted by channel attention, the following specific sub-steps are included: Perform discrete cosine transform processing on the input feature sequence to obtain discrete cosine transform coefficients; Sequentially pass the discrete cosine transform coefficients through a second fully connected layer, a Dropout layer, and ReLU activation processing to obtain output features after passing through the ReLU activation function; Sequentially pass the output features after passing through the ReLU activation function through a third fully connected layer, Sigmoid, and layer normalization processing to obtain channel weights after layer normalization; Perform weighted fusion of the input feature sequence and the channel weights after layer normalization to obtain features weighted by channel attention.

4. The AI-based intelligent lesion detection method in virtual hysteroscopy surgery training according to claim 3, wherein The input feature sequence is processed by discrete cosine transform to obtain discrete cosine transform coefficients, and the following relational expressions exist in the corresponding process: ; Among them, represents the -th coefficient in the -th channel after the input feature sequence undergoes discrete cosine transform, represents the -th element in the -th channel of the input feature, represents the sequence length, represents the cosine function, represents pi, represents the index of the channel, represents the index of the coefficient, represents the index of the element; In the step of successively passing the discrete cosine transform coefficients through the second fully connected layer, Dropout layer, and ReLU activation processing to obtain the output features after passing through the ReLU activation function, the following relational expressions exist in the corresponding process: ; Among them, represents the output features of the second fully connected layer, represents the discrete cosine transform coefficients, represents the weight matrix of the second fully connected layer, represents the bias vector of the second fully connected layer, represents the output features after passing through the Dropout layer, represents performing the Dropout operation, represents the dropout rate of Dropout, represents the output features after passing through the ReLU activation function, represents being processed by the ReLU activation function; In the step of successively passing the output features after passing through the ReLU activation function through the third fully connected layer, Sigmoid, and layer normalization processing to obtain the channel weights after layer normalization, the following relational expressions exist in the corresponding process: ; Among them, represents the output features of the third fully connected layer, represents the weight matrix of the third fully connected layer, represents the bias vector of the third fully connected layer, represents the channel weight, represents being processed by the Sigmoid function, represents the channel weight after layer normalization, represents the layer normalization operation; In the step of weighted fusion of the input feature sequence and the channel weights after layer normalization to obtain the features weighted by channel attention, the following relational expressions exist in the corresponding process: ; Among them, represents the feature weighted by channel attention, represents element-wise multiplication.

5. The AI-based intelligent lesion detection method in virtual hysteroscopy surgery training according to claim 4, wherein, In step 3, an interactive perception feature approximation module is constructed based on the feature difference calculation mechanism, fully connected layer, and ReLU activation function, and the features weighted by channel attention are processed by the interactive perception feature approximation module to obtain interactive perception features, which specifically include the following sub-steps: Based on the features weighted by channel attention, calculate the feature difference between adjacent frames to obtain the difference features; Concatenate all the features from the first frame to the penultimate frame with the difference features to obtain the concatenated feature tensor; Successively pass the concatenated feature tensor through the fourth fully connected layer and ReLU activation function processing to obtain interactive perception features.

6. The AI-based intelligent lesion detection method in virtual hysteroscopy surgery training according to claim 5, wherein Based on the features weighted by channel attention, calculate the feature difference between adjacent frames to obtain the difference features, and the following relational expressions exist in the corresponding process: ; Among them, represents the differential feature, represents all features from the second frame to the last frame, represents all features from the first frame to the penultimate frame; In the step of concatenating all the features from the first frame to the penultimate frame with the difference features to obtain the concatenated feature tensor, the following relational expressions exist in the corresponding process: ; Among them, represents the concatenated feature tensor, represents the feature concatenation operation, represents the dimension of concatenation; In the step of successively passing the concatenated feature tensor through the fourth fully connected layer and ReLU activation function processing to obtain interactive perception features, the following relational expressions exist in the corresponding process: ; Among them, represents the interactive perception feature, represents the weight matrix of the fourth fully connected layer, represents the transpose, represents the bias vector of the fourth fully connected layer.

7. The AI-based intelligent lesion detection method in virtual hysteroscopy surgery training according to claim 6, wherein In step 4, an interactive-guided attention mechanism module is constructed based on the attention mechanism and linear layer, and the surgical stage probability, the features weighted by channel attention, and the interactive perception features are processed by the interactive-guided attention mechanism module to obtain the features weighted by attention, which specifically include the following sub-steps: Repeat the surgical stage probability a certain number of times in the time dimension to obtain the surgical stage probability repeated in the time dimension, and perform context fusion processing on the surgical stage probability repeated in the time dimension and the features weighted by channel attention to obtain the visually context-fused features; Successively pass the context-fused visual features through the fifth fully connected layer and ReLU activation function processing to obtain the features after dimensionality reduction mapping; Successively pass the interactive perception features through the linear layer and Sigmoid activation function processing to obtain the attention weights; Perform weighted fusion of the attention weights and the features after dimensionality reduction mapping to obtain the features weighted by attention.

8. The AI-based intelligent lesion detection method in virtual hysteroscopy surgery training according to claim 7, wherein Repeat the surgical stage probability a certain number of times in the time dimension to obtain the repeated surgical stage probability in the time dimension. Perform context fusion processing on the repeated surgical stage probability in the time dimension and the features weighted by channel attention to obtain the visually characterized after context fusion. There are the following relational expressions in the corresponding process: ; Among them, represents the surgical stage probability after repetition in the time dimension, represents the index of the number of repetitions, represents the total number of repetitions, represents the visual features after context fusion; In the step of sequentially passing the visually characterized after context fusion through the fifth fully connected layer and the ReLU activation function to obtain the characterized after dimensionality reduction mapping, there are the following relational expressions in the corresponding process: ; Among them, represents the feature after dimensionality reduction mapping, represents the weight matrix of the fifth fully connected layer, represents the bias vector of the fifth fully connected layer; In the step of sequentially passing the interaction-aware characterized through a linear layer and the Sigmoid activation function to obtain the attention weight, there are the following relational expressions in the corresponding process: ; Among them, represents the output of the linear layer, represents the weight matrix of the linear layer, represents the bias vector of the linear layer, represents the attention weight; In the step of performing weighted fusion on the attention weight and the characterized after dimensionality reduction mapping to obtain the characterized after attention weighting, there are the following relational expressions in the corresponding process: ; Among them, represents the feature after attention weighting.

9. The AI-based intelligent lesion detection method in virtual hysteroscopy surgery training according to claim 8, wherein, In step 5, sequentially pass the characterized after attention weighting through a Transformer encoder and average pooling processing to obtain the final characterized representation, which specifically includes the following sub-steps: Use the Transformer encoder to encode the characterized after attention weighting to obtain the encoded characterized. There are the following relational expressions in the corresponding process: ; Among them, represents the encoded feature, indicating that it has been encoded by the Transformer encoder; Perform average pooling processing on the encoded characterized in the time series dimension to obtain the final characterized representation. There are the following relational expressions in the corresponding process: ; Among them, represents the value of the -th element of the final feature representation of the -th sample in the batch, represents the index of the batch dimension, represents the index of the time series dimension, represents the index of the feature dimension.

10. The AI-based intelligent lesion detection method in virtual hysteroscopic surgery training according to claim 9, wherein In step 6, sequentially pass the final characterized representation through the ReLU activation function, the fully connected layer, and the Sigmoid activation function to respectively obtain the lesion probability and the uncertainty score. It specifically includes the following sub-steps: Sequentially pass the final characterized representation through the ReLU activation function, the sixth fully connected layer, and the Sigmoid activation function to obtain the lesion probability. There are the following relational expressions in the corresponding process: ; Among them, represents the output after being processed by the ReLU activation function, represents the final feature representation, represents the output of the sixth fully connected layer, represents the weight matrix of the sixth fully connected layer, represents the bias vector of the sixth fully connected layer, represents the lesion probability; Sequentially pass the final characterized representation through the seventh fully connected layer and the Sigmoid activation function to obtain the uncertainty score. There are the following relational expressions in the corresponding process: ; Among them, represents the output of the seventh fully connected layer, represents the weight matrix of the seventh fully connected layer, represents the bias vector of the seventh fully connected layer, represents the uncertainty score.

Citation Information

Patent Citations

  • Method and terminal for detecting time series data exception by combining attention mechanism and LSTM (Long Short Term Memory)

    CN115983087A

  • Dynamic gesture recognition method based on hand key point and double-layer bidirectional LSTM network

    CN117576783A