Gastroenterology department operation decision-making auxiliary system based on machine vision

Through a cascading global-local context coding structure and a nested Transformer encoder, combined with deformable attention and gradient guidance module, iterative progressive optimization and multi-layer LSTM network are designed, which solves the problems of missegment, lesion boundary distortion and abnormal identification accuracy in surgical image segmentation, and realizes accurate lesion recognition and abnormal detection.

CN120279345AInactive Publication Date: 2025-07-08THE SIXTH MEDICAL CENT OF THE CHINESE PEOPLES LIBERATION ARMY GENERAL HOSPITAL
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510768179.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing surgical image segmentation technology has problems such as poor global context understanding ability, weak lesion boundary aberration and recognition ability, and low recognition accuracy of abnormal areas, resulting in unstable missegmentation, missegmentation and abnormal judgment.

Method used

The cascading global-local context coding structure is adopted to combine a nested Transformer encoder and a local boundary perception path, and a deformable attention mechanism and gradient-guided edge enhancement module are introduced, an iterative progressive optimization mechanism and a multi-layer stacked LSTM network are designed, and the local and global attention mechanisms are combined for accurate identification and abnormal detection of lesion areas.

Benefits of technology

It improves the accuracy of lesion contour recognition and segmentation accuracy, realizes accurate quantification and visualization prompts of lesion abnormalities, and enhances the ability to respond to sudden abnormalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279345A_ABST
    Figure CN120279345A_ABST
Patent Text Reader

Abstract

The invention relates to the field of machine vision, in particular to a digestive system department operation decision auxiliary system based on machine vision, which comprises an image acquisition subsystem, an image preprocessing subsystem, a target detection and segmentation subsystem, a semantic recognition subsystem and a visual prompt subsystem, according to the method, a cascaded global-local context coding structure is introduced, and a nested Transform encoder and a local boundary sensing path are combined in a segmentation network, so that joint modeling of global anatomical semantics and local boundary details is realized, and the accuracy of focus contour recognition is improved; according to the invention, an iterative progressive optimization mechanism is designed, and the preliminary mask is refined layer by layer by a plurality of refined segmentation units, so that the mask is refined step by step; according to the method, local and global attention mechanisms are introduced through a multi-layer stacked LSTM network and an optimization mechanism, and meanwhile, accurate quantification and visual prompt of focus anomalies are realized by combining anomaly detection and structured output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine vision, and specifically refers to a decision-making assistance system for digestive endoscopy surgery based on machine vision. Background Art

[0002] In recent years, with the development of computer vision technology, the segmentation and analysis of medical surgical images have become an important research direction for intelligent assisted diagnosis and treatment. However, in the existing surgical image segmentation technologies, the CNN structure or the single-path Transformer structure is used for segmentation, which has problems such as poor global context understanding ability and insufficient differentiation of complex tissue structures, resulting in mis-segmentation and missed-segmentation; most of the existing segmentation methods adopt a one-time output method, which has problems such as distortion of the lesion boundary and weak recognition ability, and it is easy to lose the target on the tiny lesions in the approximate background area; at present, the recognition of abnormal regions in medical images is mostly based on static feature discrimination, lacking temporal feature modeling, with the problem of low abnormal judgment accuracy, unstable performance in real surgeries, and slow response to sudden abnormal changes. Summary of the Invention

[0003] In view of the above situation, the present invention provides a decision-making assistance system for digestive endoscopy surgery based on machine vision. Aiming at the situation of mis-segmentation and missed-segmentation in the surgical image segmentation technology, the present invention introduces a cascaded global-local context encoding structure, combines a nested Transformer encoder and a local boundary perception path in the segmentation network, realizes the joint modeling of global anatomical semantics and local boundary details, and uses a deformable attention mechanism and a gradient-guided edge enhancement module to improve the accuracy of lesion contour recognition; aiming at the problems of lesion boundary distortion and weak recognition ability in the segmentation method, the present invention designs an iterative progressive optimization mechanism, and uses multiple refined segmentation units to gradually refine the initial mask layer by layer. In each round, the original features and semantic context are fused to gradually refine the mask, realizing a "coarse to fine" continuous enhancement strategy, and improving the accuracy and continuity of segmentation; aiming at the problem of low judgment accuracy in the existing abnormal region recognition method, the present invention uses a multi-layer stacked LSTM network and an optimization mechanism, introduces local and global attention mechanisms, focuses on abnormal key features, enhances the discrimination ability, and at the same time combines abnormal detection and structured output to realize the accurate quantification and visual prompt of lesion abnormalities.

[0004] A decision-making assistance system for digestive endoscopy surgery based on machine vision provided by the present invention includes an image acquisition subsystem, an image preprocessing subsystem, a target detection and segmentation subsystem, a semantic recognition subsystem, and a visual prompt subsystem, and specifically includes the following:

[0005] The image acquisition subsystem obtains surgical images in real time by using a high-frame-rate micro-endoscope;

[0006] The image preprocessing subsystem standardizes and enhances the surgical image using multi-scale histogram equalization and CLAHE enhancement strategies to obtain the processed image;

[0007] The target detection and segmentation subsystem includes a deep learning model composed of an anchor box-based multi-scale attention fusion detection network and a cascaded global-local context encoding image segmentation network to identify the lesion area and contour segmentation of the processed image, obtaining the image target detection result;

[0008] The semantic recognition subsystem is based on the image target detection result, uses graph neural network and medical knowledge graph for semantic understanding, obtaining surgical risk classification and lesion location information;

[0009] The visualization prompt subsystem superimposes the semantic recognition result on the surgical image in the form of augmented reality.

[0010] Furthermore, the target detection and segmentation subsystem identifies the lesion area and contour segmentation of the processed image, specifically including the following steps:

[0011] Step S1: Multi-scale feature extraction, constructing an anchor box-based multi-scale attention fusion detection network, inputting the processed image to obtain a multi-scale fusion feature map, which contains anchor boxes;

[0012] Step S2: Lesion candidate region detection, performing target classification and bounding box regression on the region corresponding to each anchor box, determining the position and category of the preliminary lesion candidate region, and obtaining the lesion candidate region through non-maximum suppression processing;

[0013] Step S3: Cascaded global-local context encoding segmentation, constructing a global-local context modeling segmentation network based on the Transformer structure, performing contour segmentation on the lesion candidate region to obtain a preliminary segmentation mask for each candidate lesion region;

[0014] Step S4: Iterative segmentation, introducing a progressive segmentation strategy, composed of multiple refined segmentation units in cascade, iteratively optimizing the preliminary segmentation mask obtained in step S3 in turn to obtain a set of masks;

[0015] Step S5: Target detection output, by analyzing the contour of each mask in the set of masks, extracting the bounding box parameters to obtain a set of preliminary lesion detection boxes;

[0016] Step S6: Structured output. A multi - feature fusion mechanism is used to screen and fuse the results of the preliminary lesion detection box set. An anomaly detection algorithm is embedded to obtain the final lesion detection box set. The position information, class label, and anomaly flag of each lesion detection box in it are output in a structured format, mapped to the multi - scale fusion feature map, and the image object detection result is generated.

[0017] Further, step S3 specifically includes the following steps:

[0018] Step S31: Build a network. Build a 4 - layer nested Transformer encoder, each layer contains a standard multi - head self - attention module and a cross - scale self - attention module. Introduce the deformable attention mechanism, and extract the context semantic information of the lesion candidate region through the residual skip connection between each layer.

[0019] Step S32: Build a local boundary perception path. Build a local convolutional path composed of 3 convolutional blocks and multi - level high - resolution upsampling, which is parallel to the 4 - layer nested Transformer encoder. Add a gradient - guided edge enhancement module and a multi - scale fusion skip connection to generate a local perception feature map.

[0020] Step S33: Fusion decoder. Build a fusion decoder containing a spatial attention module and a channel attention module. The context semantic information obtained in step S31 is used as the global path input, and the local perception feature map obtained in step S32 is used as the local path input. Concatenate the features of the global path and the local path, and fuse the prediction results through a dynamic weight weighting mechanism to output a preliminary segmentation mask.

[0021] Further, step S4 specifically includes the following steps:

[0022] Step S41: Input initialization. Denote the preliminary segmentation mask of step S3 as , which is used as the starting mask for iteration;

[0023] Step S42: Initial round of refined segmentation. In the first refined segmentation unit, fuse the multi - scale fusion feature map obtained in step S1 and the context semantic information obtained in step S3 as the input, and perform refinement processing on to generate the first - round optimized mask ;

[0024] Step S43: Progressive optimization. Based on the first - round optimized mask , repeat the operation of step S42 in the second refined segmentation unit to generate the second - round optimized mask ;

[0025] Step S44: Iterative optimization. Set the number of iterations, and repeat step S44 until the number of iterations is reached to obtain a mask set 。

[0026] Further, in step S6, the anomaly algorithm uses a multi-layer memory optimization anomaly recognition algorithm, which specifically includes the following steps:

[0027] Step S61: Construct a stacked LSTM network. The LSTM network includes an input gate, a forget gate, and an output gate, and outputs a multi-scale hidden state sequence, including time points;

[0028] Step S62: Enhance the memory gating mechanism. For the input gate, forget gate, and output gate in the 4-layer stacked LSTM network, perform gated enhancement reconstruction and introduce a gated memory selector;

[0029] Step S63: Anomaly feature focusing. Introduce a local window attention mechanism and a global context attention mechanism to construct an anomaly feature focusing module, and perform weighted enhancement on the time points in the multi-scale hidden state sequence;

[0030] Step S64: Hyperparameter optimization. Design a mutation branch heuristic optimization algorithm, combine the global population drift and local mutation search strategies, and iteratively optimize the hyperparameters of the stacked LSTM network;

[0031] Step S65: Anomaly score generation. Based on the sequence residuals between the multi-scale hidden state sequence and the predicted values, construct a dynamic error distribution modeling module, introduce a sliding window statistic and an adaptive quantile threshold discrimination strategy, and calculate the anomaly probability score at each moment;

[0032] Step S66: Multi-scale fusion. Fuse the anomaly probability scores and the output results of the multi-layer LSTM, introduce an attention weighting strategy to achieve multi-scale representation reconstruction, output the fused anomaly detection results, and construct a set of final lesion detection frames.

[0033] The beneficial effects achieved by the present invention using the above solution are as follows:

[0034] (1) Aiming at the problems of mis-segmentation and missed-segmentation in the surgical image segmentation technology, the present invention introduces a cascaded global-local context encoding structure, combines a nested Transformer encoder and a local boundary-aware path in the segmentation network, realizes the joint modeling of global anatomical semantics and local boundary details, and uses a deformable attention mechanism and a gradient-guided edge enhancement module to improve the accuracy of lesion contour recognition;

[0035] (2) Aiming at the problems of lesion boundary distortion and weak recognition ability existing in the segmentation method, the present invention designs an iterative progressive optimization mechanism. Multiple refined segmentation units are used to gradually refine the initial mask layer by layer. In each round, the original features and semantic context are fused to gradually refine the mask, realizing a "coarse to fine" continuous enhancement strategy and improving the accuracy and continuity of segmentation.

[0036] (3) Aiming at the problem of low judgment accuracy existing in the existing abnormal area recognition method, the present invention uses a multi-layer stacked LSTM network and an optimization mechanism, introduces local and global attention mechanisms, focuses on abnormal key features, enhances the discrimination ability, and combines anomaly detection with structured output to achieve accurate quantification and visual prompt of lesion abnormalities. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a subsystem connection diagram of a decision-making assistance system for digestive medicine surgery provided by the present invention;

[0038] Figure 2 It is a schematic diagram for lesion area recognition and contour segmentation;

[0039] Figure 3 It is a schematic diagram of step S4.

[0040] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the scope of protection of the present invention.

[0042] In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention.

[0043] Example 1, refer to Figure 1, a decision-making assistance system for digestive medicine surgery based on machine vision provided by the present invention, includes an image acquisition subsystem, an image preprocessing subsystem, a target detection and segmentation subsystem, a semantic recognition subsystem, and a visualization prompt subsystem, specifically including the following:

[0044] The image acquisition subsystem obtains surgical images in real time by using a high-frame-rate microendoscope;

[0045] The image preprocessing subsystem uses multi-scale histogram equalization and CLAHE enhancement strategies to standardize and enhance the surgical images, obtaining processed images;

[0046] The target detection and segmentation subsystem includes a deep learning model composed of an anchor box-based multi-scale attention fusion detection network and a cascaded global-local context encoding image segmentation network to identify lesion regions and contour segmentation of the processed images, obtaining image target detection results;

[0047] The semantic recognition subsystem is based on the image target detection results, uses graph neural networks and medical knowledge graphs for semantic understanding, obtaining surgical risk classification and lesion location information;

[0048] The visualization prompt subsystem superimposes the semantic recognition results on the surgical images in an augmented reality manner.

[0049] Embodiment 2, refer to Figure 2 , based on the above embodiment, the target detection and segmentation subsystem identifies lesion regions and contour segmentation of the processed images, specifically including the following steps:

[0050] Step S1: Multi-scale feature extraction, constructing an anchor box-based multi-scale attention fusion detection network, inputting the processed images, obtaining a multi-scale fusion feature map, which contains anchor boxes;

[0051] Step S2: Lesion candidate region detection, performing target classification and bounding box regression on the regions corresponding to each anchor box, determining the positions and categories of the preliminary lesion candidate regions, and obtaining the lesion candidate regions through non-maximum suppression processing;

[0052] Step S3: Cascaded global-local context encoding segmentation, constructing a global-local context modeling segmentation network based on the Transformer structure, performing contour segmentation on the lesion candidate regions, obtaining a preliminary segmentation mask for each candidate lesion region;

[0053] Step S4: Iterative segmentation, introducing a progressive segmentation strategy, composed of multiple refined segmentation units in cascade, iteratively optimizing the preliminary segmentation mask obtained in Step S3 in sequence, obtaining a set of masks;

[0054] Step S5: Object detection output. By performing contour analysis on each mask in the mask set, bounding box parameters are extracted to obtain a preliminary set of lesion detection boxes.

[0055] Step S6: Structured output. A multi-feature fusion mechanism is used to screen and fuse the results of the preliminary set of lesion detection boxes, and an anomaly detection algorithm is embedded to obtain the final set of lesion detection boxes. The position information, class label, and anomaly identification of each lesion detection box are output in a structured format and mapped to a multi-scale fusion feature map to generate the image object detection result.

[0056] Embodiment 3. This embodiment is based on the above embodiment. Step S3 specifically includes the following steps:

[0057] Step S31: Build a network. Build a 4-layer nested Transformer encoder, with each layer containing a standard multi-head self-attention module and a cross-scale self-attention module. Introduce a deformable attention mechanism, and extract the context semantic information of the lesion candidate region through residual skip connections between layers.

[0058] Step S32: Build a local boundary perception path. Build a local convolutional path consisting of 3 convolutional blocks and multi-level high-resolution upsampling, parallel to the 4-layer nested Transformer encoder. Add a gradient-guided edge enhancement module and multi-scale fusion skip connections to generate a local perception feature map.

[0059] Step S33: Fusion decoder. Build a fusion decoder containing a spatial attention module and a channel attention module. The context semantic information obtained in Step S31 is used as the input for the global path, and the local perception feature map obtained in Step S32 is used as the input for the local path. Concatenate the features of the global path and the local path, and fuse the prediction results through a dynamic weight weighting mechanism to output a preliminary segmentation mask.

[0060] In this embodiment, Step S32 constructs a parallel local perception path, which includes:

[0061] 3 local convolutional blocks, each containing convolution, normalization, and activation operations, for extracting local texture features of the boundary region.

[0062] A multi-level high-resolution upsampling module for maintaining the segmentation accuracy and restoring the spatial resolution.

[0063] Integrate a gradient-guided edge enhancement module to guide the model to focus on the high-frequency edge region with image gradient information, improving the boundary recognition ability.

[0064] Set up a multi-scale skip connection mechanism to fuse features at different resolution levels to generate a local perception feature map with fine edge perception ability.

[0065] The fusion decoder in step S33 specifically includes:

[0066] Spatial attention module: enhancing the model's response to key spatial regions;

[0067] Channel attention module: highlighting highly discriminative semantic channels;

[0068] Using a cascaded structure to fuse the global path and the local path, guiding the decoder to comprehensively predict using these two types of features;

[0069] Introducing a dynamic weight weighting mechanism to adaptively adjust the global-local feature fusion ratio according to the task difficulty and feature response, and finally output a preliminary segmentation mask image.

[0070] Example 4, refer to Figure 3 , based on the above example, step S4 specifically includes the following steps:

[0071] Step S41: Input initialization, the preliminary segmentation mask of step S3 is denoted as , serving as the starting mask for iteration;

[0072] Step S42: Initial round of refined segmentation. In the first refined segmentation unit, fuse the multi-scale fusion feature map obtained in step S1 and the context semantic information obtained in step S3 as the input, and perform refinement processing on to generate the first-round optimized mask ;

[0073] Step S43: Progressive optimization. Based on the first-round optimized mask , repeat the operation of step S42 in the second refined segmentation unit to generate the second-round optimized mask ;

[0074] Step S44: Iterative optimization. Set the number of iterations, and repeat step S44 until the number of iterations is reached to obtain the mask set .

[0075] In this example, taking the intraoperative real-time obtained liver tumor ultrasound image as the input image, perform multi-round refined optimization segmentation processing using the preliminary segmentation mask generated in step S3. The specific steps are as follows:

[0076] Step S41: Select a frame of two-dimensional cross-sectional image from the intraoperative ultrasound image sequence with a resolution of 512×512. After obtaining the preliminary segmentation mask through step S3, perform a preliminary localization of the tumor area. Although the main tumor area is covered, there are obvious serrations and artifacts at the boundary, and there are also some background mis-segmentations;

[0077] Step S42: Build the first refined segmentation unit and perform the following operations:

[0078] Input fusion feature construction: Extract the multi-scale fusion feature map constructed in Step S1 and fuse the context features output by the nested Transformer encoder in Step S3;

[0079] Fusion module design: Form fusion features by feature splicing, 1×1 convolution dimensionality reduction, and SE attention module to focus on the activation region;

[0080] Refinement network structure: Build a refine block composed of two 3×3 convolutions, BN, ReLU, and residual structure, and input the preliminary segmentation mask and fusion features together;

[0081] Mask update: Output a new mask through the refine block, with smoother boundaries and significantly reduced mis-segmented regions;

[0082] Step S43: Continue with the second-round optimization and build the second refine block;

[0083] Step S44: Iterative optimization. In each round of iteration, introduce the residual map of the previous round and update the dynamic fusion weights.

[0084] Example 5. Based on the above example, in Step S6, the anomaly algorithm uses a multi-layer memory optimization anomaly recognition algorithm, which specifically includes the following steps:

[0085] Step S61: Build a stacked LSTM network. Design a 4-layer stacked LSTM network. The LSTM network includes an input gate, a forget gate, and an output gate, and outputs a multi-scale hidden state sequence, including time points;

[0086] Step S62: Enhance the memory gating mechanism. Perform gating enhancement reconstruction on the input gate, forget gate, and output gate in the 4-layer stacked LSTM network, and introduce a gating memory selector;

[0087] Step S63: Anomaly feature focusing. Introduce a local window attention mechanism and a global context attention mechanism, build an anomaly feature focusing module, and perform weighted enhancement on the time points in the multi-scale hidden state sequence;

[0088] Step S64: Hyperparameter optimization. Design a mutation branch heuristic optimization algorithm, combine the global population drift and local mutation search strategies, and perform iterative optimization on the hyperparameters of the stacked LSTM network;

[0089] Step S65: Abnormal score generation. Based on the sequence residuals between the multi-scale hidden state sequences and the predicted values, a dynamic error distribution modeling module is constructed. A sliding window statistics and adaptive quantile threshold discrimination strategy are introduced to calculate the abnormal probability scores at each moment.

[0090] Step S66: Multi-scale fusion. The abnormal probability scores and the output results of the multi-layer LSTM are fused. An attention weighting strategy is introduced to achieve the reconstruction of multi-scale representations, and the fused abnormal detection results are output to construct the final set of lesion detection boxes.

[0091] In this embodiment, the sample code used is as follows:

[0092] import torch

[0093] import torch.nn as nn

[0094] import torch.nn.functional as F

[0095] # Step S61: Construct a stacked LSTM network (with residual connections)

[0096] class ResidualLSTM(nn.Module):

[0097] def __init__(self, input_size, hidden_size, num_layers):

[0098] super().__init__()

[0099] self.num_layers = num_layers

[0100] self.hidden_size = hidden_size

[0101] self.lstm_layers = nn.ModuleList(

[0102] nn.LSTM(input_size if i == 0 else hidden_size, hidden_size, batch_first=True)

[0103] for i in range(num_layers) )

[0105] def forward(self, x):

[0106] outputs = []

[0107] for i, lstm in enumerate(self.lstm_layers):

[0108] residual = x

[0109] x, _ = lstm(x)

[0110] if residual.shape == x.shape:

[0111] x = x + residual # Residual connection

[0112] outputs.append(x)

[0113] return outputs # Output the hidden state of each layer

[0114] # Step S62: Enhance the memory gating mechanism (simplified to a gate selector)

[0115] class GateSelector(nn.Module):

[0116] def __init__(self, hidden_size):

[0117] super().__init__()

[0118] self.gate = nn.Sequential(

[0119] nn.Linear(hidden_size, hidden_size),

[0120] nn.Sigmoid() )

[0122] def forward(self, h):

[0123] gate_weight = self.gate(h)

[0124] return h * gate_weight

[0125] # Step S63: Abnormal feature focusing (local + global attention mechanism)

[0126] class AttentionFusion(nn.Module):

[0127] def __init__(self, hidden_size):

[0128] super().__init__()

[0129] self.query = nn.Linear(hidden_size, hidden_size)

[0130] self.key = nn.Linear(hidden_size, hidden_size)

[0131] self.value = nn.Linear(hidden_size, hidden_size)

[0132] def forward(self, x):

[0133] Q = self.query(x)

[0134] K = self.key(x)

[0135] V = self.value(x)

[0136] attn_score = torch.bmm(Q, K.transpose(1, 2)) / (x.size(-1) **0.5)

[0137] attn_weights = torch.softmax(attn_score, dim=-1)

[0138] return torch.bmm(attn_weights, V)

[0139] # Step S65 + S66: Dynamic error score modeling + Multi-scale fusion

[0140] def compute_anomaly_score(preds, targets, window=5):

[0141] residual = torch.abs(preds - targets)

[0142] scores = []

[0143] for i in range(window, residual.size(1)):

[0144] hist = residual[:, i - window:i, :].mean(dim=1)

[0145] cur = residual[:, i:i+1, :]

[0146] score = (cur > hist).float().mean(dim=-1, keepdim=True)

[0147] scores.append(score)

[0148] return torch.cat(scores, dim=1)

[0149] # Model construction

[0150] input_size = 10

[0151] hidden_size = 32

[0152] num_layers = 3

[0153] seq_len = 40

[0154] batch_size = 2

[0155] # Initialize model components

[0156] lstm_net = ResidualLSTM(input_size, hidden_size, num_layers)

[0157] gate_module = GateSelector(hidden_size)

[0158] attn_module = AttentionFusion(hidden_size)

[0159] # Input simulation

[0160] x = torch.randn(batch_size, seq_len, input_size)

[0161] # Steps S61 - S63: Forward Propagation and Fusion

[0162] hidden_states = lstm_net(x)

[0163] gated_states = [gate_module(h) for h in hidden_states]

[0164] attn_outputs = [attn_module(h) for h in gated_states]

[0165] fused = sum(attn_outputs) / len(attn_outputs)

[0166] # Simulate Predictions and Targets, Calculate Anomaly Scores (S65 - S66)

[0167] y_true = torch.randn_like(fused)

[0168] anomaly_scores = compute_anomaly_score(fused, y_true)

[0169] print("Anomaly score shape:", anomaly_scores.shape) # Should be [batch_size, seq_len - window].

[0170] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non - exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.

[0171] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

[0172] The above describes the present invention and its implementation manners. Such description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. In summary, if those of ordinary skill in the art are inspired by it and, without departing from the gist of the present invention, design similar structural manners and embodiments to this technical solution without creative efforts, they shall fall within the protection scope of the present invention.

Claims

1. A decision-making assistance system for digestive medicine surgery based on machine vision, including an image acquisition subsystem, an image preprocessing subsystem, a target detection and segmentation subsystem, a semantic recognition subsystem, and a visualization prompt subsystem, specifically including the following: The image acquisition subsystem obtains surgical images in real time; The image preprocessing subsystem performs standardization and enhancement processing on the surgical images to obtain processed images; The target detection and segmentation subsystem includes a deep learning model composed of an anchor box-based multi-scale attention fusion detection network and an image segmentation network with cascaded global-local context encoding to identify the lesion area and contour segmentation of the processed image, and obtain the image target detection result; The semantic recognition subsystem performs semantic understanding based on the image target detection result to obtain surgical risk classification and lesion localization information; The visualization prompt subsystem superimposes the semantic recognition result on the surgical image in the form of augmented reality.

2. The decision-making assistance system for digestive medicine surgery based on machine vision according to claim 1, wherein: The target detection and segmentation subsystem performs lesion area recognition and contour segmentation on the processed image, specifically including the following steps: Step S1: Multi-scale feature extraction, construct an anchor box-based multi-scale attention fusion detection network, input the processed image, and obtain a multi-scale fusion feature map, which contains anchor boxes; Step S2: Lesion candidate area detection, perform target classification and bounding box regression on the area corresponding to each anchor box, and obtain the lesion candidate area through non-maximum suppression processing; Step S3: Cascaded global-local context encoding segmentation, construct a global-local context modeling segmentation network based on the Transformer structure, perform contour segmentation on the lesion candidate area, and obtain a preliminary segmentation mask for each candidate lesion area; Step S4: Iterative segmentation, introduce a progressive segmentation strategy, which is composed of multiple refined segmentation units in cascade, and iteratively optimize the preliminary segmentation mask obtained in step S3 in turn to obtain a mask set; Step S5: Target detection output, extract the bounding box parameters by performing contour analysis on each mask in the mask set to obtain a preliminary lesion detection box set; Step S6: Structured output, use a multi-feature fusion mechanism to screen and fuse the results of the preliminary lesion detection box set, embed an anomaly detection algorithm, obtain the final lesion detection box set, output the position information, class label, and anomaly identification of each lesion detection box in the structured format, map it to the multi-scale fusion feature map, and generate the image target detection result.

3. The decision-making assistance system for digestive medicine surgery based on machine vision according to claim 2, wherein: Step S3 specifically includes the following steps: Step S31: Construct a network, construct a 4-layer nested Transformer encoder to extract the context semantic information of the lesion candidate area; Step S32: Construct a local boundary-aware path, construct a local convolutional path, parallel to the 4-layer nested Transformer encoder, add a gradient-guided edge enhancement module and a multi-scale fusion skip connection to generate a local perception feature map; Step S33: Fusion decoder. Construct a fusion decoder that includes a spatial attention module and a channel attention module. The context semantic information obtained in step S31 is used as the input for the global path, and the local perception feature map obtained in step S32 is used as the input for the local path. Concatenate the features of the global path and the local path, and fuse the prediction results through a dynamic weight weighting mechanism to output a preliminary segmentation mask.

4. The decision-making assistance system for digestive medicine surgery based on machine vision according to claim 2, characterized in that: In step S6, the anomaly algorithm uses a multi-layer memory optimization anomaly recognition algorithm, which specifically includes the following steps: Step S61: Construct a stacked LSTM network. Design a 4-layer stacked LSTM network. The LSTM network includes an input gate, a forget gate, and an output gate, and outputs a multi-scale hidden state sequence, including time points. Step S62: Enhance the memory gating mechanism. Reconstruct the gating of the input gate, forget gate, and output gate in the 4-layer stacked LSTM network, and introduce a gating memory selector. Step S63: Anomaly feature focusing. Introduce a local window attention mechanism and a global context attention mechanism to construct an anomaly feature focusing module to weight and enhance the time points in the multi-scale hidden state sequence. Step S64: Hyperparameter optimization. Design a mutation branch heuristic optimization algorithm, combine the global population drift and the local mutation search strategy, and iteratively optimize the hyperparameters of the stacked LSTM network. Step S65: Anomaly score generation. Based on the sequence residuals between the multi-scale hidden state sequence and the predicted values, construct a dynamic error distribution modeling module, introduce a sliding window statistics and an adaptive quantile threshold discrimination strategy, and calculate the anomaly probability score at each moment. Step S66: Multi-scale fusion. Fuse the anomaly probability scores and the output results of the multi-layer LSTM, introduce an attention weighting strategy to achieve multi-scale representation reconstruction, output the fused anomaly detection results, and construct a set of final lesion detection frames.

Citation Information

Patent Citations

  • Aero-engine gas path fault diagnosis method based on LSTM in combination with attention mechanism

    CN113158537A

  • Semantic segmentation method of two-stage iteration context boundary feedback network

    CN115147604A

  • Colon polyp segmentation method and device and storage medium

    CN116542921A

  • Semantic segmentation method and device based on context cascade and multi-scale feature refinement

    CN116543155A

  • Water supply pipeline operation data anomaly detection method

    CN116842323A