An optimal lens intelligent prediction method and system for open surgery
By setting up multiple cameras during open surgery and combining deep learning technology and the Time-Block module to process video features, the problems of data redundancy and unstable perspective switching in traditional methods are solved, and seamless switching of multiple perspectives and complete recording of key information are achieved.
Patent Information
- Application Number
- CN202510066764.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-01-16
AI Technical Summary
Existing multi-camera recording methods have data redundancy and fail to fully utilize time series features in open surgery, resulting in the loss of key image information. Traditional single-view cameras find it difficult to achieve high-speed and seamless switching of perspectives in complex surgical environments.
A multi-camera setup is used, combined with a deep convolutional neural network and an object detection network to extract image and semantic features. Time-Block modules and residual connection technology are used to process long time series data, and a softmax layer is used to select the optimal perspective.
It achieves seamless switching of multiple perspectives in complex surgical environments, ensures complete recording and analysis of key information, and improves the accuracy and stability of perspective prediction.
Smart Images

Figure CN119992417B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of intelligent medical treatment and the field of computer vision, and particularly relates to a best lens intelligent prediction method and system for open surgery. BACKGROUND
[0002] In a complex open surgery environment, clear observation of human tissues and surgical instruments is crucial for improving the safety and efficiency of surgery. With the rapid development of artificial intelligence technology, traditional surgical video recording methods have been unable to meet the needs of modern surgical education and medical collaboration. This is mainly because the bodies of doctors and nurses inevitably block the view of the surgical area during the surgery process, resulting in the loss of key image information.
[0003] Although the existing multi-camera recording method can record the surgery process from multiple angles, it often produces a large amount of data redundancy. In addition, the existing multi-camera lens selection prediction method generally ignores the time sequence characteristics, mainly focusing on the analysis of single-frame or short-term features, and fails to fully utilize the strong correlation between consecutive frames. Time information plays a crucial role in handling occlusions and optimizing view switching.
[0004] Therefore, there is an urgent need for a method that can intelligently predict the best view switching opportunity in a complex open surgery scenario. This method should have the ability to achieve high-speed, seamless view switching while ensuring image quality, thereby ensuring that key information during the surgery process is recorded and analyzed completely and clearly. SUMMARY
[0005] In order to solve the limitations of single-view cameras in open surgery scene recording in the prior art, especially the problem of video information loss due to object occlusion during the surgery process, the present application proposes a best lens intelligent prediction method for open surgery, which specifically includes the following steps:
[0006] At least 6 cameras are set up at different angles in the open surgery scene, image features are extracted from video frames through a deep convolutional neural network, and semantic features are extracted from video frames through a target detection network;
[0007] All image features and semantic features extracted at the same time step are spliced together as a joint feature vector at that moment;
[0008] Through a feature dimension reduction module, a two-layer feedforward neural network is used to map the joint feature vector to a lower-dimensional vector space, obtaining a reduced latent feature vector;
[0009] The latent feature vector is input into a network stacked with multiple Time-Block modules to extract a time feature vector, and layer normalization processing is performed;
[0010] The normalized time feature vector is input into a softmax layer to obtain a camera label probability distribution, and the camera with the highest probability is pushed as the best view angle.
[0011] Further, the deep convolutional neural network is a ResNet-18 network pre-trained on an ImageNet dataset.
[0012] Further, the target detection model is a YOLOv5s target detection model trained on a private open surgery target detection dataset.
[0013] Further, the processing process of each Time-Block module on the latent feature vector X obtained after joint input and dimension reduction of the multi-view video image features and semantic features includes:
[0014] The low-dimensional feature representation is input into a globally connected multi-head self-attention mechanism to extract corresponding context features;
[0015] The obtained context features are subjected to multi-scale convolution operation to extract features at different scales, and the features input into the Time-Block module are subjected to residual connection and layer normalization processing;
[0016] The normalized features are then spliced with the features processed by the feedforward neural network, and then subjected to once again normalization processing, and the output of the last normalization is taken as the output of the Time-Block module.
[0017] The application also proposes an open surgery-oriented best shot intelligent prediction system for realizing an open surgery-oriented best shot intelligent prediction method, which comprises a deep convolutional neural network, a target detection network, a splicing module, a time feature extraction network and a classifier.
[0018] The deep convolutional neural network is implemented by a pre-trained ResNet-18 network and is used to extract video features from video frames;
[0019] The target detection network is implemented by a trained YOLOv5s target detection model and is used to extract semantic features from video frames;
[0020] The splicing module is used to splice the video features and the semantic features together as a joint feature vector;
[0021] The time feature extraction network is a feature extraction network obtained by stacking a plurality of Time-Block modules, and is used to extract a time feature vector from the joint feature vector;
[0022] A classifier is used to classify the time feature vectors of the multiple cameras to select the camera with the best view angle from the multiple cameras.
[0023] Compared with the prior art, the scheme has the following beneficial effects:
[0024] 1、The present application solves the problem that the traditional single-view camera may miss key surgical information by synchronously shooting multiple cameras, ensuring comprehensive recording of the entire surgical process, especially in complex surgeries or scene occlusion, the application of multiple views improves the integrity of information;
[0025] 2、The present application processes feature information of different modalities of video features and semantic information by combining deep learning technology (such as ResNet-18 and YOLOv5s), the system can effectively extract key information in the surgical process and analyze complex multi-modal data. This makes the system have stronger adaptability and can handle various different surgical scenes;
[0026] 3、The TimeBlock module and residual connection technology are used to solve the problem of gradient disappearance in long time series data, effectively capture long-term dependencies in the surgical process, and improve the stability and accuracy of the view prediction model. BRIEF DESCRIPTION OF DRAWINGS
[0027] Fig. 1 A best shot intelligent prediction method flowchart for open surgery is provided in the present application;
[0028] Fig. 2 A view selection algorithm structure diagram based on time series prediction is provided in the present application. DETAILED DESCRIPTION
[0029] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0030] The present application proposes a best shot intelligent prediction method for open surgery, as shown in Figs. 1-2 , specifically including the following steps:
[0031] At least 6 angle cameras are arranged in the open surgery scene, image features are extracted from video frames by a deep convolutional neural network, and semantic features are extracted from video frames by a target detection network;
[0032] All image features and semantic features of the same time step are spliced together as a joint feature vector of the moment;
[0033] Through the feature dimension reduction module, a two-layer feedforward neural network is used to map the joint feature vector to a lower latitude vector space to obtain a latent feature vector after dimension reduction;
[0034] The latent feature vector is input into a network stacked with multiple Time-Block modules to obtain a time feature vector, and layer normalization processing is performed;
[0035] The normalized time feature vector is input into a softmax layer to obtain a camera label probability distribution, and the camera with the highest probability is pushed as the best view.
[0036] The best lens intelligent prediction method for open surgery can provide comprehensive visual monitoring during the operation to ensure that the operation process under different angles and perspectives is captured completely.
[0037] In this embodiment, the multi-view surgical video record is obtained by preprocessing the surgical video data, the best surgical view is selected from six views in the form of a one-hot vector, and the best surgical view selection dataset is obtained by selecting and labeling the best view by professional senior physicians.
[0038] As an optional implementation, the best lens intelligent prediction method for open surgery includes the following steps:
[0039] Step 1: The multi-view surgical video collected is used to extract video features using a pre-trained deep convolutional neural network ResNet-18, and semantic feature information of the video frame is extracted using a pre-trained object detection model YoloV5s model.
[0040] Specifically, assuming that the video feature is a vector of length N v , the semantic feature is a vector of length N S , and after splicing according to the feature dimension, a vector of length N S +N vjoint feature vector, which contains both visual information and semantic information of the recognized objects in the video. Formed as follows: N concatenate = [N S , N v ].
[0041] Step 2: Perform latent space mapping on the joint feature N concatenate generated in step 1, from which a high-density feature representation N concatenate is extracted, which can effectively capture the key dynamic change information in the surgical process. The high-density feature representation vector generated in step 2 is input into the prediction module. The high-dimensional, sparse input features are mapped to a low-dimensional, dense latent space, compressing the feature space, retaining the most important information, and reducing computational complexity. The fused joint feature is a high-dimensional vector N concatenate which contains both video features and semantic features. This embodiment uses an Embedding layer to map this high-dimensional feature to a low-dimensional latent space representation N embedding , the process of which is as follows:
[0042] N embedding = Embedding(N concatenate ) = W e N concatenate + b e
[0043] where Embedding(·) represents the embedding layer, W e represents the weight matrix of the embedding layer, and b e represents the bias vector of the embedding layer.
[0044] Step 3: Through the stacking of multiple Time-Block modules, the long-term dependencies and dynamic change information in the surgical process can be effectively captured, improving the accuracy and real-time performance of the best camera lens selection prediction. The Time-Block module accepts the low-dimensional feature representation N embedding processed by the Embedding layer, and then processes these features in parallel through the Multi-Head Self-Attention mechanism to capture complex dependencies and dynamic change information between different time steps.
[0045] Specifically, the process of extracting the corresponding context features through the globally connected Multi-Head Self-Attention mechanism includes:
[0046] Construct a query matrix, a key matrix, and a value matrix from the feature N embedding , which includes:
[0047] Q = W Q×N embedding ,K=W K ×N embedding ,V=W V ×N embedding
[0048] where W Q ,W K ,W V are three different learnable matrices;
[0049] In this embodiment, Q, K, V are divided into H parts, that is, the features of each sub-attention module; As an optional implementation method, the query vector, the key vector, and the value vector of the sub-attention can be updated by using a feedforward network. Taking the query vector of the hth attention head as an example, the update process is represented as:
[0050] Q'(h) = W(h)Q + Q(h)
[0051] where W(h) is a learnable matrix;
[0052] Using the updated query vector, key vector, and value vector to update the attention, the output of the hth attention head module is represented as:
[0053]
[0054] where Attention(h) represents the output of the hth attention head module; Q'(h), K'(h), V'(h) are the updated values of the query vector, the key vector, and the value vector respectively; dk represents the dimension of the key vector; softmax(·) represents the softmax function; (·) T represents the transpose of a vector or a matrix;
[0055] The updated query vector, key vector, and value vector are concatenated. Taking the query vector as an example, the concatenated query matrix is represented as:
[0056] Q' = Concat(Q'(1), …, Q'(H))
[0057] The output of the multi-head attention mechanism is represented as:
[0058] MultiHead(Q', K', V') = Concat(Attention(1), …, Attention(H))Wo
[0059] MultiHead(Q', K', V') denotes the multi-head attention mechanism, Q' denotes the query vector of the multi-head attention mechanism, K' denotes the key vector of the multi-head attention mechanism, and V' denotes the value vector of the multi-head attention mechanism; Wo denotes the weight matrix of the final linear transformation; h e {1, 2,..., H}, H is the number of heads in the multi-head attention mechanism; W(h) is a learnable weight matrix used to adjust the influence of the global query on each head.
[0060] Next, these context features are passed to a feedforward neural network to further extract high-level features through fully connected layers and nonlinear activation functions such as ReLU. To enhance the stability and training efficiency of the model, residual connections are introduced after each sub-module by directly adding the input to the output to ensure effective information transmission and avoid the problem of gradient vanishing. Layer normalization is then performed on the features to reduce internal covariate shift and promote rapid convergence of the model. This process is represented as:
[0061]
[0062]
[0063] X" = LayerNorm(X' + FFN(X'))
[0064] wherein, MultiscaleInception represents the result of the multi-scale convolution operation; X' represents the features output by the first layer normalization operation in the Time-Block module, which are obtained by performing feedforward neural network operations and residual connections on
[0065] The multiple Time-Block modules are stacked in order to form a deep time series modeling network, allowing the entire system to extract more complex and abstract temporal features layer by layer and fully capture long-term dependencies and dynamic changes during the surgical process.
[0066] Step 4: The features output by the multiple Time-Block modules are processed through layer normalization to obtain a time feature vector Z. The feature vector Z is input into a softmax layer for classification, which calculates the probability score of the camera label, represented as:
[0067]
[0068] wherein, P(y = n | X ″ ) represents the probability of the camera label being n given the features X″ (For simplicity of representation, the embodiment adopts X ″ represents the probability of selecting the lens label n, y represents the category or selection of the camera label, and n represents the specific category of the camera label; W n represents the learnable weight vector related to the camera label n, and b n represents the bias term related to the camera label n.
[0069] Step 5: During the training process, the model is optimized by weighted cross-entropy loss. The loss function used when training the network stacked by Time-Block modules and the softmax layer is represented as:
[0070]
[0071] where L weighted为 is the loss function; N is the total number of samples in the data set; C is the number of different categories, i.e. the number of shots; w c is the weight representing different categories, used to handle the case of class imbalance; y i,c is the true label of the i-th sample in the c-th category; p i,c is the probability value of the corresponding category output by the model.
[0072] The application also proposes an optimal lens intelligent prediction system for open surgery, which is used to realize an optimal lens intelligent prediction method for open surgery, including a deep convolutional neural network, a target detection network, a splicing module, a time feature extraction network and a classifier, wherein:
[0073] The deep convolutional neural network is realized by a pre-trained ResNet-18 network, which is used to extract video features from video frames;
[0074] The target detection network is realized by a trained YOLOv5s target detection model, which is used to extract semantic features from video frames;
[0075] The splicing module is used to splice the video features and semantic features together as a joint feature vector;
[0076] The time feature extraction network is a feature extraction network obtained by stacking multiple Time-Block modules, which is used to extract a time feature vector from the joint feature vector;
[0077] The classifier is used to classify according to the time feature vectors of multiple cameras, and select the camera with the best view angle from the multiple cameras.
[0078] While embodiments of the application have been shown and described, it is to be understood that the embodiments described are merely exemplary of the principles and application of the present application. Numerous modifications and adaptions can be effected without departing from the spirit and scope of the present application, which is not limited to the exact construction and arrangement described. It is intended, therefore, to cover all modifications and adaptions that fall within the scope of the claims and their equivalents.
Claims
1. An intelligent prediction method for optimal shots in open surgery, characterized by: The specific steps include: Set up cameras from at least six angles in an open surgical scene, extract image features from video frames using a deep convolutional neural network, and extract semantic features from video frames using an object detection network; All the image features and semantic features extracted at the same time step are concatenated together as the joint feature vector of that time step; Through the feature dimensionality reduction module, a two-layer feedforward neural network is used to map the joint feature vector to a lower-dimensional vector space to obtain the reduced-dimensional latent feature vector; The latent feature vector is input into a network with multiple Time-Block modules stacked to extract the time feature vector; each Time-Block module jointly inputs the multi-view video image features and semantic features to obtain the latent feature vector after dimensionality reduction. The processing process includes: Input the low-dimensional feature representation into the globally connected multi-head self-attention mechanism to extract the corresponding context features; Perform multi-scale convolution operations on the obtained context features, extract features at different scales, and add the features extracted at different scales to obtain multi-scale convolution features; The multi-scale convolution features are processed by the first-layer feedforward network, and then a residual connection is performed. Then, the features are processed by a normalization layer and used as the output of the first-layer feedforward network. The output of the first-layer feedforward network is processed by the second-layer feedforward network, and then a residual connection is performed. Then, it is processed by a normalization layer and used as the output of the Time-Block module. The time feature vector is passed through the softmax layer to obtain the probability distribution of each camera label, and the camera with the highest probability is pushed as the best view.
2. The method for intelligently predicting the best shot for open surgery according to claim 1, wherein: The deep convolutional neural network is the ResNet-18 network pre-trained on the ImageNet dataset.
3. The method for intelligently predicting the best shot for open surgery according to claim 1, wherein: The object detection network is a YOLOv5s object detection model trained on a private open surgical object detection dataset.
4. The method for intelligently predicting the best shot for open surgery according to claim 1, wherein: The process of extracting corresponding context features through the globally connected multi-head self-attention mechanism includes: in, represents the multi-head attention mechanism, represents the concatenation of all sub-attention head query vectors in the multi-head attention mechanism, represents the concatenation of all sub-attention head key vectors in the multi-head attention mechanism, Represents the concatenation of all sub-attention head value vectors in the multi-head attention mechanism; represents the h-th sub-attention module, h∈{1,2,…,H}, H is the number of heads in the multi-head attention mechanism, represents the query vector of the h-th sub-attention module, represents the key vector of the h-th sub-attention module, Represents the value vector of the h-th sub-attention module; The weight matrix representing the final linear transformation; represents the dimension of the key vector; represents the softmax function; Indicates the transpose of a vector or matrix; Represents a splicing operation.
5. The method for intelligently predicting the best shot for open surgery according to claim 4, characterized in that: The output of the Time-Block module is expressed as: in, Represents multi-scale convolution operations the result; Represents the features of the first layer normalization output in the Time-Block module; represents a feedforward network; Represents the features of the normalized output of the second layer.
6. The method for intelligently predicting the best shot for open surgery according to claim 1, wherein: The process of obtaining the probability distribution of each camera label through the softmax layer includes: in, Represents a given feature Select the Camera tab The probability of y represents the category or choice of the camera label, Indicates the specific category of the camera tag; Indicates the camera label The associated learnable weight vector, Indicates the camera label Related bias terms.
7. The method for intelligently predicting the best shot for open surgery according to claim 1, wherein: The loss function used when training a network stacked by Time-Block modules and a softmax layer is expressed as: in, is the loss function; N is the total number of samples in the dataset; C is the number of different categories, that is, the number of shots; Represents the weights of different categories, used to deal with category imbalance; Represents the true label of the i-th sample in the c-th category; Output the probability value of the corresponding category for the model.
8. An intelligent prediction system for optimal shots in open surgery, characterized by: The method for implementing the optimal shot intelligent prediction method for open surgery as described in claim 1 comprises a deep convolutional neural network, a target detection network, a splicing module, a temporal feature extraction network, and a classifier, wherein: A deep convolutional neural network, implemented using a pre-trained ResNet-18 network, is used to extract image features from video frames; The object detection network is implemented using the pre-trained YOLOv5s object detection model to extract semantic features from video frames; The splicing module is used to splice the image features and semantic features together as a joint feature vector; The time feature extraction network is a feature extraction network formed by stacking multiple Time-Block modules, which is used to extract the time feature vector from the joint feature vector; A classifier is used to perform classification based on the time feature vectors of multiple cameras and select a camera with the best viewing angle from the multiple cameras.
Citation Information
Patent Citations
Switching method and device of dual-lens motion camera and computer equipment
CN118574002A
Method for searching for optimum view angle in multi-view angle environment and computer program product thereof
TW201427413A