Optimal lens intelligent prediction method and system for open surgical operation

By setting up multiple cameras in open surgical scenarios, using deep convolutional neural networks and object detection networks to extract features, and combining Time-Block modules and residual connection technology, high-speed and seamless switching of viewing angles in complex surgical scenarios are achieved, solving the problem of information loss caused by occlusion of a single-view camera, and ensuring the integrity of key information during the surgical process.

CN119992417AActive Publication Date: 2025-05-13CHONGQING INST OF GREEN & INTELLIGENT TECH CHINESE ACAD OF SCI
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510066764.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

In an open surgical environment, single-view cameras lose video information due to object occlusion. The existing multi-camera recording methods generate data redundancy and fail to fully utilize the time series features, so they cannot effectively predict the optimal viewing angle switching timing.

Method used

Multiple cameras are used to perform synchronous shooting in open surgical scenes. Images and semantic features are extracted through deep convolutional neural networks and object detection networks, combined with Time-Block module and residual connection technology, time features are extracted and the best viewing angle is selected through softmax layer.

Benefits of technology

It realizes high-speed and seamless switching of perspectives in complex open surgical scenarios, ensuring that key information during the surgical process is recorded and analyzed in a complete and clear manner, and improving the accuracy and stability of perspective prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992417A_ABST
    Figure CN119992417A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent medical treatment and the field of computer vision, and particularly relates to an optimal lens intelligent prediction method and system for an open type surgical operation, and the method comprises the steps: setting at least six-angle cameras in an open type surgical operation scene, extracting video features from the video frames through a deep convolutional neural network, and extracting semantic features from the video frames through a target detection network; splicing the extracted video features and semantic features together as a joint feature vector; inputting the joint feature vector into a network formed by stacking a plurality of Time-Block modules, and extracting to obtain a time feature vector; and enabling the normalized time feature vector to pass through a softmax layer to obtain label probability distribution of each camera, and pushing the camera with the highest probability as an optimal view angle. According to the invention, the capability of high-speed and seamless view angle switching can be realized, so that key information in the operation process can be completely and clearly recorded and analyzed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent medical technology and computer vision, and in particular relates to an optimal lens intelligent prediction method and system for open surgical operations. Background Art

[0002] In a complex open surgical environment, clear observation of human tissue and surgical instruments is essential to improving the safety and efficiency of surgery. With the rapid development of artificial intelligence technology, traditional surgical video recording methods can no longer meet the needs of modern surgical education and medical collaboration. This is mainly because the bodies of doctors and nurses will inevitably block the view of the surgical area during surgery, resulting in the loss of key image information.

[0003] Although existing multi-camera recording methods can record surgical procedures from multiple angles, they often generate a large amount of data redundancy. In addition, existing multi-camera lens selection prediction methods generally ignore time series features, mainly focusing on the analysis of single frames or short-term features, and fail to fully utilize the strong correlation between consecutive frames. Temporal information plays a vital role in handling occlusions and optimizing perspective switching.

[0004] Therefore, there is an urgent need for a method that can intelligently predict the best time to switch perspectives in complex open surgery scenarios. This method should be able to achieve high-speed, seamless switching of perspectives while ensuring image quality, thereby ensuring that key information during surgery can be fully and clearly recorded and analyzed. Summary of the invention

[0005] In order to solve the limitations of the existing single-view camera in recording open surgical scenes, especially the problem of video information loss caused by object occlusion during surgery, the present invention proposes an optimal shot intelligent prediction method for open surgical operations, which specifically includes the following steps:

[0006] Set up cameras at at least 6 angles in an open surgical scene, extract image features from video frames through a deep convolutional neural network, and extract semantic features from video frames through an object detection network;

[0007] All image features and semantic features extracted at the same time step are concatenated together as the joint feature vector at that moment;

[0008] Through the feature dimension reduction module, a two-layer feedforward neural network is used to map the joint feature vector to a lower-dimensional vector space to obtain the reduced-dimensional potential feature vector;

[0009] The potential feature vector is input into a network with multiple Time-Block modules stacked to extract the time feature vector, and then the layer is normalized.

[0010] The normalized time feature vector is passed through the softmax layer to obtain the probability distribution of each camera label, and the camera with the highest probability is pushed as the best viewing angle.

[0011] Furthermore, the deep convolutional neural network is a ResNet-18 network pre-trained on the ImageNet dataset.

[0012] Furthermore, the target detection model is a YOLOv5s target detection model trained on a private open surgical target detection dataset.

[0013] Furthermore, each Time-Block module processes the potential feature vector X obtained after the dimensionality reduction of the multi-view video image features and semantic features by combining the inputs, including:

[0014] Input the low-dimensional feature representation into the globally connected multi-head self-attention mechanism to extract the corresponding contextual features;

[0015] The obtained context features are subjected to multi-scale convolution operations to extract features at different scales. The features of the Time-Block module are input for residual connection and then layer normalization is performed.

[0016] The normalized features are then concatenated with the features processed by the feedforward neural network, and then normalized again. The output of the last normalization is used as the output of the Time-Block module.

[0017] The present invention also proposes an optimal shot intelligent prediction system for open surgery, which is used to implement an optimal shot intelligent prediction method for open surgery, including a deep convolutional neural network, a target detection network, a splicing module, a time feature extraction network and a classifier, wherein:

[0018] A deep convolutional neural network, implemented using a pre-trained ResNet-18 network, is used to extract video features from video frames;

[0019] The object detection network is implemented using the trained YOLOv5s object detection model to extract semantic features from video frames;

[0020] A concatenation module, used to concatenate video features and semantic features together as a joint feature vector;

[0021] The time feature extraction network is a feature extraction network obtained by stacking multiple Time-Block modules, which is used to extract the time feature vector from the joint feature vector;

[0022] A classifier is used to classify according to the time feature vectors of multiple cameras, and select a camera with the best viewing angle from the multiple cameras.

[0023] Compared with the prior art, the solution of the present invention has the following beneficial effects:

[0024] 1. The present invention uses multiple cameras to shoot synchronously, which solves the problem that traditional single-view cameras may miss key surgical information, ensuring a comprehensive record of the entire surgical process. Especially in the case of complex surgery or scene occlusion, the application of multiple perspectives improves the integrity of the information;

[0025] 2. The present invention combines deep learning technology (such as ResNet-18 and YOLOv5s) to process the feature information of different modes of video features and semantic information. The system can effectively extract key information during the operation and analyze complex multimodal data. This makes the system more adaptable and can handle various different surgical scenarios;

[0026] 3. The TimeBlock module and residual connection technology are used to solve the problem of gradient vanishing in long time series data, effectively capture the long-term dependencies in the surgical process, and improve the stability and accuracy of the view prediction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is a flow chart of an optimal lens intelligent prediction method for open surgery according to the present invention;

[0028] Figure 2 This is a structural diagram of the perspective selection algorithm based on time series prediction of the present invention. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0030] The present invention proposes an intelligent prediction method for the best shot for open surgery. Figures 1-2 , specifically including the following steps:

[0031] Set up cameras at at least 6 angles in an open surgical scene, extract image features from video frames through a deep convolutional neural network, and extract semantic features from video frames through an object detection network;

[0032] All image features and semantic features extracted at the same time step are concatenated together as the joint feature vector at that moment;

[0033] Through the feature dimension reduction module, a two-layer feedforward neural network is used to map the joint feature vector to a lower-dimensional vector space to obtain the reduced-dimensional potential feature vector;

[0034] The potential feature vector is input into a network with multiple Time-Block modules stacked to extract the time feature vector, and then the layer is normalized.

[0035] The normalized time feature vector is passed through the softmax layer to obtain the probability distribution of each camera label, and the camera with the highest probability is pushed as the best viewing angle.

[0036] The present invention provides an optimal lens intelligent prediction method for open surgical operations, which is used to provide comprehensive visual monitoring during the operation to ensure that the operation process at different angles and viewing angles is completely captured. The present invention requires at least 6 cameras to be installed at different positions on the shadowless lamp in the operating room, and through multi-view synchronous shooting, it is ensured that every detail of the operation process can be recorded and analyzed in real time.

[0037] In this embodiment, the captured surgical video data is preprocessed to obtain multi-view surgical video records, and the best surgical perspective is selected from six perspectives in the form of a unique hot vector. Professional senior physicians select and annotate the best perspective to obtain the best surgical perspective selection data set. The data set is used to train an end-to-end time series prediction network for multi-angle camera selection in open surgery, and the network is used to predict the video perspective selection sequence for a period of time in the future by inputting historical video data.

[0038] As an optional implementation, this embodiment implements an optimal shot intelligent prediction method for open surgery, including the following steps:

[0039] Step 1: For the collected multi-view surgical videos, use the pre-trained deep convolutional neural network ResNet-18 to extract its video features, and use the pre-trained target detection model YoloV5s model to extract the semantic feature information of the video frames. The extracted features are fused early to form a joint feature representation, and the visual features extracted from ResNet-18 and the semantic features from YoloV5s are spliced.

[0040] Specifically, assuming that the video feature is a length of N v A semantic feature is a vector of length N. S vector, and concatenate them according to the feature dimension to obtain a vector of length N S +N vThese features contain the visual information of the video and the semantic information of the objects recognized in the video. The following representation is formed: N concatenate =[N S ,N v ].

[0041] Step 2: For the joint feature N generated in step 1 concatenate Perform latent space mapping and extract high-density feature representation N from it concatenate , which can effectively capture the key dynamic change information during the operation. The high-density feature representation vector generated by the step is input into the prediction module. The high-dimensional, sparse input features are mapped to a low-dimensional, dense latent space, compressing the feature space, retaining the most important information, and reducing the computational complexity. The fused joint feature is a high-dimensional vector N concatenate , the high-dimensional vector contains video features and semantic features. This embodiment uses the Embedding layer to map this high-dimensional feature to a low-dimensional latent space representation N embedding , the process is:

[0042] N embedding =Embedding(N concatenate )=W e N concatenate +b e

[0043] Among them, Embedding(·) represents the embedding layer, W e represents the weight matrix of the embedding layer, b e Represents the bias vector of the embedding layer.

[0044] Step 3: By stacking multiple Time-Block modules, the long-term dependencies and dynamic change information during the surgery can be effectively captured, and the accuracy and real-time performance of the prediction of the best lens selection for multi-view cameras can be improved. The Time-Block module receives the low-dimensional feature representation N after processing by the Embedding Layer. embedding , and then these features are processed in parallel through a multi-head self-attention mechanism to capture the complex dependencies and dynamic change information between different time steps.

[0045] Specifically, the process of first extracting the corresponding context features through the globally connected multi-head self-attention mechanism includes:

[0046] By feature N embedding Construct query matrix, key matrix, and value matrix, including:

[0047] Q=W Q×N embedding ,K=W K ×N embedding ,V=W V ×N embedding

[0048] Among them, W Q , W K , W V are three different learnable matrices;

[0049] In this embodiment, Q, K, and V are divided into H parts, namely, the features of each sub-attention module; as an optional implementation method, the query vector, key vector, and value vector of the sub-attention can be updated using a feedforward network. Taking the query vector of the hth attention head as an example, the update process is expressed as:

[0050] Q′(h)=W(h)Q+Q(h)

[0051] Among them, W(h) is a learnable matrix;

[0052] Using the updated query vector, key vector, and value vector to update attention, the output of the h-th attention head module is expressed as:

[0053]

[0054] Where Attention(h) represents the output of the h-th attention head module; Q′(h), K′(h), and V′(h) are the updated values ​​of the query vector, key vector, and value vector, respectively; dk represents the dimension of the key vector; softmax(·) represents the softmax function; (·) T Indicates the transpose of a vector or matrix;

[0055] The updated query vector, key vector, and value vector are concatenated. Taking the query vector as an example, the concatenated query matrix is ​​expressed as:

[0056] Q′=Concat(Q′(1),…,Q′(H))

[0057] The output of the multi-head attention mechanism is expressed as:

[0058] MultiHead(Q′,K′,V′)=Concat(Attention(1),…,Attention(H))Wo

[0059] Among them, MultiHead(Q′,K′,V′) represents a multi-head attention mechanism, Q′ represents the query vector of the multi-head attention mechanism, K′ represents the key vector of the multi-head attention mechanism, and V′ represents the value vector of the multi-head attention mechanism; Wo represents the weight matrix of the final linear transformation; h∈{1,2,…,H}, H is the number of heads in the multi-head attention mechanism; W(h) is a learnable weight matrix used to adjust the impact of the global query on each head.

[0060] Next, these contextual features are passed to the feedforward neural network to further extract high-level features through the fully connected layer and nonlinear activation function (such as ReLU). In order to enhance the stability and training efficiency of the model, a residual connection is introduced after each submodule. By adding the input directly to the output, the effective transmission of information is ensured to avoid the gradient vanishing problem. The layer normalization standardizes the features, reduces the internal covariate shift, and promotes the rapid convergence of the model. The process is expressed as:

[0061]

[0062]

[0063] X″=LayerNorm(X′+FFN(X′))

[0064] in, represents the result of the multiscale convolution operation MultiscaleInception; X′ represents the output feature of the first layer normalization operation in the Time-Block module. The results are obtained by layer normalization after feedforward neural network operation and residual connection; LayerNorm(·) represents the characteristics of the output of the second layer normalization operation in the Time-Block module; FFN(·) represents the feedforward network; X″ represents the characteristics of the second layer normalization output.

[0065] Multiple Time-Block modules are stacked in sequence to form a deep time series modeling network, which enables the entire system to extract more complex and abstract time features layer by layer and comprehensively capture the long-term dependencies and dynamic changes during the surgery.

[0066] Step 4: The features output by multiple Time-Block modules are normalized to obtain the time feature vector Z. The feature vector Z is input into the softmax layer for classification. This layer calculates the probability score of the camera label, which is expressed as:

[0067]

[0068] Where P(y=n|X ″ ) represents a given feature X″ (In order to simplify the representation, this embodiment uses X ″ represents the probability of selecting lens label n under the features output by multiple Time-Block modules), y represents the category or selection of camera labels, and n represents the specific category of camera labels; W n represents the learnable weight vector associated with camera label n, b n Represents the bias term associated with camera label n.

[0069] Step 5: During training, the model is optimized using weighted cross entropy loss. The loss function used for training a network stacked by Time-Block modules and a softmax layer is expressed as:

[0070]

[0071] Among them, L weighted为 Loss function; N is the total number of samples in the data set; C is the number of different categories, that is, the number of shots; w c To represent the weights of different categories, it is used to deal with the situation of category imbalance; y i,c represents the true label of the i-th sample in the c-th category; p i,c Output the probability value of the corresponding category for the model.

[0072] The present invention also proposes an optimal shot intelligent prediction system for open surgery, which is used to implement an optimal shot intelligent prediction method for open surgery, including a deep convolutional neural network, a target detection network, a splicing module, a time feature extraction network and a classifier, wherein:

[0073] A deep convolutional neural network, implemented using a pre-trained ResNet-18 network, is used to extract video features from video frames;

[0074] The object detection network is implemented using the trained YOLOv5s object detection model to extract semantic features from video frames;

[0075] A concatenation module, used to concatenate video features and semantic features together as a joint feature vector;

[0076] The time feature extraction network is a feature extraction network obtained by stacking multiple Time-Block modules, which is used to extract the time feature vector from the joint feature vector;

[0077] A classifier is used to classify according to the time feature vectors of multiple cameras, and select a camera with the best viewing angle from the multiple cameras.

[0078] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An intelligent prediction method for the best shot in open surgery, characterized in that: The specific steps include: Set up cameras at at least 6 angles in an open surgical scene, extract image features from video frames through a deep convolutional neural network, and extract semantic features from video frames through an object detection network; All image features and semantic features extracted at the same time step are concatenated together as the joint feature vector at that moment; Through the feature dimension reduction module, a two-layer feedforward neural network is used to map the joint feature vector to a lower-dimensional vector space to obtain the reduced-dimensional potential feature vector; The potential feature vector is input into a network with multiple Time-Block modules stacked to extract the time feature vector; The time feature vector is passed through the softmax layer to obtain the probability distribution of each camera label, and the camera with the highest probability is pushed as the best viewing angle.

2. The best shot intelligent prediction method for open surgery according to claim 1, characterized in that: The deep convolutional neural network is a ResNet-18 network pre-trained on the ImageNet dataset.

3. The best shot intelligent prediction method for open surgery according to claim 1, characterized in that: The object detection model is a YOLOv5s object detection model trained on a private open surgical object detection dataset.

4. The best shot intelligent prediction method for open surgery according to claim 1, characterized in that: Each Time-Block module processes the potential feature vector X obtained by reducing the dimensionality of the multi-view video image features and semantic features by combining the inputs: Input the low-dimensional feature representation into the globally connected multi-head self-attention mechanism to extract the corresponding contextual features; The obtained context features are subjected to multi-scale convolution operations to extract features at different scales, and the features extracted at different scales are added together to obtain multi-scale convolution features; The multi-scale convolutional features are processed by the first layer of feedforward network and then subjected to a residual connection. The features are then processed by a normalization layer and used as the output of the first layer of feedforward network. The output of the first-layer feedforward network is processed by the second-layer feedforward network, and then a residual connection is performed, and then it is processed by a normalization layer as the output of the Time-Block module.

5. The best shot intelligent prediction method for open surgery according to claim 4, characterized in that: The process of extracting corresponding context features through the globally connected multi-head self-attention mechanism includes: MultiHead(Q,K,V)=Concat(Attention(1),…,Attention(H))Wo Q = Concat(Q(1),…,Q(H)) Among them, MultiHead(Q,K,V) represents a multi-head attention mechanism, Q represents the concatenation of query vectors of all sub-attention heads in the multi-head attention mechanism, K represents the concatenation of key vectors of all sub-attention heads in the multi-head attention mechanism, and V represents the concatenation of value vectors of all sub-attention heads in the multi-head attention mechanism; Attention(h) represents the h-th sub-attention module, h∈{1,2,…,H}, H is the number of heads in the multi-head attention mechanism, Q(h) represents the query vector of the h-th sub-attention module, K(h) represents the value of the h-th sub-attention module, and V(h) represents the value of the h-th sub-attention module; Wo represents the weight matrix of the final linear transformation; dk represents the dimension of the key vector; softmax(·) represents the softmax function; (·) T Indicates the transpose of a vector or matrix; Concat(·) indicates a concatenation operation.

6. The best shot intelligent prediction method for open surgery according to claim 4, characterized in that: The output of the Time-Block module is expressed as: X″=LayerNorm(X′+FFN(X′)) in, represents the result of the multi-scale convolution operation MultiscaleInception; X′ represents the characteristics of the output of the first layer normalization operation in the Time-Block module; LayerNorm(·) represents the characteristics of the output of the second layer normalization operation in the Time-Block module; FFN(·) represents the feedforward network; X″ represents the characteristics of the second layer normalization output.

7. The best shot intelligent prediction method for open surgery according to claim 4, characterized in that: The process of obtaining the probability distribution of each camera label through the softmax layer includes: Where P(y=n|X″) represents the probability of selecting lens label n given feature X″, y represents the category or selection of camera label, and n represents the specific category of camera label; W n represents the learnable weight vector associated with camera label n, b n Represents the bias term associated with camera label n.

8. The best shot intelligent prediction method for open surgery according to claim 1, characterized in that: The loss function used when training a network stacked by Time-Block modules and a softmax layer is expressed as: Among them, L weighted is the loss function; N is the total number of samples in the data set; C is the number of different categories, that is, the number of shots; w c To represent the weights of different categories, it is used to deal with the situation of category imbalance; y i,c represents the true label of the i-th sample in the c-th category; p i,c Output the probability value of the corresponding category for the model.

9. An intelligent prediction system for optimal shots in open surgery, characterized in that: The method for realizing the best shot intelligent prediction method for open surgery as described in claim 1 comprises a deep convolutional neural network, a target detection network, a splicing module, a temporal feature extraction network and a classifier, wherein: A deep convolutional neural network, implemented using a pre-trained ResNet-18 network, is used to extract video features from video frames; The object detection network is implemented using the trained YOLOv5s object detection model to extract semantic features from video frames; A concatenation module, used to concatenate video features and semantic features together as a joint feature vector; The time feature extraction network is a feature extraction network obtained by stacking multiple Time-Block modules, which is used to extract the time feature vector from the joint feature vector; A classifier is used to classify according to the time feature vectors of multiple cameras, and select a camera with the best viewing angle from the multiple cameras.

Citation Information

Patent Citations

  • Monocular depth prediction algorithm based on multi-scale progressive interaction and aggregation cross attention features

    CN116485860A

  • Open type operation scene graph automatic generation method, system, equipment and medium

    CN117746294A

  • Switching method and device of dual-lens motion camera and computer equipment

    CN118574002A

  • Method for searching for optimum view angle in multi-view angle environment and computer program product thereof

    TW201427413A