An operating room intelligent monitoring method and system based on monitoring video recognition
By using a multi-camera array and deep learning technology, an intelligent monitoring system for the operating room is constructed to achieve panoramic monitoring of the surgical process and real-time anomaly detection. This solves the problems of low regulatory efficiency and insufficient intelligent analysis in existing operating room monitoring systems, and improves the efficiency and reliability of operating room safety management.
Patent Information
- Application Number
- CN202510221700.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-02-27
Smart Images

Figure CN120147924B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical safety supervision, and particularly relates to a surgical room intelligent monitoring method and system based on monitoring video recognition. BACKGROUND
[0002] At present, surgical room safety supervision mainly relies on manual patrol and fixed camera video playback, and there are problems such as low supervision efficiency, poor real-time performance, and delayed response to abnormal events. The traditional surgical room monitoring system can only perform basic video acquisition and storage, and cannot intelligently analyze and real-time early warn the key information such as personnel behavior, medical instrument use, and surgical process standardization during the operation. At the same time, the existing surgical room management system lacks the dynamic monitoring ability of the whole operation process, and it is difficult to discover and intervene potential medical safety hazards in time.
[0003] In the prior art, part of the research attempts to use computer vision technology to analyze the surgical room scene, but there are still great limitations in target detection accuracy, multi-target real-time tracking, and surgical process modeling in complex environments. In addition, due to the lack of systematic knowledge guidance mechanism and intelligent abnormal recognition method, it is difficult to accurately evaluate and timely early warn the standardization of surgical operation, and it is unable to meet the actual needs of modern surgical room safety management. SUMMARY
[0004] In view of the problems such as low manual supervision efficiency, delayed response to abnormal events, and lack of intelligent analysis ability of the existing surgical room supervision system, the present application is proposed.
[0005] Therefore, the problem to be solved by the present application is how to use computer vision and deep learning technology to realize intelligent, automatic, and real-time monitoring in the surgical room environment, timely discover and early warn the illegal operation and safety hazards in the operation process, and improve the efficiency and reliability of the surgical room safety management.
[0006] To solve the above technical problems, the present application provides the following technical solutions:
[0007] In a first aspect, the embodiments of the present application provide a surgical room intelligent monitoring method based on monitoring video recognition, which comprises: arranging a surgical room camera array to collect surgical process videos by region block; extracting video frame features and calculating an optical flow field through a residual neural network; fusing multiple video streams based on a feature alignment module guided by the optical flow to construct a surgical region panorama sequence; inputting the surgical region panorama sequence into a double-branch target detection network, extracting medical staff positions and patient body position data through a main branch, extracting surgical instrument features through an auxiliary branch by an attention mechanism, and constructing a surgical scene model containing personnel spatial positions and instrument use areas based on a multi-view mapping algorithm; obtaining medical staff posture data by applying a skeleton key point extraction network based on the personnel position sequence in the surgical scene model, performing trajectory tracking based on Kalman filtering in combination with the instrument use areas in the scene model, and constructing a node-weighted surgical process dynamic graph; inputting the surgical process dynamic graph into a time series graph convolution network, calculating a dynamic spatial correlation matrix through a multi-head attention mechanism, fusing a surgical type encoding vector, and generating a multi-dimensional surgical progress state report; comparing the progress state report with a surgical specification library by using a knowledge distillation algorithm, dividing abnormal events into different levels according to the degree of violation and safety risks, and recording a surgical monitoring log.
[0008] As a preferred scheme of the surgical room intelligent monitoring method based on monitoring video recognition, the method comprises the following steps: constructing the surgical region panorama sequence comprises the following steps; arranging a plurality of high-definition cameras to construct a surgical room monitoring array, dividing the surgical room into a plurality of monitoring areas, and collecting multiple original video data streams; pre-processing each video data stream to obtain standardized RGB three-channel video data; inputting the RGB three-channel video data into a residual neural network, extracting multi-scale features through a plurality of residual blocks, each residual block containing a convolution layer and a short circuit connection, and outputting a video feature vector; calculating motion information between adjacent video frames by using an optical flow algorithm, constructing an optical flow map, and fusing the optical flow map and the video feature vector to obtain dynamic feature data; constructing a feature alignment module based on camera calibration parameters, spatially resampling the dynamic feature data to form perspective alignment features; performing feature concatenation on the perspective alignment features under each camera perspective, performing feature selection and fusion by using a channel attention and a spatial attention weighting strategy, reconstructing an image through a deconvolution network, and outputting the surgical region panorama sequence.
[0009] As a preferred scheme of the operating room intelligent monitoring method based on monitoring video recognition provided by the application, wherein: the step of constructing the operating scene model containing the spatial position of personnel and the use area of instruments based on the multi-view mapping algorithm comprises the following steps: dividing the panoramic sequence of the operating area into grid units, extracting a target detection feature vector, and constructing input data of a feature pyramid network; performing feature extraction on the input data through a main branch convolutional layer, each convolutional block containing a normalization layer and an activation function, forming a personnel target feature map; inputting the personnel target feature map into a region proposal network, screening candidate regions through non-maximum suppression, and outputting the position coordinates of medical staff and the posture parameters of a patient; performing channel attention weighting on the input data of an auxiliary branch, distributing channel weight coefficients, highlighting the features of surgical instruments, and forming an instrument feature map; performing region detection on the instrument feature map, extracting local response values, and determining the use area and class label of surgical instruments; calculating the spatial mapping relationship based on the camera calibration parameters, converting the position coordinates of medical staff, the posture parameters of a patient, and the use area of instruments to a unified coordinate system, and constructing an operating scene model.
[0010] As a preferred scheme of the operating room intelligent monitoring method based on monitoring video recognition provided by the application, wherein: the step of constructing the operating scene model containing the spatial position of personnel and the use area of instruments based on the multi-view mapping algorithm comprises the following steps: dividing the panoramic sequence of the operating area into grid units, extracting a target detection feature vector, and constructing input data of a feature pyramid network; performing feature extraction on the input data through a main branch convolutional layer, each convolutional block containing a normalization layer and an activation function, forming a personnel target feature map; inputting the personnel target feature map into a region proposal network, screening candidate regions through non-maximum suppression, and outputting the position coordinates of medical staff and the posture parameters of a patient; performing channel attention weighting on the input data of an auxiliary branch, distributing channel weight coefficients, highlighting the features of surgical instruments, and forming an instrument feature map; performing region detection on the instrument feature map, extracting local response values, and determining the use area and class label of surgical instruments; calculating the spatial mapping relationship based on the camera calibration parameters, converting the position coordinates of medical staff, the posture parameters of a patient, and the use area of instruments to a unified coordinate system, and constructing an operating scene model.
[0011] As a preferred scheme of the operating room intelligent monitoring method based on monitoring video recognition provided by the application, the generation of the multi-dimensional operation progress state report comprises the following steps: dividing the operation flow dynamic graph into a plurality of time windows in time sequence, sampling the node features in each time window in time sequence, and constructing a graph sequence sample matrix; performing feature extraction on the graph sequence sample matrix through a time series graph convolution network, wherein a multilayer perceptron is used to update the node features, a convergence layer is used to integrate neighborhood information, and a graph time sequence feature is output; applying a multi-head attention mechanism to the graph time sequence feature, projecting the feature to a query space, a key space, and a value space, calculating an attention score, and obtaining a dynamic correlation strength between nodes; constructing a dynamic space correlation matrix based on the dynamic correlation strength, wherein the matrix elements represent the interaction degree between node pairs, and a normalized correlation matrix is obtained through normalization processing; encoding the operation type information into a vector representation, mapping the operation type vector to the feature space through a type embedding matrix, and performing feature fusion on the normalized correlation matrix to obtain a fusion feature tensor; decoupling the fusion feature tensor in time and space dimensions, respectively extracting the time sequence variation feature and the spatial distribution feature of the operation progress, and constructing a multi-dimensional feature combination; constructing a multi-level medical expert knowledge tree, hierarchically encoding the expert experience, and establishing a knowledge-guided evaluation rule library; using the evaluation rule library to perform dynamic evaluation of the multi-dimensional feature combination guided by expert knowledge, combining the current operation stage to calculate the importance weight of the step, and generating an operation progress score index; integrating the operation progress score index according to a preset template, adding a timestamp and scene description information, and outputting a multi-dimensional operation progress state report.
[0012] As a preferred scheme of the operating room intelligent monitoring method based on monitoring video recognition provided by the application, the generation of the multi-dimensional operation progress state report comprises the following steps: dividing the operation flow dynamic graph into a plurality of time windows in time sequence, sampling the node features in each time window in time sequence, and constructing a graph sequence sample matrix; performing feature extraction on the graph sequence sample matrix through a time series graph convolution network, wherein a multilayer perceptron is used to update the node features, a convergence layer is used to integrate neighborhood information, and a graph time sequence feature is output; applying a multi-head attention mechanism to the graph time sequence feature, projecting the feature to a query space, a key space, and a value space, calculating an attention score, and obtaining a dynamic correlation strength between nodes; constructing a dynamic space correlation matrix based on the dynamic correlation strength, wherein the matrix elements represent the interaction degree between node pairs, and a normalized correlation matrix is obtained through normalization processing; encoding the operation type information into a vector representation, mapping the operation type vector to the feature space through a type embedding matrix, and performing feature fusion on the normalized correlation matrix to obtain a fusion feature tensor; decoupling the fusion feature tensor in time and space dimensions, respectively extracting the time sequence variation feature and the spatial distribution feature of the operation progress, and constructing a multi-dimensional feature combination; constructing a multi-level medical expert knowledge tree, hierarchically encoding the expert experience, and establishing a knowledge-guided evaluation rule library; using the evaluation rule library to perform dynamic evaluation of the multi-dimensional feature combination guided by expert knowledge, combining the current operation stage to calculate the importance weight of the step, and generating an operation progress score index; integrating the operation progress score index according to a preset template, adding a timestamp and scene description information, and outputting a multi-dimensional operation progress state report.
[0013] As a preferred scheme of the operating room intelligent monitoring method based on monitoring video recognition provided by the application, wherein: the step of generating the anomaly detection matrix comprises the following steps: dimension alignment is performed on the monitoring feature vector and the standard feature template, the two groups of features are mapped to the same feature space to obtain a uniformly represented feature pair; the feature pair is processed through a double-flow network, a main branch calculates a similarity score of the feature pair, and an auxiliary branch extracts a difference feature of the feature pair to obtain a feature matching measurement value; a time sequence sliding window is established, the feature matching measurement value is sampled, and a time-varying deviation feature is extracted through a long short-term memory network; a behavior sequence deviation, an instrument operation deviation and a personnel interaction deviation of the time-varying deviation feature are calculated to obtain a deviation feature group; and the deviation feature group is reorganized into a three-dimensional tensor structure according to a time axis, a space axis and a feature axis to generate the anomaly detection matrix.
[0014] In a second aspect, the embodiments of the application provide an operating room intelligent monitoring system based on monitoring video recognition, which comprises a video acquisition and processing module for arranging an operating room camera array to collect operating process videos by region block, extracting video frame features through a residual neural network and calculating an optical flow field, and fusing multiple video streams based on a feature alignment module guided by the optical flow to construct an operating area panorama sequence; a target detection and recognition module for inputting the operating area panorama sequence into a double-branch target detection network, extracting medical staff position and patient body position data through a main branch, extracting surgical instrument features through an auxiliary branch by an attention mechanism, and constructing a surgical scene model containing personnel spatial position and instrument use area based on a multi-view mapping algorithm; a behavior tracking analysis module for obtaining medical staff posture data by applying a skeleton key point extraction network based on the personnel position sequence in the surgical scene model, performing trajectory tracking based on Kalman filtering in combination with the instrument use area in the scene model, and constructing a node-weighted surgical process dynamic graph; a state evaluation generation module for inputting the surgical process dynamic graph into a time sequence graph convolution network, calculating a dynamic spatial correlation matrix through a multi-head attention mechanism, fusing a surgical type encoding vector, and generating a multi-dimensional surgical progress state report; and an anomaly monitoring and recording module for comparing the progress state report with a surgical specification library by using a knowledge distillation algorithm, dividing abnormal events into different levels according to the violation degree and safety risk, and recording a surgical monitoring log.
[0015] The beneficial effects of the present application are: the present application successfully solves the technical problem of incomplete coverage of a single view in a traditional operating room monitoring system through a multi-camera array partition collection strategy, significantly improving the monitoring coverage of the operating area. The designed double-branch network architecture and attention mechanism realize the precise differentiated extraction of medical staff and surgical instrument features, effectively improving the accuracy of target detection, and the innovative design of the skeleton key point extraction network and three-dimensional convolution block realizes high-precision capture and analysis of the fine actions of medical staff. The developed timing diagram convolution network and multi-head attention mechanism successfully realize dynamic feature extraction and depth correlation analysis of the surgical process, and through the construction of a hierarchical medical expert knowledge tree, expert experience is effectively converted into executable evaluation rules. The proposed three-branch encoding network and spatiotemporal attention mechanism significantly improve the extraction ability and template generation accuracy of surgical specification features, and the innovative combination of the double-flow comparison network and the long short-term memory network realizes comprehensive capture of instantaneous and long-term abnormal patterns in the surgical process, significantly improving the sensitivity of abnormal detection. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0017] Fig. 1 The overall flowchart of the operating room intelligent monitoring method based on monitoring video recognition.
[0018] Fig. 2 The operating area panoramic sequence construction flowchart of the operating room intelligent monitoring method based on monitoring video recognition.
[0019] Fig. 3 The surgical process dynamic graph construction flowchart of the operating room intelligent monitoring method based on monitoring video recognition. DETAILED DESCRIPTION
[0020] In order to make the above-mentioned purposes, features and advantages of the present application more apparent and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings of the specification.
[0021] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, but the present application can also be implemented in other ways different from the description, and those skilled in the art can make similar generalizations without departing from the connotation of the present application, therefore the present application is not limited to the specific embodiments disclosed below.
[0022] Secondly, the "one embodiment" or "embodiment" referred to herein is intended to mean a specific feature, structure, or characteristic under at least one implementation of the application. The "in one embodiment" appearing in various places in this specification are not all referring to the same embodiment nor are they mutually exclusive of other embodiments.
[0023] Embodiment 1, refer to Fig. 1-3 For the first embodiment of the application, the embodiment provides a surgery room intelligent monitoring method based on monitoring video recognition, and the overall flow chart is as Fig. 1 shown.
[0024] In the prior art, the surgery room video monitoring system has the following problems: the traditional monitoring system usually uses a single camera for video acquisition, which leads to incomplete coverage of the visual angle and makes it difficult to achieve effective monitoring of the whole scene of the operation; the existing video processing method has poor adaptability to the complex lighting environment in the operating room, which is easy to produce image distortion and affects the monitoring effect; at the same time, such system lacks intelligent recognition ability for the behaviors of medical staff and the use of surgical instruments, and cannot realize automatic monitoring and analysis of the operation process. In addition, due to the failure to fully utilize medical expert knowledge for evaluation and guidance, there is obvious deficiency in abnormal behavior recognition and risk assessment.
[0025] To solve the above technical problems, the application provides a surgery room intelligent monitoring method based on monitoring video recognition, which comprises the following steps:
[0026] S1: arranging a surgery room camera array to collect operation process videos by region, extracting video frame features and calculating an optical flow field through a residual neural network, fusing multiple video streams based on a light flow guided feature alignment module, and constructing a surgery region panorama sequence.
[0027] Specifically, the surgery region panorama sequence construction flow chart is as Fig. 2 shown, comprising the following steps:
[0028] S1.1: arranging multiple high-definition cameras to construct a surgery room monitoring array, dividing the surgery room into multiple monitoring regions, and collecting multiple original video data streams.
[0029] S1.2: performing video frame denoising, illumination normalization and size unification processing on each video data stream to obtain standardized RGB three-channel video data.
[0030] It should be noted that the video frame denoising uses a bilateral filtering algorithm to remove Gaussian noise and impulse noise, the illumination normalization compensates for uneven illumination through an adaptive histogram equalization method, and the size unification resamples all video frames to a preset resolution to ensure the consistency of feature extraction.
[0031] S1.3: input the RGB three-channel video data into the residual neural network, extract multi-scale features through multiple residual blocks, each residual block contains a convolution layer and a short circuit connection, and output a video feature vector.
[0032] S1.4: calculate the motion information between adjacent video frames by using the optical flow algorithm, construct an optical flow map, and fuse the optical flow map with the video feature vector to obtain dynamic feature data.
[0033] S1.5: construct a feature alignment module based on the camera calibration parameters, spatially resample the dynamic feature data, and form perspective alignment features.
[0034] S1.6: concatenate the perspective alignment features under each camera perspective, use a channel attention and spatial attention weighting strategy for feature selection and fusion, reconstruct the image through a deconvolution network, and output a panoramic sequence of the surgical area.
[0035] Specifically, the perspective alignment features collected by multiple cameras are concatenated in the channel dimension to construct a multi-perspective feature tensor, and batch normalization is used to obtain a standardized feature mapping; a channel attention module is constructed, global average pooling and global maximum pooling are used to extract feature statistics in parallel, a two-layer fully connected network is used to generate a channel weight vector, and the channel weight vector is multiplied with the standardized feature mapping to obtain a channel weighted feature; a spatial attention module is constructed, the maximum value and average value of the channel weighted feature are calculated along the channel dimension, the maximum value feature map and the average value feature map are spliced, and a spatial attention map is generated through a convolution layer and a Sigmoid activation function, and the spatial attention map is multiplied with the channel weighted feature to obtain a double-weighted feature; a deconvolution network with a multi-layer U-Net structure is built, each layer contains an upsampling module, a convolution layer and an activation function, the double-weighted feature is input into the deconvolution network, and the image spatial resolution is restored through layer-by-layer upsampling and feature reconstruction; a skip connection is set in the deconvolution network, the intermediate layer features in the encoding stage are fused with the corresponding layer features in the decoding stage, and a reconstructed feature map is output; an image post-processing module is applied to the reconstructed feature map, including edge enhancement and contrast adjustment, and a clear surgical area image is generated, and the continuous frame images are arranged in time sequence to form a panoramic sequence of the surgical area.
[0036] It is worth noting that the multi-camera array partition acquisition strategy overcomes the limitations of a single perspective, significantly improving the monitoring coverage of the surgical area; the multi-level video preprocessing mechanism enhances the adaptability of the system to complex lighting environments, effectively ensuring the image quality; the residual neural network combined with the optical flow feature design improves the accuracy of dynamic scene feature extraction; the fusion strategy of the feature alignment module and the double attention mechanism realizes the precise integration of multi-perspective information, significantly improving the restoration quality of the panoramic sequence.
[0037] S2: input the panoramic sequence of the surgical area into the double-branch target detection network, the main branch extracts the medical staff position and patient body position data, the auxiliary branch extracts the surgical instrument features through the attention mechanism, and a surgical scene model containing the spatial positions of the medical staff and the use area of the instruments is constructed based on a multi-view mapping algorithm.
[0038] Specifically, the following steps are included:
[0039] S2.1: divide the panoramic sequence of the surgical area into grid cells, extract a target detection feature vector, and construct input data for a feature pyramid network.
[0040] S2.2: extract features from the input data through the main branch convolutional layer, each convolutional block containing a normalization layer and an activation function, and form a medical staff target feature map.
[0041] S2.3: input the medical staff target feature map into a region proposal network, screen the candidate regions through non-maximum suppression, and output the medical staff position coordinates and patient body posture parameters.
[0042] The medical staff target feature map is input into a region proposal module, the region proposal module contains an anchor box set of a preset size, generates a plurality of candidate boxes based on the feature map position, calculates the target scores of the candidate boxes through a regression network, a anchor box regression branch is constructed, the branch predicts the center point offset and size scaling coefficient of each candidate box through four convolutional layers, performs coordinate transformation on the predicted box to obtain a corrected target box, a classification branch is constructed, the classification branch uses three convolutional layers to identify the features in the corrected target box, and outputs the medical staff class probability and patient class probability of the target box; calculate the detection frame confidence based on the target score and the class probability, screen the overlapping detection frames using the non-maximum suppression algorithm, and retain the detection frames with the highest confidence and an overlap rate lower than a set threshold; extract the center coordinates of the medical staff detection frame, establish a mapping matrix from the image coordinates to the actual space of the operating room, convert the two-dimensional image coordinates into three-dimensional space coordinates according to the mapping matrix, and output the spatial position parameters of the medical staff; apply a posture estimation sub-network to the patient detection frame, the sub-network includes three layers of depth separable convolutional layers, extracts the human body skeleton key point coordinates through a heat map response, calculates the trunk inclination angle and limb joint angle based on the key point coordinates, and outputs the patient body posture parameters.
[0043] S2.4: perform channel attention weighting on the input data of the auxiliary branch, assign channel weight coefficients, highlight the features of the surgical instruments, and form an instrument feature map.
[0044] It should be noted that the input data of the auxiliary branch is compressed by the feature compression unit, and the spatial compression is performed on each feature channel to obtain a channel description vector; a feature weight learning module is constructed, which includes two layers of feature transformation networks, the first layer performs feature compression on the channel description vector, and the second layer restores the compressed features to the original feature dimension to output initial feature weights; the initial feature weights are subjected to standardization processing, and the weight values are mapped to a preset interval to generate standardized weight coefficients; a feature enhancement matrix is constructed based on the standardized weight coefficients, the dimension of the feature enhancement matrix matches the feature dimension of the input data, and the input data is subjected to feature enhancement operation; the enhanced features are combined with the original input features through a feature fusion module, the feature distribution is adjusted through a feature balancing unit, and enhanced feature data is output; the enhanced feature data is subjected to spatial information integration by a feature integration network, the feature integration network uses a preset convolution unit to extract local feature correlation, and outputs a device feature map after nonlinear transformation processing.
[0045] S2.5: Region detection is performed on the device feature map, local response values are extracted, and the use area and class label of the surgical instrument are determined.
[0046] In specific implementation, a target detection network is constructed, the network includes multiple layers of feature extraction units, each layer of feature extraction is followed by a feature normalization unit and a nonlinear transformation unit, and multi-scale feature representations are extracted on the device feature map; a preset detection frame is set for each scale level of the feature representation, detection frame parameters are determined based on the size distribution of the surgical instrument to generate a detection frame position mapping; the feature representation is input into a detection processing network, which is divided into a position prediction branch and a class prediction branch, the position prediction branch predicts the position parameters and scale parameters of the detection frame, and the class prediction branch outputs the class probability distribution of each type of surgical instrument; the predicted frame is decoded, the target frame coordinates are calculated according to the detection frame position mapping and the predicted parameters, the probability mapping function is used to process the class probability distribution to obtain the instrument type of each target frame; the comprehensive score is calculated based on the class probability and position confidence of the detection frame, the detection frame filtering algorithm is executed on the overlapping detection frames, the locally optimal detection result is retained, and the use area and class label of the surgical instrument are output.
[0047] S2.6: The spatial mapping relationship is calculated based on the camera calibration parameters, the medical staff position coordinates, the patient body posture parameters and the instrument use area are converted to a unified coordinate system, and a surgical scene model is constructed.
[0048] Notably, this step has the following technical effects: the innovative design of the double-branch network architecture and attention mechanism realizes the differentiated feature extraction of medical staff and surgical instruments, significantly improving the target detection accuracy; the application of the feature enhancement matrix makes the system have better adaptability to targets of different scales, greatly reducing the detection false negative rate; the introduction of the multi-view mapping algorithm realizes high-precision spatial positioning, providing reliable guarantee for accurate modeling of the surgical scene; the design of the depth separable convolution layer and the heat map response, combined with the attention mechanism enhanced surgical instrument detection, forms a complete surgical scene monitoring solution.
[0049] S3: Based on the personnel position sequence in the surgical scene model, the skeleton key point extraction network is applied to obtain the medical staff posture data, combined with the instrument use area in the scene model to perform trajectory tracking based on Kalman filtering, and a node weighted surgical process dynamic graph is constructed.
[0050] Specifically, the surgical process dynamic graph construction process is as shown in the flowchart of Fig. 3 , which includes the following steps:
[0051] S3.1: Extract the medical staff position sequence from the surgical scene model, construct a time sequence sliding window, time-align the medical staff position data in each sliding window, and construct a behavior analysis sample.
[0052] S3.2: Input the behavior analysis sample into the skeleton key point extraction network, extract human joint features through a multi-layer feature encoder, and form a posture skeleton feature map.
[0053] S3.3: Perform spatio-temporal convolution operation on the posture skeleton feature map, extract human posture change features, fuse joint motion information through self-attention mechanism, and output posture feature sequence.
[0054] The three-dimensional convolution block is constructed based on the pose skeleton feature map, wherein the convolution block contains multiple layers of convolution kernels, and the convolution kernels slide in the spatial dimension and the time dimension respectively to extract the space-time domain features and output an initial pose sequence; a spatial topology graph is established for the human body joint nodes in the initial pose sequence, the connection relationship between the joint nodes is defined according to the anatomical relationship, and a pose feature adjacency matrix is constructed; a graph convolution module is used to perform feature transmission on the pose feature adjacency matrix, and each node updates its own features based on the features of adjacent nodes to generate pose topology features; a multi-level self-attention unit is established, each unit contains a query matrix, a key-value matrix and a numerical matrix, the attention score is calculated through matrix multiplication, and the joint node attention weight is output; the pose topology features are multiplied by the joint node attention weight to highlight the motion features of the key joint nodes, and weighted pose features are formed; a time sequence fusion module is established, a long short-term memory network is used to perform time sequence modeling on the weighted pose features, and the pose change law between continuous frames is captured; the time sequence modeling result is subjected to feature dimension reduction, the high-dimensional features are mapped to a target feature space through a fully connected layer, and a standardized pose vector is generated; the standardized pose vector is combined and arranged according to the time sequence, time sequence position coding is added, and a pose feature sequence containing human joint motion information is output.
[0055] S3.4: Extract the surgical instrument use area and class label from the surgical scene model, apply the Kalman filter algorithm to predict the instrument position change, and obtain the instrument motion sequence.
[0056] S3.5: Establish a space-time association model of the pose feature sequence and the instrument motion sequence, calculate the person-instrument interaction strength, and construct an initial process feature map.
[0057] Specifically, the pose feature sequence is divided into pose segments according to a time window, the spatial coordinates of the human body joint nodes are extracted for each pose segment, and a person trajectory matrix is constructed; an instrument trajectory feature map is established based on the instrument motion sequence, which contains instrument position coordinates, moving direction and speed parameters, and forms an instrument motion descriptor; a space-time interaction measure model is constructed, which includes a distance calculation unit and a direction matching unit, and calculates the spatial overlap degree of the person trajectory and the instrument trajectory; a multi-layer perceptron network is established, the person trajectory matrix and the instrument motion descriptor are input, the person-instrument pairing features are extracted, and an interaction candidate is generated; the interaction candidate is screened by a probability threshold, low-probability interactions are removed, and valid person-instrument interaction pairs are retained to form an interaction relationship graph; the correlation degree between nodes is calculated based on the interaction relationship graph, the interaction strength is weighted by using a distance decay function, and an interaction strength matrix is output; the interaction strength matrix is combined with the time sequence context information, the node representation is learned through a graph embedding network, and a graph feature vector is constructed; a graph reconstruction module is used to restore the graph feature vector to a node-edge structure, the node represents the person and the instrument, and the edge represents the interaction relationship, and an initial process feature map is constructed.
[0058] S3.6: Calculate the importance score of the nodes in the initial procedure feature graph, assign weight parameters based on the temporal correlation and spatial dependence between nodes, and construct a node-weighted surgical procedure dynamic graph.
[0059] Further, a feature vector matrix is constructed for the nodes of the initial procedure feature graph, which includes node type, position information and interaction relationship, forming a node description table; a bidirectional temporal propagation network is established, which includes a forward propagation layer and a backward propagation layer, to extract the temporal dependence features between nodes and generate a temporal correlation matrix; a spatial attention module is used to calculate the spatial distance between node pairs, and the distance data is input into a Gaussian kernel function to output a spatial dependence coefficient; a node importance evaluator is constructed, which integrates node degree centrality, feature clustering coefficient and interaction frequency to calculate node importance score; an edge weight calculation model is established based on the temporal correlation matrix and the spatial dependence coefficient to quantify the connection strength between node pairs and generate an edge weight vector; the node importance score and the edge weight vector are combined and normalized to obtain standardized weight parameters, and a weight distribution table is constructed; a graph update module is designed to apply the weight distribution table to the initial procedure feature graph to update node attributes and edge attributes, forming a weighted feature graph; the weighted feature graph is temporally sampled to establish a dynamic graph sequence, and a timestamp is added to output the node-weighted surgical procedure dynamic graph.
[0060] Notably, this step has the following technical effects: the innovative design of the skeleton key point extraction network and the three-dimensional convolution block accurately captures the fine actions of medical staff; the introduction of the multi-level self-attention mechanism enhances the feature expression of key nodes, improving the accuracy of pose recognition; the application of the spatio-temporal interaction measure model and the graph embedding network enables the system to accurately model the person-device interaction relationship; the design of the bidirectional temporal propagation network and the spatial attention module forms a complete surgical procedure dynamic representation scheme; the optimization design of node importance evaluation and edge weight calculation accurately identifies and quantifies the importance of key nodes in the surgical procedure.
[0061] S4: Input the surgical procedure dynamic graph into the temporal graph convolution network, calculate the dynamic spatial correlation matrix through the multi-head attention mechanism, and fuse the surgical type encoding vector to generate a multi-dimensional surgical process status report.
[0062] Specifically, the following steps are included:
[0063] S4.1: Divide the surgical procedure dynamic graph into multiple time windows according to the time sequence, and perform temporal sampling on the node features within each time window to construct a graph sequence sample matrix.
[0064] S4.2: Feature extraction on the graph sequence sample matrix by a temporal graph convolutional network, in which a multi-layer perceptron is used to update the node features in the graph convolutional layer, and a pooling layer is used to integrate the neighborhood information to output the graph temporal features.
[0065] S4.3: A multi-head attention mechanism is applied to the graph temporal features to project the features into the query space, key space, and value space, calculate the attention scores, and obtain the dynamic correlation strength between nodes.
[0066] S4.4: A dynamic spatial correlation matrix is constructed based on the dynamic correlation strength, and the matrix elements represent the interaction degree between node pairs. The normalized correlation matrix is obtained through normalization processing.
[0067] S4.5: The surgery type information is encoded into a vector representation, and a type embedding matrix is used to map the surgery type vector to the feature space. The fusion feature tensor is obtained by feature fusion between the normalized correlation matrix and the feature tensor.
[0068] S4.6: The fusion feature tensor is decoupled in the time and space dimensions, and the temporal variation features and spatial distribution features of the surgery process are extracted respectively to construct a multi-dimensional feature combination.
[0069] S4.7: A multi-level medical expert knowledge tree is constructed, the expert experience is hierarchically encoded, and a knowledge-guided evaluation rule base is established.
[0070] In this embodiment, the surgery experience data of multiple medical experts is collected, the experience data is classified and labeled, and the expert knowledge sample set is constructed according to the surgery type, operation specification, and safety requirements. Based on the expert knowledge sample set, a hierarchical knowledge structure is established, which includes the surgery process layer, the operation specification layer, and the safety monitoring layer, forming an initial knowledge tree. The nodes of the initial knowledge tree are semantically encoded, the text knowledge is converted into numerical features using a word vector model, and a knowledge feature matrix is generated. A knowledge graph network is constructed, which connects the nodes in the knowledge feature matrix according to the logical relationship, learns the node representation through the graph embedding algorithm, and outputs the knowledge vector. The knowledge vector is classified using a hierarchical clustering algorithm, the knowledge clusters are divided based on the similarity threshold, and a knowledge category index table is established. A rule reasoning module is designed, which includes a condition judgment unit and a result derivation unit, converts the expert experience into formal rules, and generates a rule set. The rule set is prioritized, the weight coefficients are assigned based on the rule applicability and importance, and a rule weight matrix is constructed. The knowledge category index table and the rule weight matrix are integrated to establish the knowledge rule mapping relationship, and a hierarchical evaluation rule base is output.
[0071] S4.8: The multi-dimensional feature combination is dynamically evaluated by the expert knowledge guided by the evaluation rule base, the step importance weight is calculated combined with the current surgery stage, and the surgery process score index is generated.
[0072] In particular implementation, the temporal features and spatial features are extracted from the multi-dimensional feature combination to establish a feature analysis module, which classifies the features according to a preset feature category standard to form a feature category mapping table; a surgical stage recognizer is constructed, which performs sliding window analysis based on the temporal feature sequence, judges the current surgical stage using a state transition model, and outputs a stage identifier; a rule matching module is established by retrieving a rule subset corresponding to the current surgical stage from an evaluation rule library, which calculates the matching degree between the features and the rules to generate a rule applicability matrix; a dynamic weight calculation unit is designed, which assigns importance coefficients to different evaluation dimensions according to the surgical stage identifier and the rule applicability matrix to construct a weight distribution vector; a multi-layer evaluation network is established, which inputs the feature category mapping table and the weight distribution vector into each evaluation layer to calculate single-dimensional scores through weighted combination and outputs scored features; a fuzzy comprehensive evaluation method is used to fuse the scored features, a score grade division standard is established, a threshold boundary is determined according to the score distribution, and a score interval is generated; a score generator is designed, which maps the single-dimensional scores to the corresponding score interval and adjusts the scores in combination with expert rules to form normalized scores; the normalized scores are subjected to temporal smoothing processing to eliminate score fluctuations, a confidence marker is added, and a surgical progress score indicator is output.
[0073] S4.9: The surgical progress score indicator is integrated according to a preset template, time stamp and scene description information are added, and a multi-dimensional surgical progress state report is output.
[0074] Notably, this step has the following technical effects: the innovative design of the temporal graph convolution network and the multi-head attention mechanism realizes dynamic feature extraction and correlation analysis of the surgical procedure; the construction of the multi-level medical expert knowledge tree successfully converts expert experience into formal evaluation rules; the application of the multi-layer evaluation network and the dynamic weight calculation improves the scientificity of the evaluation results; the design of the feature analysis module and the surgical stage recognizer realizes accurate evaluation of the surgical progress; the introduction of the fuzzy comprehensive evaluation method enhances the reliability of the scores, providing important decision support for intelligent monitoring in the operating room.
[0075] S5: The knowledge distillation algorithm is used to compare the progress state report with the surgical specification library, and the abnormal events are divided into different levels according to the degree of violation and safety risk, and a surgical monitoring log containing time stamp and abnormal level is recorded.
[0076] Specifically, the following steps are included:
[0077] S5.1: The multi-dimensional surgical progress state report is decomposed into behavior sequence features, instrument operation features and personnel interaction features to generate a monitoring feature vector.
[0078] S5.2: Extract the standard operation specification corresponding to the operation type from the operation specification library, establish a teacher model through the knowledge distillation network, and output the specification feature template.
[0079] Specifically, the standard operation specification in the operation specification library is divided into operation step sequences according to the time sequence structure, the action description features, instrument configuration features and personnel configuration features are extracted for each operation step, and the specification feature sequence is constructed; the three-branch coding network is used to process the specification feature sequence, the first branch encodes the action semantic information, the second branch encodes the instrument combination information, and the third branch encodes the personnel configuration information to obtain the specification coding vector; the specification coding vector is enhanced based on the space-time attention mechanism, the time sequence dependence strength and the space correlation strength between steps are calculated, and the attention enhanced features are generated; the teacher model of the knowledge distillation network is constructed, the multi-layer perceptron is used to perform nonlinear transformation on the attention enhanced features, and the standard feature distribution of each step is output in the form of soft label; the feature template generator is constructed based on the standard feature distribution of the teacher model, the feature distribution is mapped to the specification feature space, the time sequence information and the space information are integrated, and the specification feature template is output.
[0080] S5.3: Perform feature matching on the monitoring feature vector and the specification feature template, calculate the behavior deviation degree, and generate an anomaly detection matrix.
[0081] Optionally, the monitoring feature vector and the specification feature template are dimensionally aligned, the two groups of features are mapped to the same feature space through the feature projection network to obtain the uniformly represented feature pair; the double-flow comparison network is constructed, the cosine similarity score of the feature pair is calculated in the main flow branch, the local difference features of the feature pair are extracted in the auxiliary flow branch, and the feature matching measurement value is output; the time sequence sliding comparison window is established, the time sequence sampling is performed on the feature matching measurement value, and the long short-term memory network is used to extract the time-varying deviation features of the feature sequence; a multi-scale deviation calculation module is designed to calculate the behavior sequence deviation, the instrument operation deviation and the personnel interaction deviation respectively, and a deviation feature group is constructed; the deviation feature group is reorganized into a three-dimensional tensor structure, and the tensor dimensions correspond to the time axis, the space axis and the feature axis respectively, and an anomaly detection matrix is output.
[0082] It should be noted that the feature projection network adopts a three-layer fully connected layer structure, the first layer input dimension is the original feature dimension, the hidden layer dimension is 512, and the output layer dimension is 256, and each layer is followed by a batch normalization layer and a ReLU activation function; in the double-flow comparison network, the main flow branch calculates the feature matching degree using cosine similarity, and the auxiliary flow branch uses a 1x1 convolution to extract local differences, with 64 convolution kernels and a step size of 1; the window size of the time sequence sliding comparison window is set to 10 frames, and the sliding step size is 5 frames; the long short-term memory network includes two layers of LSTM units, with a hidden layer dimension of 128, and the bias terms of the input and output gates are initialized to 1.0; the multi-scale deviation calculation module sets weight coefficients for the three types of deviation features, with a behavior sequence deviation weight of 0.4, an instrument operation deviation weight of 0.3, and a personnel interaction deviation weight of 0.3.
[0083] In specific implementation, the training of the feature projection network adopts a contrast learning strategy, the loss function uses InfoNCEloss, the temperature parameter is set to 0.07, the Adam optimizer is used for parameter update, and the learning rate is set to 0.0001; the double-flow comparison network is constructed, the main flow branch calculates the cosine similarity score of the feature pair, the similarity threshold is set to 0.75, the auxiliary flow branch extracts the local difference features of the feature pair, the difference features are reduced to 64-dimensional vectors through a 1x1 convolution layer and a global average pooling layer, and the feature matching metric value is output; the time sequence sliding comparison window is established, the feature matching metric value is time sequence sampled, the sampling frequency is 10Hz, the long short-term memory network is used to extract the time-varying deviation features of the feature sequence, and the time-varying deviation features have a dimension of 128; the multi-scale deviation calculation module is designed to calculate the behavior sequence deviation, the instrument operation deviation, and the personnel interaction deviation, respectively, the deviation calculation adopts the Euclidean distance metric, and the deviation feature group is constructed; the deviation feature group is reorganized into a three-dimensional tensor structure, the tensor dimension is [T, S, F], T is the time dimension length, S is the space dimension length, and F is the feature dimension length, the tensor dimensions correspond to the time axis, the space axis, and the feature axis respectively, and the abnormality detection matrix is output.
[0084] S5.4: Construct a multi-level anomaly evaluation model based on the abnormality detection matrix, classify the abnormal events according to the violation degree threshold and the safety risk coefficient, and output the risk level data.
[0085] Further, the anomaly detection matrix is unfolded along the time dimension and the space dimension to extract multi-dimensional abnormal features including behavior abnormality degree, instrument operation abnormality degree and personnel interaction abnormality degree; a three-layer abnormality evaluation network is constructed, the input layer receives the multi-dimensional abnormal features, the hidden layer adopts a residual connection structure to fuse the multi-dimensional abnormal information, and the output layer generates an abnormality degree score; a hierarchical threshold discriminator is established, a rule violation degree benchmark threshold is set according to the type of surgery, and the threshold is dynamically adjusted in combination with the safety risk coefficient of the surgery stage; a difference coefficient of the abnormality degree score and the adjusted threshold is calculated, the abnormal events are classified and labeled according to a preset risk level interval, and a risk level index is generated; the risk level index is associated with the surgery scene information, an abnormal type identifier and a danger degree marker are added, and risk level data containing multi-level classification results are output.
[0086] S5.5: The risk level data is associated with a system timestamp to generate an abnormal event descriptor.
[0087] S5.6: The abnormal event descriptor is written into a monitoring database according to the priority order of the abnormal level to generate a surgery monitoring log containing a timestamp and an abnormal level.
[0088] Preferably, the present application adopts a three-branch coding network and a space-time attention mechanism design, which significantly improves the extraction ability of surgical specification features, making the generation of specification templates more accurate; the innovative architecture of the double-flow comparison network combined with the long short-term memory network enables the system to capture both instantaneous and long-term abnormal patterns, significantly improving the sensitivity of anomaly detection; the application of the three-layer abnormality evaluation network and the dynamic threshold adjustment mechanism enables the system to adaptively evaluate the risk according to the type and stage characteristics of the surgery, significantly improving the accuracy of abnormal classification; through the design of abnormal event priority ordering and timestamp association, the system realizes efficient abnormal event tracking and recording, providing strong support for surgery quality control and safety management.
[0089] Further, the embodiment also provides a surgery room intelligent monitoring system based on monitoring video recognition, comprising a video acquisition and processing module, which is used for arranging a surgery room camera array to acquire surgery process videos by region division, extracting video frame features and calculating an optical flow field through a residual neural network, fusing multiple video streams based on a feature alignment module guided by the optical flow, and constructing a surgery region panorama sequence; a target detection and recognition module, which is used for inputting the surgery region panorama sequence into a double-branch target detection network, extracting medical staff positions and patient body position data through a main branch, extracting surgery instrument features through an auxiliary branch by an attention mechanism, and constructing a surgery scene model containing personnel spatial positions and instrument use areas based on a multi-view mapping algorithm; a behavior tracking and analysis module, which is used for obtaining medical staff posture data by applying a skeleton key point extraction network based on personnel position sequences in the surgery scene model, performing trajectory tracking based on Kalman filtering in combination with the instrument use areas in the scene model, and constructing a node-weighted surgery process dynamic graph; a state evaluation generation module, which is used for inputting the surgery process dynamic graph into a time series graph convolution network, calculating a dynamic spatial correlation matrix through a multi-head attention mechanism, fusing a surgery type encoding vector, and generating a multi-dimensional surgery progress state report; and an abnormality monitoring and recording module, which is used for comparing the progress state report with a surgery specification library by using a knowledge distillation algorithm, dividing abnormal events into different levels according to violation degrees and safety risks, and recording surgery monitoring logs.
[0090] In summary, the present application successfully solves the technical problem of incomplete coverage of a single view angle in a traditional surgery room monitoring system by using a multi-camera array partition acquisition strategy, significantly improving the monitoring coverage rate of the surgery region. The double-branch network architecture and attention mechanism designed by the present application achieve precise and differentiated extraction of medical staff and surgery instrument features, effectively improving the accuracy of target detection. The innovative design of the skeleton key point extraction network and the three-dimensional convolution block realizes high-precision capture and analysis of the fine actions of medical staff. The developed time series graph convolution network and multi-head attention mechanism successfully achieve dynamic feature extraction and depth correlation analysis of the surgery process, and effectively convert expert experience into executable evaluation rules by constructing a hierarchical medical expert knowledge tree. The three-branch encoding network and spatio-temporal attention mechanism significantly improve the extraction ability of surgery specification features and the accuracy of template generation, and the innovative combination of the double-flow comparison network and the long short-term memory network realizes comprehensive capture of instantaneous and long-term abnormal patterns in the surgery process, greatly improving the sensitivity of abnormality detection.
[0091] Embodiment 2, refer to Fig. 1-3 For the second embodiment of the present application, the embodiment provides a surgery room intelligent monitoring method based on monitoring video recognition. In order to verify the beneficial effects of the present application, economic benefit calculation and simulation experiments are used for scientific demonstration.
[0092] To verify the effectiveness and practicability of the method, the research team conducted a simulation experiment in the general surgery operating room of a certain third-grade hospital for three months. The experiment environment was built by using a Hikvision DS-2CD6984F panoramic network camera to build an operating room monitoring array. Four cameras were arranged in the operating room to cover the operating table area, the anesthesia workstation area, the instrument table area and the overall environment. The camera acquisition resolution was set to 1920x1080 pixels, the frame rate was 30fps, and the video compression was performed using H.264 encoding format. The experimental data collection selected 100 cases of laparoscopic cholecystectomy surgery as the research object, including 80 cases of conventional surgery and 20 cases of complication surgery cases.
[0093] The experimental platform was based on NVIDIA Tesla V100 GPU for deep learning model training and inference, and the server was configured with Intel Xeon Gold 6248R CPU and 256GB RAM. The algorithm implementation used PyTorch1.9.0 framework and CUDA11.1 environment. In the video preprocessing stage, the spatial domain standard deviation of bilateral filtering was set to 75, the value domain standard deviation was set to 75, and the window size was 5x5 pixels. The adaptive histogram equalization was used for illumination normalization, the block size was set to 8x8 pixels, and the contrast limit threshold was set to 3.0.
[0094] The residual neural network used ResNet-50 as the basic architecture, and the pre-trained model was migrated on the ImageNet dataset. The initial learning rate was set to 0.001, the cosine annealing strategy was used for learning rate adjustment, and the training batch size was 32. The optical flow algorithm selected PWC-Net, and the multi-scale feature pyramid layer number was set to 5, and the minimum feature map size was 32x32 pixels. The spatial transformation network in the feature alignment module used bicubic interpolation for resampling, and the estimation of the transformation matrix was optimized using the iterative least squares method.
[0095] In the target detection stage, the backbone network of the dual-branch network used Feature Pyramid Network structure, and the anchor box scale ratio was set to [0.5, 1, 2], and the aspect ratio was set to [0.5, 1, 2]. The IoU threshold of non-maximum suppression was set to 0.5, and the detection confidence threshold was set to 0.7. The pose estimation subnetwork used HRNet-W32 architecture, and the key point detection heat map resolution was 64x64 pixels, and the Gaussian kernel standard deviation of the heat map generation was set to 2.0 pixels.
[0096] In order to objectively evaluate the performance advantages of the method, the research team compared three kinds of mainstream operating room monitoring methods, as shown in Table 1:
[0097] Table 1 Performance comparison of different operating room monitoring methods
[0098] Evaluation index Traditional single-view monitoring Multi-modal fusion method The method of the present application Medical staff detection accuracy (%) 78.5 89.7 94.2 Surgical instrument recognition accuracy (%) 72.3 86.5 91.8 Abnormal behavior detection rate (%) 65.8 82.9 88.6 False positive rate (%) 15.3 8.5 5.2 System delay (ms) 850 450 280 Scene coverage rate (%) 75.0 85.0 95.5
[0099] From the comparison data in Table 1, it can be seen that the method of the present application is significantly better than the prior art solution in various key performance indicators. In particular, in terms of detection accuracy of medical staff and identification accuracy of surgical instruments, it has reached a high precision level of 94.2% and 91.8% respectively, which is 15.7 and 19.5 percentage points higher than the traditional single-view monitoring solution. In terms of abnormal behavior detection, the detection rate of the method of the present application reaches 88.6%, and the false positive rate is only 5.2%, which reflects excellent monitoring reliability.
[0100] The experimental results show that, in terms of dynamic modeling of surgical procedures, the method of the present application successfully identifies the key step transition points in 93 out of 100 surgical cases, with a timing accuracy of 93.5%. In terms of abnormal event detection, the analysis of 20 cases of complications shows that the system detects potential risk signs an average of 42 seconds in advance, providing sufficient response time for medical staff. The knowledge distillation model performs well in normative evaluation, with an average absolute error of only 0.085 in the analysis of 1000 randomly sampled surgical operation segments, indicating that the system has high evaluation accuracy and stability.
[0101] After 3 months of experimental verification, the method of the present application has shown stable and reliable performance in the actual environment of the operating room. The system's average processing delay is kept within 280 milliseconds, meeting the needs of real-time monitoring. The multi-camera array arrangement scheme makes the system's coverage rate of the surgical area reach 95.5%, effectively solving the blind area problem of traditional single-view monitoring. In addition, the system performs well under different lighting conditions, with a detection performance fluctuation of no more than 3% in the morning, noon and evening, showing strong environmental adaptability.
[0102] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit it. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the present application, which should be covered by the scope of the claims of the present application.
Claims
1. An operating room intelligent monitoring method based on monitoring video recognition, characterized in that: The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area.
2. The operating room intelligent monitoring method based on monitoring video recognition of claim 1, wherein: The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method based on a panoramic sequence of a surgical area. The application relates to a surgical process state report generation method Input the RGB three-channel video data into a residual neural network, extract multi-scale features through multiple residual blocks, each residual block containing a convolution layer and a short circuit connection, and output a video feature vector; Adopt an optical flow algorithm to calculate the motion information between adjacent video frames, construct an optical flow map, and fuse the optical flow map with the video feature vector to obtain dynamic feature data; Construct a feature alignment module based on camera calibration parameters, spatially resample the dynamic feature data, and form perspective alignment features under each camera perspective; Cascade the perspective alignment features under each camera perspective, adopt a weighting strategy of channel attention and spatial attention for feature selection and fusion, reconstruct the image through a deconvolution network, and output a panoramic sequence of the surgical area.
3. The monitoring video recognition based operating room intelligent monitoring method of claim 1, wherein: The construction of a surgical scene model containing personnel spatial positions and instrument use areas based on a multi-perspective mapping algorithm includes the following steps: Divide the panoramic sequence of the surgical area into grid cells, extract target detection feature vectors, and construct input data for a feature pyramid network; Extract features from the input data through a main branch convolution layer, each convolution block containing a normalization layer and an activation function, and form a personnel target feature map; Input the personnel target feature map into a region proposal network, screen candidate regions through non-maximum suppression, and output medical staff position coordinates and patient body posture parameters; Perform channel attention weighting on the input data of the auxiliary branch, assign channel weight coefficients, highlight the features of surgical instruments, and form an instrument feature map; Perform region detection on the instrument feature map, extract local response values, and determine the use area and class label of the surgical instrument; Calculate the spatial mapping relationship based on the camera calibration parameters, convert the medical staff position coordinates, patient body posture parameters, and instrument use area to a unified coordinate system, and construct a surgical scene model.
4. The monitoring video recognition based operating room intelligent monitoring method of claim 1, wherein: The construction of a node-weighted surgical procedure dynamic graph includes the following steps: Extract the medical staff position sequence from the surgical scene model, construct a time series sliding window, time-align the medical staff position data in each sliding window, and construct a behavior analysis sample; Input the behavior analysis sample into a skeleton key point extraction network, extract human joint features through multiple layers of feature encoders, and form a posture skeleton feature map; Perform spatio-temporal convolution operation on the posture skeleton feature map, extract human posture change features, fuse node motion information through a self-attention mechanism, and output a posture feature sequence; Extract the surgical instrument use area and class label from the surgical scene model, apply the Kalman filter algorithm to predict the instrument position change, and obtain the instrument motion sequence; Establish a spatio-temporal correlation model between the posture feature sequence and the instrument motion sequence, calculate the personnel-instrument interaction intensity, and construct an initial procedure feature graph; Calculate the importance score of each node in the initial procedure feature graph, assign weight parameters based on the time series correlation and spatial dependence between nodes, and construct a node-weighted surgical procedure dynamic graph.
5. The monitoring video recognition based operating room intelligent monitoring method of claim 1, wherein: The comparison of the progress status report and the surgical specification library using the knowledge distillation algorithm includes the following steps: Decompose the multi-dimensional surgical progress status report into behavior sequence features, instrument operation features, and personnel interaction features, and generate a monitoring feature vector; Extract the standard operation specification corresponding to the operation type from the operation specification library, establish a teacher model through a knowledge distillation network, and output a specification feature template; Match the monitoring feature vector with the specification feature template, calculate the behavior deviation degree, and generate an anomaly detection matrix; Based on the anomaly detection matrix, a multi-level anomaly evaluation model is constructed, and according to the violation degree threshold and the safety risk coefficient, the abnormal events are classified, and the risk level data is output; Associate the risk level data with the system timestamp, and generate the abnormal event descriptor; According to the priority order of the abnormal level, the abnormal event descriptor is written into the monitoring database, and the operation monitoring log containing the timestamp and the abnormal level is generated.
6. The operating room intelligent monitoring method based on monitoring video recognition of claim 5, wherein: The generation of the anomaly detection matrix includes the following steps: Align the dimensions of the monitoring feature vector and the specification feature template, map the two groups of features to the same feature space, and obtain the uniformly represented feature pair; Process the feature pair through a double-flow network, calculate the similarity score of the feature pair through the main branch, and extract the difference features of the feature pair through the auxiliary branch to obtain the feature matching metric value; Establish a time sequence sliding window, sample the feature matching metric value, and extract the time-varying deviation features through a long short-term memory network; Calculate the behavior sequence deviation, instrument operation deviation and personnel interaction deviation of the time-varying deviation features to obtain the deviation feature group; Reorganize the deviation feature group into a three-dimensional tensor structure according to the time axis, space axis and feature axis to generate the anomaly detection matrix.
7. An operating room intelligent monitoring system based on monitoring video recognition, based on the operating room intelligent monitoring method based on monitoring video recognition in any of claims 1-6, characterized in that: Further comprising, A video acquisition and processing module for arranging a camera array in the operating room to collect operation process videos by region, extracting video frame features through a residual neural network and calculating an optical flow field, fusing multiple video streams based on an optical flow guided feature alignment module, and constructing a panoramic sequence of the operation area; A target detection and identification module for inputting the panoramic sequence of the operation area into a double-branch target detection network, extracting medical staff position and patient body position data through the main branch, and extracting surgical instrument features through the auxiliary branch by attention mechanism, and constructing a surgical scene model containing personnel spatial position and instrument use area based on a multi-view mapping algorithm; A behavior tracking analysis module for obtaining medical staff posture data based on the personnel position sequence in the surgical scene model by applying a skeleton key point extraction network, and combining the instrument use area in the scene model to perform trajectory tracking based on Kalman filtering, and constructing a node weighted surgical process dynamic graph; A state evaluation generation module for inputting the surgical process dynamic graph into a time series graph convolution network, calculating a dynamic spatial correlation matrix through a multi-head attention mechanism, and fusing a surgical type encoding vector to generate a multi-dimensional surgical progress state report; An abnormality monitoring and recording module for comparing the progress state report with the operation specification library using a knowledge distillation algorithm, classifying abnormal events into different levels according to the violation degree and safety risk, and recording the operation monitoring log.
Citation Information
Patent Citations
Single-view-angle three-dimensional human skeleton key point detection method and device, equipment and medium
CN115482481A
Video description generation method and system based on video space-time scene graph fusion reasoning
CN117370604A