Sterile operation detection method based on deep learning algorithm
Through deep learning algorithms and sparse reconstruction technology, movable targets in the sterile area of the operating room are monitored in real time, solving the real-time problem of intraoperative sterile operation detection, reducing the risk of postoperative infection, and improving detection accuracy and environmental adaptability.
Patent Information
- Application Number
- CN202510866004.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-26
AI Technical Summary
The existing technology lacks real-time performance and cannot effectively monitor and prevent violations of sterile operations by doctors during surgery, resulting in a high risk of postoperative infection.
The sterile operation detection method based on deep learning algorithm is adopted to identify the sterile area of the operating room through the camera device, a touch prediction model is constructed, and the motion trajectory of the movable target is detected in real time. Combined with sparse reconstruction algorithm and dictionary update technology, the recognition accuracy and environmental adaptability are improved.
Real-time detection of sterile operations is achieved, the risk of postoperative infection is reduced, and the monitoring accuracy of sterile operations and the ability to adapt to complex environments is improved.
Smart Images

Figure CN120375296A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image recognition, and in particular relates to a sterile operation detection method based on a deep learning algorithm. Background Art
[0002] The heterogeneity of infections complicates the epidemiological study of SSI: the incidence rates vary greatly among different surgeries, hospitals, surgeons, and patients. The increasing number of minimally invasive (laparoscopic) surgeries has led to a decrease in the incidence of SSI. For example, in patients undergoing cholecystectomy, the SSI rate after laparoscopic surgery is 1.1%, while the SSI rate after open surgery is 4%. Similarly, in patients with acute appendicitis, the SSI incidence of minimally invasive surgery is 2%, and that of open surgery is 8%. Moreover, a considerable number of intraoperative infections are caused by doctors or assistants inadvertently violating the aseptic principle. To avoid violating aseptic operations, hospitals often give preoperative reminders to doctors, but this measure lacks real-time performance, so intraoperative infections still occur. Summary of the Invention
[0003] In view of this, the present invention aims to propose a sterile operation detection method based on a deep learning algorithm, in order to solve at least one of the above partial technical problems.
[0004] To achieve the above object, the technical solution of the present invention is realized as follows: The first aspect of the present invention proposes a sterile operation detection method based on a deep learning algorithm, and the method includes: Continuously collecting images in the operating room using a camera device, identifying the sterile area in the operating room, and obtaining the segmentation mask of the movable target within the sterile area; Determining whether each pixel point of the image is a moving pixel through a preset threshold; Obtaining the foreground area through the segmentation mask, expanding the foreground area outward to obtain the surrounding area, calculating the proportion of moving pixels in the foreground area and the surrounding area, and generating a motion area ratio matrix; Constructing a touch prediction model, training the touch prediction model using the motion area ratio matrix, and completing real-time sterile operation detection through the trained touch prediction model.
[0005] Further, the recognition accuracy of the movable target is improved by constructing an overcomplete dictionary, and the construction process of the overcomplete dictionary is as follows: Taking the features of the movable target in the sterile area as the template dictionary, each template represents the features of the target in a certain state, and the features are the edges and optical flow patterns of the movable target; Sampling with a Gaussian distribution to generate a candidate target set, forming a dictionary matrix, and performing dimensionality reduction projection on the dictionary atoms using principal component analysis to reduce the computational complexity.
[0006] Further, a sparse reconstruction algorithm is used to improve the recognition ability of movable targets in different environments, and the recognition threshold of movable targets is reduced when conditions are met. The sparse reconstruction algorithm is as follows: Represent the target to be detected as a sparse linear combination of templates in the dictionary: , ; where z t is the target to be detected in the current frame, is the template dictionary, and c t is the sparse coefficient vector, obtained by minimizing the error. argmin means to find the c that minimizes the entire objective function, t , is the square of the L2 norm, λ is the regularization coefficient, is the L1 norm, that is, the sum of the absolute values of all elements.
[0007] Further, by defining the coherence index between dictionary atoms, the dictionary atoms with high correlation are screened to improve the efficiency of the sparse reconstruction algorithm.
[0008] Further, when the characteristics of the movable target change, the template dictionary is dynamically updated. The update process is as follows: Update the template dictionary every 10 frames. During the use of the sparse reconstruction algorithm, record the sparse coefficients of each frame, and screen the dictionary atoms with the smallest weight according to the sparse coefficients; Calculate the similarity between the target to be detected and the atoms in the dictionary: ; where z t is the target to be detected in the current frame, d gi is the atom in the dictionary, is the square of the L2 norm.
[0009] If the minimum similarity is lower than the preset threshold, trigger dictionary update, add the target to be detected to the dictionary to replace the dictionary atom with the smallest weight, and update it to a new template; otherwise, keep the dictionary unchanged.
[0010] Further, the process of determining whether each pixel point of the image is a moving pixel through the preset threshold is as follows: Introduce the motion sparsity S to exclude background noise or local optical flow jitter: S = number of moving pixels / total number of pixels in the neighborhood window; Introduce historical consistency H to avoid accidental jitter: ; Set a binary mask for each pixel point: , where P is a threshold judgment parameter, , L is the optical flow value, and L t is the optical flow value at time t, and L t-1 is the optical flow value at the previous moment. w1, w2, and w3 are weight parameters, σ is the determination threshold for moving and non-moving objects, and n is the sum of the number of pixel frames within a specified time.
[0011] Further, the foreground area gradually shrinks inward according to a certain pixel value: ; The surrounding area gradually expands outward according to a certain pixel value: ; Combining the images of the previous and next 10 frames, a motion area ratio matrix with p + q rows and 21 columns is obtained. Among them, a r is the proportion of moving pixels in the foreground area, and b r is the proportion of moving pixels in the surrounding area. a ij is the pixel point in the foreground area, and b ij is the pixel point in the surrounding area. m ij is a binary mask for determining whether a pixel point is a moving pixel.
[0012] Further, the touch prediction model receives the image, the motion area ratio matrix, and the segmentation mask, extracts features using independent convolutional encoders respectively, and performs feature fusion in the channel dimension; The fused features are processed by six layers of bidirectional LSTM to capture the dynamic information in the time series, and are classified and predicted through a two-layer multi-layer perceptron; The touch prediction model uses binary cross-entropy as the loss function and is optimized using the Adam optimizer.
[0013] The second aspect of the present invention proposes a server, including at least one processor, and a memory communicatively connected to the processor. The memory stores instructions executable by the at least one processor. When the instructions are executed by the processor, the at least one processor executes the method as described in the first aspect.
[0014] The third aspect of the present invention proposes a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the method as described in the first aspect is implemented.
[0015] Compared with the prior art, the aseptic operation detection method based on the deep learning algorithm of the present invention has the following beneficial effects: By dividing the aseptic area and the movable target and detecting the motion trajectory of the movable target, real-time aseptic operation detection is completed. Description of the Drawings
[0016] The accompanying drawings, which form a part of the present invention, are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings: Figure 1 It is a schematic diagram of the working process of the aseptic operation detection method described in the embodiment of the present invention. Detailed implementation manners
[0017] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0018] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise stated, the meaning of "a plurality" is two or more.
[0019] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "mounted", "connected", "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.
[0020] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0021] Embodiment 1: An aseptic operation detection method based on a deep learning algorithm, including: S1. Continuously collect images in the operating room using a camera device, identify the aseptic area in the operating room, and obtain the segmentation mask of the movable target in the aseptic area; S2. Determine whether each pixel point of the image is a moving pixel through a preset threshold; S3. Obtain the foreground region through the segmentation mask, expand the foreground region outward to obtain the surrounding region, calculate the proportion of moving pixels in the foreground region and the surrounding region, and generate a motion region ratio matrix; S4. Construct a touch prediction model, and use the motion region ratio matrix to train the touch prediction model, and complete real-time aseptic operation detection through the trained touch prediction model.
[0022] In step S1 of this embodiment, the imaging device includes six cameras, including two surgical area monitoring cameras and four operating room global monitoring cameras; The surgical area monitoring cameras are set on two surgical lights or on both sides of the main surgical light to achieve real-time monitoring of key areas such as the aseptic area of the operating table, the doctor and his hands; The global monitoring cameras are arranged at the four corners of the operating room to ensure full coverage of the dynamic changes in the operating room. These six cameras are connected to the computing terminal through a wireless network, enabling real-time data transmission and processing.
[0023] Step S1 adopts a region extraction algorithm based on the YOLOv10 convolutional neural network to identify aseptic areas such as the operating table, the aseptic area on the surface of the operating table, the doctor, the doctor's hand, and the doctor's chest area, as well as other objects; the strong tracking Kalman filtering algorithm is used to track movable objects (or people) to generate motion trajectories. For the detected movable objects, a segmentation model based on U-net is used to obtain the segmentation mask (SegmentationMask).
[0024] In order to improve the detection accuracy of dynamic targets, overcomplete dictionaries are constructed for features such as hands and instruments in the surgical aseptic area, and sparse representation methods are used to achieve high-precision reconstruction. In order to improve the recognition ability of the YOLO model in complex environments, this embodiment combines a sparse reconstruction algorithm, and for identifications that meet the sparse reconstruction conditions, the recognition threshold is lowered.
[0025] The recognition accuracy of movable targets is improved by constructing an overcomplete dictionary. The construction process of the overcomplete dictionary is as follows: The features of movable targets in the aseptic area are used as the template dictionary, and each template represents the features of the target in a certain state. The features are the edges and optical flow patterns of the movable targets; The candidate target set is generated by Gaussian distribution sampling to form a dictionary matrix, and principal component analysis is used to perform dimensionality reduction projection on the dictionary atoms to reduce the computational complexity.
[0026] The sparse reconstruction algorithm is used to improve the recognition ability of movable targets in different environments, and the recognition threshold for movable targets is reduced when the conditions are met. The sparse reconstruction algorithm is as follows: Represent the target to be detected as a sparse linear combination of templates in the dictionary: , ; where z t is the target to be detected in the current frame, is the template dictionary, c t is the sparse coefficient vector, obtained by minimizing the error, and argmin means to find the c that minimizes the entire objective function t , is the square of the L2 norm, λ is the regularization coefficient, is the L1 norm, i.e., the sum of the absolute values of all elements.
[0027] By defining the coherence index between dictionary atoms, select the dictionary atoms with high correlation to improve the efficiency of the sparse reconstruction algorithm.
[0028] For two atoms of the dictionary d i and d j , the coherence index is expressed as: ; Furthermore, perform coherence regularization, that is, minimize the coherence between dictionary atoms while optimizing the dictionary. The calculation process is: ; where λ is the weight hyperparameter of the coherence regularization term, used to control the balance between coherence and reconstruction error; x i is the coefficient representing the signal y i under the atom d j ; D is the new dictionary; use the subsequent coherence parameter and the previous reconstruction error to comprehensively obtain the minimum value.
[0029] When the characteristics of the moving target change, dynamically update the template dictionary. The update process is: Update the template dictionary every 10 frames. During the use of the sparse reconstruction algorithm, record the sparse coefficients of each frame, and select the dictionary atom with the smallest weight according to the sparse coefficients; among them, the sparse coefficient is: the coefficient of the recognized object mapped to the atom in the dictionary, specifically c t .
[0030] Calculate the similarity between the target to be detected and the atoms in the dictionary: ; where z t is the target to be detected in the current frame, d gi is the atom in the dictionary, is the square of the L2 norm.
[0031] If the minimum similarity is lower than the preset threshold, the dictionary is updated, the target to be detected is added to the dictionary to replace the dictionary atom with the smallest weight, and it is updated to a new template; otherwise, the dictionary remains unchanged.
[0032] The process of determining whether each pixel of the image is a moving pixel through the preset threshold is as follows: Introduce the motion sparsity S to exclude background noise or local optical flow jitter: S = number of moving pixels / total number of pixels in the neighborhood window; Introduce historical consistency H to avoid occasional jitter: ; Set a binary mask for each pixel to determine whether the pixel is stationary (0) or moving (1): , where P is the threshold judgment parameter, , L is the optical flow value, L t is the optical flow value at time t, L t-1 is the optical flow value at the previous moment, w1, w2, w3 are weight parameters, σ is the judgment threshold for moving and non-moving objects, and n is the sum of the number of pixel frames within a specified time.
[0033] The foreground area gradually shrinks inward according to a certain pixel value: ; The surrounding area gradually expands outward according to a certain pixel value: ; Combining the images of the previous and next 10 frames, a motion area ratio matrix with p + q rows and 21 columns is obtained, where a r is the proportion of moving pixels in the foreground area, b r is the proportion of moving pixels in the surrounding area, a ij is the pixel in the foreground area, b ij is the pixel in the surrounding area, m ij is the binary mask for determining whether the pixel is a moving pixel.
[0034] The touch prediction model receives the image, the motion area ratio matrix, and the segmentation mask, extracts features using independent convolutional encoders respectively, and performs feature fusion in the channel dimension; The fused features are processed by six layers of bidirectional LSTM to capture the dynamic information in the time series, and classification prediction is performed through two layers of multi-layer perceptrons; The touch prediction model uses binary cross-entropy as the loss function and is optimized using the Adam optimizer.
[0035] In this embodiment, the touch prediction model includes: Input layer: Directly captured RGB image: Provides appearance information of movable and immovable objects; Optical flow: Used to capture motion information of human hands and objects between adjacent frames; Foreground mask: A multi-channel binary mask used to indicate regions of different objects, including: mask of the sterile area, sterile area on the operating table surface, doctor, doctor's hands, and the doctor's chest area.
[0036] Data preprocessing: Perform union processing on all recognized objects, crop the image region, and encode the input data: The RGB image, optical flow information, and temporal sequence of the motion region are respectively input into independent encoder branches; RGB and optical flow encoders: Consist of 5 convolutional modules, and each convolutional module contains the following layers: 3×3 convolutional layer: Used to extract local features; ReLU activation layer: Introduces non-linearity to help the model learn complex features; LayerNorm layer: Normalizes the features of each layer to improve training stability; 2×2 Max-Pooling layer: Downsamples the feature map to reduce the computational amount; A dropout structure is set in the third and fifth convolutional modules: The foreground mask passes through a separate encoder to process the spatial relationship between the hand and the object; Foreground mask encoder: Consists of 3 convolutional layers, followed by a ReLU activation layer for each layer, and the output is used to encode the positional relationship between the sterile area, the hands of the operator, the surgical area, etc. and other objects.
[0037] It should be noted that each object is recognized separately, and the recognition results of each object are combined as the foreground mask; The said motion region ratio matrix only corresponds to each moving object, that is, the motion region ratio is not constructed for the foreground mask; The function of the said motion region ratio matrix is: to characterize the collision feature (during the collision process, the collided object will deform, and the collision feature is the deformation information).
[0038] Furthermore, the above sterile operation detection mainly identifies the collision between the sterile area and the contaminated area. Therefore, it is necessary to extract the foreground mask formed by the fusion of each region and extract the collision, and obtain the collision between the sterile and contaminated areas through the collision and the recognition fusion of each region.
[0039] Feature fusion: Concatenate (concat) the features extracted from the foreground mask and the motion region ratio matrix with the features of the RGB and optical flow encoders in the channel dimension.
[0040] LSTM layer: The fused features are processed by a 6-layer bidirectional LSTM (Long Short-Term Memory).
[0041] Multi-Layer Perceptron (MLP): The features output by the LSTM layer are fed into a 2-layer multi-layer perceptron, with two hidden layers. The first hidden layer has 300 parameters, the second hidden layer has 150 parameters, and there is a final output layer.
[0042] Loss function: Use binary cross-entropy loss as the loss function and use the adam optimizer for optimization.
[0043] Embodiment 2: A server includes at least one processor and a memory communicatively connected to the processor. The memory stores instructions executable by the at least one processor. When the instructions are executed by the processor, the at least one processor executes the method as described in Embodiment 1.
[0044] Embodiment 3: A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the method as described in Embodiment 1.
[0045] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the present embodiments, and they should all be covered by the scope of the claims and the description of the present invention.
[0046] The above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A sterile operation detection method based on a deep learning algorithm, characterized in that, The method includes: continuously collecting images in the operating room using a camera device, identifying the sterile area in the operating room, and obtaining a segmentation mask of the movable object within the sterile area; judging whether each pixel point of the image is a moving pixel through a preset threshold; obtaining a foreground area through the segmentation mask, expanding the foreground area outwards to obtain a surrounding area, calculating the proportion of moving pixels within the foreground area and the surrounding area, and generating a motion ratio matrix; constructing a touch prediction model based on the motion ratio matrix, training the touch prediction model using the collected motion area ratio matrix, and completing real-time sterile operation detection through the trained touch prediction model.
2. The aseptic operation detection method based on a deep learning algorithm according to claim 1, characterized in that Improving the recognition accuracy of the movable object by constructing an overcomplete dictionary. The construction process of the overcomplete dictionary is as follows: using the features of the movable object in the sterile area as a template dictionary, where each template represents the features of the object in a certain state, and the features are the edges and optical flow of the movable object; generating a candidate object set by Gaussian distribution sampling to form a dictionary matrix, and performing dimensionality reduction projection on the dictionary atoms using principal component analysis.
3. The aseptic operation detection method based on a deep learning algorithm according to claim 2, wherein Using a sparse reconstruction algorithm to improve the recognition ability of the movable object in different environments, and reducing the recognition threshold of the movable object when conditions are met. The sparse reconstruction algorithm is as follows: representing the object to be detected as a sparse linear combination of the templates in the dictionary: , ; Among them, z t is the target to be detected in the current frame, is the template dictionary, c t is the sparse coefficient vector, obtained by minimizing the error, and argmin means to find c that minimizes the entire objective function t , is the square of the L2 norm, and λ is the regularization coefficient, is the L1 norm, that is, the sum of the absolute values of all elements.
4. The sterile operation detection method based on a deep learning algorithm according to claim 3, characterized in that: by defining a coherence index between dictionary atoms, screening out dictionary atoms with high correlation to improve the efficiency of the sparse reconstruction algorithm.
5. The aseptic operation detection method based on a deep learning algorithm according to claim 3, characterized in that, When the features of the movable object change, dynamically update the template dictionary. The update process is as follows: updating the template dictionary every 10 frames. During the use of the sparse reconstruction algorithm, record the sparse coefficients of each frame, and screen out the dictionary atom with the smallest weight according to the sparse coefficients; calculating the similarity between the object to be detected and the atoms in the dictionary: ; where z t is the target to be detected in the current frame, d gi is the atom in the dictionary, is the square of the L2 norm; if the minimum similarity is lower than the preset threshold, trigger dictionary update, add the object to be detected to the dictionary to replace the dictionary atom with the smallest weight, and update it as a new template; otherwise, keep the dictionary unchanged.
6. The aseptic operation detection method based on the deep learning algorithm according to claim 1, wherein The process of judging whether each pixel point of the image is a moving pixel through a preset threshold is as follows: introducing motion sparsity S to exclude background noise or local optical flow jitter: S = number of moving pixel points / total number of pixel points within the neighborhood window; introducing historical consistency H to avoid accidental jitter: ; setting a binary mask for each pixel point: , where P is a threshold judgment parameter, , L is the optical flow value, and L t is the optical flow value at time t, and L t-1 is the optical flow value at the previous moment, w1, w2, and w3 are weight parameters, σ is the judgment threshold for moving and non-moving objects, and n is the sum of the number of pixel frames within a specified time.
7. The sterile operation detection method based on a deep learning algorithm according to claim 1, characterized in that: the foreground area gradually shrinks inwards according to a certain pixel value: ; the surrounding area gradually expands outwards according to a certain pixel value: ; Combine the images of the previous and next 10 frames to obtain a motion region ratio matrix with p + q rows and 21 columns, where a r is the proportion of moving pixels in the foreground region, and b r is the proportion of moving pixels in the surrounding region. a ij is the pixel point in the foreground region, and b ij is the pixel point in the surrounding region. m ij is a binary mask for determining whether a pixel point is a moving pixel.
8. The sterile operation detection method based on a deep learning algorithm according to claim 1, characterized in that: the touch prediction model receives the image, the motion area ratio matrix, and the segmentation mask, extracts features using independent convolutional encoders respectively, and performs feature fusion in the channel dimension; the fused features are processed by six layers of bidirectional LSTM to capture the dynamic information in the time series, and classification prediction is performed through two layers of multi-layer perceptron. The touch prediction model uses binary cross-entropy as the loss function and uses the Adam optimizer for parameter optimization.
9. A server, characterized in that: It includes at least one processor and a memory communicatively connected to the processor. The memory stores instructions executable by the at least one processor. When the instructions are executed by the processor, the at least one processor is caused to execute the method according to any one of claims 1-8.
10. A computer-readable storage medium stores a computer program, characterized in that: When the computer program is executed by a processor, it implements the method according to any one of claims 1-8.
Citation Information
Patent Citations
Personnel access detection method, device and equipment based on monitoring video
CN111310733A
Sterile surgery operation prompting method, system and device
CN115394415A
Moving target detection method based on visual brain network background modeling
CN116844234A
Intraoperative medical image operation method, system and equipment and storage medium
CN116974369A
Medical image intelligent processing method and system based on deep learning
CN119722620A
Cited By
Surgical sterile operation real-time monitoring method based on sensor
CN121354037A