Image processing system in orthopedic anesthesia operation based on optimized sedation management and area block
By combining data processing of state images and ultrasound images, the administration rate of sedative drugs is dynamically adjusted, which solves the problem of unstable efficacy of sedative drugs in regional block anesthesia, realizes real-time monitoring and optimized management of the patient's sedation status, and improves the accuracy and safety of sedation management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG WENDENG WHOLE BONE YANTAI HOSPITAL CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-21
AI Technical Summary
In existing regional block anesthesia techniques for orthopedic surgery, the efficacy of intraoperative sedative drugs is unstable, making it difficult to continuously cover the entire surgical process, and there are problems with patient discomfort or movement during the operation.
By combining intraoperative images of the patient and ultrasound images, and employing modules for data acquisition, image preprocessing, feature extraction, multimodal fusion, and visualization interaction, the sedation drug administration rate is dynamically adjusted to achieve real-time monitoring and optimized management of the patient's sedation status.
It improves the accuracy and robustness of sedation depth assessment, achieves second-level feature updates, meets online monitoring needs, ensures medication safety and precise control, and reduces the incidence of intraoperative adverse events.
Smart Images

Figure CN121904002A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of surgical assistance technology, specifically to an image processing system for orthopedic anesthesia based on optimized sedation management and regional blockade. Background Technology
[0002] Regional block involves injecting local anesthetic directly into the peripheral nerves or nerve plexuses to block nerve conduction, achieving painlessness during surgery and postoperative analgesia. Compared to general anesthesia, regional block has the following advantages:
[0003] Reduce systemic drug dosage and decrease the risk of respiratory depression and nausea and vomiting;
[0004] Keep the patient awake or mildly sedated during the operation to facilitate neurological function monitoring;
[0005] It significantly prolongs the postoperative analgesia time and promotes rapid recovery.
[0006] CN108784836A discloses an image processing system for orthopedic anesthesia based on optimized sedation management and regional blockade. The image processing system includes: an ultrasound image acquisition module, an operation status detection module, a main control module, an image processing module, a 3D navigation module, a storage module, and a display module. The image processing module eliminates the need for manual lesion identification by the surgeon, improving image recognition efficiency and reducing labor costs by eliminating the need for surgeons to possess strong lesion identification skills. Simultaneously, the 3D navigation module overlays virtual images of the surgical site's bones and the specific 3D images of the surgical instruments' positions within the patient's body onto the surgeon's real field of vision, providing real-time navigation during surgery. This significantly accelerates the surgeon's operation, enabling safer, more accurate, and more efficient completion of the surgery.
[0007] The aforementioned patents identify the location of the patient's lesion before surgery, enabling the surgeon to complete the operation accurately, efficiently, and quickly. Preoperative sedation can alleviate the patient's anxiety and fear when entering the operating room and reduce hemodynamic fluctuations during the induction period. However, preoperative sedation has certain limitations, as most short-acting drugs have a half-life of only tens of minutes, making it difficult to cover the entire operation.
[0008] During the surgery, as the intensity of operations such as incision, traction, and grinding changes, the patient may experience discomfort or movement. Even if the regional blockade is sufficient, there may be occasional "blockade blind spots" or a decrease in drug efficacy, requiring intraoperative sedation / analgesia. Therefore, maintaining moderate sedation (OAA / S≈2–3 or BIS≈60–80) can ensure patient cooperation and prevent mismovement and memory loss.
[0009] Therefore, this application proposes an image processing system that optimizes sedation management for patients during regional nerve block surgery by combining images of the patient's intraoperative status and ultrasound images. Summary of the Invention
[0010] One of the objectives of this invention is to provide an intraoperative image processing system for orthopedic anesthesia based on optimized sedation management and regional blockade, which combines intraoperative patient status images and ultrasound images to determine changes in the patient's sedation status during surgery, thereby assisting in the management of the patient's sedation.
[0011] To achieve the above objectives, the present invention provides the following technical solution: an image processing system for orthopedic anesthesia based on optimized sedation management and regional blockade, comprising:
[0012] The data acquisition module acquires intraoperative patient status images and ultrasound images, connects the status image and ultrasound image data sources to the PTP network, and aligns the timing errors between the status image and ultrasound image.
[0013] The image processing module preprocesses the patient status image and ultrasound image, suppresses image noise, and enhances the contrast and pseudo-color of the status image and ultrasound image.
[0014] The feature extraction module extracts state features from the state image and ultrasound features from the ultrasound image, and constructs a state feature set and an ultrasound feature set. When determining the patient's sedation state, the two feature sets are used to verify each other.
[0015] The multimodal fusion model divides the state feature set and ultrasound feature set into multiple time windows, constructs a time-series vector, maps each image feature to a unified dimension through a feature encoder, dynamically integrates different modalities, and outputs the fused features to the model to determine the patient's sedation status.
[0016] The sedation feedback module determines the dosing rate per second based on the patient's sedation status and the implemented sedation protocol.
[0017] The visualization and interaction module displays patient status characteristics and ultrasound features through images and can be used to control the status of the drug delivery pump.
[0018] In one or more embodiments of the present invention, the state images include: facial and eye videos, pupil dynamic imaging, and thermal imaging, and the ultrasound images include: cardiac ultrasound, vascular Doppler, transcranial Doppler, lung ultrasound, and airway ultrasound.
[0019] In one or more embodiments of the present invention, the step of time-series alignment of the state image data source and the ultrasound image data source through the PTP network is as follows:
[0020] Select a highly stable OCXO as the PTP Grandmaster, deploy two levels of Boundary Clock switches, and divide the image data source and PTP control flow into different VLANs and configure 802.1p priority respectively;
[0021] The Grandmaster and Boundary Clock switches are interconnected via a dual-rack redundant link, and the image data VLAN and PTP control VLAN are connected on the link. The Ethernet interface of each acquisition device is connected to the corresponding switch port, and 802.1Q tagging is enabled.
[0022] Enable IEEE 1588 v2 hardware timestamps on each data source device Frame Grabber, set the master and slave devices (master device PtpSlaveOnly=false, slave device PtpSlaveOnly=true), and activate Chunk Timestamp output in the vendor configuration tool.
[0023] On the switch, check the master offset and mean path delay metrics to ensure that the error is within the sub-microsecond to microsecond range. On the host side, you can use ptp4l or device logs to track the latency jitter of each device and check the alignment of the ChunkTimestamp with the global clock.
[0024] In one or more embodiments of the present invention, the image preprocessing steps are as follows:
[0025] The status image and ultrasound image are denoised in parallel. The facial / eye video, pupil dynamics and thermal images are filtered in the spatiotemporal domain to remove sensor noise and suppress patient motion artifacts. The ultrasound image is denoised by combining anisotropic diffusion filtering and wavelet domain thresholding to suppress speckle noise while preserving the edge details of nerve and blood vessel tissue.
[0026] Contrast enhancement was performed on both types of images. The state image was enhanced by applying CLAHE adaptive histogram equalization to highlight facial muscle texture, pupil edges, and thermal spot temperature gradients. The ultrasound image was enhanced by multi-scale Retinex or gradient domain enhancement algorithms to improve the grayscale differences of low-contrast tissue structures.
[0027] Grayscale or thermal images are mapped to pseudocolor, thermal images are mapped to temperature gradients, and facial thermal images can intuitively distinguish between sympathetic and parasympathetic activity; ultrasound echo intensity is displayed through grayscale pseudocolor, with different echo amplitudes corresponding to different hues, enhancing the visualization of nerve bundles, blood vessels, and drug diffusion areas.
[0028] All image frames are subjected to linear or logarithmic gamma correction, and smoothing and frame interpolation are performed in time. The enhanced state features are displayed simultaneously with the ultrasound segmentation / pseudocolor layer.
[0029] In one or more embodiments of the present invention, state features are extracted from the state image to construct a state feature set:
[0030] Automatically detect and align facial poses in each state frame image, crop out the ROIs of eyes, forehead, and nose and lips, and uniformly scale each ROI to a fixed resolution, while linearly normalizing pixel or temperature values.
[0031] Using a preset duration as a sliding window, the preprocessed frame sequence is sampled at equal intervals, with each time window containing the same number of frames. The balance between continuity and responsiveness is achieved through window overlap.
[0032] Within each time window, facial expression action unit intensity, eyelid opening and closing index and blink rate are extracted for the facial ROI. The pupil diameter and contraction / dilution speed are obtained using segmentation and tracking algorithms. Local temperature difference and change rate are calculated from thermal imaging to form multi-dimensional visual and dynamic features.
[0033] For each extracted indicator, calculate the mean, variance, extreme values, and slope of change statistics, and concatenate the statistics into a high-dimensional feature vector in a predetermined order;
[0034] The feature vectors corresponding to each time window are arranged into an n×d feature matrix in time sequence. The matrix columns are normalized using Z-score and then stored in a circular buffer so that the multimodal fusion and real-time classification modules can call them at any time.
[0035] In one or more embodiments of the present invention, ultrasound features are extracted from ultrasound images to construct an ultrasound feature set:
[0036] Based on clinical annotations or deep learning models, the system automatically locates and segments the tissue or lesion region to be analyzed. For dynamic ultrasound that requires frame-by-frame tracking, it combines optical flow or block-matching-based motion estimation algorithms to achieve continuous updates of the ROI.
[0037] Within each ROI, a grayscale histogram and its statistics, including mean, variance, skewness, and kurtosis, are calculated to capture the overall echo intensity distribution characteristics, providing basic indicators for subsequent texture and morphology analysis.
[0038] Contrast, energy, correlation, and entropy indices are extracted using gray-level co-occurrence matrix, gray-level running length matrix, or local binary mode to reveal the internal microstructure and speckle distribution characteristics of tissues.
[0039] Based on ROI binarization and contour detection, the area, perimeter, compactness, and aspect ratio parameters of the target region are measured. At the same time, the edge gradient, curvature, and edge density are calculated to reflect the smoothness and complexity of the lesion boundary.
[0040] Discrete wavelet transform or Fourier transform is applied to ROI images to extract energy distribution and spectral entropy in different frequency bands, capture multi-scale details of ultrasonic echoes, and use subband energy or spectral centroid as additional features.
[0041] For multi-frame or Doppler ultrasound sequences, calculate flow velocity distribution, blood flow time, intensity curve parameters, tissue motion velocity and strain rate indices to provide temporal information for assessing vascular status and tissue stiffness;
[0042] The ultrasound features are concatenated into a high-dimensional vector in a predetermined order. Each dimension is standardized by Z-score and stored in the feature matrix to obtain an ultrasound feature set consisting of the number of samples × the feature dimension.
[0043] In one or more embodiments of the present invention, a temporal vector is constructed and mapped to a unified dimension via a feature encoder:
[0044] Under a sliding window with fixed duration and step size, the state features and ultrasound features are aligned and sampled, then segmented separately. Each window contains the same number of time points and outputs a shape of [M, L, D]. state ] and [M, L, D ultra Preliminary temporal fragments;
[0045] The state features are standardized with zero mean and unit variance, and missing values are retained and filled using linear interpolation or previous values; the original texture or depth features of the ultrasound image are extracted first, the spatial resolution is unified, and the channels are normalized and linearly mapped to a new interval according to fixed rules.
[0046] Lightweight encoders were designed for two modalities: a Medium Level Logic (MLP) with hidden layers was used for state features, and a small CNN was used for ultrasound features, mapping their respective inputs to the same embedding dimension D. embed Based on modal differences, select independent or parameter-shared encoders, and add Norm and Dropout to each layer to improve stability;
[0047] The encoder outputs two sets of equal-length sequences E state =[e1,…,e M ] and E ultra =[u1,…,u M (Each vector has a dimension of D) embed Then, a fusion vector s is generated by simple concatenation or weighted linear fusion. i Finally, a multimodal time series sequence S=[s1,…,s2] is obtained for use in subsequent time series networks. M];
[0048] In selecting hyperparameters, a balance needs to be struck between the receptive field and computational cost—a longer window length L can capture slower changes but reduces real-time performance, while a higher overlap rate results in more samples.
[0049] In one or more embodiments of the present invention, different modalities are dynamically integrated, and the fused features are output as a model:
[0050] At each time step t, take the state embedding e. t and ultrasound embedded u t Construct a joint query Q=[e t ;u t The attention score α is calculated using multi-head self-attention, along with key-value pairs K and V (both [e; u] or derived from a specialized linear mapping). (m) t. Each head m corresponds to a set of learnable modal weights, ultimately generating a weighted fusion vector s. t ;
[0051] Let the sequence {s1,…,s} M The input encoder captures short-term fluctuations and long-term dependencies, and the output context embedding h is used. t Residual connections are added between layers to maintain information flow and stabilize training with LayerNorm;
[0052] embedding h in context t Apply a two-branch head:
[0053] Discrete grading: Fully connected → Softmax predicts four levels of sedation (awake / lightly sedated / moderate / deeply sedated), using cross-entropy as the loss;
[0054] Continuous exponent: Fully connected → linear output sedation exponent, truncated with upper and lower bounds and with mean squared error as loss, multi-task joint training (weighted cross-entropy + MSE), select the maximum value of Softmax or regression value to map to state label for inference.
[0055] In one or more embodiments of the present invention, the dosing rate per second is determined based on the patient's sedation status and the implemented sedation protocol:
[0056] The system reads the patient's current sedation indicators and preset targets in real time, while retaining the drug administration rate from the previous moment. Users need to preset the PID control gain and the upper and lower limits of the drug administration rate, and configure the maximum rate variation range.
[0057] The error is calculated, the rate increment is calculated using the PID formula, and finally the new rate is clipped with upper and lower limits and smoothed to generate the final dosing rate, which is then sent to the infusion pump.
[0058] When the absolute value of the error exceeds a predefined threshold, an alarm is triggered, and the PID parameters are adaptively tuned periodically or as needed to cope with the dynamic changes in the patient's physiological state.
[0059] In one or more embodiments of the present invention, patient condition characteristics and ultrasound characteristics are visualized:
[0060] Subscribe to patient status features and ultrasound feature interfaces from the backend multimodal fusion service, and transmit them uniformly in JSON or Protocol Buffer format;
[0061] Define the main view area and side panel area on the front end. The main view is used to display time series curves and heat maps.
[0062] Facial thermal imaging, pupil diameter, and facial expression unit state features are mapped to multiple polylines or heat map bars, with different indicators highlighted by color and line width.
[0063] Real-time rendering of temporal variation curves of ultrasonic texture feature intensity or subband energy, with the ability to switch between frequency domain and time domain views;
[0064] Achieve linkage between status feature and ultrasound feature charts: After selecting a time interval on any curve or image, the corresponding interval in another view will be highlighted simultaneously.
[0065] Through the above technical solution, the present invention has the following beneficial effects:
[0066] 1. This application integrates preprocessing, feature extraction, multimodal fusion, visualization, and closed-loop drug delivery control of state images and ultrasound images, and integrates multi-source information such as facial expressions, pupil dynamics, thermal imaging, and ultrasound texture to improve the accuracy and robustness of sedation depth assessment.
[0067] 2. By employing a sliding window and efficient coding mapping, second-level feature updates and fusion are achieved to meet the needs of online monitoring. Through algorithms such as spatiotemporal filtering, wavelet threshold denoising, and CLAHE contrast enhancement, noise is effectively suppressed while preserving key tissue and physiological details.
[0068] 3. Based on PID or model predictive control algorithms, the dosing rate is automatically calculated and issued, and manual fine-tuning and operation logs are recorded to ensure medication safety and precise control.
[0069] 4. Real-time display of multi-channel time-series curves, pseudo-color heatmaps, and comparison with original frames, supporting mouse hover, region zoom, and threshold alerts to facilitate clinical decision-making.
[0070] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings. Attached Figure Description
[0071] Figure 1 This is a schematic diagram of the image processing system of the present invention. Detailed Implementation
[0072] The following describes several embodiments of the present invention with reference to the accompanying drawings. For clarity, many practical details will be described in the following description. However, it should be understood that these practical details are not intended to limit the invention. That is, in some embodiments of the invention, these practical details are not essential. And features of different embodiments may be interchanged if feasible.
[0073] Unless otherwise defined, all terms used herein (including technical and scientific terms) have their ordinary meanings, which are understandable to those skilled in the art. Furthermore, the definitions of the foregoing terms in commonly used dictionaries should be interpreted in the context of this specification as having the meaning consistent with the relevant field of this invention. Unless specifically defined, these terms will not be construed as having idealized or overly formal meanings.
[0074] The following explains the relationships and terms used in this application:
[0075] Sedation management refers to the entire process of achieving a predetermined depth of sedation in patients during surgery or treatment by selecting, administering, and monitoring sedative drugs, while maintaining stable vital signs, safety, and comfort.
[0076] Patient's response to verbal stimuli, pain stimuli, ocular manifestations, facial expressions, changes in vital signs, and EEG / BIS characteristics under different sedation states:
[0077] Awake: Can open eyes independently and accurately execute simple commands; blinks frequently with normal pupil diameter; expressive, can smile or nod; normal respiration, heart rate, and blood pressure; high frequency, low amplitude; BIS 90–100;
[0078] Mild sedation: eyes open and can answer or nod when called by name; eyes open when tapped on the shoulder; no obvious avoidance of painful stimuli; blinking frequency slightly decreased, pupillary reflex normal; facial expression relatively calm, glabellar muscles relaxed; breathing slightly slowed, blood pressure slightly decreased; mild theta wave increase; BIS 75–90;
[0079] Moderate sedation: eyes open when called loudly or repeatedly; head turns or hands are raised when lightly touched; purposeful avoidance of needle pricks or pressure; low blinking frequency, occasional long eye closures; decreased facial tension, monotonous expression; further decrease in respiratory rate, small fluctuations in blood pressure; predominant theta waves; BIS 60–75.
[0080] Deep sedation: no response to language, only brief eye opening in response to shaking or bright light stimuli; brief reflexes in response to bright light or shaking; primitive avoidance movements in response to strong stimuli (needle prick); prolonged eye closure, pupil constriction; expressionless face, smooth brow; significant respiratory depression, requiring airway support; increased delta waves; BIS 40–60.
[0081] General anesthesia: No verbal or visual response; no response to light touch or painful stimuli; eyes completely closed, pupils not reflecting light; completely expressionless; loss of spontaneous ventilation, large circulatory fluctuations, requiring mechanical ventilation support; EEG inhibition; BIS <40.
[0082] The above explanation does not fully encompass the relationship definition given in this application, but only represents a part of it.
[0083] This invention provides an intraoperative image processing system for orthopedic anesthesia based on optimized sedation management and regional blockade, which acquires intraoperative images of the patient's status and ultrasound images, and determines changes in the patient's sedation status after processing the status images and ultrasound images.
[0084] The image processing system includes:
[0085] The data acquisition module acquires intraoperative patient status images and ultrasound images, connects the status image and ultrasound image data sources to the PTP network, and aligns the timing errors between the status image and ultrasound image.
[0086] The image processing module preprocesses the patient status image and ultrasound image, suppresses image noise, and enhances the contrast and pseudo-color of the status image and ultrasound image.
[0087] The feature extraction module extracts state features from the state image and ultrasound features from the ultrasound image, and constructs a state feature set and an ultrasound feature set. When determining the patient's sedation state, the two feature sets are used to verify each other.
[0088] The multimodal fusion model divides the state feature set and ultrasound feature set into multiple time windows, constructs a time-series vector, maps each image feature to a unified dimension through a feature encoder, dynamically integrates different modalities, and outputs the fused features to the model to determine the patient's sedation status.
[0089] The sedation feedback module determines the dosing rate per second based on the patient's sedation status and the implemented sedation protocol.
[0090] The visualization and interaction module displays patient status characteristics and ultrasound features through images and can be used to control the status of the drug delivery pump.
[0091] In the above implementation scheme, since different sedation states will be reflected in the patient's physical response, by determining the patient's status images and ultrasound images during the operation, it is possible to reflect the changes in the patient's sedation state, thereby analyzing what kind of sedation state the patient is in (light sedation, moderate sedation, deep sedation or anesthesia level), so that the sedation plan can be adjusted according to the patient's response to ensure the stability of sedation during the operation.
[0092] In one embodiment, the status images include: facial and eye video, pupil dynamic imaging, and thermal imaging, and the ultrasound images include: cardiac ultrasound, vascular Doppler, transcranial Doppler, lung ultrasound, and airway ultrasound.
[0093] Among them, facial and eye videos are captured by a fixed camera to capture micro-expressions, eyelid opening and closing, and blinking frequency changes;
[0094] Dynamic pupil imaging measures pupil diameter and light reflection using an infrared or femtosecond camera.
[0095] Thermal imaging uses infrared thermal imagers to monitor changes in the dimensionality of the skin on the face and / or extremities, indirectly reflecting sympathetic / parasympathetic activity.
[0096] Echocardiography reflects changes in left ventricular function and hemodynamics in real time, helping to determine whether the depth is too deep (reduced cardiac output) or too shallow (sudden increase in heart rate / blood pressure).
[0097] Vascular Doppler monitoring of fluctuations in peripheral blood perfusion and vascular resistance can alert to sympathetic / parasympathetic imbalance.
[0098] Transcranial Doppler ultrasound can assess changes in cerebral blood flow and help prevent the risk of cognitive impairment or postoperative delirium.
[0099] Lung ultrasound provides real-time warnings of respiratory depression, atelectasis, and diaphragmatic fatigue, helping to prevent hypoventilation or pulmonary complications.
[0100] Airway Doppler assessment of airway patency and intubation safety.
[0101] By fusing multiple sensors, including facial / eye video, dynamic pupil imaging, thermal imaging, and transcranial Doppler ultrasound, the system can assess sedation depth in real time and drive a closed-loop drug delivery system to dynamically adjust the dosage of sedative drugs, thereby assisting in sedation management, implementing different sedation strategies for different patients, and reducing the incidence of intraoperative adverse events.
[0102] In one embodiment, the steps for time-series alignment of the state image data source and the ultrasound image data source when accessing the PTP network are as follows:
[0103] Select a highly stable OCXO as the PTP Grandmaster, deploy two levels of Boundary Clock switches, and divide the image data source and PTP control flow into different VLANs and configure 802.1p priority respectively;
[0104] The Grandmaster and Boundary Clock switches are interconnected via a dual-rack redundant link, and the image data VLAN and PTP control VLAN are connected on the link. The Ethernet interface of each acquisition device is connected to the corresponding switch port, and 802.1Q tagging is enabled to ensure that PTP packets and video data streams go their separate ways.
[0105] Enable IEEE 1588 v2 hardware timestamps on each data source device Frame Grabber, set the master and slave devices (master device PtpSlaveOnly=false, slave device PtpSlaveOnly=true), and activate Chunk Timestamp output in the vendor configuration tool.
[0106] On the switch, check the master offset and mean path delay metrics to ensure that the error is within the sub-microsecond to microsecond range. On the host side, you can use ptp4l or device logs to track the latency jitter of each device and check the alignment of the ChunkTimestamp with the global clock.
[0107] In the above implementation scheme, the data acquisition module reads image frames across VLANs according to the hardware timestamps of each device, aligns the facial / eye, pupil, thermal imaging and ultrasound data according to the global timestamp, and determines the patient's sedation state by retrieving the changes in the corresponding images of different devices within the same time period when classifying the patient's sedation state.
[0108] In one embodiment, the image preprocessing steps are as follows:
[0109] The status image and ultrasound image are denoised in parallel. The facial / eye video, pupil dynamics and thermal images are filtered in the spatiotemporal domain to remove sensor noise and suppress patient motion artifacts. The ultrasound image is denoised by combining anisotropic diffusion filtering and wavelet domain thresholding to suppress speckle noise while preserving the edge details of nerve and blood vessel tissue.
[0110] Contrast enhancement was performed on both types of images. The state image was enhanced by applying CLAHE adaptive histogram equalization to highlight facial muscle texture, pupil edges, and thermal spot temperature gradients. The ultrasound image was enhanced by multi-scale Retinex or gradient domain enhancement algorithms to improve the grayscale differences of low-contrast tissue structures.
[0111] Grayscale or thermal images are mapped to pseudocolor, thermal images are mapped to temperature gradients, and facial thermal images can intuitively distinguish between sympathetic and parasympathetic activity; ultrasound echo intensity is displayed through grayscale pseudocolor, with different echo amplitudes corresponding to different hues, enhancing the visualization of nerve bundles, blood vessels, and drug diffusion areas.
[0112] All image frames are subjected to linear or logarithmic gamma correction, and smoothing and frame interpolation are performed in time. The enhanced state features are displayed simultaneously with the ultrasound segmentation / pseudocolor layer.
[0113] In the above implementation scheme, image details are preserved to the greatest extent while noise is effectively suppressed, and contrast and visualization effects are improved, so as to provide accurate and stable image input for subsequent multimodal fusion analysis and calming assessment.
[0114] In one embodiment, state features are extracted from the state image to construct a state feature set:
[0115] Automatically detect and align facial poses in each state frame image, crop out the ROIs of eyes, forehead, and nose and lips, and uniformly scale each ROI to a fixed resolution, while linearly normalizing pixel or temperature values.
[0116] Using a preset duration as a sliding window, the preprocessed frame sequence is sampled at equal intervals, with each time window containing the same number of frames. The balance between continuity and responsiveness is achieved through window overlap.
[0117] Within each time window, facial expression action unit intensity, eyelid opening and closing index and blink rate are extracted for the facial ROI. The pupil diameter and contraction / dilution speed are obtained using segmentation and tracking algorithms. Local temperature difference and change rate are calculated from thermal imaging to form multi-dimensional visual and dynamic features.
[0118] For each extracted indicator, calculate the mean, variance, extreme values, and slope of change statistics, and concatenate the statistics into a high-dimensional feature vector in a predetermined order;
[0119] The feature vectors corresponding to each time window are arranged into an n×d feature matrix in time sequence. The matrix columns are normalized using Z-score and then stored in a circular buffer so that the multimodal fusion and real-time classification modules can call them at any time.
[0120] In this implementation scheme, n is the number of time windows and d is the vector dimension. The multidimensional state features of facial muscle micro-movements, eyelid movements, pupil dynamics, and facial thermal distribution are transformed into a structured feature set, which can provide stable and accurate data input for subsequent judgment of the patient's sedation status.
[0121] In one embodiment, ultrasound features are extracted from ultrasound images to construct an ultrasound feature set:
[0122] Based on clinical annotations or deep learning models, the system automatically locates and segments the tissue or lesion region to be analyzed. For dynamic ultrasound that requires frame-by-frame tracking, it combines optical flow or block-matching-based motion estimation algorithms to achieve continuous updates of the ROI.
[0123] Within each ROI, a grayscale histogram and its statistics, including mean, variance, skewness, and kurtosis, are calculated to capture the overall echo intensity distribution characteristics, providing basic indicators for subsequent texture and morphology analysis.
[0124] Contrast, energy, correlation, and entropy indices are extracted using gray-level co-occurrence matrix, gray-level running length matrix, or local binary mode to reveal the internal microstructure and speckle distribution characteristics of tissues.
[0125] Based on ROI binarization and contour detection, the area, perimeter, compactness, and aspect ratio parameters of the target region are measured. At the same time, the edge gradient, curvature, and edge density are calculated to reflect the smoothness and complexity of the lesion boundary.
[0126] Discrete wavelet transform or Fourier transform is applied to ROI images to extract energy distribution and spectral entropy in different frequency bands, capture multi-scale details of ultrasonic echoes, and use subband energy or spectral centroid as additional features.
[0127] For multi-frame or Doppler ultrasound sequences, calculate flow velocity distribution, blood flow time, intensity curve parameters, tissue motion velocity and strain rate indices to provide temporal information for assessing vascular status and tissue stiffness;
[0128] The ultrasound features are concatenated into a high-dimensional vector in a predetermined order. Each dimension is standardized by Z-score and stored in the feature matrix to obtain an ultrasound feature set consisting of the number of samples × the feature dimension.
[0129] In this implementation, grayscale normalization is performed on the filtered image to map pixel values to a uniform range, ensuring that images from different devices and acquisition parameters have comparability and stable input. For dynamic ultrasound sequences, optical flow or block matching motion estimation algorithms can be added to achieve ROI tracking between consecutive frames. Accurate ROI segmentation can focus on key structures and avoid the impact of background interference on feature extraction.
[0130] Based on ROI binarization and contour detection, the morphological parameters of the target region are calculated. By using edge detection algorithms such as first-order gradient or second-order curvature, the smoothness and complexity features of the boundary are extracted. Morphological and edge indicators can reflect the geometric shape information of tissue or lesion, which is particularly important for identifying the boundary of masses or irregular lesions.
[0131] In one embodiment, a temporal vector is constructed and mapped to a uniform dimension via a feature encoder:
[0132] Under a sliding window with fixed duration and step size, the state features and ultrasound features are aligned and sampled, then segmented separately. Each window contains the same number of time points and outputs a shape of [M, L, D]. state ] and [M, L, D ultra Preliminary temporal fragments;
[0133] The state features are standardized with zero mean and unit variance, and missing values are retained and filled using linear interpolation or previous values; the original texture or depth features of the ultrasound image are extracted first, the spatial resolution is unified, and the channels are normalized and linearly mapped to a new interval according to fixed rules.
[0134] Lightweight encoders were designed for two modalities: a Medium Level Logic (MLP) with hidden layers was used for state features, and a small CNN was used for ultrasound features, mapping their respective inputs to the same embedding dimension D. embed Based on modal differences, select independent or parameter-shared encoders, and add Norm and Dropout to each layer to improve stability;
[0135] The encoder outputs two sets of equal-length sequences E state =[e1,…,e M ] and E ultra =[u1,…,u M (Each vector has a dimension of D) embed Then, a fusion vector s is generated by simple concatenation or weighted linear fusion. i Finally, a multimodal time series sequence S=[s1,…,s2] is obtained for use in subsequent time series networks. M ].
[0136] In the selection of hyperparameters, a balance needs to be struck between receptive field and computational cost. A longer window length L can capture slower changes but reduces real-time performance, and a higher overlap rate means more samples. The depth and hidden width of the encoding network also need to be balanced between representational capability and computational cost, while the expansion of splicing dimensions can be controlled by linear dimensionality reduction.
[0137] In this implementation, the original segmented feature shape is [M, L, D]. state ] and [M, L, D ultra The entire process is divided into M time windows. Within each time window, L frames are extracted at the original sampling rate. Each frame is represented by D... state / D ultra A dimensional state / ultrasonic eigenvector representation.
[0138] Among them, the small CNN kernel size is 3*3, the number of layers is 3-5, and the output channels are D. embed The width of the MLP hidden layer is 2*D state Output D embed .
[0139] Normalization and Dropout are two complementary techniques. Normalization focuses on stabilizing training and accelerating convergence, while Dropout focuses on regularization and preventing overfitting.
[0140] In one embodiment, different modalities are dynamically integrated, and the fused features are output as the model:
[0141] At each time step t, take the state embedding e. t and ultrasound embedded u t Construct a joint query Q=[e t ;u t The attention score α is calculated using multi-head self-attention, along with key-value pairs K and V (both [e; u] or derived from a specialized linear mapping). (m) t. Each head m corresponds to a set of learnable modal weights, ultimately generating a weighted fusion vector s. t K (key) is obtained by linear transformation of the input features and is used to perform a dot product with the query vector Q to calculate the attention weights and determine "which positions I should pay attention to". V (value) is obtained by another linear transformation of the same input features x and carries the actual information of that position. Finally, the values are summed according to the attention weights to form the output.
[0142] Let the sequence {s1,…,s} M The input encoder captures short-term fluctuations and long-term dependencies, and the output context embedding h is used. t Residual connections are added between layers to maintain information flow and stabilize training with LayerNorm;
[0143] embedding h in context t Apply a two-branch head:
[0144] Discrete grading: Fully connected → Softmax predicts four levels of sedation (awake / lightly sedated / moderate / deeply sedated), using cross-entropy as the loss;
[0145] Continuous exponent: Fully connected → linear output sedation exponent, truncated with upper and lower bounds and with mean squared error as loss, multi-task joint training (weighted cross-entropy + MSE), select the maximum value of Softmax or regression value to map to state label for inference.
[0146] In the above implementation scheme, h t After completing the time-level dynamic fusion, the fusion vector s is... t The results of context modeling integrate sequential dependency information.
[0147] LayerNorm is a technique used in neural networks to normalize the feature dimensions, which can stabilize training, accelerate convergence, and maintain consistent normalization results even with small batches or single samples.
[0148] Softmax transforms the logits vector output by the model in the classification head into a probability distribution, making the sum of the probabilities of all classes equal to 1, thus allowing it to be directly used for multi-class decision-making.
[0149] In one embodiment, the dosing rate per second is determined based on the patient's sedation status and the implemented sedation protocol:
[0150] The system reads the patient's current sedation indicators and preset targets in real time, while retaining the drug administration rate from the previous moment. Users need to preset the PID control gain and the upper and lower limits of the drug administration rate, and configure the maximum rate variation range.
[0151] The error is calculated, the rate increment is calculated using the PID formula, and finally the new rate is clipped with upper and lower limits and smoothed to generate the final dosing rate, which is then sent to the infusion pump.
[0152] When the absolute value of the error exceeds a predefined threshold, an alarm is triggered, and the PID parameters are adaptively tuned periodically or as needed to cope with the dynamic changes in the patient's physiological state.
[0153] In this implementation scheme, the changes in the drug administration rate are determined by calculating the resulting errors, thereby ensuring the stability of the patient during the surgical procedure.
[0154] In one embodiment, patient condition characteristics and ultrasound features are visualized:
[0155] Subscribe to patient status features and ultrasound feature interfaces from the backend multimodal fusion service, and transmit them uniformly in JSON or Protocol Buffer format;
[0156] Define the main view area and side panel area on the front end. The main view is used to display time series curves and heat maps.
[0157] Facial thermal imaging, pupil diameter, and facial expression unit state features are mapped to multiple polylines or heat map bars, with different indicators highlighted by color and line width.
[0158] Real-time rendering of temporal variation curves of ultrasonic texture feature intensity or subband energy, with the ability to switch between frequency domain and time domain views;
[0159] Achieve linkage between status feature and ultrasound feature charts: After selecting a time interval on any curve or image, the corresponding interval in another view will be highlighted simultaneously.
[0160] Although the present invention has been disclosed in conjunction with the above embodiments, it is not intended to limit the present invention. Any person skilled in the art can make various modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be defined by the appended claims.
Claims
1. An intraoperative image processing system for orthopedic anesthesia based on optimized sedation management and regional blockade, characterized in that: include: The data acquisition module acquires intraoperative patient status images and ultrasound images, connects the status image and ultrasound image data sources to the PTP network, and aligns the timing errors between the status image and ultrasound image. The image processing module preprocesses the patient status image and ultrasound image, suppresses image noise, and enhances the contrast and pseudo-color of the status image and ultrasound image. The feature extraction module extracts state features from the state image and ultrasound features from the ultrasound image, respectively, and constructs a state feature set and an ultrasound feature set. The multimodal fusion model divides the state feature set and ultrasound feature set into multiple time windows, constructs a time-series vector, maps each image feature to a unified dimension through a feature encoder, dynamically integrates different modalities, and outputs the fused features to the model to determine the patient's sedation status. The sedation feedback module determines the dosing rate per second based on the patient's sedation status and the implemented sedation protocol. The visualization and interaction module displays patient status characteristics and ultrasound features through images and can be used to control the status of the drug delivery pump.
2. The intraoperative image processing system for orthopedic anesthesia based on optimized sedation management and regional blockade as described in claim 1, characterized in that, Status images include: facial and eye videos, pupil dynamic imaging, and thermal imaging; ultrasound images include: cardiac ultrasound, vascular Doppler, transcranial Doppler, lung ultrasound, and airway ultrasound.
3. The intraoperative image processing system for orthopedic anesthesia based on optimized sedation management and regional blockade as described in claim 2, characterized in that, The steps for time-series alignment of the status image data source and the ultrasound image data source into the PTP network are as follows: Select a highly stable OCXO as the PTP Grandmaster, deploy two levels of Boundary Clock switches, and divide the image data source and PTP control flow into different VLANs and configure 802.1p priority respectively; The Grandmaster and Boundary Clock switches are interconnected via a dual-rack redundant link, and the image data VLAN and PTP control VLAN are connected on the link. The Ethernet interface of each acquisition device is connected to the corresponding switch port, and 802.1Q tagging is enabled. Enable IEEE 1588 v2 hardware timestamps on each data source device, Frame Grabber, set up master and slave devices, and activate Chunk Timestamp output; On the switch, check the master offset and mean path delay metrics to ensure that the error is within the sub-microsecond to microsecond range. On the host side, you can use ptp4l or device logs to track the latency jitter of each device and check the alignment of the Chunk Timestamp with the global clock.
4. The intraoperative image processing system for orthopedic anesthesia based on optimized sedation management and regional blockade as described in claim 3, characterized in that, The image preprocessing steps are as follows: The status image and ultrasound image are denoised in parallel. The facial / eye video, pupil dynamics and thermal images are filtered in the spatiotemporal domain to remove sensor noise and suppress patient motion artifacts. The ultrasound image is denoised by combining anisotropic diffusion filtering and wavelet domain thresholding to suppress speckle noise while preserving the edge details of nerve and blood vessel tissue. Contrast enhancement was performed on both types of images. The state image was enhanced by applying CLAHE adaptive histogram equalization to highlight facial muscle texture, pupil edges, and thermal spot temperature gradients. The ultrasound image was enhanced by multi-scale Retinex or gradient domain enhancement algorithms to improve the grayscale differences of low-contrast tissue structures. Grayscale or thermal images are mapped to pseudocolor, thermal images are mapped to temperature gradients, and facial thermal images can intuitively distinguish between sympathetic and parasympathetic activity; ultrasound echo intensity is displayed through grayscale pseudocolor, with different echo amplitudes corresponding to different hues, enhancing the visualization of nerve bundles, blood vessels, and drug diffusion areas. All image frames are subjected to linear or logarithmic gamma correction, and smoothing and frame interpolation are performed in time. The enhanced state features are displayed simultaneously with the ultrasound segmentation / pseudocolor layer.
5. The intraoperative image processing system for orthopedic anesthesia based on optimized sedation management and regional blockade as described in claim 4, characterized in that, Extract state features from the state image and construct a state feature set: Automatically detect and align facial poses in each state frame image, crop out the ROIs of eyes, forehead, and nose and lips, and uniformly scale each ROI to a fixed resolution, while linearly normalizing pixel or temperature values. Using a preset duration as a sliding window, the preprocessed frame sequence is sampled at equal intervals, with each time window containing the same number of frames. The balance between continuity and responsiveness is achieved through window overlap. Within each time window, facial expression action unit intensity, eyelid opening and closing index and blink rate are extracted for the facial ROI. Segmentation and tracking algorithms are used to obtain pupil diameter and constriction / diffuse speed. Local temperature difference and rate of change are calculated from thermal imaging to form multi-dimensional visual and dynamic features. For each extracted indicator, calculate the mean, variance, extreme values, and slope of change statistics, and concatenate the statistics into a high-dimensional feature vector in a predetermined order; The feature vectors corresponding to each time window are arranged into an n×d feature matrix in time sequence. The matrix columns are normalized using Z-score and then stored in a circular buffer so that the multimodal fusion and real-time classification modules can call them at any time.
6. The intraoperative image processing system for orthopedic anesthesia based on optimized sedation management and regional blockade as described in claim 5, characterized in that, Extract ultrasound features from ultrasound images and construct an ultrasound feature set: Based on clinical annotations or deep learning models, the system automatically locates and segments the tissue or lesion region to be analyzed. For dynamic ultrasound that requires frame-by-frame tracking, it combines optical flow or block-matching-based motion estimation algorithms to achieve continuous updates of the ROI. Within each ROI, a grayscale histogram and its statistics, including mean, variance, skewness, and kurtosis, are calculated to capture the overall echo intensity distribution characteristics, providing basic indicators for subsequent texture and morphology analysis. Contrast, energy, correlation, and entropy indices are extracted using gray-level co-occurrence matrix, gray-level running length matrix, or local binary mode to reveal the internal microstructure and speckle distribution characteristics of tissues. Based on ROI binarization and contour detection, the area, perimeter, compactness, and aspect ratio parameters of the target region are measured. At the same time, the edge gradient, curvature, and edge density are calculated to reflect the smoothness and complexity of the lesion boundary. Discrete wavelet transform or Fourier transform is applied to ROI images to extract energy distribution and spectral entropy in different frequency bands, capture multi-scale details of ultrasonic echoes, and use subband energy or spectral centroid as additional features. For multi-frame or Doppler ultrasound sequences, calculate flow velocity distribution, blood flow time, intensity curve parameters, tissue motion velocity and strain rate indices to provide temporal information for assessing vascular status and tissue stiffness; The ultrasound features are concatenated into a high-dimensional vector in a predetermined order. Each dimension is standardized by Z-score and stored in the feature matrix to obtain an ultrasound feature set consisting of the number of samples × the feature dimension.
7. The intraoperative image processing system for orthopedic anesthesia based on optimized sedation management and regional blockade as described in claim 6, characterized in that, Construct temporal vectors and map them to a unified dimension via a feature encoder: Under a sliding window with fixed duration and step size, the state features and ultrasound features are aligned and sampled, then segmented separately. Each window contains the same number of time points and outputs a shape of [M, L, D]. state ] and [M, L, D ultra Preliminary temporal fragments; The state features are standardized with zero mean and unit variance, and missing values are retained and filled using linear interpolation or previous values. The original texture or depth features of the ultrasound image are extracted first, the spatial resolution is unified, and the channels are normalized and linearly mapped to a new interval according to fixed rules. Lightweight encoders were designed for two modalities: a Medium Level Logic (MLP) with hidden layers was used for state features, and a small CNN was used for ultrasound features, mapping their respective inputs to the same embedding dimension D. embed Based on modal differences, select independent or parameter-shared encoders, and add Norm and Dropout to each layer to improve stability; The encoder outputs two sets of equal-length sequences E state =[e1,…,e M ] and E ultra =[u1,…,u M Then, a fusion vector s is generated by simple concatenation or weighted linear fusion. i Finally, a multimodal time series sequence S=[s1,…,s2] is obtained for use in subsequent time series networks. M ].
8. The intraoperative image processing system for orthopedic anesthesia based on optimized sedation management and regional blockade as described in claim 7, characterized in that, Dynamically integrate different modalities and output the fused features as the model: At each time step t, take the state embedding e. t and ultrasound embedded u t Construct a joint query Q=[e t ;u t The attention score α is calculated using multi-head self-attention, along with the key-value pair K and V. (m) t, where each head m corresponds to a set of learnable modal weights, ultimately generating a weighted fusion vector s. t ; Let the sequence {s1,…,s} M The input encoder captures short-term fluctuations and long-term dependencies, and the output context embedding h is used. t Residual connections are added between layers to maintain information flow and stabilize training with LayerNorm; embedding h in context t Apply a two-branch head: Discrete grading: Fully connected → Softmax predicts four levels of calming states, using cross-entropy as the loss; Continuous exponent: Fully connected → linear output sedation exponent, truncated with upper and lower bounds and with mean squared error as loss, multi-task joint training, selecting the maximum value of Softmax or regression value to map to state labels for inference.
9. The intraoperative image processing system for orthopedic anesthesia based on optimized sedation management and regional blockade as described in claim 8, characterized in that, The dosing rate per second was determined based on the patient's sedation status and the implemented sedation protocol: The system reads the patient's current sedation indicators and preset targets in real time, while retaining the drug administration rate from the previous moment. Users need to preset the PID control gain and the upper and lower limits of the drug administration rate, and configure the maximum rate variation range. The error is calculated, the rate increment is calculated using the PID formula, and finally the new rate is clipped with upper and lower limits and smoothed to generate the final dosing rate, which is then sent to the infusion pump. When the absolute value of the error exceeds a predefined threshold, an alarm is triggered, and the PID parameters are adaptively tuned periodically or as needed to cope with the dynamic changes in the patient's physiological state.
10. The intraoperative image processing system for orthopedic anesthesia based on optimized sedation management and regional blockade as described in claim 9, characterized in that, Visualize and display patient condition characteristics and ultrasound features: Subscribe to patient status features and ultrasound feature interfaces from the backend multimodal fusion service and transmit them uniformly in JSON or Protocol Buffer format; Define the main view area and side panel area on the front end. The main view is used to display time series curves and heat maps. Facial thermal imaging, pupil diameter, and facial expression unit state features are mapped to multiple polylines or heat map bars, with different indicators highlighted by color and line width. Real-time rendering of temporal variation curves of ultrasonic texture feature intensity or subband energy, with the ability to switch between frequency domain and time domain views; Achieve linkage between status feature and ultrasound feature charts: After selecting a time interval on any curve or image, the corresponding interval in another view will be highlighted simultaneously.
Citation Information
Patent Citations
Image processing system in orthopedic anesthesia based on optimized sedation management and regional blocking
CN108784836A