A method and system for remote scheduling of intraoperative multi-modal medical images
Patent Information
- Application Number
- CN202511283516.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2045-09-09
AI Technical Summary
[0002]在现有的医学图像处理技术中,术中图像的调取是常见操作然而,手术过程具有高度的复杂性和动态性,不同手术阶段所产生的图像在特征表现上往往存在较大的差异,且这些差异并非是简单直观、易于区分的
[0020] This invention provides real-time monitoring of surgical progress through surgical stage identification. Simultaneously, it precisely filters images related to the current surgical state using a semantic feature matrix, solving the challenges of multimodal image fusion and surgical stage identification. Furthermore, by combining dynamically adjusted displayed images and non-contact gesture control, surgeons can efficiently acquire the necessary images, reducing surgical interruptions and ensuring smooth operation. Additionally, through interactive scheduling between the non-contact gesture control model and preset C-arm position coordinates, surgeons can operate the image display in a more natural and convenient manner. An embedded reinforcement learning mechanism and surgeon image retrieval trajectory optimization of the initial scheduling results, combined with real-time surgeon feedback to update the displayed images, further enhance the personalization and adaptability of the image display. In summary, this invention solves the problem of inaccurate surgical stage identification during surgical image retrieval, which hinders the fulfillment of surgeons' image display needs.
Smart Images

Figure CN121306482B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image processing technology, and more specifically, to a method and system for remote scheduling of intraoperative multimodal medical images. Background Technology
[0002] In existing medical image processing technologies, retrieving intraoperative images is a common operation. However, the surgical process is highly complex and dynamic, and images from different surgical stages often exhibit significant differences in feature representation. These differences are not simple, intuitive, or easily distinguishable. Current technologies are not precise or comprehensive enough in image feature extraction and analysis; they only focus on some superficial and singular features, failing to effectively capture the deep, multi-dimensional features closely related to the surgical stage within surgical images. This results in the inability to accurately identify the surgical stage during image retrieval, making it difficult to meet the surgeon's image display needs.
[0003] Therefore, there is an urgent need for a remote scheduling method and system for intraoperative multimodal medical images, which solves the problem of not being able to accurately identify the surgical stage during the process of calling up surgical images and making it difficult to meet the surgeon's image display needs. Summary of the Invention
[0004] The purpose of this invention is to provide a method and system for remote scheduling of intraoperative multimodal medical images to improve the aforementioned problems. To achieve the above objective, the technical solution adopted by this invention is as follows:
[0005] In a first aspect, this application provides a method for remote scheduling of intraoperative multimodal medical images, comprising:
[0006] Acquire multimodal medical images and surgeon images retrieval trajectory;
[0007] The multimodal medical images are input into a multi-input spatiotemporal convolutional neural network, and surgical stage recognition is performed by combining an attention mechanism and a gated recurrent unit, and the intraoperative state category is output.
[0008] Based on the multimodal medical images, image nodes are extracted and a graph is constructed. The image nodes are then enhanced and embedded through the correlation relationships of all graphs to obtain a semantic feature matrix.
[0009] Candidate images of the current surgical state in the semantic feature matrix are selected based on the intraoperative state category, and a scheduling score model is constructed by calculating the scheduling score of each candidate image.
[0010] Based on the scheduling scoring model and the current intraoperative behavior status, the displayed image is dynamically adjusted, and interactive scheduling is performed in combination with the non-contact gesture control model and the preset C-arm position coordinates to obtain the initial scheduling result.
[0011] The initial scheduling result is optimized based on the embedded reinforcement learning mechanism and the surgeon's image retrieval trajectory. The displayed image is updated in conjunction with the surgeon's real-time feedback, and the optimized image retrieval result is output.
[0012] Secondly, this application also provides a remote scheduling system for intraoperative multimodal medical images, comprising:
[0013] The acquisition module is used to acquire the retrieval trajectory of multimodal medical images and surgeon images;
[0014] The recognition module is used to input the multimodal medical image into a multi-input spatiotemporal convolutional neural network, combine an attention mechanism and a gated recurrent unit to recognize the surgical stage, and output the intraoperative state category.
[0015] The construction module is used to extract image nodes and construct a graph based on the multimodal medical image, and to enhance the embedding of image nodes through the correlation relationship of all graphs to obtain a semantic feature matrix;
[0016] The calculation module is used to filter candidate images of the current surgical state in the semantic feature matrix according to the intraoperative state category, and to build a scheduling score model by calculating the scheduling score of each candidate image.
[0017] The adjustment module is used to dynamically adjust the displayed image based on the scheduling scoring model and the current intraoperative behavior status, and to perform interactive scheduling by combining the non-contact gesture control model with the preset C-arm position coordinates to obtain the initial scheduling result.
[0018] The update module is used to optimize the initial scheduling results based on the embedded reinforcement learning mechanism and the surgeon's image retrieval trajectory, update the displayed images in conjunction with the surgeon's real-time feedback, and output the optimized image retrieval results.
[0019] The beneficial effects of this invention are as follows:
[0020] This invention provides real-time monitoring of surgical progress through surgical stage identification. Simultaneously, it precisely filters images related to the current surgical state using a semantic feature matrix, solving the challenges of multimodal image fusion and surgical stage identification. Furthermore, by combining dynamically adjusted displayed images and non-contact gesture control, surgeons can efficiently acquire the necessary images, reducing surgical interruptions and ensuring smooth operation. Additionally, through interactive scheduling between the non-contact gesture control model and preset C-arm position coordinates, surgeons can operate the image display in a more natural and convenient manner. An embedded reinforcement learning mechanism and surgeon image retrieval trajectory optimization of the initial scheduling results, combined with real-time surgeon feedback to update the displayed images, further enhance the personalization and adaptability of the image display. In summary, this invention solves the problem of inaccurate surgical stage identification during surgical image retrieval, which hinders the fulfillment of surgeons' image display needs.
[0021] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing embodiments of the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a schematic diagram of the remote scheduling method for intraoperative multimodal medical images as described in an embodiment of the present invention;
[0024] Figure 2 This is a schematic diagram of the structure of the remote scheduling device for intraoperative multimodal medical images described in an embodiment of the present invention.
[0025] The diagram is labeled as follows: 800, remote scheduling device for intraoperative multimodal medical images; 801, processor; 802, memory; 803, multimedia component; 804, I / O interface; 805, communication component. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0027] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0028] Example 1:
[0029] This embodiment provides a method for remote scheduling of intraoperative multimodal medical images.
[0030] See Figure 1 The figure shows that the method includes steps S1 to S6, including:
[0031] S1: Obtain the trajectory of multimodal medical images and surgeon images;
[0032] In this step, the multimodal medical images include preoperative images, intraoperative digital subtraction angiography images, catheter data sequences, and vascular imaging images.
[0033] S2: Input the multimodal medical image into a multi-input spatiotemporal convolutional neural network, combine the attention mechanism and gated recurrent unit to identify the surgical stage, and output the intraoperative state category;
[0034] To clarify the specific method for obtaining intraoperative status categories, step S2 includes S21 to S24, specifically:
[0035] S21: The multimodal medical image is preprocessed to obtain a preprocessed medical image;
[0036] To clarify the specific acquisition method of preprocessed medical images, step S21 includes S211 to S216, specifically:
[0037] S211: Standardize the preoperative multimodal images and intraoperative digital subtraction angiography images in the multimodal medical images to obtain standardized images;
[0038] In this step, the standardization process includes spatial resolution unification, image cleaning, and registration preprocessing.
[0039] The process includes: unifying the spatial resolution of the images to ensure that all images have the same resolution; cleaning the images to remove noise and invalid information; and performing registration preprocessing to align the preoperative and intraoperative images spatially.
[0040] S212: Extract features from the catheter data sequence in the multimodal medical images to obtain catheter behavior features;
[0041] In this step, the catheter data sequence in the multimodal medical images is spatiotemporally aligned with a preset timestamp to ensure data synchronization in time and space. Next, the rate of change of velocity per unit time is calculated based on the aligned catheter data sequence, and a one-dimensional propulsion velocity tensor is generated. Finally, feature extraction is performed using the one-dimensional propulsion velocity tensor to obtain the behavioral features of the catheter.
[0042] S213: Resample the vascular imaging image in the multimodal medical image to obtain a resampled vascular imaging image;
[0043] In this step, the resampling process includes adjusting the image resolution and interpolation to ensure the applicability of the vascular imaging images in subsequent 3D reconstruction, thus solving the problem of inconsistent resolution of vascular imaging images.
[0044] S214: Perform three-dimensional reconstruction processing on the resampled vascular imaging image based on the traveling cube algorithm to generate a three-dimensional vascular model of the patient;
[0045] In this step, the patient's three-dimensional vascular model is used to visually display the patient's vascular structure, providing support for surgical planning and navigation, and solving the problem of generating a three-dimensional vascular model from two-dimensional vascular imaging images.
[0046] S215: Extract anatomical features from the patient's three-dimensional vascular model to generate the patient's vascular anatomical features;
[0047] In this step, anatomical feature extraction provides more detailed vascular information for surgical planning.
[0048] S216: Based on the standardized image, the catheter behavior characteristics, and the patient's vascular anatomy characteristics, a preprocessed medical image is constructed.
[0049] In this step, a preprocessed medical image containing rich information is generated by constructing multiple features, solving the problem of how to integrate multiple features into a single preprocessed image.
[0050] S22: Input the preprocessed medical image into a multi-input spatiotemporal convolutional neural network, and perform one-dimensional convolution operations on the catheter advancement behavior tensor and the surgeon's action tensor in the preprocessed medical image to extract static behavior features;
[0051] In this step, a one-dimensional convolution operation is used to extract static features related to surgical behavior from time-series data.
[0052] S23: Based on the attention mechanism and gating loop unit, the surgeon's continuous behavior is memorized to obtain the surgeon's dynamic behavioral characteristics;
[0053] In this step, the surgeon's continuous behavior is weighted based on the attention mechanism, and the time-series data of the surgeon's behavior is processed by the gated recurrent unit (GRU) to capture dynamic changes. The dynamic changes of the surgeon's continuous behavior patterns are captured by the attention mechanism and the gated recurrent unit (GRU) to extract more representative dynamic behavior features.
[0054] S24: Based on the static behavioral characteristics and the surgeon's dynamic behavioral characteristics, predict the surgical state category of the current surgical state and output the intraoperative state category.
[0055] In this step, the static behavioral features and the surgeon's dynamic behavioral features are fused together. The fused features are then used to predict the surgical state category, and the intraoperative state category is output. By accurately predicting the current surgical state using the static behavioral features and the surgeon's dynamic behavioral features, the surgical progress can be monitored in real time, thus solving the problem of automatic surgical state identification.
[0056] S3: Based on the multimodal medical images, extract image nodes and construct a graph. Enhance the embedding of image nodes through the correlation relationships of all graphs to obtain a semantic feature matrix;
[0057] In this step, the semantic feature matrix is used to accurately filter images related to the current surgical state, solving the problems of multimodal image fusion and surgical stage identification.
[0058] To clarify the specific method for obtaining the semantic feature matrix, step S3 includes S31 to S34, specifically:
[0059] S31: Extract image nodes from the multimodal medical image based on a convolutional neural network to obtain an image node set;
[0060] In this step, the features of the image node include grayscale statistical features, positional encoding, structural labels, and image semantic vectors;
[0061] The gray-level statistical features are used to reflect the gray-level distribution of the image, the positional encoding is used to represent the spatial position of the image node in the image, the structural label is used to label the anatomical structure type of the image node, and the image semantic vector is used to extract semantic features through a pre-trained model.
[0062] S32: Based on the different relationships between the image node sets, establish corresponding edges, and combine each image node and edge to obtain a graph;
[0063] In this step, edges are established based on the relationships between image nodes to construct the graph. The rules for edge establishment are as follows:
[0064] When two images belong to the same anatomical region but are in different modalities, establish modal-corresponding edges;
[0065] When two images are frequently activated in the same intraoperative state, a state co-occurrence edge is established based on the state transition co-occurrence frequency.
[0066] When spatial continuity exists between images, spatial adjacency edges are established; using the above rules, a graph is systematically constructed.
[0067] The graph is constructed using the above rules. Each node in the graph represents an image node, and the edges represent the relationships between nodes.
[0068] S33: Dynamically assign weights to the graph attention network based on the attention mechanism to obtain an optimized graph attention network;
[0069] In this step, the attention mechanism dynamically adjusts the weights according to the importance of the image nodes, better captures the key relationships between nodes, improves the representation capability of the graph, and solves the problem of inconsistent node importance in the graph.
[0070] S34: Perform context-enhanced embedding on each graph according to the optimized graph attention network to obtain the image semantic feature matrix.
[0071] In this step, nodes in each graph are embedded according to the optimized graph attention network to enhance the contextual information of the nodes. The embedded node features are combined into a semantic feature matrix to obtain the image semantic feature matrix.
[0072] S4: Based on the intraoperative state category, candidate images of the current surgical state in the semantic feature matrix are selected, and a scheduling score model is constructed by calculating the scheduling score of each candidate image;
[0073] To clarify the specific method for obtaining the scheduling scoring model, step S4 includes S41 to S44, specifically:
[0074] S41: Filter the current surgical state in the semantic feature matrix according to the intraoperative state category to obtain candidate images;
[0075] In this step, the filtering operation includes: extracting feature vectors corresponding to intraoperative state categories, calculating the similarity between each image in the semantic feature matrix and the current state category, and filtering candidate images based on a similarity threshold. This filtering process is used to quickly locate images most relevant to the current surgical state, reducing the computational load of subsequent processing.
[0076] S42: Input the surgeon's image retrieval trajectory into the gated loop unit for processing, capture the surgeon's long-term operational preferences through time-series coding, and generate a surgeon preference embedding vector;
[0077] In this step, the surgeon's image retrieval trajectory is a feature sequence of historical retrieval images. This trajectory is input into a single-layer gated recurrent unit (GRU) to capture the surgeon's short-term operational preferences, and into a multi-layer gated recurrent unit (GRU) to capture the surgeon's long-term operational preferences. The short-term and long-term operational preference vectors are fused to generate a surgeon preference embedding vector, which is used to understand the surgeon's behavioral patterns, thus solving the problem of how to extract long-term preference information from the surgeon's historical operations.
[0078] The gated loop unit (GRU) includes a single-layer gated loop unit and a multi-layer gated loop unit.
[0079] S43: Calculate the scheduling score based on the current state vector in the intraoperative state category, the surgeon's preference embedding vector, and the candidate image to obtain the matching score between the current state and the surgeon's preference;
[0080] To clarify the specific method for obtaining the matching score between the current state and the surgeon's preferences, step S43 includes steps S431 to S433, specifically:
[0081] S431: After dimensional alignment of the current state vector, the surgeon's preference embedding vector, and the semantic feature vector of the candidate image, the fusion vector is obtained;
[0082] In this step, dimensional alignment is performed based on the current state vector, the surgeon's preference embedding vector, and the semantic feature vector of the candidate image. This dimensional alignment is used to ensure the dimensional consistency of features from different sources. By concatenating the aligned vectors, a fused vector is obtained.
[0083] S432: Input the fused vector into the converter model, calculate the semantic association weights of the fused vector through a multi-head attention mechanism, and obtain the semantic association feature matrix of the candidate image;
[0084] In this step, the multi-head attention mechanism is used to capture the semantic associations between different parts of the fused vector, generating a more representative semantic association feature matrix.
[0085] The semantic association weights include the semantic relevance among state, preference, and image.
[0086] Preferably, the converter model is a Transformer model.
[0087] S433: Based on the semantic association feature matrix of the candidate images, and combined with the preset scoring rules, calculate the scheduling score for each candidate image to obtain the matching score between the current state and the surgeon's preference.
[0088] In this step, the scoring rules are used to quickly evaluate the suitability of candidate images.
[0089] S44: Based on a preset matching threshold, the matching scores between the current state and the surgeon's preference are ranked and constructed, and the scheduling scoring model is output.
[0090] In this step, the scheduling scoring model is constructed to quickly evaluate the applicability of candidate images.
[0091] Preferably, the arrangement is in descending order, and the scheduling and scoring model includes the image index, layer type, and suggested display area, which is used to quickly evaluate the applicability of candidate images.
[0092] S5: Based on the scheduling scoring model and the current intraoperative behavior status, dynamically adjust the displayed image, and perform interactive scheduling with the non-contact gesture control model and the preset C-arm position coordinates to obtain the initial scheduling result;
[0093] In this step, the dynamic adjustment of the displayed image and the non-contact gesture control are combined to enable the surgeon to efficiently acquire the required images, reduce surgical interruption time, and ensure the smoothness of the surgery; and the non-contact gesture control model interacts and schedules with the preset C-arm position coordinates, allowing the surgeon to operate the image display in a more natural and convenient way.
[0094] To clarify the specific method for obtaining the initial scheduling result, step S5 includes S51 to S56, specifically:
[0095] S51: Obtain historical blood vessel angle data;
[0096] S52: Based on the scheduling and scoring model and the current intraoperative behavior status, dynamically adjust the displayed image to obtain the target displayed image;
[0097] In this step, the candidate image with the highest score is selected based on the scheduling and scoring model, and the display position and size of the image are adjusted to the target area according to the current intraoperative behavior status.
[0098] Preferably, the target area includes a central main display area, a right-side AI results area, and a bottom prompt area.
[0099] S53: Determine the type of the target display image, and adjust the display layout according to the image type to obtain the adjusted display layout;
[0100] In this step, the target display image is type-determined. When the target display image is a 3D model layer, the interactive module of the rotating display device is automatically activated to support the user's 3D perspective operation. When the target display image is a keyframe sequence, it is automatically aligned to the nearest key event node based on the time tag to achieve precise timing matching.
[0101] S54: Based on machine learning algorithms, train historical vascular angle data and current patient 3D vascular model to generate the optimal angle training model;
[0102] In this step, machine learning algorithms learn the optimal vascular angle information from historical data and the current patient model.
[0103] S55: Analyze the angles in the vascular imaging image, train the model with the best angle to recommend the angle layer, and obtain the angle recommendation layer;
[0104] In this step, the angle recommendation layer provides the surgeon with the best vascular angle information.
[0105] S56: Based on the non-contact gesture control model and the angle recommendation layer, the adjustment display layout and the preset C-arm position coordinates are interactively scheduled to obtain the initial scheduling result.
[0106] In this step, the non-contact gesture control model and the recommended values in the angle recommendation layer are transmitted to the intraoperative control panel, and the adjustment of the display layout and the preset C-arm position coordinates are interactively scheduled to achieve automatic preliminary angle adjustment and viewpoint switching that conforms to the current anatomical structure. The interactive scheduling can effectively reduce the intraoperative angle trial and error time, reduce the frequency of contrast agent injection, reduce X-ray exposure time, improve the overall operation efficiency, and reduce the radiation dose to the surgeon and the patient.
[0107] The operator can choose from the following options: maintain the system's recommended angle; fine-tune the angle using gestures or control panel buttons; or ignore the recommendation and set a custom angle.
[0108] The non-contact gesture control model is used to layout the display area, monitor the surgeon's behavior in real time after retrieving images, and adjust it based on the surgeon's gesture commands. Specific monitored behaviors include:
[0109] Whether to click or view the image (considered as following).
[0110] Whether the image is viewed for an extended period of time (considered valid use).
[0111] Whether to actively switch / replace the image (considered a negative preference).
[0112] Should the system recommendation be ignored and other images be retrieved directly (considered as model mismatch)?
[0113] S6: Optimize the initial scheduling results based on the embedded reinforcement learning mechanism and the surgeon's image retrieval trajectory, update the displayed images in conjunction with the surgeon's real-time feedback, and output the optimized image retrieval results.
[0114] To clarify the specific method for obtaining the optimized image retrieval results, step S6 includes S61 to S65, specifically:
[0115] S61: Optimize the initial scheduling result based on the current intraoperative behavior state and the surgeon's image retrieval trajectory using the embedded reinforcement learning mechanism to obtain an optimized scheduling result;
[0116] In this step, based on the current intraoperative behavior state and the surgeon's image retrieval trajectory, a reinforcement learning algorithm is used to update the scheduling strategy. The initial scheduling result is adjusted through the optimized scheduling strategy to obtain an optimized scheduling result.
[0117] S62: Train the behavior feedback pool based on the optimized scheduling results to obtain the optimized feedback pool;
[0118] In this step, the optimized scheduling results are fed back to the behavior feedback pool. The behavior feedback pool is trained using a machine learning algorithm to optimize its model parameters, resulting in an optimized behavior feedback pool for subsequent scheduling optimization.
[0119] S63: Optimize the parameters in the scheduling model using the policy gradient method to obtain an optimized scheduling model;
[0120] In this step, the scheduling model includes a Transformer classifier. The parameters in the scheduling model are updated according to the policy gradient method to obtain an optimized scheduling model, which is used for subsequent image scheduling.
[0121] S64: Update the gated loop unit based on the optimized feedback pool and optimized scheduling model to obtain the surgeon's preference trend;
[0122] In this step, the optimized feedback pool and optimized scheduling model are input into the gated recurrent unit (GRU), the parameters of the GRU are updated through backpropagation, and the surgeon's preference trend is obtained through the updated GRU.
[0123] S65: Based on the surgeon's preference trend, and combined with the surgeon's real-time feedback, update the displayed image and output the optimized image retrieval result.
[0124] In this step, real-time feedback from the surgeon is obtained through non-contact gesture control. The layout and content of the displayed image are adjusted based on the surgeon's preferences and this feedback. The displayed image is then updated according to the adjusted layout and content, and the optimized image is output to the operating table for the surgeon's use, resulting in an optimized image retrieval result. This enhances the personalization and adaptability of the image display.
[0125] Example 2:
[0126] This embodiment provides a remote scheduling system for intraoperative multimodal medical images, the system comprising:
[0127] The acquisition module is used to acquire the retrieval trajectory of multimodal medical images and surgeon images;
[0128] The recognition module is used to input the multimodal medical image into a multi-input spatiotemporal convolutional neural network, combine an attention mechanism and a gated recurrent unit to recognize the surgical stage, and output the intraoperative state category.
[0129] To clarify the specific methods for obtaining the identification module, the following are included:
[0130] The first processing unit is used to preprocess the multimodal medical image to obtain a preprocessed medical image;
[0131] To clarify the specific acquisition method of the first processing unit, the following are included:
[0132] The standardization subunit is used to standardize the preoperative multimodal images and intraoperative digital subtraction angiography images in the multimodal medical images to obtain standardized images;
[0133] The first extraction subunit is used to extract features from the catheter data sequence in the multimodal medical image to obtain catheter behavior features;
[0134] The resampling subunit is used to resample the vascular imaging image in the multimodal medical image to obtain a resampled vascular imaging image.
[0135] The reconstruction subunit is used to perform three-dimensional reconstruction processing on the resampled vascular imaging image based on the traveling cube algorithm to generate a three-dimensional vascular model of the patient.
[0136] The second extraction subunit is used to extract anatomical features from the patient's three-dimensional vascular model and generate the patient's vascular anatomical features.
[0137] A sub-unit is constructed based on the standardized image, the catheter behavior features, and the patient's vascular anatomy features to obtain a preprocessed medical image.
[0138] The convolutional unit is used to input the preprocessed medical image into a multi-input spatiotemporal convolutional neural network, and to perform one-dimensional convolution operations on the catheter advancement behavior tensor and the surgeon's action tensor in the preprocessed medical image to extract static behavioral features.
[0139] The memory unit is used to memorize the surgeon's continuous behavior based on the attention mechanism and the gating loop unit, so as to obtain the surgeon's dynamic behavioral characteristics;
[0140] The prediction unit is used to predict the surgical state category based on the static behavioral characteristics and the surgeon's dynamic behavioral characteristics, and output the intraoperative state category.
[0141] The construction module is used to extract image nodes and construct a graph based on the multimodal medical image, and to enhance the embedding of image nodes through the correlation relationship of all graphs to obtain a semantic feature matrix;
[0142] The calculation module is used to filter candidate images of the current surgical state in the semantic feature matrix according to the intraoperative state category, and to build a scheduling score model by calculating the scheduling score of each candidate image.
[0143] To clarify the specific methods for obtaining the calculation module, the following are included:
[0144] The filtering unit is used to filter the current surgical state in the semantic feature matrix according to the intraoperative state category to obtain candidate images;
[0145] The second processing unit is used to input the surgeon's image retrieval trajectory into the gated loop unit for processing, and to capture the surgeon's long-term operational preferences through time-series coding to generate a surgeon preference embedding vector.
[0146] The scoring unit is used to calculate the scheduling score based on the current state vector in the intraoperative state category, the surgeon's preference embedding vector, and the candidate image, to obtain the matching score between the current state and the surgeon's preference.
[0147] The sorting unit is used to sort and construct the matching scores between the current state and the surgeon's preference based on a preset matching threshold, and output the scheduling scoring model.
[0148] The adjustment module is used to dynamically adjust the displayed image based on the scheduling scoring model and the current intraoperative behavior status, and to perform interactive scheduling by combining the non-contact gesture control model with the preset C-arm position coordinates to obtain the initial scheduling result.
[0149] The update module is used to optimize the initial scheduling results based on the embedded reinforcement learning mechanism and the surgeon's image retrieval trajectory, update the displayed images in conjunction with the surgeon's real-time feedback, and output the optimized image retrieval results.
[0150] It should be noted that the specific manner in which each module performs its operation in the apparatus described in the above embodiments has been described in detail in the embodiments of the method, and will not be elaborated here.
[0151] Example 3:
[0152] Corresponding to the above method embodiments, this embodiment also provides a remote scheduling device for intraoperative multimodal medical images. The remote scheduling device for intraoperative multimodal medical images described below and the remote scheduling method for intraoperative multimodal medical images described above can be referred to in correspondence.
[0153] Figure 2 This is a block diagram illustrating a remote scheduling device 800 for intraoperative multimodal medical images according to an exemplary embodiment. Figure 2 As shown, the remote scheduling device 800 for intraoperative multimodal medical images may include a processor 801 and a memory 802. The remote scheduling device 800 may also include one or more of a multimedia component 803, an I / O interface 804, and a communication component 805.
[0154] The processor 801 controls the overall operation of the remote scheduling device 800 for intraoperative multimodal medical images to complete all or part of the steps in the aforementioned remote scheduling method for intraoperative multimodal medical images. The memory 802 stores various types of data to support the operation of the remote scheduling device 800 for intraoperative multimodal medical images. This data may include, for example, instructions for any application or method operating on the remote scheduling device 800 for intraoperative multimodal medical images, as well as application-related data such as contact data, sent and received messages, images, audio, video, etc. The memory 802 can be implemented using any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 802 or transmitted via the communication component 805. The audio component also includes at least one speaker for outputting audio signals. I / O interface 804 provides an interface between processor 801 and other interface modules, such as keyboards, mice, and buttons. These buttons can be virtual or physical. Communication component 805 is used for wired or wireless communication between the remote scheduling device 800 for intraoperative multimodal medical images and other devices. Wireless communication includes, for example, Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination thereof. Therefore, the corresponding communication component 805 may include a Wi-Fi module, a Bluetooth module, or an NFC module.
[0155] In an exemplary embodiment, the remote scheduling device 800 for intraoperative multimodal medical images may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the aforementioned remote scheduling method for intraoperative multimodal medical images.
[0156] Example 4:
[0157] Corresponding to the above method embodiments, this embodiment also provides a medium. The medium described below can be referred to in conjunction with the above-described method for remote scheduling of intraoperative multimodal medical images.
[0158] A medium storing a computer program, which, when executed by a processor, implements the steps of the method for remote scheduling of intraoperative multimodal medical images as described in the above method embodiments.
[0159] The medium can specifically be any medium capable of storing program code, such as a USB flash drive, external hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0160] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0161] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for remote scheduling of intraoperative multimodal medical images, characterized in that, include: Acquire multimodal medical images and surgeon images retrieval trajectory; The multimodal medical images are input into a multi-input spatiotemporal convolutional neural network, and surgical stage recognition is performed by combining an attention mechanism and a gated recurrent unit, and the intraoperative state category is output. Based on the multimodal medical images, image nodes are extracted and a graph is constructed. The image nodes are then enhanced and embedded through the correlation relationships of all graphs to obtain a semantic feature matrix. Candidate images of the current surgical state in the semantic feature matrix are selected based on the intraoperative state category, and a scheduling score model is constructed by calculating the scheduling score of each candidate image. The specific methods for obtaining the scheduling scoring model include: Candidate images are obtained by filtering the current surgical state in the semantic feature matrix according to the intraoperative state category; The surgeon's image recall trajectory is input into a gated loop unit for processing. The surgeon's long-term operational preferences are captured through time-series coding, and a surgeon preference embedding vector is generated. The fusion vector is obtained by aligning the current state vector in the intraoperative state category, the surgeon's preference embedding vector, and the semantic feature vector of the candidate image according to their dimensions. The fused vector is input into the converter model, and the semantic association weights of the fused vector are calculated through a multi-head attention mechanism to obtain the semantic association feature matrix of the candidate image. Based on the semantic association feature matrix of the candidate images, and combined with the preset scoring rules, a scheduling score is calculated for each candidate image to obtain the matching score between the current state and the surgeon's preference. The matching scores between the current state and the surgeon's preference are ranked and constructed based on a preset matching threshold, and a scheduling scoring model is output. Based on the scheduling scoring model and the current intraoperative behavior status, the displayed image is dynamically adjusted, and interactive scheduling is performed in combination with the non-contact gesture control model and the preset C-arm position coordinates to obtain the initial scheduling result. The initial scheduling result is optimized based on the embedded reinforcement learning mechanism and the surgeon's image retrieval trajectory. The displayed image is updated in conjunction with the surgeon's real-time feedback, and the optimized image retrieval result is output.
2. The method for remote scheduling of intraoperative multimodal medical images according to claim 1, characterized in that, The multimodal medical images are input into a multi-input spatiotemporal convolutional neural network, which combines an attention mechanism with a gated recurrent unit to identify surgical stages and output intraoperative state categories, including: The multimodal medical images are preprocessed to obtain preprocessed medical images; The preprocessed medical image is input into a multi-input spatiotemporal convolutional neural network. The catheter advancement behavior tensor and the surgeon's action tensor in the preprocessed medical image are subjected to one-dimensional convolution operations to extract static behavioral features. Based on the attention mechanism and gating loop unit, the surgeon's continuous behavior is memorized to obtain the surgeon's dynamic behavioral characteristics; Based on the static behavioral characteristics and the surgeon's dynamic behavioral characteristics, the current surgical state category is predicted, and the intraoperative state category is output.
3. The method for remote scheduling of intraoperative multimodal medical images according to claim 2, characterized in that, The multimodal medical images are preprocessed to obtain preprocessed medical images, including: The preoperative multimodal images and intraoperative digital subtraction angiography images in the multimodal medical images are standardized to obtain standardized images; Feature extraction is performed on the catheter data sequence in the multimodal medical images to obtain catheter behavior features; The vascular imaging images in the multimodal medical images are resampled to obtain resampled vascular imaging images; The resampled vascular imaging image is reconstructed using the traveling cube algorithm to generate a three-dimensional vascular model of the patient. Anatomical features are extracted from the patient's three-dimensional vascular model to generate the patient's vascular anatomical features; Preprocessed medical images are constructed based on the standardized images, catheter behavior features, and patient vascular anatomy features.
4. The method for remote scheduling of intraoperative multimodal medical images according to claim 3, characterized in that, Based on the aforementioned scheduling scoring model and the current intraoperative behavioral state, the displayed image is dynamically adjusted. Interactive scheduling is performed using a non-contact gesture control model and preset C-arm position coordinates to obtain initial scheduling results, including: Obtain historical vascular angle data; Based on the scheduling and scoring model and the current intraoperative behavior status, the displayed image is dynamically adjusted to obtain the target displayed image; The target display image is type-determined, and the display layout is adjusted based on the image type to obtain the adjusted display layout; The optimal angle training model is generated by training historical vascular angle data and current patient 3D vascular model based on machine learning algorithms. The angles in the vascular imaging images are analyzed, and the model is trained to recommend layers for the analyzed angles by using the optimal angles, thus obtaining the angle recommendation layer; Based on the non-contact gesture control model and the angle recommendation layer, the adjustment display layout and the preset C-arm position coordinates are interactively scheduled to obtain the initial scheduling result.
5. A remote scheduling system for intraoperative multimodal medical images, characterized in that, include: The acquisition module is used to acquire the retrieval trajectory of multimodal medical images and surgeon images; The recognition module is used to input the multimodal medical image into a multi-input spatiotemporal convolutional neural network, combine an attention mechanism and a gated recurrent unit to recognize the surgical stage, and output the intraoperative state category. The construction module is used to extract image nodes and construct a graph based on the multimodal medical image, and to enhance the embedding of image nodes through the correlation relationship of all graphs to obtain a semantic feature matrix; The calculation module is used to filter candidate images of the current surgical state in the semantic feature matrix according to the intraoperative state category, and to build a scheduling score model by calculating the scheduling score of each candidate image. The calculation module includes: The filtering unit is used to filter the current surgical state in the semantic feature matrix according to the intraoperative state category to obtain candidate images; The second processing unit is used to input the surgeon's image retrieval trajectory into the gated loop unit for processing, and to capture the surgeon's long-term operational preferences through time-series coding to generate a surgeon preference embedding vector. The scoring unit is used to calculate the scheduling score based on the current state vector in the intraoperative state category, the surgeon's preference embedding vector, and the candidate image, to obtain the matching score between the current state and the surgeon's preference. The scoring unit includes: The current state vector, the surgeon's preference embedding vector, and the semantic feature vector of the candidate image are aligned in dimensions and then concatenated to obtain a fusion vector; The fused vector is input into the converter model, and the semantic association weights of the fused vector are calculated through a multi-head attention mechanism to obtain the semantic association feature matrix of the candidate image. Based on the semantic association feature matrix of the candidate images, and combined with the preset scoring rules, a scheduling score is calculated for each candidate image to obtain the matching score between the current state and the surgeon's preference. The sorting unit is used to sort and construct the matching scores between the current state and the surgeon's preference based on a preset matching threshold, and output the scheduling scoring model. The adjustment module is used to dynamically adjust the displayed image based on the scheduling scoring model and the current intraoperative behavior status, and to perform interactive scheduling by combining the non-contact gesture control model with the preset C-arm position coordinates to obtain the initial scheduling result. The update module is used to optimize the initial scheduling results based on the embedded reinforcement learning mechanism and the surgeon's image retrieval trajectory, update the displayed images in conjunction with the surgeon's real-time feedback, and output the optimized image retrieval results.
6. The remote scheduling system for intraoperative multimodal medical images according to claim 5, characterized in that, The identification module includes: The first processing unit is used to preprocess the multimodal medical image to obtain a preprocessed medical image; The convolutional unit is used to input the preprocessed medical image into a multi-input spatiotemporal convolutional neural network, and to perform one-dimensional convolution operations on the catheter advancement behavior tensor and the surgeon's action tensor in the preprocessed medical image to extract static behavioral features. The memory unit is used to memorize the surgeon's continuous behavior based on the attention mechanism and the gating loop unit, so as to obtain the surgeon's dynamic behavioral characteristics; The prediction unit is used to predict the surgical state category based on the static behavioral characteristics and the surgeon's dynamic behavioral characteristics, and output the intraoperative state category.
7. The remote scheduling system for intraoperative multimodal medical images according to claim 6, characterized in that, The first processing unit includes: The standardization subunit is used to standardize the preoperative multimodal images and intraoperative digital subtraction angiography images in the multimodal medical images to obtain standardized images; The first extraction subunit is used to extract features from the catheter data sequence in the multimodal medical image to obtain catheter behavior features; The resampling subunit is used to resample the vascular imaging image in the multimodal medical image to obtain a resampled vascular imaging image. The reconstruction subunit is used to perform three-dimensional reconstruction processing on the resampled vascular imaging image based on the traveling cube algorithm to generate a three-dimensional vascular model of the patient. The second extraction subunit is used to extract anatomical features from the patient's three-dimensional vascular model and generate the patient's vascular anatomical features. A sub-unit is constructed based on the standardized image, the catheter behavior features, and the patient's vascular anatomy features to obtain a preprocessed medical image.
Citation Information
Patent Citations
System and method for intraoperatory determining image alignment
CN115705655A
Real-time data visualization system in neurosurgery
CN120388691A