Language-based hybrid operating room interaction method and apparatus
By acquiring the surgeon's speech-to-text, recognizing intent, and controlling the device to acquire images, combined with a multi-scale embedding layer and a language-guided attention calibration module for surgical target recognition, the problem of multi-task parallel processing in endovascular interventional surgery is solved, achieving efficient device control and target recognition, and improving the consistency and efficiency of surgical operations.
Patent Information
- Application Number
- CN202510509868.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-04-22
AI Technical Summary
In endovascular interventional surgery, doctors face complex and high-intensity tasks, which can lead to a loss of focus, increase the difficulty of the surgery and affect the outcome. Existing single-task assistance methods have failed to effectively solve the problem of multi-task parallel processing.
By acquiring the surgeon's speech-to-text, recognizing surgical intent, controlling surgical equipment to acquire images, and combining image interpretation for surgical target recognition, the system achieves efficient collaboration through multi-scale embedding layers and a language-guided attention calibration module, enabling both equipment control and target recognition.
It reduces the difficulty of surgical procedures, improves the consistency and efficiency of surgical procedures, reduces the workload of doctors, and enhances surgical outcomes.
Smart Images

Figure CN120600024B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a language-based composite operating room interaction method and apparatus. Background Technology
[0002] Vascular diseases are a prevalent type of disease worldwide, seriously threatening human life and health. With the advancement of medical technology, endovascular interventional surgery has become a common and important means of treating vascular diseases due to its significant advantages such as minimal invasiveness and good patient prognosis.
[0003] During endovascular interventional procedures, surgeons face complex and demanding tasks. They must not only precisely deliver various interventional instruments to the lesion site, but also adjust the angiography angle in real time according to the progress of the procedure to obtain clear images of the lesion area. Simultaneously, they must read and analyze the angiographic images in real time to accurately determine the extent of the lesion, the course of the blood vessels, and other crucial information. These tasks are interconnected and mutually influential, requiring surgeons to maintain a high level of concentration simultaneously. However, human energy is finite, and multitasking can easily lead to a dispersion of attention, making it difficult to perform each task perfectly. This not only increases the difficulty of the procedure but may also affect the surgical outcome and even endanger the patient's life. Summary of the Invention
[0004] This invention provides a language-based composite operating room interaction method and device to address the problem in existing technologies where high-intensity and complex surgical procedures are complicated by numerous tasks and tight schedules, while doctors have limited energy, leading to increased surgical difficulty and affecting surgical outcomes. By using language-based surgical target recognition, it can provide assistance to doctors during the surgical process, thereby reducing the difficulty of surgical operations and improving the consistency and efficiency of surgical operations.
[0005] This invention provides a language-based composite operating room interaction method, comprising:
[0006] Obtain the surgical transcript corresponding to the surgeon's speech in a hybrid operating room;
[0007] Based on the surgical transcription text, surgical intent is identified to obtain the surgeon's intent.
[0008] Based on the surgeon's intent, the surgical equipment in the hybrid operating room is controlled to acquire surgical patient images corresponding to the surgeon's intent in the surgical transcription text;
[0009] Based on the surgical patient image, the surgeon's intent, and the surgical transcription text, surgical target identification is performed to obtain surgical target identification results; wherein, the surgical target identification results include surgical site identification results and / or surgical instrument identification results.
[0010] According to a language-based composite operating room interaction method provided by the present invention, the surgeon's intent includes device control intent;
[0011] The step of controlling the surgical equipment in the hybrid operating room to acquire surgical patient images corresponding to the surgeon's intent in the surgical transcription text, based on the surgeon's intent, includes:
[0012] Based on the device control intent and the surgical transcription text, device control instructions are generated, and based on the device control instructions, the surgical equipment in the hybrid operating room is controlled to acquire corresponding surgical patient images.
[0013] The surgical patient image corresponds to the image position indicated by the device control command, or the indicated image position and irradiation angle.
[0014] According to a language-based composite operating room interaction method provided by the present invention, the surgeon's intent also includes an image interpretation intent;
[0015] The surgical target identification based on the surgical patient image, the surgeon's intent, and the surgical transcription text, to obtain the surgical target identification result, includes:
[0016] Based on the image interpretation intent and the surgical transcription text, image interpretation text is generated, and text embedding is performed on the image interpretation text to obtain text features;
[0017] Image block embedding is performed on the surgical patient images to obtain image features of multiple image blocks;
[0018] Based on the text features and the image features of each image block, surgical target recognition is performed to obtain the surgical target recognition result.
[0019] According to a language-based composite operating room interaction method provided by the present invention, the surgical target recognition based on the text features and the image features of each image block to obtain the surgical target recognition result includes:
[0020] Based on the text features and the image features of each image block, spatial similarity matching is performed point by point to obtain a spatial similarity map.
[0021] Surgical target identification is performed based on the spatial similarity map to obtain the surgical target identification result.
[0022] According to a language-based composite operating room interaction method provided by the present invention, the step of surgical target recognition based on the spatial similarity map to obtain surgical target recognition results includes:
[0023] The surgical patient images are feature-encoded to obtain image feature maps, and the image feature maps are then downsampled.
[0024] Based on the downsampled image feature map and the spatial similarity map, surgical target recognition is performed to obtain the surgical target recognition result.
[0025] According to a language-based composite operating room interaction method provided by the present invention, the step of embedding image blocks into the surgical patient image to obtain image features of multiple image blocks includes:
[0026] By using a multi-scale embedding layer, image block embedding is performed on the surgical patient images to obtain image features of multiple image blocks;
[0027] The multi-scale embedding layer contains multiple dilated convolutional layers, each with the same kernel size but different expansion rates.
[0028] According to a language-based composite operating room interaction method provided by the present invention, the step of recognizing surgical intent based on the surgical transcribed text to obtain the surgeon's intent includes:
[0029] Based on the sign language corpus, the surgical transcription text is calibrated with surgical vocabulary to obtain calibrated transcription text;
[0030] Based on the calibrated transcribed text, surgical intent is identified to obtain the surgeon's intent.
[0031] The terminology corpus includes target image positions corresponding to various surgeries and surgical sites, target irradiation angles under each target image position, and names of surgical instruments corresponding to various surgeries and surgical sites.
[0032] The present invention also provides a language-based composite operating room interaction device, comprising:
[0033] The transcription text acquisition unit is used to acquire the surgical transcription text corresponding to the surgeon's speech in the composite operating room;
[0034] A surgical intent recognition unit is used to recognize the surgical intent based on the surgical transcription text to obtain the surgeon's intent.
[0035] The surgical equipment control unit is used to control the surgical equipment in the hybrid operating room to acquire surgical patient images corresponding to the surgeon's intent in the surgical transcription text, based on the surgeon's intent.
[0036] The surgical target recognition unit is used to perform surgical target recognition based on the surgical patient image, the surgeon's intention, and the surgical transcription text, and obtain the surgical target recognition result.
[0037] The surgical target identification results include surgical site identification results and / or surgical instrument identification results.
[0038] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the language-based composite operating room interaction method as described above.
[0039] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the language-based composite operating room interaction method as described above.
[0040] The present invention provides a language-based hybrid operating room interaction method and apparatus, which acquires surgical transcription text corresponding to the surgeon's speech in the hybrid operating room, performs surgical intent recognition based on the surgical transcription text to obtain the surgeon's intent, controls surgical equipment in the hybrid operating room to acquire surgical patient images corresponding to the surgeon's intent in the surgical transcription text, and performs surgical target recognition based on the surgical patient images, the surgeon's intent, and the surgical transcription text to obtain the surgical target recognition result. This overcomes the shortcomings of traditional solutions, which involve numerous tasks in complex surgical processes and limited surgeons' energy, leading to increased surgical difficulty and affecting surgical outcomes. By using the surgeon's intraoperative speech to achieve surgical equipment control and surgical target recognition during the surgical process, it can provide assistance to the surgeon during the operation, thereby reducing the difficulty of the surgical operation and improving the consistency and efficiency of the surgical operation. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0042] Figure 1 This is a flowchart illustrating the language-based composite operating room interaction method provided by the present invention;
[0043] Figure 2 This is a general framework diagram of the language-based composite operating room interaction method provided by the present invention;
[0044] Figure 3 This is a schematic diagram of the structure of the multi-scale embedding layer provided by the present invention;
[0045] Figure 4 This is an example diagram of the contents of the terminology corpus provided by the present invention;
[0046] Figure 5 This is a schematic diagram of the structure of the language-based composite operating room interaction device provided by the present invention;
[0047] Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0049] Vascular diseases, prevalent globally, pose a serious threat to human health. With advancements in medical technology, endovascular interventional surgery, with its minimally invasive approach and favorable patient outcomes, has become a common and crucial method for treating vascular diseases. However, endovascular interventional surgery presents surgeons with complex and demanding tasks. This includes precisely selecting and delivering various interventional instruments to the lesion site, adjusting angiography angles as the procedure progresses to obtain clear images of the lesion area, and simultaneously reading and analyzing the angiographic images in real time to accurately determine the extent of the lesion and the vessel's course. These tasks are interconnected and mutually influential, requiring surgeons to maintain a high level of concentration simultaneously. However, human energy is finite, and multitasking can easily lead to distraction, making it difficult to perform each task perfectly. This not only increases the difficulty of the procedure but can also affect the outcome and even endanger the patient's life.
[0050] In recent years, the application of artificial intelligence (AI) technology in the medical field has become increasingly widespread, bringing new opportunities to solve the aforementioned problems. Currently, some AI methods have been applied to the auxiliary stages of endovascular interventional surgery. For example, lesion measurement technology can quickly and accurately measure parameters such as the size and shape of lesions, providing data support for doctors to formulate surgical plans; vascular segmentation technology can precisely separate blood vessels from complex medical images, helping doctors to observe vascular structures more clearly. However, most of these AI methods are limited to assisting a specific sub-task and lack comprehensive solutions that can efficiently coordinate multiple tasks. Therefore, although these single-task assistance methods improve the execution efficiency of specific tasks to some extent, since doctors need to handle multiple tasks simultaneously during surgery, optimizing a single task does not fundamentally solve the problem of the heavy workload of doctors multitasking.
[0051] Therefore, developing a solution for surgical target recognition in a hybrid operating room that can integrate multiple tasks and achieve efficient collaboration is of great practical significance and has broad application prospects for improving the efficiency of interventional surgery, reducing surgical risks, and improving patient prognosis.
[0052] To address this, the present invention provides a language-based interactive method for hybrid operating rooms, aimed at target recognition during surgical procedures within a hybrid operating room. Figure 1 This is a flowchart illustrating the language-based composite operating room interaction method provided by the present invention, as shown below. Figure 1 As shown, the method includes:
[0053] Step 110: Obtain the surgical transcription text corresponding to the surgeon's speech in the hybrid operating room;
[0054] Step 120: Based on the surgical transcription text, perform surgical intent recognition to obtain the surgeon's intent;
[0055] Step 130: Based on the surgeon's intent, control the surgical equipment in the hybrid operating room to acquire surgical patient images corresponding to the surgeon's intent in the surgical transcription text;
[0056] Step 140: Based on the surgical patient image, the surgeon's intent, and the surgical transcription text, perform surgical target recognition to obtain the surgical target recognition result;
[0057] The surgical target identification results include surgical site identification results and / or surgical instrument identification results.
[0058] Specifically, before performing target identification during the surgical process in the hybrid operating room, i.e., surgical target identification, it is necessary to first determine the object of identification, i.e., the basis for identification. Considering the rich information contained in language, the content of intraoperative communication between doctors can usually reflect the current surgical stage, surgical goals, etc. Therefore, in this embodiment of the invention, the doctor's voice during the surgical process can be used as the basis for identification. That is, the surgeon's voice during the ongoing surgical process in the hybrid operating room can be acquired. Here, the surgeon's voice can be obtained by directly targeting the surgeon's voice, or it can be separated from the surgeon's voice collected from the entire surgical voice in the hybrid operating room. For example, the surgeon's voice can be determined from the mixed surgical voice through methods such as voiceprint separation and role recognition.
[0059] The device used for collecting voice can be a microphone or recorder carried by the surgeon, or a sound pickup component installed in a fixed position in the hybrid operating room, such as a high-definition camera, or a movable audio acquisition device placed in the hybrid operating room, such as a panoramic camera. This embodiment of the invention does not specifically limit the type of device.
[0060] Here, a hybrid operating room refers to a modern surgical space that integrates traditional surgical procedures and interventional procedures. By combining a traditional surgical operating room with advanced imaging diagnostic and interventional treatment equipment, it enables multidisciplinary collaboration and one-stop intraoperative diagnosis and treatment. It is suitable for complex and high-risk surgical scenarios (such as surgical scenarios in the fields of neurology, cardiology, and vascular surgery), and can provide more efficient and precise solutions for complex surgeries, thereby reducing the surgical risks of various complex surgeries and saving surgical time.
[0061] Surgeon voice does not refer to the voice of a specific surgeon, but rather to the voice of surgeons involved in the surgical procedure who perform tasks closely related to the surgery, such as adjusting surgical equipment, interpreting patient images, and controlling surgical instruments. These surgeons may be a single surgeon (the lead surgeon) or multiple surgeons (lead surgeon and assistant surgeons). In cases with multiple surgeons, surgeon voice can include snippets of speech from a single surgeon, as well as conversations between them.
[0062] After obtaining the surgeon's voice, the text content can be obtained from it in this embodiment of the invention, that is, the surgical transcription text corresponding to the surgeon's voice can be obtained. Specifically, the surgeon's voice can be transcribed by voice transcription technology, a locally deployed voice transcription engine, or a large language model, to transcribe the content in the surgeon's voice into text, thereby obtaining the surgical transcription text.
[0063] It should be noted that if the original audio of the surgery is collected from the entire operating room, and if the surgeon's voice has been transcribed during the process of acquiring the voice, such as through voiceprint separation, role recognition, and voice transcription, then the transcribed text can be directly determined based on the marked role. If the voice has not been transcribed, and the surgeon's voice is obtained directly through voiceprint feature recognition and separation, then voice transcription is still required to obtain the transcribed text.
[0064] Furthermore, after obtaining the surgical transcription text, in this embodiment of the invention, surgical intent recognition can be performed based on this surgical transcription text to determine the surgeon's surgical intent, i.e., the surgeon's intent. Since surgeons in a hybrid operating room often face multiple tasks, such as not only precisely delivering surgical instruments but also adjusting imaging angles in real time and interpreting patient images, the surgical intent recognition obtained from the surgical transcription text often includes multiple intents, such as the intent to control surgical equipment, the intent to interpret patient images, and the intent to change surgical instruments.
[0065] Once the surgeon's specific intentions are obtained, in this embodiment of the invention, various specific surgical tasks can be executed according to these intentions. By executing these surgical tasks, target recognition of the surgical process in the composite operating room can be achieved, and surgical target recognition results, such as surgical site recognition results and surgical instrument recognition results, can be obtained. Based on this, medical assistance can be provided to the surgeon (such as delivering surgical instruments and helping to interpret patient images), thereby reducing the surgeon's workload, reducing the difficulty of operation, and improving the surgical effect.
[0066] Specifically, this process can begin by controlling the surgical equipment within the hybrid operating room, based on the surgeon's intent, to acquire images of the patient undergoing surgery. This equipment can be image acquisition devices such as a C-Arm X-ray Equipment or a digital subtraction angiography (DSA) machine. More specifically, based on the surgeon's intent and the surgical transcription text, the surgical equipment can be positioned according to the surgeon's desired image location and irradiation angle to acquire images of the patient, thus obtaining surgical patient images that match the surgeon's intent and the surgical transcription text.
[0067] Furthermore, surgical target identification can be performed based on the surgical patient image to obtain the surgical identification result of the surgical procedure in the composite operating room. Specifically, this can be achieved by combining the surgeon's intent and the surgical transcription text to interpret the surgical patient image and identify surgical targets, such as surgical instruments to be delivered, surgical sites to be segmented from the surgical patient image, etc. In this way, the surgical target identification result of the surgical procedure can be obtained, which may include surgical site identification result and / or surgical instrument identification result.
[0068] Here, the surgical site identification result in the surgical target identification result can be the image of the surgical site that the surgeon needs to specifically view and analyze from the surgical patient's images, such as the vascular segmentation image in vascular interventional surgery, the lesion site image in tumor interventional surgery, etc.; the surgical instrument identification result can be the surgical instruments that the surgeon needs to use in the current stage of the operation in the hybrid operating room, such as guidewires, catheters, stents, balloons, etc.
[0069] In this embodiment of the invention, the surgeon's surgical intent can be identified through the surgeon's voice during the operation, thereby enabling multiple functions such as automatically adjusting the angiography angle, automatically extracting the best exposed blood vessels, and automatically monitoring surgical instruments. It can also provide corresponding assistance according to the surgeon's needs at the current stage of the operation, thereby achieving efficient interaction in the hybrid operating room.
[0070] This invention provides a language-based interactive method for a hybrid operating room. It acquires surgical transcription text corresponding to the surgeon's speech in the hybrid operating room, performs surgical intent recognition based on the transcription text to obtain the surgeon's intent, controls surgical equipment in the hybrid operating room to acquire surgical patient images corresponding to the surgeon's intent in the transcription text, and performs surgical target recognition based on the surgical patient images, surgeon's intent, and the transcription text to obtain the surgical target recognition result. This overcomes the shortcomings of traditional solutions, which involve numerous tasks during complex surgical procedures and limited surgeon energy, leading to increased surgical difficulty and affecting surgical outcomes. By using the surgeon's intraoperative speech to achieve surgical equipment control and surgical target recognition during the surgical process, it can provide assistance to the surgeon during the operation, thereby reducing the difficulty of the surgical operation and improving the consistency and efficiency of the surgical procedure.
[0071] Based on the above embodiments, the surgeon's intent includes the intent to control the device;
[0072] Step 130 includes:
[0073] Based on the device control intent and surgical transcription text, device control instructions are generated, and based on the device control instructions, the surgical equipment in the hybrid operating room is controlled to acquire corresponding surgical patient images.
[0074] Among them, the surgical patient image corresponds to the image position indicated by the equipment control command, or the indicated image position and irradiation angle.
[0075] Specifically, the process of controlling the surgical equipment in the hybrid operating room to acquire surgical patient images corresponding to the surgeon's intent in the surgical transcription text, based on the surgeon's intent, may include:
[0076] When surgical intent is identified based on the surgical transcript and it is determined that the surgeon intends to control the surgical equipment (i.e., the surgeon's intent includes the intent to control the equipment), the surgeon's intraoperative language will be translated into control instructions for the surgical equipment to acquire patient images.
[0077] In detail, this process can involve generating device control instructions based on the device control intent and the surgical transcription text. That is, by combining the device control intent with the surgical transcription text, the surgeon's intraoperative language is translated into instructions to control the surgical equipment, such as controlling the surgical equipment to a certain imaging position and a certain irradiation angle under that position. These device control instructions can then be used to control the surgical equipment in the hybrid operating room to acquire images of the surgical patient. For example, based on the device control instructions, the surgical equipment, such as a C-arm X-ray machine, can be rotated to the imaging position described in the surgeon's intraoperative language, or to the specific irradiation angle under that image position, and then patient images are acquired, such as angiographic images.
[0078] Based on the above embodiments, the surgeon's intent also includes the intent to interpret images;
[0079] Step 140 includes:
[0080] Based on the image interpretation intent and surgical transcription text, image interpretation text is generated, and text embedding is performed on the image interpretation text to obtain text features;
[0081] Image block embedding is performed on surgical patient images to obtain image features of multiple image blocks;
[0082] Surgical target recognition results are obtained by recognizing surgical targets based on text features and image features of each image block.
[0083] Specifically, the process of identifying the surgical target based on the surgical patient's image, the surgeon's intent, and the transcribed surgical text to obtain the surgical target identification result includes:
[0084] When surgical intent is identified based on the surgical transcription text, and it is determined that the surgeon intends to interpret the patient's images (i.e., the surgeon's intent includes the intent to interpret the images), the surgeon's intraoperative language is converted into prior language information to activate relevant parts of the surgical patient image during the surgical target recognition process, ultimately yielding the surgical target recognition result.
[0085] In detail, this could involve generating image interpretation instructions based on the intended interpretation of the images and the surgical transcription text. That is, by combining the intended interpretation with the surgical transcription text, the surgeon's intraoperative language is converted into image interpretation text, which is then used to interpret the surgical patient's images. Specifically, the image interpretation text can be a type of language prior information. Using this language prior as input allows the surgical target recognition task based on the surgical patient's images to rely more on intraoperative language rather than visual information.
[0086] After generating the image interpretation text, in this embodiment of the invention, the LGAC Module (Language-Guided Attention Calibrated Module) can be introduced to interpret the surgical patient images and finally output the surgical target recognition results. Figure 2 This is a general framework diagram of the language-based composite operating room interaction method provided by the present invention, as shown below. Figure 2 As shown, after obtaining the surgical patient image (Digital subtraction angiography image, DSA image) and the image interpretation text (Language Prior), text embedding can be performed on the image interpretation text to obtain text features. Specifically, this can be achieved by embedding language features into the input language prior information using a parameter-frozen text encoder (Contrastive Language-Image Pre-training Encoder, CLIP Encoder), thereby obtaining language features. The encoded language features can then be input into the adapter module for processing, so that it can better adapt to the surgical target recognition task and obtain text features.
[0087] Meanwhile, the input surgical patient images can be embedded using the MPE Layer (Multi-Patch Embedding Layer), which uses convolutional layers to embed the images into patches, thereby obtaining the image features of multiple patches.
[0088] Following this, guided by text features, surgical target identification can be performed using the image features of each image block to identify surgical targets, such as surgical instruments to be delivered, images of surgical sites to be segmented from surgical patient images, etc., thus obtaining the surgical target identification result of the surgical process.
[0089] Based on the above embodiments, surgical target recognition is performed based on text features and image features of each image block to obtain surgical target recognition results, including:
[0090] Based on text features and image features of each image block, spatial similarity matching is performed point by point to obtain a spatial similarity map.
[0091] Surgical target identification is performed based on spatial similarity maps to obtain surgical target identification results.
[0092] Specifically, the process of identifying the surgical target based on text features and the image features of each image block, and obtaining the surgical target identification result, includes the following steps:
[0093] After obtaining the text features, when combining them with image patch features for surgical target recognition, spatial similarity matching can be performed first using the text features and the image features of each image patch. This calculates the similarity between the text features and the image patch features at each point in space, resulting in a spatial similarity map. Specifically, this can be achieved by calculating the inner product of the image patch features and the text features point by point in space. Afterward, this spatial similarity map can be decoded to identify the surgical target, thus obtaining the surgical target recognition result.
[0094] Based on the above embodiments, surgical target recognition is performed based on spatial similarity maps to obtain surgical target recognition results, including:
[0095] The surgical patient images are feature-encoded to obtain image feature maps, and the image feature maps are then downsampled.
[0096] Surgical target recognition results are obtained by using the downsampled image feature map and spatial similarity map.
[0097] Specifically, when using spatial similarity maps for surgical target recognition, the spatial similarity map can be fused with the downsampled image feature map. The fusion method can be to multiply the two in the spatial dimension to obtain a fused feature map. Then, the fused feature map can be decoded by an image decoder to obtain the surgical target recognition result.
[0098] The image feature map is a feature map obtained by feature encoding the surgical patient image. For example, it can be a feature map obtained by inputting the surgical patient image into any existing segmentation backbone network. In the embodiment of the present invention, the segmentation backbone network will remove the last output layer on the basis of the existing structure.
[0099] In this embodiment of the invention, the proposed language-guided attention calibration module can extract regions highly correlated with language prior information from the input surgical patient image and calibrate the image features accordingly, thereby greatly improving the ability to identify surgical targets and ensuring the accuracy and reliability of surgical target identification. At the same time, the proposed multi-scale embedding layer can extract and fuse image patch features of different scales from the input surgical patient image, which can improve the effect of cross-modal information fusion.
[0100] Based on the above embodiments, image block embedding is performed on the surgical patient images to obtain image features of multiple image blocks, including:
[0101] By using a multi-scale embedding layer, image block embedding is performed on surgical patient images to obtain image features of multiple image blocks;
[0102] The multi-scale embedding layer contains multiple dilated convolutional layers, each with the same kernel size but different dilation rates.
[0103] Specifically, Figure 3 This is a schematic diagram of the structure of the multi-scale embedding layer provided by the present invention, as shown below. Figure 3 As shown, when embedding surgical patient images into image blocks using a multi-scale embedding layer, it is considered that using convolutions of a fixed size for image block embedding would lead to a lack of information exchange between image blocks, and the information that can be captured within each image block is limited. Therefore, in this embodiment of the invention, dilated convolutions of the same size but different dilation rates are proposed to achieve multi-scale image block embedding. That is, multiple spatial convolutional layers in the multi-scale embedding layer are used to embed surgical patient images into image blocks, obtaining image features of multiple image blocks. This not only allows information from different receptive fields to be fused within each image block, but also enables information exchange between adjacent image blocks, further improving the information fusion effect.
[0104] In each of the hollow convolutional layers, the size of the convolutional kernel is the same, but the dilation rate is different. That is, the size of the convolutional kernel is "K×K", and the value of "K" can be set according to actual needs, but the dilation rate is 1, 2 and 5 respectively.
[0105] Based on the above embodiments, step 120 includes:
[0106] Based on a sign language corpus, surgical vocabulary is calibrated on the surgical transcription text to obtain calibrated transcription text.
[0107] Surgical intent is identified based on calibrated transcribed text to obtain the surgeon's intent.
[0108] The manual corpus includes target image positions corresponding to various surgeries and surgical sites, target irradiation angles under each target image position, and names of surgical instruments corresponding to various surgeries and surgical sites.
[0109] Specifically, the process of obtaining the surgeon's intent by recognizing surgical intent based on the surgical transcription text can include:
[0110] Considering that surgeons have many tasks and limited time during surgery, they often use short words and terms familiar to them or used in their field to express their needs during intraoperative communication. In this case, the surgeon's intent obtained by directly identifying the intent based on the surgical transcription text may be incorrect. In order to achieve accurate intent identification and thus provide precise assistance to the surgeon during the operation, in this embodiment of the invention, before surgical intent identification, surgical vocabulary calibration can be performed on the surgical transcription text, such as typo correction, abbreviation completion, terminology explanation, etc., to obtain calibrated transcription text.
[0111] The surgical terminology calibration here can be achieved using a pre-built surgical terminology corpus. This corpus contains numerous professional surgical terms, such as the target imaging positions corresponding to various surgeries and surgical sites, the target irradiation angles under each target imaging position, and the names of surgical instruments corresponding to various surgeries and surgical sites. Among them, the target imaging position can be several (e.g., 6) commonly used imaging positions corresponding to the corresponding surgery and surgical site, and the target irradiation angle can be the optimal irradiation angle (exposure angle) under the corresponding imaging position.
[0112] Figure 4 This is an example diagram of the contents of the terminology corpus provided by the present invention, such as... Figure 4 As shown, when the surgery is an interventional procedure involving blood vessels, the pre-established surgical corpus can contain two parts of information. One part is the surgical angiography position, that is, it can contain several commonly used blood vessel angiography positions. Specifically, this can include the full name and abbreviation of the angiography position and the optimal exposure angle, such as RAO (Right Anterior Oblique) 30° + CAU (Caudal) 20° and CAU 20° for LCX (Left Circumflex Branch); CRA (Cranial) 30° and LAO (Left Anterior Oblique) 45° + CRA 20° and RAO 30° + CRA 20° for LAD (Left Anterior Descending Branch); and CRA 30° and LAO (Right Coronary Artery) for RCA (Right Coronary Artery). 30°, etc.; the other part is the name of the surgical instrument, such as Balloon, Guidewire, etc., and the standard names of the surgical instruments used can be stored in the surgical terminology corpus.
[0113] In addition, it should be noted that the pre-established medical terminology corpus and intraoperative interaction mode proposed in the embodiments of the present invention can allow surgeons to provide input in the form of voice during surgery, thereby enabling different intraoperative auxiliary functions, including surgical equipment control and surgical target recognition.
[0114] Furthermore, the method provided in this invention has achieved state-of-the-art performance after being tested on two complex coronary vessel datasets and multi-instrument datasets. It is expected to enable efficient interaction between doctors and robots or intelligent systems in future smart hybrid operating rooms, thereby improving the consistency and efficiency of surgical operations and promoting the development of interventional surgery.
[0115] The following describes the language-based hybrid operating room interaction device provided by the present invention. The language-based hybrid operating room interaction device described below can be referred to in correspondence with the language-based hybrid operating room interaction method described above.
[0116] Figure 5 This is a schematic diagram of the structure of the language-based composite operating room interaction device provided by the present invention, as shown below. Figure 5 As shown, the device includes:
[0117] The transcription text acquisition unit 510 is used to acquire the surgical transcription text corresponding to the surgeon's speech in the composite operating room.
[0118] The surgical intent recognition unit 520 is used to recognize the surgical intent based on the surgical transcription text to obtain the surgeon's intent.
[0119] The surgical equipment control unit 530 is used to control the surgical equipment in the hybrid operating room to acquire surgical patient images corresponding to the surgeon's intent in the surgical transcription text based on the surgeon's intent.
[0120] The surgical target recognition unit 540 is used to perform surgical target recognition based on the surgical patient image, the surgeon's intention, and the surgical transcription text to obtain surgical target recognition results; wherein, the surgical target recognition results include surgical site recognition results and / or surgical instrument recognition results.
[0121] The present invention provides a language-based hybrid operating room interaction device that acquires surgical transcription text corresponding to the surgeon's speech in the hybrid operating room, performs surgical intent recognition based on the surgical transcription text to obtain the surgeon's intent, controls the surgical equipment in the hybrid operating room to acquire surgical patient images corresponding to the surgeon's intent in the surgical transcription text, and performs surgical target recognition based on the surgical patient images, surgeon's intent, and surgical transcription text to obtain the surgical target recognition result. This overcomes the shortcomings of traditional solutions, which involve numerous tasks in complex surgical processes and limited surgeon's energy, leading to increased surgical difficulty and affecting surgical outcomes. By using the surgeon's intraoperative speech, the device enables control of surgical equipment and recognition of surgical targets during the surgical process, thereby providing assistance to the surgeon during the operation, reducing the difficulty of the operation, and improving the consistency and efficiency of the operation.
[0122] Based on the above embodiments, the surgeon's intent includes the intent to control the device;
[0123] The surgical equipment control unit 530 is used for:
[0124] Based on the device control intent and the surgical transcription text, device control instructions are generated, and based on the device control instructions, the surgical equipment in the hybrid operating room is controlled to acquire corresponding surgical patient images.
[0125] The surgical patient image corresponds to the image position indicated by the device control command, or the indicated image position and irradiation angle.
[0126] Based on the above embodiments, the surgeon's intent also includes the intent to interpret images;
[0127] The surgical target recognition unit 540 is used for:
[0128] Based on the image interpretation intent and the surgical transcription text, image interpretation text is generated, and text embedding is performed on the image interpretation text to obtain text features;
[0129] Image block embedding is performed on the surgical patient images to obtain image features of multiple image blocks;
[0130] Based on the text features and the image features of each image block, surgical target recognition is performed to obtain the surgical target recognition result.
[0131] Based on the above embodiments, the surgical target recognition unit 540 is used for:
[0132] Based on the text features and the image features of each image block, spatial similarity matching is performed point by point to obtain a spatial similarity map.
[0133] Surgical target identification is performed based on the spatial similarity map to obtain the surgical target identification result.
[0134] Based on the above embodiments, the surgical target recognition unit 540 is used for:
[0135] The surgical patient images are feature-encoded to obtain image feature maps, and the image feature maps are then downsampled.
[0136] Based on the downsampled image feature map and the spatial similarity map, surgical target recognition is performed to obtain the surgical target recognition result.
[0137] Based on the above embodiments, the surgical target recognition unit 540 is used for:
[0138] By using a multi-scale embedding layer, image block embedding is performed on the surgical patient images to obtain image features of multiple image blocks;
[0139] The multi-scale embedding layer contains multiple dilated convolutional layers, each with the same kernel size but different expansion rates.
[0140] Based on the above embodiments, the surgical intent recognition unit 520 is used for:
[0141] Based on the sign language corpus, the surgical transcription text is calibrated with surgical vocabulary to obtain calibrated transcription text;
[0142] Based on the calibrated transcribed text, surgical intent is identified to obtain the surgeon's intent.
[0143] The terminology corpus includes target image positions corresponding to various surgeries and surgical sites, target irradiation angles under each target image position, and names of surgical instruments corresponding to various surgeries and surgical sites.
[0144] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6As shown, the electronic device may include a processor 610, a communications interface 620, a memory 630, and a communication bus 640, wherein the processor 610, communications interface 620, and memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute a language-based hybrid operating room interaction method. This method includes: acquiring surgical transcription text corresponding to the surgeon's speech in the hybrid operating room; performing surgical intent recognition based on the surgical transcription text to obtain the surgeon's intent; controlling surgical equipment in the hybrid operating room to acquire surgical patient images corresponding to the surgeon's intent in the surgical transcription text based on the surgeon's intent; and performing surgical target recognition based on the surgical patient images, the surgeon's intent, and the surgical transcription text to obtain a surgical target recognition result; wherein the surgical target recognition result includes surgical site recognition results and / or surgical instrument recognition results.
[0145] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0146] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein when the program instructions are executed by a computer, the computer is able to execute the language-based hybrid operating room interaction method provided by the above methods, the method comprising: acquiring surgical transcription text corresponding to the surgeon's speech in the hybrid operating room; performing surgical intent recognition based on the surgical transcription text to obtain the surgeon's intent; controlling the surgical equipment in the hybrid operating room to acquire surgical patient images corresponding to the surgeon's intent in the surgical transcription text based on the surgeon's intent; performing surgical target recognition based on the surgical patient images, the surgeon's intent, and the surgical transcription text to obtain surgical target recognition results; wherein the surgical target recognition results include surgical site recognition results and / or surgical instrument recognition results.
[0147] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the language-based hybrid operating room interaction method provided by the methods described above. This method includes: acquiring surgical transcription text corresponding to the surgeon's speech in the hybrid operating room; performing surgical intent recognition based on the surgical transcription text to obtain the surgeon's intent; controlling surgical equipment in the hybrid operating room to acquire surgical patient images corresponding to the surgeon's intent in the surgical transcription text based on the surgeon's intent; and performing surgical target recognition based on the surgical patient images, the surgeon's intent, and the surgical transcription text to obtain a surgical target recognition result; wherein the surgical target recognition result includes surgical site recognition result and / or surgical instrument recognition result.
[0148] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0149] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A language-based composite operating room interaction method, characterized in that, include: Obtain the surgical transcript corresponding to the surgeon's speech in a hybrid operating room; Based on the surgical transcription text, surgical intent is identified to obtain the surgeon's intent. Based on the surgeon's intent, the surgical equipment in the hybrid operating room is controlled to acquire surgical patient images corresponding to the surgeon's intent in the surgical transcription text; Based on the surgical patient image, the surgeon's intent, and the surgical transcription text, surgical target recognition is performed to obtain surgical target recognition results; wherein, the surgical target recognition results include surgical site recognition results and / or surgical instrument recognition results; The surgeon's intent includes the intent to interpret the images; The surgical target identification based on the surgical patient image, the surgeon's intent, and the surgical transcription text, to obtain the surgical target identification result, includes: Based on the image interpretation intent and the surgical transcription text, image interpretation text is generated, and text embedding is performed on the image interpretation text to obtain text features; Image block embedding is performed on the surgical patient images to obtain image features of multiple image blocks; Based on the text features and the image features of each image block, surgical target recognition is performed to obtain the surgical target recognition result.
2. The language-based composite operating room interaction method according to claim 1, characterized in that, The surgeon's intent also includes the intent to control the equipment; The step of controlling the surgical equipment in the hybrid operating room to acquire surgical patient images corresponding to the surgeon's intent in the surgical transcription text, based on the surgeon's intent, includes: Based on the device control intent and the surgical transcription text, device control instructions are generated, and based on the device control instructions, the surgical equipment in the hybrid operating room is controlled to acquire corresponding surgical patient images. The surgical patient image corresponds to the image position indicated by the device control command, or corresponds to the image position and irradiation angle indicated by the device control command.
3. The language-based composite operating room interaction method according to claim 1, characterized in that, The surgical target recognition based on the text features and the image features of each image block, to obtain the surgical target recognition result, includes: Based on the text features and the image features of each image block, spatial similarity matching is performed point by point to obtain a spatial similarity map. Surgical target identification is performed based on the spatial similarity map to obtain the surgical target identification result.
4. The language-based composite operating room interaction method according to claim 3, characterized in that, The surgical target identification based on the spatial similarity map, to obtain the surgical target identification result, includes: The surgical patient images are feature-encoded to obtain image feature maps, and the image feature maps are then downsampled. Based on the downsampled image feature map and the spatial similarity map, surgical target recognition is performed to obtain the surgical target recognition result.
5. The language-based composite operating room interaction method according to any one of claims 1 to 4, characterized in that, The image block embedding of the surgical patient images yields image features of multiple image blocks, including: By using a multi-scale embedding layer, image block embedding is performed on the surgical patient images to obtain image features of multiple image blocks; The multi-scale embedding layer contains multiple dilated convolutional layers, each with the same kernel size but different expansion rates.
6. The language-based composite operating room interaction method according to any one of claims 1 to 4, characterized in that, The process of recognizing surgical intent based on the transcribed surgical text to obtain the surgeon's intent includes: Based on the sign language corpus, the surgical transcription text is calibrated with surgical vocabulary to obtain calibrated transcription text; Based on the calibrated transcribed text, surgical intent is identified to obtain the surgeon's intent. The terminology corpus includes target image positions corresponding to various surgeries and surgical sites, target irradiation angles under each target image position, and names of surgical instruments corresponding to various surgeries and surgical sites.
7. A voice-based interactive device for operating rooms, characterized in that, include: The transcription text acquisition unit is used to acquire the surgical transcription text corresponding to the surgeon's speech in the composite operating room; A surgical intent recognition unit is used to recognize the surgical intent based on the surgical transcription text to obtain the surgeon's intent. The surgical equipment control unit is used to control the surgical equipment in the hybrid operating room to acquire surgical patient images corresponding to the surgeon's intent in the surgical transcription text, based on the surgeon's intent. The surgical target recognition unit is used to perform surgical target recognition based on the surgical patient image, the surgeon's intention, and the surgical transcription text to obtain surgical target recognition results; wherein, the surgical target recognition results include surgical site recognition results and / or surgical instrument recognition results; The surgeon's intent includes the intent to interpret the images; The surgical target identification based on the surgical patient image, the surgeon's intent, and the surgical transcription text, to obtain the surgical target identification result, includes: Based on the image interpretation intent and the surgical transcription text, image interpretation text is generated, and text embedding is performed on the image interpretation text to obtain text features; Image block embedding is performed on the surgical patient images to obtain image features of multiple image blocks; Based on the text features and the image features of each image block, surgical target recognition is performed to obtain the surgical target recognition result.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the language-based composite operating room interaction method as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the language-based composite operating room interaction method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Interaction method of endoscope auxiliary software
CN114360524A
Blood vessel key point identification method based on voice interaction and multi-source information fusion
CN118053184A