Remote surgical image transmission method, device and surgical robot system

By acquiring display images and multimodal data during remote surgery, configuring identification strategies and determining transmission strategies, the problem of inability to ensure image quality in key areas under limited bandwidth in traditional technology is solved, and high-definition sunlight transmission under bandwidth-limited conditions is achieved, improving the visual accuracy and real-timeness of the surgery.

CN119925000BActive Publication Date: 2025-06-17SHANGHAI MICROPORT MEDBOT (GRP) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510412784.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-17
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

Traditional remote surgical image transmission cannot guarantee the image quality of key areas under limited bandwidth, resulting in low clarity.

Method used

By obtaining the displayed images and multimodal data in the user's in-service, configuring the identification strategy, processing the displayed images to obtain detection results, including the region of interest and its corresponding region parameters, and determining the transmission strategy based on the region parameters and current bandwidth conditions, ensuring high-definition transmission of key areas.

Benefits of technology

In the case of limited bandwidth, high-definition sunlight transmission in key areas is ensured, and visual accuracy on the doctor's control end and real-time operation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119925000B_ABST
    Figure CN119925000B_ABST
Patent Text Reader

Abstract

The present invention provides a remote surgical image transmission method, device and surgical robot system, which relate to the technical field of medical equipment. The present application obtains the user's intraoperative display image and multimodal data in real time; wherein the multimodal data includes the spatial information of the instrument, the surgical procedure information and the process information. According to the multimodal data, the recognition strategy is dynamically configured to select the appropriate recognition strategy in the current scene, so as to perform targeted detection on the displayed image and accurately identify the area of ​​interest that meets the real-time change requirements of the surgery. Furthermore, the present application dynamically configures the transmission strategy according to the identified regional parameters and the current bandwidth conditions, so as to ensure high-definition transmission of key areas under limited bandwidth, thereby ensuring the visual accuracy of the doctor's control terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical equipment, and in particular to a remote surgical image transmission method, device and surgical robot system. Background Art

[0002] Remote surgery is a medical operation mode based on network communication and robotic technology. Its core feature is that the doctor and the patient are in different geographical locations. The doctor remotely controls the robotic equipment to complete the operation by transmitting surgical instructions and image data in real time. During the operation, surgical instruments and anatomical areas are the key areas of concern for doctors, and the image clarity and real-time performance of these key areas need to be guaranteed. In traditional technologies, remote surgical image transmission has problems such as limited bandwidth, transmission delay, and image quality degradation, resulting in low clarity in key areas.

[0003] Therefore, how to prioritize the image quality of key areas within limited bandwidth resources has become an urgent problem to be solved. Summary of the invention

[0004] The purpose of the present invention is to provide a remote surgical image transmission method, device and surgical robot system to overcome the defect of traditional technology that the image quality of key areas cannot be guaranteed within limited bandwidth resources.

[0005] In a first aspect, the present application provides a remote surgery image transmission method, comprising:

[0006] Acquire the display image and multimodal data of the user during the operation; wherein the multimodal data includes the spatial information of the instrument, the operation procedure information and the process information;

[0007] According to the multimodal data, a recognition strategy is configured, and based on the recognition strategy, the display image is processed to obtain a detection result, wherein the detection result includes a region of interest and its corresponding region parameters;

[0008] A transmission strategy is determined according to the regional parameters and the current bandwidth situation, and based on the transmission strategy, the display image is transmitted to the doctor control terminal.

[0009] In one embodiment, configuring a recognition strategy according to the multimodal data includes:

[0010] Encoding the multimodal data to obtain multimodal features;

[0011] Analyzing the semantic relationship between the multimodal features based on the large language model to generate a semantically consistent comprehensive representation;

[0012] Determining the activation priority of each identifier according to the comprehensive representation;

[0013] The recognition strategy is obtained by selecting a recognition device whose activation priority meets the preset requirements from a preset module library, and configuring a recognition mechanism for each selected recognition device.

[0014] In one of the embodiments, the multimodal data also includes intraoperative abnormality information;

[0015] The configuring the recognition strategy according to the multimodal data further includes:

[0016] In the case where abnormal information is detected during the operation, the activation priority of each identifier is adjusted according to the abnormal information to adjust the current recognition strategy.

[0017] In one embodiment, the step of processing the displayed image to obtain a detection result based on the recognition strategy includes:

[0018] Using a pre-trained spatiotemporal feature extraction model to extract features from a single frame of the display image to obtain spatiotemporal features; wherein the spatiotemporal feature extraction model is trained based on a deep learning model;

[0019] The identifier determines the position of each target area in the displayed image according to the spatiotemporal characteristics, and evaluates the priority of each target area according to a preset rule to obtain a region of interest and its corresponding region parameters.

[0020] In one embodiment, the processing of the display image to obtain a detection result based on the recognition strategy further includes:

[0021] After determining the position of each target area in the display image, the edges of adjacent target areas are smoothed by a potential field method, and the smoothed target area is used as a region of interest and its region parameters are output.

[0022] In one embodiment, the processing of the display image to obtain a detection result based on the recognition strategy further includes:

[0023] After determining the position of each target region in the display image, the edges of the adjacent multiple target regions are smoothed by using a region growing operator, and the smoothed target regions are used as regions of interest and their region parameters are output.

[0024] In one embodiment, configuring the recognition strategy according to the multimodal data further includes:

[0025] After selecting recognizers whose activation priorities meet preset requirements from a preset module library, the retainable recognizers are determined from the selected multiple recognizers based on the confidence obtained by pre-training of each of the recognizers and a weighted decision method.

[0026] In one embodiment, the method further comprises:

[0027] After the test results are obtained, an assessment is performed based on the confidence level of the test results, and a warning signal is generated when the assessment score is lower than a preset score.

[0028] In one of the embodiments, the region parameters include size information and priority;

[0029] Determining a transmission strategy according to the regional parameters and the current bandwidth situation, and transmitting the display image to the doctor control terminal based on the transmission strategy, includes:

[0030] According to the region parameters, configure encoding parameters of the region of interest and the background region, wherein the encoding parameters include at least one of a compression rate, a bit rate, a quantization parameter, a frame rate, and a resolution;

[0031] Select the corresponding transmission protocol according to the current bandwidth situation;

[0032] The display image is encoded based on the encoding parameters, and the encoded multiple data packets are transmitted to the doctor control terminal through the selected transmission protocol.

[0033] In one embodiment, the method further comprises:

[0034] Use pre-trained prediction models to predict the bandwidth fluctuation trend of the network within a preset time;

[0035] According to the bandwidth fluctuation trend and the regional parameters, adjusting encoding parameters before transmission, and allocating bandwidth for each of the data packets;

[0036] The prediction model is obtained based on historical bandwidth data and machine learning algorithm training.

[0037] In one embodiment, the method further comprises:

[0038] After transmitting a plurality of data packets to the doctor control end, the encoding parameters and the transmission protocol are adjusted according to feedback information from the doctor control end to ensure the clarity of each region of interest.

[0039] In a second aspect, the present application provides a remote surgery image transmission device, the device comprising:

[0040] A data acquisition module, used to acquire the display image and multimodal data of the user during the operation; wherein the multimodal data includes the spatial information of the instrument, the operation procedure information and the process information;

[0041] A processing module, configured to configure a recognition strategy according to the multimodal data, and based on the recognition strategy, process the display image to obtain a detection result, wherein the detection result includes a region of interest and its corresponding region parameters;

[0042] The transmission module is used to determine a transmission strategy according to the regional parameters and the current bandwidth situation, and transmit the display image to the doctor control terminal based on the transmission strategy.

[0043] In a third aspect, the present application also provides a surgical robot system, the system comprising:

[0044] The control console is used to provide a control platform for doctors and generate control instructions according to the doctors' operations;

[0045] The surgical robot is used to perform surgery according to the control instructions, and adopt the remote surgical image transmission method described in any one of the first aspects to transmit the display images obtained during the operation to the control console.

[0046] The above-mentioned remote surgical image transmission method, device and surgical robot system have at least the following advantages:

[0047] This application obtains the user's intraoperative display images and multimodal data in real time; wherein, the multimodal data includes the spatial information of the instrument, the surgical procedure information and the process information. Based on the multimodal data, the recognition strategy is dynamically configured to select the appropriate recognition strategy in the current scenario, so as to conduct targeted detection of the displayed image and accurately identify the area of ​​interest that meets the real-time changing needs of the surgery. Furthermore, this application dynamically configures the transmission strategy based on the identified regional parameters and the current bandwidth conditions, to ensure high-definition transmission of key areas under limited bandwidth, thereby ensuring the visual accuracy of the doctor's control terminal. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a structural block diagram of a surgical robot system in one embodiment;

[0049] Figure 2 A schematic diagram of a flow chart of a remote surgery image transmission method in one embodiment;

[0050] Figure 3 A flowchart of configuring identification strategy steps in one embodiment;

[0051] Figure 4 A schematic diagram of a process for obtaining a test result in one embodiment;

[0052] Figure 5 is a regional schematic diagram of a potential energy field in one embodiment;

[0053] Figure 6is a schematic diagram of a process of region growing in one embodiment;

[0054] Figure 7 A schematic diagram of a flow chart of steps for determining a transmission strategy and transmitting a display image in one embodiment;

[0055] Figure 8 1 is a structural block diagram of a remote surgery image transmission device in one embodiment. DETAILED DESCRIPTION

[0056] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0057] While some exemplary embodiments of the present invention have been described for the purpose of illustration, it should be understood that the present invention may be implemented in other ways not specifically shown in the drawings.

[0058] See also Figure 1 In an exemplary embodiment, the present application provides a surgical robot system, including: a control console and a surgical robot. The surgical robot is located at the local side, and the control console is located at the remote side. The local side refers to the side where the patient is located, and the remote side refers to the side where the doctor is located.

[0059] The control console is used to provide a control platform and generate control instructions based on the doctor's operations.

[0060] A surgical robot is used to perform surgery according to control instructions and transmit the display images obtained during the operation to the control console.

[0061] Specifically, the surgical robot performs the operation based on the doctor's control instructions on the remote side. During the operation, the surgical robot obtains the user's display image and multimodal data in real time; among which, the multimodal data includes the spatial information of the instrument, the surgical procedure information and the process information. The surgical robot also configures the recognition strategy based on the above multimodal data, and based on the recognition strategy, processes the display image to obtain the detection result. The detection result includes the region of interest (ROI) and its corresponding regional parameters. According to the regional parameters and the current bandwidth situation, the transmission strategy is determined, and based on the transmission strategy, the display image is transmitted to the cloud server through the network.

[0062] The remote side obtains the display image from the cloud server through the network, transmits it to the stereo monitor on the control console through optical fiber, and displays it to the doctor.

[0063] In the above-mentioned surgical robot system, the surgical robot obtains the user's intraoperative display images and multimodal data in real time; wherein the multimodal data includes the spatial information of the instrument, the surgical procedure information and the process information. Based on the multimodal data, the recognition strategy is dynamically configured to select the appropriate recognition strategy in the current scenario, so as to conduct targeted detection of the displayed image and accurately identify the area of ​​interest that meets the real-time changing needs of the surgery. Furthermore, the present application dynamically configures the transmission strategy based on the identified regional parameters and the current bandwidth conditions, so as to ensure high-definition transmission of key areas under limited bandwidth, thereby ensuring the visual accuracy of the doctor's control terminal.

[0064] In an exemplary embodiment, the present application provides a remote surgery image transmission method. Figure 1 This is explained using the surgical robot in the example.

[0065] See also Figure 2 , Figure 2 The figure is a flowchart of a remote surgery image transmission method of this embodiment, which specifically includes the following steps:

[0066] Step S202, obtaining the display image and multimodal data of the user during the operation; wherein the multimodal data includes the spatial information of the instrument, the operation procedure information and the process information.

[0067] Specifically, the images displayed during the operation include endoscopic images, which are collected in real time by the endoscope configured by the surgical robot and displayed to the doctor's control terminal. Since the operation requires precise operation, the endoscopic images need to be transmitted to the doctor's control terminal in a low-latency manner. Furthermore, the above-mentioned displayed images can also include infrared fluorescence imaging, 3D stereoscopic vision images, or intraoperative ultrasound images, which can help doctors understand the actual situation of the patient in real time and help doctors perform operations more accurately.

[0068] Furthermore, multimodal data refers to data in multiple dimensions of a patient during surgery.

[0069] The spatial information of the instrument is the spatial variation of the instrument used in this operation, reflecting the position and posture changes of the instrument in the workspace, including translation data, rotation data, three-dimensional position, posture and movement speed. The encoder unifies it into the feature dimension of the language model.

[0070] The surgical procedure information is the category of the surgery, such as prostate surgery, pancreatic surgery, etc., which is converted into language features and used as the surgical category feature data of this surgery.

[0071] The progress information is obtained based on the relevant images of the surgery, such as the intraoperative visual field map, and is used to infer the sub-stage of the current surgery. Exemplarily, the embodiment of the present application performs dual-channel processing through the visual model DinoV2 and the SigLIP module, and extracts visual features from the image to obtain the above-mentioned progress information, wherein the model DinoV2 is used for visual space representation, and the model SigLIP is used for visual texture representation. Furthermore, through the above-mentioned visual space representation and visual texture representation, the surgical instruments can be tracked in real time and key anatomical tissues can be identified, so as to distinguish the sub-stage of the surgery at the current moment. Exemplarily, for radical prostatectomy, it usually includes operation stages such as nerve stripping and vascular stripping, and the above-mentioned nerve stripping and vascular stripping are the sub-stages of the surgery.

[0072] Step S204, configuring a recognition strategy according to the multimodal data, and based on the recognition strategy, processing the displayed image to obtain a detection result, the detection result including the region of interest and its corresponding region parameters.

[0073] Specifically, in the above multimodal data, the spatial relationship between the surgical instrument and the target tissue can be determined according to the spatial information of the instrument; the key anatomical structure can be screened according to the procedure information and the preset surgical process template; the area of ​​interest can be dynamically adjusted according to the process information. Based on the above information, the key operation area of ​​the operation at the current moment can be inferred, so as to determine the identifier corresponding to each key operation area, realize targeted detection of the key area, and provide doctors with high-precision images. The combination of the above multiple identifiers and the recognition mechanism of each identifier are the recognition strategy. It should be understood that different identifiers are independent dedicated algorithms. For example, the target detector can adopt the YOLO series to quickly locate the tip of the surgical instrument. The fine segmentation model can adopt the U-Net algorithm for pixel-level segmentation of blood vessels, nerves or diseased tissues. Each identifier has differences in accuracy, speed, and applicable scenarios, and needs to be intelligently selected according to the real-time context. In this embodiment, the relevant identifiers are stored in the module library in advance for selection. In practical applications, new algorithms can be dynamically loaded as needed to achieve modular expansion.

[0074] Furthermore, based on the above recognition strategy, different identifiers are used to identify different areas of the displayed image to obtain multiple ROI areas and their corresponding area parameters. The area parameters include size information and priority. The size information includes bounding box, pixel area, relative area ratio and semantic mask; the identifier assigns a priority to each ROI area, which is used to determine the subsequent encoding quality, transmission order and bandwidth allocation. Furthermore, the area parameters also include the confidence of the ROI area, which is used to indicate the reliability of the result.

[0075] Step S206, determining a transmission strategy according to the regional parameters and the current bandwidth situation, and transmitting the display image to the doctor control terminal based on the transmission strategy.

[0076] Specifically, in a remote surgery environment, the embodiment of the present application dynamically adjusts the transmission strategy according to the bandwidth situation. For example, when the bandwidth is sufficient, the complete display image can be transmitted to show the doctor global information. In the case of limited bandwidth and multiple ROI areas, the transmission strategy can be determined based on priority and size; for example, each ROI area is independently encoded and transmitted, and priority is given to transmission while ensuring the compression quality of high-priority and large-size ROI areas, while reducing the compression quality of background data or extending the transmission to ensure visual clarity at the doctor's control end.

[0077] The above-mentioned remote surgical image transmission method obtains the user's intraoperative display images and multimodal data in real time; wherein the multimodal data includes the spatial information of the instrument, the surgical procedure information and the process information. Based on the multimodal data, the recognition strategy is dynamically configured to select the appropriate recognition strategy in the current scenario, so as to perform targeted detection on the displayed image and accurately identify the ROI area that meets the real-time change requirements of the surgery. Furthermore, the present application dynamically configures the transmission strategy based on the identified regional parameters and the current bandwidth conditions, so as to ensure high-definition transmission of key areas under limited bandwidth and ensure the visual accuracy of the doctor's control terminal.

[0078] See also Figure 3 Optionally, according to the multimodal data, a recognition strategy is configured, including:

[0079] Step S302: encode the multimodal data to obtain multimodal features.

[0080] Step S304: parse the semantic relationship between the multimodal features based on the large language model to generate a semantically consistent comprehensive representation.

[0081] Step S306: determining the activation priority of each identifier according to the comprehensive representation.

[0082] Step S308, selecting an identifier whose activation priority meets the preset requirements from the preset module library, and configuring a recognition mechanism for each selected identifier to obtain a recognition strategy.

[0083] Specifically, the present application is based on multi-dimensional detection data, using the understanding and reasoning capabilities of a large language model to generate a recognition strategy. First, the multimodal data is encoded to obtain multimodal features with space, language, and vision. Exemplarily, the large language model in the embodiment of the present application adopts an LLM model. LLM uses its powerful natural language understanding capabilities to analyze the semantic relationship between different modalities, and combines it with an external knowledge base or training data for reasoning, and finally outputs a semantically consistent comprehensive representation.

[0084] Furthermore, the task of the recognizer is to determine the most critical detection area in the current environment, so it is necessary to analyze the activation order of different recognizers to ensure the optimal allocation of detection resources. Specifically, according to the preset rules and the comprehensive representation generated by the above LLM, the activation priority of all possible recognizers is calculated. Among them, the preset rules are determined based on medical knowledge and experience. For different types of surgery, the key areas are different. For example, laparoscopic cholecystectomy focuses on the gallbladder, bile duct, liver blood vessels, etc.; gastrointestinal surgery focuses on intestinal mucosa, anastomosis, etc. For different sub-stages of remote surgery, the key areas are also different. For example, in the instrument insertion stage, the focus is on the instrument entrance and vascular injury area; in the cutting and suturing stage, the focus is on monitoring the surgical area and blood flow status. In addition, for abnormal events such as bleeding and tissue damage during surgery, it is also necessary to focus on the bleeding point or anastomosis area. According to the above preset rules and multimodal data, LLM can calculate the key areas that may appear in different scenarios and set activation priorities for each recognizer corresponding to the key area.

[0085] It should be understood that the above activation priority is not a single numerical ranking, but includes a multi-dimensional scoring vector. For example, it can be scored from multiple dimensions such as clinical criticality, timing urgency, resource consumption cost, and scene adaptability. Weights are dynamically assigned to each dimension according to the actual scenario, and the weighted sum is used to obtain the final activation priority of each ROI area. In practical applications, the recognition strategy can be that multiple identifiers run in parallel, and different recognition mechanisms are assigned to different identifiers. For example, for critical areas of surgical instruments and anatomical features, high-priority identifiers are used and high computing resources are allocated; for non-critical areas, low-priority identifiers can be used and batch processed.

[0086] It should be noted that in actual use, multiple identifier combinations can be pre-set according to preset rules and prior experience to meet different application scenarios, and the appropriate identifier combination can be selected according to the multimodal data during the operation. Furthermore, it is also possible to select a pre-stored identifier combination in the initial stage of the operation, and adjust the recognition strategy based on the identifier combination according to the feedback from the doctor's control terminal during the operation to improve the real-time and accuracy of the surgical scene.

[0087] The above remote surgical image transmission method analyzes and generates the activation priority of different identifiers based on the comprehensive information output by LLM, and determines the most suitable recognition strategy in the current scenario. According to the priority, the corresponding identifier is activated to achieve targeted detection of the surgical area, serving the real-time operation guidance of the remote surgical robot. The entire process embodies the fusion processing of multimodal data with space, language, and vision. Through LLM and priority decoding, the adaptive identifier is dynamically selected, which improves the accuracy and pertinence of detection in surgical scenarios.

[0088] Optionally, the multimodal data also includes abnormal information during the operation; then, configuring the recognition strategy according to the multimodal data also includes:

[0089] In the case where abnormal information is detected during the operation, the activation priority of each identifier is adjusted according to the abnormal information to adjust the current recognition strategy.

[0090] For example, if abnormal information such as heavy bleeding is detected during surgery, the original priority is temporarily adjusted, the segmentation model corresponding to the bleeding area is forcibly activated, and non-urgent tasks are suspended, such as suspending the recognition of background areas, etc. By timely adjusting the current recognition strategy, doctors can focus on important areas. It should be noted that the background area changes according to different surgical categories and stages. For example, in laparoscopic cholecystectomy, the gallbladder, bile duct, and liver blood vessels are ROI areas, and the remaining areas can be regarded as background areas or non-ROI areas.

[0091] Optionally, configuring a recognition strategy according to the multimodal data further includes:

[0092] After selecting recognizers whose activation priorities meet preset requirements from a preset module library, based on the confidence obtained by pre-training of each recognizer and a weighted decision method, the retainable recognizers are determined from the selected multiple recognizers to obtain a recognition strategy.

[0093] Specifically, confidence is a value that measures the degree of certainty of the model's output results. The value range is usually between 0 and 1. The higher the confidence, the more accurate the detection result. For a recognizer, its confidence is usually determined by the statistical characteristics of the model on the training data. If there are recognizers with similar functions in the recognizer combination, different recognizers may give different detection results. In this case, decisions can be made based on confidence and weighted decision methods to determine the retained recognizers and improve the stability and accuracy of recognition.

[0094] Furthermore, if multiple identifiers generate policy conflicts, a final judgment is made through preset rules.

[0095] See also Figure 4Optionally, based on the recognition strategy, processing the displayed image to obtain a detection result includes:

[0096] Step S402, using a pre-trained spatiotemporal feature extraction model to extract features from a single frame of the displayed image to obtain spatiotemporal features; wherein the spatiotemporal feature extraction model is trained based on a deep learning model.

[0097] In step S404, the identifier determines the position of each target area in the display image according to the spatiotemporal characteristics, and evaluates the priority of each target area according to a preset rule to obtain the region of interest and its corresponding region parameters.

[0098] Specifically, the embodiment of the present application adopts a target detection algorithm based on shape and texture features, combined with deep learning models such as convolutional neural networks, to achieve an effective fusion of spatial feature extraction and temporal dynamic modeling, accurately locate the end of the surgical instrument and key anatomical tissue, and ensure the stability of the detection results in the time dimension.

[0099] Exemplarily, the present application embodiment uses a pre-trained CNN network to extract features from a single frame of video images to obtain a high-level representation of surgical instruments and key anatomical structures. The CNN network used may include but is not limited to ResNet, EfficientNet or MobileNet. Its purpose is to identify the shape, edge and texture features of surgical instruments; and to extract structural information of anatomical tissues, such as blood vessels, nerves, tumors, etc.

[0100] Furthermore, in view of the temporal dynamic characteristics of the surgical video, this embodiment provides three implementation solutions:

[0101] (1) Spatiotemporal feature extraction method based on CNN+Transformer. After CNN extracts the spatial features of a single frame, the Transformer model is used for global temporal modeling. The Transformer's self-attention mechanism is used to analyze long-term dependencies to ensure the continuity of the motion trajectory of surgical instruments and tissues. This method can effectively reduce the gradient vanishing problem of the traditional RNN structure, improve the stability of temporal information modeling, and is suitable for highly complex surgical scenarios.

[0102] (2) Spatiotemporal feature extraction based on CNN+LSTM. After CNN extracts single-frame features, the feature sequence is input into LSTM for temporal modeling. LSTM is suitable for short-term time-dependent tasks, such as capturing the short-term motion patterns of robot actuators. Due to the long-term dependency problem of LSTM, its performance on long-sequence tasks may be slightly weaker than that of Transformer.

[0103] (3) Spatiotemporal feature extraction based on 3D CNN. 3D CNN (such as C3D and I3D) is used to directly perform spatiotemporal integrated modeling on the input video sequence. 3D convolution operation can simultaneously capture inter-frame motion information and spatial features, reducing the complexity of the model structure. It is suitable for scenarios with sufficient computing resources, but may require more computing optimization compared to the Transformer method.

[0104] The above remote surgical image transmission method uses a target detection algorithm based on shape and texture features, combined with deep learning models such as convolutional neural networks, to achieve an effective fusion of spatial feature extraction and temporal dynamic modeling, and accurately locate the end of the surgical instrument and key anatomical tissues. The detection results obtained based on the above spatiotemporal features ensure stability in the time dimension.

[0105] Optionally, based on the recognition strategy, processing the display image to obtain the detection result further includes:

[0106] After determining the position of each target area in the display image, the edges of adjacent target areas are smoothed by the potential field method, and the smoothed target area is used as the region of interest and its region parameters are output.

[0107] Specifically, there may be discontinuous boundaries between adjacent regions, and the ROIs between different frames may change in a jumpy manner, resulting in visual unevenness; in addition, the boundaries between adjacent ROIs may show irregular, jagged changes, affecting subsequent accurate analysis. Based on this, the embodiment of the present application smoothes the size edges of adjacent ROI regions by means of a potential energy field to further determine the ROI region.

[0108] See also Figure 5 , a positive virtual potential field is established in the ROI area, and an attractive field can be formed with the center. A negative virtual potential field is established in the non-ROI area, and a repulsive potential field is formed with the center. For the repulsive potential field, the closer to the area, the greater the virtual repulsion, and the magnitude of the repulsion is proportional to the square of the distance, and the magnitude can be adjusted through parameters. The opposite is true for the gravitational field. Furthermore, the potential field can establish and integrate different ROI area weights to assign appropriate clarity coefficients to all pixels in the image. The closer to the ROI area, the greater the virtual gravitation, and the magnitude of the gravitation is proportional to the distance, and the corresponding pixels will be displayed more clearly.

[0109] The calculation formula of the repulsive force Fc is as follows:

[0110]

[0111] Among them, Uc is the repulsion ratio adjustment parameter, P real is the position of ROI (such as surgical instrument marker), PGc P real The distance from the pixel point in the non-ROI area, P0 is the set distance threshold parameter, when the instrument tip P real To P Gc If the distance is greater than this parameter, no repulsion will be generated, which means that the pixel belongs to the range of ROI clear processing.

[0112] The above-mentioned remote surgical image transmission method adopts potential field fusion technology to associate the ROI area clarity coefficient with the distance, realize pixel-level fine control, make the ROI area edge balanced and fit the anatomical structure, realize natural transition, and improve the subsequent processing accuracy.

[0113] Optionally, based on the recognition strategy, processing the display image to obtain the detection result further includes:

[0114] After determining the position of each target region in the display image, the edges of multiple adjacent target regions are smoothed by a region growing operator, and the smoothed target region is used as a region of interest and its region parameters are output.

[0115] Specifically, in another embodiment, a region growing operator may be used to uniformly process the edges of multiple adjacent ROIs to obtain smoother region boundaries that are more consistent with the edges of complex anatomical morphologies.

[0116] See also Figure 6, if the identifier used to identify surgical instruments is recorded as the first identifier, the identifier used to identify joint tissue is recorded as the second identifier, and the identifier used to identify abnormal conditions is recorded as the third identifier, then the ROI center point extracted from the first identifier, the ROI center point extracted from the second identifier, and the ROI center point extracted from the third identifier (such as bleeding point) are added to the seed point queue. By taking a seed point from the seed point queue in turn, traverse its neighborhood pixels. For each neighborhood pixel, determine whether it meets the growth conditions according to the defined texture growth criterion. The texture growth criterion mainly uses texture feature descriptors (such as energy, contrast, entropy and other features extracted by the gray level co-occurrence matrix GLCM) to measure the texture similarity between pixels. For the edge of the instrument, its texture is usually different from the surrounding tissue. The difference in texture features between the pixel to be grown and the grown area can be calculated. If the difference is within a certain range, growth is allowed. If it meets the requirements, the pixel is added to the current growth area and added to the seed point queue; if it does not meet the requirements, the pixel is skipped and the above steps are repeated until the seed point queue is empty and the growth of a region is completed. Finally, the boundaries of the grown regions are smoothed by using morphological operations (such as dilation, erosion, opening, closing, etc.) to make the marked region boundaries more natural. For example, small holes inside the region are filled by closing operations, and small burrs on the boundaries are removed by opening operations. The region boundaries include the edge of the instrument, the gallbladder, and local bleeding points.

[0117] The above remote surgery image transmission method selects a central point inside each ROI area and grows from the center to the boundary. The boundary can be processed in a fine-grained manner, so that a smooth transition area is formed between different ROIs.

[0118] Furthermore, the detection results can be processed by combining potential field with region growing. For example, region growing can be used for local optimization to ensure the integrity of ROI boundaries and smooth transition between adjacent ROIs; then potential field can be used for global optimization to smooth ROI changes in the time dimension and improve video stability.

[0119] Optionally, the remote surgery image transmission method further includes:

[0120] After the test results are obtained, an evaluation is performed based on the confidence level of the test results, and a warning signal is generated when the evaluation score is lower than the preset score.

[0121] Specifically, the detection results include two types of areas, one is the surgical instrument area, and the other is the anatomical part area. Both of the above categories include multiple ROI areas. The confidence of the detection result refers to whether the currently detected ROI area is credible, and the confidence can be obtained in a variety of ways. Exemplarily, the confidence acquisition method may include model output confidence, temporal stability analysis, historical data statistics, fusion confidence of multiple identifiers, etc. Among them, the model output confidence is directly output by the identifier that generates the ROI area. Temporal stability analysis means that in the video stream, ROI is generated for each frame. If the ROI is unstable in multiple consecutive frames, the confidence is reduced. Historical data statistics means that according to historical statistics, the false detection rate of a certain identifier is high. If the confidence of the current frame is much lower than the historical mean, it may be a false detection. The fusion confidence of multiple identifiers refers to weighted fusion of multiple identifiers. If the confidence after fusion is low, it may be a false detection. Using the above method, a warning signal can be generated when the evaluation score is low to remind the operator to pay attention and avoid affecting the surgical process.

[0122] Furthermore, while generating the warning signal, the priority of each ROI area is adjusted to the initial priority based on the preset backup plan, and adjustments are made on this basis in the future. For example, the priority of the center area of ​​the screen is set to the highest.

[0123] See also Figure 7 Optionally, the regional parameters include size information and priority; then, according to the regional parameters and the current bandwidth situation, a transmission strategy is determined, and based on the transmission strategy, the display image is transmitted to the doctor control terminal, including:

[0124] Step S702: configuring encoding parameters of the region of interest and the background region according to the region parameters, wherein the encoding parameters include at least one of compression rate, bit rate, quantization parameter, frame rate, and resolution.

[0125] Step S704: Select a corresponding transmission protocol according to the current bandwidth situation.

[0126] Step S706, encode the display image based on the encoding parameters, and transmit the encoded multiple data packets to the doctor control terminal through the selected transmission protocol.

[0127] Specifically, the region parameters include size information and priority. The size information includes bounding box, pixel area, relative area ratio and semantic mask; the identifier assigns a priority to each ROI region, which is used to determine the subsequent encoding quality, transmission order and bandwidth allocation. Depending on the region parameters, the corresponding encoding parameters can be configured for the image. For example, for regions with high priority and large size, high bit rate, low compression and high-definition resolution are used to ensure that key parts are clearly visible; for regions with low priority and small size, the bit rate can be appropriately reduced to avoid bandwidth waste; for background areas, high compression rate and low bit rate are used to reduce bandwidth occupancy.

[0128] Coding parameters are used for video compression and quality control. The compression rate is used to determine the degree of data compression; the bit rate refers to the amount of data encoded per second in the video, which is used to determine the picture quality; the quantization parameter (QP) is used to control the data accuracy when encoding the video. The lower the QP value, the higher the image quality and the larger the file size; the frame rate refers to the number of frames played per second, which is used to determine the video smoothness and storage size; the resolution refers to the pixel size of the video image, which is used to determine the visual clarity. By properly selecting the above coding parameters, the clarity of key areas can be ensured when bandwidth is limited.

[0129] In remote surgery, the choice of transmission protocol is crucial and must meet the requirements of low latency, high reliability, data integrity, dynamic bandwidth scheduling, etc. For example, if the current bandwidth is sufficient, a low-latency protocol can be selected to ensure the real-time performance of the ROI area and maintain the current allocation. If the current bandwidth is insufficient, a high-compression protocol can be selected to save bandwidth.

[0130] It should be understood that before transmitting the display image, a coding algorithm is first used to compress the display image obtained by the surgical robot into multiple coding blocks or frames as transmission units, and then the transmission units are transmitted to the remote side through cloud transmission. After receiving the transmission units, the remote side reassembles these transmission units and finally restores the complete image or video stream.

[0131] The encoding is based on the high-quality encoding strategy for the ROI area. For example, the encoding quality of the ROI area can be improved by reducing the quantization parameter QP value of the ROI area. Furthermore, for the non-ROI area, the resolution can be reduced, the frame rate can be reduced, or the bit rate can be compressed at a low rate. The compression ratio can be dynamically adjusted according to the bandwidth situation.

[0132] Optionally, the embodiment of the present application also monitors the channel bandwidth in real time. If the bandwidth is sufficient, the real-time performance of the video is prioritized, the transmission delay is reduced, the current encoding parameters and bandwidth allocation are maintained, and high-quality transmission in the ROI area is ensured. If the bandwidth is insufficient, a high compression rate can be used to save bandwidth while maintaining video availability.

[0133] For example, when it is detected that the bandwidth is reduced and the clarity of the ROI area still needs to be ensured, the embodiment of the present application provides the following two implementation methods for the ROI area:

[0134] 1) Reduce the quantization parameter (QP value) of the ROI area: The quantization parameter QP directly affects the compression ratio and image quality of the video. Reducing the QP can improve the image quality. For example, in H.265 encoding, the QP can be dynamically adjusted for the ROI area so that the ROI area can still maintain high definition under bandwidth-constrained conditions.

[0135] 2) Increase the encoding bit rate in the ROI area: By increasing the encoding bit rate in the ROI area, compression loss is reduced and the ability to retain picture details is improved.

[0136] For non-ROI areas, in order to save bandwidth, different compression strategies are adopted to reduce the amount of data transmission without affecting the overall visual experience. The present application embodiment provides three implementation methods:

[0137] 1) Down-resolution processing: Down-resolution encoding is used for non-ROI areas, such as from 1080p to 720p, to reduce the amount of transmitted data.

[0138] 2) Reduce the frame rate: When the bandwidth is extremely limited, you can reduce the frame rate of the non-ROI area, for example, from 30FPS to 15FPS, to reduce the bitrate burden.

[0139] 3) Low bit rate compression strategy: Use a higher QP value in the non-ROI area to increase the video compression ratio so that it occupies less bandwidth. At the same time, use the variable bit rate (VBR) method to dynamically adjust the bit rate according to the available bandwidth to avoid image distortion.

[0140] The above remote surgical image transmission method configures the encoding strategy according to the regional parameters of the ROI area, and selects the corresponding transmission protocol according to the current bandwidth situation. Under the condition of limited bandwidth, the transmission quality of the key area is guaranteed, ensuring the visual clarity of the doctor's control end.

[0141] Optionally, the remote surgery image transmission method further includes:

[0142] A pre-trained prediction model is used to predict the bandwidth fluctuation trend of the network within a preset time. According to the bandwidth fluctuation trend and regional parameters, the encoding parameters are adjusted before transmission, and bandwidth is allocated to each data packet. The prediction model is trained based on historical bandwidth data and machine learning algorithms.

[0143] Specifically, historical broadband data includes the base station load and bandwidth of the target hospital. The machine learning algorithm is used to predict the bandwidth fluctuation trend, which can ensure the broadband priority of key operating rooms and the low-bandwidth periods during remote surgery. By adjusting the encoding parameters and allocating bandwidth in advance, the visual clarity and real-time performance of the doctor's control end can be ensured.

[0144] Furthermore, machine learning algorithms can be selected according to the actual application scenarios. For example, for scenarios with seasonal fluctuations, time series models can be selected; for complex scenarios with multiple variables such as surgical flow, number of remote users, base station load, etc., random forest algorithms can be selected.

[0145] The above-mentioned remote surgical image transmission method obtains the network bandwidth status of the current video signal in real time, judges the bandwidth trend in advance through the pre-trained prediction model, and makes more appropriate pre-adjustment strategies according to the characteristics of long-term or short-term fluctuations of the bandwidth, thereby improving the stability and adaptability of the system.

[0146] Optionally, the remote surgery image transmission method further includes:

[0147] After transmitting multiple data packets to the doctor control end, the encoding parameters and transmission protocol are adjusted according to the feedback information from the doctor control end to ensure the clarity of each region of interest.

[0148] Specifically, the above feedback information can be automatically calculated based on the quality assessment algorithm and fed back to the local side, or it can be manually indicated by the operator. For example, evaluation indicators can be pre-constructed to determine whether the doctor's screen has problems such as blur, delay, and freeze by verifying the evaluation indicators. Among them, the evaluation indicators include: peak signal-to-noise ratio, structural similarity index, packet loss rate, and latency. Furthermore, options such as "blurred screen" and "adjust clarity" can be pre-set on the operation interface. If the selection instruction of the above options is received, it will be fed back to the local side. In addition, the operator can also input the above instructions by voice.

[0149] After receiving the above instructions, the local side adjusts the encoding parameters and transmission protocol to improve the clarity of the ROI area. For example, when the bandwidth is reduced, the compression rate of the non-ROI area continues to be reduced to ensure the bandwidth of the ROI area. In the above manner, the encoding parameters and transmission protocol can be adjusted in real time according to the feedback information of the remote side to improve the clarity of the ROI area.

[0150] The above remote surgical image transmission method obtains multimodal data with space, language, and vision in real time, uses a large language model combined with multimodal data to jointly determine the call and priority of the recognizer, and then generates a recognition strategy. The recognition strategy is used to process the displayed image, and the obtained ROI area meets the real-time change requirements of the surgery.

[0151] Furthermore, the present application also screens the generated recognition strategies according to the confidence and weighted decision method of each recognizer to determine the types and mechanisms of the recognizers that are ultimately retained. During the operation, the priority of the recognizers can be dynamically triggered according to the emergency situation, and the corresponding recognizers can be forcibly triggered, which is highly flexible.

[0152] Furthermore, the present application also uses a deep learning model to extract spatiotemporal features from a single frame of the displayed image, thereby achieving an effective fusion of spatial feature extraction and temporal dynamic modeling, enabling the identifier to accurately locate the end of the surgical instrument and key anatomical tissue, and ensure the stability of the detection results in the time dimension.

[0153] Furthermore, the present application also uses potential energy fields and / or regional growth to smooth adjacent ROI areas in the detection results, so that the edges of the ROI areas are balanced and fit the anatomical structure, which can better cover the irregular edges of complex instruments and anatomical structures in endoscopic surgery scenarios, thereby improving the subsequent processing accuracy.

[0154] Furthermore, after obtaining the test results, the present application performs an evaluation based on the confidence level of the test results and generates a warning signal when the evaluation score is lower than a preset score, to alert the operator to avoid affecting the surgical process.

[0155] Furthermore, the present application also dynamically adjusts the encoding parameters according to the regional parameters, and selects the corresponding transmission protocol according to the current bandwidth conditions. Based on the encoding parameters and transmission protocols, the doctor control terminal can receive a display screen with low latency and high accuracy.

[0156] Furthermore, the present application also monitors the channel bandwidth in real time, and adjusts the encoding parameters and transmission protocol in real time according to the broadband situation, so that the ROI area can still maintain high definition under bandwidth-limited conditions. In addition, the present application also adjusts the encoding parameters and transmission protocol in real time according to the feedback information from the doctor's control end to ensure the visual accuracy of the doctor's control end.

[0157] Furthermore, the present application also trains a prediction model based on historical bandwidth data and machine learning algorithms. The prediction model can be used to predict bandwidth fluctuation trends, adjust encoding parameters and allocate bandwidth before transmission, and further ensure high definition of the ROI area under bandwidth-constrained conditions.

[0158] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0159] Based on the same inventive concept, an embodiment of the present application also provides a remote surgical image transmission device, which is suitable for the above-mentioned remote surgical image transmission method. The implementation solution provided by the device to solve the problem is similar to the implementation solution recorded in the above-mentioned method. Therefore, the specific limitations in one or more device embodiments provided below can be referred to the limitations on the method above, and will not be repeated here.

[0160] See also Figure 8 In one embodiment, the remote surgery image transmission device includes: a data acquisition module, a processing module and a transmission module.

[0161] The data acquisition module is used to obtain the display images and multimodal data of the user during surgery; wherein the multimodal data includes the spatial information of the instrument, the surgical procedure information and the process information.

[0162] The processing module is used to configure the recognition strategy according to the multimodal data, and based on the recognition strategy, process the display image to obtain the detection result, and the detection result includes the region of interest and its corresponding region parameters.

[0163] The transmission module is used to determine the transmission strategy according to the regional parameters and the current bandwidth situation, and transmit the display image to the doctor control terminal based on the transmission strategy.

[0164] Optionally, the processing module is also used to encode multimodal data to obtain multimodal features; parse the semantic relationship between multimodal features based on a large language model to generate a semantically consistent comprehensive representation; determine the activation priority of each recognizer based on the comprehensive representation; determine the selected recognizer and the recognition mechanism of each recognizer based on the activation priority to obtain a recognition strategy.

[0165] Optionally, the multimodal data also includes abnormal information during the operation; the processing module is further used to adjust the activation priority of each identifier according to the abnormal information when abnormal information is detected during the operation, so as to adjust the current recognition strategy.

[0166] Optionally, the processing module is also used to select identifiers whose activation priorities meet preset requirements from a preset module library, and then determine the retainable identifiers from the selected multiple identifiers based on the confidence obtained from the pre-training of each identifier and a weighted decision method to obtain a recognition strategy.

[0167] Optionally, the processing module is also used to use a pre-trained spatiotemporal feature extraction model to extract features from a single frame of the displayed image to obtain spatiotemporal features; wherein the spatiotemporal feature extraction model is trained based on a deep learning model; the identifier determines the position of each target area in the displayed image based on the spatiotemporal features, and evaluates the priority of each target area according to preset rules to obtain the area of ​​interest and its corresponding area parameters.

[0168] Optionally, the processing module is further used to, after determining the position of each target area in the display image, smooth the edges of adjacent target areas by a potential field method, take the smoothed target area as the region of interest and output its region parameter.

[0169] Optionally, the processing module is further used to, after determining the position of each target area in the display image, smooth the edges of multiple adjacent target areas through a region growing operator, use the smoothed target area as a region of interest and output its region parameter.

[0170] Optionally, the processing module is further used to evaluate the test results according to the confidence level after obtaining the test results, and generate a warning signal when the evaluation score is lower than a preset score.

[0171] Optionally, the transmission module is also used to configure the encoding parameters of the region of interest and the background region according to the regional parameters, wherein the encoding parameters include at least one of compression rate, bit rate, quantization parameter, frame rate, and resolution; select the corresponding transmission protocol according to the current bandwidth situation; encode the display image based on the encoding parameters, and transmit the encoded multiple data packets to the doctor control terminal through the selected transmission protocol.

[0172] Optionally, the transmission module is also used to use a pre-trained prediction model to predict the bandwidth fluctuation trend of the network within a preset time; according to the bandwidth fluctuation trend and regional parameters, the encoding parameters are adjusted before transmission, and bandwidth is allocated to each data packet; wherein the prediction model is trained based on historical bandwidth data and a machine learning algorithm.

[0173] Optionally, the transmission module is also used to adjust the encoding parameters and transmission protocol according to feedback information from the doctor control end after transmitting multiple data packets to the doctor control end, so as to ensure the clarity of each region of interest.

[0174] The above-mentioned remote surgical image transmission device acquires multimodal data with space, language, and vision in real time, and uses a large language model combined with multimodal data to jointly determine the call and priority of the recognizer, and then generates a recognition strategy. The recognition strategy is used to process the displayed image, and the obtained ROI area meets the real-time change requirements of the surgery.

[0175] Furthermore, the present application also screens the generated recognition strategies according to the confidence and weighted decision method of each recognizer to determine the types and mechanisms of the recognizers that are ultimately retained. During the operation, the priority of the recognizers can be dynamically triggered according to the emergency situation, and the corresponding recognizers can be forcibly triggered, which is highly flexible.

[0176] Furthermore, the present application also uses a deep learning model to extract spatiotemporal features from a single frame of the displayed image, thereby achieving an effective fusion of spatial feature extraction and temporal dynamic modeling, enabling the identifier to accurately locate the end of the surgical instrument and key anatomical tissue, and ensure the stability of the detection results in the time dimension.

[0177] Furthermore, the present application also uses potential energy fields and / or regional growth to smooth adjacent ROI areas in the detection results, so that the edges of the ROI areas are balanced and fit the anatomical structure, which can better cover the irregular edges of complex instruments and anatomical structures in endoscopic surgery scenarios, thereby improving the subsequent processing accuracy.

[0178] Furthermore, after obtaining the test results, the present application performs an evaluation based on the confidence level of the test results and generates a warning signal when the evaluation score is lower than a preset score, to alert the operator to avoid affecting the surgical process.

[0179] Furthermore, the present application also dynamically adjusts the encoding parameters according to the regional parameters, and selects the corresponding transmission protocol according to the current bandwidth conditions. Based on the encoding parameters and transmission protocols, the doctor control terminal can receive a display screen with low latency and high accuracy.

[0180] Furthermore, the present application also monitors the channel bandwidth in real time, and adjusts the encoding parameters and transmission protocol in real time according to the broadband situation, so that the ROI area can still maintain high definition under bandwidth-limited conditions. In addition, the present application also adjusts the encoding parameters and transmission protocol in real time according to the feedback information from the doctor's control end to ensure the visual accuracy of the doctor's control end.

[0181] Furthermore, the present application also trains a prediction model based on historical bandwidth data and machine learning algorithms. The prediction model can be used to predict bandwidth fluctuation trends, adjust encoding parameters and allocate bandwidth before transmission, and further ensure high definition of the ROI area under bandwidth-constrained conditions.

[0182] Each module in the above remote surgery image transmission device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.

[0183] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0184] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0185] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A remote surgery image transmission method, characterized in that: The method comprises: Acquire the display image and multimodal data of the user during the operation; wherein the multimodal data includes the spatial information of the instrument, the operation procedure information and the process information; encode the multimodal data to obtain the multimodal features; Analyzing the semantic relationship between the multimodal features based on the large language model to generate a semantically consistent comprehensive representation; According to the comprehensive representation, determining the activation priority of each recognizer to select the recognizer and its corresponding recognition mechanism to obtain a recognition strategy; Based on the recognition strategy, the display image is processed to obtain a detection result, wherein the detection result includes a region of interest and its corresponding region parameters; A transmission strategy is determined according to the regional parameters and the current bandwidth situation, and based on the transmission strategy, the display image is transmitted to the doctor control terminal.

2. The method according to claim 1, characterized in that The determining of the activation priority of each identifier to select the identifier and its corresponding recognition mechanism to obtain the recognition strategy includes: The recognition strategy is obtained by selecting a recognition device whose activation priority meets the preset requirements from a preset module library, and configuring a recognition mechanism for each selected recognition device.

3. The method according to claim 2, characterized in that The multimodal data also includes abnormal information during surgery; The configuring the recognition strategy according to the multimodal data further includes: In the case where abnormal information is detected during the operation, the activation priority of each identifier is adjusted according to the abnormal information to adjust the current recognition strategy.

4. The method according to claim 2, characterized in that: The step of processing the displayed image to obtain a detection result based on the recognition strategy includes: Using a pre-trained spatiotemporal feature extraction model to extract features from a single frame of the display image to obtain spatiotemporal features; wherein the spatiotemporal feature extraction model is trained based on a deep learning model; The identifier determines the position of each target area in the displayed image according to the spatiotemporal characteristics, and evaluates the priority of each target area according to a preset rule to obtain a region of interest and its corresponding region parameters.

5. The method according to claim 4, characterized in that The processing of the display image to obtain a detection result based on the recognition strategy further includes: After determining the position of each target area in the display image, the edges of adjacent target areas are smoothed by a potential field method, and the smoothed target area is used as a region of interest and its region parameters are output.

6. The method according to claim 4, characterized in that The step of processing the display image to obtain a detection result based on the recognition strategy further includes: After determining the position of each target region in the display image, the edges of the adjacent multiple target regions are smoothed by using a region growing operator, and the smoothed target regions are used as regions of interest and their region parameters are output.

7. The method according to claim 2, characterized in that The configuring the recognition strategy according to the multimodal data further includes: After selecting recognizers whose activation priorities meet preset requirements from a preset module library, the retainable recognizers are determined from the selected multiple recognizers based on the confidence obtained by pre-training of each of the recognizers and a weighted decision method.

8. The method according to claim 1, characterized in that The method further comprises: After the test results are obtained, an assessment is performed based on the confidence level of the test results, and a warning signal is generated when the assessment score is lower than a preset score.

9. The method according to claim 1, characterized in that: The area parameters include size information and priority; Determining a transmission strategy according to the regional parameters and the current bandwidth situation, and transmitting the display image to the doctor control terminal based on the transmission strategy, includes: According to the region parameters, configure encoding parameters of the region of interest and the background region, wherein the encoding parameters include at least one of a compression rate, a bit rate, a quantization parameter, a frame rate, and a resolution; Select the corresponding transmission protocol according to the current bandwidth situation; The display image is encoded based on the encoding parameters, and the encoded multiple data packets are transmitted to the doctor control terminal through the selected transmission protocol.

10. The method according to claim 9, characterized in that The method further comprises: Use pre-trained prediction models to predict the bandwidth fluctuation trend of the network within a preset time; According to the bandwidth fluctuation trend and the regional parameters, adjusting encoding parameters before transmission, and allocating bandwidth for each of the data packets; The prediction model is obtained based on historical bandwidth data and machine learning algorithm training.

11. The method according to claim 9, characterized in that The method further comprises: After transmitting a plurality of data packets to the doctor control end, the encoding parameters and the transmission protocol are adjusted according to feedback information from the doctor control end to ensure the clarity of each region of interest.

12. A remote surgery image transmission device, characterized in that: The device comprises: A data acquisition module, used to acquire the display image and multimodal data of the user during the operation; wherein the multimodal data includes the spatial information of the instrument, the operation procedure information and the process information; A processing module is used to encode the multimodal data to obtain multimodal features; analyze the semantic relationship between the multimodal features based on a large language model to generate a semantically consistent comprehensive representation; determine the activation priority of each recognizer according to the comprehensive representation to select the recognizer and its corresponding recognition mechanism to obtain a recognition strategy, and based on the recognition strategy, process the display image to obtain a detection result, wherein the detection result includes a region of interest and its corresponding region parameters; The transmission module is used to determine a transmission strategy according to the regional parameters and the current bandwidth situation, and transmit the display image to the doctor control terminal based on the transmission strategy.

13. A surgical robot system, characterized in that: The system comprises: The control console is used to provide a control platform for doctors and generate control instructions according to the doctors' operations; A surgical robot is used to perform surgery according to the control instructions, and adopt the remote surgery image transmission method described in any one of claims 1-11 to transmit the display image obtained during the operation to the control console.

Citation Information

Patent Citations

  • Computer program, learning model generation method, and information processing device

    CN117956939A

  • Deep learning adaptive operation video coding optimization method and system

    CN118433386A