Remote medical video transmission system for clinical real scene

By constructing the anatomical matrix and motion matrix, the ADW matrix is ​​generated for unified hierarchical encoding, which solves the problem of video quality degradation in high-motion or high-complex surgical scenarios in the existing technology, and realizes more efficient telemedicine video transmission and a more stable telemedicine system.

CN120111240AInactive Publication Date: 2025-06-06南昌大学第一附属医院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510295714.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN120111240A_ABST
    Figure CN120111240A_ABST
Patent Text Reader

Abstract

The invention discloses a remote medical video transmission system for clinical live-action, and relates to the technical field of video transmission. Comprising an anatomical feature analysis module which constructs an anatomical matrix based on medical image data and performs spatial priority division on a video area according to a clinical diagnosis and treatment scene; the instrument motion detection module is used for dynamically tracking the surgical instrument and the operation process, constructing a motion matrix, and optimizing coding weights among different frames by using a dynamic adjustment factor of a time dimension; the ADW matrix construction module is used for generating an ADW matrix in combination with the anatomical matrix and the motion matrix, performing unified layered coding based on the ADW matrix, and setting different coding precisions for different weight regions; and the parameter linkage adjustment module is used for optimizing the coding level and the transmission rate in combination with the movement state of the instrument to adapt to the dynamic change of different operation scenes. According to the invention, the overall quality and efficiency of telemedicine video transmission are improved, and the coding quality of the telemedicine video is refined and optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video transmission, in particular to a remote medical video transmission system for clinical real scenes. Background Art

[0002] In recent years, the rapid development of telemedicine technology has greatly promoted the balanced allocation of medical resources, especially in application scenarios such as remote consultation, surgical guidance, and remote ultrasound, where high-definition, low-latency real-time video transmission has become one of the key technologies. At present, clinical real-life telemedicine video transmission mainly relies on traditional video coding technologies, such as compression algorithms such as H.264, H.265, and AV1, to reduce data volume and improve transmission efficiency. However, conventional video coding technology is mainly optimized for general video data, and fails to fully consider the special needs of clinical medical scenarios, especially in high-precision scenarios such as live surgery and minimally invasive interventional surgery. Traditional coding methods often have problems such as loss of key details and insufficient dynamic adjustment capabilities.

[0003] The existing telemedicine video transmission technology has the following main shortcomings, such as the failure to fully utilize medical imaging data for video area optimization coding. Existing methods usually adopt fixed coding strategies and cannot dynamically adjust the coding level of key surgical areas, resulting in video quality degradation and even mosaic or blurring in high-motion or high-complexity scenes. Traditional coding methods lack accurate modeling of surgical instrument motion characteristics and cannot effectively deal with the rapid movement or partial stillness of surgical instruments, thus affecting the observation experience of remote operators. In addition, most of the existing video coding optimization schemes are based on image complexity or motion characteristics for optimization alone, but fail to comprehensively consider the synergy of anatomical features and motion features, resulting in inefficient allocation of coding resources. Especially in remote minimally invasive surgery, the motion trajectory of surgical instruments and the texture complexity of the target anatomical structure have a direct impact on video coding. Summary of the invention

[0004] In view of the problems existing in the existing telemedicine video transmission process, the present invention proposes a telemedicine video transmission system for clinical real scenes.

[0005] Therefore, the problem to be solved by the present invention is how to provide more stable and efficient technical support for remote medical surgery guidance and remote diagnosis and treatment.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: The present invention provides a remote medical video transmission system for clinical real scenes, which includes an anatomical feature analysis module, which is used to construct an anatomical matrix based on medical image data, and to divide video areas into spatial priorities according to clinical diagnosis and treatment scenarios; an instrument motion detection module, which is used to dynamically track surgical instruments and operation processes, construct a motion matrix, and optimize the coding weights between different frames using a dynamic adjustment factor in the time dimension; an ADW matrix construction module, which is used to combine the anatomical matrix and the motion matrix to generate an ADW matrix, and perform unified hierarchical coding based on the ADW matrix, and set different coding accuracies for different weight areas; a parameter linkage adjustment module, which is used to optimize the coding level and transmission rate in combination with the instrument motion state, and adapt to the dynamic changes of different surgical scenarios.

[0007] As a preferred solution of the remote medical video transmission system for clinical scenes described in the present invention, the anatomical matrix is ​​a two-dimensional or three-dimensional weight matrix constructed based on medical imaging data, in which each pixel corresponds to a weight value for indicating the importance of the area.

[0008] As a preferred solution of the telemedicine video transmission system for clinical real scenes described in the present invention, the spatial priority division includes: using a multi-scale pyramid model to divide the video screen into different levels, such as high, medium, and low priority areas, and the resolution and frame rate of each layer are different; and using scale-invariant feature transformation or optical flow analysis to calculate the stability of key areas between video frames, and adjust the priority according to dynamic changes.

[0009] As a preferred solution of the telemedicine video transmission system for clinical real scenes described in the present invention, the construction of the motion matrix includes: calculating the movement speed, acceleration and direction change of the instrument; the horizontal and vertical coordinates correspond to the pixel coordinates in the video frame respectively; setting the motion area weight; normalizing the movement speed, acceleration and direction change, and calculating the weight factors respectively, and weighted summing the three weight factors to obtain the motion weight factor; taking the product of the motion area weight and the motion weight factor as the initial coding weight; according to the motion matrix and the historical trajectory, calculating the dynamic adjustment factor of the time dimension, dynamically optimizing the initial coding weights at different time points, and obtaining the final coding weight.

[0010] As a preferred solution of the remote medical video transmission system for clinical real scenes described in the present invention, the ADW matrix construction module includes an ADW matrix generation unit, which combines the anatomical matrix and the motion matrix to generate an ADW matrix and a unified hierarchical coding unit, performs hierarchical coding based on the ADW matrix, and sets different coding accuracies for different weight areas.

[0011] As a preferred solution of the remote medical video transmission system for clinical real scenes described in the present invention, the generation of the ADW matrix includes processing the anatomical matrix and the motion matrix through three stages of mutual mapping, weight adjustment, and feature alignment, so that the features of the anatomical matrix and the motion matrix are deeply combined in the spatial dimension and the temporal dimension.

[0012] As a preferred solution of the remote medical video transmission system for clinical real scenes described in the present invention, the operation process of the unified hierarchical coding unit includes two stages: coding adaptability analysis and frame internal and external correlation calculation.

[0013] As a preferred solution of the remote medical video transmission system for clinical real scenes described in the present invention, the coding adaptability analysis includes: calculating the spatial complexity and motion complexity of each area according to the priority divided areas; dynamically adjusting the coding strategy according to the spatial complexity and motion complexity.

[0014] As a preferred solution of the remote medical video transmission system for clinical scenes described in the present invention, the optimization of the coding level includes: for the current frame and the previous frame, calculating the similarity of adjacent frames; extracting the motion vector field of the surgical instrument area and calculating the local optical flow change; dynamically adjusting the coding level according to the similarity and the local optical flow change.

[0015] As a preferred solution of the remote medical video transmission system for clinical scenes described in the present invention, the transmission rate includes: establishing a mapping relationship between the coding level and the transmission rate: if the coding level is increased, the transmission rate is increased; if the coding level is reduced, the transmission rate is reduced; according to the current network available bandwidth, ensure that the adjusted transmission rate does not exceed the network carrying capacity.

[0016] The present invention provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program instructions are executed by the processor, the steps of the remote medical video transmission system for clinical real scenes as described in the first aspect of the present invention are implemented.

[0017] The present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program instructions are executed by a processor, the steps of the remote medical video transmission system for clinical real scenes as described in the first aspect of the present invention are implemented.

[0018] The beneficial effects of the present invention are as follows: the present invention combines the anatomical matrix and the motion matrix, generates an ADW matrix through mapping, and performs unified hierarchical encoding based on the matrix, so that the key medical area obtains higher encoding accuracy, while the encoding resources of the background area or irrelevant area are appropriately reduced, thereby improving the overall quality and efficiency of telemedicine video transmission, and the encoding quality of the telemedicine video is refined and optimized. While ensuring high-precision encoding of the key surgical area, it also effectively reduces bandwidth occupancy and improves the stability and adaptability of the telemedicine system. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0020] Figure 1 The framework diagram of the telemedicine video transmission system for clinical real scenes.

[0021] Figure 2 Schematic diagram of the ADW matrix building block for the telemedicine video transmission system for clinical real-life scenarios. DETAILED DESCRIPTION

[0022] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0023] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0025] Example 1 Reference Figure 1~Figure 2 , which is the first embodiment of the present invention, and provides a remote medical video transmission system for clinical real scenes, including: The anatomical feature analysis module 100 is used to construct an anatomical matrix based on medical imaging data and to divide the video areas into spatial priorities according to the clinical diagnosis and treatment scenarios; the instrument motion detection module 200 is used to dynamically track surgical instruments and operation processes, construct a motion matrix, and use the dynamic adjustment factor of the time dimension to optimize the coding weights between different frames, thereby improving the timing transmission stability of key areas; the ADW matrix construction module 300 is used to combine the anatomical matrix and the motion matrix to generate an ADW matrix, and perform unified hierarchical coding based on the ADW matrix, set different coding accuracies for different weight areas, improve the coding quality of key areas, and reduce the transmission load of secondary areas; the parameter linkage adjustment module 400 establishes hierarchical coding parameter linkage adjustment rules, optimizes the coding level and transmission rate in combination with the instrument motion state, adapts to the dynamic changes of different surgical scenarios, and ensures the stability and adaptability of telemedicine video transmission.

[0026] During telemedicine video transmission, the clarity of medical images and the transmission priority of key areas directly affect the doctor's diagnosis and surgical guidance. However, traditional video coding technology treats different areas in a relatively balanced manner and cannot perform differentiated optimization based on the importance of different parts in medical scenarios. Therefore, analysis and division based on anatomical features is the basis of telemedicine video coding, which can provide accurate regional weights for subsequent coding, improve the clarity of key areas, and reduce the coding redundancy of non-key areas, thereby optimizing transmission efficiency.

[0027] It should be noted that medical imaging data has highly specialized anatomical structures, such as surgical areas, diseased tissues, organ boundaries, etc. The clarity of these parts plays a key role in medical diagnosis and surgical planning. However, the video content in the surgical scene usually contains a lot of irrelevant information, such as surrounding background, non-surgical parts, etc. If the same encoding weight is used for all areas, it will not only occupy bandwidth resources, but also may cause the key areas to lose important details under high compression ratios. Therefore, it is necessary to establish an anatomical matrix to spatially prioritize the video areas so that the encoding process can distinguish between key areas and non-key areas, thereby optimizing the video quality under limited bandwidth conditions. Therefore, in an embodiment of the present application, the anatomical feature analysis module 100 specifically includes: Anatomical Matrix It is a two-dimensional or three-dimensional weight matrix constructed based on medical imaging data, in which each pixel (or voxel) corresponds to a weight value to indicate the importance of the area. Its construction process includes the following steps: Medical imaging data comes from MRI (magnetic resonance imaging), CT (computed tomography), ultrasound images or intraoperative endoscopic videos; the acquired DICOM format medical images are standardized and converted to make them suitable for subsequent processing; Gaussian filtering or adaptive filtering methods are used to remove noise in the image to improve image quality, and histogram equalization or Laplace enhancement technology is used to improve the visibility of key tissues.

[0028] Based on convolutional neural networks (CNN) or image gradient-based edge detection methods, key anatomical structures in medical images can be identified. Semantic segmentation networks such as U-Net and DeepLabV3+ can be used to distinguish surgical sites, lesion areas, and surrounding normal tissues.

[0029] Optionally, different weights can be assigned to different anatomical areas according to surgical or diagnostic requirements, including: surgical target areas (such as lesions, tissue cutting areas), surgical instrument interaction areas (such as clamping sites, cutting edges), background and irrelevant areas (such as operating beds, surrounding non-target tissues).

[0030] After constructing the anatomical matrix, it is necessary to further spatially prioritize the video areas according to the surgical scenario. The specific steps are as follows: Use a multi-scale pyramid model to divide the video screen into different levels, such as high-, medium-, and low-priority areas. The resolution and frame rate of each layer are different. Use scale-invariant feature transformation or optical flow analysis to calculate the stability of key areas between video frames, and adjust the priority according to their dynamic changes.

[0031] For example, if scale-invariant feature transform (SIFT) is used for key area matching, the specific process includes: For two consecutive frames, the scale-invariant feature transform (SIFT) is used to extract the key point set, where each key point contains information such as position, scale, and direction; then, the Euclidean distance of the matching key points is calculated; if the Euclidean distance is less than the set distance threshold, the point is considered to be matched successfully. The total number of matching points can be used to measure stability: first calculate the total number of key points in the current frame and the next frame, respectively, denoted as m and n; find the number of key point pairs that are successfully matched between the two frames, determine the larger value of m and n, and divide the number of key points that are successfully matched by the larger value of m and n. The ratio obtained is the stability.

[0032] Furthermore, the priority is adjusted based on stability, with stable areas having higher priorities and unstable areas having lower priorities.

[0033] It should be made clear that the priority of the same area in adjacent frames should remain relatively stable to avoid fluctuations in encoding quality due to temporary occlusion; if an area is affected by the operation of the instrument (such as clamping, cutting), the encoding priority of the area should be dynamically increased.

[0034] The anatomical feature analysis module provides accurate regional weight reference for the encoding of telemedicine videos by constructing an anatomical matrix and performing spatial priority division based on medical imaging data.

[0035] In the process of telemedicine video transmission, the motion trajectory of surgical instruments and their interaction with anatomical structures are crucial for doctors' remote guidance and intraoperative decision-making. However, due to the complexity of surgical scenes, traditional video encoding methods are difficult to adaptively optimize the motion characteristics of instruments, resulting in key frame loss, motion blur or waste of encoding resources.

[0036] The motion characteristics of surgical instruments have the following features: the movement speed and direction of surgical instruments change at any time, affecting the visualization of surgical operations; the movement of instruments is usually concentrated in the surgical target area, and it is necessary to increase the encoding priority for this area; surgical operations usually consist of multiple continuous actions (such as clamping, cutting, suturing, etc.), and it is necessary to ensure a smooth transition between adjacent frames to avoid visual discontinuities caused by unstable frame rate; during the operation, the instrument may partially occlude the anatomical structure, and the encoding strategy needs to be optimized through a detection mechanism to reduce the impact of occlusion on video transmission.

[0037] In order to meet these challenges, the present invention constructs a module that can detect the motion of surgical instruments in real time, construct a motion matrix, and introduce a time dimension adjustment factor to ensure the coding stability of key areas on the time axis, thereby improving the quality and timing consistency of telemedicine videos. In the embodiment of the present application, the instrument motion detection module 200 specifically includes: Deep learning target detection algorithms, such as YOLOv8 and Faster R-CNN, are used to detect surgical instruments in intraoperative video frames and annotate their categories, bounding box locations, and confidence levels.

[0038] Optionally, in order to further improve the robustness of detection, a multimodal method that fuses color, shape, and lighting features can be used to avoid recognition errors caused by factors such as lighting changes and instrument overlap.

[0039] Furthermore, the optical flow method is used to calculate the movement direction and speed of surgical instruments in continuous frames, and the Kalman filter is used for trajectory prediction to reduce detection errors and improve the continuity of the motion trajectory. The trajectory points of surgical instruments are recorded to construct a motion trajectory dataset for subsequent optimization of the encoding strategy.

[0040] It should be noted that the optical flow method is a computer vision technology commonly used to calculate pixel motion in an image sequence. The core idea is to infer the direction and speed of an object's motion by analyzing pixel changes between adjacent frames. Currently, the Lucas-Kanade optical flow method and the Farneback optical flow method are the two most commonly used methods, the former is suitable for sparse feature point tracking, and the latter is suitable for dense optical flow calculation. In surgical scenes, the movement of surgical instruments is more complex, and it is necessary to accurately track the trajectory of the instrument. Therefore, the present invention uses sparse optical flow to track feature points. In the present invention, in telemedicine video transmission, the motion trajectory of surgical instruments is an important factor affecting video encoding. Since the movement of instruments is usually fast and fine, directly using ordinary frame difference methods may cause tracking instability or excessive errors. Therefore, the use of the optical flow method can effectively calculate the direction and speed of motion of instruments in continuous frames, providing basic data for subsequent motion trajectory prediction and coding optimization.

[0041] Kalman filtering is a recursive algorithm used to estimate the state of a dynamic system. It can smoothly predict the state (position, speed, etc.) of a target in the presence of noise. Since the optical flow calculation process may be affected by factors such as changes in ambient light and deformation of the instrument, large errors may occur in the motion estimation of some frames. Directly using the original optical flow data for encoding optimization may affect the stability of telemedicine videos due to jitter or errors. Therefore, trajectory prediction through Kalman filtering can effectively remove abnormal data, making the motion trajectory of surgical instruments smoother, thereby improving the encoding quality. Kalman filtering can also smooth errors caused by factors such as changes in light and occlusion during the optical flow calculation process, thereby avoiding jitter in the motion trajectory and improving the stability of the encoded data.

[0042] Preferably, the motion matrix is ​​composed of a two-dimensional weight matrix, which represents the distribution of the encoded weights of the surgical instrument at different time points, and the construction includes the following steps: According to the motion trajectory data set, the movement speed of the instrument is first calculated. The calculation process is as follows: record the coordinates of the instrument at the current moment and the coordinates of the previous moment; calculate the horizontal position difference between the current moment and the previous moment, that is, the change in the horizontal coordinate; calculate the vertical position difference between the current moment and the previous moment, that is, the change in the vertical coordinate; square the above two changes respectively and add them together; take the square root of the result of the addition to obtain the displacement distance between the two moments; calculate the ratio of the displacement distance to the time interval between frames, which is the movement speed.

[0043] Secondly, calculate the acceleration of the device. The calculation process is as follows: calculate the difference between the current speed and the previous speed; calculate the ratio of the speed change to the time interval between frames, which is the acceleration; use the acceleration information to determine whether the device is in a stable operation or sudden motion state, and adjust the encoding strategy accordingly.

[0044] It is best to determine the direction change of the instrument. The calculation process is as follows: calculate the change in the vertical coordinate between the current moment and the previous moment; calculate the change in the horizontal coordinate between the current moment and the previous moment; calculate the ratio of the change in the vertical coordinate to the change in the horizontal coordinate, and take the inverse tangent value to obtain the direction angle at the current moment; the degree of change in the direction angle is used to measure the uncertainty of the instrument operation. The greater the change, the greater the instability of the operation, and the need to improve the encoding accuracy of this area.

[0045] Set the horizontal and vertical coordinates of the motion matrix to correspond to the pixel coordinates in the video frame respectively; set the weight of the motion area, which is determined according to the detection result of the device; calculate the weight adjustment factor of the motion area, which is calculated based on the speed, acceleration and direction change. The specific method is as follows: The speed, acceleration and direction change are normalized and assigned different weight values; the ratio between the current speed and the set maximum speed is calculated to obtain the speed weight factor; the ratio between the current acceleration and the set maximum acceleration is calculated to obtain the acceleration weight factor; the ratio between the current direction change and the set maximum direction change is calculated to obtain the direction change weight factor; the above three weight factors are weighted and summed to obtain the motion weight factor, and the initial coding weight is obtained by calculating the product of the motion area weight and the motion weight factor.

[0046] Furthermore, since surgical instruments usually move continuously in a short period of time and have a short static time, directly using the same encoding strategy for all frames may cause inter-frame blurring of moving objects, so it is necessary to dynamically adjust the inter-frame weights based on historical trajectories. And if the frame rate or encoding accuracy is not properly increased when the instrument moves at high speed, it may cause the remote doctor to be unable to accurately judge the surgical operation, affecting the accuracy of remote guidance. By calculating the dynamic adjustment factor, the encoding quality can be improved at high speed and the bandwidth usage can be reduced at low speed or static. Specifically, according to the motion matrix and historical trajectory, the dynamic adjustment factor of the time dimension is calculated: the absolute change between the current moment's movement speed and the previous moment's movement speed is calculated, and divided by the set maximum speed value to obtain the normalized speed change ratio, and then the absolute change between the current moment's direction angle and the previous moment's direction angle is calculated, and divided by the set maximum angle value to obtain the normalized direction change ratio. Set two weight parameters, multiply the normalized speed change ratio and direction change ratio by the corresponding weight parameters respectively, and calculate the sum of the above two weighted ratios. The sum is the dynamic adjustment factor of the time dimension.

[0047] According to the dynamic adjustment factor, the initial coding weights at different time points are dynamically optimized: the product of the motion area weight and the dynamic adjustment factor is calculated, and appropriate adjustments are made based on the basic weight value to obtain the final coding weight; if the acceleration of the device increases (i.e., the movement is violent), the coding accuracy of the area is improved to reduce motion blur; if the movement of the device is relatively stable, the coding weight of the area is reduced to save bandwidth and generate the final coding weight.

[0048] The present invention further utilizes the dynamic adjustment factor of the time dimension to optimize the encoding strategy based on the historical trajectory. The present invention constructs the time dimension adjustment factor by calculating the speed change rate and direction change rate of adjacent frames, so that the encoding strategy can be dynamically adjusted with the movement state of the instrument, which can effectively improve the quality of remote medical videos and ensure the continuity and clarity of remote surgical operations.

[0049] In the process of telemedicine video transmission, it is necessary to ensure high-quality display of key surgical areas and optimize bandwidth usage to avoid excessive coding resources in secondary areas. Therefore, it is necessary to jointly optimize the anatomical features and surgical instrument movements in the video so that the system can dynamically adjust the coding accuracy to ensure the clarity and smoothness of telemedicine videos. Therefore, this module constructs anatomy-instrument dynamic weight matrix based on anatomical matrix and motion matrix to optimize the spatiotemporal joint coding strategy.

[0050] In traditional telemedicine video coding, the coding weights of all regions are usually uniform or simply weighted, without considering the complexity of the surgical scene. As a result, if the coding quality of the key areas of the anatomical structure (such as lesions, surgical approaches, etc.) is insufficient, the doctor may not be able to accurately identify the tissue characteristics, thereby affecting the accuracy of remote diagnosis and treatment or surgical guidance. If the coding quality of the motion area of ​​the surgical instrument is insufficient, it may cause blurred operations or broken frames, making remote guidance difficult. And because the surgical process is dynamic, the areas of interest at different time points may be different, so it is necessary to dynamically adjust the coding strategy according to the motion information, and traditional methods often cannot adapt to this demand. In order to solve the above problems, this module jointly optimizes the anatomical matrix and the motion matrix, dynamically adjusts the video coding accuracy according to the spatial and temporal dimensions, improves the visual quality of the key areas, and reduces the transmission load of the secondary areas. In the embodiment of the present application, the ADW matrix construction module 300 specifically includes: the construction of the ADW matrix is ​​based on the anatomical matrix and the motion matrix, and is generated by interactive mapping and iterative optimization, so that the coding accuracy can adapt to the changes in the surgical scene.

[0051] The ADW matrix construction module 300 includes an ADW matrix generation unit 301, which combines the anatomical matrix and the motion matrix to generate an ADW matrix for subsequent encoding processing; a unified hierarchical encoding unit 302, which performs hierarchical encoding based on the ADW matrix and sets different encoding precisions for different weight areas to improve the encoding quality of key areas and reduce the transmission load of secondary areas.

[0052] The ADW matrix generating unit comprises: Specifically, the anatomical matrix and the motion matrix are processed through three stages: mutual mapping, weight adjustment, and feature alignment, so that the features of the two are deeply combined in the spatial and temporal dimensions, and finally an ADW matrix suitable for telemedicine video optimization is formed. The specific operations are as follows: The anatomical matrix contains the spatial distribution information of the medical imaging area; the motion matrix describes the dynamic information of the surgical instrument in the time series; the mapping relationship is used for feature fusion, so that the regional priority of the anatomical matrix affects the dynamic weight of the motion matrix, and the changing trend of the motion matrix corrects the weight distribution of the anatomical matrix.

[0053] It should be noted that directly superimposing two matrices will cause the weights of certain areas to be weakened or amplified, so the weights must be influenced by each other through mapping relationships. High-priority areas in the anatomical matrix (such as surgical fields and key anatomical parts) need to be strengthened in the motion matrix so that these areas maintain high encoding accuracy when the motion changes.

[0054] The feature deviation after mapping is calculated to determine the coupling degree between anatomical features and motion features, and the weights are adjusted through iterative optimization.

[0055] For example, when a high-weight anatomical region changes greatly in the motion matrix, it means that the region should be given a higher dynamic coding weight; when the anatomical region corresponding to a frequent motion region has a low weight, it means that there may be feature mismatch and adjustment is needed. If the coupling degree is found to be insufficient, the weight adjustment is required to optimize the feature matching degree of the ADW matrix.

[0056] The weights can be gradually adjusted for areas with low coupling to avoid distortion caused by one-time adjustment. Overall adjustments can be made based on multiple frames of data to ensure that the adjusted weights still meet the requirements of the surgical scenario. The weights can be continuously adjusted in an iterative manner until the feature deviation is reduced to a reasonable range.

[0057] Furthermore, a feature transformation strategy is used to align anatomical features and motion features at the encoding level to avoid optimization failure due to scale mismatch. The feature dimension is reduced through principal component analysis to make the distribution of the ADW matrix more stable. The distribution of anatomical information and motion information is different. If they are directly merged, some areas will lose information or be over-enhanced. Therefore, feature alignment can ensure the optimization effect.

[0058] It should be noted that the present invention enables the movement trend of surgical instruments to affect the weight adjustment of the anatomical matrix through mutual mapping, improves the utilization rate of coding resources, and avoids the influence of motion blur on key parts on surgical operations. In the mapping process, anatomical features and motion features may be mismatched, so it is necessary to calculate feature deviations and continuously adjust weights through iterative optimization until the feature matching degree reaches the optimal state. If fixed weights are directly used to fuse anatomical and motion information, some important areas may be ignored, or coding may be unstable due to information conflicts. Therefore, this method gradually adjusts the coupling degree to ensure that the optimized ADW matrix can meet the needs of the surgical scene without causing distortion or bandwidth waste due to too fast adjustment. It also ensures that the anatomical information and motion information are consistent at the coding level through feature alignment. At the same time, dimensionality reduction processing reduces redundant information, making the optimization more efficient and stable.

[0059] The unified hierarchical coding unit specifically includes: Preferably, the operation process of unified layered coding includes two stages: coding adaptability analysis and frame internal and external correlation calculation.

[0060] Directly dividing the image into fixed levels will lead to uneven distribution of coding resources. Therefore, we first conduct coding adaptability analysis to ensure that the level distribution meets the image complexity. Specifically, the coding adaptability analysis includes: according to the priority classification in the anatomical feature analysis module, the surgical target area is set as a high priority area; the surgical instrument interaction area is set as a medium priority area; and the background and irrelevant areas are set as low priority areas. Calculate the spatial complexity of each area and motion complexity ,For example: in, It is the image brightness gradient, which measures the texture complexity.

[0061] If both the spatial complexity and motion complexity are high, the coding level will be increased first to ensure that the details of the area are clear; if the spatial complexity is low but the motion complexity is high, the coding level will be increased only when the motion is intense to avoid wasting bandwidth; if the spatial complexity is high but the motion complexity is low, high accuracy will be maintained in high-complexity areas within the frame, but inter-frame coding resources will be reduced.

[0062] If both the spatial complexity and motion complexity are high, the coding level will be increased first to ensure that the details of the area are clear; if the spatial complexity is low but the motion complexity is high, the coding level will be increased only when the motion is intense to avoid wasting bandwidth; if the spatial complexity is high but the motion complexity is low, high accuracy will be maintained in high-complexity areas within the frame, but inter-frame coding resources will be reduced.

[0063] Furthermore, the frame internal and external correlation calculation includes calculating the similarity of adjacent frames. : in, Pixel In the Frame and The similarity between frames, For the Frame image at pixel point The intensity value at For the Frame image at pixel point The intensity value at .

[0064] like If it is greater than the threshold, the encoding bit rate of the area is reduced to reduce repeated information; if it is less than the threshold, the encoding accuracy of the area is increased to retain more details.

[0065] Static tissue areas have complex textures but almost no motion. In traditional coding methods, these areas are often given higher coding resources, resulting in bandwidth waste. In the present invention, such areas still maintain high quality during intra-frame coding, but inter-frame coding resources can be appropriately reduced to improve coding efficiency.

[0066] In surgical videos, there is a high degree of similarity between many adjacent frames, but traditional encoding methods fail to fully utilize this feature, resulting in redundant information occupying bandwidth. The present invention calculates the similarity of adjacent frames and dynamically adjusts the encoding bit rate according to the similarity to improve the encoding efficiency.

[0067] In the embodiment of the present application, the parameter linkage adjustment module 400 specifically includes: Combined with the coding level mapping rule of the unified hierarchical coding unit, the real-time motion state of the surgical instrument is obtained at the same time, such as motion speed, rotation angle, etc. It should be noted that the instrument area within the surgical field of view is focused on and the coding level of the area is dynamically adjusted.

[0068] Since traditional coding strategies cannot adapt to the violent movement of surgical instruments, which may cause blurring or distortion of key parts, local coding optimization can enhance the coding quality of the areas around surgical instruments, ensuring that doctors can clearly observe the operation of the instruments.

[0069] Setting the coding enhancement factor of the instrument area depends on the real-time motion state of the surgical instrument and the similarity of adjacent frames. For example, for the current frame and the previous frame, the similarity of adjacent frames is calculated. If the similarity is high, it means that the current frame has changed little and there is no need to significantly adjust the coding level. If the similarity is low, it means that the changes between frames are large, and the coding level may need to be increased to reduce blur and loss of details. The motion vector field of the surgical instrument area is then extracted, and the local optical flow changes are calculated. The difference in optical flow vectors between pixels in adjacent frames indicates the degree of motion change in the surgical instrument area between two adjacent frames. If the optical flow changes significantly, it means that the surgical instrument has undergone significant displacement or rotation, which may cause blurred or distorted images, and the coding level needs to be increased to ensure video quality.

[0070] Furthermore, if the similarity between adjacent frames is high and the local optical flow changes little, it means that the picture content is basically stable, and the encoding level can be reduced to reduce bandwidth usage. However, if there is sudden noise in the current frame (such as instantaneous lighting changes or equipment occlusion), the similarity analysis results need to be filtered to avoid misjudgment and resulting in reduced image quality.

[0071] If the similarity is high but the local optical flow changes dramatically, it means that most of the picture is stable and only the surgical instrument area is moving rapidly. In this case, instead of adjusting the overall coding level, the regional coding enhancement strategy is adopted to improve the coding quality only in the area around the surgical instrument, reduce motion blur and edge artifacts, and avoid unnecessary coding enhancement of the background area.

[0072] If the similarity between adjacent frames is low and the local optical flow changes greatly, it means that the overall content of the picture has changed significantly, such as perspective switching, rapid movement of instruments, etc. In this case, the overall coding level should be increased to ensure that the surgical picture can quickly adapt to changes and reduce image distortion. However, this adjustment needs to be combined with the network bandwidth to avoid transmission delays or data loss due to too fast an increase in the coding level.

[0073] It should be noted that directly increasing or decreasing the coding level may cause sudden changes in the picture, so a smooth transition strategy is introduced during the adjustment process to make the coding level change more natural and avoid unnecessary flickering or discontinuity during video transmission. For example, before adjusting the coding level, the average optical flow change trend within a certain time window is calculated. If the change rate is fast, the coding level is increased step by step instead of a large adjustment at one time.

[0074] Optionally, after the coding level is adjusted, the similarity and optical flow changes of the next frame are evaluated immediately. If the ideal effect is still not achieved after adjustment, fine-tuning is performed. For example, if blurring still exists after the coding level is increased, local coding is appropriately strengthened; if bandwidth occupancy is not significantly reduced after the coding level is lowered, consider adjusting the transmission rate.

[0075] Preferably, by establishing a mapping relationship between the coding level and the transmission rate, the stability of the transmission can be improved while ensuring the image quality: if the coding level is increased, such as , then increase the transmission rate to ensure the stability of high-definition video; if the encoding level is reduced, such as , then reduce the transmission rate to avoid bandwidth waste, and ensure that the adjusted transmission rate does not exceed the network carrying capacity according to the current network available bandwidth. is the current level, and They are the previous level and the next level respectively.

[0076] The prior art usually adopts global coding optimization, that is, a unified coding strategy is adjusted for the entire video frame. However, in surgical videos, the background information is usually relatively stable, while the surgical instrument area changes dramatically. Therefore, global adjustment may lead to bandwidth waste or unnecessary image quality improvement. The present invention proposes a local coding enhancement strategy for the surgical instrument area, that is, when most of the picture is stable, but the surgical instrument area moves rapidly (adjacent frames have high similarity but the local optical flow changes dramatically), only the coding level of the area around the surgical instrument is increased, while the background area is kept at a low coding level. Through local coding enhancement, unnecessary bandwidth overhead can be reduced, while avoiding sudden changes or visual flicker caused by the overall coding level adjustment, making the video quality more stable. If the coding level adjustment is too drastic, it may cause video flickering, incoherence between frames, and even affect the doctor's judgment. Therefore, the present invention introduces a smooth transition strategy. When adjusting the coding level, the average optical flow change trend within a certain time window is calculated. If the change rate is fast, a step-by-step adjustment strategy is adopted instead of a one-time large-scale modification of the coding level. It not only improves the visual quality of telemedicine videos, but also optimizes data transmission efficiency, ensuring that doctors can clearly and accurately observe the operation of surgical instruments during remote surgery.

[0077] This embodiment also provides a computer device, which is suitable for a remote medical video transmission system for clinical real scenes, including a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute computer executable instructions to implement the remote medical video transmission system for clinical real scenes as proposed in the above embodiment.

[0078] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.

[0079] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the remote medical video transmission system for clinical real scenes as proposed in the above embodiment is implemented.

[0080] In summary, the present invention combines the anatomical matrix and the motion matrix, generates an ADW matrix through mapping, and performs unified hierarchical encoding based on the matrix, so that the key medical area obtains higher encoding accuracy, while the encoding resources of the background area or irrelevant area are appropriately reduced, thereby improving the overall quality and efficiency of telemedicine video transmission, and the encoding quality of the telemedicine video is refined and optimized. While ensuring high-precision encoding of the key surgical areas, it also effectively reduces bandwidth occupancy and improves the stability and adaptability of the telemedicine system.

[0081] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A telemedicine video transmission system for clinical real scenes, characterized in that: include: The anatomical feature analysis module is used to construct an anatomical matrix based on medical imaging data and to divide the video areas into spatial priorities according to clinical diagnosis and treatment scenarios; The instrument motion detection module is used to dynamically track surgical instruments and operation processes, construct motion matrices, and optimize the coding weights between different frames using dynamic adjustment factors in the time dimension; An ADW matrix construction module is used to combine the anatomical matrix and the motion matrix to generate an ADW matrix, and perform unified hierarchical coding based on the ADW matrix, setting different coding precisions for different weight regions; The parameter linkage adjustment module is used to optimize the coding level and transmission rate based on the movement status of the instrument to adapt to the dynamic changes of different surgical scenarios.

2. The remote medical video transmission system for clinical real scene according to claim 1, characterized in that: The anatomical matrix is ​​a two-dimensional or three-dimensional weight matrix constructed based on medical imaging data, in which each pixel corresponds to a weight value, which is used to indicate the importance of the area.

3. The remote medical video transmission system for clinical real scene according to claim 1, characterized in that: The spatial prioritization includes: A multi-scale pyramid model is used to divide the video into different levels, with different resolutions and frame rates for each level. Use scale-invariant feature transformation or optical flow analysis to calculate the stability of key regions between video frames and adjust the priorities according to dynamic changes.

4. The remote medical video transmission system for clinical real scene according to claim 1, characterized in that: The construction of the motion matrix includes: Calculate the movement speed, acceleration and direction change of the device; the horizontal and vertical coordinates correspond to the pixel coordinates in the video frame; set the weight of the movement area; Normalize the motion speed, acceleration and direction change, calculate the weight factors respectively, and perform weighted summation of the three weight factors to obtain the motion weight factor; The product of the motion region weight and the motion weight factor is used as the initial encoding weight; According to the motion matrix and historical trajectory, the dynamic adjustment factor of the time dimension is calculated, and the initial coding weights at different time points are dynamically optimized to obtain the final coding weights.

5. The remote medical video transmission system for clinical real scene according to claim 4, characterized in that: The ADW matrix construction module includes an ADW matrix generation unit, which combines the anatomical matrix and the motion matrix to generate the ADW matrix; a unified hierarchical coding unit, which performs hierarchical coding based on the ADW matrix and sets different coding precisions for different weight areas.

6. The remote medical video transmission system for clinical real scene according to claim 1, characterized in that: The generation of the ADW matrix includes three stages of processing: mutual mapping, weight adjustment, and feature alignment of the anatomical matrix and the motion matrix, so that the features of the anatomical matrix and the motion matrix are deeply combined in the spatial dimension and the temporal dimension.

7. The remote medical video transmission system for clinical real scene according to claim 6, characterized in that: The operation process of the unified hierarchical coding unit includes two stages: coding adaptability analysis and frame internal and external correlation calculation.

8. The remote medical video transmission system for clinical real scene according to claim 7, characterized in that: The encoding suitability analysis includes: According to the priority divided areas, calculating the spatial complexity and motion complexity of each area; The encoding strategy is dynamically adjusted according to the spatial complexity and motion complexity.

9. The remote medical video transmission system for clinical real scene according to claim 1, characterized in that: The optimization of the coding level includes: for the current frame and the previous frame, calculating the similarity of adjacent frames; extracting the motion vector field of the surgical instrument area and calculating the local optical flow change; The encoding level is dynamically adjusted according to the similarity and the local optical flow change.

10. The remote medical video transmission system for clinical real scene according to claim 9, characterized in that: The transmission rates include: Establish a mapping relationship between coding level and transmission rate: If the coding level is increased, the transmission rate will increase; if the coding level is decreased, the transmission rate will decrease; Based on the current available network bandwidth, ensure that the adjusted transmission rate does not exceed the network carrying capacity.