Mixed reality surgery training collaboration method and system supporting multi-role interaction

By implementing dynamic allocation of multi-role permissions, precise virtual-real registration, haptic feedback, and intelligent bandwidth allocation in the mixed reality surgical training system, the problem of low efficiency in multi-role collaboration has been solved, and the accuracy, security, and real-time performance of collaborative training have been improved.

CN121742712BActive Publication Date: 2026-04-28YUNNAN NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YUNNAN NORMAL UNIV
Filing Date
2026-02-28
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing mixed reality surgical training systems suffer from low efficiency in multi-role collaboration, fixed permissions that prevent flexible adjustments, insufficient registration accuracy between virtual models and real-time anatomical structures, lack of tactile feedback in gesture operations, and unreasonable data transmission security and network bandwidth allocation, all of which affect the accuracy and real-time performance of collaborative guidance.

Method used

By establishing a dynamic multi-role permission allocation mechanism based on the surgical stage, a dynamic virtual-real registration system based on multimodal fusion, a gesture interaction mapping enhanced by haptic feedback, a layered encrypted data security transmission, and bandwidth adaptive allocation based on deep Q networks, the system achieves automatic identification and permission adaptation, accurate registration and haptic feedback, differentiated encrypted data transmission, and intelligent bandwidth allocation for the surgical stage.

Benefits of technology

It significantly improves the accuracy, security, and real-time performance of mixed reality surgical training collaboration, ensures the flexibility of permissions and the stability of the collaborative system, enhances the immersiveness of interaction and operational efficiency, and guarantees the security of data and the quality of critical data streams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121742712B_ABST
    Figure CN121742712B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of medical data processing, and discloses a mixed reality surgery training collaboration method and system supporting multi-role interaction. The method comprises the following steps: realizing surgery stage identification and dynamically adjusting multi-role permissions through weighted fusion of videos and instrument features, realizing virtual-actual accurate registration through non-rigid interpolation deformation, enhancing interactive immersion through spatial distance and motion parameter driven haptic feedback, guaranteeing data security through differential encryption and chain hash storage, and intelligently allocating bandwidth and dynamically adjusting model precision through a deep Q network. The application solves the problems of fixed multi-role collaboration permissions, insufficient virtual-actual registration accuracy, lack of feedback for gesture operation, unsafe data transmission and unreasonable network bandwidth allocation in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical data processing technology, and in particular to a mixed reality surgical training collaboration method and system that supports multi-role interaction. Background Technology

[0002] Existing mixed reality surgical training systems overlay virtual 3D models onto the real surgical environment using head-mounted displays. Doctors observe the virtual models to assist in surgical operations. Some systems support remote experts to view the surgical footage via video calls and provide voice guidance. In addition, existing systems are equipped with basic gesture recognition functions, allowing doctors to control the display and hiding of virtual models through hand movements. These systems provide technical support for surgical training and remote consultations to a certain extent.

[0003] However, existing technologies suffer from low efficiency in multi-role collaboration. This is mainly manifested in the fixed operating permissions of different roles at different stages of surgery, making it impossible to flexibly adjust the intervention permissions of experts at critical stages. At the same time, the registration accuracy between the virtual model and the real-time anatomical structure is insufficient. When the patient's organs deform due to breathing and surgical operations during the operation, the virtual model cannot adjust synchronously, resulting in misalignment between the virtual and real, which affects the accuracy of guidance. In addition, gesture operation lacks tactile feedback, making it difficult for doctors to accurately control the rotation and scaling of the virtual model. Furthermore, there are security risks in the data transmission process and unreasonable network bandwidth allocation. When multiple data streams compete for limited bandwidth, the quality of critical video streams deteriorates.

[0004] Further analysis reveals that the core problem of multi-role collaboration stems from the inability to dynamically identify current operational needs and adaptively adjust the permissions of each role based on the surgical stage. This requires the system to have intelligent recognition capabilities during the surgical stage. However, even if dynamic permission allocation is achieved, the misalignment between the virtual model and the actual anatomical structure will still lead to the failure of collaborative guidance. Essentially, this requires establishing a precise registration mechanism between the preoperative model and the real-time intraoperative image. After registration, if doctors lack tactile feedback when operating the virtual model, they cannot perceive the force and contact status, which will seriously affect the accuracy of interaction and immersion. The lack of tactile feedback further exposes the problem of the lack of closed-loop feedback between gesture recognition and model interaction. At the same time, the massive amount of sensitive medical data generated by multi-role collaboration faces the risk of privacy leakage during network transmission, and data of different security levels are not subject to differentiated encryption protection. In addition, the limited network bandwidth cannot guarantee the transmission quality of key data streams when carrying multiple high-definition video streams, 3D model data, and voice communication due to the lack of an intelligent bandwidth allocation mechanism. This unreasonable bandwidth allocation directly affects the clarity of the images received by remote experts and the smoothness of model loading, ultimately restricting the practicality of the entire multi-role collaborative training system. Summary of the Invention

[0005] This application provides a mixed reality surgical training collaboration method and system supporting multi-role interaction. It addresses the problems of fixed multi-role collaborative permissions, insufficient virtual-real registration accuracy, lack of feedback for gesture operations, insecure data transmission, and unreasonable network bandwidth allocation in existing technologies by establishing a dynamic multi-role permission allocation mechanism based on surgical stage recognition, a dynamic virtual-real registration system based on multimodal fusion, gesture interaction mapping enhanced with haptic feedback, secure data transmission with layered encryption, and adaptive bandwidth allocation based on deep Q-networks. This solution significantly improves the accuracy, security, and real-time performance of mixed reality surgical training collaboration.

[0006] Firstly, this application provides a mixed reality surgical training collaboration method supporting multi-role interaction, the mixed reality surgical training collaboration method supporting multi-role interaction comprising:

[0007] Step S1: Extract endoscopic video features and instrument sequence features, perform weighted fusion, output the surgical stage probability distribution through a classifier, determine the current surgical stage when the maximum probability value exceeds the threshold, and generate a multi-role permission matrix based on the current surgical stage.

[0008] Step S2: Match the anatomical landmarks on the endoscope with the corresponding points on the preoperative model, select valid point pairs with matching distances that meet the requirements as deformation control points, perform non-rigid interpolation deformation on the vertex coordinates of the model, and obtain a registration model that is aligned with the real-time anatomical structure.

[0009] S3 step: Detect user gestures to obtain operation type and motion parameters, calculate the spatial distance between the gesture position and the surface of the registration model, determine the tactile feedback type based on the distance range and motion speed, and drive the actuator to output the corresponding vibration mode;

[0010] Step S4: Identify the data type of the transmission and mark the security level. Apply the corresponding encryption method to the data of different levels. Combine the multi-role permission matrix and the current surgical stage to determine the access request. Hash the access record and store it in a chain.

[0011] Step S5: Collect network bandwidth and buffer status to construct a state vector, select the bitrate adjustment strategy for each data stream through a deep Q network, and adjust the grid density of the registration model according to bandwidth conditions and user view distance.

[0012] Secondly, this application provides a mixed reality surgical training collaborative system that supports multi-role interaction, the mixed reality surgical training collaborative system supporting multi-role interaction comprising:

[0013] The weighting module is used to extract endoscopic video features and instrument sequence features, perform weighted fusion, output the surgical stage probability distribution through a classifier, determine the current surgical stage when the maximum probability value exceeds the threshold, and generate a multi-role permission matrix based on the current surgical stage.

[0014] The filtering module is used to match anatomical landmarks on the endoscope with corresponding points on the preoperative model, filter valid point pairs with matching distances that meet the requirements as deformation control points, and perform non-rigid interpolation deformation on the vertex coordinates of the model to obtain a registration model that is aligned with the real-time anatomical structure.

[0015] The output module is used to detect the user's gesture to obtain the operation type and motion parameters, calculate the spatial distance between the gesture position and the surface of the registration model, determine the tactile feedback type based on the distance range and motion speed, and drive the actuator to output the corresponding vibration mode.

[0016] The storage module is used to identify the data type of the transmission and mark the security level, apply corresponding encryption methods to data of different levels, and combine the multi-role permission matrix and the current surgical stage to determine the access request, and then hash the access record and store it in a chain.

[0017] The adjustment module is used to collect network bandwidth and buffer status to construct a state vector, select the bitrate adjustment strategy for each data stream through a deep Q network, and adjust the grid density of the registration model according to bandwidth conditions and user view distance.

[0018] Thirdly, a mixed reality surgical training collaboration device supporting multi-role interaction is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor invokes the instructions in the memory to cause the mixed reality surgical training collaboration device supporting multi-role interaction to execute the aforementioned mixed reality surgical training collaboration method supporting multi-role interaction.

[0019] Fourthly, a computer-readable storage medium is provided, wherein instructions are stored therein, which, when executed on a computer, cause the computer to perform the aforementioned mixed reality surgical training collaboration method supporting multi-role interaction.

[0020] The technical solution provided in this application extracts endoscopic video features and instrument sequence features, performs weighted fusion, and outputs a probability distribution of surgical stages via a classifier. When the maximum probability value exceeds a threshold, the current surgical stage is determined. Based on the current surgical stage, a multi-role permission matrix is ​​generated. This technical solution achieves automatic identification of surgical stages and dynamic adaptation of permissions, overcoming the problem of low collaborative efficiency caused by fixed permission configurations in existing technologies. In particular, the weighted fusion mechanism organically combines visual information with instrument usage information, making stage identification more accurate and reliable. The dynamic adjustment of permissions based on stages ensures that remote experts can intervene in time to provide annotation and guidance during critical surgical stages, avoiding the problems of insufficient or excessive permissions, and significantly improving the flexibility and targeting of multi-role collaboration. By matching anatomical landmarks on the endoscopic image with corresponding points on the preoperative model, selecting valid point pairs with matching distances that meet the requirements as deformation control points, and performing non-rigid interpolation deformation on the vertex coordinates of the model to obtain a registered model aligned with the real-time anatomical structure, this technical solution solves the problem of misalignment between virtual and real models caused by organ deformation during surgery. In particular, non-rigid interpolation deformation can adapt to the complex deformation of local organs, ensuring that the virtual model always maintains a precise correspondence with the actual anatomical structure. This provides accurate spatial reference for annotation and guidance by remote experts, avoids guidance deviations caused by registration errors, and significantly improves the accuracy of collaborative training. This technical solution, which detects user gestures to obtain operation type and motion parameters, calculates the spatial distance between the gesture position and the registered model surface, determines the tactile feedback type based on the distance range and motion speed, and drives the actuator to output the corresponding vibration mode, makes up for the lack of physical feedback in existing gesture interactions. This allows doctors to perceive tactile information such as proximity, contact, and resistance when operating virtual models, significantly enhancing the immersiveness of the interaction and the precise control of the operation. In particular, the feedback type determination based on spatial distance and motion parameters achieves delicate tactile simulation, enabling doctors to perform complex operations such as model rotation and scaling more naturally, improving the realism and operational efficiency of virtual surgical training.

[0021] By identifying the data type of transmission and marking its security level, applying corresponding encryption methods to different levels of data, and combining a multi-role permission matrix with the current surgical stage to determine access requests, a multi-layered protection system for secure data transmission is established by hashing and chaining access records. This overcomes the problems of insufficient security and lax access control in existing medical data transmission technologies. In particular, the differentiated encryption strategy adopts the national cryptographic algorithms SM4, AES-256, and ChaCha20 respectively to meet the different security requirements of core privacy data, sensitive medical data, and real-time operation data. This ensures both data confidentiality and real-time requirements. The dual access control based on the permission matrix and surgical stage ensures the legality and compliance of data access. The immutable audit chain built by the chained hash storage provides complete traceability of data access behavior, significantly improving the data security and trustworthiness of the collaborative system. This technical solution, which constructs a state vector by collecting network bandwidth and buffer status data, selects the bitrate adjustment strategy for each data stream using a deep Q-network, and adjusts the grid density of the registration model based on bandwidth conditions and user viewing distance, solves the problem of degraded quality of key data streams caused by unreasonable network bandwidth allocation in existing technologies. In particular, the deep Q-network can learn the complex mapping relationship between network state and bitrate strategy, realizing intelligent bandwidth allocation for multiple data streams. It automatically adjusts video resolution, frame rate, and model accuracy when network conditions fluctuate, prioritizing the clarity of the main endoscopic video stream to ensure that remote experts can obtain high-quality surgical images. At the same time, it dynamically adjusts the model grid density based on user viewing distance, reducing the data transmission burden while ensuring visual effects. This significantly improves the stability and smoothness of the collaborative system under limited bandwidth and ensures the reliability of real-time collaboration among multiple roles. Attached Figure Description

[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram of an embodiment of a mixed reality surgical training collaboration method supporting multi-role interaction in this application.

[0024] Figure 2 This is a schematic diagram illustrating the dynamic adjustment relationship between network bandwidth and video bitrate in an embodiment of this application;

[0025] Figure 3 This is a schematic diagram of an embodiment of a mixed reality surgical training collaboration system that supports multi-role interaction in this application.

[0026] Figure 4 This is a schematic block diagram of the structure of a mixed reality surgical training collaborative device that supports multi-role interaction in an embodiment of the present invention. Detailed Implementation

[0027] This application provides a mixed reality surgical training collaborative method and system that supports multi-role interaction. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0028] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of the mixed reality surgical training collaboration method supporting multi-role interaction in this application includes:

[0029] Step S1: Extract endoscopic video features and instrument sequence features, perform weighted fusion, output the surgical stage probability distribution through a classifier, determine the current surgical stage when the maximum probability value exceeds the threshold, and generate a multi-role permission matrix based on the current surgical stage.

[0030] Step S2: Match the anatomical landmarks on the endoscope with the corresponding points on the preoperative model, select valid point pairs with matching distances that meet the requirements as deformation control points, perform non-rigid interpolation deformation on the vertex coordinates of the model, and obtain a registration model that is aligned with the real-time anatomical structure.

[0031] S3 step: Detect user gestures to obtain operation type and motion parameters, calculate the spatial distance between the gesture position and the surface of the registration model, determine the tactile feedback type based on the distance range and motion speed, and drive the actuator to output the corresponding vibration mode;

[0032] S4 step: Identify the data type of the transmission and mark the security level. Apply the corresponding encryption method to the data of different levels. Combine the multi-role permission matrix and the current surgical stage to judge the access request. Hash the access record and store it in a chain.

[0033] S5 step: Collect network bandwidth and buffer state to construct state vector, select bitrate adjustment strategy for each data stream through deep Q network, and adjust grid density of registration model according to bandwidth conditions and user view distance.

[0034] It is understood that the executing entity of this application can be a mixed reality surgical training collaboration system supporting multi-role interaction, or it can be a terminal or a server; the specific implementation is not limited here. This application's embodiments use a server as an example for illustration.

[0035] Specifically, the original video frames are adjusted to a standard size of 224×224 pixels and input into the feature extraction network for processing. The feature extraction network employs a combination of convolutional and pooling layers. The convolutional layers perform convolution operations by sliding convolution kernels across the image to extract local feature patterns. The pooling layers reduce the dimensionality of the feature map, decreasing computation while preserving key feature information. After multiple convolutional and pooling operations, a 512-dimensional image feature vector is finally output. This vector contains a numerical representation of visual information such as organ morphology, blood vessel distribution, and tissue color in the endoscopic image. Instrument sequence feature extraction uses RFID tags to identify the type of surgical instrument currently in use. The identified instrument type is converted into a one-hot encoded form. One-hot encoding is implemented by creating a vector with a length equal to the total number of instrument types, setting 1 at the corresponding instrument type position, and setting 0 at the remaining positions, thus forming the instrument sequence feature vector. The feature fusion process involves weighting and concatenating image feature vectors and instrument sequence feature vectors according to preset weight coefficients. The video feature weight is set to 0.7, and the instrument feature weight is set to 0.3. The specific operation of weighted concatenation is to multiply each element of the image feature vector by 0.7 and each element of the instrument sequence feature vector by 0.3, then concatenate the two weighted vectors end-to-end to form a fused feature vector. This fused feature vector is input into three fully connected layers for classification. The first fully connected layer contains 256 neurons, the second layer contains 128 neurons, and the third layer contains 7 neurons corresponding to the seven surgical stages. The calculation process of the fully connected layers involves multiplying the input vector by the weight matrix, adding a bias term, and then performing a nonlinear transformation using the ReLU activation function. The output layer uses the Softmax function to convert the neuron output values ​​into a probability distribution. The Softmax function is calculated by taking the exponent of each neuron's output value and then dividing by the sum of the exponents of all neuron output values ​​to obtain the probability values ​​for the seven stages, with a sum of 1. When the probability value of a certain stage exceeds a threshold of 0.85, that stage is determined as the current surgical stage. The multi-role permission matrix is ​​generated based on a pre-defined role-operation mapping table for the current surgical stage. This table stores the permission values ​​for each operation for each role at different surgical stages. A permission value of 1 indicates the operation is allowed, while a permission value of 0 indicates the operation is prohibited. For example, in the tumor resection stage, the surgeon's CT viewing permission is set to 1, model rotation permission is set to 1, and annotation addition permission is set to 1. The remote expert's annotation addition permission is elevated to 1, granting temporary annotation permission. All operational permissions for the observing trainees are set to 0, granting them only viewing permissions.

[0036] Preoperative model generation begins with reading the DICOM file from the CT data. The DICOM file contains continuous CT slice images and their spatial location information. The moving cubes algorithm traverses the three-dimensional voxel space of the CT data, detecting the presence of organ surfaces within each voxel cell. This detection method involves determining if any of the eight vertices of the voxel cell has a CT value greater than a set threshold while the CT values ​​of adjacent vertices are less than the threshold. If this condition exists, it indicates that an organ surface has passed through that voxel cell. Based on the relationship between the CT values ​​of the eight vertices of the voxel cell and the threshold, the moving cubes algorithm queries a predefined triangular facet configuration table to determine the vertex positions and normal vectors of the generated triangular facets within that voxel cell. After traversing all voxel cells, all generated triangular facets are combined to form a triangular mesh model. This model contains vertex coordinate arrays and normal vector arrays, constituting the preoperative model. The extraction of anatomical landmarks from endoscopic images begins with grayscale processing of the current frame image. The RGB three-channel color image is converted to a single-channel grayscale image using the formula: grayscale value equals red channel value multiplied by 0.299, green channel value multiplied by 0.587, and blue channel value multiplied by 0.114. The grayscale image is then subjected to Gaussian filtering to remove noise. Gaussian filtering uses a 5×5 convolution kernel, with the weights at each position in the kernel calculated according to a two-dimensional Gaussian distribution. The preprocessed image is then input into the Harris corner detection algorithm. This algorithm calculates the corner response value for each pixel in the image. The corner response value is calculated based on the covariance matrix of the image gradient. The eigenvalues ​​of the covariance matrix reflect the degree of gradient change of the pixel in two orthogonal directions. When both eigenvalues ​​are large, the point is considered a corner. At blood vessel bifurcation points and organ edges, the image grayscale changes drastically, resulting in higher corner response values. Pixels with corner response values ​​exceeding a set threshold are selected to form a set of candidate landmark coordinates. Gradient direction histogram feature descriptors are extracted within a 16×16 pixel neighborhood of each candidate landmark coordinate. The neighborhood is divided into 4×4 sub-regions, and gradient histograms in 8 directions are calculated for each sub-region, forming a 128-dimensional feature vector. Feature matching calculates the Euclidean distance between each feature descriptor in the endoscopic image anatomical landmark set and the feature descriptor of the pre-calibrated corresponding point in the preoperative model. The Euclidean distance is calculated by subtracting the corresponding elements of the two 128-dimensional feature vectors, squaring the result, summing the results, and taking the square root. Matching point pairs with an Euclidean distance less than a threshold of 0.6 are selected as deformation control points. A thin-plate spline deformation function is constructed based on the coordinates of the deformation control points. The deformation function includes an affine transformation part and a nonlinear deformation part. The nonlinear deformation part uses radial basis functions, which are in the form of the square of the distance multiplied by the natural logarithm of the distance. The weighting coefficients of the radial basis function are solved by minimizing the registration error. The registration error is defined as the sum of squared Euclidean distances between the coordinates of the control points of the deformed preoperative model and the coordinates of the corresponding landmark points on the endoscopic screen, plus a regularization term to control the smoothness of the deformation.After obtaining the weight coefficients, the thin-plate spline deformation function is applied to the coordinates of all vertices in the preoperative model for interpolation deformation calculation, and the new coordinates of each vertex are calculated to form a registration model aligned with the real-time anatomical structure.

[0037] User gesture detection is based on hand images captured by the depth and infrared cameras built into the mixed reality headset. The hand tracking module processes the images to extract the 3D coordinates of key points on the hand skeleton. The hand skeleton model contains 21 key points, covering the joints of the palm and five fingers. Gesture type recognition is based on the spatial relationship between key points. The criteria for a pinch gesture are: the Euclidean distance between the tips of the thumb and index finger is less than 20 mm; the criteria for a translation gesture are: the displacement of the center point of the palm is greater than 10 mm between consecutive frames while the fingers remain relatively stationary; the criteria for a rotation gesture are: the angle change of the vector from the wrist to the tip of the middle finger is greater than 5 degrees between consecutive frames; and the criteria for a zoom gesture are: the rate of change of the distance between the center points of both palms exceeds a set threshold between consecutive frames. The recognized gesture type is recorded as an operation type identifier. Motion parameters are extracted from the consecutive frame data of the gesture. The rotational angular velocity is calculated by dividing the difference in rotation angle between the current frame and the previous frame by the time interval; the zoom speed is calculated by dividing the change in the distance between the hands by the time interval. Spatial distance calculation first obtains the 3D coordinates of the fingertip, then iterates through the coordinates of all vertices on the registered model surface, calculating the Euclidean distance between the fingertip coordinates and each vertex coordinate. The Euclidean distance is calculated by subtracting the corresponding components of the 3D coordinates, summing the squares, and then taking the square root. The minimum value among all distances is selected as the spatial distance between the gesture position and the model surface. Haptic feedback type determination is based on spatial distance and motion parameters. When the spatial distance is greater than 20 mm, it is determined to be proximity feedback; when the spatial distance is less than 5 mm, it is determined to be contact feedback; when the operation type is identified as rotation and the rotational angular velocity exceeds 30 degrees per second, it is determined to be resistance feedback; when the model is scaled to 50% or 200% of its original size, it is determined to be over-limit feedback. The feedback parameter table stores the vibration frequency and amplitude values ​​corresponding to different feedback types. The vibration frequency for proximity feedback is set to 100 Hz and the amplitude to 0.5 G; the vibration frequency for contact feedback is set to 175 Hz and the amplitude to 1.5 G; and the vibration frequency for resistance feedback is set to 150 Hz and the amplitude is dynamically calculated based on the rotational angular velocity. The tactile actuator driver generates a PWM control signal based on the acquired vibration frequency and amplitude values. The frequency of the PWM signal corresponds to the vibration frequency, and the duty cycle corresponds to the amplitude. The actuator drives the linear resonant actuator to generate the corresponding mechanical vibration according to the PWM signal.

[0038] Transmitted data type identification categorizes data by analyzing its content and format characteristics. Regular expression matching is used to identify patient identity information; regular expressions define text patterns for information such as name, ID number, and medical record number. Pattern matching is performed on the text content of the transmitted data, and when a match is successful, the data is marked as core privacy data. File header tag identification identifies medical images and reports. DICOM files contain specific file header tags; by reading the tag information in the file header, it is determined whether it is a CT image. Pathology reports are usually in PDF or DOC format, and the filename or content contains keywords such as "pathology" or "diagnosis." Encoding format identification identifies video and audio streams. H.264 and H.265 encoded video streams have specific encoding header identifiers, and Opus encoded audio streams have specific frame structures; the encoding format is determined by parsing the header information of the data packets. Classified data is labeled with a security level: core privacy data is marked as Level 1, sensitive medical data as Level 2, and real-time operational data as Level 3. The encryption algorithm selects the corresponding encryption method based on the security level identifier. Level 1 data uses the national standard SM4 algorithm, a block cipher that divides the plaintext into 128-bit blocks and uses a 128-bit key for 32 rounds of iterative encryption. Each round includes operations such as XOR, S-box substitution, and linear transformation. Level 2 data uses the AES-256 algorithm, which divides the plaintext into 128-bit blocks and uses a 256-bit key for 14 rounds of iterative encryption. The encryption process includes operations such as byte substitution, row shifting, column mixing, and round key addition. A digital watermark is embedded in the encrypted CT image data. The watermark embedding process first divides the CT image into 8×8 pixel blocks, performs a discrete cosine transform on each block to obtain frequency domain coefficients, and selects the intermediate frequency coefficients to modify the watermark bits. The modification method is to multiply the coefficient value by 1 and add the embedding strength multiplied by the watermark bit value. Level 3 data uses the ChaCha20 algorithm, a stream cipher algorithm that uses a 256-bit key, a 96-bit nonce, and a 64-bit counter to generate a keystream. The keystream is XORed with the plaintext to obtain the ciphertext. Access request judgment extracts the user's data access request information, including the user role identifier, the requested resource identifier, and the operation type. Based on the user role identifier, a multi-role permission matrix is ​​queried to obtain the permission value for the corresponding resource and operation. A permission value of 1 indicates access is allowed, and a permission value of 0 indicates access is denied. Simultaneously, the current surgical stage identifier is checked; some operations are only allowed in specific surgical stages. By combining the permission value and stage conditions, it is determined whether the access conditions are met. Access record hash storage concatenates the various information of the access record into a string. The string content includes the user identifier, resource identifier, operation type, and timestamp. The SHA-256 hash algorithm is applied to the concatenated string to calculate the hash value. The SHA-256 algorithm converts input data of arbitrary length into a 256-bit hash value.The hash value of the current record is concatenated with the hash value of the previous record. The SHA-256 algorithm is then applied to the concatenated result to calculate a new hash value, forming a chain structure. The hash value is stored as block data on the blockchain nodes. The blockchain adopts a consortium blockchain architecture, where multiple nodes jointly maintain the ledger. New blocks are added to the chain through a consensus mechanism, with a block generation interval of 5 seconds. Each block contains 100 to 200 access records.

[0039] Network status acquisition obtains the current available bandwidth value through the operating system's network interface. Available bandwidth is measured by sending test data packets within a short period, and the bandwidth value is calculated by dividing the amount of successfully transmitted data by the time. Buffer occupancy is obtained by querying the video decoder's buffer status; buffer occupancy equals the used buffer size divided by the total buffer size. The 3D model data transmission queue length records the number of model data packets waiting to be transmitted. Voice communication jitter is calculated by statistically analyzing the variance of the arrival time interval of audio data packets; jitter reflects the stability of network transmission. The stage code for the current surgical stage is obtained from the identification results of step S1; the seven surgical stages are coded as 1 to 7. The number of online users counts the number of user terminals currently connected to the collaborative system. The network packet loss rate is calculated by dividing the difference between the total number of sent data packets and the number of successfully received data packets by the total number of sent packets. The average round-trip time is measured by sending ICMP messages, from the time of sending to the time of receiving a response. The above nine parameters are combined into a state vector, with each element of the state vector corresponding to a parameter value. The forward propagation calculation of the deep Q-network inputs the state vector into the network's input layer, which contains nine neurons corresponding to a nine-dimensional state vector. The first fully connected layer contains 128 neurons. The output value of each neuron is equal to the dot product of the input vector and the neuron's weight vector, plus the bias value. The dot product is calculated by multiplying corresponding elements and then summing them. The output value is processed by the ReLU activation function, which sets negative values ​​to 0 and positive values ​​to remain unchanged. The second fully connected layer contains 64 neurons, receiving the output of the first layer as input and undergoing the same weight multiplication, addition, and activation processing. The number of neurons in the output layer equals the number of bitrate adjustment strategy combinations. The output layer directly outputs the Q-value corresponding to each strategy combination. The Q-value represents the expected cumulative reward obtained by choosing that strategy in the current state. The strategy combination with the largest Q-value is selected as the current bitrate adjustment strategy. The bitrate adjustment strategy includes multiple parameters: the resolution parameters for the main video stream have options such as 1920×1080 and 1280×720; the frame rate parameters have options such as 30 frames per second and 15 frames per second; the bitrate parameters for the auxiliary video stream have options such as 3Mbps, 2Mbps, and 1Mbps; and the transmission bandwidth parameters for the 3D model data have options such as 1Mbps and 0.5Mbps. The encoding configuration of each data stream is adjusted according to the parameter values ​​in the strategy. The video encoder re-encodes the video stream based on the resolution and frame rate parameters, and the transmission control module allocates network bandwidth based on the bitrate parameters. Mesh density adjustment first acquires the pose data of the user's head-mounted display (HUD), which includes the HUD's position coordinates and orientation vector. The Euclidean distance between the HUD position and the center point of the registration model is calculated as the viewing distance. When the viewing distance is less than 50 cm and the available bandwidth is greater than 2 Mbps, a high-precision mesh is selected. The high-precision mesh contains 500,000 triangular faces.When the viewing distance is between 50 cm and 100 cm or the available bandwidth is between 1 Mbps and 2 Mbps, a medium-precision grid is selected, containing 200,000 triangular faces. When the viewing distance is greater than 100 cm or the available bandwidth is less than 1 Mbps, a low-precision grid is selected, containing 50,000 triangular faces. A mesh simplification algorithm processes the registration model. The simplification algorithm selects edges to be folded based on the cost function of edge folding. The cost function calculates the geometric error introduced after folding the edge, prioritizing the folding of edges with lower costs. This edge folding operation is repeated until the target number of faces is reached.

[0040] For example, in a thoracoscopic lobectomy training session, the endoscopic camera captured images containing lung tissue and blood vessels. The video frames were input into a feature extraction network to generate a 512-dimensional image feature vector. Simultaneously, an RFID tag identified that the surgeon was using an ultrasonic scalpel, with the instrument type coded as a one-hot vector. The two feature vectors were weighted and concatenated with weights of 0.7 and 0.3, respectively, and then fused. The resulting feature vector was input into a fully connected classification network. The Softmax output showed a probability of 0.91 for the tumor resection stage, exceeding the threshold. The system determined that the tumor resection stage was currently underway and automatically adjusted the permission matrix to grant annotation permissions to remote experts. Preoperative CT data was used to generate a triangular mesh model containing lung lobes and blood vessels using a moving cubes algorithm. After grayscale conversion and Gaussian filtering of the current endoscopic frame, Harris corner detection detected a landmark point with a corner response value of 0.82 at the pulmonary vessel bifurcation. A 128-dimensional feature descriptor for this point was extracted, and its Euclidean distance with the pre-labeled vessel bifurcation feature descriptor in the preoperative model was calculated to be 0.54, less than the threshold. This matching point pair was used as a deformation control point. Based on 15 extracted deformation control points, the thin-plate spline deformation function calculates the weighting coefficients of the radial basis function, performing non-rigid deformation on all vertices of the preoperative model to generate a registration model aligned with the real-time lung tissue morphology. The surgeon wears haptic gloves, and hand tracking detects a pinching gesture between the thumb and forefinger. The distance between the fingertip coordinates and the nearest vertex on the registration model surface is 3 mm, less than the 5 mm threshold, indicating a contact feedback type. The system queries the feedback parameter table and obtains a vibration frequency of 175 Hz and an amplitude of 1.5 G, driving the linear resonant actuator in the glove to output the corresponding vibration, allowing the surgeon to feel the tactile sensation of grasping the virtual model. When the surgeon rotates the model to view the hilar structure, the system detects a rotational angular velocity of 45 degrees per second, exceeding the 30 degrees per second threshold, indicating a resistance feedback type. The vibration frequency is adjusted to 150 Hz with an amplitude calculated based on the angular velocity of 1.35 G, providing a tactile simulation of rotational resistance. Data transmission during surgery includes the patient's medical record number, intraoperative CT images, and endoscopic video stream. Regular expression matching identifies the medical record number as core privacy data (Level L1), DICOM file headers identify CT images as sensitive medical data (Level L2), and H.265 encoding identifies the video stream as real-time operational data (Level L3). The medical record number is encrypted using the SM4 algorithm, with a 128-bit key established via the SM2 key exchange protocol. Plaintext blocks are iterated and encrypted 32 times to generate ciphertext. The CT images are encrypted using the AES-256 algorithm, with a 256-bit key derived from the user's password through PBKDF2 iterations 100,000 times. After encryption, a 64-bit watermark containing the user ID and timestamp is embedded in the DCT field. The video stream is encrypted using the ChaCha20 algorithm, with a 256-bit key and a 96-bit nonce generating a keystream that is XORed with the video data.A remote expert initiates a request to access CT images. The system extracts the remote expert's role identifier and queries the permission matrix, finding that the role's CT viewing permission value is 1 during the tumor resection stage. The current stage is determined to be tumor resection, and the access request is granted, generating an access record containing the user ID, CT image resource ID, viewing operation, and current timestamp. The record information is concatenated into a string, and a SHA-256 hash value is calculated to obtain a 256-bit hash. This hash value is concatenated with the hash value of the previous record and then recalculated using SHA-256. The new hash value is stored as block data, and blockchain nodes add the new block to the chain within 5 seconds using the PBFT consensus algorithm. During the surgery, network status monitoring shows that the current available bandwidth is 3.2 Mbps, the main video stream buffer occupancy rate is 0.65, the auxiliary video stream buffer occupancy rate is 0.42, the model data queue length is 8, the voice jitter value is 12 milliseconds, the current surgical stage encoding is 5, the number of online users is 4, the packet loss rate is 0.02, and the average round-trip latency is 25 milliseconds. These nine parameters form the state vector input to the deep Q-network. The input layer receives a nine-dimensional vector. The 128 neurons in the first fully connected layer calculate the dot product of the input vector and its respective weight vector, add a bias, and then activate using ReLU. The 64 neurons in the second layer receive the output from the first layer and continue calculation. The 25 neurons in the output layer output the Q-values ​​of each strategy combination. The strategy with the highest Q-value is the main video at a bitrate of 6 Mbps to maintain a resolution of 1920×1080, the auxiliary video at a bitrate of 2 Mbps, and the model data at a bandwidth of 1 Mbps. The video encoder adjusts the encoding configuration according to the strategy parameters, and the transmission control module allocates bandwidth resources. The head-mounted display pose data shows that the distance between the user's position and the center point of the registration model is 60 cm, and the available bandwidth is 3.2 Mbps, which is greater than 2 Mbps. However, since the distance of 60 cm is between 50 and 100 cm, the system selects a medium-precision mesh. The mesh simplification algorithm performs edge folding operations on the registration model, calculates the geometric error cost of folding each edge, and prioritizes folding edges with lower costs, gradually simplifying the 500,000-face patch to 200,000 faces. The simplified model is then transmitted to the user's head-mounted display for rendering and display.

[0041] In one specific embodiment, step S1 includes:

[0042] After adjusting the single-frame image of the endoscope video stream to a preset size, it is input into a feature extraction network containing convolutional and pooling layers for layer-by-layer convolutional operations and feature dimensionality reduction to obtain a fixed-dimensional image feature vector.

[0043] Identify the device type at each time step in the device usage sequence, map each device type to a one-hot encoded form, and obtain the device sequence feature vector;

[0044] The image feature vector and the instrument sequence feature vector are concatenated element-wise according to the video feature weights and instrument feature weights to obtain the fused feature vector.

[0045] The fused feature vectors are sequentially transformed and activated by three fully connected layers. The output layer is calculated by the Softmax function to obtain the probability distribution of seven stages: preoperative preparation, incision, separation, hemostasis, tumor resection, suturing and postoperative examination. When the maximum probability exceeds the set threshold, the corresponding stage is determined as the current surgical stage.

[0046] Based on the role-operation mapping table preset in the current surgical stage index, the surgeon, assistant surgeon, remote expert, and observer are respectively assigned permission values ​​for CT viewing, model rotation, annotation addition, video control, and data access, generating a multi-role permission matrix.

[0047] Specifically, single-frame image adjustment of the endoscopic video stream first reads the pixel data of the original video frame. The resolution of the original image is typically 1920×1080 pixels, containing three color channels: RGB. Image adjustment uses a bilinear interpolation algorithm to scale the original image to a preset size of 224×224 pixels. The bilinear interpolation calculation process involves mapping the coordinates of each pixel position in the target image back to the coordinates in the original image, and taking the weighted average of the four nearest pixels around that coordinate as the target pixel value. The weights are calculated inversely proportional to the distance between the coordinate and the centers of the four pixels. The adjusted image is then input into a feature extraction network, which contains a combination of multiple convolutional and pooling layers. The convolutional layers process by sliding fixed-size convolutional kernels across the image. The convolutional kernels are typically 3×3 or 5×5 matrices containing learnable weight parameters. At each sliding position, the convolutional kernel is element-wise multiplied with the pixel values ​​of the corresponding region in the image, summed, and then a bias value is added to obtain the value at the corresponding position in the output feature map. Convolutional operations extract local feature patterns from an image. Shallow convolutional layers extract low-level features such as edges and textures, while deep convolutional layers extract high-level semantic features such as organ morphology and blood vessel distribution. Pooling layers reduce the dimensionality of the feature map. Max pooling selects the maximum value within a local region of the feature map as the output, while average pooling calculates the average value within that region as the output. The pooling window is typically 2×2 with a stride of 2, halving the size of the feature map. Through layer-by-layer processing of convolutions and pooling, the spatial size of the image gradually decreases while the number of feature channels gradually increases. Finally, global average pooling or flattening operations convert the feature map into a one-dimensional vector, resulting in a 512-dimensional image feature vector.

[0048] The instrument usage sequence records the types of instruments used at each time point during the surgery. Instrument type identification is achieved through RFID tag readers. Each surgical instrument is attached to an RFID tag, which stores a unique identifier for the instrument type. When the doctor retrieves an instrument, the reader automatically reads the tag information and identifies the type of instrument currently in use, such as an electrosurgical unit, hemostat, stapler, or retractor. Instrument types are mapped to one-hot encoding, a vector representation method. Assuming the total number of possible instrument types used during surgery is N, a vector of length N is created, where all positions except the position corresponding to the current instrument type are set to 0. For example, if the electrosurgical unit is numbered 1, the hemostat is numbered 2, and the stapler is numbered 3, when the electrosurgical unit is identified, the one-hot encoded vector is [1, 0, 0, ..., 0], and when the hemostat is identified, the vector is [0, 1, 0, ..., 0]. This encoding method converts category information into a numerical vector, maintaining equidistant spacing between different instrument types and avoiding the introduction of sequential relationships. The instrument sequence feature vector is the one-hot encoded vector of the instrument type at the current moment.

[0049] The weighted concatenation of image feature vectors and instrument sequence feature vectors begins with element-wise weighting of both vectors. Each element of the image feature vector is multiplied by a video feature weight, set to 0.7, indicating that image information dominates the surgical stage assessment. Each element of the instrument sequence feature vector is multiplied by an instrument feature weight, set to 0.3, indicating that instrument usage information serves as an auxiliary assessment criterion. The weighting calculation involves multiplying each element of the vector by its corresponding weight value, resulting in two new weighted vectors. The concatenation operation joins the two weighted vectors end-to-end to form a longer vector. Assuming the image feature vector has a dimension of 512 and the instrument sequence feature vector has a dimension of 20, the concatenated fused feature vector has a dimension of 532. This fused feature vector comprehensively incorporates both visual and instrument usage information, with both types of information contributing to subsequent stage classifications according to their predetermined weights.

[0050] The fused feature vector input is used to classify and identify surgical stages using three fully connected layers. The first fully connected layer contains 256 neurons, each connected to all elements of the fused feature vector. The neuron's computation involves multiplying each element of the input vector by its corresponding weight parameter, summing all products, and adding the neuron's bias parameter to obtain the neuron's pre-activation output value. This pre-activation output value undergoes a non-linear transformation using the ReLU activation function. The ReLU function is defined as output equal to input when the input value is greater than 0, and output 0 when the input value is less than or equal to 0. This non-linearity enhances the network's expressive power. The 256 neurons in the first fully connected layer output form a 256-dimensional intermediate feature vector. The second fully connected layer contains 128 neurons, receiving the 256-dimensional output from the first layer as input, and performing the same weight multiplication and addition, bias addition, and ReLU activation processing to output a 128-dimensional higher-level feature vector. The third fully connected layer contains 7 neurons, corresponding to the seven surgical stages: preoperative preparation, incision, separation, hemostasis, tumor resection, suturing, and postoperative examination. Each neuron in the third layer also undergoes weight multiplication and addition, and bias addition, but ReLU activation is not used; the output value is directly output. The seven values ​​of the output layer are converted into a probability distribution using the Softmax function. The calculation steps of the Softmax function are: first, take the natural exponent of the output value of each neuron; then, calculate the sum of all exponent values; finally, divide each exponent value by the sum to obtain a normalized probability value. The sum of the seven probability values ​​equals 1, and each probability value represents the likelihood of the current input corresponding to that surgical stage. The maximum value among the seven probability values ​​is selected. When the maximum probability value exceeds a set threshold of 0.85, the current surgery is determined to be in the corresponding stage, and that stage is identified as the current surgical stage.

[0051] Once the current surgical stage is determined, a pre-defined role-operation mapping table is indexed based on the stage identifier. This mapping table is a predefined data structure that stores the permission configuration rules for each user role and operation at different surgical stages. The mapping table is organized as a three-dimensional matrix: the first dimension corresponds to the surgical stage, the second to the user role, and the third to the operation type. User roles include four categories: surgeon, assistant surgeon, remote expert, and observer / trainee. Operation types include five categories: CT viewing, model rotation, annotation addition, video control, and data retrieval. Each position in the mapping table stores a permission value; a permission value of 1 indicates that the role is allowed to perform the operation at that stage, while a permission value of 0 indicates that execution is prohibited. Based on the stage number of the current surgical stage, the permission configuration data for that stage is extracted from the mapping table. The four user roles are traversed, and permission values ​​for the five operations are assigned to each role, forming a multi-role permission matrix. The multi-role permission matrix is ​​a 4×5 two-dimensional matrix, where rows correspond to user roles, columns correspond to operation types, and matrix elements are permission values. For example, during the hemostasis phase, the surgeon's CT viewing permission is temporarily set to 0, as this phase requires focus on hemostasis. Model rotation permission is set to 1, annotation addition permission to 1, video control permission to 1, and data access permission to 1. The remote expert's annotation addition permission is upgraded from the usual 0 to 1, granting temporary annotation permissions for remote guidance of the hemostasis operation. Other permissions are configured according to the mapping table. The permission values ​​for assistant doctors and observing trainees are also assigned according to the mapping table configuration for this phase. The generated multi-role permission matrix is ​​used in subsequent data access control to determine whether a user's access request meets the permission conditions for the current phase.

[0052] In one specific embodiment, step S2 includes:

[0053] The moving cube algorithm was used to extract the organ surface from the preoperative CT data, generating a triangular mesh containing vertex coordinates and normal vectors to obtain the preoperative model;

[0054] The feature detector is applied to the current frame of the endoscopic video stream to extract anatomical landmarks of blood vessel bifurcation points and organ edges, and the feature descriptor of each landmark is calculated to obtain the set of anatomical landmarks in the endoscopic image.

[0055] Each feature descriptor in the endoscopic image anatomical landmark set is compared with the feature descriptor of the corresponding pre-calibrated point in the preoperative model using Euclidean distance calculation. Matching point pairs with distance values ​​less than a set threshold are filtered out to obtain a set of effective point pairs as deformation control points.

[0056] Based on deformation control points, a thin plate spline deformation function is constructed. After calculating the weight coefficients of the radial basis function, non-rigid interpolation deformation calculation is performed on the coordinates of all vertices in the preoperative model to obtain a registration model aligned with the real-time anatomical structure.

[0057] Specifically, the process of extracting organ surface data using the moving cubes algorithm from preoperative CT data begins with reading the DICOM format CT scan file. The DICOM file contains two-dimensional CT slice images arranged in layers, along with the spatial location and pixel spacing information of each slice. By reading this information, a three-dimensional voxel space is constructed. Each voxel in the voxel space stores the CT value at its corresponding location, and the CT value reflects the density information of the tissue. The moving cubes algorithm traverses each cube cell in the voxel space. Each cube cell consists of eight adjacent voxel vertices. The algorithm checks the relationship between the CT values ​​of these eight vertices and a preset threshold. The threshold is usually set as the boundary between the CT values ​​of the organ tissue and the surrounding tissues. For example, the threshold for lung tissue is set to -600 HU, and the threshold for liver tissue is set to 50 HU. When the CT values ​​of some vertices in a cube cell are greater than a threshold while the CT values ​​of others are less than a threshold, the organ surface is determined to pass through that cube cell. An 8-bit binary code is formed based on the relationship between the CT values ​​of the eight vertices and the threshold. This code serves as an index to query a predefined table of 256 triangular facet configurations. The configuration table stores the topological structure of the triangular facets that should be generated within the cube cell in each case, including the positions of the vertices of the triangular facets on the cube edges. The precise positions of the triangular facet vertices are calculated using linear interpolation. Interpolation is performed between the two endpoints of a cube edge based on the ratio of the distance between the CT values ​​of the endpoints and the threshold. The interpolation calculation involves adding the edge's direction vector to the coordinates of the edge's starting point and multiplying it by an interpolation coefficient. The interpolation coefficient is equal to the difference between the starting point's CT value and the threshold divided by the difference between the starting and ending point's CT values. After calculating the coordinates of each triangular facet vertex, a normal vector is calculated based on the coordinates of the three vertices of the triangular facet. The normal vector is calculated using the cross product of two edge vectors: the second vertex minus the first vertex, and the third vertex minus the first vertex. The cross product result is normalized to obtain a unit normal vector, which points to the outer surface of the organ. After traversing all the cube units, the vertex coordinates and normal vectors of all generated triangular faces are stored in an array. The vertex coordinate array records the three-dimensional spatial position of each vertex, and the normal vector array records the normal vector direction of each vertex. These data together constitute the triangular mesh representation of the preoperative model.

[0058] The current frame of the endoscopic video stream is used to extract anatomical landmarks using a feature detector. First, the RGB color image is converted to grayscale, transforming the three-channel color information into single-channel brightness information. A weighted averaging method is used: the red channel is multiplied by 0.299, the green channel by 0.587, and the blue channel by 0.114 to obtain the grayscale value. The weighting coefficients are determined based on the human eye's sensitivity to different colors. The grayscale image is then smoothed using a Gaussian filter. The Gaussian filter uses a 5×5 convolution kernel, with the highest weight at the center and decreasing towards the edges according to a Gaussian distribution. The filtering operation applies the kernel to each pixel, multiplies the pixel value within the kernel's coverage area by the corresponding weight, sums the results, and replaces the original value of the center pixel. Gaussian filtering eliminates random noise in the image while preserving edge features. The preprocessed image is input into the Harris corner detection algorithm. The Harris algorithm calculates the corner response value of each pixel in the image. The corner response value reflects the gradient intensity of that pixel in two orthogonal directions. The calculation process first calculates the horizontal and vertical gradients of the image. The horizontal gradient is calculated using the Sobel operator, which is a 3×3 convolution kernel that performs weighted differences on the left and right neighbors of the pixel. The vertical gradient is also calculated using the Sobel operator to perform weighted differences on the upper and lower neighbors. A structure tensor matrix is ​​constructed for each pixel location. The elements of the matrix are the weighted sum of the gradient products within the pixel's neighborhood window, specifically the sum of the squares of the horizontal gradient, the sum of the squares of the vertical gradient, and the sum of the products of the horizontal and vertical gradients. The corner response value is calculated based on the determinant and trace of the structure tensor matrix. The determinant reflects the product of the gradients in the two directions, and the trace reflects the sum of the gradients in the two directions. The corner response value is equal to the determinant minus the square of the trace multiplied by an empirical coefficient of 0.04. At blood vessel bifurcation points and organ edges, image grayscale changes drastically in both directions, with corner response values ​​significantly higher than those in flat and edge areas. Pixels with corner response values ​​exceeding a set threshold are selected to form a candidate landmark coordinate set. A 16×16 pixel neighborhood window is extracted around each candidate landmark coordinate, and this window is divided into 4×4 sub-regions. A gradient direction histogram is calculated for each sub-region, and the gradient direction is quantized into 8 directional intervals. The cumulative sum of gradient magnitudes within each directional interval is calculated to form an 8-dimensional direction histogram. The histograms of the 16 sub-regions are connected to form a 128-dimensional feature descriptor. The feature descriptor is robust to image rotation, scaling, and illumination changes. The candidate landmark coordinates are associated with and stored with their corresponding feature descriptors to form a set of anatomical landmarks in the endoscopic image.

[0059] The feature descriptors in the endoscopic image's anatomical landmark set are matched with the feature descriptors of pre-calibrated corresponding points in the preoperative model. These pre-calibrated corresponding points are key anatomical locations marked on the 3D model preoperatively through expert interaction or automated algorithms. Each pre-calibrated point records its 3D coordinates and corresponding feature descriptor. The feature descriptor is generated by extracting feature descriptors from a slice or multi-planar reconstructed image of the preoperative CT data as a 2D image at the marked point location. The matching process iterates through each landmark in the endoscopic image's landmark set, taking its 128-dimensional feature descriptor and calculating the Euclidean distance between it and the feature descriptors of all pre-calibrated points in the preoperative model. The Euclidean distance is calculated as the square root of the sum of the squared differences between corresponding elements of two 128-dimensional vectors. This sum of squared differences equals the square of the difference between the first element of the first vector and the first element of the second vector, plus the square of the difference between the second element of the first vector and the second element of the second vector, and so on, accumulating 128 terms. Finally, the square root of the accumulated sum is taken to obtain the Euclidean distance. A smaller Euclidean distance value indicates that the two feature descriptors are more similar, and the corresponding anatomical location is more likely to be the same point. Matching point pairs with an Euclidean distance less than a set threshold of 0.6 are selected. This threshold is set based on a balance between the discriminative power of the feature descriptors and the accuracy of the matching. Too small a threshold results in an insufficient number of matching point pairs, while too large a threshold introduces false matches. The selected matching point pairs include the two-dimensional image coordinates of the landmark points in the endoscopic view and the three-dimensional model coordinates of the corresponding points in the preoperative model. These matching point pairs constitute the effective point pair set as deformation control points.

[0060] Deformation control points are used to construct the deformation function of thin-plate splines. Thin-plate splines are an interpolation method that calculates the deformation at any location in space using the constraints and smoothness constraints of the control points. The deformation function consists of an affine transformation part and a nonlinear deformation part. The affine transformation part describes the overall rotation, translation, and scaling, while the nonlinear deformation part describes local bending and twisting. The nonlinear deformation part uses radial basis functions, which are in the form of the square of the distance multiplied by the natural logarithm of the distance. For any point in 3D space, the Euclidean distance from that point to each deformation control point is calculated. The square of the distance multiplied by the logarithm of the distance is then multiplied by the weight coefficient corresponding to that control point. The weighted radial basis function values ​​of all control points are summed and added to the affine transformation part to obtain the deformed coordinates of that point. The weighting coefficients are solved by establishing a system of linear equations. The constraint of this system is that the output value of the deformation function at each deformation control point equals the target position of that control point. In other words, the coordinates of the control points in the preoperative model, calculated by the deformation function, equal the coordinates of the corresponding landmark points in the endoscopic image. The two-dimensional image coordinates of the landmark points in the endoscopic image are back-projected into three-dimensional space using a perspective projection model. Back-projection requires the camera's intrinsic and extrinsic parameter matrices. The intrinsic parameter matrix contains the focal length and principal point coordinates, while the extrinsic parameter matrix contains the camera's pose. The back-projection calculation maps two-dimensional pixel coordinates to three-dimensional ray directions, and combines depth information or stereo matching to determine the three-dimensional coordinates. The linear equation system also includes smoothness constraints. These constraints control the curvature of the deformation through a regularization term to avoid excessive bending. The regularization term is the square of the L2 norm of the weighting coefficients multiplied by a regularization parameter set to 0.01. Solving the linear equation system yields the weighting coefficients and affine transformation parameters corresponding to all deformation control points. The coordinates of all vertices in the preoperative model are subjected to non-rigid interpolation deformation calculation using deformation functions. Each vertex of the model is traversed, and the vertex coordinates are input into the deformation function. The distance from the vertex to all deformation control points is calculated. The radial basis function value is calculated based on the distance, multiplied by the weight coefficients and summed. The linear and translation parts of the affine transformation are added to obtain the deformed vertex coordinates. The deformed vertex coordinates replace the original coordinates. After all vertex coordinates are updated, a registration model aligned with the real-time anatomical structure is formed.

[0061] In one specific embodiment, a feature detector is applied to the current frame of the endoscopic video stream to extract anatomical landmarks of blood vessel bifurcation points and organ edges, and a feature descriptor is calculated for each landmark to obtain a set of anatomical landmarks in the endoscopic image, including:

[0062] The current frame of the endoscope video stream is grayscaled and Gaussian filtered to eliminate image noise, resulting in a preprocessed image.

[0063] The multi-scale Harris corner detection algorithm is applied to the preprocessed image to detect the pixel positions where the corner response value exceeds the set threshold in the blood vessel intersection area and organ contour edge, and to obtain a set of candidate marker coordinates.

[0064] Extract the gradient direction histogram within the neighborhood of each candidate marker point's coordinates, quantize the gradient magnitude and direction into a fixed-dimensional feature vector, and obtain the feature descriptor for each marker point;

[0065] The coordinates of candidate landmarks and their corresponding feature descriptors are associated and stored to obtain a set of anatomical landmarks in the endoscopic image.

[0066] Specifically, the grayscale processing of the current frame of the endoscopic video stream converts the RGB three-channel color image into a single-channel grayscale image. The conversion calculates the grayscale value of each pixel by multiplying the red channel value by 0.299, adding the green channel value by 0.587, and adding the blue channel value by 0.114. The weighting coefficients reflect the human eye's sensitivity to different colors, with the green channel having the highest weight because the human eye is most sensitive to green, followed by red, and then blue. The grayscale image then undergoes Gaussian filtering using a 5×5 convolution kernel. The kernel center has the highest weight, decreasing towards the edges according to a two-dimensional Gaussian distribution. The center weight is 0.375, the weight of the eight adjacent positions is 0.061, the weight of the twelve outer positions is 0.01, and the weight of the four corner positions is 0.002. The total weight is normalized to 1. The filtering operation places a convolution kernel at each pixel position in the image, multiplies the 25 pixel values ​​covered by the kernel with the weights at the corresponding positions of the kernel, and sums the 25 products to replace the original gray value of the center pixel. The filtering is completed by traversing all pixel positions in the image. Gaussian filtering smooths the image and eliminates random noise. Noise usually manifests as abnormal gray values ​​of isolated pixels. After neighborhood weighted averaging, abnormal values ​​are diluted by surrounding normal values. At the same time, the Gaussian distribution characteristics of the convolution kernel preserve the edge structure of the image. The gray value difference between adjacent pixels at the edge position is large. After weighted averaging, the edge still maintains a certain gradient change, resulting in a preprocessed image with noise eliminated.

[0067] The multi-scale Harris corner detection algorithm performs corner detection on image pyramids at different scales. The image pyramids are constructed by downsampling the original preprocessed image layer by layer. The first layer is the original resolution image; the second layer averages every 2×2 pixel blocks of the first layer image to generate an image with half the resolution; the third layer also downsamples the second layer image, constructing pyramids of 3 to 5 layers. Harris corner response values ​​are calculated on each image layer. The calculation process first calculates the horizontal and vertical gradients of the image. The horizontal gradient uses the Sobel operator, which is a 3×3 convolution kernel. The left column of the horizontal kernel is -1, -2, -1; the middle column is 0, 0, 0; and the right column is 1, 2, 1. The convolution operation calculates the horizontal gradient value at each pixel location. For the vertical gradient, the Sobel kernel has 1, 2, 1 on the top row, 0, 0, 0 on the middle row, and -1, -2, -1 on the bottom row, calculating the vertical gradient value. For each pixel location, a structure tensor matrix is ​​constructed. The elements in the first row and first column of the matrix are the sum of the squared horizontal gradients within the 5×5 neighborhood window of that pixel. The elements in the first row and second column, and the elements in the second row and first column, are the sum of the products of the horizontal and vertical gradients. The elements in the second row and second column are the sum of the squared vertical gradients. The determinant of the structure tensor matrix is ​​equal to the product of the elements in the first row and first column and the square of the elements in the second row and second column. The trace is equal to the sum of the elements in the first row and first column and the second row and second column. The corner response value is equal to the determinant minus the square of the trace multiplied by an empirical coefficient of 0.04. In the vascular intersection region, blood vessels intersect in two directions, and the image grayscale changes significantly in both the horizontal and vertical directions. Both the horizontal and vertical gradients are large, resulting in a large determinant and trace value in the structure tensor matrix, but the determinant grows faster, and the corner response value is large. At the edge of an organ contour, one side is organ tissue and the other side is a cavity or other tissue, with a significant difference in grayscale. The gradient is small along the edge direction and large perpendicular to the edge direction, and the corner response value is larger at the edge corners. Pixels in each image layer whose corner response values ​​exceed a set threshold are selected. The threshold is set based on the image contrast and noise level, and is usually five percent of the dynamic range of the response values. The corner positions detected in each layer are mapped back to the original resolution coordinates, and after merging and deduplication, a set of candidate marker coordinates is obtained.

[0068] For each candidate marker point, a feature descriptor is extracted from its coordinates. A 16×16 pixel square neighborhood region is extracted around the marker point, and this neighborhood region is divided into 16 sub-regions of size 4×4 pixels each. A gradient direction histogram is calculated within each sub-region, with the gradient direction ranging from 0 to 360 degrees divided into eight directional intervals, each covering a 45-degree range. The first interval is from 0 to 45 degrees, the second from 45 to 90 degrees, and so on. Each pixel within a sub-region is assigned to its corresponding interval based on its gradient direction. The gradient magnitude of that pixel is accumulated into the histogram value of the corresponding interval. The gradient magnitude is the square root of the sum of the squared horizontal and squared vertical gradients. The gradient direction is the arctangent of the vertical gradient divided by the horizontal gradient, with the arctangent ranging from -90 degrees to 90 degrees, adjusted to the 0 to 360-degree range based on the signs of the horizontal and vertical gradients. Each of the 16 sub-regions generates an 8-dimensional gradient direction histogram. The 16 histograms are connected in the row and column order of the sub-regions to form a 128-dimensional feature vector. The feature vector is normalized so that the Euclidean norm of the vector is equal to 1. The normalization is calculated by taking the square root of the sum of the squares of all elements of the vector as the norm. Each element is divided by the norm to obtain the normalized feature vector. Normalization eliminates the influence of illumination changes and obtains the feature descriptor of each marker point.

[0069] Candidate landmark coordinates and their corresponding feature descriptors are stored in a structured data format. Each landmark is recorded as a data structure, containing its two-dimensional pixel coordinates (x and y) in the image, corner response values, the image pyramid level of the landmark, and a 128-dimensional feature descriptor array. All candidate landmark data structures are stored in a dynamic array or list, forming a set of endoscopic landmarks. The landmark set is sorted in descending order of corner response values; landmarks with higher response values ​​have more significant features and higher matching reliability. In subsequent matching processes, landmarks with higher response values ​​are prioritized in the sorted set.

[0070] In one specific embodiment, step S3 includes:

[0071] The hand tracking module of the mixed reality headset detects the three-dimensional coordinates of key points of the user's hand bones, and identifies the type of gesture such as pinching, translation, rotation or scaling based on the spatial relationship of the key points of the fingers, and obtains the operation type identifier.

[0072] Calculate the Euclidean distance between the three-dimensional coordinates of the fingertip in the user's gesture and the coordinates of each vertex on the registered model surface, and select the minimum distance value as the spatial distance between the gesture position and the model surface;

[0073] When the spatial distance is greater than the first distance threshold, it is determined to be a proximity feedback type; when the spatial distance is less than the second distance threshold, it is determined to be a contact feedback type; when the operation type is identified as rotation and the rotational angular velocity exceeds the velocity threshold, it is determined to be a resistance feedback type, thus obtaining the tactile feedback type.

[0074] The system queries a preset feedback parameter table based on the tactile feedback type to obtain the corresponding vibration frequency and amplitude values, and drives the tactile actuator to output a vibration mode according to the vibration frequency and amplitude values.

[0075] Specifically, the hand tracking module of the mixed reality headset captures real-time image data of the user's hand using a built-in depth camera and an infrared camera. The depth camera emits infrared structured light or time-of-flight beams, calculating the depth value of each pixel based on the phase difference or time difference of the reflected light to form a depth image. The infrared camera captures a grayscale image of the hand. The hand tracking algorithm detects the hand region on the depth and grayscale images. Hand region recognition is based on the continuity of depth values ​​and grayscale texture features. Pixels with continuously varying depth values ​​within a certain range are clustered as candidate hand regions, and skin color features are detected within these candidate regions for further confirmation. After detecting the hand region, the algorithm extracts the 3D coordinates of key points of the hand skeleton. The hand skeleton model contains 21 key points, corresponding to the center of the palm, wrist, four joints of the thumb, four joints of the index finger, four joints of the middle finger, four joints of the ring finger, and four joints of the little finger. Keypoint detection employs a deep learning model. The model takes as input the stitched features of a depth image and a grayscale image, and outputs the two-dimensional coordinates of 21 keypoints in the image plane. Combined with the depth values ​​of the corresponding positions in the depth image, the two-dimensional coordinates are converted into three-dimensional spatial coordinates. This conversion uses the camera's intrinsic parameter matrix, which includes the focal length and principal point coordinates. The x-component of the three-dimensional coordinates is equal to the two-dimensional coordinate x minus the principal point x, multiplied by the depth value, and divided by the focal length. The y-components and z-components are calculated similarly, yielding the three-dimensional coordinates of the hand skeleton keypoints in the head-mounted display coordinate system. Gesture type recognition is based on the spatial relationship between keypoints. The criterion for a pinch gesture is that the Euclidean distance between the thumb tip keypoint and the index finger tip keypoint is less than a distance threshold. The Euclidean distance is calculated by subtracting the corresponding components of the three-dimensional coordinates of the two keypoints, summing the squares, and then taking the square root. When the distance is less than 20 millimeters, it is determined to be a pinch gesture. The criteria for a translation gesture are: the displacement of the palm center keypoint between consecutive frames is greater than a displacement threshold, and the finger keypoints remain relatively stationary relative to the palm center. The displacement is calculated by subtracting the palm center coordinates from the previous frame's coordinates from the current frame's coordinates. A translation gesture is defined as a gesture with a displacement magnitude greater than 10 mm and a distance between the finger keypoints and the palm center that changes less than 5 mm between consecutive frames. The criteria for a rotation gesture are: the angle between the vector from the wrist keypoint to the middle fingertip keypoint between consecutive frames changes greater than an angle threshold. The angle is calculated by dividing the vector dot product by the product of the vector magnitudes and taking the inverse cosine. A rotation gesture is defined as a gesture with an angle change greater than 5 degrees. The rotation axis direction is recorded as the cross product direction of the vectors from the two frames. The scaling gesture requires both hands. The criteria are: the rate of change of the distance between the palm center keypoints of the left and right hands between consecutive frames exceeds a rate of change threshold. The rate of change is calculated as the current frame distance minus the previous frame distance divided by the previous frame distance. A scaling gesture is defined as a gesture with an absolute rate of change greater than 10%. The identified gesture types are recorded as operation type identifiers: pinch is 1, translation is 2, rotation is 3, and zoom is 4.

[0076] The 3D coordinates of the fingertips in user gestures are determined based on the gesture type. Pinch and touch operations focus on the coordinates of the index fingertip, while rotation and translation operations focus on the coordinates of the center of the palm. The 3D coordinates of the corresponding key points are extracted from the key point data of the hand skeleton. The coordinates of each vertex on the registered model surface are stored in a vertex coordinate array, the length of which is equal to the number of vertices in the model. Each vertex in the vertex coordinate array is traversed, and the Euclidean distance between the fingertip coordinate and the vertex coordinate is calculated. The Euclidean distance is calculated by subtracting the square of the x-coordinate of the fingertip from the square of the vertex x-coordinate, adding the square of the difference in the y-coordinate, adding the square of the difference in the z-coordinate, and taking the square root of the sum. The minimum distance value is selected from all the distance values ​​between the vertex and the fingertip. The vertex corresponding to the minimum distance value is the point on the model surface that is closest to the fingertip. This minimum distance value is used as the spatial distance between the gesture position and the model surface. The spatial distance reflects the proximity of the user's hand to the virtual model. The smaller the distance, the closer the hand is to the model surface. A distance of zero means that the fingertip is exactly on the model surface.

[0077] The determination of haptic feedback type is based on spatial distance and operation type identifier. First, the relationship between the spatial distance and a first distance threshold is determined. The first distance threshold is set to 20 mm. When the spatial distance is greater than 20 mm, the hand is still some distance from the model surface but has entered the interaction range, which is determined to be proximity feedback, indicating that the user's hand is approaching the virtual object. When the spatial distance is less than a second distance threshold, it is determined to be contact feedback, which is set to 5 mm. When the distance is less than 5 mm, the hand is very close to or in contact with the model surface, and contact feedback simulates the tactile sensation of a finger touching an object surface. When the operation type identifier is 3, indicating a rotation gesture, the rotational angular velocity is calculated. The rotational angular velocity is equal to the change in rotational angle between consecutive frames divided by the time interval. The time interval is calculated based on the frame rate of the headset. At a frame rate of 90 frames per second, the time interval is 11.1 milliseconds. The angle change is obtained from the change in the vector angle during the gesture recognition process. When the rotational angular velocity exceeds the velocity threshold of 30 degrees per second, it is determined to be resistance feedback, which simulates the feeling of resistance when rotating the virtual object. When the model scales to its limit, it is considered an over-limit feedback type. Limit sizes include a minimum size of 50% of the original size and a maximum size of 200% of the original size. The current model size is recorded using a scaling factor; if the scaling factor is less than or equal to 0.5 or greater than or equal to 2, it is considered an over-limit, and the over-limit feedback alerts the user that the scaling limit has been reached. The haptic feedback type is determined by the priority order of conditional judgments: first, it checks if it is an over-limit feedback; second, it checks if it is resistance feedback; third, it checks if it is contact feedback; and finally, it checks if it is proximity feedback, thus obtaining the haptic feedback type for the current frame.

[0078] The feedback parameter table is a pre-stored data structure that records vibration parameters corresponding to different tactile feedback types. Each feedback type is associated with a set of vibration frequency and amplitude values. The proximity feedback type corresponds to a vibration frequency of 100 Hz and an amplitude of 0.5 G; the contact feedback type corresponds to a vibration frequency of 175 Hz and an amplitude of 1.5 G; the resistance feedback type corresponds to a vibration frequency of 150 Hz. The amplitude value is dynamically calculated based on the rotational angular velocity using the formula: base amplitude 0.8 G plus the angular velocity divided by 100. The angular velocity is in degrees per second. The over-limit feedback type corresponds to a pulse vibration mode, with a single pulse lasting 100 milliseconds and an amplitude of 2 G, a pulse interval of 50 milliseconds, and three consecutive pulses. Based on the determined tactile feedback type, the feedback parameter table is consulted to extract the corresponding vibration frequency and amplitude values. The tactile actuator is a linear resonant actuator that receives control signals to drive the internal mass block to vibrate. The control signal is generated using pulse width modulation (PWM). The frequency of the PWM signal is set to the vibration frequency value, and the duty cycle of the PWM signal is set to map the amplitude value to the range of 0 to 100%. The amplitude value is in gigabit (G), with 0.5G corresponding to a 25% duty cycle, 1.5G to a 75% duty cycle, and 2G to a 100% duty cycle. The haptic actuator generates mechanical vibrations of corresponding frequency and intensity based on the frequency and duty cycle of the PWM control signal. The vibrations are transmitted through the actuator housing to the surface of the haptic glove that contacts the user's skin, providing haptic feedback. Different frequencies and amplitudes of vibration produce different tactile sensations: low-frequency, low-amplitude vibrations produce a gentle sense of approach; high-frequency, medium-amplitude vibrations produce a distinct sense of contact; medium-frequency, high-amplitude vibrations produce a sense of resistance; and pulsed vibrations produce a sense of alertness. Haptic feedback enhances the user's immersion in the virtual model and improves their precise control.

[0079] In one specific embodiment, step S4 includes:

[0080] Content analysis is performed on the data to be transmitted. Patient identity information is matched with regular expressions and marked as core privacy data. CT images and pathology reports are identified by file header tags and marked as sensitive medical data. Video streams and audio streams are identified by encoding format and marked as real-time operation data, resulting in classified data with security level labels.

[0081] Different encryption algorithms are applied to the classified data according to the security level label. The core privacy data is encrypted end-to-end using the national cryptographic SM4 algorithm, the sensitive medical data is encrypted using the AES-256 algorithm and embedded with a digital watermark, and the real-time operation data is encrypted using the ChaCha20 algorithm to obtain the encrypted data stream.

[0082] After receiving a user's data access request, extract the user role identifier and operation type from the request, combine the permission value of the corresponding role in the multi-role permission matrix and the stage identifier of the current surgical stage, determine whether the access request meets the permission conditions, and allow access and generate an access record when the conditions are met.

[0083] The user identifier, resource identifier, operation type, and timestamp in the access record are hashed. The hash value is then concatenated with the hash value of the previous record and hashed again before being stored in a chain to the blockchain node.

[0084] Specifically, the content analysis of the data to be transmitted uses regular expressions to match patient identity information. The regular expressions define the name pattern as a continuous sequence of two to four Chinese characters, and the ID number pattern as an 18-digit or 17-digit number followed by the letter X. The data text content is scanned character by character, and when a string matching the pattern is found, the data is marked as core privacy data and recorded as L1 security level. For file header label recognition, the first 132 bytes of the file are read, and the offset from 128 to 131 bytes is checked for DICOM format (using the magic number DICM). The label area is then read to search for patient name labels (group 0010, element 0010). If the label exists, it is marked as sensitive medical data (L2 level). For encoding format recognition, the first few bytes of the data packet are read to search for NAL unit headers. If SPS or PPS units with type values ​​of 7 or 8 are detected, it is determined to be an H.264 video stream. The first byte is read, the TOC configuration is parsed, and the Opus specification is matched to determine it as an Opus audio stream, marked as real-time operational data (L3 level). After classification, the data is accompanied by a corresponding security level identifier.

[0085] The encryption algorithm selects the encryption method according to the security level. Level 1 data uses the SM4 algorithm, dividing the plaintext into 128-bit blocks and performing 32 rounds of iterative transformation using a 128-bit key. Each round includes key addition, S-box substitution, and linear transformation. The key is established between the communicating parties via the SM2 elliptic curve protocol. After exchanging temporary public keys, the parties calculate a shared key as the SM4 key. Level 2 data uses the AES-256 algorithm, with 128-bit blocks and a 256-bit key, performing 14 rounds of iteration. Each round includes byte substitution, row shifting, column mixing, and key addition. A digital watermark is embedded in the encrypted CT image. The image is divided into 8×8 pixel blocks for DCT transformation. The coefficients at the intermediate frequency coefficient positions are modified according to the watermark bit; when the watermark bit is 1, the coefficient is multiplied by 1.1, and when it is 0, it is multiplied by 0.9. After modification, the image is converted back to the spatial domain using inverse DCT. The watermark content, including the user ID and timestamp, is converted into a 64-bit binary sequence and embedded. Level 3 data uses the ChaCha20 algorithm. The core function takes a 256-bit key, a 96-bit nonce, and a 64-bit counter as input, performs 20 rounds of QuarterRound transformation, and outputs a 512-bit key stream block. The key stream is XORed byte by byte with the plaintext to obtain the ciphertext. ChaCha20 has a fast computation speed and is suitable for real-time video stream encryption.

[0086] The data access request permission determination extracts the user identifier and operation type from the request. Based on the user identifier, it queries the user role mapping table to return the role identifier. The operation type is coded from 1 to 5, representing CT viewing, model rotation, annotation addition, video control, and data retrieval, respectively. Based on the current surgical stage identifier and role identifier, it indexes a multi-role permission matrix. The matrix index is calculated by multiplying the stage number by the number of roles and adding the role number. It queries the matrix element to determine if the permission value is 1. If the permission value is 1, access is allowed and an access record is generated. The record contains the user identifier, resource identifier, operation type, and time field in Unix timestamp format.

[0087] The hash chain storage of access records concatenates four fields—user identifier, resource identifier, operation type, and timestamp—into a string using vertical bars as separators. This string is then converted into a byte sequence and input into the SHA-256 algorithm. After padding the input, the SHA-256 algorithm divides the data into 512-bit blocks. Each block undergoes 64 rounds of compression to update eight 32-bit working variables. After 64 rounds, the working variables are concatenated to output a 256-bit hash value. The hash value of the current record is concatenated with the hash value of the previous record and then SHA-256 is calculated again. The new hash value is stored as block data on the blockchain nodes. The blockchain adopts a consortium blockchain architecture with multiple nodes maintaining the ledger. New blocks are added to the chain using the PBFT consensus algorithm, forming an immutable access audit chain.

[0088] In one specific embodiment, step S5 includes:

[0089] The system collects real-time data on available network bandwidth, endoscope main video stream buffer occupancy, auxiliary view video stream buffer occupancy, 3D model data transmission queue length, voice communication jitter, current surgical stage coding, number of online users, network packet loss rate, and average round-trip latency. These parameters are then combined into a nine-dimensional state vector.

[0090] The state vector is input into a deep Q-network for forward propagation calculation. The deep Q-network contains three fully connected layers. The input data is processed by linear transformation and activation function in sequence. The output layer generates the Q value of each data stream bitrate adjustment strategy combination. The strategy combination with the largest Q value is selected as the bitrate adjustment strategy.

[0091] Based on the parameter values ​​in the bitrate adjustment strategy, the resolution and frame rate of the endoscope's main video stream are adjusted, the bitrate of the auxiliary view video stream is adjusted, and the transmission bandwidth allocation of the 3D model data is adjusted to obtain the bitrate configuration parameters for each data stream.

[0092] Obtain the user's head-mounted display's viewing direction and the distance value of the registration model. When the distance value is less than the first distance threshold and the available bandwidth is greater than the first bandwidth threshold, select a high-precision grid. When the distance value is between the first and second distance thresholds or the available bandwidth is between the first and second bandwidth thresholds, select a medium-precision grid. When the distance value is greater than the second distance threshold or the available bandwidth is less than the second bandwidth threshold, select a low-precision grid. Adjust the grid density of the registration model according to the selection results.

[0093] Specifically, real-time network status parameters are acquired through the operating system network interface and application layer monitoring modules. The current available bandwidth is measured by sending test data packets. Data packets of known size are sent within a short time window, and the amount of successfully transmitted data is recorded. The bandwidth value is obtained by dividing the data amount by the transmission time. Buffer occupancy is queried from the video decoder's buffer status. The occupancy rate of the endoscope's main video stream is equal to the used buffer size divided by the total buffer size. The occupancy rate of the auxiliary viewpoint video stream is calculated using the same method. An occupancy rate close to 1 indicates the buffer is about to overflow, while an occupancy rate close to 0 indicates insufficient buffer data. The 3D model data transmission queue length is used to count the number of model data packets waiting to be transmitted. An increase in queue length indicates that the transmission speed cannot keep up with the data generation speed. Voice communication jitter is calculated by analyzing the variance of the arrival time interval of audio data packets. The standard deviation of the difference sequence of arrival times of consecutive data packets is the jitter value. A large jitter value indicates unstable network transmission. The stage code of the current surgical stage is obtained from the recognition result of step S1, with the seven surgical stages coded from 1 to 7. The number of online users is counted to determine the number of user terminals currently connected to the collaborative system. Network packet loss rate is calculated by subtracting the number of successfully received packets from the total number of packets sent, and then dividing by the total number of packets sent. Average round-trip time (RTT) is calculated by sending ICMP echo request packets, recording the time from sending to receiving the echo response, and averaging the measurements multiple times. Nine parameters are combined in a fixed order to form a nine-dimensional state vector. The first element of the vector is the bandwidth value, the second element is the main video buffer occupancy rate, and so on up to the ninth element, which is the average RTT.

[0094] The forward propagation computation of a deep Q-network takes a nine-dimensional state vector as input, and the input layer contains nine neurons with nine parameters. The first fully connected layer contains 128 neurons, each fully connected to the nine neurons in the input layer. The output value of a neuron is calculated as the dot product of the input vector and the neuron's weight vector, plus a bias value. The dot product is the sum of corresponding element-wise multiplications. The nine input values ​​are multiplied by the nine weight values, summed, and then the bias is added. The output value is processed by the ReLU activation function, which sets negative values ​​to zero and preserves positive values. After activation, the output forms a 128-dimensional intermediate vector. The second fully connected layer contains 64 neurons, receiving the 128-dimensional output from the first layer. Each neuron calculates the dot product of the 128 inputs and 128 weights, plus a bias, and after ReLU activation, outputs a 64-dimensional vector. The number of neurons in the third output layer equals the total number of bitrate adjustment strategy combinations. Each strategy combination includes the Cartesian product of the main video stream's resolution option multiplied by its frame rate option, the auxiliary video stream's bitrate option, and the model's data bandwidth option. Assuming there are 2 main video resolutions, 2 frame rates, 3 auxiliary video bitrates, and 2 model bandwidth options, the total number of strategy combinations is 2 x 2 x 3 x 2 = 24. The output layer contains 24 neurons. Each neuron calculates the dot product of 64 inputs and 64 weights plus a bias, directly outputting the value without an activation function. The output value is the Q-value of that strategy combination. The Q-value represents the expected cumulative reward for choosing that strategy in the current state. The maximum Q-value is found by iterating through the 24 Q-values, and the strategy combination corresponding to the maximum Q-value is used as the currently executed bitrate adjustment strategy.

[0095] The bitrate configuration parameters are adjusted based on the parameter values ​​in the strategy combination. The strategy combination encoding parses out the main video resolution parameter, frame rate parameter, auxiliary video bitrate parameter, and model bandwidth parameter. When the main video resolution parameter is 0, it is set to 1280×720 pixels; when the parameter is 1, it is set to 1920×1080 pixels. After receiving the resolution setting, the video encoder reconfigures the encoding parameters, scaling the original video frames to the target resolution using a bilinear interpolation algorithm. When the frame rate parameter is 0, it is set to 15 frames per second; when the parameter is 1, it is set to 30 frames per second. The encoder adjusts the encoding time interval according to the frame rate parameter: 66.7 milliseconds for 15 frames per second and 33.3 milliseconds for 30 frames per second. The auxiliary video bitrate parameter is set to 1 Mbps, 2 Mbps, and 3 Mbps when it is 0, 1, or 2, respectively. The bitrate controls the bitrate output by the encoder and is achieved by adjusting the quantization parameter; a larger quantization parameter results in a higher compression ratio and a lower bitrate. When the model bandwidth parameter is 0, 0.5 Mbps of bandwidth is allocated; when the parameter is 1, 1 Mbps of bandwidth is allocated. The transmission control module limits the transmission rate of model data according to the bandwidth allocation, using a token bucket algorithm for control. Tokens are generated at the bandwidth rate, and sending data packets consumes tokens. When tokens are insufficient, data packets are queued. After the bitrate configuration parameters of each data stream are updated, the encoder and transmission module operate according to the new parameters.

[0096] Mesh density adjustment first acquires the user's head-mounted display (HUD) pose data, including position coordinates and orientation quaternions. Position coordinates are calculated from the fusion of the HUD's IMU sensors and an external positioning system, while orientation quaternions are calculated from the IMU's gyroscope and accelerometer data. The registration model's center point coordinates are the average of all vertex coordinates. The Euclidean distance between the HUD's position coordinates and the model's center point coordinates is calculated; this distance is equal to the square root of the sum of the squares of the differences in the three coordinate axes, and is used as the viewing distance. The relationship between the viewing distance and a first and second distance threshold is determined. The first threshold is set to 50 cm, and the second to 100 cm. Simultaneously, the relationship between available bandwidth and a first and second bandwidth threshold is determined. The first bandwidth threshold is set to 2 Mbps, and the second to 1 Mbps. When the viewing distance is less than 50 cm and the available bandwidth is greater than 2 Mbps, indicating close-range user observation of the model and good network conditions, a high-precision mesh is selected. This high-precision mesh contains 500,000 triangular faces. When the viewing distance is between 50 cm and 100 cm, or the available bandwidth is between 1 Mbps and 2 Mbps, a medium-precision grid is selected, containing 200,000 triangular faces. When the viewing distance is greater than 100 cm, or the available bandwidth is less than 1 Mbps, or the user is observing from a distance or the network conditions are poor, a low-precision grid is selected, containing 50,000 triangular faces. The grid simplification algorithm performs edge folding operations on the registration model. Edge folding merges two vertices of an edge into one vertex and deletes two adjacent triangular faces. Each edge fold calculates the geometric error cost, which is equal to the square of the distance from the new vertex to the original surface. Edges with smaller costs are folded first, and edge folding is repeated until the number of faces reaches the target accuracy level.

[0097] Figure 2 This is a schematic diagram illustrating the dynamic adjustment relationship between network bandwidth and video bitrate in an embodiment of this application. Figure 2This diagram illustrates the adaptive network bandwidth allocation process based on a deep Q-network in this embodiment. The horizontal axis represents the time progression of the surgical training process, and the vertical axis represents the bitrate or bandwidth value. The solid dotted curve represents the real-time changes in available network bandwidth, the solid square dotted curve represents the bitrate adjustment strategy for the main endoscopic video stream, and the dashed triangular dotted curve represents the bitrate adjustment strategy for the auxiliary viewpoint video stream. It is clearly observed that when available bandwidth decreases between 15 and 30 seconds, the deep Q-network automatically reduces the main video stream bitrate from 8 Mbps to 4 Mbps, while simultaneously reducing the auxiliary video stream bitrate from 3 Mbps to 1 Mbps to ensure transmission stability. When the bandwidth gradually recovers after 35 seconds, the system promptly increases the bitrate of both video streams to restore image quality. This demonstrates that the intelligent bandwidth allocation mechanism based on the deep Q-network in this application can dynamically adjust the bitrate configuration parameters of each data stream according to the network status, prioritizing the quality of critical video streams and ensuring the smoothness and reliability of multi-role collaborative training.

[0098] The above describes the mixed reality surgical training collaboration method supporting multi-role interaction in the embodiments of this application. The following describes the mixed reality surgical training collaboration system supporting multi-role interaction in the embodiments of this application. Please refer to [link to relevant documentation]. Figure 3 One embodiment of the mixed reality surgical training collaboration system supporting multi-role interaction in this application includes:

[0099] The weighting module is used to extract endoscopic video features and instrument sequence features, perform weighted fusion, output the surgical stage probability distribution through a classifier, determine the current surgical stage when the maximum probability value exceeds the threshold, and generate a multi-role permission matrix based on the current surgical stage.

[0100] The filtering module is used to match anatomical landmarks on the endoscope with corresponding points on the preoperative model, filter valid point pairs with matching distances that meet the requirements as deformation control points, and perform non-rigid interpolation deformation on the vertex coordinates of the model to obtain a registration model that is aligned with the real-time anatomical structure.

[0101] The output module is used to detect the user's gesture to obtain the operation type and motion parameters, calculate the spatial distance between the gesture position and the surface of the registration model, determine the tactile feedback type based on the distance range and motion speed, and drive the actuator to output the corresponding vibration mode.

[0102] The storage module is used to identify the data type of the transmission and mark the security level, apply corresponding encryption methods to data of different levels, and combine the multi-role permission matrix and the current surgical stage to determine the access request, and then hash the access record and store it in a chain.

[0103] The adjustment module is used to collect network bandwidth and buffer status to construct a state vector, select the bitrate adjustment strategy for each data stream through a deep Q network, and adjust the grid density of the registration model according to bandwidth conditions and user view distance.

[0104] Figure 3 The mixed reality surgical training collaboration system supporting multi-role interaction in the embodiments of the present invention will be described in detail from the perspective of modular functional entities. The mixed reality surgical training collaboration device supporting multi-role interaction in the embodiments of the present invention will be described in detail from the perspective of hardware processing.

[0105] Reference Figure 4 This invention also provides a mixed reality surgical training collaboration device that supports multi-role interaction. This device can be a server, and its internal structure can be as follows: Figure 4 As shown, this multi-role interactive mixed reality surgical training collaborative device includes a processor, memory, display screen, input device, network interface, and database connected via a system bus. The processor, designed as a computer, provides computing and control capabilities. The memory of this multi-role interactive mixed reality surgical training collaborative device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of this multi-role interactive mixed reality surgical training collaborative device stores the data corresponding to this embodiment. The network interface of this multi-role interactive mixed reality surgical training collaborative device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the above-described method.

[0106] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the mixed reality surgical training collaborative device supporting multi-role interaction on which the present invention is applied.

[0107] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the mixed reality surgical training collaboration method supporting multi-role interaction.

[0108] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0109] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a mixed reality surgical training collaborative device (which may be a personal computer, server, or network device, etc.) supporting multi-role interaction to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0110] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A mixed reality surgical training collaborative method supporting multi-role interaction, characterized in that, The method includes: Step S1: Extract endoscopic video features and instrument sequence features, perform weighted fusion, output the surgical stage probability distribution through a classifier, determine the current surgical stage when the maximum probability value exceeds the threshold, and generate a multi-role permission matrix based on the current surgical stage. Step S2: Match the anatomical landmarks on the endoscope with the corresponding points on the preoperative model, select valid point pairs with matching distances that meet the requirements as deformation control points, perform non-rigid interpolation deformation on the vertex coordinates of the model, and obtain a registration model that is aligned with the real-time anatomical structure. S3 step: Detect user gestures to obtain operation type and motion parameters, calculate the spatial distance between the gesture position and the surface of the registration model, determine the tactile feedback type based on the distance range and motion speed, and drive the actuator to output the corresponding vibration mode; Step S4: Identify the data type of the transmission and mark the security level. Apply the corresponding encryption method to the data of different levels. Combine the multi-role permission matrix and the current surgical stage to determine the access request. Hash the access record and store it in a chain. Step S5: Collect network bandwidth and buffer status to construct a state vector. Select the bitrate adjustment strategy for each data stream using a deep Q-network. Adjust the grid density of the registration model based on bandwidth conditions and user view distance. This includes: real-time collection of current available network bandwidth, endoscope main video stream buffer occupancy rate, auxiliary view video stream buffer occupancy rate, 3D model data transmission queue length, voice communication jitter value, stage coding of the current surgical stage, number of online users, network packet loss rate, and average round-trip latency. Combine these parameters into a nine-dimensional state vector. Input the state vector into a deep Q-network for forward propagation calculation. The deep Q-network contains three fully connected layers, which sequentially perform linear transformations and activation function processing on the input data. The output layer generates a Q-value for the combination of bitrate adjustment strategies for each data stream. The strategy combination with the largest Q value is selected as the bitrate adjustment strategy. Based on the parameter values ​​in the bitrate adjustment strategy, the resolution and frame rate of the endoscope's main video stream, the bitrate of the auxiliary viewpoint video stream, and the transmission bandwidth allocation of the 3D model data are adjusted to obtain the bitrate configuration parameters for each data stream. The distance between the user's head-mounted display's viewing direction and the registration model is obtained. When the distance value is less than a first distance threshold and the available bandwidth is greater than a first bandwidth threshold, a high-precision grid is selected. When the distance value is between the first and second distance thresholds or the available bandwidth is between the first and second bandwidth thresholds, a medium-precision grid is selected. When the distance value is greater than the second distance threshold or the available bandwidth is less than the second bandwidth threshold, a low-precision grid is selected. The grid density of the registration model is adjusted according to the selection results.

2. The mixed reality surgical training collaboration method supporting multi-role interaction according to claim 1, characterized in that, Step S1 includes: After adjusting the single-frame image of the endoscope video stream to a preset size, it is input into a feature extraction network containing convolutional and pooling layers for layer-by-layer convolutional operations and feature dimensionality reduction to obtain a fixed-dimensional image feature vector. Identify the device type at each time step in the device usage sequence, map each device type to a one-hot encoded form, and obtain the device sequence feature vector; The image feature vector and the instrument sequence feature vector are concatenated element-wise according to the video feature weight and the instrument feature weight to obtain the fused feature vector. The fused feature vector is sequentially transformed and activated by three fully connected layers. The output layer is calculated by the Softmax function to obtain the probability distribution of seven stages: preoperative preparation, incision, separation, hemostasis, tumor resection, suturing and postoperative examination. When the maximum probability exceeds the set threshold, the corresponding stage is determined as the current surgical stage. Based on the preset role-operation mapping table of the current surgical stage index, the surgeon, assistant surgeon, remote expert, and observer are respectively assigned permission values ​​for CT viewing, model rotation, annotation addition, video control, and data access, generating a multi-role permission matrix.

3. The mixed reality surgical training collaboration method supporting multi-role interaction according to claim 1, characterized in that, Step S2 includes: The moving cube algorithm was used to extract the organ surface from the preoperative CT data, generating a triangular mesh containing vertex coordinates and normal vectors to obtain the preoperative model; The feature detector is applied to the current frame of the endoscopic video stream to extract anatomical landmarks of blood vessel bifurcation points and organ edges, and the feature descriptor of each landmark is calculated to obtain the set of anatomical landmarks in the endoscopic image. Each feature descriptor in the endoscopic image anatomical landmark point set is compared with the feature descriptor of the pre-calibrated corresponding point in the preoperative model using Euclidean distance calculation. Matching point pairs with distance values ​​less than a set threshold are filtered out to obtain a set of effective point pairs as deformation control points. Based on the deformation control points, a thin plate spline deformation function is constructed. After calculating the weight coefficients of the radial basis function, non-rigid interpolation deformation calculation is performed on the coordinates of all vertices in the preoperative model to obtain a registration model aligned with the real-time anatomical structure.

4. The mixed reality surgical training collaboration method supporting multi-role interaction according to claim 3, characterized in that, The endoscopic video stream is processed by applying a feature detector to extract anatomical landmarks at blood vessel bifurcation points and organ edges in the current frame, and a feature descriptor is calculated for each landmark to obtain a set of anatomical landmarks in the endoscopic image, including: The current frame of the endoscope video stream is grayscaled and Gaussian filtered to eliminate image noise, resulting in a preprocessed image. The preprocessed image is subjected to a multi-scale Harris corner detection algorithm to detect the pixel positions where the corner response value exceeds a set threshold in the blood vessel intersection area and organ contour edge, and to obtain a set of candidate marker coordinates. Extract the gradient direction histogram within the neighborhood of each candidate marker point's coordinates, quantize the gradient magnitude and direction into a fixed-dimensional feature vector, and obtain the feature descriptor for each marker point; The candidate marker coordinates and their corresponding feature descriptors are associated and stored to obtain a set of anatomical markers in the endoscopic image.

5. The mixed reality surgical training collaboration method supporting multi-role interaction according to claim 1, characterized in that, Step S3 includes: The hand tracking module of the mixed reality headset detects the three-dimensional coordinates of key points of the user's hand bones, and identifies the type of gesture such as pinching, translation, rotation or scaling based on the spatial relationship of the key points of the fingers, and obtains the operation type identifier. Calculate the Euclidean distance between the three-dimensional coordinates of the fingertip in the user's gesture and the coordinates of each vertex on the surface of the registration model, and select the minimum distance value as the spatial distance between the gesture position and the model surface; When the spatial distance is greater than the first distance threshold, it is determined to be a proximity feedback type; when the spatial distance is less than the second distance threshold, it is determined to be a contact feedback type; when the operation type is identified as rotation and the rotational angular velocity exceeds the velocity threshold, it is determined to be a resistance feedback type, thus obtaining the tactile feedback type. According to the tactile feedback type, a preset feedback parameter table is queried to obtain the corresponding vibration frequency value and amplitude value, and the tactile actuator is driven to output a vibration mode according to the vibration frequency value and amplitude value.

6. The mixed reality surgical training collaboration method supporting multi-role interaction according to claim 1, characterized in that, Step S4 includes: Content analysis is performed on the data to be transmitted. Patient identity information is matched with regular expressions and marked as core privacy data. CT images and pathology reports are identified by file header tags and marked as sensitive medical data. Video streams and audio streams are identified by encoding format and marked as real-time operation data, resulting in classified data with security level labels. Different encryption algorithms are applied to the classified data according to the security level identifier. The core privacy data is encrypted end-to-end using the national cryptographic SM4 algorithm, the sensitive medical data is encrypted using the AES-256 algorithm and embedded with a digital watermark, and the real-time operation data is encrypted using the ChaCha20 algorithm to obtain an encrypted data stream. After receiving a user's data access request, the user role identifier and operation type in the request are extracted. Combined with the permission value of the corresponding role in the multi-role permission matrix and the stage identifier of the current surgical stage, it is determined whether the access request meets the permission conditions. If the conditions are met, access is allowed and an access record is generated. The user identifier, resource identifier, operation type, and timestamp in the access record are hashed. The hash value is then concatenated with the hash value of the previous record and hashed again. The hash value is then stored in a chain to the blockchain node.

7. A mixed reality surgical training collaborative system supporting multi-role interaction, characterized in that, For implementing the mixed reality surgical training collaboration method supporting multi-role interaction as described in any one of claims 1-6, the mixed reality surgical training collaboration system supporting multi-role interaction comprises: The weighting module is used to extract endoscopic video features and instrument sequence features, perform weighted fusion, output the surgical stage probability distribution through a classifier, determine the current surgical stage when the maximum probability value exceeds the threshold, and generate a multi-role permission matrix based on the current surgical stage. The filtering module is used to match anatomical landmarks on the endoscope with corresponding points on the preoperative model, filter valid point pairs with matching distances that meet the requirements as deformation control points, and perform non-rigid interpolation deformation on the vertex coordinates of the model to obtain a registration model that is aligned with the real-time anatomical structure. The output module is used to detect the user's gesture to obtain the operation type and motion parameters, calculate the spatial distance between the gesture position and the surface of the registration model, determine the tactile feedback type based on the distance range and motion speed, and drive the actuator to output the corresponding vibration mode. The storage module is used to identify the data type of the transmission and mark the security level, apply corresponding encryption methods to data of different levels, and combine the multi-role permission matrix and the current surgical stage to determine the access request, and then hash the access record and store it in a chain. The adjustment module is used to collect network bandwidth and buffer status to construct a state vector, select the bitrate adjustment strategy for each data stream through a deep Q network, and adjust the grid density of the registration model according to bandwidth conditions and user view distance.

8. A mixed reality surgical training collaborative device supporting multi-role interaction, characterized in that, The method includes a memory and a processor, the memory storing a computer program that can run on the processor, and the processor executing the computer program to implement the mixed reality surgical training collaborative method supporting multi-role interaction as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it causes the processor to execute the mixed reality surgical training collaboration method supporting multi-role interaction as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-person cooperative surgery virtual practical training system based on multi-view behavior identification

    CN116661600A

  • Cesarean section training method and system based on mixed reality

    CN119559838A