Agent-based medical instrument control method, server, and storage medium
Patent Information
- Application Number
- CN202611015486.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-09
- Publication Date
- 2026-09-25
AI Technical Summary
然而,此类方案仅对单一信号进行独立阈值判定,未将力信号与位移信号在时间维度上进行关联分析,无法识别受力变化与位移响应之间的因果时序关系,导致对潜在风险事件的感知存在明显滞后
[0006]再一方面,本发明实施例还提供一种计算机存储介质,所述计算机存储介质存储有机器可执行指令,服务器的处理器从所述计算机可读存储介质读取所述机器可执行指令,所述处理器执行所述机器可执行指令,使得所述服务器执行上述的基于智能体的医疗仪器控制方法。
Smart Images

Figure CN122805375A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and more specifically, to a medical instrument control method, server, and storage medium based on intelligent agents. Background Technology
[0002] In invasive medical procedures (such as catheter-based interventional surgery and endoscopic minimally invasive surgery), the precision of the operation directly affects patient safety and surgical outcomes. Currently, doctors mainly rely on visual image guidance to manually operate medical instruments. During the operation, the mechanical interaction information between the instrument tip and human tissue is usually only presented on the display screen in the form of simple numerical values. Doctors need to rely on experience to judge and adjust the force and direction of operation in real time, which places extremely high demands on the operator's experience level and on-site reaction ability.
[0003] In existing technologies, some solutions attempt to collect force and displacement data by adding force and displacement sensors to the instrument's end effector and trigger alarms based on preset thresholds to indicate abnormal states. However, such solutions only perform independent threshold determinations on a single signal, failing to correlate force and displacement signals over time. This makes it impossible to identify the causal time sequence relationship between force changes and displacement responses, resulting in a significant lag in the perception of potential risk events. Other solutions employ a closed-loop control strategy with fixed rules, directly adjusting the applied force at the instrument's end effector based on the current force value. However, this strategy only focuses on the instantaneous state, ignoring the precursory waveform characteristics before force abrupt changes and the recovery process of displacement response after such abrupt changes. It cannot extract effective decision-making basis from the complete event evolution chain, thus making it difficult to achieve refined force-position coordinated control in complex organizational environments. In addition, existing solutions typically lack cross-event memory mechanisms, and the handling of each stress-induced abrupt change event is independent of each other. They fail to utilize the aftereffects of historical events to assist in intervention decisions for current events, resulting in a lack of continuity and adaptability in control strategies. Under complex conditions such as sudden changes in tissue stiffness or instrument jamming, they are prone to excessive force application or operational instability. Summary of the Invention
[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a medical instrument control method based on an intelligent agent, the method comprising: Acquire the instrument end force feedback signal stream and instrument end displacement feedback signal stream generated by the target medical instrument during the execution of the invasive target medical operation, wherein the instrument end force feedback signal stream and the instrument end displacement feedback signal stream have a synchronous acquisition relationship in the time dimension; Force abrupt event detection processing is performed on the force feedback signal stream at the end of the instrument to determine the occurrence time of the force abrupt event and extract the precursor waveform segment and the aftereffect waveform segment corresponding to each force abrupt event. The occurrence time of each force abrupt event is used as the forced segmentation boundary point for displacement response segmentation processing of the displacement feedback signal stream at the end of the instrument. For each sudden force event, the displacement feedback signal stream at the end of the instrument is divided into a displacement response front segment and a displacement response back segment, with the occurrence time of the sudden force event as the boundary. The end displacement retraction parameter and end displacement offset direction parameter of the displacement response front segment are extracted, and the end displacement recovery parameter and end displacement recovery direction parameter of the displacement response back segment are extracted. The precursor waveform segment of each force mutation event, the end displacement retraction parameter, and the end displacement offset direction parameter are used together as the precursor perception input of the instrument control intervention decision-making agent. The precursor pattern encoder in the instrument control intervention decision-making agent performs precursor pattern encoding processing on the precursor perception input to generate a precursor pattern encoding vector. The intervention strategy mapping network in the instrument control intervention decision-making agent uses the precursor pattern encoding vector and the aftereffect pattern encoding vector corresponding to the previous force mutation event as the joint decision basis to perform intervention strategy mapping processing to generate the instrument end force-position hybrid control command corresponding to the current force mutation event. The aftereffect pattern encoding vector is obtained by the aftereffect pattern encoder in the instrument control intervention decision-making agent performing aftereffect pattern encoding processing on the aftereffect waveform segment of the force mutation, the end displacement recovery parameter, and the end displacement recovery direction parameter. The instrument end force-position hybrid control command is used to adjust the applied force magnitude and displacement direction of the target medical instrument end.
[0005] Furthermore, embodiments of the present invention also provide a server, comprising: A processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to perform the above-described agent-based medical instrument control method by executing the machine-executable instructions.
[0006] In another aspect, embodiments of the present invention also provide a computer storage medium storing machine-executable instructions. A processor of a server reads the machine-executable instructions from the computer-readable storage medium and executes the machine-executable instructions, causing the server to execute the aforementioned agent-based medical instrument control method.
[0007] Based on the above, by synchronously acquiring the force feedback signal stream and displacement feedback signal stream at the end of the instrument, and taking the occurrence time of the force mutation event as the forced segmentation boundary point of the displacement response signal, this invention decouples each force mutation event into three stages: precursor, response, and aftereffect. It extracts multi-dimensional feature parameters such as precursor waveform segments, displacement shrinkage and offset direction, and displacement recovery and recovery direction, forming a structured feature input covering the entire life cycle of the event. Furthermore, this invention introduces an intelligent agent for instrument control intervention decision-making. This agent encodes the precursor perception input through a precursor mode encoder and the aftereffect characteristics of the previous event through an aftereffect mode encoder. Using both as the basis for joint decision-making, a force-position hybrid control command is generated through an intervention strategy mapping network, thereby establishing cross-event contextual memory. This ensures that the current control decision not only depends on the real-time precursor characteristics of the current event but also incorporates the recovery mode information of the previous event, achieving temporal continuity and adaptive evolution of the control strategy. This allows the agent to perceive risk trends in advance based on the weak characteristics of the precursor waveform within a very short time window of force abrupt change, accurately locate the abnormal direction by combining the recoil and offset characteristics of the displacement response, and dynamically adjust the current force magnitude and displacement direction by referring to the recovery mode of historical events. This achieves a shift in control paradigm from passive response to active prediction in the complex mechanical environment of invasive operations, effectively reducing the risk of tissue damage caused by control lag or strategy isolation. At the same time, the unified output of the force-position hybrid control command avoids strategy conflicts between force control and displacement control, significantly improving operational safety and accuracy. Attached Figure Description
[0008] Figure 1 This is a schematic diagram of the execution flow of the medical instrument control method based on intelligent agents provided in an embodiment of the present invention.
[0009] Figure 2 This is a schematic diagram of the interface of the mobile terminal monitoring application provided in an embodiment of the present invention.
[0010] Figure 3 This is another schematic diagram of the interface of the mobile terminal monitoring application provided in this embodiment of the invention. Detailed Implementation
[0011] Figure 1This is a flowchart illustrating a medical instrument control method based on an intelligent agent according to an embodiment of the present invention. It should be noted that in this embodiment, the acquisition, transmission, and storage of all instrument end-effector force feedback signal streams, instrument end-effector displacement feedback signal streams, tissue impedance feedback signal streams, operator active control torque signal streams, and treatment energy output power sequences are performed within the scope authorized by the medical institution. Patient privacy data has been anonymized and desensitized before collection. Data transmission is encrypted using a transport layer security protocol, and storage is encrypted using an advanced encryption standard algorithm. This embodiment uses a percutaneous liver tumor ablation procedure performed by a puncture ablation medical instrument as an application scenario. The puncture ablation medical instrument is equipped with a six-dimensional force sensor and an electromagnetic displacement tracking sensor at its end-effector. The force sensor collects the axial and lateral forces experienced by the instrument end-effector during the puncture process at a preset sampling frequency, while the displacement tracking sensor simultaneously collects the displacement coordinates of the instrument end-effector in three-dimensional space.
[0012] Step S110: Acquire the instrument end force feedback signal stream and instrument end displacement feedback signal stream generated by the target medical instrument during the invasive target medical operation. The instrument end force feedback signal stream and the instrument end displacement feedback signal stream are synchronously acquired in the time dimension.
[0013] The computing environment is connected to the force sensor signal conditioning module and displacement tracking sensor signal conditioning module of the target medical instrument via a high-speed serial data bus. The data frame output by the force sensor signal conditioning module includes a timestamp field and a force component array field. The force component array field is a three-dimensional floating-point array, with the three components representing the axial force component, the first lateral force component, and the second lateral force component, all in Newtons. The data frame output by the displacement tracking sensor signal conditioning module includes a timestamp field and a displacement coordinate array field. The displacement coordinate array field is a three-dimensional floating-point array, with the three components representing the horizontal, vertical, and depth coordinates of the spatial coordinate system, all in millimeters. The timestamp fields of both data frames are timed by the same clock source, and the timestamp format is a Unix timestamp, in milliseconds.
[0014] The computing environment stores the received force sensor data frames and displacement tracking sensor data frames into two circular buffers, respectively. Each element of the circular buffer is a structure containing a timestamp value and an array of sampled values. Using the timestamp as the alignment key, the computing environment pairs force and displacement sampled frames from the two circular buffers whose timestamp differences are within a preset synchronization tolerance range. The paired data structure is a structure containing a timestamp field, a force component array field, and a displacement coordinate array field. All paired data are arranged in ascending order of timestamp to form a synchronously acquired data sequence. The force component array sequence is extracted from the synchronously acquired data sequence to form the instrument's end-effector force feedback signal stream, and the displacement coordinate array sequence is extracted to form the instrument's end-effector displacement feedback signal stream. The data structure of the instrument's end-effector force feedback signal stream is a two-dimensional floating-point array. The first dimension of the array corresponds to the sampling time index, and the second dimension corresponds to the three force components, with the dimension being Newtons. The data structure of the instrument's end-effector displacement feedback signal stream is also a two-dimensional floating-point array. The first dimension of the array corresponds to the sampling time index, and the second dimension corresponds to the three displacement coordinate components, with the dimension being millimeters.
[0015] Step S120: Perform force change event detection processing on the force feedback signal stream at the end of the instrument, determine the occurrence time of the force change event, and extract the precursor waveform segment and the aftereffect waveform segment of the force change corresponding to each force change event. The occurrence time of each force change event is used as the forced segmentation boundary point for displacement response segmentation processing of the displacement feedback signal stream at the end of the instrument.
[0016] Step S121: Perform sliding window differential processing on the force feedback signal stream at the instrument end, calculate the force change rate sequence between adjacent sampling points, perform sign flip detection processing on the force change rate sequence, and identify the set of sign flip points in the force change rate sequence where the positive change rate suddenly changes to the negative change rate or vice versa.
[0017] The computational environment performs sliding window differential processing on the axial force component sequence of the force feedback signal stream at the instrument's end. The sliding window size is two sampling points. For the axial force component values corresponding to two adjacent sampling times, the difference between the value at the later sampling time and the value at the previous sampling time is calculated, and then divided by the time interval between the two sampling times to obtain the rate of change of force between the adjacent sampling points. The dimension of the rate of change of force is Newtons per second. This calculation is repeated for all adjacent sampling point pairs to obtain the rate of change of force sequence, which is a one-dimensional floating-point array, and the array length is 1 less than the total number of original sampling points.
[0018] The computing environment performs sign reversal detection on the force rate of change sequence. It iterates through adjacent force rate of change values in the sequence. If the preceding value is positive and the following value is negative, a sign reversal point (from positive to negative) is marked at that adjacent position. Conversely, if the preceding value is negative and the following value is positive, a sign reversal point (from negative to positive) is marked at that adjacent position. All marked sign reversal points are sorted in ascending order by timestamp, forming a sign reversal point set. Each element in the set is a structure containing a timestamp and a reversal type identifier, where the identifier is either a positive-to-negative constant or a negative-to-normal constant.
[0019] Step S122: Select symbol flip points that meet the preset flip amplitude threshold condition and flip speed threshold condition from the symbol flip point set as candidate occurrence times of force mutation events. For each candidate occurrence time, search in reverse along the time axis for the moment when the force change rate in the force feedback signal stream at the end of the instrument first enters the stable fluctuation range as the starting time of the force mutation precursor of the force mutation event.
[0020] For each symbol inversion point in the symbol inversion point set, the computational environment calculates the corresponding inversion amplitude. The inversion amplitude is the absolute value of the difference between the maximum value of the rate of change of force within a preset window length before the inversion point and the minimum value of the rate of change of force within a preset window length after the inversion point, measured in Newtons per second. This inversion amplitude is compared with a preset inversion amplitude threshold, also measured in Newtons per second. The inversion velocity at this point is calculated as the inversion amplitude divided by the inversion duration, which is the time it takes for the rate of change of force to recover from its extreme value to a stable value after the inversion point, measured in seconds. This inversion velocity is compared with a preset inversion velocity threshold, measured in Newtons per second squared. If both the inversion amplitude and velocity exceed the preset inversion amplitude and velocity thresholds, the symbol inversion point is marked as a candidate occurrence time for a force mutation event.
[0021] For each candidate occurrence time, the computational environment scans the force change rate sequence in reverse along the time axis. The local standard deviation of the force change rate sequence within this scanning window is calculated. If the local standard deviation is less than a preset stationarity threshold, the force change rate is considered to have entered a stationary fluctuation range. The end time of the stationary fluctuation range is the starting time of the force mutation precursor for this force mutation event. The starting time of the force mutation precursor is the moment when the local standard deviation first falls below the stationarity threshold during the scanning process.
[0022] Step S123: Search along the time axis for the moment when the rate of change of force in the force feedback signal stream at the end of the instrument re-enters the stable fluctuation range as the end time of the force mutation aftereffect of the force mutation event, and extract the segment of the force feedback signal stream at the end of the instrument between the start time of the force mutation precursor and the candidate occurrence time as the waveform segment of the force mutation precursor.
[0023] Starting from the candidate occurrence time, the computational environment scans the force change rate sequence along the time axis in the forward direction and calculates the local standard deviation in the same way as in step S122. The moment when the force change rate re-enters the stable fluctuation range is found, which is the end time of the aftereffect of the force mutation.
[0024] The computational environment extracts a subsequence of axial force component samples from the force feedback signal stream at the instrument's end, between the start time of the force abrupt change precursor and the candidate occurrence time. This subsequence is the waveform segment of the force abrupt change precursor. The data structure of the waveform segment of the force abrupt change precursor is a one-dimensional floating-point array, where each element is the axial force component value at that sampling time, with the dimension in Newtons. The length of the array is equal to the difference between the candidate occurrence time index and the start time index of the force abrupt change precursor, plus 1.
[0025] Step S124: Extract the instrument end force feedback signal stream segment between the candidate occurrence time and the end time of the force mutation effect as the force mutation effect waveform segment.
[0026] The computational environment extracts a subsequence of axial force component samples from the force feedback signal stream at the instrument's end, between the candidate occurrence time and the end time of the force abrupt change effect. This subsequence is the waveform segment of the force abrupt change effect. The data structure of the waveform segment of the force abrupt change effect is the same as that of the waveform segment of the precursor waveform segment.
[0027] Step S130: For each sudden force event, the displacement feedback signal stream at the end of the instrument is divided into a displacement response front segment and a displacement response back segment, with the occurrence time of the sudden force event as the boundary. The end displacement retraction parameter and end displacement offset direction parameter of the displacement response front segment are extracted, and the end displacement recovery parameter and end displacement recovery direction parameter of the displacement response back segment are extracted.
[0028] Step S131: Using the occurrence time of the sudden force event as the dividing boundary, divide the displacement feedback signal stream at the end of the instrument into a displacement response front segment located before the occurrence time and a displacement response back segment located after the occurrence time.
[0029] The computational environment uses the timestamps of the candidate occurrence times of the force abrupt change events determined in step S122 as the segmentation boundaries. All displacement sampling frames in the instrument end displacement feedback signal stream with timestamps less than the segmentation boundaries are assigned to the pre-displacement segment of the displacement response, while all displacement sampling frames with timestamps greater than the segmentation boundaries are assigned to the post-displacement segment. Both the pre-displacement and post-displacement segments are structure arrays containing timestamp values and three-dimensional displacement coordinates.
[0030] Step S132: Perform extreme point search processing on the displacement components of the instrument end in the three-dimensional coordinates of the space in the displacement response phase, determine the maximum offset of the instrument end in each spatial dimension from the original travel path, and take the largest modulus among the maximum offsets in each spatial dimension as the end displacement retraction parameter.
[0031] The computational environment acquires the displacement direction vector of the instrument's end effector along its original path before the occurrence of a sudden force change event during the displacement response phase. This direction vector is obtained by normalizing the displacement coordinate difference from the start of the displacement response phase to the segmentation boundary. For the three-dimensional displacement coordinates at each sampling moment in the displacement response phase, the vertical distance from the coordinate to the original path line is calculated. The vertical distance is calculated by taking the magnitude of the cross product of the vector formed by the coordinate at the segmentation boundary and the original path direction vector. The vertical distances at all sampling moments in the displacement response phase constitute a sequence. The computational environment uses a sliding window peak detection algorithm to search for a maximum point in the vertical distance sequence, and takes the maximum value among all maxima as the end effector displacement retraction parameter. If no maximum point exists in the displacement response phase, the maximum value in the vertical distance sequence is taken as the end effector displacement retraction parameter. The dimension of the end effector displacement retraction parameter is millimeters.
[0032] Step S133: Extract the cosine components of the displacement direction in each spatial dimension at the moment when the instrument tip reaches the corresponding time of the tip displacement retraction parameter, and construct the tip displacement offset direction parameter.
[0033] The displacement coordinates at the sampling time corresponding to the environmental positioning end-point displacement retraction parameter are calculated. The difference vector between this coordinate and the displacement coordinate at the segmentation boundary time is calculated, and this difference vector is normalized using the L2 norm. Normalization is performed by dividing each component of the difference vector by its L2 norm. The components of the normalized vector in each spatial dimension are the displacement direction cosine components. These three displacement direction cosine components constitute the end-point displacement offset direction parameter. The data structure of the end-point displacement offset direction parameter is a ternary floating-point array containing horizontal, vertical, and depth direction cosine components, each component being dimensionless.
[0034] Step S134: Perform trend tracking processing on the displacement recovery process of the instrument end in the latter part of the displacement response, and determine the total displacement regression of the instrument end from the maximum offset position to the original travel path in the latter part of the displacement response as the end displacement recovery parameter.
[0035] The computational environment extracts the trend of the instrument's end-effector displacement coordinates after the segmentation boundary time in the displacement response phase. The projection length of the difference vector between the displacement coordinates at each sampling time and the displacement coordinates at the segmentation boundary time is calculated along the original travel path direction. The projection length is the dot product of the difference vector and the original travel path direction vector. Trend analysis is performed on this projection length sequence to find the moment when the projection length first recovers to a preset proportion range of the projection length before the force abrupt change event; this moment is the displacement recovery completion moment. The end-effector displacement recovery parameter is equal to the projection length at the displacement recovery completion moment minus the projection length at the segmentation boundary time, with units of millimeters.
[0036] Step S135: Extract the cosine components of the displacement regression direction in each spatial dimension of the instrument end during the displacement recovery process to form the end displacement recovery direction parameters.
[0037] The computational environment extracts the difference vector between the displacement coordinates at the moment of displacement recovery completion and the displacement coordinates at the segmentation boundary. This difference vector is then normalized using the L2 norm. The components of the normalized vector in each spatial dimension are the cosine components of the displacement regression direction. The three cosine components of the displacement regression direction constitute the end-effector displacement recovery direction parameter. The data structure of the end-effector displacement recovery direction parameter is the same as that of the end-effector displacement offset direction parameter.
[0038] Step S140: The precursor waveform segment, end displacement retraction parameter, and end displacement offset direction parameter corresponding to each force mutation event are used as the precursor perception input of the instrument control intervention decision-making agent. The precursor pattern encoder in the instrument control intervention decision-making agent performs precursor pattern encoding processing on the precursor perception input to generate a precursor pattern encoding vector. The intervention strategy mapping network in the instrument control intervention decision-making agent uses the precursor pattern encoding vector and the aftereffect pattern encoding vector corresponding to the previous force mutation event as the joint decision basis to perform intervention strategy mapping processing to generate the instrument end force-position hybrid control command corresponding to the current force mutation event. The aftereffect pattern encoding vector is obtained by the aftereffect pattern encoder in the instrument control intervention decision-making agent performing aftereffect pattern encoding processing on the force mutation aftereffect waveform segment, end displacement recovery parameter, and end displacement recovery direction parameter. The instrument end force-position hybrid control command is used to adjust the applied force magnitude and displacement direction of the target medical instrument end.
[0039] The instrument control intervention decision-making agent is a decision-making model based on a deep neural network, comprising three sub-networks: a precursor pattern encoder, a follow-up pattern encoder, and an intervention strategy mapping network. The precursor pattern encoder is a feature encoding network based on a one-dimensional convolutional neural network, containing three one-dimensional convolutional layers and a global average pooling layer. The kernel size and stride of each one-dimensional convolutional layer are preset to an odd integer value. The first, second, and third one-dimensional convolutional layers contain a preset number of kernels. Each convolutional layer is followed by a linear activation function with leakage correction and layer normalization. The global average pooling layer averages the feature map output from the last convolutional layer along the time dimension, outputting a fixed-length feature vector. The input to the precursor pattern encoder is a one-dimensional vector concatenated from the precursor waveform segment of the sudden force change, the end displacement retraction parameter, and the end displacement offset direction parameter. The output of the precursor mode encoder is the precursor mode encoding vector, and the dimension of the precursor mode encoding vector is equal to the number of convolutional kernels in the third convolutional layer.
[0040] The network structure of the aftereffect mode encoder is exactly the same as that of the precursor mode encoder, containing the same number and configuration of one-dimensional convolutional layers and global average pooling layers. The input to the aftereffect mode encoder is a one-dimensional vector concatenated from the waveform segment resulting from the abrupt change in force, the end-displacement recovery parameter, and the end-displacement recovery direction parameter. The output of the aftereffect mode encoder is the aftereffect mode encoded vector, which has the same dimension as the precursor mode encoded vector.
[0041] The intervention strategy mapping network is a multilayer perceptron based on a fully connected neural network, consisting of two fully connected hidden layers and one output layer. The input dimension of the first fully connected hidden layer is equal to the sum of the dimensions of the precursor pattern encoding vector and the aftereffect pattern encoding vector, and the output dimension is the preset number of hidden layer units. The input dimension of the second fully connected hidden layer is the preset number of hidden layer units, and the output dimension is also the preset number of hidden layer units. Each hidden layer is followed by a linear activation function with leakage correction and layer normalization. The output layer is a fully connected layer with an input dimension equal to the preset number of hidden layer units and an output dimension of 4. The four components of the output vector correspond to the applied force adjustment, the horizontal displacement direction adjustment, the vertical displacement direction adjustment, and the depth displacement direction adjustment, respectively. The applied force adjustment is in Newtons, and the three displacement direction adjustments are dimensionless increments of the direction cosine.
[0042] The computing environment feeds the precursor sensing input of the current force abrupt change event into the precursor mode encoder to generate the current precursor mode encoding vector. The computing environment reads the aftereffect mode encoding vector corresponding to the previous force abrupt change event from the memory cache. If the current force abrupt change event is the first force abrupt change event in this operation, the previous aftereffect mode encoding vector is replaced with an all-zero vector. The computing environment concatenates the current precursor mode encoding vector and the previous aftereffect mode encoding vector along the feature dimension, and inputs the concatenated vector into the intervention strategy mapping network. After layer-by-layer computation, the intervention strategy mapping network outputs a hybrid force-position control command containing four components at the instrument end effector.
[0043] The computing environment simultaneously sends the aftereffect waveform segment of the current stress change event, the end displacement recovery parameter, and the end displacement recovery direction parameter to the aftereffect mode encoder to generate the current aftereffect mode encoding vector, which is stored in the memory cache for use in the decision-making of the next stress change event.
[0044] The method may further include: step S210: collecting historical instrument end force feedback signal streams and historical instrument end displacement feedback signal streams recorded in historical invasive target medical operations, and performing force mutation event detection processing on the historical instrument end force feedback signal streams to obtain a set of historical force mutation events and their corresponding historical force mutation precursor waveform segments and historical force mutation aftereffect waveform segments.
[0045] The computing environment reads multiple historical data records of invasive target medical procedures from the medical operation history database. Each data record contains an operation identifier string, a historical instrument end-effector force feedback signal stream, and a historical instrument end-effector displacement feedback signal stream. The data formats of the historical instrument end-effector force feedback signal stream and displacement feedback signal stream are the same as those defined in step S110. The computing environment processes the historical instrument end-effector force feedback signal stream in each data record according to the force mutation event detection and processing flow of steps S121 to S124, obtaining a set of historical force mutation events, as well as the historical force mutation precursor waveform segment and historical force mutation aftereffect waveform segment corresponding to each historical force mutation event.
[0046] Step S220: Based on the occurrence time of each historical force abrupt change event in the set of historical force abrupt change events, perform displacement response segmentation processing on the historical instrument end displacement feedback signal stream to obtain historical end displacement retraction parameters, historical end displacement offset direction parameters, historical end displacement recovery parameters, and historical end displacement recovery direction parameters.
[0047] For each historical force abrupt change event, the computing environment processes the displacement response segmentation and parameter extraction process in steps S131 to S135 on the corresponding historical instrument end displacement feedback signal stream at the time of occurrence, and obtains the historical end displacement retraction parameter, historical end displacement offset direction parameter, historical end displacement recovery parameter, and historical end displacement recovery direction parameter.
[0048] Step S230: Combine the historical stress change precursor waveform segment, historical end displacement shrinkage parameter, and historical end displacement offset direction parameter corresponding to each historical stress change event into training precursor input samples, and combine the historical stress change aftereffect waveform segment, historical end displacement recovery parameter, and historical end displacement recovery direction parameter corresponding to the same historical stress change event into training aftereffect input samples.
[0049] The computational environment constructs training sample pairs for each historical stress abrupt change event. The training precursor input sample is a one-dimensional floating-point array, composed of three components: an array of historical stress abrupt change waveform segments, a historical end-displacement retraction parameter, and a historical end-displacement offset direction parameter. The training post-effect input sample is a one-dimensional floating-point array, composed of three components: an array of historical stress abrupt change post-effect waveform segments, a historical end-displacement recovery parameter, and a historical end-displacement recovery direction parameter.
[0050] Step S240: Obtain the instrument end force-position hybrid control command actually applied by the operator when the corresponding historical force change event occurs as the training target command.
[0051] The computing environment extracts the instrument end effector force-position hybrid control command actually applied by the operator at the moment of each historical force abrupt change event from the operator's control records in the medical operation history database. The structure of this control command is the same as the control command output in step S140, including the applied force adjustment amount and three displacement travel direction adjustment amounts.
[0052] Step S250: Input the training precursor input sample into the precursor pattern encoder of the instrument control intervention decision agent to generate the training precursor pattern encoding vector, and input the training aftereffect input sample into the aftereffect pattern encoder of the instrument control intervention decision agent to generate the training aftereffect pattern encoding vector.
[0053] The computational environment feeds the training precursor input samples into the precursor pattern encoder. After passing through three one-dimensional convolutional layers for convolution, activation, and normalization, the samples are then processed by a global average pooling layer to generate the training precursor pattern encoding vector. The training post-effect input samples are then fed into the post-effect pattern encoder, which generates the training post-effect pattern encoding vector using the same processing flow.
[0054] Step S260: Input the training precursor pattern encoding vector and the training aftereffect pattern encoding vector corresponding to the previous historical force mutation event into the intervention strategy mapping network of the instrument control intervention decision agent to generate training prediction force-potential hybrid control instructions.
[0055] The computational environment obtains the training post-effect pattern encoding vector corresponding to the previous historical force mutation event from the training sequence. If the current event is the first event in the sequence, an all-zero vector is used. The training precursor pattern encoding vector and the previous training post-effect pattern encoding vector are concatenated and input into the intervention policy mapping network. After layer-by-layer calculation through two fully connected hidden layers and mapping through the output layer, a training prediction force-potential hybrid control command is generated.
[0056] Step S270: Calculate the instruction deviation between the training prediction force-position hybrid control instruction and the training target instruction, and adjust the internal connection weights of the precursor mode encoder, the aftereffect mode encoder and the intervention strategy mapping network with the goal of reducing the instruction deviation, until the instruction deviation converges to the preset allowable deviation range.
[0057] The computational environment calculates the mean squared error (MSE) loss between the training predicted force-position hybrid control command and the training target command. The MSE loss is calculated by squared the difference between the predicted and target values for each of the four components of the command, summing the four squares, and then dividing by 4. The computational environment uses an adaptive moment estimation optimizer, employing the MSE loss as the loss function. The error backpropagation algorithm is used to sequentially calculate the gradient of the loss with respect to the weight and bias parameters of each layer in the intervention policy mapping network, the aftereffect pattern encoder, and the precursor pattern encoder, and updates the parameters of each layer based on the gradients. Preset initial learning rate and batch size values are used during training. Training is performed in multiple epochs. Training stops and the model parameters are saved when the MSE loss on the validation set no longer decreases or falls below a preset tolerance range for several consecutive epochs.
[0058] The method may further include: step S310: after deploying the instrument control intervention decision-making agent to the target medical instrument, real-time acquisition of the real-time stress mutation precursor waveform segment, real-time end displacement retraction parameter, and real-time end displacement offset direction parameter corresponding to the latest stress mutation event in the current invasive target medical operation.
[0059] After the instrument control intervention decision-making agent completes the training in step S270 and is deployed to the embedded computing unit of the target medical instrument, the computing environment extracts the real-time stress mutation precursor waveform segment, real-time end displacement retraction parameter, and real-time end displacement offset direction parameter from the latest stress mutation event during real-time operation, following the processing flow from steps S120 to S135. The extracted data structure is the same as the corresponding data structure in the training phase.
[0060] Step S320: Input the real-time stress change precursor waveform segment, real-time end displacement retraction parameter, and real-time end displacement offset direction parameter into the instrument control intervention decision-making intelligent agent, and the instrument control intervention decision-making intelligent agent generates the instrument end force-position hybrid control command for the current stress change event.
[0061] Following the forward inference process of step S140, the computing environment sends the real-time precursor sensing input to the precursor mode encoder to generate a real-time precursor mode encoding vector. It then reads the real-time aftereffect mode encoding vector of the previous force-induced abrupt change event from the memory cache, concatenates the two vectors, and inputs them into the intervention strategy mapping network. The network outputs a hybrid force-position control command for the current force-induced abrupt change event at the instrument's end effector. The computing environment then sends this control command to the force-position actuator controller of the target medical instrument via the control bus.
[0062] Step S330: Within the preset observation window after executing the force-position hybrid control command at the instrument end, continuously monitor whether a new force mutation event that meets the conditions for force mutation event detection and processing reappears in the force feedback signal stream at the instrument end.
[0063] The computing environment starts timing from the moment the control command is executed, and continuously monitors the force feedback signal flow at the instrument end according to the force change event detection conditions in steps S121 to S122 within a preset observation window. If a sign flip point that meets the flip amplitude threshold condition and the flip speed threshold condition is detected within the observation window, a new force change event is determined to have occurred.
[0064] Step S340: If a new stress mutation event is detected within the preset observation window, extract the waveform segment of the new stress mutation effect corresponding to the new stress mutation event, and perform waveform similarity comparison processing between the waveform segment of the new stress mutation effect and the waveform segment of the stress mutation effect of the current stress mutation event.
[0065] If a new stress mutation event is detected within the preset observation window, the computing environment extracts the waveform segment corresponding to the new stress mutation event. The computing environment calculates the dynamic time warping distance between the waveform segment corresponding to the new stress mutation and the waveform segment corresponding to the current stress mutation event in step S140. The dynamic time warping distance is calculated by constructing a cumulative distance matrix between the two waveform segments and finding a warped path from the starting point to the ending point of the matrix through a dynamic programming algorithm, such that the cumulative sum of the Euclidean distances between the two waveform sample values corresponding to each point on the path is minimized. The smaller the dynamic time warping distance value, the higher the morphological similarity between the two waveform segments.
[0066] Step S350: When the waveform similarity comparison processing result indicates that the waveform similarity between the waveform segment after the new force mutation and the waveform segment after the current force mutation event is lower than the preset similarity lower limit, it is determined that the control effect of the instrument end force-position hybrid control command on the current force mutation event has not reached the expected level.
[0067] The computational environment compares the dynamic time warp distance value with a preset similarity lower limit threshold, which is a positive floating-point number. If the dynamic time warp distance value is greater than the preset similarity lower limit threshold, meaning the distance between the two waveform segments is large and the similarity is low, then the control effect is determined to be unsatisfactory.
[0068] Step S360: Based on the waveform segment of the precursor of the new force change event and the end displacement parameters, regenerate the modified instrument end force-position hybrid control command, and replace the original instrument end force-position hybrid control command with the modified instrument end force-position hybrid control command to perform force-position hybrid control.
[0069] The computational environment extracts the precursor waveform segment, end displacement retraction parameter, and end displacement offset direction parameter corresponding to the new force abrupt change event, and regenerates the corrected instrument end force-position hybrid control command according to the forward inference process in step S140. The computational environment overwrites the corrected command into the instruction register of the force-position actuator controller of the target medical instrument via the control bus, replacing the original control command that did not meet expectations. The force-position actuator controller adjusts the applied force magnitude and displacement direction according to the corrected command.
[0070] The method may further include: step S410: during the process of multiple medical instruments performing the same type of invasive target medical operation, collecting the set of historical force mutation events uploaded by each medical instrument and the historical instrument end force-position hybrid control command corresponding to each historical force mutation event.
[0071] The computing environment collects historical data from the operation log storage server of multiple medical instruments of the same model. The collected historical data includes the operation identifier of each operation record, the set of historical force mutation events, and the historical instrument end force-position hybrid control command applied by the operator for each historical force mutation event.
[0072] Step S420: Extract the waveform envelope features and waveform spectral distribution features of the force change rate sequence from the historical force change precursor waveform segments of each historical force change event as the force change precursor pattern fingerprint.
[0073] For each historical stress-induced abrupt change event, the computational environment calculates its first-order difference sequence as the stress change rate sequence. The upper and lower envelopes of the stress change rate sequence are calculated, obtained by fitting local maxima and minima using cubic spline interpolation. The waveform envelope features include the peak amplitude of the upper envelope and the valley amplitude of the lower envelope. A Fast Fourier Transform is performed on the stress change rate sequence to obtain the spectral amplitude distribution, which is a complex array. The waveform spectral distribution features are arrays consisting of the amplitude values of the first predetermined number of frequency components in the spectral amplitude distribution. The waveform envelope feature array and the waveform spectral distribution feature array are concatenated into a one-dimensional floating-point array, which is the stress-induced abrupt change pattern fingerprint.
[0074] Step S430: Cluster the set of historical stress-induced mutation events based on the stress-induced mutation precursor pattern fingerprint to obtain stress-induced mutation event clusters with similar stress-induced mutation precursor pattern fingerprints.
[0075] The computing environment uses the precursor pattern fingerprints of all historical stress-induced abrupt changes as the input feature vector set and employs a density-based spatial clustering algorithm for clustering. The density-based spatial clustering algorithm identifies density-connected sample clusters in the feature space based on preset neighborhood radii and preset minimum sample values. In the clustering results, each cluster represents a stress-induced abrupt change event cluster, and the historical stress-induced abrupt change events within the cluster share similar precursor pattern fingerprints.
[0076] Step S440: Perform instruction statistical processing on the historical instrument end force-position hybrid control instructions corresponding to each historical force-change event within each force-change event cluster to determine the distribution range and concentration trend of control instructions corresponding to each force-change event cluster.
[0077] The computational environment calculates the arithmetic mean and standard deviation of the four components of the historical instrument end-force-potential hybrid control commands within each cluster of force-induced abrupt events. The central tendency of the control commands is represented by a vector composed of the arithmetic means of the four components. The distribution interval of the control commands is determined by adding or subtracting a preset multiple of the standard deviation of the arithmetic mean of each component.
[0078] Step S450: Construct a mapping table between the fingerprint of the precursor pattern of stress mutation and the trend of the control command set, and load the mapping table into the initial mapping layer of the intervention strategy mapping network of the instrument control intervention decision agent.
[0079] The computing environment constructs a mapping table, which is an associative array. The key is an integer value representing a cluster identifier, and the values are structures containing the fingerprint vector of the precursor pattern of the cluster centers and the trend vector in the set of control instructions. When training the instrument-controlled intervention decision agent, the trend vector in the set of control instructions in the mapping table is used as the initial value of the bias parameter of the output layer of the intervention strategy mapping network, and the fingerprint vector of the precursor pattern of the cluster centers is used as the initial value of the bias parameter of the last convolutional layer of the precursor pattern encoder, thus completing the initial loading of historical experience.
[0080] The method may further include: step S510: obtaining a sequence of preoperative planning path points generated by the preoperative planning system before the target medical instrument performs the invasive target medical operation, the sequence of preoperative planning path points containing the three-dimensional spatial coordinates of multiple planning path points arranged in the order of operation.
[0081] The computing environment obtains the preoperative planning path point sequence from the preoperative planning system through the medical image archiving and communication system interface. The preoperative planning path point sequence is a JSON array, and each element in the array is a planning path point object, containing an index field and a three-dimensional coordinate field. The three-dimensional coordinate field contains horizontal coordinate values, vertical coordinate values, and depth coordinate values, all in millimeters.
[0082] Step S520: After the instrument control intervention decision-making agent generates a hybrid control command for the instrument end force and position for each force mutation event, the displacement direction adjustment amount carried in the hybrid control command is extracted. The actual displacement direction of the instrument end indicated by the displacement direction adjustment amount is compared with the planned path direction corresponding to the current operation step in the preoperative planned path point sequence to calculate the spatial angle deviation between the actual displacement direction of the instrument end and the planned path direction.
[0083] The computational environment extracts three displacement direction adjustment components from the instrument's end-effector force-position hybrid control command. These three components constitute the direction cosine vector of the actual displacement direction. The computational environment determines the planned path point corresponding to the current operation step. The planned path direction is represented by the normalized direction vector pointing from the current planned path point to the next planned path point. The spatial angle deviation is equal to the angle between the actual displacement direction vector and the planned path direction vector. The angle is calculated by taking the dot product of the two vectors and using the inverse cosine of the dot product, with the dimension in degrees.
[0084] Step S530: Obtain the real-time three-dimensional spatial coordinates of the target medical instrument tip during the invasive target medical procedure, and calculate the remaining distance deviation between the real-time three-dimensional spatial coordinates and the next planned path point in the preoperative planned path point sequence.
[0085] The computational environment extracts the real-time 3D spatial position coordinates from the latest sampled frame of the displacement feedback signal stream at the instrument's end. The remaining distance deviation is equal to the Euclidean distance between the real-time 3D spatial position coordinates and the coordinates of the next planned path point, calculated as the square root of the sum of the squares of the differences between the components of the two 3D coordinate vectors, with the dimension in millimeters.
[0086] Step S540: Input the spatial angle deviation and the remaining distance deviation into the path correction priority arbitration logic. The path correction priority arbitration logic determines whether the priority control target of the current operation step is to reduce the spatial angle deviation or the remaining distance deviation, and outputs the priority control target identifier.
[0087] The path correction priority arbitration logic is a rule-based and fuzzy inference-based decision-making logic. It fuzzifies the spatial angle deviation and the remaining distance deviation into multiple fuzzy sets, including small, medium, and large sets. The fuzzification process involves calculating the membership degree of the input value to each fuzzy set, using triangular and trapezoidal membership functions. Based on a pre-defined fuzzy inference rule table, and combining the fuzzy sets of spatial angle deviation and remaining distance deviation, the priority control target is inferred. The fuzzy inference rule table contains multiple rules, each stating that if both the spatial angle deviation and the remaining distance deviation belong to a certain fuzzy set, then the priority control target is a certain target identifier. The Mamdani fuzzy inference method and the centroid method are used for defuzzification to obtain the priority control target identifier. The value of the priority control target identifier is either a direction priority constant or a distance priority constant.
[0088] Step S550: When the priority control target indicator indicates that the spatial angle deviation should be reduced first, a directional correction control component is generated. The directional correction control component is used to correct the displacement direction adjustment in the force-position hybrid control command at the end of the instrument.
[0089] When the priority control target identifier is set to a direction priority constant, the computational environment calculates the direction correction control component. The direction correction control component is a three-dimensional vector with the same dimension as the displacement direction adjustment amount. This vector points towards the planned path direction, and its magnitude is proportional to the spatial angle deviation. The computational environment performs a weighted summation of the direction correction control component and the original displacement direction adjustment amount. The weighting coefficients are dynamically determined based on the fuzzy set membership degree of the spatial angle deviation.
[0090] Step S560: When the priority control target identifier indicates that the remaining distance deviation should be reduced first, a speed correction control component is generated. The speed correction control component is used to correct the applied force adjustment in the instrument end force-position hybrid control command to change the propulsion speed of the instrument end.
[0091] When the priority control target identifier is set to a distance priority constant, the computational environment calculates the velocity correction control component. The velocity correction control component is a scalar value representing the increment of the applied force adjustment. The magnitude of the velocity correction control component is proportional to the remaining distance deviation. The computational environment arithmetically adds the velocity correction control component to the original applied force adjustment.
[0092] Step S570: Superimpose the direction correction control component or the speed correction control component onto the instrument end force-position hybrid control command to obtain the instrument end force-position hybrid control command after path correction.
[0093] Based on the type of the priority control target identifier, the computing environment superimposes the directional correction control component of step S550 or the velocity correction control component of step S560 onto the corresponding component of the instrument end-effector force-position hybrid control command. The superimposed command is the instrument end-effector force-position hybrid control command after path correction. The computing environment then sends this command to the force-position actuator controller.
[0094] The method may further include: step S610: during the invasive target medical operation performed by the target medical instrument, the tissue impedance feedback signal flow of the biological tissue contacted by the end of the target medical instrument is collected in real time, and the tissue impedance feedback signal flow and the force feedback signal flow at the end of the instrument have a time-synchronized acquisition relationship.
[0095] The computational environment acquires the tissue impedance feedback signal stream at the same sampling frequency as the force sensor through an impedance measurement loop formed by the contact between the terminal electrode of the target medical instrument and the biological tissue. The data structure of the tissue impedance feedback signal stream is a one-dimensional floating-point array, where each element is the biological tissue impedance value measured at that sampling moment, in ohms. Each impedance sampling data frame contains a timestamp field and an impedance value field. The timestamp field is timed by the same clock source as the force sensor, ensuring time synchronization with the force feedback signal stream at the instrument's end.
[0096] Step S620: Perform impedance layer identification processing on the tissue impedance feedback signal stream, and determine the set of tissue layer boundary moments by detecting the moments when the impedance value in the tissue impedance feedback signal stream undergoes a step change.
[0097] The computational environment performs impedance stratification identification processing on the tissue impedance feedback signal stream. This impedance stratification identification process employs a sliding window-based change-point detection algorithm. The computational environment performs a sliding window scan on the impedance feedback signal stream with a preset window length, calculating the mean difference of impedance values within two adjacent sliding windows. If the absolute value of the mean difference exceeds a preset impedance step threshold, an impedance step point is marked at the boundary between the two windows. The impedance step threshold is preset based on the range of impedance differences among different tissue types in the target biological tissue. The timestamps corresponding to all marked impedance step points are arranged in ascending order, forming a set of tissue stratification boundary times. This set of tissue stratification boundary times is an array of timestamp values.
[0098] Step S630: Perform time correlation matching processing on each organizational stratification boundary moment in the set of organizational stratification boundary moments and the occurrence time of the stress mutation event, identify the stress mutation events that coincide with each organizational stratification boundary moment in the set of organizational stratification boundary moments in time, and mark them as organizational crossing type stress mutation events, and mark the stress mutation events that do not coincide with any organizational stratification boundary moment as non-organization crossing type stress mutation events.
[0099] The computing environment calculates the time difference between the occurrence time of each stress-induced abrupt change event detected in step S120 and the time differences between the occurrence times of each tissue layer boundary in the tissue layer boundary time set. If the absolute value of the time difference between the occurrence time of a stress-induced abrupt change event and the nearest tissue layer boundary time is less than a preset time coincidence threshold, the stress-induced abrupt change event is marked as a tissue-crossing stress-induced abrupt change event. The time coincidence threshold is measured in milliseconds. Stress-induced abrupt change events that do not meet the time coincidence condition are marked as non-tissue-crossing stress-induced abrupt change events. The event type attribute field of each stress-induced abrupt change event records its marking result.
[0100] Step S640: For tissue-crossing force mutation events, extract the impedance change amplitude and direction in the tissue impedance feedback signal stream before and after the occurrence of the tissue-crossing force mutation event, generate tissue type change features, input the tissue type change features into the intervention strategy mapping network of the instrument control intervention decision agent, perform tissue type adaptive weighted correction processing on the instrument end force-position hybrid control command generated by the intervention strategy mapping network, and obtain the instrument end force-position hybrid control command based on tissue identification.
[0101] For each tissue-crossing force abrupt change event, the computational environment extracts the average impedance within a preset forward window before the event's occurrence as the pre-crossing impedance value, and extracts the average impedance within a preset backward window after the event's occurrence as the post-crossing impedance value. The impedance change amplitude is equal to the absolute value of the difference between the post-crossing impedance value and the pre-crossing impedance value, with the dimension ohms. The impedance change direction is indicated by the sign of the difference, taking either positive or negative values. The impedance change amplitude and direction together constitute the tissue type transition feature, which is a structure containing an amplitude field and a direction field.
[0102] The computational environment encodes tissue type transition features into a two-dimensional tissue type transition encoding vector. The amplitude field, after normalization, serves as the first dimension of the vector, and the direction field is mapped to the second dimension. An tissue type adaptive weighting layer is inserted before the output layer of the intervention strategy mapping network. This weighting layer receives the hidden vector output from the second hidden layer of the intervention strategy mapping network and the tissue type transition encoding vector, and calculates the tissue type adaptive weighting coefficients through a gating mechanism. The gating mechanism is a fully connected layer with the same output dimension as the hidden vector, and the output values are mapped to the range of 0 to 1 using a sigmoid activation function. The hidden vector is then multiplied element-wise by these weighting coefficients to obtain the tissue type adaptively weighted hidden vector, which is then fed into the output layer to generate a tissue-based instrument end-effector force-position hybrid control command.
[0103] Step S650: For non-organic traversal type force mutation events, maintain the instrument end force-potential hybrid control command generated by the intervention strategy mapping network unchanged.
[0104] For events labeled as non-tissue-crossing force mutation events, the computing environment skips the tissue type adaptive weighted correction process in step S640 and directly uses the instrument end force-position hybrid control command generated by the intervention strategy mapping network as the final output.
[0105] For example, the method may further include: step S710: during the invasive target medical operation performed by the target medical instrument, obtaining the movable spatial boundary model of the end effector of the target medical instrument in free space, wherein the movable spatial boundary model is obtained by preoperative medical image reconstruction and defines the three-dimensional spatial range in which the end effector of the target medical instrument is allowed to move.
[0106] The computational environment obtains the movable spatial boundary model from the preoperative planning system. The data structure of the movable spatial boundary model is a three-dimensional binary mask array, where the three dimensions correspond to the horizontal, vertical, and depth directions in the surgical area's spatial coordinate system, respectively. Voxels with a value of 1 represent safe areas where the instrument's end effector is allowed to enter, while voxels with a value of 0 represent prohibited areas or areas outside the anatomical structure. This binary mask array is generated from preoperative computed tomography or magnetic resonance imaging images using three-dimensional reconstruction and surgical area segmentation algorithms.
[0107] Step S720: Extract the displacement direction adjustment amount carried in the instrument end force-position hybrid control command generated by the instrument control intervention decision-making agent, and predict the expected position coordinates of the instrument end in the next operation step based on the displacement direction adjustment amount and the current three-dimensional spatial position coordinates of the instrument end.
[0108] The computational environment extracts the current three-dimensional spatial coordinates from the displacement feedback signal stream at the instrument's end effector, denoted as the current coordinate vector. It then extracts the three displacement direction adjustment components carried in the force-position hybrid control command at the instrument's end effector, forming a displacement direction adjustment vector. The expected position coordinates of the instrument's end effector are equal to the current coordinate vector plus the product of the displacement direction adjustment vector and a preset operation step length. The magnitude of the displacement direction adjustment vector is preset to the operation step length, with dimensions in millimeters. The expected position coordinates of the instrument's end effector are a three-dimensional floating-point vector, with each component having dimensions in millimeters.
[0109] Step S730: Perform spatial inclusion relationship determination processing on the expected arrival position coordinates of the instrument terminal and the movable space boundary model to determine whether the expected arrival position coordinates of the instrument terminal are within the three-dimensional space defined by the movable space boundary model.
[0110] The computational environment converts the physical space coordinates of the expected arrival position of the instrument's end effector into voxel index coordinates of the movable space boundary model. The conversion is achieved by subtracting the origin coordinates of the movable space boundary model from the physical space coordinates, dividing by the voxel size, and rounding to obtain the voxel index 3D coordinates. The computational environment then uses these voxel index coordinates to query the corresponding value in the 3D binary mask array of the movable space boundary model. If the retrieved value is 1, the expected arrival position of the instrument's end effector is determined to be within the 3D space defined by the movable space boundary model. If the retrieved value is 0 or the voxel index coordinates exceed the array boundary, the expected arrival position of the instrument's end effector is determined to be outside the 3D space.
[0111] Step S740: When the output of the spatial inclusion relationship determination process indicates that the expected arrival position coordinates of the instrument end are outside the three-dimensional space range defined by the movable space boundary model, the displacement travel direction adjustment amount is subjected to directional constraint projection processing. The expected arrival position coordinates of the instrument end are projected onto the boundary surface of the movable space boundary model along the shortest distance direction to obtain the projected safe expected arrival position coordinates.
[0112] When step S730 determines that the expected arrival position coordinates of the instrument's end effector exceed the 3D spatial range, the computational environment performs directional constraint projection processing. In the binary mask array of the movable space boundary model, the computational environment uses the exceeding voxel index coordinates as the starting point and employs a 3D Euclidean distance transformation algorithm to calculate the shortest spatial distance from these coordinates to all voxels with a value of 1. The distance transformation algorithm outputs the voxel index coordinates of the nearest safe voxel, which are the projected safe expected arrival position coordinates. The computational environment then uses the projected safe expected arrival position coordinates to calculate the physical space coordinates.
[0113] Step S750: Calculate the corrected displacement direction adjustment amount based on the projected expected safe arrival position coordinates, and replace the original displacement direction adjustment amount with the corrected displacement direction adjustment amount and write it into the instrument end force-position hybrid control command.
[0114] The computational environment calculates the difference vector between the projected safe arrival position coordinates and the current coordinate vector. This difference vector is then normalized using the L2 norm, and the normalized vector is the corrected displacement direction adjustment. The three components of the corrected displacement direction adjustment replace the corresponding three displacement direction adjustment components in the original instrument end-effector force-position hybrid control command, and are written into the command data structure.
[0115] Step S760: When the output of the spatial inclusion relationship determination process indicates that the expected position coordinates of the instrument end are within the three-dimensional space defined by the movable space boundary model, the original displacement direction adjustment amount remains unchanged.
[0116] When step S730 determines that the expected position coordinates of the instrument end are within the range of three-dimensional space, the computing environment does not perform any correction operation, maintains the original displacement direction adjustment amount unchanged, and directly outputs the original instrument end force-position hybrid control command.
[0117] The method may further include: step S810: obtaining a set of historical force mutation events uploaded by the same type of target medical instruments deployed in different medical institutions when performing invasive target medical operations, and the historical instrument end force-position hybrid control command applied by the operator for each historical force mutation event.
[0118] The computing environment queries the data lake of the cross-institutional medical instrument operation data center for historical operation records of the target medical instrument of the same model. The query criteria are that the instrument model string and the operation type string are equal. The result set returned by the query includes the operation identifier, the set of historical force mutation events, and the historical instrument end force-position hybrid control command corresponding to each historical force mutation event. The data structure of the historical instrument end force-position hybrid control command is the same as the command structure output in step S140.
[0119] Step S820: Extract waveform morphology features from the waveform segments of the precursors of each historical stress change event to obtain waveform morphology feature vectors including waveform rise slope, waveform fall slope, waveform peak-to-valley amplitude ratio, and waveform duration.
[0120] The computing environment extracts waveform morphology features from the historical waveform segments preceding each historical force abrupt change event. The waveform rise slope is calculated as the ratio of the difference in axial force components from the start time to the peak force time to the time difference in the waveform segment preceding the force abrupt change, expressed in Newtons per second. The waveform fall slope is calculated as the ratio of the difference in axial force components from the peak force time to the end time to the time difference in the waveform segment preceding the force abrupt change, expressed in Newtons per second. The waveform peak-to-valley amplitude ratio is the ratio of the maximum to the minimum axial force component values in the waveform segment preceding the force abrupt change, and is dimensionless. The waveform duration is the time difference between the end and start times of the waveform segment preceding the force abrupt change, expressed in milliseconds. The four values—waveform rise slope, waveform fall slope, waveform peak-to-valley amplitude ratio, and waveform duration—are combined into a four-dimensional floating-point array, which constitutes the waveform morphology feature vector.
[0121] Step S830: Perform cluster analysis on the waveform morphology feature vectors of all historical stress-induced abrupt events to obtain multiple waveform morphology categories of stress-induced abrupt events, and calculate the category center waveform morphology feature vector for each category of stress-induced abrupt events.
[0122] The computing environment uses the waveform morphology feature vectors of all historical stress-induced abrupt events as input and performs cluster analysis using the K-means clustering algorithm. The parameter K in the K-means clustering algorithm is a preset number of clusters. The K-means clustering algorithm iteratively performs two steps: cluster allocation and center update, until the change in the position of the cluster center is less than a preset convergence threshold. After clustering, K waveform morphology categories of stress-induced abrupt events are obtained. For each category, the arithmetic mean of all waveform morphology feature vectors within that category is calculated; this mean vector is the cluster center waveform morphology feature vector of that stress-induced abrupt event waveform morphology category.
[0123] Step S840: For each category of pre-stress waveform morphology, collect all historical instrument end force-position hybrid control commands corresponding to all historical stress-induced events belonging to that category, forming a category-specific control command set. Perform statistical distribution analysis on the historical instrument end force-position hybrid control commands in each category-specific control command set, and extract the optimal control command vector corresponding to the pre-stress waveform morphology of that category.
[0124] For each pre-stress abrupt change waveform morphology category obtained from clustering in step S830, the computing environment collects all historical instrument end-effector force-position hybrid control commands corresponding to all historical stress abrupt change events within that category, forming a category-specific control command set. This category-specific control command set is an array of command vectors, where each command vector contains an applied force adjustment and three displacement direction adjustments. The computing environment calculates the arithmetic mean of the four components of all command vectors in the category-specific control command set, and uses this average vector as the optimal control command vector corresponding to that pre-stress abrupt change waveform morphology category.
[0125] Step S850: Construct a mapping table between the waveform morphology categories of the pre-stress mutation and the optimal control command vector, and load the mapping table into the initial weight configuration of the intervention strategy mapping network of the instrument control intervention decision agent.
[0126] The computational environment constructs a mapping table, which is an associative array. The keys are integer values representing the waveform morphology category identifiers of the precursor to a force mutation, and the values are the optimal control command vectors. During the training and initialization phase of the instrument control intervention decision agent, the optimal control command vectors for each waveform morphology category of the precursor to a force mutation in the mapping table are used as the prior initialization values for the bias parameters of the output layer of the intervention strategy mapping network. The waveform morphology feature vectors of the category centers corresponding to the category identifiers are used as the bias prior values for the fully connected layer of the precursor mode encoder.
[0127] Step S860: When the instrument control intervention decision-making agent generates a hybrid control command for the instrument end force and position in response to a real-time force change event, the optimal control command vector corresponding to the current force change precursor waveform morphology category is extracted from the control command vector output by the intervention strategy mapping network as the control command generation benchmark. The control command generation benchmark and the control command vector generated by the instrument control intervention decision-making agent in real-time reasoning are weighted and fused to obtain a hybrid control command for the instrument end force and position that incorporates historical experience.
[0128] During the real-time inference phase, the computing environment extracts the waveform morphology feature vectors of historical precursor waveform segments of the current real-time stress abrupt change event. It calculates the Euclidean distance between this feature vector and the category center waveform morphology feature vectors of each stress abrupt change precursor waveform morphology category, and selects the stress abrupt change precursor waveform morphology category with the smallest distance as the current stress abrupt change precursor waveform morphology category. The optimal control command vector corresponding to this stress abrupt change precursor waveform morphology category is retrieved from the mapping table and used as the control command generation benchmark. The weighted fusion processing method involves weighting the control command generation benchmark and the control command vector output in real-time by the intervention strategy mapping network according to preset historical experience weight coefficients and real-time inference weight coefficients, with the sum of the historical experience weight coefficients and the real-time inference weight coefficients being 1. The fused vector is the instrument end-effector force-position hybrid control command that incorporates historical experience.
[0129] The method may further include: step S910: during the process of the instrument control intervention decision-making agent generating instrument end force-position hybrid control command for multiple consecutive force mutation events, recording the force mutation precursor waveform segment corresponding to each force mutation event and the applied force adjustment amount in the instrument end force-position hybrid control command generated by the instrument control intervention decision-making agent for the force mutation event.
[0130] During the real-time operation of the instrument control intervention decision-making agent, the computing environment maintains a queue of force adjustment history records. Each element of the queue is a structure containing a force change event identifier field, a force change precursor waveform segment field, and an applied force adjustment amount field. The value of the applied force adjustment amount field is the first component of the instrument end-effector force-potential hybrid control command generated in step S140, with dimensions in Newtons.
[0131] Step S920: Extract the applied force adjustment amount of the previous force change event and the peak force amplitude in the waveform segment of the preceding force change event of the next force change event, and construct a correlation sample pair between the applied force adjustment and the subsequent peak force.
[0132] The computational environment iterates through the historical records of force adjustment events, traversing the records of two adjacent force abrupt change events. It extracts the applied force adjustment value from the preceding force abrupt change event and the peak force amplitude value from the preceding waveform segment of the subsequent force abrupt change event. The peak force amplitude value is the maximum value of the axial force component in the waveform segment. A force adjustment value and a peak force amplitude value are combined to form a correlation sample pair. All correlation sample pairs constructed from adjacent event pairs are collected into an array.
[0133] Step S930: Perform trend analysis on the force adjustment effect of the associated sample pairs of force adjustment and subsequent force peak values accumulated from multiple consecutive force abrupt events, and determine the suppression response curve of the applied force adjustment amount on the amplitude of the force peak value of the subsequent force abrupt event.
[0134] The computational environment uses the array of associated sample pairs accumulated in step S920, with the applied force adjustment amount as the independent variable and the peak force amplitude as the dependent variable, to fit the inhibition response curve using a locally weighted regression algorithm. The kernel function for the locally weighted regression is a trigonometric kernel function, and the bandwidth parameter is the reciprocal of the square root of the number of associated sample pairs. The fitted inhibition response curve is a univariate continuous function defined within the range of applied force adjustment amounts, and the function value represents the expected subsequent peak force amplitude under a given applied force adjustment amount.
[0135] Step S940: Identify the effective adjustment range and over-adjustment range of the force adjustment amount applied by the instrument control intervention decision-making agent under the current operating environment based on the inhibition response curve.
[0136] The computational environment calculates the first derivative of the suppression response curve. Within the range of applied force adjustment, the intervals where the first derivative is negative and its absolute value is greater than a preset effective slope threshold are marked as effective adjustment intervals. The intervals where the absolute value of the first derivative is less than or equal to the preset effective slope threshold, or where the first derivative is positive, are marked as over-adjustment intervals. An effective adjustment interval indicates that within this range, increasing the applied force adjustment effectively reduces the subsequent peak force amplitude. An over-adjustment interval indicates that within this range, increasing the applied force adjustment has a weak effect on reducing the subsequent peak force amplitude or has the opposite effect.
[0137] Step S950: When the force adjustment amount generated by the instrument control intervention decision agent in response to the current force change event falls into the over-adjustment range, the force adjustment amount truncation correction process is triggered to limit the force adjustment amount to the upper limit boundary value of the effective adjustment range.
[0138] The computational environment checks whether the currently generated applied force adjustment value falls within the over-adjustment range. If it does, the applied force adjustment value is replaced with the upper boundary value of the effective adjustment range. The upper boundary value is the applied force adjustment value at the boundary between the effective adjustment range and the over-adjustment range. If it falls within the effective adjustment range, the original applied force adjustment value remains unchanged.
[0139] Step S960: Extract the displacement direction adjustment amount from the instrument end force-position hybrid control command generated by the instrument control intervention decision-making agent in response to the current force change event, and perform force-position coupling consistency verification processing on the displacement direction adjustment amount and the applied force adjustment amount.
[0140] The computational environment extracts three displacement direction adjustment components from the current instrument end-effector force-position hybrid control command. Force-position coupling consistency verification is performed based on a preset force-position coordination mapping relationship, which is a lookup table. The lookup table uses the applied force adjustment value as the key and the allowable range of the displacement direction adjustment value as the value. The allowable range defines the minimum and maximum values of each component of the displacement direction adjustment value under the applied force adjustment value. The computational environment checks whether each component value of the displacement direction adjustment value falls within the allowable range corresponding to the applied force adjustment value. If all three components fall within the allowable range, the force-position coupling consistency verification passes. If any component does not fall within the allowable range, the force-position coupling consistency verification fails.
[0141] Step S970: When the output of the force-position coupling consistency verification process indicates that the displacement direction adjustment amount and the applied force adjustment amount after truncation correction processing do not meet the preset force-position coordination mapping relationship, the displacement direction adjustment amount is synchronously corrected so that the corrected displacement direction adjustment amount and the corrected applied force adjustment amount meet the force-position coordination mapping relationship. The displacement direction adjustment amount after synchronous correction processing and the applied force adjustment amount after truncation correction processing are written together into the force-position hybrid control command at the instrument end, so as to obtain the force-position coordinated instrument end force-position hybrid control command.
[0142] When the force-position coupling consistency check fails, the computational environment performs synchronous correction processing on each component of the displacement direction adjustment. The synchronous correction process involves truncating the values of components exceeding the allowable range to the nearest boundary value within that range. Components within the allowable range are left unchanged. The three synchronously corrected displacement direction adjustment components, along with the truncated force adjustment value from step S950, are written into the instrument end-effector force-position hybrid control command, replacing the original command component values, resulting in a force-position coordinated instrument end-effector force-position hybrid control command. The computational environment then sends this command to the force-position actuator controller.
[0143] The method may further include: step S1010: during the invasive target medical operation performed by the target medical instrument, acquiring the operator-active control torque signal flow applied by the operator to the end of the target medical instrument through the control handle.
[0144] The computing environment acquires the operator's active control torque signal stream through the torque sensor built into the control handle. The data frame from the control handle's torque sensor includes a timestamp field and a torque component array field. The torque component array field is a three-dimensional floating-point array, with the three components representing the axial torque component, the first lateral torque component, and the second lateral torque component, all measured in Newton-meters. The timestamp of the data frame uses the same clock source as the force sensor and displacement sensor. The data structure of the operator's active control torque signal stream is a two-dimensional floating-point array, where the first dimension corresponds to the sampling time index, and the second dimension corresponds to the three torque components.
[0145] Step S1020: Perform human-computer interaction intent analysis on the operator's active control torque signal stream and the instrument end force feedback signal stream. By comparing the temporal correspondence between the direction change sequence of the operator's active control torque signal stream and the waveform segment of the force change precursor of the force change event in the instrument end force feedback signal stream, identify the operator's predictive intervention actions before the force change event occurs.
[0146] The computing environment detects directional changes in the operator's active control torque signal flow, extracting the moments and directions of significant deflections in the control torque direction. A significant deflection is determined when the angle between the control torque direction vectors of adjacent sampling points exceeds a preset directional change threshold. The computing environment performs a temporal correlation comparison between these significant deflection moments and the occurrence times of force-induced abrupt events detected in step S120. If a significant deflection moment occurs between the start and occurrence times of a force-induced abrupt event's precursor, and the deflection direction is consistent with the force direction change trend in the precursor waveform segment, then the control action corresponding to this significant deflection moment is determined to be a predictive intervention action by the operator.
[0147] Step S1030: When an operator's predictive intervention action is detected, the direction and amplitude of the control torque of the operator's predictive intervention action are extracted, and the direction and amplitude of the control torque are converted into the operator's intention control component. The operator's intention control component is then fused with the instrument end force position hybrid control command generated by the instrument control intervention decision-making agent through human-machine collaborative fusion processing. The operator's intention control component and the instrument end force position hybrid control command are weighted and superimposed through a preset human-machine collaborative weight allocation strategy to obtain the human-machine collaborative instrument end force position hybrid control command.
[0148] The computational environment extracts the direction vector and amplitude of the control torque at the moment of the operator's predictive intervention. After normalizing the direction vector, it is multiplied by a preset torque-to-position control coefficient to obtain the operator's intended force control component and the operator's intended displacement direction control component. The dimension of the operator's intended force control component is Newton, and the dimension of the operator's intended displacement direction control component is a dimensionless direction cosine increment. The human-machine collaboration weight allocation strategy presets operator weight coefficients and agent weight coefficients, the sum of which is 1. The weighted superposition method involves multiplying the operator's intended control component by the operator weight coefficient, and adding the instrument end-effector force-position hybrid control command generated by the instrument control intervention decision agent multiplied by the agent weight coefficient, to obtain the human-machine collaborative instrument end-effector force-position hybrid control command.
[0149] Step S1040: After the stress mutation event ends, collect the stress mutation aftereffect waveform segment corresponding to the stress mutation event, and extract the aftereffect decay rate index of the stress mutation aftereffect waveform segment.
[0150] After the force abrupt change event, the computing environment acquires the waveform segment of the aftereffect extracted in step S124. The aftereffect decay rate is calculated by fitting an exponential decay function to the waveform segment of the aftereffect. The form of the exponential decay function is that the force value equals the initial peak force multiplied by a negative decay coefficient of the natural constant e multiplied by a power of time. The fitting uses a nonlinear least squares method to obtain the decay coefficient, which is the aftereffect decay rate index, with the dimension of second.
[0151] Step S1050: Compare the aftereffect decay rate index with the preset aftereffect decay rate benchmark. When the aftereffect decay rate is faster than the aftereffect decay rate benchmark, increase the value of the operator weight coefficient in the human-machine collaborative weight allocation strategy; when the aftereffect decay rate is slower than the aftereffect decay rate benchmark, decrease the value of the operator weight coefficient in the human-machine collaborative weight allocation strategy.
[0152] The computing environment compares the aftereffect decay rate index with a preset aftereffect decay rate benchmark value. The preset benchmark value is the average aftereffect decay rate of similar force-induced abrupt events statistically obtained from historical operational data. If the aftereffect decay rate is greater than the benchmark value multiplied by a preset rapid decay judgment coefficient, meaning the aftereffect decay is faster than the benchmark, it indicates that the operator's predictive intervention effectively accelerated force recovery, and the operator's weight coefficient is increased by a preset weight adjustment step size. If the aftereffect decay rate is less than the benchmark value divided by a preset slow decay judgment coefficient, meaning the aftereffect decay is slower than the benchmark, it indicates that the operator's predictive intervention failed to effectively accelerate force recovery, and the operator's weight coefficient is decreased by a preset weight adjustment step size. The updated operator weight coefficient and agent weight coefficient are stored in the collaborative weight configuration area in memory.
[0153] Step S1060: Apply the dynamically adjusted operator weight coefficient to the human-machine collaborative fusion processing of subsequent force mutation events, and update the human-machine collaborative weight allocation strategy.
[0154] The computing environment writes the updated operator weight coefficient from step S1050 into the human-machine collaborative weight allocation strategy configuration. When a subsequent force mutation event triggers the human-machine collaborative fusion processing in step S1030, the updated operator weight coefficient and agent weight coefficient are used for weighted superposition calculation to achieve online adaptation to the operator's control style.
[0155] The method may further include: step S1110: after the invasive target medical operation is completed, obtain the occurrence time sequence of all force mutation events recorded throughout the entire invasive target medical operation and the instrument end force-position hybrid control command corresponding to each force mutation event.
[0156] After detecting the operation completion event, the computing environment reads the sequence of occurrence times of all force abrupt events from the operation process data log file. The sequence of occurrence times is an array of timestamps arranged in ascending order of time. At the same time, it reads the instrument end-effector force-position hybrid control command corresponding to each force abrupt event. The data structure of the command includes the applied force adjustment amount and the adjustment amounts of three displacement travel directions.
[0157] Step S1120: Based on the temporal distribution density of the occurrence time sequence, the invasive target medical operation is divided into operation stages. The time period that is continuous in time and the frequency of the force mutation event is higher than the preset frequency threshold is divided into the first operation segment, and the time period that the frequency of the force mutation event is lower than the preset frequency threshold is divided into the second operation segment.
[0158] The computing environment performs temporal distribution density analysis on the occurrence time series. A sliding window method is used, with a window width equal to a preset time length and a sliding step size of half the window width. Within each window, the occurrence frequency of stress abrupt events is counted, and this frequency is divided by the window width to obtain the stress abrupt event frequency for that window, measured in Hertz. The computing environment merges consecutive windows with frequencies above a preset frequency threshold into a first operating segment, and consecutive windows with frequencies below the preset frequency threshold into a second operating segment. Each operating segment includes a start timestamp field and an end timestamp field.
[0159] Step S1130: For each first operation segment, extract the applied force adjustment sequence and displacement direction adjustment sequence from the instrument end force-position hybrid control command corresponding to all force change events in the first operation segment. Perform statistical analysis on the applied force adjustment sequence to calculate the average control intensity of the applied force adjustment in the first operation segment, and generate the applied force control intensity label of the first operation segment based on the average control intensity.
[0160] For each first operating segment, the computing environment collects the applied force adjustment amounts for all force-induced abrupt events within that segment, forming an applied force adjustment amount sequence. The arithmetic mean of this sequence is calculated; this value is the mean control intensity, measured in Newtons. Based on the preset intensity range within which the mean control intensity falls, the computing environment assigns a force control intensity label to each first operating segment. The preset intensity ranges include low-intensity, medium-intensity, and high-intensity ranges, corresponding to low-intensity, medium-intensity, and high-intensity labels, respectively.
[0161] Step S1140: Perform direction adjustment frequency statistical analysis on the displacement direction adjustment sequence, calculate the direction change frequency of the displacement direction adjustment in the first operation segment, and generate the direction control complexity label of the first operation segment based on the direction change frequency.
[0162] For each first operation segment, the computing environment collects the displacement direction adjustment amounts of all force-induced abrupt events within that segment, forming a displacement direction adjustment amount sequence. The direction change frequency is calculated by counting the number of times the angle between two adjacent displacement direction adjustment amounts in the sequence exceeds a preset direction change judgment threshold, and then dividing by the time length of that segment, with the dimension being Hertz. Based on the preset frequency range into which the direction change frequency falls, the computing environment assigns a direction control complexity label to the first operation segment. The preset frequency range includes a low-frequency range, a medium-frequency range, and a high-frequency range, with corresponding direction control complexity labels of low complexity, medium complexity, and high complexity, respectively.
[0163] Step S1150: Associate and label the force control intensity label and direction control complexity label of each first operation segment with the corresponding planned path area in the preoperative planned path point sequence to generate a planned path area annotation map carrying operation difficulty annotation information. Feed the planned path area annotation map back to the preoperative planning system to provide operation difficulty reference information for the preoperative path planning of subsequent similar invasive target medical operations.
[0164] For each first operational segment, the computational environment locates the planned path area covered by that segment within the preoperative planned path point sequence. This area is determined by the starting and ending planned path point numbers. The computational environment then writes force control intensity and direction control complexity labels into the metadata of this area in the planned path area annotation map. The planned path area annotation map is a JSON object containing an array of path point numbers and an array of operational difficulty labels. The computational environment feeds back the planned path area annotation map to the preoperative planning system via a data interface. The preoperative planning system overlays this annotation map onto the 3D anatomical model, using different colors or grayscale values to represent the operational difficulty level of different areas for subsequent surgical planning reference.
[0165] The method may further include: step S1210: after the instrument control intervention decision-making agent generates the instrument end force-position hybrid control command and drives the target medical instrument end to execute it, continuously monitor the instrument end displacement feedback signal flow, and extract the actual displacement response trajectory of the target medical instrument end during the execution of the instrument end force-position hybrid control command.
[0166] After the force-position hybrid control command is issued at the instrument's end, the computing environment collects displacement coordinate data from the displacement feedback signal stream at a preset trajectory tracking sampling period to form the actual displacement response trajectory. The actual displacement response trajectory is a three-dimensional coordinate point array, where each element is a structure containing horizontal, vertical, and depth coordinate values, all in millimeters.
[0167] Step S1220: Compare the actual displacement response trajectory with the expected displacement trajectory indicated by the displacement direction adjustment amount in the instrument end force-position hybrid control command, and calculate the trajectory tracking deviation vector between the actual displacement position and the expected displacement position at each sampling time. Perform deviation pattern analysis on the trajectory tracking deviation vector to determine whether the trajectory tracking deviation vector belongs to a systematic deviation pattern or a random deviation pattern. A systematic deviation pattern is when the trajectory tracking deviation vector maintains a consistent deviation direction at multiple consecutive sampling times, while a random deviation pattern is when the trajectory tracking deviation vector exhibits irregular directional jumps at multiple consecutive sampling times.
[0168] The expected displacement trajectory is calculated starting from the initial position of the instrument's end effector before the command is executed, with the direction indicated by the displacement travel direction adjustment as the direction of movement, and the step size as the operation step length. The trajectory tracking deviation vector is a three-dimensional vector obtained by subtracting the expected displacement position coordinates from the actual displacement position coordinates. The computational environment performs deviation pattern analysis on the trajectory tracking deviation vector sequence. This analysis involves calculating the angle sequence between adjacent vectors in the trajectory tracking deviation vector sequence. If the angle sequence is less than a preset angle consistency threshold for multiple consecutive sampling times, it is determined to be a systematic deviation pattern. If the angle sequence frequently shows jumps greater than the angle consistency threshold, it is determined to be a random deviation pattern.
[0169] Step S1230: When it is determined that the trajectory tracking deviation vector belongs to the systematic deviation mode, calculate the mean value of the deviation direction and the mean value of the deviation amplitude of the systematic deviation mode, and generate the zero-position compensation parameters of the displacement actuator based on the mean value of the deviation direction and the mean value of the deviation amplitude.
[0170] When a systematic deviation pattern is identified, the computational environment calculates the mean direction vector and mean magnitude of all vectors in the deviation vector sequence of the trajectory tracking. The mean direction vector is the normalized vector obtained by the arithmetic mean of the unit direction vectors of each deviation vector. The mean deviation magnitude is the arithmetic mean of the magnitudes of each deviation vector. The zero-position compensation parameter is a three-dimensional vector whose direction is opposite to the mean deviation direction vector, and its magnitude is equal to the mean deviation magnitude. The dimension of the zero-position compensation parameter is millimeters.
[0171] Step S1240: Write the zero-position compensation parameter into the control parameters of the target medical instrument displacement actuator, and perform offset correction on the zero-position reference of the displacement actuator.
[0172] The computing environment writes the zero-position compensation parameters into the zero-position reference offset register of the displacement actuator of the target medical instrument via the control bus. In subsequent displacement control, the displacement actuator adds the zero-position compensation parameters to the commanded displacement coordinates before executing position servo control, thereby correcting systematic displacement deviations.
[0173] Step S1250: When the trajectory tracking deviation vector is determined to be a random deviation mode, the deviation amplitude fluctuation range of the random deviation mode is calculated, and the output gain coefficient of the intervention strategy mapping network in the instrument control intervention decision agent for the displacement direction adjustment amount is adjusted according to the deviation amplitude fluctuation range. The adjusted output gain coefficient is then applied to the displacement direction adjustment amount generation process of subsequent force change events.
[0174] When the deviation is determined to be random, the difference between the maximum and minimum magnitudes of each vector in the trajectory tracking deviation vector sequence is calculated in the computational environment. This difference represents the deviation amplitude fluctuation range, measured in millimeters. The output gain coefficient is adjusted as follows: if the deviation amplitude fluctuation range exceeds a preset upper limit of fluctuation tolerance, the weight parameters of the three components corresponding to the displacement direction adjustment in the output layer of the intervention strategy mapping network are multiplied by a decay coefficient less than 1, reducing the output amplitude of the displacement adjustment. If the deviation amplitude fluctuation range is less than a preset lower limit of fluctuation tolerance, the corresponding weight parameters are multiplied by a gain coefficient greater than 1, increasing the output amplitude of the displacement adjustment. The adjusted output gain coefficient is stored in the output layer weight parameters of the intervention strategy mapping network and applied to the subsequent displacement direction adjustment generation process for sudden force events.
[0175] The method may further include: step S1310: during the invasive target medical operation performed by the target medical instrument, monitoring the time interval sequence between adjacent force change events in the force feedback signal stream at the end of the instrument.
[0176] When the computing environment detects each new force-induced abrupt change event in step S120, it calculates the time difference between the occurrence time of the event and the occurrence time of the previous force-induced abrupt change event, and appends this time difference to the time interval sequence. The time interval sequence is a one-dimensional floating-point array with the unit of seconds.
[0177] Step S1320: Perform time interval change trend analysis on the time interval sequence. When the time interval sequence is detected to show a gradually shortening convergence trend, determine that the current operating area is a region with dense force mutations. After determining that the current operating area is a region with dense force mutations, extract the precursor mode encoding vector generated by the precursor mode encoder for the most recently completed force mutation event and the aftereffect mode encoder generated by the aftereffect mode encoder for the force mutation event in the instrument control intervention decision agent.
[0178] The computing environment employs linear regression analysis on the time interval sequence to calculate the regression slope of the time interval sequence with respect to the sampling sequence number. If the regression slope is negative and its absolute value is greater than a preset convergence threshold, it is determined that the time interval sequence exhibits a gradually shortening convergence trend, and the current operating region is a region with dense stress mutations. The computing environment extracts the precursor mode encoding vector and aftereffect mode encoding vector corresponding to the most recently completed stress mutation event from the memory cache. Both vectors are the outputs of the precursor mode encoder and aftereffect mode encoder in step S140.
[0179] Step S1330: Perform vector concatenation processing on the precursor mode encoding vector and the aftereffect mode encoding vector to generate the regional stress mutation feature vector of the current operation area. Perform similarity matching processing on the regional stress mutation feature vector with the regional stress mutation feature vectors of various operation areas recorded in historical invasive target medical operations. Retrieve the historical operation area with the highest similarity to the regional stress mutation feature vector of the current operation area from the historical records.
[0180] The computational environment concatenates the precursor pattern encoding vector and the aftereffect pattern encoding vector along the feature dimension. The concatenated vector is the regional stress mutation feature vector. The computational environment reads the regional stress mutation feature vectors of each historical operational region from the historical operational region feature library and calculates the cosine similarity between the current regional stress mutation feature vector and each historical regional stress mutation feature vector. The cosine similarity is calculated by dividing the dot product of the two vectors by the product of their L2 norms. The historical operational region with the highest cosine similarity is selected as the matching result.
[0181] Step S1340: Extract the historical instrument end force-position hybrid control command sequence corresponding to the historical operation area, and statistically analyze the historical control mean sequence of the applied force adjustment amount and the historical control mean sequence of the displacement direction adjustment amount in the historical operation area. Based on the historical control mean sequence, perform bias adjustment processing on the output of the intervention strategy mapping network of the instrument control intervention decision agent in the current operation area, and bring the applied force adjustment amount generated by the intervention strategy mapping network closer to the historical control mean sequence. After the target medical instrument leaves the area of dense force change, cancel the bias adjustment processing and restore the original output mapping relationship of the intervention strategy mapping network.
[0182] The computational environment extracts the historical average sequences of applied force adjustment and displacement direction adjustment corresponding to the matched historical operating regions from historical data. The bias adjustment is performed by superimposing a bias adjustment value onto the output layer bias parameters of the intervention strategy mapping network. The applied force component of the bias adjustment value is the difference between the historical control mean and the current network output applied force adjustment value multiplied by a preset convergence coefficient. The displacement component is calculated similarly. When the regression slope of the time interval sequence returns to a positive value or its absolute value is less than the convergence threshold, it is determined that the target medical instrument has left the region of dense force mutation, the bias adjustment value is revoked, and the output layer bias parameters are restored to their original values.
[0183] The method may further include: step S1410: during the process of the instrument control intervention decision-making agent generating instrument end force-position hybrid control command for multiple consecutive force mutation events, recording the pre-force mutation waveform segment and post-force mutation waveform segment corresponding to each force mutation event.
[0184] During real-time operation, the computing environment maintains a queue of waveform records for force abrupt events. Each element in the queue is a structure containing a force abrupt event identifier field, a precursor waveform segment field, and a follow-up waveform segment field. The data structure for the precursor waveform segment is a one-dimensional floating-point array, where each element represents an axial force component value in Newtons. The storage format for the follow-up waveform segment is the same as that for the precursor waveform segment.
[0185] Step S1420: Extract the waveform segments of the aftereffect of the force mutation in the previous force mutation event and the waveform segments of the precursor of the force mutation in the next force mutation event from two adjacent force mutation events. Perform waveform splicing and connection processing on the waveform segments of the aftereffect of the force mutation in the previous force mutation event and the waveform segments of the precursor of the force mutation in the next force mutation event to construct a continuous force fluctuation evolution trajectory across force mutation events.
[0186] The computational environment retrieves records of two adjacent abrupt force events from the waveform recording queue. It extracts the waveform segments representing the aftereffects of the abrupt force event and the waveform segments representing the precursors of the abrupt force event. The two arrays are then concatenated chronologically, with their time axes aligned at the same sampling interval, and the start time of the latter array immediately following the end time of the former. The concatenated array represents the continuous force fluctuation evolution trajectory across the abrupt force events. This trajectory is a one-dimensional floating-point array, with a length equal to the sum of the lengths of the two original arrays. The array elements represent axial force components in Newtons.
[0187] Step S1430: Perform force wave propagation pattern recognition processing on the continuous force wave evolution trajectory. By analyzing the attenuation coefficient and phase offset of the force amplitude in the continuous force wave evolution trajectory as it propagates from the aftereffect stage of a single force abrupt event to the precursor stage of the next force abrupt event, determine the spatial propagation characteristics of the force disturbance when the target medical instrument terminal advances in the tissue medium.
[0188] The computational environment performs force wave propagation pattern recognition processing on the continuous force wave evolution trajectory. The peak force amplitude in the waveform segment following a sudden force change in the previous event is denoted as A1, and the peak force amplitude in the waveform segment preceding a sudden force change in the next event is denoted as A2. The attenuation coefficient is calculated as the quotient of A2 divided by A1, and is dimensionless. The phase offset is calculated as the time difference between the peak force moment of the waveform segment preceding a sudden force change in the next event and the peak force moment of the waveform segment following a sudden force change in the previous event, in seconds. The attenuation coefficient and phase offset are used as spatial propagation characteristic parameters of the force disturbance as the target medical instrument tip propagates through the current tissue medium. The spatial propagation characteristic parameters are a structure containing attenuation coefficient and phase offset values.
[0189] Step S1440: Construct a tissue medium mechanical response prediction model based on the spatial propagation characteristics of the force disturbance. The tissue medium mechanical response prediction model is used to predict the occurrence time and waveform shape of the precursor waveform segment of the force disturbance of the future force disturbance event based on the aftereffect waveform segment of the current force disturbance event.
[0190] The tissue media mechanical response prediction model is a sequence prediction model based on a recurrent neural network. This model receives a waveform segment of the post-stress abrupt change in the current stress event as the input sequence. This input sequence is encoded by a long short-term memory (LSM) network layer, with the hidden state dimension of the LSM layer pre-defined as Hpred. The hidden state vector of the encoder at the last time step serves as the encoded representation of the tissue media mechanical state. This encoded representation, concatenated with spatial propagation feature parameters, is input to a decoder. The decoder also employs an LSM network structure to autoregressively generate a predicted sequence of waveform segments preceding future stress abrupt changes. The predicted sequence includes an array of predicted waveforms and the predicted occurrence time. The tissue media mechanical response prediction model is pre-trained on historical operational data, using pairs of event segments preceding and following historical continuous stress fluctuation trajectories as training sample pairs.
[0191] Step S1450: Connect the tissue medium mechanical response prediction model to the precursor mode encoder of the instrument control intervention decision-making agent. When the deviation between the time of occurrence of the future force mutation event predicted by the tissue medium mechanical response prediction model and the time of occurrence of the force mutation event actually detected in the force feedback signal stream at the end of the instrument exceeds the preset prediction deviation tolerance range, it is determined that the mechanical properties of the tissue medium currently in contact have changed.
[0192] The computational environment continuously runs the tissue media mechanical response prediction model in real-time. Whenever a new stress abrupt change event is detected in step S120, a waveform segment of the aftereffect of the stress abrupt change event from the previous event is input into the tissue media mechanical response prediction model, and the model outputs a predicted value for the time of occurrence of the new event. The absolute value of the time difference between the predicted and actual occurrence times is calculated. If this absolute value exceeds a preset prediction deviation tolerance range, it is determined that the mechanical properties of the currently contacted tissue media have changed. The unit of measurement for the prediction deviation tolerance range is seconds.
[0193] Step S1460: When it is determined that the mechanical properties of the tissue medium have changed, the instrument control intervention decision agent is triggered to adaptively adjust the internal connection weights of the precursor pattern encoder and the intervention strategy mapping network. The adaptive adjustment aims to reduce the tolerance range of prediction bias.
[0194] When an anomaly in the mechanical properties of the tissue medium is detected, the computational environment triggers an online adaptive adjustment process for the instrument control intervention decision-making agent. This online adaptive adjustment employs a few-round fine-tuning strategy, using training sample pairs corresponding to several recently completed force mutation events as the fine-tuning dataset. Gradient descent is used to update all internal connection weights of the precursor pattern encoder, aftereffect pattern encoder, and intervention strategy mapping network. The optimizer used for gradient descent updating is an adaptive moment estimation optimizer, with a learning rate set to a fine-tuning learning rate value lower than the initial learning rate in step S270. Fine-tuning training is performed for a preset number of rounds, with the training objective being to minimize the mean squared error loss between the trained predicted force-potential hybrid control command and the actual control command applied by the operator. After adaptive adjustment is complete, the prediction parameters of the tissue medium mechanical response prediction model are updated. At this point, the instrument control intervention decision-making agent has adapted to the new tissue medium mechanical properties.
[0195] The method may further include: step S1510: during the execution of an invasive target medical operation, acquiring the therapeutic energy output power sequence applied to the tissue by the energy output module loaded at the end of the target medical instrument, wherein the therapeutic energy output power sequence and the force feedback signal flow at the end of the instrument are acquired synchronously in time.
[0196] The computing environment acquires the therapeutic energy output power sequence at the same sampling frequency as the force sensor through the power monitoring interface of the target medical instrument's energy output module. The energy output module includes therapeutic energy application devices such as radiofrequency ablation electrodes or laser fibers. The data structure of the therapeutic energy output power sequence is a one-dimensional floating-point array, where each element is the energy output power value at that sampling moment, measured in watts. Each power sampling data frame contains a timestamp field and a power value field; the timestamp field uses the same clock source as the force sensor.
[0197] Step S1520: Perform cross-modal correlation analysis on the treatment energy output power sequence and the force feedback signal stream at the instrument end to identify the correlation pattern of the influence of changes in energy output power in the treatment energy output power sequence on the frequency and amplitude of force mutation events in the force feedback signal stream at the instrument end.
[0198] The computational environment aligns the therapeutic energy output power sequence and the force feedback signal stream at the instrument's end by timestamps, constructing an aligned data table of power and force. The computational environment performs correlation analysis using the therapeutic energy output power as the independent variable and the statistical characteristics of force mutation events as the dependent variable. The statistical characteristics include the frequency of occurrence, calculated as the number of force mutation events occurring within a preset time window (in Hertz), and the mean amplitude of force mutations, calculated as the arithmetic mean of the peak force values of the force mutation events within that window (in Newtons). A cross-correlation function is used to calculate the changing trends of the frequency and mean amplitude of force mutation events under different power value ranges, identifying the moderating effect of changes in energy output power on the statistical characteristics of force mutation events. The influencing correlation pattern is a set of correlation rules, each rule containing the power range, the direction of frequency change, and the direction of amplitude change.
[0199] Step S1530: Determine the energy output power adjustment window in the therapeutic energy output power sequence that can suppress the occurrence of stress mutation events based on the influence correlation pattern. The energy output power adjustment window defines the timing and magnitude of energy output power adjustment.
[0200] Based on the influence correlation patterns identified in step S1520, the computing environment filters out power value ranges that can reduce the frequency of stress mutation events or the average amplitude of stress mutations. The adjustment timing of the energy output power adjustment window is defined as a Boolean condition, triggered when the frequency of stress mutation events exceeds a preset frequency upper limit threshold or the average amplitude of stress mutations exceeds a preset amplitude upper limit threshold. The adjustment amplitude is defined as the difference between the target power value and the current power value, where the target power value is the power value within the aforementioned power value range that optimizes the reduction in frequency or amplitude. The energy output power adjustment window is a structure containing a trigger condition field and a target power value field.
[0201] Step S1540: When the instrument control intervention decision-making agent generates the instrument end force-position hybrid control command, it simultaneously queries the energy output power adjustment window. If the current operation step falls within the energy output power adjustment window, it generates a treatment energy output power accompanying adjustment command. The treatment energy output power accompanying adjustment command includes the energy output power target value that is executed in conjunction with the instrument end force-position hybrid control command.
[0202] After generating the instrument end-effector force-position hybrid control command in step S140, the computing environment queries whether the trigger conditions for the energy output power adjustment window are met based on the frequency and average amplitude of the current force abrupt change events. If the trigger conditions are met, the corresponding target power value is extracted, and a treatment energy output power accompanying adjustment command is generated. The data structure of the treatment energy output power accompanying adjustment command is a JSON object containing a target power value field, with the target power value measured in watts.
[0203] Step S1550: The treatment energy output power adjustment command and the instrument end force-position hybrid control command are time-aligned and packaged to generate a force-position energy joint control command package, and the force-position energy joint control command package is sent to the force-position actuator and energy output module of the target medical instrument for synchronous execution.
[0204] The computing environment aligns the timestamp of the treatment energy output power adjustment command with the timestamp of the instrument end-effector force-position hybrid control command, using the later timestamp of the two commands as the joint execution timestamp. The computing environment packages the four components of the instrument end-effector force-position hybrid control command and the target power value of the treatment energy output power adjustment command into a force-position energy joint control command package. This force-position energy joint control command package is a JSON object containing the execution timestamp field, the applied force adjustment amount field, the horizontal adjustment amount field for the displacement direction field, the vertical adjustment amount field for the displacement direction field, the depth adjustment amount field for the displacement direction field, and the treatment energy output power target value field. The computing environment synchronously sends this command package to the force-position actuator controller and the energy output module controller via the control bus. The force-position actuator controller and the energy output module controller parse the corresponding fields and execute the commands.
[0205] In one embodiment, Figure 2 This is a schematic diagram of the interface of the mobile terminal monitoring application provided in an embodiment of the present invention. Figure 2 The left image in the image corresponds to the real-time monitoring dashboard interface. Figure 2 The right image in the image corresponds to the details interface of the sudden force event.
[0206] During the execution of the agent-based medical instrument control method described in steps S110 to S140, the computing environment pushes the force feedback signal stream at the instrument end, the displacement feedback signal stream at the instrument end, the result of force change event detection and processing, and the generation status of the force-position hybrid control command at the instrument end to the mobile terminal in real time. Figure 2The top of the real-time monitoring dashboard interface shown in the left figure displays the connection status of the target medical instrument. The middle of the interface simultaneously displays the force feedback signal flow and displacement feedback signal flow at the instrument's end in the form of dual-waveform real-time monitoring cards. The two waveform curves are strictly aligned in the time dimension, and the synchronous acquisition relationship in the time dimension described in step S110 is indicated by a dashed line between the waveforms. When a force mutation event is detected in step S120, a force mutation event warning card pops up below the waveform card. The warning card displays the time of occurrence of the force mutation event and shows the precursor waveform segment of the force mutation extracted in step S123 and the aftereffect waveform segment extracted in step S124 in a scaled-down waveform format. Below the warning card, a parameter value panel is displayed in a layout of displacement response pre-segment and displacement response post-segment, respectively presenting the end displacement retraction parameter and end displacement offset direction parameter extracted in steps S132 to S133, and the end displacement recovery parameter and end displacement recovery direction parameter extracted in steps S134 to S135. The bottom of the interface displays the adjustment amount of the applied force generated in step S140 in the form of a force indicator bar, and the adjustment amount of the displacement direction in the form of a direction compass.
[0207] Operator click Figure 2 After the force change event warning card shown in the left image, the mobile application redirects to... Figure 2 The right-hand image shows the details interface for the force mutation event. The top of this interface displays the event number and occurrence time of the force mutation event, and labels indicating whether the event belongs to the tissue-crossing or non-tissue-crossing force mutation event marked in step S630. The middle of the interface displays the complete waveform of the instrument terminal force feedback signal stream obtained from steps S121 to S124, from the start of the force mutation precursor to the end of the force mutation effect. Vertical lines mark the occurrence times as forced segmentation boundary points on the waveform, and different background colors distinguish the waveform segments of the force mutation precursor and the force mutation effect. The force change rate sequence curve and the marked positions of the sign reversal point set are also overlaid. Below the waveform, a table displays all displacement response parameters extracted from steps S132 to S135. The bottom of the interface displays the values of the applied force adjustment and displacement direction adjustment generated in step S140 in a circular dashboard format.
[0208] In one embodiment, Figure 3 This is another schematic diagram of the interface of the mobile terminal monitoring application provided in an embodiment of the present invention. Figure 3 The left image in the image corresponds to the regulatory effectiveness monitoring interface. Figure 3 The right image in the diagram corresponds to the path correction monitoring interface.
[0209] During the execution of the online adjustment and correction process described in steps S310 to S360, the mobile terminal application switches to... Figure 3 The left image shows the control effectiveness monitoring interface. The top of the interface displays a status indicator indicating that the current instrument end-effector force-position hybrid control command generated by the instrument control intervention decision-making agent in step S320 has been executed. The middle of the interface displays the remaining duration of the preset observation window described in step S330 using a circular countdown progress component, and continuously displays the instrument end-effector force feedback signal flow using a real-time waveform monitoring area. When a new force mutation event is detected within the preset observation window in step S330, the location of the new force mutation event is marked with a flashing marker on the waveform. The bottom of the interface displays the visualization results of the waveform similarity comparison processing described in step S340. The waveform segment of the force mutation effect after-effect extracted in step S124 is compared with the waveform segment of the new force mutation effect after-effect extracted in step S340 in the form of an overlay waveform diagram, and the percentage value of the waveform similarity is displayed below the waveform diagram. When step S350 determines that the waveform similarity is lower than the preset lower limit of similarity and the control effect does not meet expectations, the similarity value is highlighted in a warning color, and a warning text indicating that the control effect does not meet expectations is displayed on the interface. At the same time, the operation button for generating the force-position hybrid control command at the end of the correction instrument as described in step S360 is displayed.
[0210] During the execution of the path correction and safety boundary constraint process described in steps S510 to S570 and S710 to S760, the mobile terminal application switches to Figure 3The path correction monitoring interface is shown in the right figure. The main body of the interface is a three-dimensional view area. Within the view, the movable spatial boundary model described in step S710 is displayed in a semi-transparent three-dimensional outline. This movable spatial boundary model is obtained from preoperative medical image reconstruction and defines the three-dimensional spatial range within which the end effector of the target medical instrument is allowed to move. Inside the model, the sequence of preoperative planned path points obtained in step S510 is displayed in the form of curves, and the three-dimensional spatial coordinates of each planned path point are marked on the path. The real-time three-dimensional spatial position coordinates of the instrument end are displayed as a highlighted sphere along the path. A red directional arrow leading from the sphere indicates the actual displacement direction of the instrument end, and a blue dashed arrow indicates the planned path direction. The spatial angle deviation calculated in step S520 is marked between the two arrows. The remaining distance deviation calculated in step S530 is marked between the sphere and the next planned path point with a dashed line. The bottom of the interface displays the arbitration results of the path correction priority arbitration logic in step S540, showing the priority control target identifier in the form of a pop-up card, and listing the suggested adjustment values of the direction correction control component generated in step S550 and the velocity correction control component generated in step S560. The bottom of the interface displays the final parameters of the instrument end-effector force-position hybrid control command generated in step S570 after path correction. Simultaneously, the interface displays the judgment results of the spatial inclusion relationship determination process in step S730 in a status indicator bar. When the expected arrival position coordinates of the instrument end-effector exceed the three-dimensional space range defined by the movable space boundary model, the projected safe expected arrival position coordinates described in step S740 and the corrected displacement travel direction adjustment amount calculated in step S750 are displayed in a warning color.
[0211] In an exemplary embodiment, a server is provided, which may be a terminal, server, etc., and its internal structure includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, near-field communication, or other technologies. When the computer program is executed by the processor, it implements a medical instrument control method based on intelligent agents. The display unit is used to form a visually visible image and may be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device can be a touch layer covering the display screen, or a button, trackball, or touchpad set on the server casing, or an external keyboard, touchpad, or mouse, etc.
[0212] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.
Claims
1. A medical instrument control method based on intelligent agents, characterized in that, The method includes: Acquire the force feedback signal stream and displacement feedback signal stream at the end of the instrument generated by the target medical instrument during the invasive target medical procedure; Force abrupt event detection processing is performed on the force feedback signal stream at the end of the instrument to determine the occurrence time of the force abrupt event and extract the precursor waveform segment and the aftereffect waveform segment corresponding to each force abrupt event. The occurrence time of each force abrupt event is used as the forced segmentation boundary point for displacement response segmentation processing of the displacement feedback signal stream at the end of the instrument. For each sudden force event, the displacement feedback signal stream at the end of the instrument is divided into a displacement response front segment and a displacement response back segment, with the occurrence time of the sudden force event as the boundary. The end displacement retraction parameter and end displacement offset direction parameter of the displacement response front segment are extracted, and the end displacement recovery parameter and end displacement recovery direction parameter of the displacement response back segment are extracted. The precursor waveform segment of each force mutation event, the end displacement retraction parameter, and the end displacement offset direction parameter are used together as the precursor perception input of the instrument control intervention decision-making agent. The precursor pattern encoder in the instrument control intervention decision-making agent performs precursor pattern encoding processing on the precursor perception input to generate a precursor pattern encoding vector. The intervention strategy mapping network in the instrument control intervention decision-making agent uses the precursor pattern encoding vector and the aftereffect pattern encoding vector corresponding to the previous force mutation event as the joint decision basis to perform intervention strategy mapping processing to generate the instrument end force-position hybrid control command corresponding to the current force mutation event.
2. The medical instrument control method based on intelligent agents according to claim 1, characterized in that, The process of detecting and processing the force feedback signal stream at the instrument's end to determine the occurrence time of the force abrupt change event and extracting the precursor waveform segment and the aftereffect waveform segment corresponding to each force abrupt change event includes: The force feedback signal stream at the end of the instrument is subjected to sliding window differential processing to calculate the force change rate sequence between adjacent sampling points. The force change rate sequence is subjected to sign flip detection processing to identify the set of sign flip points in the force change rate sequence where the positive change rate suddenly changes to the negative change rate or vice versa. From the set of symbol flip points, select symbol flip points that meet the preset flip amplitude threshold condition and flip speed threshold condition as candidate occurrence times of force mutation events. For each candidate occurrence time, search in reverse along the time axis for the moment when the force change rate in the force feedback signal stream at the end of the instrument first enters the stable fluctuation range as the starting time of the force mutation precursor of the force mutation event. The moment when the rate of change of force in the force feedback signal stream at the end of the instrument re-enters the stable fluctuation range is searched along the time axis in the positive direction is taken as the end time of the aftereffect of the force mutation event. The segment of the force feedback signal stream at the end of the instrument between the start time of the force mutation precursor and the candidate occurrence time is extracted as the waveform segment of the force mutation precursor. Extract the instrument end force feedback signal stream segment between the candidate occurrence time and the end time of the force mutation effect as the force mutation effect waveform segment.
3. The medical instrument control method based on intelligent agents according to claim 1, characterized in that, For each sudden force event, the instrument's end-effector displacement feedback signal stream is divided into a pre-displacement response segment and a post-displacement response segment, with the occurrence time of the sudden force event as the boundary. The end-effector displacement retraction parameter and end-effector displacement offset direction parameter of the pre-displacement response segment are extracted, and the end-effector displacement recovery parameter and end-effector displacement recovery direction parameter of the post-displacement response segment are extracted, including: Using the occurrence time of the force abrupt change event as the dividing boundary, the displacement feedback signal stream at the end of the instrument is divided into the displacement response front segment located before the occurrence time and the displacement response back segment located after the occurrence time; The extreme point search process is performed on the displacement components of the instrument end in the three-dimensional coordinate system in the displacement response phase to determine the maximum offset of the instrument end in each spatial dimension from the original travel path, and the largest modulus among the maximum offsets in each spatial dimension is taken as the end displacement retraction parameter. Extract the cosine components of the displacement direction of the instrument tip in each spatial dimension at the time corresponding to the tip displacement retraction parameter, and construct the tip displacement offset direction parameter. The displacement recovery process of the instrument end in the latter part of the displacement response is subjected to trend tracking processing, and the total displacement regression of the instrument end from the maximum offset position to the original travel path in the latter part of the displacement response is determined as the end displacement recovery amount parameter. The displacement regression direction cosine components of the instrument end in each spatial dimension during the displacement recovery process are extracted to form the end displacement recovery direction parameters.
4. The medical instrument control method based on intelligent agents according to claim 1, characterized in that, The method further includes: Collect historical instrument end force feedback signal streams and historical instrument end displacement feedback signal streams recorded in historical invasive target medical operations, and perform the force mutation event detection processing on the historical instrument end force feedback signal streams to obtain a set of historical force mutation events and their corresponding historical force mutation precursor waveform segments and historical force mutation aftereffect waveform segments. Based on the occurrence time of each historical force abrupt change event in the set of historical force abrupt change events, the displacement response segmentation processing is performed on the historical instrument end displacement feedback signal stream to obtain historical end displacement retraction parameters, historical end displacement offset direction parameters, historical end displacement recovery parameters, and historical end displacement recovery direction parameters; For each historical stress change event, the waveform segment of the precursor of the historical stress change, the parameter of the historical end displacement recoil, and the parameter of the historical end displacement offset direction are combined into training precursor input samples. For the same historical stress change event, the waveform segment of the aftereffect of the historical stress change, the parameter of the historical end displacement recovery, and the parameter of the historical end displacement recovery direction are combined into training aftereffect input samples. The instrument end force-position hybrid control command actually applied by the operator when the corresponding historical force change event occurs is used as the training target command; The training precursor input samples are input into the precursor pattern encoder of the instrument control intervention decision-making agent to generate a training precursor pattern encoding vector, and the training after-effect input samples are input into the after-effect pattern encoder of the instrument control intervention decision-making agent to generate a training after-effect pattern encoding vector. The training precursor pattern encoding vector and the training aftereffect pattern encoding vector corresponding to the previous historical force mutation event are input into the intervention strategy mapping network of the instrument control intervention decision-making agent to generate training prediction force-position hybrid control instructions. Calculate the instruction deviation between the training prediction force-position hybrid control instruction and the training target instruction, and adjust the internal connection weights of the precursor mode encoder, the aftereffect mode encoder and the intervention strategy mapping network with the goal of reducing the instruction deviation, until the instruction deviation converges to a preset allowable deviation range.
5. The medical instrument control method based on intelligent agents according to claim 1, characterized in that, The method further includes: After the instrument control intervention decision-making agent is deployed to the target medical instrument, it collects in real time the real-time force mutation precursor waveform segment, real-time end displacement retraction parameter, and real-time end displacement offset direction parameter corresponding to the latest force mutation event in the current invasive target medical operation. The real-time stress change precursor waveform segment, real-time end displacement retraction parameter, and real-time end displacement offset direction parameter are input into the instrument control intervention decision-making intelligent body, which then generates the instrument end force-position hybrid control command for the current stress change event. Within a preset observation window after executing the force-position hybrid control command at the instrument end, continuously monitor whether a new force mutation event that meets the force mutation event detection and processing conditions reappears in the force feedback signal stream at the instrument end. If a new stress mutation event is detected within the preset observation window, the waveform segment of the new stress mutation effect corresponding to the new stress mutation event is extracted, and the waveform segment of the new stress mutation effect is compared with the waveform segment of the stress mutation effect of the current stress mutation event for waveform similarity comparison. When the waveform similarity comparison processing result indicates that the waveform similarity between the new stress mutation effect waveform segment and the stress mutation effect waveform segment of the current stress mutation event is lower than the preset similarity lower limit, it is determined that the control effect of the instrument end force-position hybrid control command on the current stress mutation event has not reached the expected level. Based on the waveform segment of the precursor of the new force change event and the end displacement parameters, a new modified instrument end force-position hybrid control command is generated, and the original instrument end force-position hybrid control command is replaced with the modified instrument end force-position hybrid control command for force-position hybrid control.
6. The medical instrument control method based on intelligent agents according to claim 1, characterized in that, The method further includes: During the same type of invasive target medical operation performed by multiple medical instruments, the set of historical force mutation events uploaded by each medical instrument and the historical instrument end force-position hybrid control command corresponding to each historical force mutation event are collected. The waveform envelope features and waveform spectral distribution features of the force change rate sequence in the historical force change precursor waveform segment of each historical force change event are extracted as the force change precursor pattern fingerprint; Based on the stress mutation precursor pattern fingerprint, the set of historical stress mutation events is clustered to obtain stress mutation event clusters with similar stress mutation precursor pattern fingerprints. For each cluster of force-induced abrupt events, the historical instrument end force-position hybrid control commands corresponding to each historical force-induced abrupt event are statistically processed to determine the distribution range and concentration trend of control commands corresponding to each cluster of force-induced abrupt events. A mapping table is constructed between the fingerprint of the precursor pattern of force mutation and the central trend of the control command set, and the mapping table is loaded into the initial mapping layer of the intervention strategy mapping network of the instrument control intervention decision agent.
7. The medical instrument control method based on intelligent agents according to claim 1, characterized in that, The method further includes: The preoperative planning path point sequence generated by the preoperative planning system before the invasive target medical instrument performs the invasive target medical operation is obtained. The preoperative planning path point sequence contains the three-dimensional spatial coordinates of multiple planning path points arranged in the order of operation. After the instrument control intervention decision-making agent generates a hybrid control command for the instrument end force and position for each force mutation event, the displacement direction adjustment amount carried in the hybrid control command is extracted. The actual displacement direction of the instrument end indicated by the displacement direction adjustment amount is compared with the planned path direction corresponding to the current operation step in the preoperative planned path point sequence to calculate the spatial angle deviation between the actual displacement direction of the instrument end and the planned path direction. The real-time three-dimensional spatial coordinates of the target medical instrument tip during the invasive target medical operation are obtained, and the remaining distance deviation between the real-time three-dimensional spatial coordinates and the next planned path point in the preoperative planned path point sequence is calculated. The spatial angle deviation and the remaining distance deviation are input into the path correction priority arbitration logic. The path correction priority arbitration logic determines whether the priority control target of the current operation step is to reduce the spatial angle deviation or the remaining distance deviation, and outputs the priority control target identifier. When the priority control target identifier indicates that the spatial angle deviation should be reduced first, a directional correction control component is generated. The directional correction control component is used to correct the displacement direction adjustment in the instrument end force-position hybrid control command. When the priority control target identifier indicates that the remaining distance deviation should be reduced first, a speed correction control component is generated. The speed correction control component is used to correct the applied force adjustment in the instrument end force-position hybrid control command to change the propulsion speed of the instrument end. The direction correction control component or the velocity correction control component is superimposed on the instrument end force-position hybrid control command to obtain the instrument end force-position hybrid control command after path correction.
8. The medical instrument control method based on intelligent agents according to claim 1, characterized in that, The method further includes: During the invasive target medical operation performed by the target medical instrument, the tissue impedance feedback signal flow of the biological tissue contacted by the end of the target medical instrument is collected in real time, and the tissue impedance feedback signal flow and the force feedback signal flow at the end of the instrument are time-synchronized. The tissue impedance feedback signal stream is subjected to impedance layer identification processing, and the set of tissue layer boundary moments is determined by detecting the moments when the impedance value in the tissue impedance feedback signal stream undergoes a step change. Each tissue stratification boundary moment in the set of tissue stratification boundary moments is matched with the occurrence time of the stress mutation event. The stress mutation events that coincide with each tissue stratification boundary moment in the set of tissue stratification boundary moments are identified and marked as tissue-crossing stress mutation events. The stress mutation events that do not coincide with any tissue stratification boundary moment are marked as non-tissue-crossing stress mutation events. For the tissue-crossing force mutation event, the impedance change amplitude and direction in the tissue impedance feedback signal stream before and after the occurrence of the tissue-crossing force mutation event are extracted to generate tissue type change features. The tissue type change features are input into the intervention strategy mapping network of the instrument control intervention decision agent. The instrument end force-position hybrid control command generated by the intervention strategy mapping network is subjected to tissue type adaptive weighting correction processing to obtain the instrument end force-position hybrid control command based on tissue recognition. For the aforementioned non-organizational traversal-type force abrupt change event, the instrument end force-position hybrid control command generated by the intervention strategy mapping network remains unchanged.
9. A server, characterized in that, include: processor; A machine-readable storage medium for storing machine-executable instructions of the processor; The processor is configured to execute the agent-based medical instrument control method of any one of claims 1 to 8 by executing the machine-executable instructions.
10. A computer storage medium, characterized in that, The computer storage medium stores machine-executable instructions, the server's processor reads the machine-executable instructions from the computer-readable storage medium, and the processor executes the machine-executable instructions, causing the server to perform the agent-based medical instrument control method as described in any one of claims 1 to 8.