Force feedback control method, apparatus, device, medium and product

By integrating force feedback devices and force data prediction models into the surgical robot system, and predicting interactive forces based on visual images, the problem of operators relying on experience to judge force is solved. This achieves precise force feedback control without hardware modification, improving surgical safety and accuracy.

CN122297122APending Publication Date: 2026-06-30AGIBOT MEDTECH (SUZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610392769.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-27
Publication Date
2026-06-30

Smart Images

  • Figure CN122297122A_ABST
    Figure CN122297122A_ABST
Patent Text Reader

Abstract

This application discloses a force feedback control method, device, equipment, medium, and product. The method includes: acquiring an interactive image reflecting the interaction state between a target instrument and target tissue; inputting the interactive image into a force data prediction model and outputting predicted interactive force data; wherein the predicted interactive force data includes at least one interactive force vector; generating a force feedback control command based on the predicted interactive force data; and sending the force feedback control command to a force feedback device to drive the force feedback device to output a force on the main control arm that matches the interactive force vector. This improves the safety and reliability of tactile feedback, thereby avoiding misleading user operations and ensuring the continuity and safety of the surgical procedure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of surgical robot technology, and in particular to a force feedback control method, device, equipment, medium and product. Background Technology

[0002] Currently, surgical robots generally adopt a master-slave control architecture, mainly consisting of a doctor's console and a patient-side surgical cart. The doctor's console is equipped with a master control arm and a monitor, allowing the operator to control and observe the surgical field in real time; the patient-side surgical cart carries detachable surgical instruments and an endoscope for acquiring images of the surgical field, with the images acquired by the endoscope being transmitted to the monitor on the console in real time.

[0003] During surgery, the operator relies primarily on visual information from the monitor screen combined with personal experience to judge the force applied. However, this method is highly dependent on subjective judgment and is prone to causing tissue damage or operational errors due to improper force control. Summary of the Invention

[0004] This application provides a force feedback control method, device, equipment, medium, and product to improve the safety and reliability of tactile feedback, thereby avoiding misleading the operator and ensuring the continuity and safety of the surgical procedure.

[0005] According to one aspect of this application, a force feedback control method is provided, applied to a master-slave surgical robot system, the surgical robot system including a master control arm and at least one slave robotic arm controlled by the master control arm, the master control arm integrating a force feedback device, and the slave robotic arm carrying a target instrument, the method comprising: Acquire interactive images that reflect the interaction between the target instrument and the target tissue; The interactive image is input into the force data prediction model, and the predicted interactive force data is output; wherein, the predicted interactive force data includes at least one interactive force vector; Based on the predicted interactive force data, a force feedback control command is generated and sent to the force feedback device to drive the force feedback device to output a force on the main control arm that matches the interactive force vector.

[0006] Optionally, force feedback control commands for maintaining or attenuating the force are generated, including: Using the historical interaction force data output at the previous moment as the initial value, a force feedback control command that decays to a preset threshold is generated within a preset time period; or, Using the historical interaction force data output at the previous moment as a constant value, a force feedback control command is generated to maintain the constant force. The "previous time" refers to the most recent time before the current time when the preset reliability condition was met.

[0007] Optionally, multiple training samples can be obtained, including: Acquire multiple sets of interaction sample images of the surgical instruments interacting with tissues, as well as the raw interaction force data of the surgical instrument ends; Correction and coordinate transformation are performed on the original interaction force data to obtain interaction force sample data in the end coordinate system of the surgical instrument; Each group of interactive sample images is paired with the interactive force sample data to form multiple training samples.

[0008] Optionally, pairing each group of interactive sample images with the interactive force sample data to form multiple training samples includes: For each group of interactive sample images, based on the shooting time and preset time window of the interactive sample images in the current group, the force selection range is determined, and the interactive force sample data whose acquisition time is within the force selection range is paired with the interactive sample images; The training samples are determined based on each of the interaction sample images in the current group and its paired interaction sample images.

[0009] Optionally, the target loss function includes a data fitting loss function and a physical constraint function. The step of determining the total loss value based on the target loss function, according to the interaction force sample data in the training samples and the estimated interaction force data, includes: Based on the data fitting loss function, error analysis is performed on the interaction force sample data in the training samples and the estimated interaction force data to obtain the data fitting loss value. Based on the physical constraint function, a physical rationality analysis is performed on the estimated interaction force data to obtain the physical constraint loss value. The total loss value is determined based on the data fitting loss value and the physical constraint loss value.

[0010] According to another aspect of this application, a force feedback control device is provided, configured in a master-slave surgical robot system, the surgical robot system including a master control arm and at least one slave robotic arm controlled by the master control arm, the master control arm integrating a force feedback device, and the slave robotic arm carrying a target instrument; the device includes: The interactive image acquisition module is used to acquire interactive images that reflect the interaction state between the target instrument and the target tissue. The predictive interaction force data determination module is used to input the interaction image into the force data prediction model and output the predicted interaction force data; wherein the predicted interaction force data includes at least one interaction force vector. The instruction output module is used to generate a force feedback control instruction based on the predicted interactive force data, and send the force feedback control instruction to the force feedback device to drive the force feedback device to output a force on the main control arm that matches the interactive force vector.

[0011] According to another aspect of this application, an electronic device is provided, the electronic device comprising: At least one processor; and a memory communicatively connected to said at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the force feedback control method described in any embodiment of this application.

[0012] According to another aspect of this application, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the force feedback control method described in any embodiment of this application.

[0013] According to another aspect of this application, a computer program product is provided, including a computer program that, when executed by a processor, implements the force feedback control method as described in any embodiment of this application.

[0014] The technical solution of this application embodiment predicts the force data of interactive images based on a force data prediction model, outputting predicted interactive force data including at least one interactive force vector. Based on the predicted interactive force data, a force feedback control command is generated and sent to a force feedback device, driving the force feedback device to output a force on the main control arm that matches the interactive force vector. This solves the problem in the prior art that the interaction force can only be judged by the operator's personal experience, which can easily lead to tissue damage or operational errors due to improper force control. Without changing the hardware structure of the surgical instruments (e.g., without adding force sensors), the force feedback control command is generated by calling the force data prediction model, thereby driving the force feedback device to output a force on the main control arm. This ensures that the force feedback device can accurately provide feedback to the operator on the force currently applied to the instrument, realizing tactile feedback to the operator, thereby reducing the risk of operator misoperation and ensuring the continuity and safety of master-slave operation of the instrument.

[0015] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the doctor's console structure shown in an embodiment of this application; Figure 2 This is a schematic diagram of the patient trolley structure shown in an embodiment of this application; Figure 3 This is a schematic diagram of any of the robotic arm structures shown in the embodiments of this application; Figure 4 This is a flowchart of a force feedback control method provided according to an embodiment of this application; Figure 5 This is a flowchart of a force feedback control method provided according to an embodiment of this application; Figure 6 This is a flowchart of a force feedback control method provided according to an embodiment of this application; Figure 7 This is a flowchart of a force feedback control method provided according to an embodiment of this application; Figure 8 This is a flowchart of a force feedback control method provided according to an embodiment of this application; Figure 9 This is a flowchart of the force feedback control method provided according to the embodiments of this application; Figure 10 This is a flowchart of the force feedback control method provided according to the embodiments of this application; Figure 11 This is a structural schematic diagram of a force feedback control device according to an embodiment of this application; Figure 12 This is a schematic diagram of the structure of an electronic device that implements the force feedback control method of the embodiments of this application. Detailed Implementation

[0018] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0019] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0020] It should be noted that the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information involved in the technical solution disclosed herein all comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken to prevent unauthorized access to user personal information data and to maintain user personal information security and network security. It should also be noted that the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information involved in the technical solution disclosed herein are all conducted with the user's knowledge and consent, and comply with relevant privacy protection regulations.

[0021] Before describing this technical solution, its application scenarios can be introduced first. The technical solution provided in this embodiment can be applied to any scenario that requires force feedback during the interaction between medical devices and biological tissues.

[0022] In existing technologies, operators primarily rely on two-dimensional visual images acquired by the endoscope and personal experience to indirectly infer the force applied to the tissue. This visual-compensated tactile approach makes it difficult for operators to accurately judge the magnitude of the contact force, easily leading to tissue damage due to improper force control, or operational errors such as instrument slippage and insecure sutures due to insufficient force, thus affecting operational safety. Furthermore, to achieve force feedback to the operator's hand, traditional solutions typically require installing force sensors at the instrument's end to acquire force information, which not only necessitates hardware modifications to the instrument but also increases system costs.

[0023] To address the aforementioned issues, this application proposes a real-time force feedback technology solution based on visual and mechanical mapping. This solution eliminates the need for force sensors at the end of surgical instruments. It calculates the mechanical information generated by the interaction between the instrument end and the target object and precisely maps it to the main control arm of the doctor's console, thereby reconstructing the operator's tactile perception channel.

[0024] By applying this technical solution, operators can clearly perceive minute changes in force as if they were directly holding the instrument, achieving precise control over the operating force, thereby effectively avoiding the risk of damage to the target object and improving the safety and accuracy of the operation.

[0025] In one example scenario, during operations requiring extremely high force perception, such as clamping fragile blood vessels or nerve tissue, performing delicate suturing and tissue dissection, the technical solution provided in this application can generate and feed back dynamic interactive force vectors applied to the target object by the instrument in real time. This allows the operator to realistically perceive subtle force fluctuations, ensuring precise control in complex operations. More importantly, since this solution does not rely on additional hardware force sensors and requires no changes to the existing hardware structure, it can achieve the effect of force feedback from the main hand. Therefore, this solution can be applied to most existing robot systems whose hardware structure does not contain force sensors, enabling the implementation of main hand force feedback functionality without hardware upgrade costs.

[0026] To more clearly illustrate the implementation details of this technical solution, a typical laparoscopic surgical robot system will be used as an example below. It should be noted that the force feedback control method provided in this application is not only applicable to this example system, but can also be widely applied to various surgical robot systems with a master-slave architecture.

[0027] A laparoscopic surgical robotic system typically consists of three parts: a surgeon's console, a patient carriage, and an imaging platform. The doctor's console serves as the operator's central control point. Seated at the console, the operator observes two-dimensional or three-dimensional images of the surgical area transmitted from the imaging platform, and issues commands to the main control arm, thereby remotely controlling the movement of the secondary robotic arms on the patient trolley. To achieve force feedback, a force feedback device can be integrated into the main control arm. For example, the motor driving the main control arm handle (i.e., the master hand) can function as the force feedback device. This motor can apply a reverse torque to the handle in real time based on the dynamic interactive force vector calculated by this technical solution, allowing the operator to truly perceive the force experienced when the instrument's end interacts with the tissue.

[0028] The patient trolley, located beside the patient's bed, is the system's execution terminal. It features a robotic arm that simulates the function of a human arm for support and positioning. Various instruments can be attached to the end of the robotic arm to simulate the human hand and reproduce the flexible movements of the wrist. Furthermore, the system has a physiological tremor filtering function, effectively eliminating the operator's natural hand tremors and improving the precision and stability of the operation.

[0029] The imaging platform is responsible for acquiring and processing visual information. It obtains surgical field images through a laparoscope (or "endoscope") placed inside the patient's body and transmits high-definition, three-dimensional, real-time images to the doctor's console, providing the operator with an immersive surgical view.

[0030] With the aforementioned system architecture, laparoscopic surgical robots are widely used in complex surgeries such as abdominal, thoracic, and general surgeries. The force feedback control method proposed in this application, which requires no additional hardware modifications, will further enhance the safety and accuracy of the system during delicate operations.

[0031] Figure 1 This is a schematic diagram of the doctor's console structure shown in an embodiment of this application. Figure 1 As shown, the doctor's control console includes a doctor's trolley chassis 101, a foot pedal movement adjustment component 102, at least one main control arm 103, a stereo monitor 104, a stereo monitor rotation adjustment component 105, a stereo monitor lifting adjustment component 106, and a handrail lifting adjustment component 107. The system includes a doctor's trolley chassis 101 for supporting and securing the entire doctor's console; a foot pedal adjustment component 102 mounted on the chassis for easy foot movement or locking of the trolley; a main control arm 103 providing an interface for outputting control commands to drive the slave robotic arms and target instruments on the patient trolley; a stereoscopic monitor 104 with an operator observation window consisting of two eyepieces for presenting a three-dimensional image of the surgical area; a stereoscopic monitor rotation adjustment component 105 for adjusting the rotation angle of the stereoscopic monitor 104, thereby changing the tilt angle of the eyepieces; a stereoscopic monitor height adjustment component 106 for adjusting the height of the stereoscopic monitor 104, thereby changing the height of the eyepieces; and an armrest height adjustment component 107 for adjusting the height of the armrests, providing comfortable arm support for the operator during operation. Through these adjustable structures, the doctor's console can accommodate users of different heights and operating habits, improving operational comfort and adaptability.

[0032] Figure 2 This is a schematic diagram of a patient trolley structure as illustrated in an embodiment of this application. The patient trolley typically consists of a chassis 201, a column 202, multiple slave robotic arms 203 connected to the column, and one or more instrument manipulators 204 located at the ends of support assemblies for each slave robotic arm. Target instruments and / or endoscopes are detachably mounted on the instrument manipulators 204. Each instrument manipulator 204 is used to support one or more target instruments and / or endoscopes operated on at the surgical site within the patient's body. The instrument manipulators 204 can control the associated surgical instruments in various ways with one or more mechanical degrees of freedom (e.g., all six Cartesian degrees of freedom, or five or fewer Cartesian degrees of freedom). Typically, through mechanical constraints or software limitations, the instrument manipulators drive the instruments to rotate around a center of motion that remains fixed relative to the patient. This center is typically located where the instrument enters the body wall, referred to as the "discent point" or "fixed point."

[0033] like Figure 2 As shown, the patient trolley is equipped with four robotic arms, including one endoscopic arm and three instrument arms. Figure 3 This is a schematic diagram of any robotic arm structure shown in an embodiment of this application. Figure 3 As shown, the robotic arm has 8 degrees of freedom (8-DOF). Figure 3 In the diagram, each dashed line and arrowed line represents a degree of freedom, labeled sequentially as Degree of Freedom 1, Degree of Freedom 2, Degree of Freedom 3, Degree of Freedom 4, Degree of Freedom 5, Degree of Freedom 6, Degree of Freedom 7, and Degree of Freedom 8. Each degree of freedom corresponds to a joint. The required Cartesian degrees of freedom for the end effector are only 6 (3 positional degrees of freedom and 3 orientation degrees of freedom), thus each end effector has redundant degrees of freedom (e.g., 2 redundant degrees of freedom), forming a redundant structure. This means that, with the end effector pose unchanged, the end effector can present countless configurations (i.e., combinations of joint angles). When the end effector pose is fixed, various configurations can be generated by adjusting the joint angles; the solution space formed by these configurations is the position null space.

[0034] An imaging platform typically consists of an image capture unit (commonly an endoscope) and one or more video displays. In some laparoscopic surgical robots, the endoscope tip integrates an imaging sensor (such as a CCD or CMOS) responsible for converting optical images inside the patient's body into electrical signals. These signals are then transmitted to the imaging platform's main unit via photoelectric conversion and a transmission link. After image processing algorithms such as enhancement and noise reduction, a high-definition surgical field image is finally displayed on the video monitor. This not only provides real-time visual feedback to the operator but also allows assistants and other medical staff to observe simultaneously.

[0035] In terms of teleoperation force transmission, the system typically employs a master-slave teleoperation architecture: commands issued by the operator at the doctor's console drive a remote-controlled motor, and the resulting driving force is transmitted via a precision transmission system to the end effector of the target instrument, thereby causing the instrument to complete the corresponding action. Such systems offer high spatial flexibility; the control unit (input device) can be deployed in locations far from the patient, whether within the same operating room, in different rooms, or even across cities to achieve remote surgery. Given that remote manipulation, remote control, and telepresence technologies are mature in this field, their specific system architecture and component details will not be elaborated upon here.

[0036] Based on the above system architecture, the force feedback control method provided in the embodiments of this application will be described in detail below.

[0037] Figure 4This is a flowchart illustrating a force feedback control method according to an embodiment of this application. This embodiment is applicable to any situation requiring force feedback during the interaction between an instrument and biological tissue. The force feedback control method can be applied to a master-slave surgical robot system. The surgical robot system includes a master control arm and at least one slave robotic arm controlled by the master control arm. The master control arm integrates a force feedback device, and the slave robotic arm carries the target instrument. This method can be executed by a force feedback control device, which can be implemented in hardware and / or software. It should be noted that the force feedback control device can be a computing unit integrated into a mobile carriage (such as a robot main control console), a standalone electronic device, or a functional module externally attached to the carriage or robotic arm base. Regardless of its physical form, the device can establish a data connection with the master control arm and the slave robotic arm to determine the interactive force data and drive the force feedback device on the master control arm to perform force feedback.

[0038] like Figure 4 As shown, the force feedback control method includes the following steps: S301. Acquire an interactive image reflecting the interaction state between the target instrument and the target tissue.

[0039] In this context, the target instrument refers to the execution tool used to perform physical interactive operations such as clamping, cutting, suturing, and peeling. For example, the target instrument can be a surgical instrument; it can also be other instruments with master-slave teleoperation capabilities (such as tooling instruments) to achieve master-slave teleoperation on an industrial robot.

[0040] The target tissue refers to the object that comes into direct contact with the instrument, is subjected to force, or is within the vicinity of the interaction range. For example, the target object can be a living organism or a teaching model (plastic sphere, silicone model, etc.) used for operator skill training. The interaction state characterizes the dynamic physical relationship between the target instrument and the target tissue at the moment of contact and during the continuous action process. The interaction image can be visual data reflecting the interaction state, and its form can be a single-frame static image or multiple frames of continuous time-series images. For example, raw image frames acquired in real time by a visual sensor (such as an industrial camera, depth camera, or endoscope) can be used as interaction images; or, after preprocessing the acquired raw images, the processed images can be used as interaction images to characterize the operational scenario and contact state at the current moment or within a specific time window.

[0041] In this embodiment, a visual acquisition module, such as an endoscope or a high-resolution visual sensor, can be deployed at or near the tip of the target instrument. When the target instrument interacts with the target tissue, the visual acquisition module can capture the microscopic texture, three-dimensional geometry, and color distribution information of the tissue surface in real time or periodically, and simultaneously output a continuous video stream. Keyframes can be extracted from this video stream as the basic interactive image, or at least one of the following preprocessing operations can be performed: denoising, region of interest (ROI) cropping, size normalization, smoothing filtering, contrast enhancement, color correction, and white balance adjustment, to generate an enhanced interactive image with a high signal-to-noise ratio.

[0042] S302. Input the interactive image into the force data prediction model and output the predicted interactive force data; the predicted interactive force data includes at least one interactive force vector.

[0043] The force data prediction model aims to infer the real-time mechanical state of biological tissues based on visual data. This model can be a data-driven model based on a deep learning architecture (such as at least one of Convolutional Neural Networks (CNN), Graph Neural Networks (GNN), or Vision Transformer (ViT)), or a hybrid model incorporating physical priors (such as a physics-driven model combined with finite element analysis). For example, a data-driven model can be trained using paired datasets of images and mechanics, enabling the model to learn the nonlinear mapping relationship between tissue visual features (such as indentation depth, texture stretching, and surface wrinkle morphology) and mechanical parameters, thus obtaining the force data prediction model. This allows the force data prediction model to invert the interactive mechanical state based on real-time visual input data.

[0044] Predicted interaction force data is used to characterize the interaction forces between the target instrument and tissue, and may include interaction force vectors with magnitude and direction. These force vectors accurately describe the force distribution in a three-dimensional coordinate system and can be decomposed into normal pressure perpendicular to the contact surface and tangential frictional force parallel to the contact surface, including triaxial force components and triaxial moment components, fully reflecting the six-dimensional force / moment state at the contact point. Furthermore, the predicted interaction force data may also include information such as the stress distribution range of the contact surface, the coordinates of the resultant force application point, and the force variation trend over time.

[0045] In one implementation, a single-frame or multi-frame interactive image can be input into a pre-trained force data prediction model. The model internally captures visual semantic features (such as tissue deformation gradient fields, contact edge curvature, and texture optical flow features) from the interactive images through feature extraction operations. Based on fully connected layers or regression heads, it analyzes the mapping relationship between visual semantic features and mechanical vectors, outputting predicted interactive force data.

[0046] In another implementation, geometric features (such as contact area, maximum indentation depth, deformation volume, and instrument intrusion angle) can be extracted from images based on computer vision algorithms or networks. These geometric features are then used as input vectors for a force data prediction model. The force data prediction model combines the pre-defined or online-identified tissue physical properties (such as elastic modulus, Poisson's ratio, and viscosity coefficient) to analyze the geometric features, calculate theoretical contact force data, and obtain predicted interaction force data.

[0047] In another implementation, a sequence of consecutive multi-frame interactive images can be input into a force data prediction model with spatiotemporal processing capabilities (such as a spatiotemporal network combining CNN with LSTM, GRU, or Transformer). The force data prediction model can analyze the static deformation features of a single-frame interactive image and the rate, acceleration, and texture flow trajectory of tissue deformation between consecutive interactive images to obtain dynamic deformation features. Then, force data recognition is performed based on the static and dynamic deformation features to obtain predicted interactive force data. The predicted interactive force data output in this approach can include instantaneous interactive force vectors and the force change trend at the next moment.

[0048] Based on single-frame or multi-frame time-series interactive images, the above method can accurately analyze the local deformation characteristics of tissues (such as compression, tension, and shear effects) and the real-time pose changes of instruments, thereby accurately identifying and quantifying the dynamic mechanical state during the contact process, and improving the prediction accuracy and real-time performance of subsequent force feedback control.

[0049] S303. Based on the predicted interaction force data, generate a force feedback control command and send the force feedback control command to the force feedback device to drive the force feedback device to output a force on the main control arm that matches the interaction force vector.

[0050] Force feedback control commands refer to drive signals generated by processing predicted interactive force data through proportional gain adjustment, filtering and denoising, and / or mapping transformation, conforming to the communication protocol of the force feedback device. These commands can manifest as control parameters such as motor target torque values, winding current setpoints, or position corrections, aiming to transform virtual mechanical data into executable physical actions. Force feedback devices can be electromechanical actuators with active force perception reproduction capabilities, capable of actively applying resistance, thrust, or high-frequency vibrations in three-dimensional space based on force feedback control commands, thereby simulating realistic physical contact textures. The main control arm is the input device for the operator to hold and manipulate; the operator remotely controls the movement of the target instrument by applying displacement or rotation to it.

[0051] Force feedback devices refer to interactive hardware that provides tactile feedback, capable of transmitting physical forces, torques, or vibrations to the operator to simulate tactile perception. Examples of force feedback devices include, but are not limited to: wearable devices (such as tactile gloves), handheld controllers (such as force feedback handles or interactive pens), master-slave operating terminals (such as the master control terminal of a robot hand or robotic arm), etc. Those skilled in the art should understand that any device capable of outputting controllable mechanical signals to the human body (including hands, limbs, torso, or feet) through force feedback is within the scope of protection of the force feedback devices described in this application.

[0052] The applied force refers to the actual physical force exerted by the force feedback device on the end effector of the main control arm. The vector direction of this force is consistent with the predicted tissue reaction force (or mirrored / mapped in the same direction according to operating habits), and its magnitude can be consistent with the magnitude of the interactive force vector, or scaled to fit the optimal perceptual range of the human hand, thereby providing the operator with a realistic tactile sensation of touching biological tissue. Optionally, at least one interactive force vector may include three-dimensional force components and three-dimensional torque components, and its output frequency can be synchronized with the interactive image acquisition frequency, such as above 30Hz; this interactive force vector is defined in the actuator coordinate system of the instrument end effector.

[0053] In this embodiment, the interactive force vector in the predicted interactive force data can be calculated into the target torque value required by each joint motor of the force feedback device. Based on this target torque value, a force feedback control command is generated. Alternatively, the interactive force vector can be multiplied by a preset force gain coefficient according to actual needs to adjust the sensitivity of the force feedback; then, based on the adjusted interactive force vector, the target torque value required by each joint motor of the force feedback device is calculated, and a force feedback control command is generated. The generated force feedback control command can be sent to the drive motor of the force feedback device via a real-time communication bus. The drive motor outputs an electromagnetic torque corresponding to the direction of the interactive force vector, which is then converted into physical resistance or thrust acting on the main control arm. For example, when it is predicted that the distal target instrument is subjected to an upward supporting force from the tissue, the main control arm is controlled to generate a downward damping sensation, and the magnitude of the force increases linearly with the increase of the predicted value, ensuring that the force sensation perceived by the operator's hand and the force state of the distal tissue are highly matched in terms of time synchronization and proportional consistency.

[0054] Furthermore, when the interactive force vector exceeds a preset risk threshold (indicating potential tissue damage), an additional repulsive force component can be superimposed on the force feedback control command. This setting increases the motion resistance of the force feedback device in the risk direction, limiting further displacement of the main control arm and preventing the operator from applying excessive force due to misoperation or loss of feel, thereby ensuring the safety of the interactive operation process.

[0055] The technical solution of this application embodiment predicts the force data of interactive images based on a force data prediction model, outputting predicted interactive force data including at least one interactive force vector. Based on the predicted interactive force data, a force feedback control command is generated and sent to a force feedback device, driving the force feedback device to output a force on the main control arm that matches the interactive force vector. This solves the problem in the prior art that the interaction force can only be judged by the operator's personal experience, which can easily lead to tissue damage or operational errors due to improper force control. Without changing the hardware structure of the surgical instruments (e.g., without adding force sensors), the force feedback control command is generated by calling the force data prediction model, thereby driving the force feedback device to output a force on the main control arm. This ensures that the force feedback device can accurately provide feedback to the operator on the force currently applied to the instrument, realizing tactile feedback to the operator, thereby reducing the risk of operator misoperation and ensuring the continuity and safety of master-slave operation of the instrument.

[0056] Figure 5 This is a flowchart of a force feedback control method according to an embodiment of this application. Based on the aforementioned embodiments, the force data prediction model may include a feature extraction layer, a feature fusion layer, and a force data prediction layer. Accordingly, in the process of inputting interactive images into the force data prediction model and outputting predicted interactive force data, feature extraction can be performed on multiple interactive images based on the feature extraction layer to obtain image features corresponding to the multiple interactive images; cross-frame feature fusion can be performed on multiple image features based on the feature fusion layer to obtain fused features characterizing the interaction process between the target device and the target tissue at multiple time points; force prediction can be performed on the fused features based on the force data prediction layer to obtain the predicted interactive force data. Specific implementation methods can be found in the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.

[0057] like Figure 5 As shown, the method specifically includes the following steps: S401. Acquire an interactive image reflecting the interaction state between the target instrument and the target tissue.

[0058] For further details regarding this step, please refer to S301 above; it will not be repeated here.

[0059] S402. Perform feature extraction on multiple interactive images based on the feature extraction layer to obtain image features corresponding to the multiple interactive images.

[0060] The feature extraction layer is a network layer used to extract features from the input data. For example, it can consist of components such as convolutional layers, pooling layers, or attention modules, aiming to convert high-dimensional image data into feature vectors or feature maps rich in semantic information. Multiple interactive images can be a group of image frames acquired sequentially over time during the contact between the instrument and tissue, reflecting dynamic features of the tissue over time, such as continuous deformation and displacement trajectory. Image features refer to the mathematical representations (such as high-dimensional tensors) obtained after processing by the feature extraction layer, capable of expressing deformation attributes such as edge direction, texture roughness, deformation amplitude, and motion vectors.

[0061] In one implementation, for each interactive image, a feature extraction layer (such as a convolutional neural network (CNN) or a visual encoder) can be used to independently analyze the texture details, edge contours, and spatial deformations of the interactive image, compressing and mapping high-dimensional pixel data into dense image features. This approach can generate a unique feature representation for each input, accurately capturing the instantaneous interactive state at the current moment.

[0062] In another implementation, multiple interactive images can be treated as a multidimensional data matrix containing spatial resolution and temporal sequence. A feature extraction layer (such as a convolutional kernel) is then used to simultaneously slide and scan in both the spatial and temporal domains, capturing spatial texture changes and dynamic deformation trends between consecutive interactive images, and outputting image features that fuse spatiotemporal information. This approach not only allows image features to reflect static contact morphology but also to express dynamic mechanical processes such as minute vibrations and rapid deformation rates, enhancing the model's ability to perceive dynamic interactions.

[0063] In another implementation, the feature extraction layer can include parallel spatial and temporal flow branches: the spatial flow branch processes interactive images that meet preset conditions, extracting high-resolution static features such as static texture, color, and geometric morphology. The temporal flow branch processes dynamic features such as motion vector fields, velocity distribution, and deformation rates based on optical flow maps or inter-frame difference maps determined from multiple interactive images. By performing feature stitching or weighted fusion on the static and dynamic features, these two types of heterogeneous features are integrated to generate comprehensive image features that possess both rich detail and dynamic perception capabilities, fully characterizing the complex interaction states between instruments and tissues.

[0064] For example, the feature extraction layer can be built based on a convolutional neural network or a Vision Transformer (ViT), which can efficiently extract high-dimensional semantic image features, including the dynamic deformation process of the tissue (such as stretching speed and displacement acceleration), the motion trend and precise pose and attitude angle of the instrument, as well as tissue texture and micro-deformation structure, so as to accurately estimate the interaction force data between the instrument and the tissue at the current moment.

[0065] It should be noted that the force data prediction model may also include an input layer. The input layer receives a continuous sequence of N interactive images (e.g., 16 frames). The input layer feeds the received interactive images into the feature extraction layer to perform feature extraction, obtaining image features corresponding to multiple interactive images.

[0066] S403. Based on the feature fusion layer, perform cross-frame feature fusion on multiple image features to obtain fused features that characterize the interaction process between the target instrument and the target tissue at multiple time points.

[0067] The feature fusion layer is a processing module used to integrate image features from multiple sources or at multiple time points. It aims to associate, weight, and aggregate features from different time steps through operations such as attention mechanisms, recurrent neural network (RNN) units, or temporal convolutions to generate fused features. These fused features can comprehensively characterize the dynamic interaction between the target device and tissue across multiple time points, from initial contact and continuous force application to deformation feedback. For example, the feature fusion layer can be composed of at least one of a Long Short-Term Memory (LSTM) network, a gated recurrent unit (GRU), a self-attention module, and a temporal convolutional network (TCN).

[0068] In one implementation, image features arranged chronologically are sequentially input into a feature fusion layer based on LSTM or GRU. Gating units such as forget gates, input gates, and output gates are used to dynamically filter and retain key deformation information from historical moments, while filtering out irrelevant background noise. As time progresses, the network's hidden state is continuously updated, fusing current features with accumulated historical memory features to output fused features. This approach can capture long-term dependencies and viscoelastic hysteresis effects during tissue stress processes.

[0069] In another implementation, the entire temporal image feature sequence can be input into a feature fusion layer based on a self-attention mechanism. By calculating the similarity between features at any two time points in the image feature sequence, a relevance weight matrix (attention score) reflecting the degree of association is generated. Then, the features at all time points are weighted and summed to obtain the fused features. This method can dynamically aggregate global information of the sequence, giving high weight to historical features highly correlated with the current state for key retention, while suppressing irrelevant noise. This allows the generated fused features to accurately characterize nonlinear abrupt changes or periodic mechanical oscillation patterns in the interaction process, improving the accuracy of force recognition for complex dynamic changes.

[0070] In another implementation, a feature fusion layer can be used to perform weighted summation and nonlinear transformation on multiple image features within a local time window to output fused features. This method can extract local trends of feature changes over time (such as high-frequency dynamic information like acceleration and jerk), allowing the generated fused features to retain fine local dynamic details and fully integrate the temporal context information of the neighborhood.

[0071] To accurately extract the dynamic deformation during the interaction between the target device and the tissue, the feature fusion layer may include a temporal offset module and an attention module. A specific implementation of performing cross-frame feature fusion on multiple image features based on the feature fusion layer to obtain fused features characterizing the interaction process between the target device and the target tissue at multiple time points can be as follows: Based on the temporal offset module, the feature values ​​of some channels in multiple image features are offset along the time dimension to obtain an enhanced feature sequence; the enhanced feature sequence is input to the attention module, which outputs the spatiotemporal correlation weights between multiple frame features in the enhanced feature sequence; weighted fusion processing is performed on the enhanced feature sequence according to the spatiotemporal correlation weights to obtain the fused features.

[0072] The Temporal Shift Module (TSM) aims to construct an enhanced feature sequence containing motion information from adjacent frames by directionally shifting the image feature sequence along the time axis. Each channel corresponds to a specific semantic information or physical property (such as edge texture, color distribution, motion vectors, or high-frequency noise). Spatiotemporal correlation weights are used to quantify the degree of dependence between features at different times and spatial locations in the enhanced feature sequence. High weights indicate a strong causal relationship (such as the current deformation being caused by contact several frames ago), while low weights represent weak correlations or noise.

[0073] In this embodiment, the temporal offset module can divide the feature channels into at least one subset of left-shifted channels, right-shifted channels, and hold channels. Specifically, the method for offsetting the feature values ​​of some channels in the image features along the time dimension can be as follows: For the channels belonging to the left-shifted channel subset, their feature values ​​in the time dimension are shifted forward by one time step. That is, the feature values ​​of this channel at the current time t are replaced with the corresponding feature values ​​at the previous time t-1. This operation allows the network at the current time to see some information from the previous frame, thereby capturing the motion trend from the past to the present (such as the deformation direction of the tissue). For the channels belonging to the right-shifted channel subset, their feature values ​​in the time dimension are shifted backward by one time step. That is, the feature values ​​of this channel at the current time t are replaced with the corresponding feature values ​​at the next time t+1. For the remaining channels belonging to the hold channel subset, their feature values ​​remain unchanged at the current time t, ensuring that the network can still directly obtain the current spatial texture and instantaneous state. This differential offset operation enables the feature vector of each spatial location to fuse information from multiple time points. Without significantly increasing computational complexity, it endows the network with pixel-level motion perception capabilities similar to optical flow, which can accurately capture subtle dynamics such as local stretching and shear deformation when biological tissue is subjected to force, thereby improving the estimation accuracy of instantaneous force values ​​and force change rates.

[0074] Furthermore, the enhanced feature sequence can be input into a self-attention module (based on Transformer). This module adaptively learns and calculates the spatiotemporal correlation weights between frame features at any two time points through query, key, and value mapping, effectively capturing long-distance spatiotemporal dependencies and identifying the historical offset features that contribute the most to the current state. Using the calculated spatiotemporal correlation weights, weighted summation and aggregation are performed on the frame features in the multi-scale enhanced feature sequence to generate the final fused feature. This fused feature not only preserves local fine dynamic details but also incorporates long-term accumulated stress information, achieving a full-band characterization of the mechanical behavior of complex viscoelastic organizations. It can respond to instantaneous rapid deformations and remember long-term interactive progressive states.

[0075] This processing method, which cascades the temporal offset module with the attention mechanism, can explicitly encode the dynamic evolution information of the time dimension into the feature space. Through adaptive weighting, it achieves the decoupling and recombination of spatiotemporal information, enabling the generated fused features to accurately distinguish between the real mechanical state and environmental noise, thereby improving the accuracy, robustness and generalization ability of force data prediction.

[0076] S404. Based on the force data prediction layer, force prediction is performed on the fused features to obtain predicted interactive force data; the predicted interactive force data includes at least one interactive force vector.

[0077] Among them, the force data prediction layer is a module used to nonlinearly map high-dimensional visual semantic information into specific mechanical parameters. Its structure can be flexibly designed as a multilayer perceptron (MLP), decoder, or hybrid neural network, etc.

[0078] In one implementation, a global pooling operation can be performed on the fused features based on the force data prediction layer, compressing them into a fixed-length global feature vector to eliminate spatial redundancy and preserve the overall semantics. This global feature vector is then input into a network consisting of multiple fully connected layers with activation functions (such as ReLU or GELU). Through successive nonlinear transformations, the abstract visual encoding is decoded into specific mechanical physical quantities, and the output layer generates predicted interactive force data representing the overall force state.

[0079] In another implementation, the force data prediction layer can be a network structure consisting of multiple cascaded fully connected layers. The fused features are input into the force data prediction layer, and based on its internal multi-layered fully connected layers, the high-dimensional fused features obtained through spatiotemporal fusion are mapped to output data, resulting in predicted interactive force data. For example, the interactive force vectors in the predicted interactive force data include the force / torque state in the current time frame of the instrument's end effector coordinate system.

[0080] It should be noted that the above model structure is merely an exemplary implementation for predicting interactive force data. In practical applications, any model architecture capable of mapping interactive image features to mechanical parameters (including but not limited to other deep learning variants or hybrid models) falls within the scope of this embodiment.

[0081] S405. Based on the predicted interaction force data, generate a force feedback control command and send the force feedback control command to the force feedback device to drive the force feedback device to output a force on the main control arm that matches the interaction force vector.

[0082] The technical solution provided in this embodiment, by performing feature extraction on multiple interactive images, can selectively filter out non-mechanically relevant factors such as illumination noise and background interference, accurately capture the deformation patterns and dynamic evolution processes of tissues, and transform the original visual data into image features with high semantic density. This reduces the computational load of the force data prediction model and improves the model's sensitivity and robustness to weak mechanical signals. Furthermore, by performing cross-frame fusion on the features of multiple frames through a feature fusion layer, discrete instantaneous visual features are transformed into continuous dynamic process features. The generated fusion features not only retain rich spatial details but also reflect the temporal state of tissue deformation, such as rate, direction, and historical cumulative effects. This enhances the model's understanding of complex mechanical properties of biological soft tissues, such as nonlinearity and viscoelasticity, enabling it to effectively distinguish between instantaneous noise and real mechanical responses. Consequently, it improves the robustness and prediction accuracy of the force data prediction model in dynamic interactive scenarios. Through the force data prediction layer, the interactive force state between the target instrument and the tissue can be accurately identified with high confidence, and the interactive force vector containing direction and magnitude is output. This ensures the accuracy of the interactive force data output and the accuracy of the robot system's force feedback control, thereby improving the safety of the interaction process between the instrument and the tissue.

[0083] Figure 6 This is a flowchart of a force feedback control method according to an embodiment of this application. Based on the foregoing embodiments, the force data prediction model includes a style prediction network and a data prediction network. Furthermore, the interactive image, reference style image, and target noise image can be input into the style prediction network to obtain predicted style information; the data prediction network is then invoked to perform prediction on the predicted style information to obtain predicted interactive force data. Specific implementation details can be found in the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.

[0084] like Figure 6 As shown, the method specifically includes the following steps: S501. Acquire an interactive image reflecting the interaction state between the target instrument and the target tissue.

[0085] S502. Input the interactive image, the reference style image, and the target noise image into the style prediction network to obtain the predicted style information.

[0086] The reference style image can be a standardized ideal image sample pre-generated through manual labeling or a combination of various traditional image processing algorithms (such as edge extraction, contrast enhancement, layer separation, etc.). This image can present an ideal state image of the key features required for the force prediction task. For example, the reference style image can be a structured template enhanced from historical interaction images, such as highlighting instrument and tissue outlines, enhancing edge contrast, preserving gloss texture, or a guide image with at least one of the following based on occlusion relationships and mask annotation; it can also be an image after cropping or blurring non-ROI regions to highlight the core area of ​​interest and suppress background interference.

[0087] The target noise image can be pre-configured or randomly generated image data containing random noise, artifacts, or specific interference patterns. For example, the target noise image can be a pixel matrix that conforms to a Gaussian or uniform distribution. Unlike the semantic information of the actual interactive image or the style guidance of the reference image, the target noise image itself does not contain any structured semantics and mainly serves as a perturbation signal to introduce randomness.

[0088] In this embodiment, the role of the style prediction network is to reconstruct predicted style information that conforms to a specific operational interaction style by performing progressive denoising, iterative optimization, or feature mapping on the target noisy image, under the constraints of a reference style image and guided by the features of the actual interaction image. The output form of this predicted style information may include style vectors, style parameters, or style feature maps, which retain the geometric structure of the original interaction image and incorporate a high-dimensional representation of the specific operational interaction style. It should be noted that the texture changes, edge enhancements, or color gradients generated by this stylization process can be strongly correlated with physical phenomena such as tissue deformation and contact pressure distribution, thereby effectively assisting in the perception of mechanical states.

[0089] In one implementation, the style prediction network can extract content features of the interactive image and style statistical features (such as channel mean and variance) of a reference style image, based on a shared or independent encoder. The style prediction network can then convert the target noisy image into a fine-tuned offset for the style statistical features and correct these features. Normalization (such as adaptive instance normalization) is performed on the corrected style statistical features and content features to generate predicted style information. This approach achieves fine-grained control over the generated style while enhancing the randomness and diversity of style expression through noise injection.

[0090] In another implementation, the style prediction network can be pre-trained on a large-scale style image dataset to enable it to generate high-quality style images. Specifically, the interactive image, the reference style image, and the target noisy image can all be used as input data for the style prediction network. The network then performs a style transfer operation on the target noisy image by fusing the content features of the interactive image with the style features of the reference style image, outputting predicted style information.

[0091] To achieve dynamic fusion of style and content under noise guidance and generate predicted style information with reference style features, interactive content semantics, and random diversity, the following steps are taken during the process of inputting the interactive image, reference style image, and target noise image into the style prediction network and outputting predicted style information: First coding features corresponding to the target noise image, second coding features corresponding to the interactive image, and third coding features corresponding to the reference style image are determined. Feature concatenation is then performed on the coding sub-features at the same spatial location among the first, second, and third coding features to obtain concatenated features. Cross-attention processing is then performed on the first coding features and the concatenated features to obtain attention features. Finally, denoising is performed on the attention features to obtain the predicted style information.

[0092] In this embodiment, the style prediction network can achieve fine-grained alignment of noise, content, and style features based on a multimodal coding and attention fusion mechanism. Specifically, the style prediction network can encode the target noisy image, the interaction image, and the reference style image respectively, extracting a first encoded feature (noise feature) containing a random perturbation prior, a second encoded feature (content feature) carrying the real scene content and structure, and a third encoded feature (style feature) containing clues about the target operation interaction or operation style. In order to fuse multimodal features in a local space, the sub-features at the same spatial position of the first, second, and third encoded features can be concatenated to construct a concatenated feature that simultaneously covers the noise context, content structure, and style texture. A cross-attention mechanism is adopted, using the first encoded feature (noise prior) as the query and the concatenated feature as the key and value, so that the first encoded feature can adaptively focus on and aggregate effective information (such as interaction content structure and style texture) in the concatenated feature, generating semantically rich attention features. The attention features are denoised to remove invalid random components, and the accurate predicted style information (such as predicted style images or high-dimensional style features) is reconstructed step by step.

[0093] For example, the image prediction model can receive a reference style image set, the actual acquired interactive image, and an initial noisy image (such as Gaussian noise, i.e., the target noisy image) as input data. After encoder mapping, the above three types of encoded features are obtained; concatenated features are obtained by concatenating features at the same position; the first encoded features are mapped to the content and style space using an attention mechanism to generate attention features; after denoising and decoding, the attention features are reconstructed into predicted style information that conforms to the desired style.

[0094] The above method generates predictive style information that is similar to the reference style by integrating the visual norms of the reference style image, the structural constraints of the interactive image, and the randomness introduced by noise. This ensures that key interactive information (such as tissue deformation and instrument boundaries) is not distorted, improves the expressiveness of the image, and thus assists in subsequent mechanical state perception, improving the efficiency and accuracy of predicting interactive force data.

[0095] S503. Call the data prediction network to perform prediction on the prediction style information to obtain prediction interaction force data; the prediction interaction force data includes at least one interaction force vector.

[0096] The data prediction network aims to establish a mapping relationship between specific visual style expressions and mechanical values. This network can be optimized using a training set containing style image samples and their corresponding mechanical annotations. Optionally, the model architecture of the data prediction network can be a regression analysis model, a multilayer perceptron (MLP), a convolutional neural network (CNN), or a Transformer-based sequence prediction model.

[0097] Specifically, the generated predicted style information can be input into the data prediction network. Based on the learned mapping rules between style and mechanics, the network analyzes the implicit physical properties, such as the tissue stiffness corresponding to texture changes and the contact pressure corresponding to edge deformation, and then derives the appropriate predicted interaction force data.

[0098] To further improve the accuracy of temporal dynamic perception, the data prediction network can also be a cascaded architecture consisting of a feature extraction layer, a feature fusion layer, and a force data prediction layer.

[0099] When the predicted style information is a predicted style image, each frame of the image can be encoded through the feature extraction layer to obtain high-dimensional image features; the feature fusion layer (integrating temporal convolution or attention mechanism) is used to perform cross-frame fusion of these multi-time image features to generate fused features that can characterize the dynamic interaction process between the target device and the tissue in the continuous time dimension; the force data prediction layer performs regression analysis on the fused features to output the final force data.

[0100] When the predicted style information is a predicted style vector, the data prediction network can be simplified into a cascaded architecture consisting of a feature fusion layer and a force data prediction layer. In this case, the predicted style vector can be used as a pre-extracted image feature. Based on the feature fusion layer, cross-frame feature fusion is performed on the predicted style vectors at multiple time points to obtain fused features that characterize the spatiotemporal changes of the interaction process. Based on the force data prediction layer, force prediction is performed on the fused features to obtain the predicted interaction force data.

[0101] S504. Based on the predicted interaction force data, generate a force feedback control command and send the force feedback control command to the force feedback device to drive the force feedback device to output a force on the main control arm that matches the interaction force vector.

[0102] The technical solution provided in this embodiment decouples content and style by inputting interactive images, reference style images, and target noise images into a style prediction network. The target noise image enhances the randomness and diversity of the predicted style information, enabling the predicted style information to retain key interactive operation information in the interactive image while possessing high style expressiveness. Furthermore, by calling the data prediction network to perform mechanical inference on the predicted style information, the network fully utilizes the mechanically relevant features (such as texture enhancement in stress concentration areas) explicitly amplified during stylization processing. This allows the network to accurately capture minute deformations and stress signs, ensuring the accuracy of the generated predicted interactive force data. This data maintains high confidence even under complex conditions such as drastic changes in lighting and individual tissue differences, thereby improving the accuracy and safety of force-tactile feedback.

[0103] Figure 7 This is a flowchart of a force feedback control method provided according to an embodiment of this application. Based on the aforementioned embodiments, the force data prediction model further includes an uncertainty prediction layer. The uncertainty prediction layer performs confidence assessment on the fused features to obtain uncertainty attributes. When the uncertainty attributes meet preset reliability conditions, a force feedback control command is generated based on the predicted interactive force data. Specific implementation details can be found in the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.

[0104] like Figure 7 As shown, the method specifically includes the following steps: S601. Acquire an interactive image reflecting the interaction state between the target instrument and the target tissue.

[0105] S602. Perform feature extraction on multiple interactive images based on the feature extraction layer to obtain image features corresponding to the multiple interactive images.

[0106] S603. Based on the feature fusion layer, perform cross-frame feature fusion on multiple image features to obtain fused features that characterize the interaction process between the target instrument and the target tissue at multiple time points.

[0107] S604. Based on the force data prediction layer, force prediction is performed on the fused features to obtain predicted interaction force data; based on the uncertainty prediction layer, confidence assessment is performed on the fused features to obtain uncertainty attributes.

[0108] Typically, force data prediction models output their confidence scores along with the predicted interaction force data. However, the confidence scores output by the model may be biased (e.g., overconfidence), meaning that a high confidence score does not necessarily represent high accuracy. To address this issue (i.e., high confidence scores do not match actual accuracy), an uncertainty attribute can be generated through an uncertainty prediction layer. This uncertainty attribute characterizes the reliability of the confidence scores output by the model. The uncertainty attribute is negatively correlated with the reliability of the confidence score: a higher uncertainty attribute value indicates a less reliable confidence score for the predicted interaction force data, meaning the model lacks confidence in the current prediction or is at risk of overconfidence; a lower uncertainty attribute value indicates a more reliable confidence score for the predicted interaction force data. In other words, even if the model outputs high-confidence predicted interaction force data, a high corresponding uncertainty attribute value may mean that the high confidence score does not accurately reflect the prediction accuracy.

[0109] For example, the model might provide a high-confidence prediction of interaction force data, but this high confidence level could be incorrect (i.e., the model is overconfident). The uncertainty attribute is used to determine whether this high confidence level truly reflects the accuracy of the predicted interaction force data, preventing misleading force feedback due to the model's incorrect assessment of confidence.

[0110] In one implementation, when the fused feature is input to the uncertainty prediction layer, this layer can employ a multiple sampling strategy based on random perturbation (e.g., enabling random feature discarding, parameter noise injection, or Monte Carlo sampling) to perform multiple forward propagations on the same fused feature, thereby generating a set of interaction force data samples. The uncertainty attribute can be quantified by calculating the statistical dispersion index (such as variance or entropy) of this set of interaction force data samples. If the values ​​in this set of interaction force data are close to each other (i.e., low dispersion), it indicates that the model's understanding of the current fused feature is highly consistent. In this case, the generated uncertainty attribute value is low, indicating that the predicted interaction force data output by the model and its accompanying confidence level are highly reliable. Conversely, if the values ​​in this set of interaction force data differ greatly (i.e., high dispersion), it indicates that there is disagreement or ambiguity in the model's internal understanding of the current feature. In this case, the generated uncertainty attribute value is high. In this high uncertainty scenario, even if the model outputs a specific predicted value and a corresponding high confidence score, the confidence level is still judged as unreliable (i.e., there is a risk of overconfidence). Therefore, by quantifying the dispersion of the prediction results, the uncertainty attribute can truly reflect the stability of the model's cognition, thereby verifying and correcting the credibility of the original confidence index and preventing downstream decisions from being misled by incorrect confidence information.

[0111] In another implementation, the uncertainty prediction layer can be constructed as an integrated architecture containing multiple sub-models. These sub-models have the same network structure but are trained with different initial weights or based on different subsets of training data. When fused features are input into this uncertainty prediction layer, all sub-models perform inference in parallel, outputting their respective predicted interaction force data and corresponding confidence scores. The uncertainty attribute can be comprehensively calculated by statistically analyzing the dispersion (e.g., standard deviation) and consistency of the confidence scores of the interaction force data output by each sub-model. Specifically, if the interaction force data predicted by each sub-model is highly similar and the confidence scores are consistent, the generated uncertainty attribute value is low, indicating that the confidence score of the current prediction result is highly reliable. Conversely, if the interaction force data predicted by the sub-models differs greatly, or if there is a large deviation in the confidence score assessment of the same predicted interaction force data, a high uncertainty attribute value is generated, thus warning that the current prediction confidence score is unreliable.

[0112] In a specific network architecture, both the force data prediction layer and the uncertainty prediction layer can be connected after the feature fusion layer used to output fused features.

[0113] For example, a dual-branch regression head can be constructed through multiple fully connected layers: the force data prediction branch (i.e., the force data prediction layer) is used to output a 6-dimensional continuous interactive force vector (represented as the force and torque state in the coordinate system of the end effector of the instrument at the current moment); the uncertainty prediction layer (i.e., the uncertainty prediction layer) simultaneously outputs an uncertainty attribute scalar, which is used to quantify the reliability of the confidence level of the above-predicted interactive force vector.

[0114] By generating uncertainty attributes representing the reliability of confidence levels based on an uncertainty prediction layer, the credibility of the model's confidence assessment itself is identified, addressing the overconfidence problem that models easily exhibit in off-training data or complex dynamic interactions. This allows for dynamic adjustment of the force feedback control strategy based on the uncertainty attribute. For example, high-precision force feedback operations are executed when the uncertainty attribute is below a preset threshold (i.e., the confidence level is reliable); when the uncertainty is high, a safe and conservative operation is switched to improve the accuracy of the force feedback operation and avoid operational risks caused by erroneous predictions from the force data prediction layer.

[0115] S605. When the uncertainty attribute meets the preset reliability conditions, generate force feedback control commands based on the predicted interactive force data.

[0116] In this embodiment, an uncertainty attribute can be compared with a preset reliability condition; wherein, the preset reliability condition is that the value of the uncertainty attribute is less than or equal to a preset safety threshold.

[0117] In response to the determination that the uncertainty attribute meets the preset reliability condition (i.e., the characterization model prediction is in a high confidence state), a force feedback control command is generated based on the predicted interactive force data, and the force feedback device is driven to perform the corresponding force control operation.

[0118] Conversely, if the uncertainty attribute is determined not to meet the preset reliability condition (i.e., the characterization model prediction is in a low confidence or unreliable state), a safety protection strategy can be triggered, including but not limited to pausing force feedback output, switching to passive impedance control mode, or issuing a warning signal, thereby effectively avoiding control instability caused by inaccurate prediction data.

[0119] Optionally, when the uncertainty attribute does not meet the preset reliability conditions, at least one of the following security processing operations shall be performed: Generate force feedback control commands to maintain or attenuate the force, and send the force feedback control commands to the force feedback device to drive the force feedback device to output the corresponding force; The warning message is displayed on the main control device where the main control arm is located.

[0120] The force feedback control command refers to the safety control signal generated when a potential force feedback risk is detected. It aims to drive the force feedback device to output zero force, maintain the current state, or decay the force according to a specific curve, thereby eliminating false resistance or misleading thrust caused by prediction deviations. The main control device, as the operating terminal (such as a robot's main control console), can integrate a main control arm, a computing unit, and a display interface. The display interface is used to present endoscopic videos or 3D models and display warning messages. Optionally, the warning messages include operation guidance information or force feedback status description information. Operation guidance information can be used to characterize and guide the operator on how to perceive the interaction force between the instrument and tissue, such as "Please switch to manual mode," "Please adjust the camera angle," or "Please pause the operation." Force feedback status description information refers to a description of the current force feedback status, such as low confidence in the force feedback data, automatic disabling of tactile output, or the use of a safety decay mode, helping the operator understand why the system changes its force perception.

[0121] When the uncertainty attribute fails to meet the preset reliability conditions, the force feedback control command currently output to the force feedback device can be smoothly reduced to zero within a preset time period based on an exponential decay algorithm, preventing sudden force loss from causing operator hand loss of control. The force feedback device can also be driven into a high-damping lock mode or kept stationary, limiting the operator's control of the main control arm's rapid movement and preventing misoperation when data is unreliable. Alternatively, tiered processing can be performed based on the degree to which the uncertainty attribute exceeds the preset reliability conditions. Optionally, if the exceedance of the preset reliability conditions is small, the predicted interactive force data can be significantly attenuated (e.g., retained by 10%), generating a force feedback control command and displaying a force feedback status description on the display interface; if the exceedance is large, a zero-force force feedback control command is generated, cutting off the force feedback and displaying a strong warning message containing detailed operation guidance information. The content of the warning message can also be dynamically updated according to changes in the uncertainty attribute, such as changing from force feedback weakening to force feedback disabling and then back to normal. It can also output warning prompts on the main control device's display interface according to preset display formats (such as highlighting, scrolling, modal dialog boxes), including force feedback status description information, clearly informing the operator that the current tactile data is unreliable and the force feedback has been automatically reset to zero; or prompting that low confidence has been detected, please adjust the viewing angle or pause the operation, and continue after the system recovers.

[0122] By combining the force feedback attenuation / locking operation with visual warning information, erroneous force perception interference can be eliminated, ensuring that operators can promptly understand the system status and adjust their operational behavior when data is unreliable, thereby guaranteeing the safety of remote operation or virtual training processes.

[0123] In this embodiment, the methods for generating force feedback control commands to maintain or attenuate the force include at least two of the following: One implementation method is to use the historical interaction force data output at the previous moment as the initial value, and generate a force feedback control command that decays to a preset threshold within a preset time period.

[0124] Here, "previous time" refers to the most recent time before the current time when the preset reliability conditions were met. "Preset duration" is the time window for implementing the force feedback attenuation strategy. "Preset threshold" is the termination value of the attenuation process, which can be set to zero or a very small safety damping force.

[0125] When the uncertainty attribute does not meet the preset reliability condition, historical interaction force data from the previous moment when the preset reliability condition was met can be obtained as the initial value. A smooth-transition force feedback control command curve is generated within a preset duration, causing the force to gradually decrease from the initial value to a preset threshold. Specifically, within the preset duration, a linear decay can be applied to the initial force, generating a force feedback control command curve that linearly and uniformly decreases from the initial value to the preset threshold. For example, if the initial value is 5 Newtons and the preset duration is 2 seconds, the force feedback control command decreases by 2.5 Newtons per second until it reaches zero. This setting simulates the physical process of uniform resistance dissipation, avoiding abrupt output jumps due to data interruption and preventing involuntary hand tremors or misoperations caused by sudden changes in reaction force.

[0126] It can also generate a force feedback control command curve that exponentially decays from the initial value to a preset threshold within a preset time period. The advantage of exponential decay is that it can quickly eliminate most potentially misleading force feedback when uncertainty occurs, and then retain a faint tactile cues, improving the smoothness and safety of force feedback.

[0127] Another approach is to use the historical interaction force data output from the previous moment as a constant value to generate a force feedback control command to maintain the force constant.

[0128] When the uncertainty attribute does not meet the preset reliability conditions, historical interaction force data from the previous moment can be acquired and used as a constant value to continuously generate force feedback control commands that maintain a constant force. Throughout the subsequent unreliable time period, regardless of how the operator moves the main control arm, the force feedback device outputs a constant force consistent with this historical data. This constant force method effectively maintains the physical continuity of force output, avoiding sudden changes or gaps in force perception due to data loss. This ensures the continuity of interactive operation while preventing misjudgments by the operator due to abnormal tactile feedback, thus ensuring the safety and stability of the interaction process.

[0129] The technical solution provided in this embodiment performs quantitative analysis on the fused features through an uncertainty prediction layer, outputs an uncertainty attribute that characterizes the reliability of the confidence level of the predicted interactive force data, and generates force feedback control commands based on the predicted interactive force data only when the uncertainty attribute meets the preset reliability conditions. This solves the risk of misoperation caused by the model's blind confidence in unknown working conditions in the prior art, avoids misjudgment leading to operational misguidance or potential tissue damage, and improves the safety of tactile feedback.

[0130] Figure 8 This is a flowchart of a force feedback control method according to an embodiment of this application. Based on the foregoing embodiments, a force data prediction model can also be pre-trained. Specific implementation details can be found in the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.

[0131] like Figure 8 As shown, the method specifically includes the following steps: S701. Acquire multiple training samples; wherein, the training samples include interaction sample images reflecting the interaction state between surgical instruments and tissues, and corresponding interaction force sample data.

[0132] To improve model accuracy, a large-scale and diverse training sample set can be constructed. Each training sample consists of a pair of interaction sample images and interaction force sample data.

[0133] Interactive sample images can be acquired by a vision acquisition module (such as an endoscope or camera), reflecting the visual state when surgical instruments (forceps, scissors, etc.) come into contact with tissue (biological tissue or tissue model). Interactive force sample data can be simultaneously measured by a high-precision force sensor (such as a six-dimensional force / torque sensor), containing the magnitude, direction, and torque information of the interaction forces, characterizing the real physical force situation.

[0134] In practice, while controlling surgical instruments to perform grasping, traction, or other operations on isolated tissues or biomimetic models, the vision acquisition module can be activated simultaneously to acquire images and read real-time readings from the force sensor at the end of the instrument. Interactive sample images acquired at the same time are matched with interaction force sample data to form high-fidelity training samples for realistic interactive scenarios. For example, when surgical instruments interact with tissue containing embedded blocks, the interaction state caused by abrupt changes in mechanical response (such as increased resistance at hard points) can be accurately captured (e.g., instrument micro-vibration, abnormal tissue deformation), resulting in interactive sample images.

[0135] Digital instrument and tissue models can also be built in a virtual environment, using a simulation engine to preset complex interactive scenarios (such as different tissue stiffness, multi-angle entry, abnormal adhesion, etc.). Through automated scripts running the simulation, the engine can simultaneously render interactive sample images including lighting and deformation textures, and output accurate interactive force sample data calculated by the physics engine. This method can generate massive amounts of samples covering various extreme working conditions at low cost and high efficiency, effectively solving the problems of difficult real-world data collection, high annotation costs, and difficulty in reproducing dangerous scenarios.

[0136] Typically, due to differences in installation location, sampling frequency, and coordinate system between visual sensors and force sensors, directly acquired image frames cannot directly correspond to the original force data in time and space. Furthermore, the original force data often contains noise or is affected by gravity. To solve this technical problem, systematic errors can be eliminated by aligning the spatiotemporal references of multimodal data and unifying their physical representation.

[0137] In this embodiment, the training samples are obtained as follows: Acquire multiple sets of interaction sample images of surgical instruments interacting with tissues, as well as the original interaction force data of the surgical instrument tip; perform correction and coordinate transformation on the original interaction force data to obtain interaction force sample data in the coordinate system of the surgical instrument tip; pair each set of interaction sample images with interaction force sample data to form multiple training samples.

[0138] Among them, force feedback sensors can be integrated into surgical instruments to collect raw interaction force data.

[0139] Specifically, when surgical instruments interact with tissue, the visual acquisition module at the end of the instrument can acquire images of the interaction samples, while an integrated force feedback sensor records the raw interaction force data. It should be noted that the raw interaction force data is usually an unprocessed electrical signal or physical quantity, which may contain errors such as sensor zero-point drift, temperature noise, and gravitational component interference.

[0140] Furthermore, based on pre-calibrated sensor installation parameters (including but not limited to: the sensor's installation position vector relative to the instrument tip, attitude rotation matrix, sensitive axis direction vector, and the mass centroid parameter of the instrument tip), one or more of the following combined processing methods can be performed on the raw interaction force data: removing gravity component interference from the raw interaction force data; filtering out high-frequency noise and drift signals from the raw interaction force data; and transforming the raw interaction force data to the standard coordinate system of the surgical instrument tip to obtain standardized interaction force sample data. Kalman filtering and other processing methods can also be performed on the raw interaction force data to eliminate high-frequency electromagnetic noise.

[0141] One or more frames of interactive sample images can be paired with calibrated interactive force sample data from the same time point to construct a high-precision training sample set. Image enhancement processing can also be performed on the interactive sample images, including but not limited to: cropping to focus on the target operation region (e.g., the surgical region), and size normalization to a preset resolution (e.g., ...). This includes color normalization and other processing. A training sample set is obtained by pairing the processed interactive sample images with the interactive force sample data.

[0142] This processing method, through precise spatiotemporal alignment and physical correction, effectively eliminates force measurement errors caused by sensor installation deviations and environmental interference, ensuring the accuracy and consistency of the mapping relationship between visual morphology and physical force in the training samples. This not only improves the generalization ability and robustness of subsequent force data prediction models in complex and variable operating scenarios, but also enables them to accurately reproduce real force feedback information from input images.

[0143] To eliminate the interference of sensor gravity, instrument inertia, and installation posture deviation on measurement results, and thus obtain pure mechanical signals that truly reflect the contact between the instrument and tissue, this scheme can, during the process of correcting and transforming the original interaction force data to generate standard interaction force sample data, subtract the weight and inertia of the original interaction force data based on the motion state data of the surgical instrument at the time of acquisition of the original interaction force data to obtain the true interaction force data; and transform the true interaction force data to the end coordinate system of the surgical instrument through a preset coordinate transformation matrix to obtain the interaction force sample data.

[0144] Motion state data refers to the real-time dynamic parameters of the surgical instrument acquired through kinematic calculations or an inertial measurement unit (IMU) at the same moment as the acquisition of raw interaction force data. These parameters include, but are not limited to, position, attitude, linear velocity, angular velocity, linear acceleration, and angular acceleration. Self-weight and inertial forces (including centrifugal force) refer to the non-tissue interaction interference forces generated by the instrument's own weight and accelerated motion in the sensor coordinate system. True interaction force data refers to the pure external force generated by the contact between the instrument tip and biological tissue after removing interference information from the raw interaction force data; it accurately characterizes the actual physical contact state. The coordinate transformation matrix is ​​a pre-calibrated rotation and translation matrix based on the sensor mounting posture and the geometric relationship of the instrument tip. It can be determined by the relative rotation and translation relationship between the force feedback sensor coordinate system and the surgical instrument tip coordinate system, and is used to achieve a precise mapping of the force vector from the sensor's local coordinate system to the unified surgical instrument tip coordinate system.

[0145] In this embodiment, based on the real-time motion state data of the surgical instrument at the moment of original interaction force data acquisition (such as joint acceleration, angular velocity, and posture information), combined with pre-calibrated instrument mass and center of mass parameters, the self-weight component and motion inertial force mixed in the original data can be calculated and subtracted to separate the pure external contact force as the real interaction force data. Furthermore, a preset coordinate transformation matrix can be used to accurately transform the calculated real interaction force data from the sensor local coordinate system to the surgical instrument end-effector coordinate system.

[0146] For example, sensors deployed on the joints of the robotic arm can be used to read motion state data; the self-weight and motion inertial force of the robotic arm end effector can be calculated and offset by a dynamic feedforward model to ensure that the real interaction force data is the pure interaction force between the instrument and the tissue.

[0147] For example, when a surgical robotic arm rapidly lifts an instrument upwards, the force feedback sensor detects an additional downward inertial force load. Without compensation, this load can easily be misinterpreted as pressure on the tissue. This spurious force data can be effectively eliminated by calculating the upward inertial force component in real time and subtracting it from the original interaction force data.

[0148] By using dynamic compensation and coordinate unification processing based on motion state data, the actual force situation between surgical instruments and tissues is accurately restored, eliminating the perceptual interference caused by the movement of the instruments themselves, ensuring precise synchronization of force data and visual images in the spatial reference system, and improving the accuracy and robustness of the model inferring force data from visual information.

[0149] To construct a high-quality, spatiotemporally aligned multimodal dataset and effectively eliminate temporal mismatch caused by differences in frequency or transmission delays between visual acquisition and force sensing, this scheme employs a time-window-based elastic pairing strategy to construct training samples. Specifically, for each group of interactive sample images, based on the capture time of the interactive sample images in the current group and a preset time window, the force selection range is determined, and interactive force sample data whose acquisition time falls within the force selection range are paired with the interactive sample images; training samples are determined based on each interactive sample image in the current group and its paired interactive sample images.

[0150] Here, the shooting time refers to the specific point in time when the visual acquisition module records each frame of the image. The preset time window is the sliding duration used to define the effective matching interval of force data. The force selection range is a specific time interval calculated based on the shooting time and the time window.

[0151] In this embodiment, the capture time of the current group of interactive sample images can be read, and a preset time window can be extended to the left and right sides from this center to define the force selection range. Data with capture times within this force selection range is retrieved from the continuous interactive force sample data stream. For example, if the capture time of a frame is 10 seconds and the preset time window is 4 seconds, then the force selection range is [8 seconds, 12 seconds]. In this case, all interactive force sample data whose capture times fall within this force selection range are considered candidate pairing data.

[0152] For situations where multiple interactive force data exist within the force selection range, all interactive force sample data whose acquisition time falls within the force selection range can be paired with interactive sample images. Alternatively, the interactive force sample data whose acquisition time is closest to the shooting time can be paired with the interactive sample image to minimize time deviation. Another approach is to perform statistical analysis (such as calculating the mean, variance, or time-weighted average) on all or part of the force data within the force selection range to generate force statistics that represent the force characteristics of that time period, and then pair the interactive sample image with this force statistics. This processing method not only adapts to the characteristics of sensors with different sampling frequencies but also effectively smooths high-frequency noise.

[0153] For example, high-frequency sampled interactive force data can be spatiotemporally aligned with low-frequency acquired interactive sample images. This is achieved by using the capture time of a specific interactive sample image... Based on this, select the force selection range. Interactive force sample data within. This represents the length of the preset time window. Through weighted averaging or linear interpolation algorithms, the interactive force sample data within the force selection range are processed to generate a value corresponding to the shooting time. Interactive force sample data (i.e., force vector labels) are paired with interactive sample images. This pairing process ensures that the visual deformation in the image is synchronized with the mechanical data in time.

[0154] Furthermore, each interactive sample image in the current group and its paired interactive sample image can be used as a training sample. Each group of interactive sample images and its paired force data together constitute an independent training sample. The number of training samples is consistent with the number of groups of interactive sample images. The interactive sample images within each group can contain a single frame or multiple frames.

[0155] The technical solution provided in this embodiment dynamically defines the force selection range for each group of interactive sample images based on their shooting time and a preset time window. Interactive force data whose acquisition time falls within this range are precisely paired with the images to construct training samples. This flexible alignment method based on time windows achieves precise temporal synchronization of information from different sources, such as visual data and force data, avoiding the loss of effective data due to strict zero latency and preventing incorrect matching caused by excessive time deviation. Through this efficient alignment of multi-source information, the accuracy and consistency of the physical causality between images and force states in the training samples are ensured, eliminating temporal noise interference. This allows the model to more reliably learn the intrinsic mapping law between visual deformation and mechanical feedback, thereby improving the model's generalization ability and prediction accuracy in inferring force signals from visual information.

[0156] S702. Input the interaction sample images from the training samples into the initial model and output the predicted interaction force data.

[0157] The initial model refers to the deep learning network to be trained, whose training objective is to predict the corresponding interaction force data (i.e., estimate the interaction force data, including the interaction force vector) based on the input image.

[0158] In this embodiment, interactive sample images from the training samples can be input into the initial model. Then, based on the feature extraction layer within the initial model, feature extraction is performed on multiple frames of interactive sample images to generate estimated image features for each frame. The feature fusion layer within the initial model performs cross-frame temporal fusion on the estimated image features of multiple frames to capture the spatiotemporal context information in dynamic interactive operations, obtaining estimated fused features representing motion trends. Based on the force data prediction layer within the initial model, regression analysis is performed on the estimated fused features to output estimated interactive force data.

[0159] To enhance the robustness of the initial model in complex and variable visual environments (such as different lighting, tissue texture differences, or instrument reflections) and to facilitate the decoupling of mechanically relevant geometric features from irrelevant style noise, interaction sample images, style sample images, and noise sample images are input into the style prediction network before the training samples are input into the initial model. The network outputs predicted style features, which are then used as input data for the initial model to output predicted interaction force data.

[0160] The style sample images are similar to the reference style image. The noise sample images are similar to the target noise image.

[0161] The current interaction sample image, the selected style sample image, and the generated noise sample image can be simultaneously input into the style prediction network. Internally, the network extracts the content structure of the interaction sample image through an encoder, guides texture mapping using the style sample image, and performs feature denoising by combining the noise sample image, ultimately outputting high-dimensional predicted style features. These predicted style features can then be used as input data for the initial model. The feature extraction layer within the initial model encodes these predicted style features, generating single-frame predicted image features; then, based on the feature fusion layer, cross-frame temporal fusion is performed on the multi-frame predicted image features to capture the spatiotemporal context information in dynamic interactions, resulting in predicted fused features rich in motion trends. Based on the force data prediction layer, regression analysis is performed on these fused features to output predicted interaction force data. This approach effectively eliminates the interference of imaging style differences under different interaction scenarios on force prediction.

[0162] There are two methods for training style prediction networks.

[0163] First, independent training optimization: The style prediction network can be trained independently based on the image reconstruction loss function, aiming to minimize the difference between the predicted style information and the expected style sample information, thereby improving the fidelity of style transfer. The image reconstruction loss function can be composed of a weighted sum of a content preservation loss function and a style transfer loss function. The content preservation loss function calculates the distance in the feature space between the original interactive image and the style image decoded from the predicted style features (or the interaction features and the predicted style features), constraining the style prediction network to avoid destroying key semantic structures such as instrument morphology and tissue boundaries when changing styles. The style transfer loss function calculates the difference between the stylized output image and the style sample image, constraining the style features output by the style prediction network, such as texture and color distribution, to match the reference image.

[0164] Second, end-to-end joint training: The style prediction network can be embedded into the entire force prediction process of the initial model, with the accuracy of the final estimated interaction force data as the global optimization objective. The parameters of the style prediction network and the initial model are adjusted synchronously through backpropagation gradients. This approach allows the style prediction network to learn the feature representations most conducive to force estimation, rather than simply pursuing visual similarity. Furthermore, this style prediction network can be constructed using the U-Net architecture, leveraging its skip connection characteristics to effectively preserve spatial detail information.

[0165] S703. Based on the target loss function, determine the total loss value according to the interaction force sample data and the estimated interaction force data in the training samples.

[0166] The objective loss function aims to constrain the estimated interaction force data to approximate the actual interaction force sample data. The total loss value can be used to quantify the combined deviation between the interaction force sample data and the estimated interaction force data in multidimensional physical quantities (such as six degrees of freedom, including three-dimensional force vectors and three-dimensional moment vectors). The smaller the total loss value, the higher the prediction accuracy of the model.

[0167] In this embodiment, the target loss function can be flexibly selected according to the specific task characteristics, including but not limited to: classic regression loss functions such as mean squared error (MSE), mean absolute error (MAE), Huber loss, and Log-Cosh loss, or cross-entropy loss, contrast loss, KL divergence, or structural similarity loss, etc., can also be used as needed.

[0168] Specifically, when using the mean squared error function, the numerical residuals between the six-dimensional true force / torque vectors in the interactive force sample data and the corresponding vectors in the predicted interactive force data can be calculated per degree of freedom. The residuals are then squared and summed or averaged to obtain the total loss value. This setting can apply a stronger quadratic penalty gradient to large prediction errors, not only driving the model to converge quickly to the neighborhood of the true data in the early stages of training and effectively suppressing the interference of outliers, but also guiding the network to accurately learn the complex nonlinear mapping relationship from visual features to continuous mechanical signals, thereby improving the stability and accuracy of force perception.

[0169] To ensure a high degree of match between the model's predictions and the actual labeled data, while also ensuring that the model output conforms to fundamental principles of mechanics and avoids predictions with small statistical errors but deviating from physical common sense, the objective loss function provided in this embodiment can include a data fitting loss function and a physical constraint function. This objective loss function is used to determine the total loss value based on the interaction force sample data in the training samples and the estimated interaction force data. Specific implementation methods include: Based on the data fitting loss function, error analysis is performed on the interaction force sample data and the estimated interaction force data in the training samples to obtain the data fitting loss value; based on the physical constraint function, physical rationality analysis is performed on the estimated interaction force data to obtain the physical constraint loss value; based on the data fitting loss value and the physical constraint loss value, the total loss value is determined.

[0170] The data fitting loss function (such as the mean squared error function) aims to constrain the numerical difference between the estimated interaction force data and the actual interaction force sample data; its data fitting loss value characterizes the magnitude of the prediction deviation. The physical constraint function can be a penalty term set based on mechanical safety requirements, used to evaluate the rationality and safety of the predicted interaction force data during force feedback. For example, when the estimated interaction force data deviates from safety conditions, the physical constraint loss value increases, and conversely, it approaches zero.

[0171] In practical implementation, the Euclidean distance between the predicted and actual data at each degree of freedom can be calculated based on the data fitting loss function to obtain the data fitting loss value, ensuring that the predicted values ​​are close to the actual measured values. A basic physical constraint function can be used to perform physical rationality checks on the predicted interaction force data: for example, for the contact characteristics between surgical instruments and soft tissue, a non-negative constraint can be set for the normal component of the contact force. If the model predicts a negative normal force, the physical constraint loss value is increased. It can also be checked whether the rate of change of force between consecutive frames meets the safety change conditions; if the rate of change is larger, the physical constraint loss value is increased to prevent high-frequency, severe jitter.

[0172] Furthermore, the static data fitting loss value and the dynamic physical constraint loss value can be added together by preset weighting coefficients to obtain the total loss value. This dual constraint method improves the model's force prediction accuracy while ensuring the safety, stability, and reliability of force feedback.

[0173] To ensure that the interactive force data output by the model not only numerically approximates the actual measured values ​​but also maintains consistency with motion safety in terms of temporal dynamic characteristics, thus avoiding risky force prediction results, physical constraint functions can include force continuity constraint functions and visual displacement constraint functions. The specific process of performing a physical rationality analysis on the estimated interactive force data based on these physical constraint functions to obtain the physical constraint loss value includes: Based on the force continuity constraint function, the force continuity loss value is determined according to the estimated rate of change of force in the time dimension of the interactive force data; based on the visual displacement constraint function, the visual displacement loss value is determined according to the visual motion characteristics of the interactive sample image and the estimated interactive force data; based on the force continuity loss value and the visual displacement loss value, the physical constraint loss value is determined.

[0174] The force continuity loss value aims to quantify the smoothness of the predicted interactive force data over time, penalizing abrupt changes and high-frequency jitter in the force data, thereby constraining the model to output a stable and secure smooth force data sequence. Visual motion features are quantitative information extracted from continuous interactive sample images, characterizing the interactive motion state between tissues and instruments. Specifically, these may include optical flow fields, pixel-level displacement vectors of organs and tissues, trajectory velocities of instrument ends, or local strain maps. The visual displacement loss value is used to constrain the magnitude and direction of the predicted force data to maintain physical consistency with the motion trends observed in the images. For example, in the absence of visual displacement (tissue at rest), large work forces should not be predicted, or the predicted force direction should be consistent with the direction of motion resistance.

[0175] Force continuity constraint functions can be used to constrain the smoothness of changes in estimated interactive force data between adjacent time frames, penalizing drastic changes in estimated interactive force data between adjacent frames to ensure force continuity. The force change rate reflects how quickly the force changes over time; it can be the first or higher-order derivative of the estimated interactive force data in the time dimension, describing the key factors in determining whether non-physical jumps occur in the interactive force data. For example, when the estimated interactive force data exhibits drastic oscillations or discontinuous jumps in the time series, the force continuity loss value increases significantly; conversely, when the interactive force data transitions smoothly, this loss value approaches zero.

[0176] In practical applications, predicted interaction force data from multiple consecutive frames can be extracted, and the difference norm of the force vector at adjacent time points can be calculated using a force continuity constraint function. If the interaction force data undergoes a significant abrupt change within a short period, a higher force continuity loss value is generated. The force continuity constraint function can be in the form of a second-order dynamic equation, considering not only the first-order rate of change of force but also the acceleration term of the predicted force, simulating the damping and inertial effects of the interaction between real surgical instruments and soft tissue. This allows for a more precise determination of the force continuity loss value, ensuring that the force curve is not only smooth but also conforms to dynamic response characteristics.

[0177] Visual displacement constraint functions can be used to constrain the consistency between the magnitude / direction of the estimated interactive force data and the tissue deformation or instrument displacement observed in the interactive sample image.

[0178] In practical applications, optical flow analysis or depth feature extraction is performed on the interactive sample image sequence to obtain pixel-level motion vectors or dense tissue strain fields on the tissue surface as high-precision visual motion features. A physical mapping relationship between the estimated interactive force data and visual motion features (such as optical flow amplitude or strain energy) is established using a visual displacement constraint function. For example, based on Hooke's law or a hyperelastic constitutive model, the virtual work generated by the estimated interactive force data is constrained to match the tissue strain energy observed in the image; or a monotonic mapping relationship between force and deformation is established (i.e., the greater the force, the greater the predicted tissue deformation or displacement should be). If there are cases where the predicted force is large but the optical flow shows no movement, the predicted force is zero but the optical flow shows severe deformation, or the two are inconsistent in spatial distribution and energy magnitude, the visual displacement loss value is increased.

[0179] Furthermore, the force continuity loss value and the visual displacement loss value are linearly superimposed or weighted and fused according to a preset ratio to determine the physical constraint loss value. For example, the target loss function can be expressed as: ,in, It is represented as a data fitting loss function (a function used to minimize the mean squared error (MSE) between the predictive power and the actual power). Represented as a force continuity constraint function; The weighting coefficient for the force continuity loss value; Represented as a visual displacement constraint function; The weighting coefficients are represented as the visual displacement loss values; This is expressed as the total loss value.

[0180] This application employs a force continuity constraint function to force the model to learn the temporal variation of force, suppressing abrupt changes in force prediction (such as jumps or jitters) caused by image noise or single-frame feature extraction errors. This ensures that the output interactive force data conforms to the continuous inertial characteristics of biological tissue forces. Furthermore, by utilizing a visual displacement constraint function, the visual motion features of the image (such as tissue deformation and instrument displacement) are correlated with the predicted interactive force data, enabling the model to learn the relationship between deformation and force. This avoids inconsistencies between deformation and force (e.g., the image shows deformation but the predicted force is zero, or vice versa), and enhances the model's robustness in complex occlusion or texture-deficient scenarios, allowing it to infer reasonable mechanical responses based on visual motion trends. By fusing these two losses into a physical constraint loss value, the model can output high-precision force feedback data with temporal smoothness and physical realism, improving the reliability and operational feel of the robot force feedback system.

[0181] In this embodiment, the interaction sample images from the training samples are input into the initial model. The uncertainty prediction layer in the initial model performs a confidence reliability assessment and outputs a predicted uncertainty attribute; the force data prediction layer in the initial model performs force prediction and outputs predicted interaction force data. This predicted uncertainty attribute is used to characterize the reliability of the confidence level of the predicted interaction force data output by the initial model. A larger predicted uncertainty attribute value indicates that the model believes the current scene is more complex or noisier (lower confidence); a smaller predicted uncertainty attribute value indicates that the model is more confident in the confidence level (higher confidence).

[0182] Next, to further quantify the confidence of the model's predictions, enabling the model not only to output accurate interaction force values ​​but also to adaptively assess the reliability (i.e., uncertainty) of the prediction results, thereby automatically reducing sensitivity to fitting errors and improving robustness in areas with high data noise or complex scenarios, a confidence calibration loss value can be determined based on the uncertainty loss function, according to the interaction force sample data, predicted interaction force data, and predicted uncertainty attributes in the training samples. Based on the confidence calibration loss value, the total loss value is updated to adjust the model parameters in the initial model by minimizing the updated total loss value.

[0183] The uncertainty loss function is used to assess the degree of matching between the estimated uncertainty attribute and the actual prediction error. It penalizes models with large prediction errors but low uncertainty (overconfidence) or small prediction errors but high uncertainty (overconservatism), teaching the model to accurately reflect its own predictive capabilities. The confidence calibration loss value characterizes the degree of bias in the model's confidence assessment. This value reflects not only the accuracy of the prediction but also the accuracy of the model's self-perception. The total loss value can be obtained by a weighted fusion of one or more of the following: data fitting loss (e.g., mean squared error), physical constraint loss, and confidence calibration loss.

[0184] In this embodiment, the uncertainty loss function can be used to evaluate the statistical distribution characteristics of the predicted residuals. At this point, the error between the interaction force sample data and the corresponding predicted interaction force data under multiple training samples can be calculated to obtain a residual sequence. The statistical distribution (such as quantiles, skewness, or kurtosis) of this residual sequence is compared with the expected distribution defined by the predicted uncertainty attribute. For example, the distance between these two distributions can be calculated (e.g., using Wasserstein distance or energy distance) to generate a confidence calibration loss value.

[0185] For example, when the residual sequence generated between the model's output predicted interaction force data and the actual interaction force sample data exhibits large and drastic fluctuations, it indicates that the actual error distribution has a large variance or a wide uncertainty range. However, if the uncertainty range constructed by the model's output predicted uncertainty attribute is smaller than the preset range (i.e., the model incorrectly assigns itself an excessively high confidence level), the predicted uncertainty distribution cannot cover the main probabilistic quality of the actual residuals. In this state, the actual error distribution and the predicted uncertainty distribution deviate positively in statistical characteristics, causing the distance metric to increase. At this point, the confidence level can be increased to calibrate the loss value, thereby generating a strong gradient penalty signal during backpropagation, causing the model to correct its uncertainty attribute estimation and expand the uncertainty range to match the actual error fluctuation range.

[0186] The confidence calibration loss value can be added to the total loss value according to preset weight parameters to form an updated total loss value. Alternatively, the confidence calibration loss value can replace part of the loss value (such as the data fitting loss value) in the total loss value according to preset weight parameters to form an updated total loss value. During backpropagation, this total loss value guides the model parameters to adjust both force prediction and uncertainty estimation simultaneously, enabling the model to learn to dynamically adjust the uncertainty attribute of its output according to the actual fluctuation range of the residuals, thereby achieving accurate confidence calibration.

[0187] Optionally, the uncertainty loss function can be a Gaussian negative log-likelihood loss function; based on the uncertainty loss function, the confidence calibration loss value is determined according to the interaction force sample data, the estimated interaction force data, and the estimated uncertainty attribute in the training samples, including: determining the error value between the interaction force sample data and the estimated interaction force data in the training samples according to the Gaussian negative log-likelihood loss function; and determining the confidence calibration loss value based on the estimated uncertainty attribute and the error value.

[0188] Specifically, confidence calibration calculations can be performed independently for each training sample. The difference between the interaction force sample data and the estimated interaction force data can be calculated to obtain the error value; based on the Gaussian negative log-likelihood loss function, the square of the error value is divided by the estimated uncertainty attribute to obtain the first penalty component; based on the logarithm of the estimated uncertainty attribute, the second penalty component is obtained; the first penalty component and the second penalty component are summed to obtain the confidence calibration loss value.

[0189] During training iterations, a bidirectional gradient mechanism can adaptively correct cognitive biases in the model: if the model outputs excessively low uncertainty for training samples with high error values ​​(i.e., overconfidence), the first penalty component will increase and generate a strong gradient signal, causing the model to raise its uncertainty estimate to cover the actual error and avoid blindly fitting noise; conversely, if the model outputs excessively high uncertainty for training samples with low error values ​​(i.e., overconservatism), the second penalty component will drive the model to lower its uncertainty estimate, thereby narrowing the prediction interval to improve accuracy. By minimizing this confidence level to calibrate the loss value, the model is constrained to establish a positive correlation between error magnitude and uncertainty value. That is, when the prediction is inaccurate and the uncertainty attribute is not increased accordingly, the total loss value will increase, thereby guiding backpropagation to simultaneously optimize the force data prediction layer and the uncertainty prediction layer. The updated initial model can not only output high-precision predicted interaction force data, but also simultaneously output uncertainty indicators that truly reflect the reliability of the current prediction, thereby improving the safety and robustness of force feedback interaction.

[0190] S704. With minimizing the total loss as the training objective, update the model parameters in the initial model to obtain the force data prediction model.

[0191] The model parameters refer to the set of learnable variables within the initial model, including but not limited to: connection weights, bias terms, and normalization parameters of each layer of the neural network.

[0192] In this embodiment, the model parameters in the initial model can be corrected based on the calculated total loss value using a gradient backpropagation algorithm. Specifically, the gradient of the total loss value relative to the model parameters of each layer of the initial model can be calculated. This gradient indicates the direction and magnitude of the model parameter adjustment to reduce prediction error. Optimization algorithms (such as Adam, SGD, etc.) can be used to update the model parameters in the initial model based on the gradient information. The model parameters are fine-tuned along the gradient in the opposite direction to minimize the total loss value in the next prediction. This process iterates repeatedly on training samples. Each time, one or a batch of training samples are input, the total loss value is calculated, the gradient is calculated through backpropagation, and the parameters are updated. As the number of training rounds increases, the accuracy of the model in estimating interactive force data and uncertainty improves, and the total loss value gradually converges to a stable low value. The model parameter updates can be stopped when the total loss value is less than a preset threshold, no longer decreases significantly, or reaches the maximum iteration period. At this point, the trained initial model is the completed force data prediction model.

[0193] To further ensure the stability of the training process and avoid conflicts in multi-objective optimization, a phased learning approach can be adopted, dynamically adjusting the weights of the loss function. Specifically, in the early stages of training, the weight coefficients of the physical constraint loss and confidence calibration loss can be reduced, prioritizing the minimization of the data fitting loss, allowing the model to quickly establish a basic mapping relationship from visual features to force data. After the initial model has a preliminary fitting ability, the weight coefficients of the physical constraint and confidence calibration losses can be gradually increased, guiding the model to further learn mechanical continuity, visual and force-displacement consistency, and uncertainty calibration capabilities while maintaining fitting accuracy. This progressive training method can suppress the oscillation phenomenon in the early stages of multi-task learning, ensuring a smooth transition of the model to the optimal state, thereby improving the generalization performance, prediction accuracy, and confidence reliability of the final model.

[0194] The technical solution provided in this embodiment constructs a paired training sample set containing interactive sample images and real interactive force data. Based on the training samples, the initial model is trained, enabling the initial model to learn the correlation between complex visual features (such as tissue texture and deformation details) and multidimensional mechanical signals. The training objective is to minimize the total loss between the interactive force sample data and the estimated interactive force data. The model parameters in the initial model are updated step by step to ensure the force data prediction accuracy of the force data prediction model. Based on the force data prediction model and real-time interactive images, the predicted interactive force data is accurately output, thereby providing force tactile feedback. This enables the robot or force feedback system to have force sensing capabilities without a direct force sensor configuration, while ensuring the accuracy, reliability, and safety of the tactile feedback.

[0195] As an optional embodiment of the above embodiments, specific application scenario examples are provided to enable those skilled in the art to further understand the technical solutions of the embodiments of this application. Specifically, please refer to the following detailed content.

[0196] See Figure 9 A force data prediction model can be pre-trained. This can be achieved by: during simulated surgical procedures, using a data acquisition device integrating force feedback sensors to collect multiple sets of interactive sample image sequences and corresponding instrument end force data; optionally, collecting at least 50-100 hours of effective interactive operation data (i.e., interactive sample images and interactive force data), which can cover various interactive actions such as grasping and clamping, traction and dissection, suturing (single / double arm), electrocoagulation / cutting, and no-load movement. Based on the interactive sample image sequences and interactive force data, a training sample set with paired interactive sample image sequences and interactive force sample data is obtained. Based on the mapping relationship between the interactive sample image sequences and interactive force sample data in the training sample set, an initial model (such as a machine learning model) is trained to obtain a trained force data prediction model.

[0197] Optionally, the data acquisition device can be a robot prototype with force feedback functionality. This prototype integrates a physical force feedback sensor at the wrist, lever, or connection point between the instrument and the robotic arm. The force feedback sensor can be, but is not limited to, a six-dimensional force / torque sensor, a strain gauge-based force sensing system, or a fiber optic grating sensor. Although the measurement principles of these sensors differ, this solution uses a data preprocessing step to uniformly convert the raw interactive force data collected by different sensors into standardized interactive force sample data (such as six-dimensional force / torque data) in the end-effector coordinate system (such as a Cartesian coordinate system) of the surgical instrument, along with high-precision acquisition time. The six-dimensional force / torque data may include… Force components and Torque component. This is expressed as a force along the X-axis. This is expressed as a force along the Y-axis. This is expressed as a force along the Z-axis. It is expressed as the torque of rotation about the X-axis; This is expressed as the torque of rotation about the Y-axis; It is expressed as the torque of rotation about the Z-axis.

[0198] It should be noted that although the driving source behind the sensor (robotic arm or human hand) does not directly generate the contact force between the instrument and tissue, the accelerated movement of the robotic arm during interactive operations will generate inertial force at the end of the instrument, and the instrument's own gravity will also act on the sensor. To ensure that the initial model learns the clean physical characteristics of the instrument-tissue interaction and avoids the model misinterpreting the instrument's acceleration as contact force, leading to erroneous force feedback illusions when the instrument is moving without load, a kinematic compensation module can be used to calculate and remove the inertial force and gravity components generated by the robotic arm's own movement from the original interaction force data. This eliminates the interference of the robotic arm's own movement on the measurement, obtaining pure instrument-tissue interaction force as interaction sample data. Specifically, based on gravity and inertia compensation algorithms, the motion state data of the robotic arm (i.e., kinematic parameters, such as joint acceleration) and the known mass parameters of the surgical instrument can be used to calculate and cancel the gravity and inertial force components in the sensor readings, thus obtaining interaction sample data that only reflects tissue interaction.

[0199] For example, see Figure 10 Data, including video streams and raw interaction force data, can be acquired by a surgical robot prototype equipped with force feedback sensors. Image frame extraction and enhancement processing can be performed on the video stream to obtain interaction sample images. Filtering, left-hand rotation, and coordinate transformation processing can be performed on the raw interaction force data to obtain interaction force sample images. An initial model is trained based on the pairing of interaction sample images and interaction force sample images.

[0200] While performing simulated interactive operations (such as clamping, traction, suturing, and cutting of ex vivo or simulated tissue models) on a surgical robot prototype, the system simultaneously acquires high-frame-rate surgical endoscope video streams and high-frequency instrument end-effector force data. Based on the frequency characteristics of physiological force data (typically <50Hz) and the limitations of endoscope hardware, the frame rate of the video stream is set to be no less than 30 FPS, and the sampling rate of the mechanical data is set to be no less than 200Hz. This sampling rate setting satisfies the requirement for capturing high-frequency mechanical changes (conforming to the Nyquist sampling theorem) and also achieves time alignment with the video frames through downsampling.

[0201] The simulated interactive operations can be performed directly by the operator on the prototype. The types of tissues manipulated cover common lesions in the target interactive application (such as tumors of different hardness, normal tissue, blood vessels, etc.), and can include tissue samples with abnormal embedded masses (such as calcifications, simulated tumors, suture knots). The surgical procedures performed are the standard operating procedures carried out in actual applications. This approach ensures that the data acquisition process is highly consistent with the real interactive workflow, and the acquired data naturally contains real biomechanical interaction characteristics. Moreover, no manual annotation of the mechanical data is required; it is automatically and synchronously acquired through sensors.

[0202] In this embodiment, the initial model can be a deep learning-based spatiotemporal prediction model, employing a hybrid architecture of 3D convolutional neural networks (3D-CNN) and Transformers (such as the Video Swin Transformer). A deep neural network incorporating spatial feature extraction and time series modeling can be constructed as the initial model. By minimizing the error between the predicted interaction force data output by the initial model and the interaction force sample data, and by optimizing the training in conjunction with physical constraints, a force data prediction model capable of mapping image sequences to continuous force vectors end-to-end is obtained. This force data prediction model is used to predict the interaction between instruments and tissues in real time based on visual appearance and deformation characteristics.

[0203] The initial model includes an input layer, a feature extraction layer, a feature fusion layer, and a force data prediction layer.

[0204] The input layer is used to receive interactive images; the feature extraction layer is used to extract image features (such as high-dimensional semantics or image information) from the interactive images; the feature fusion layer is used to perform cross-frame feature fusion on multiple image features and output fused features; the force data prediction layer is used to perform force prediction on the fused features to obtain predicted interactive force data.

[0205] During feature extraction, the required instrument pose, tissue texture, deformation, and other structures are often only a part of the entire image frame. When extracting information from the entire interactive image, direct feature extraction is often inefficient and limited in accuracy because the region of interest (ROI) is discretely distributed throughout the interactive image. Therefore, this application enables rapid annotation within the interactive image to help the model perform image feature extraction more accurately.

[0206] Optionally, fast annotation can be performed on the interactive image, including but not limited to processing of at least one of the following feature dimensions.

[0207] Instrument type identification: Focus on extracting the instrument outline. The difference in jaw area between different instruments affects the range of tissue clamping, which in turn determines the magnitude of force feedback (the larger the jaw area, the more tissue is clamped, and the stronger the force feedback).

[0208] Tissue deformation degree analysis: Focus on extracting tissue outline and edges. For similar tissues, the greater the deformation caused by clamping, the greater the antagonistic force feedback generated.

[0209] Tissue type identification: Image enhancement techniques such as increasing contrast and preserving gloss are used to highlight tissue texture. Different tissues have varying degrees of toughness (e.g., tendons have greater tension than fat), and under the same deformation, highly tough tissues will produce greater recovery force.

[0210] Analysis of occlusion relationship between instruments and tissues: Using depth information or layer separation technology, the relative position and occlusion level of instruments and tissues are identified, providing information for determining the direction of force vectors.

[0211] Given that predictive power data requires attention to multiple dimensions such as instrument contours, tissue deformation, material gloss, and occlusion relationships, traditional algorithms (such as edge extraction and contrast enhancement) struggle to simultaneously meet all requirements (e.g., preserving gloss while extracting edges), and manual frame-by-frame labeling is prohibitively expensive. Therefore, to improve image processing efficiency, this application pre-constructs a reference image set by manually labeling / combining traditional algorithms. This reference image set presents the ideal instrument contour, tissue contours / edges, image enhancement processes such as contrast enhancement and gloss preservation for the tissue region, instrument tissue occlusion / labeling as different regions based on occlusion relationships, and cropping or reducing sharpness for unprocessed non-critical (non-ROI) regions. These reference images (i.e., reference style images and reference sample images) represent the ideal feature representation desired for the power prediction task.

[0212] It should be noted that, due to the dynamic and ever-changing positions of instruments and tissues in a surgical setting, the goal of this approach is not to process real-time acquired images to be pixel-perfectly identical to the reference images. Instead, it utilizes the same / similar processing logic to ensure that the interactive images maintain a high degree of similarity to the reference image set in terms of feature distribution and / or visual style. In other words, by generating style-consistent enhanced images, the original image data is mapped to the feature space most easily understood by the model, thereby improving the accuracy and efficiency of force prediction.

[0213] During model training, interaction sample images from the training samples, style sample images from the reference image set, and noise sample images (such as Gaussian noise images) can be input into the style prediction network. The network outputs predicted style features, which are then used as input data for the initial model, outputting predicted interaction force data. The model parameters in the initial model are adjusted based on the predicted interaction force data and the interaction force sample data. This style feature prediction process transforms multiple tedious image processing steps into an image stylization process based on a neural network model, avoiding the cumbersome image preprocessing steps.

[0214] Finally, the trained initial model is used as a force data prediction model, which learns complex physical interactions such as machine kinematics and tissue material properties. Kinematic relationships of the implement (no Jacobian matrix calculation required): "Visual kinematics" are automatically learned through the spatiotemporal feature extraction module within the model. The displacement changes of the implement pixels in the image sequence (i.e., optical flow information) implicitly represent the motion vector of the implement in Cartesian space. The model, trained on a large amount of data, directly establishes an end-to-end mapping from "visual displacement / deformation" to "force." This approach allows the model to perceive the motion state of the implement without relying on externally input Jacobian matrices or joint angles, thereby predicting the corresponding interaction forces through vision alone, achieving decoupling from the robot hardware kinematics interface.

[0215] Tissue material properties: Without needing to preset parameters such as the elastic modulus and viscosity coefficient of the tissue, the model learns the visual deformation patterns of tissues with different hardness and disease states under stress (such as large deformation of soft tissue vs. small deformation of hard tumors), implicitly acquiring material property characteristics such as the elastic modulus of the tissue. This allows it to accurately predict the mechanical response under different tissue types. For common lesion and embedded block types that are sufficiently covered in the training set, the model can establish a robust visual-mechanical mapping and can even predict changes in mechanical properties through visual cues.

[0216] It should be noted that the object of identification in this application is the visual changes in biological tissues under stress. Whether it is clamping, compression, or traction, the force applied to tissue by the instrument will cause visible deformation, displacement, wrinkling, or changes in texture. Even if the instrument moves along the depth of the camera, the mechanical effects it produces on the tissue (such as tissue stretching or indentation) can usually be clearly captured in two-dimensional images. Similarly, when the instrument comes into contact with a hard embedded block inside the tissue, this interaction process will also produce unique, visually captureable deformation patterns on the tissue surface. A trained force data prediction model, by learning from a training sample set, can decode the corresponding interaction forces from these visual change patterns of tissues.

[0217] The initial model also includes an uncertainty prediction layer. This layer can be based on heteroscedastic uncertainty learning or evidence regression algorithms. During the training phase, a Gaussian negative log-likelihood loss function can be used to guide the initial model to reduce the prediction uncertainty attribute when the prediction is accurate, and to increase the uncertainty attribute when encountering samples that are difficult to fit (large errors). This enables the trained force data prediction model to analyze the current prediction confidence level, and in the application phase, only one forward propagation is needed to simultaneously obtain the predicted interaction force data at the current moment and the uncertainty attribute representing the prediction confidence level. This allows for the evaluation based on the uncertainty attribute to determine whether force feedback operations can be performed based on the predicted interaction force data, avoiding the risk of providing incorrect force perception under unknown conditions and improving the safety of force feedback.

[0218] Based on the above technical solutions, please refer to [link / reference]. Figure 9Force data prediction models can be deployed in surgical robots or as standalone force data prediction modules. The vision acquisition device integrated into the surgical robot can acquire interactive images reflecting the interaction between the target instrument and the target tissue. These interactive images (such as real-time image sequences) are input into the force data prediction model, which outputs predicted, continuous interactive force data (such as force / torque feedback data) and uncertainty attributes characterizing the prediction confidence level. Force feedback control commands (such as tactile signals) can be generated based on the predicted interactive force data and output to force feedback devices (such as the main control arm or operating handle of the doctor's console). The reliability of the predicted interactive force data can also be assessed, including: determining whether the uncertainty attribute meets preset reliability conditions. If the uncertainty attribute is below a preset safety threshold, it is considered reliable, and force feedback control commands are generated and sent to the force feedback device. If the uncertainty attribute is not below the preset safety threshold, it is considered unreliable, and if not, a safety strategy is activated, and safety processing operations are performed (including but not limited to cutting off force feedback, outputting zero force, or issuing audiovisual alarms).

[0219] A concrete example is provided: In surgical procedures, a trained force data prediction model can be used. Real-time endoscopic video streams from the surgical robot are continuously received. A sequence of real-time interactive images (e.g., N consecutive frames) is input into the deployed force data prediction model. The model performs forward inference, outputting an end-to-end predicted, continuous six-dimensional force / torque vector (i.e., predicted interactive force data). This process is completely independent of the surgical robot's kinematic model, joint sensor data, or real-time instrument pose estimation.

[0220] When the surgical scenario involves rare lesions not seen by the model, unknown types of embedded masses (whose visual patterns are not included in the training set), or drastic changes in lighting conditions, the model's internal recognition features will exhibit anomalies, leading to uncertainty in the force data prediction model's current prediction output. It will obviously increase.

[0221] Uncertainty threshold can be set When the reliability of the confidence level of the predicted interaction force data is assessed as unreliable (i.e., When the force feedback actuator is in operation, it may perform one or a combination of the following safety processing operations: The historical interaction force data from the previous moment is linearly decayed to zero force in a short period of time, avoiding the shock of sudden cut-off and smoothly returning control to the operator.

[0222] Maintain the output of historical interaction force data from the previous moment to prevent force value jumps.

[0223] The control panel displays warning messages indicating limited force feedback accuracy or unknown tissue type.

[0224] The advantage of this setup is that it ensures that when the model encounters unfamiliar interactive scenarios such as unknown embedded blocks, it will not provide the operator with incorrect force feedback. Instead, it returns control to the operator, who can rely on their vision and experience to make judgments, thereby ensuring operational safety.

[0225] It should be further noted that the force data prediction model provided in this application can take a two-dimensional endoscopic video stream as input without any robot-specific kinematic interface or hardware sensor, and therefore the force data prediction model can be packaged as an independent software module.

[0226] This software module can be directly deployed on surgical robot systems without hardware force sensing capabilities via software updates. It can also function as a standalone hardware component, connecting externally to the surgical robot. For example, it can connect to an imaging trolley, facilitating the transmission of force feedback control commands to the force feedback device. Furthermore, it can be provided via a cloud service platform, with a remote server performing the calculations and sending the force feedback control commands to the local robot. This reduces the cost and improves the efficiency of force feedback control performance upgrades.

[0227] Furthermore, the force feedback control function provided in this application is implemented by a software model, which can provide force feedback capability to the surgical robot system through networks, local deployment, or external means. This eliminates the need to install force sensors on the instruments, reducing consumable costs and avoiding issues such as sensor sterilization and wear.

[0228] Figure 11 This is a schematic diagram of a force feedback control device according to an embodiment of this application. The device can be configured in a master-slave surgical robot system. The surgical robot system includes a master control arm and at least one slave robotic arm controlled by the master control arm. The master control arm integrates a force feedback device, and the slave robotic arm carries a target instrument. Figure 11 As shown, the force feedback control device includes: an interactive image acquisition module 801, a predicted interactive force data determination module 802, and an instruction output module 803.

[0229] The interactive image acquisition module 801 is used to acquire interactive images reflecting the interaction state between the target instrument and the target tissue; the predicted interactive force data determination module 802 is used to input the interactive image into the force data prediction model and output predicted interactive force data; wherein the predicted interactive force data includes at least one interactive force vector; the command output module 803 is used to generate a force feedback control command based on the predicted interactive force data and send the force feedback control command to the force feedback device to drive the force feedback device to output a force on the main control arm that matches the interactive force vector.

[0230] Based on the above-mentioned device, optionally, the force data prediction model includes a feature extraction layer, a feature fusion layer, and a force data prediction layer; the predictive interactive force data determination module 802 includes: an image feature determination unit, used to perform feature extraction on multiple interactive images based on the feature extraction layer to obtain image features corresponding to the multiple interactive images; a fusion feature determination unit, used to perform cross-frame feature fusion on multiple image features based on the feature fusion layer to obtain fusion features characterizing the interaction process between the target device and the target tissue at multiple time points; and a predictive interactive force data determination unit, used to perform force prediction on the fusion features based on the force data prediction layer to obtain predicted interactive force data.

[0231] Based on the above-mentioned device, optionally, the feature fusion layer includes a temporal offset module and an attention module; the fusion feature determination unit includes: an enhanced feature sequence determination unit, used to offset the feature values ​​of some channels in multiple image features along the time dimension based on the temporal offset module to obtain an enhanced feature sequence; a spatiotemporal correlation weight determination unit, used to input the enhanced feature sequence to the attention module and output the spatiotemporal correlation weights between multiple frame features in the enhanced feature sequence; and a fusion feature determination unit, used to perform weighted fusion processing on the enhanced feature sequence according to the spatiotemporal correlation weights to obtain fused features.

[0232] Based on the above-mentioned device, optionally, the force data prediction model includes a style prediction network and a data prediction network. The device also includes: a prediction style information determination module, used to input the interactive image, the reference style image and the target noise image into the style prediction network to obtain prediction style information; and a data prediction module, used to call the data prediction network to perform prediction on the prediction style information to obtain prediction interactive force data.

[0233] Based on the above-mentioned device, optionally, the predicted style information determination module includes: a coding feature determination unit, used to determine a first coding feature corresponding to the target noise image, a second coding feature corresponding to the interactive image, and a third coding feature corresponding to the reference style image; a splicing feature determination unit, used to perform feature splicing on the coding sub-features at the same spatial position among the first coding feature, the second coding feature, and the third coding feature to obtain spliced ​​features; an attention feature determination unit, used to perform cross-attention processing on the first coding feature and the spliced ​​features to obtain attention features; and a predicted style information determination unit, used to perform denoising on the attention features to obtain predicted style information.

[0234] Based on the above-mentioned device, optionally, the force data prediction model also includes an uncertainty prediction layer; the device includes: an uncertainty attribute determination module, used to perform confidence reliability assessment on the fusion features based on the uncertainty prediction layer to obtain uncertainty attributes; wherein, the uncertainty attributes are used to characterize the reliability of the confidence of the predicted interactive force data; the fusion features are obtained based on the interactive image; and an instruction output module 803, used to generate a force feedback control instruction based on the predicted interactive force data when the uncertainty attributes meet the preset reliability conditions.

[0235] Optionally, based on the above-mentioned device, the device may further include: a safety processing module, used to perform at least one of the following safety processing operations when the uncertainty attribute does not meet the preset reliability conditions: generating a force feedback control command for maintaining or attenuating the force, sending the force feedback control command to the force feedback device to drive the force feedback device to output the corresponding force; and outputting a warning message in the display interface of the main control device where the main control arm is located; wherein the warning message includes operation guidance information or force feedback status description information.

[0236] Based on the above-mentioned device, an optional security processing module includes: The control command determination unit is used to generate a force feedback control command that decays to a preset threshold within a preset time period, using the historical interaction force data output at the previous moment as the initial value; or, using the historical interaction force data output at the previous moment as a constant value, it generates a force feedback control command to maintain the force unchanged. The previous time is the most recent time before the current time when the preset reliability conditions were met.

[0237] Optionally, based on the above-described apparatus, the apparatus may further include: The training sample determination module is used to acquire multiple training samples, including interaction sample images reflecting the interaction state between surgical instruments and tissues, and corresponding interaction force sample data. The predicted interaction force data determination module is used to input the interaction sample images from the training samples into the initial model and output the predicted interaction force data. The total loss value determination module is used to determine the total loss value based on the target loss function, the interaction force sample data from the training samples, and the predicted interaction force data. The parameter adjustment module is used to update the model parameters in the initial model with the goal of minimizing the total loss value, thereby obtaining the force data prediction model.

[0238] Based on the above-mentioned device, optionally, the training sample determination module includes: The data acquisition unit is used to acquire multiple sets of interaction sample images when surgical instruments interact with tissues, as well as the original interaction force data at the end of the surgical instruments; the data processing unit is used to perform correction and coordinate transformation on the original interaction force data to obtain interaction force sample data in the coordinate system of the end of the surgical instruments; the data pairing unit is used to pair each set of interaction sample images with interaction force sample data to form multiple training samples.

[0239] Based on the above-described apparatus, optionally, the data pairing unit includes: The pairing subunit is used to determine the force selection range for each group of interactive sample images based on the shooting time and preset time window of the interactive sample images in the current group, and to pair the interactive force sample data whose acquisition time is within the force selection range with the interactive sample images; the training sample determination subunit is used to determine the training samples based on each interactive sample image in the current group and its paired interactive sample images.

[0240] Optionally, based on the above-described apparatus, the apparatus may further include: The style feature prediction unit is used to input interaction sample images, style sample images, and noise sample images into the style prediction network and output predicted style features. The predicted style features are used as input data for the initial model to output predicted interaction force data. The style prediction network is trained based on the image reconstruction loss function, which includes a content preservation loss function and a style transfer loss function. The content preservation loss function is used to constrain the stylized output image to retain the semantic structure information of the input interaction sample image. The style transfer loss function is used to constrain the texture features of the stylized output image to match the style features of the input reference style image.

[0241] Based on the aforementioned device, optionally, the target loss function includes a data fitting loss function and a physical constraint function. The total loss value determination module includes: a data fitting loss value determination unit, used to perform error analysis on the interaction force sample data and the estimated interaction force data in the training samples based on the data fitting loss function, to obtain the data fitting loss value; a physical constraint loss value determination unit, used to perform physical rationality analysis on the estimated interaction force data based on the physical constraint function, to obtain the physical constraint loss value; and a total loss value determination unit, used to determine the total loss value based on the data fitting loss value and the physical constraint loss value.

[0242] The force feedback control device provided in this application embodiment can execute the force feedback control method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of the execution method.

[0243] Figure 12This is a schematic diagram of the structure of an electronic device implementing the force feedback control method of the embodiments of this application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the application described and / or claimed herein. In specific application scenarios, the implementation of this electronic device is diverse: it can be a stand-alone computing device, integrated into a surgical robot system, or used as a computing device in a multi-robot collaborative vehicle group.

[0244] like Figure 12 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory 12 or a random access memory 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the read-only memory 12 or a computer program loaded from storage unit 18 into the random access memory 13. The random access memory 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, read-only memory 12, and random access memory 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0245] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks. Processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, central processing unit (CPU), graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as force feedback control methods.

[0246] In some embodiments, the force feedback control method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via read-only memory 12 and / or communication unit 19. When the computer program is loaded into random access memory 13 and executed by processor 11, one or more steps of the force feedback control method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the force feedback control method by any other suitable means (e.g., by means of firmware).

[0247] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0248] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine as a standalone software package, or entirely on a remote machine or server. In the context of this application, a computer-readable storage medium may be a tangible medium that may contain or store computer programs for use by or in conjunction with an instruction execution system, apparatus, or device. Computer-readable storage media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with a user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0249] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet. The computing system can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having client-server relationships with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system, addressing the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0250] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 19, or installed from storage unit 18, or installed from read-only memory 12. When the computer program is executed by processor 11, it performs the functions defined in the methods of embodiments of this application.

[0251] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the force feedback control method provided in any embodiment of this application.

[0252] In the implementation of the computer program product, computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0253] It should be understood that the various processes shown above can be used, with steps rearranged, added, or deleted. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this application can be achieved, and this is not limited herein. The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A force feedback control method applied to a master-slave surgical robot system, the surgical robot system comprising a master control arm and at least one slave robotic arm controlled by the master control arm, wherein the master control arm integrates a force feedback device, and the slave robotic arm carries a target instrument, characterized in that, The method includes: Acquire interactive images that reflect the interaction between the target instrument and the target tissue; The interactive image is input into the force data prediction model, and the predicted interactive force data is output; wherein, the predicted interactive force data includes at least one interactive force vector; Based on the predicted interactive force data, a force feedback control command is generated and sent to the force feedback device to drive the force feedback device to output a force on the main control arm that matches the interactive force vector.

2. The method according to claim 1, characterized in that, The force data prediction model includes a feature extraction layer, a feature fusion layer, and a force data prediction layer. The step of inputting the interactive image into the force data prediction model and outputting predicted interactive force data includes: Based on the feature extraction layer, feature extraction is performed on multiple interactive images to obtain image features corresponding to the multiple interactive images; Based on the feature fusion layer, cross-frame feature fusion is performed on multiple image features to obtain fused features characterizing the interaction process between the target device and the target tissue at multiple time points; Based on the force data prediction layer, force prediction is performed on the fused features to obtain predicted interactive force data.

3. The method according to claim 2, characterized in that, The feature fusion layer includes a temporal offset module and an attention module; the cross-frame feature fusion of multiple image features based on the feature fusion layer to obtain fused features characterizing the interaction process between the target device and the target tissue at multiple time points includes: Based on the time-series offset module, the feature values ​​of some channels in multiple image features are offset along the time dimension to obtain an enhanced feature sequence; The enhanced feature sequence is input into the attention module, and the spatiotemporal correlation weights between multiple frame features in the enhanced feature sequence are output. The enhanced feature sequence is weighted and fused according to the spatiotemporal correlation weights to obtain the fused features.

4. The method according to any one of claims 1 to 3, characterized in that, The force data prediction model includes a style prediction network and a data prediction network, and the method further includes: The interactive image, the reference style image, and the target noise image are input into the style prediction network to obtain predicted style information; The data prediction network is invoked to perform prediction on the prediction style information to obtain the prediction interaction force data.

5. The method according to claim 4, characterized in that, The step of inputting the interactive image, the reference style image, and the target noise image into the style prediction network and outputting predicted style information includes: A first coding feature corresponding to the target noise image, a second coding feature corresponding to the interactive image, and a third coding feature corresponding to the reference style image are determined respectively; Feature concatenation is performed on the coded sub-features at the same spatial position in the first coded feature, the second coded feature, and the third coded feature to obtain the concatenated feature; Perform cross-attention processing on the first encoded feature and the concatenated feature to obtain attention features; The attention features are denoised to obtain the predicted style information.

6. The method according to any one of claims 1 to 3, characterized in that, The force data prediction model further includes an uncertainty prediction layer; the method further includes: The uncertainty prediction layer performs a confidence reliability assessment on the fused features to obtain uncertainty attributes; wherein, the uncertainty attributes are used to characterize the reliability of the confidence of the predicted interaction force data; the fused features are obtained based on the interaction image; When the uncertainty attribute meets the preset reliability condition, a force feedback control command is generated based on the predicted interactive force data.

7. The method according to claim 6, characterized in that, The method further includes: When the uncertainty attribute does not meet the preset reliability condition, at least one of the following security processing operations is performed: Generate a force feedback control command for maintaining or attenuating the force, and send the force feedback control command to the force feedback device to drive the force feedback device to output the corresponding force; The warning message is output on the display interface of the main control device where the main control arm is located. The warning information includes operation guidance information or force feedback status description information.

8. The method according to any one of claims 1 to 7, characterized in that, The force data prediction model was determined in the following manner: Multiple training samples are acquired; wherein, the training samples include interaction sample images reflecting the interaction state between surgical instruments and tissues, and corresponding interaction force sample data; The interaction sample images from the training samples are input into the initial model, and the estimated interaction force data is output. Based on the target loss function, the total loss value is determined according to the interaction force sample data in the training samples and the estimated interaction force data; The model parameters in the initial model are updated with the goal of minimizing the total loss value to obtain the force data prediction model.

9. The method according to claim 8, characterized in that, Before inputting the interactive sample images from the training samples into the initial model, the method further includes: The interaction sample image, style sample image, and noise sample image are input into the style prediction network, and the estimated style features are output. The estimated style features are used as the input data of the initial model to output the estimated interaction force data. The style prediction network is trained based on an image reconstruction loss function, which includes a content preservation loss function and a style transfer loss function. The content preservation loss function is used to constrain the stylized output image to retain the semantic structure information of the input interactive sample image; the style transfer loss function is used to constrain the texture features of the stylized output image to match the style features of the input reference style image.

10. A force feedback control device configured in a master-slave surgical robot system, the surgical robot system comprising a master control arm and at least one slave robotic arm controlled by the master control arm, wherein the master control arm integrates a force feedback device, and the slave robotic arm carries a target instrument; characterized in that, The device includes: The interactive image acquisition module is used to acquire interactive images that reflect the interaction state between the target instrument and the target tissue. The predictive interaction force data determination module is used to input the interaction image into the force data prediction model and output the predicted interaction force data; wherein the predicted interaction force data includes at least one interaction force vector. The instruction output module is used to generate a force feedback control instruction based on the predicted interactive force data, and send the force feedback control instruction to the force feedback device to drive the force feedback device to output a force on the main control arm that matches the interactive force vector.

11. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to said at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the force feedback control method according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the force feedback control method according to any one of claims 1-9.

13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the force feedback control method as described in any one of claims 1-9.