Image processing method and device, equipment and medium
By acquiring the pose information of the endoscope and surgical instruments, calculating the target pose change, and automatically correcting the depth image, the problem of large depth image error within the target cavity is solved, and efficient and accurate depth image generation is achieved.
Patent Information
- Application Number
- CN202511726583.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-02-17
AI Technical Summary
In existing technologies, the prediction of depth images within the target cavity has a large error, and relying on manual correction is inefficient, resulting in insufficient accuracy and reliability of depth images.
By acquiring the pose information of the endoscope's end-effector camera and the surgical instrument's end-effector, unifying them into a common coordinate system, calculating the target pose change, determining the correction coefficient, automatically correcting the initial depth image, and generating an accurate target depth image.
It improves the accuracy and reliability of depth images, avoids manual intervention, and enhances correction efficiency and consistency.
Smart Images

Figure CN121544566A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and in particular to an image processing method, apparatus, device and medium. Background Technology
[0002] In the medical field, three-dimensional models of target cavities (such as joint cavities like the knee joint cavity and ankle joint cavity, or body cavities like the thoracic cavity) can provide precise visualization of intracavitary anatomical structures for related minimally invasive surgeries, assisting doctors in precise operations and reducing surgical risks. The construction of three-dimensional models of target cavities is inseparable from the acquisition of depth images within the target cavity. Therefore, accurately acquiring depth images within the target cavity is crucial for constructing accurate three-dimensional models.
[0003] In related technologies, depth images inside the target cavity are often predicted directly using depth models. However, due to the relatively small size and complex environment inside the target cavity, the predicted depth images have large errors and can only be corrected manually. Manual correction relies on human experience and is inefficient, thus reducing the accuracy and reliability of the depth images. Summary of the Invention
[0004] The main objective of this disclosure is to provide an image processing method, apparatus, device, and medium that can improve the accuracy and reliability of depth images.
[0005] To achieve the above objectives, a first aspect of this disclosure provides an image processing method, comprising: An initial depth image of the target cavity is acquired, and the initial depth corresponding to the end of the surgical instrument is obtained from the initial depth image. The initial depth image is generated by taking an initial image of the target cavity through the end camera of the endoscope after the end of the instrument is moved to the target surface corresponding to the target cavity, and the depth of the target cavity is predicted based on the initial image. Acquire the first pose information of the end-effector camera and the second pose information of the instrument end, and acquire the third pose information of the first preset position on the endoscope and the fourth pose information of the second preset position on the surgical instrument. When the third pose information and the fourth pose information belong to a common coordinate system, the target pose change between the end camera and the end device is jointly determined based on the first pose information, the second pose information, the third pose information and the fourth pose information. The actual depth between the end-effector camera and the instrument end is determined based on the target pose change, and the correction coefficient of the initial depth image is determined based on the actual depth and the initial depth. The initial depth image is corrected based on the correction coefficient to obtain the corrected target depth image.
[0006] In some embodiments, acquiring the first pose information of the end-effector camera and the second pose information of the instrument end, and acquiring the third pose information of the first preset position on the endoscope and the fourth pose information of the second preset position on the surgical instrument, includes: The pose information of the first positioning mark pre-set on the end camera is identified to obtain the first pose information; The pose information of the second positioning mark pre-set on the end of the instrument is identified to obtain the second pose information; The pose information of the first tracker, which is pre-set at the first preset position of the endoscope, is identified to obtain the third pose information; The fourth pose information is obtained by identifying the pose information of the second tracker pre-set at the second preset position of the surgical instrument.
[0007] In some embodiments, jointly determining the target pose change between the end-effector and the instrument end based on the first pose information, the second pose information, the third pose information, and the fourth pose information includes: Based on the difference between the first pose information and the third pose information, the first pose change between the end camera and the first preset position is obtained; Based on the difference between the second pose information and the fourth pose information, the second pose change amount between the instrument end and the second preset position is obtained; Based on the difference between the third pose information and the fourth pose information, the change in the third pose between the first preset position and the second preset position is obtained; The target pose change between the end camera and the instrument end is calculated based on the first pose change, the second pose change, and the third pose change.
[0008] In some embodiments, acquiring the initial depth image within the target cavity includes: After the instrument tip is moved to the corresponding target surface inside the target cavity, an initial image is obtained by taking a picture inside the target cavity through the endoscope's end camera; The initial image is input into a pre-trained large depth model for monocular depth prediction, and the initial depth image inside the target cavity is output.
[0009] In some embodiments, the step of correcting the initial depth image based on the correction coefficient to obtain the corrected target depth image includes: When the instrument end moves to another position on the target surface, the target pose change between the end camera and the instrument end is updated again based on the second pose information of the moved instrument end, and the correction coefficient is updated based on the updated target pose change. Based on the correction coefficients corresponding to different positions of the instrument tip on the target surface, the target correction coefficients of the initial depth image are jointly determined. The initial depth image is corrected based on the target correction coefficient to obtain the corrected target depth image.
[0010] In some embodiments, after updating the correction coefficients based on the updated target pose change, the image processing method further includes: Based on the correction coefficients corresponding to different positions of the instrument tip on the target surface, the correction coefficient matrix of the initial depth image is jointly determined; Based on the correction coefficient matrix, corresponding correction processing is performed on different positions of the initial depth image to obtain the corrected target depth image.
[0011] In some embodiments, after obtaining the corrected target depth image, the image processing method further includes: When the first pose information indicates that the end-effector camera is moving, the target depth images after correction of the end-effector camera at different positions are continuously acquired during the movement of the end-effector camera. A three-dimensional model of the target cavity is constructed based on multiple consecutive corrected target depth images to obtain a three-dimensional intracavitary model of the target cavity.
[0012] To achieve the above objectives, a second aspect of this disclosure provides an image processing apparatus, comprising: An image acquisition module is used to acquire an initial depth image within the target cavity and to acquire the initial depth corresponding to the end of the surgical instrument from the initial depth image. The initial depth image is generated by taking an initial image within the target cavity through the end-effector camera after the end of the instrument has moved to the target surface corresponding to the target cavity, and by predicting the depth within the target cavity based on the initial image. The pose acquisition module is used to acquire the first pose information of the end-effector camera and the second pose information of the instrument end, and to acquire the third pose information of the first preset position on the endoscope and the fourth pose information of the second preset position on the surgical instrument. The pose relationship determination module is used to jointly determine the target pose change between the end camera and the instrument end when the third pose information and the fourth pose information belong to a common coordinate system. The correction coefficient determination module is used to determine the actual depth between the end camera and the instrument end based on the target pose change, and to determine the correction coefficient of the initial depth image based on the actual depth and the initial depth. The image processing module is used to perform correction processing on the initial depth image based on the correction coefficient to obtain the corrected target depth image.
[0013] To achieve the above objectives, a third aspect of this disclosure provides an electronic device, the electronic device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the image processing method described in the first aspect embodiment.
[0014] To achieve the above objectives, a fourth aspect of this disclosure provides a storage medium, which is a computer-readable storage medium storing a computer program that, when executed by a processor, implements the image processing method described in the first aspect of the embodiment.
[0015] This embodiment of the present disclosure, by executing an image processing method, can acquire an initial depth image within a target cavity, and obtain the initial depth corresponding to the end-effector tip of a surgical instrument from the initial depth image. The initial depth image is generated by capturing an initial image within the target cavity using an end-effector camera after the end-effector tip has moved to the target surface corresponding to the target cavity, and by predicting the depth within the target cavity based on the initial image. The method acquires first pose information of the end-effector camera and second pose information of the end-effector tip, as well as third pose information of a first preset position on the endoscope and fourth pose information of a second preset position on the surgical instrument. When the third and fourth pose information belong to a common coordinate system, the target pose change between the end-effector camera and the end-effector is jointly determined based on the first, second, third, and fourth pose information. The actual depth between the end-effector camera and the end-effector tip is determined based on the target pose change, and a correction coefficient for the initial depth image is determined based on the actual depth and the initial depth. The initial depth image is then corrected based on the correction coefficient to obtain a corrected target depth image.
[0016] Therefore, in this embodiment, after acquiring an initial depth image within the target cavity, the pose information of the endoscope's end-effector camera and the end of the surgical instrument, as well as the pose information of preset positions on the endoscope and the surgical instrument, are unified into a common coordinate system. The target pose change between the end-effector camera and the end of the instrument is then jointly calculated, thereby deriving the actual depth between them. By comparing the actual depth with the initial depth of the end of the instrument in the initial depth image, a correction coefficient is determined. This coefficient reflects the systematic error in depth prediction. Finally, the initial depth image is corrected based on the correction coefficient to generate an accurate target depth image. This process directly corrects depth errors using high-precision pose data, avoiding manual intervention based on experience. This not only improves correction efficiency but also significantly enhances the accuracy and reliability of the depth image. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of an application environment for the image processing method provided in this embodiment. Figure 2 This is a schematic flowchart of the image processing method provided in the embodiments of this disclosure; Figure 3 yes Figure 2 A flowchart further includes step S102; Figure 4 yes Figure 2 A flowchart further includes step S103; Figure 5 This is a schematic diagram of the intra-articular scene of the knee joint provided in the embodiments of this disclosure; Figure 6 yes Figure 2 A flowchart further includes step S101; Figure 7 yes Figure 2 A flowchart further includes step S105; Figure 8 yes Figure 7 A flowchart illustrating the further steps following step S501; Figure 9 yes Figure 2 A flowchart illustrating the further steps following step S105; Figure 10 This is a schematic diagram of the functional modules of the image processing apparatus provided in the embodiments of this disclosure; Figure 11 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this disclosure. Detailed Implementation
[0018] To enable those skilled in the art to better understand the solutions disclosed herein, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0019] It is understood that in the specific embodiments of this disclosure, various depth images, various pose information and related data are involved. When the above embodiments of this disclosure are applied to specific products or technologies, permission or consent from the subject is required, and the collection, use and processing of related data must comply with relevant laws, regulations and standards.
[0020] Furthermore, when the embodiments of this disclosure need to retrieve various depth images, various pose information and related data, they will obtain separate permission or separate consent for various depth images, various pose information and related data through pop-up windows or jumps to a confirmation page. After clearly obtaining separate permission or separate consent for various depth images, various pose information and related data, they will then obtain the necessary various depth images, various pose information and related data for enabling the embodiments of this disclosure to operate normally.
[0021] In this disclosure, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0022] Before providing a further detailed description of the embodiments of this disclosure, the terms and concepts used in these embodiments are explained, and they are subject to the following interpretations: Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.
[0023] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0024] In the medical field, three-dimensional models of target cavities (such as joint cavities like the knee joint cavity and ankle joint cavity, or body cavities like the thoracic cavity) can provide precise visualization of intracavitary anatomical structures for related minimally invasive surgeries, assisting doctors in precise operations and reducing surgical risks. The construction of three-dimensional models of target cavities is inseparable from the acquisition of depth images within the target cavity. Therefore, accurately acquiring depth images within the target cavity is crucial for constructing accurate three-dimensional models.
[0025] In related technologies, depth images inside the target cavity are often predicted directly using depth models. However, due to the relatively small size and complex environment inside the target cavity, the predicted depth images have large errors and can only be corrected manually. Manual correction relies on human experience and is inefficient, thus reducing the accuracy and reliability of the depth images.
[0026] In order to solve the above problems, this disclosure provides an image processing method, apparatus, device, and medium that can improve the accuracy and reliability of depth images.
[0027] Please see Figure 1 , Figure 1 A schematic diagram of the scene in which the image processing method provided in this embodiment of the disclosure is implemented includes: a terminal 11 and a server 12.
[0028] For example, server 12 can obtain an initial depth image of the target cavity from terminal 11, and obtain the initial depth corresponding to the end of the surgical instrument from the initial depth image. The initial depth image is generated by taking an initial image in the target cavity through the end-effector camera after the end of the instrument moves to the target surface corresponding to the target cavity, and by predicting the depth of the target cavity based on the initial image. Server 12 can also obtain the first pose information of the end-effector camera and the second pose information of the end of the instrument, and obtain the third pose information of the first preset position on the endoscope and the fourth pose information of the second preset position on the surgical instrument. When the third pose information and the fourth pose information belong to a common coordinate system, the target pose change between the end-effector camera and the end of the instrument is jointly determined based on the first pose information, the second pose information, the third pose information and the fourth pose information. The actual depth between the end-effector camera and the end of the instrument is determined based on the target pose change, and the correction coefficient of the initial depth image is determined based on the actual depth and the initial depth. The initial depth image is corrected based on the correction coefficient to obtain the corrected target depth image. Finally, server 12 can also send the target depth image to terminal 11.
[0029] Terminal 11 can be a mobile phone, computer, smart voice interaction device, smart wearable device, smart home appliance, vehicle terminal, etc., but is not limited to these. Terminal 11 can also execute image processing methods independently. Terminal 11 and server 12 can be directly or indirectly connected through wired or wireless communication, and this embodiment of the present disclosure does not impose any limitations.
[0030] Server 12 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Additionally, server 12 can also be a node server in a blockchain network.
[0031] It should be noted that, Figure 1 The schematic diagram of the implementation environment shown is merely an example. The scenarios described in this disclosure are intended to more clearly illustrate the technical solutions of this disclosure and do not constitute a limitation on the technical solutions provided in this disclosure. As those skilled in the art will know, with the evolution of technology and the emergence of new business scenarios, the technical solutions provided in this disclosure are also applicable to similar technical problems.
[0032] Please see Figure 2 , Figure 2This is a flowchart illustrating an image processing method provided in an embodiment of this disclosure. This image processing method can be applied to the server in the above embodiments, or can be jointly executed by a terminal and a server. The image processing method includes steps S101 to S105: Step S101: Obtain an initial depth image of the target cavity, and obtain the initial depth corresponding to the end of the surgical instrument from the initial depth image; The initial depth image is generated by taking an initial image inside the target cavity through the end-eye camera after the instrument tip moves to the target surface inside the target cavity, and the depth inside the target cavity is predicted based on the initial image. Step S102: Obtain the first pose information of the end-effector camera and the second pose information of the instrument end, and obtain the third pose information of the first preset position on the endoscope and the fourth pose information of the second preset position on the surgical instrument. Step S103: When the third pose information and the fourth pose information belong to the same coordinate system, the target pose change between the end camera and the instrument end is jointly determined based on the first pose information, the second pose information, the third pose information and the fourth pose information. Step S104: Determine the actual depth between the end-effector camera and the instrument end based on the target pose change, and determine the correction coefficient of the initial depth image based on the actual depth and the initial depth. Step S105: Correct the initial depth image based on the correction coefficient to obtain the corrected target depth image.
[0033] Regarding step S101 above, this embodiment of the disclosure can first acquire an initial depth image within the target cavity. This initial depth image can be captured by using an endoscope's end-effector camera after the surgical instrument's tip has been moved to and maintained contact with a target surface (such as a bone surface) within the target cavity. This initial image is a two-dimensional RGB or grayscale image, showing the anatomical structures and the instrument tip within the target cavity. Subsequently, depth estimation is performed on the initial image using depth prediction to generate the initial depth image. The initial depth image is a depth map, where each pixel value represents the estimated depth (relative depth value) of that point relative to the end-effector camera. Next, this embodiment of the disclosure can extract the initial depth value of the corresponding pixel position of the instrument tip from the initial depth image. The position of the instrument tip in the image can be automatically located using color recognition (e.g., using high-contrast color markers) or feature detection algorithms.
[0034] The initial depth image refers to the depth map generated from the initial image captured by the endoscope through depth prediction. It provides a relative depth estimate for each pixel but lacks an absolute scale and may contain errors. Depth prediction is the process of learning the geometric features of the scene using a pre-trained large depth model or other depth prediction models to infer the relative depth of each pixel from the two-dimensional initial image. The end-effector camera is an imaging module integrated at the end of the endoscope, used to capture real-time images of the target cavity. Its coordinate system z-axis is defined as the camera orientation perpendicular to the imaging plane. The instrument tip refers to the distal part of the surgical instrument. The target surface refers to the anatomical surface within the target cavity, such as the femoral condyle or tibial plateau in the knee joint cavity, which is the physical surface that the instrument tip contacts.
[0035] It should be noted that the embodiments of this disclosure require obtaining an initial depth reference based on image depth prediction to provide a basis for subsequent correction. Due to the confined and complex environment inside the target cavity, the depth image of the depth prediction often suffers from accumulated errors and scale uncertainties. However, the embodiments of this disclosure provide a measurable physical reference point by having the instrument tip contact the target surface, allowing the initial depth value to be compared with the actual depth value, thereby initiating the correction process.
[0036] Regarding step S102 above, this embodiment of the present disclosure can track the spatial position and attitude of the end-effector camera using a preset positioning system or positioning method, and record its pose data in a preset coordinate system, which is the first pose information. Similarly, the spatial position and attitude of the instrument end-effector can also be tracked using a preset positioning system or positioning method, and its pose data in a preset coordinate system can be recorded, which is the second pose information. Likewise, the third pose information is obtained by installing a corresponding positioning tracker at a first preset position on the endoscope, collecting the spatial position and attitude of the marker point using a positioning system or positioning method, and recording its pose data in a preset coordinate system. The first preset position is a fixed point on the endoscope for installing the positioning tracker, used to transmit the spatial pose of that position on the endoscope. The fourth pose information is obtained by acquiring the spatial position and orientation of the marker point through a positioning system or positioning method after installing the corresponding positioning tracker at the second preset position of the surgical instrument, and recording its pose data in the preset coordinate system. The second preset position is a fixed point on the surgical instrument that is preset for installing the positioning tracker and is used to transmit the spatial pose of that position on the surgical instrument.
[0037] In this context, the preset coordinate system refers to a coordinate system defined by a preset positioning system or method. The preset positioning system or method can be an optical tracking or magnetic tracking system or method. For example, in optical tracking, when using an optical navigator for pose tracking, the optical navigator obtains pose information in the preset coordinate system by tracking optical features at the end-effector camera, the end effector, the first preset position, and the second preset position. Similarly, in magnetic tracking, when using a magnetic navigator for pose tracking, the magnetic navigator obtains pose information in the preset coordinate system by tracking magnetic features at the end-effector camera, the end effector, the first preset position, and the second preset position. This embodiment uses optical tracking as an example for illustration, but it is equally applicable to magnetic tracking scenarios, and this embodiment does not impose specific limitations on this.
[0038] It should be noted that the reason this embodiment requires obtaining the aforementioned pose information is because it requires relative relationships, not absolute positions. During surgery, both the endoscope (camera) and surgical instruments are constantly moving and rotating. To improve the accuracy of depth prediction, this embodiment requires not the absolute position of the instrument tip, but its three-dimensional coordinates in the camera coordinate system, particularly its Z-axis component (i.e., depth). For example, if the tracking navigator measures the camera position as (10, 20, 30) and the direction as east, and the instrument tip position as (12, 22, 28), these two absolute positions alone are insufficient to directly determine what the instrument tip looks like in the camera's "eye." This is because if the camera is tilted, the instrument tip might be directly in front of the camera; if the camera is vertically downward, the instrument tip might be at the edge of the image. Therefore, introducing pose information can improve the appearance of the instrument tip in the camera's "eye," which is helpful for the final three-dimensional model construction.
[0039] Regarding step S103 above, this embodiment of the disclosure needs to ensure that the third pose information and the fourth pose information are in the same coordinate system. When the third pose information and the fourth pose information are in the same coordinate system, based on the principle of coordinate transformation matrix, the first pose information, the second pose information, the third pose information and the fourth pose information can be combined, and a spatial pose association between the end-effector camera and the end-effector can be established through matrix multiplication. Through the above calculation, a comprehensive parameter containing relative position and relative attitude is obtained, which is the target pose change between the end-effector camera and the end-effector.
[0040] Regarding step S104 above, this embodiment of the disclosure can extract the position component from the target pose change and calculate its projection on the z-axis of the camera coordinate system to obtain the actual depth. Simultaneously, the initial depth of the instrument end effector in the initial depth image is obtained from the above embodiment, and by comparing the actual depth and the initial depth, a correction coefficient can be determined.
[0041] It should be noted that the correction coefficient can be a simple scaling factor, which can be obtained from the ratio of the actual depth to the initial depth. Alternatively, the correction coefficient can be a depth correction function fitted by methods such as polynomial interpolation and regression analysis, for example, a correction function obtained by fitting the actual depth and the initial depth at different locations.
[0042] Regarding step S105 above, this embodiment of the disclosure can use correction coefficients to correct the depth values of pixels in the initial depth image. If the correction coefficient is a scalar, then a portion or each pixel in the initial depth image can be proportionally corrected based on the scalar. If the correction coefficient is a correction function, then the function is applied to each pixel in the initial depth image. After correction, the target depth image is a corrected depth image with high-precision absolute depth values, which is a key input for 3D reconstruction.
[0043] It should be noted that the corrected depth image eliminates the systematic errors of the depth model, providing a reliable data foundation for 3D reconstruction, thereby improving the accuracy and visualization of surgical navigation. Furthermore, the correction process can be automated, replacing manual correction and improving efficiency and consistency.
[0044] In summary, the embodiments of this disclosure, by executing the image processing method in steps S101 to S105, after acquiring the initial depth image within the target cavity, utilize the pose information of the endoscope's end-effector camera and the end of the surgical instrument, as well as the pose information of preset positions on the endoscope and the surgical instrument, to unify this information into a common coordinate system. The target pose change between the end-effector camera and the end of the instrument is then jointly calculated, thereby deriving the actual depth between them. By comparing the actual depth with the initial depth of the end of the instrument in the initial depth image, a correction coefficient is determined. This coefficient reflects the systematic error in depth prediction. Finally, the initial depth image is corrected based on the correction coefficient to generate an accurate target depth image. This process directly corrects depth errors using high-precision pose data, avoiding manual intervention dependent on experience. This not only improves correction efficiency but also significantly enhances the accuracy and reliability of the depth image.
[0045] The following is a detailed description of the further contents included in steps S101 to S105 in the embodiments of this disclosure.
[0046] Please see Figure 3 , Figure 3 yes Figure 2 The flowchart further includes step S102. In some embodiments, during the process of acquiring the first pose information of the end-effector camera and the second pose information of the instrument end, and acquiring the third pose information of the first preset position on the endoscope and the fourth pose information of the second preset position on the surgical instrument, steps S201 to S204 may be further included: Step S201: Identify the pose information of the first positioning mark pre-set on the end camera to obtain the first pose information; Step S202: Identify the pose information of the second positioning mark pre-set on the end of the instrument to obtain the second pose information; Step S203: Identify the pose information of the first tracker preset at the first preset position of the endoscope to obtain the third pose information; Step S204: Identify the pose information of the second tracker pre-set at the second preset position of the surgical instrument to obtain the fourth pose information.
[0047] In the above steps, this embodiment of the present disclosure obtains the first pose information by identifying a first positioning mark on the end camera. The first positioning mark is an optical reflective marker pre-integrated near the end camera. For example, a hemispherical reflective marker with a diameter of 2mm and a reflectivity >90% can be used, installed in a 4-point non-coplanar configuration to ensure positioning stability and accuracy. The identification process is completed by an optical navigator. The optical navigator, set in the system, captures the spatial signal of the reflective marker and calculates its position coordinates and attitude angle in a preset coordinate system in real time. This data is the first pose information and is sent to the processor for further processing.
[0048] The first pose information directly characterizes the real-time pose of the end effector camera in space, including the camera's three-dimensional coordinates and orientation (consistent with the z-axis direction of the camera coordinate system). It should be noted that the non-coplanar design of this marker avoids positioning ambiguity, and its high reflectivity ensures that it can still be accurately identified by the optical navigation system in the confined environment of the target cavity. Its output pose data provides a core benchmark for subsequent calculation of the relative positional relationship between the camera and the instrument end effector.
[0049] This embodiment of the disclosure obtains second pose information by identifying a second positioning mark at the end of the device. The second positioning mark is a dual-mode identification component, comprising two parts: one is a medical-grade color-coated marking point (such as neodymium-doped phosphor material) at the working end of the device, which forms a high-contrast feature in the endoscopic image; the other is an optical reflective marking point at the end of the device, which adopts the same hemispherical reflective structure as the first positioning mark.
[0050] For example, the second positioning mark in the embodiments of this disclosure may be different from the first positioning mark. The second positioning mark may be a marker point with a more prominent color so that the position of the instrument end can be accurately found in the image. This disclosure does not impose specific limitations on this.
[0051] Furthermore, this identification process employs a fusion of optical and visual methods. The optical navigator tracks the spatial signals of reflective markers, while the feature detection algorithm of the endoscope image locates the colored markers. After fusing the two data, the spatial pose of the instrument's end effector, i.e., the second pose information, is calculated. This information directly reflects the real-time position of the instrument's end effector (the physical endpoint in contact with the target surface). Its core function is to work in conjunction with the first pose information to provide a direct physical pose reference for deriving the actual depth of the end-effector and the instrument's end effector.
[0052] This embodiment of the disclosure obtains third pose information by identifying a first tracker at a first preset position of the endoscope. The first preset position is a certain position of the proximal end (handheld operation part) of the endoscope, and the first tracker is an optical tracker integrated therein. The tracker may include multiple non-coplanar optical reflective markers and may pre-calibrate the pose with the end camera. This embodiment of the disclosure does not impose specific limitations on this aspect.
[0053] Furthermore, the identification process is executed in real time by the optical navigator. By continuously capturing the reflected light signal of the first tracker, its pose data is converted into a common coordinate system to form the third pose information. This information represents the reference pose of the endoscope as a whole in space. Its core value lies in establishing the spatial reference anchor point of the endoscope, providing a key transition basis for unifying the pose of the end camera and the end device to the same coordinate system.
[0054] This embodiment of the disclosure obtains fourth pose information by using a second tracker that identifies a second preset position of the surgical instrument. The second preset position is the proximal end (non-working end) of the surgical instrument, and the second tracker is an optical tracker installed thereon. Furthermore, its pose association with the positioning mark at the end of the instrument can be established through a calibration process.
[0055] Furthermore, the identification process is completed synchronously by the optical navigator. By tracking the reflective marker signal of the second tracker, its pose data is mapped to a common coordinate system to obtain the fourth pose information. This information reflects the overall spatial reference pose of the surgical instrument. Its core function is to work with the third pose information to achieve spatial pose alignment between the endoscope and the surgical instrument, thus clearing the obstacles of coordinate system inconsistency for subsequent joint calculation of the target pose change of the end camera and the instrument end.
[0056] Please see Figure 4 , Figure 4 yes Figure 2 The flowchart further includes step S103. In some embodiments, the process of jointly determining the target pose change between the end-effector and the instrument end based on the first pose information, the second pose information, the third pose information, and the fourth pose information may further include steps S301 to S304: Step S301: Based on the difference between the first pose information and the third pose information, obtain the first pose change amount between the end camera and the first preset position; Step S302: Based on the difference between the second pose information and the fourth pose information, the second pose change between the instrument end and the second preset position is obtained; Step S303: Based on the difference between the third pose information and the fourth pose information, obtain the change in the third pose between the first preset position and the second preset position; Step S304: Calculate the target pose change between the end-effector camera and the instrument end based on the first pose change, the second pose change, and the third pose change.
[0057] Please see Figure 5 , Figure 5 This is a schematic diagram of the intra-articular scene of the knee joint provided in the embodiments of this disclosure, which is combined with... Figure 5 right Figure 4 The steps in the document will be explained. Figure 5 In one example scenario, the knee joint cavity is used as the target cavity, and optical tracking is employed. An end-effector camera, labeled Camera A in the diagram, is attached to the end of the endoscope. An optical navigator emits a tracking signal from above, capturing optical reflective markers C on the endoscope and D on the surgical instrument to obtain their spatial pose information. Camera A (with a first positioning marker) at the end of the endoscope is used to capture an initial image of the target cavity (such as the surface of the knee joint bone) and generate an initial depth map. A special color marker B (with a second positioning marker) at the end of the surgical instrument serves as a physical reference point in contact with the target surface. By unifying the pose information of these markers into a common coordinate system, the target pose changes of the end-effector camera and the end of the instrument can be jointly calculated, thereby correcting the initial depth image.
[0058] Regarding the above steps, this embodiment of the disclosure calculates the difference between the first pose information and the third pose information to obtain the change in the first pose between the end-effector and the endoscope at the first preset position. Here, since the first pose information is the end-effector's (corresponding to...) Figure 5 The real-time pose of camera A in space, and the third pose information is the first preset position of the endoscope (corresponding to...). Figure 5 The real-time pose of the optical reflective marker point C in the image is determined, and the two are rigidly correlated through preoperative calibration, i.e., a rigid transformation matrix. .
[0059] This embodiment of the disclosure obtains the second pose change between the instrument tip and the second preset position of the surgical instrument by calculating the difference between the second pose information and the fourth pose information. The second pose information is the instrument tip (corresponding to...) Figure 5The real-time pose of the special color marker point B in the image, and the fourth pose information is the second preset position of the surgical instrument (corresponding to...). Figure 5 The real-time pose of the optical reflective marker point D; both are rigidly correlated through preoperative calibration, i.e., a rigid transformation matrix. .
[0060] This embodiment of the disclosure obtains the change in third pose between the first preset position and the second preset position by calculating the pose difference between the third pose information and the fourth pose information. The third pose information and the fourth pose information are both calculated in real time by the optical navigation device, and their relative relationship can be directly calculated from the difference in pose vectors. This relative relationship is the change in third pose, expressed using a rigid transformation matrix. express.
[0061] Finally, based on the first pose change, the second pose change, and the third pose change, this embodiment calculates the target pose change between the end-effector and the instrument end effector using the coordinate transformation chain rule, denoted as . The pose transformation formula for the target pose change is: .
[0062] Among them, the target pose change is a complete pose information that includes position and attitude differences. Its core function is to provide a data foundation for subsequent calculation of actual depth. By extracting the position vector magnitude in this matrix, the straight-line distance between the end-effector camera and the end of the instrument can be obtained, which is the actual depth. The coordinate transformation chain rule is the key to ensuring the accuracy of pose calculation. Its principle is based on the superposition of rigid body motion, so as to determine the pose relationship of each component through coordinate transformation, thereby providing a unique and reliable physical reference standard for subsequent depth correction and avoiding errors caused by human intervention.
[0063] Please see Figure 6 , Figure 6 yes Figure 2 A flowchart further includes step S101. In some embodiments, the process of acquiring the initial depth image within the target cavity may further include steps S401 to S402: Step S401: After the instrument tip moves to the corresponding target surface inside the target cavity, an initial image is obtained by taking a picture inside the target cavity through the end-eye camera. Step S402: Input the initial image into the pre-trained large depth model for monocular depth prediction and output the initial depth image inside the target cavity.
[0064] To achieve the above steps, this embodiment requires acquiring an initial image to generate an initial depth image. Specifically, after the endpiece of the surgical instrument is moved to a target surface within the target cavity, such as the femoral condyle in the knee joint cavity, the talus plateau in the ankle joint cavity, or the pleural surface in the thoracic cavity, and contact is maintained, an end-effector camera integrated into the endoscope is used to capture images of the area containing the target anatomical structures and the endpiece of the instrument within the target cavity, thus obtaining an initial image. This initial image is a two-dimensional image, which can be in RGB or grayscale format, and must clearly show the special markings on the endpiece of the instrument.
[0065] Among them, the end-effector camera is the imaging device installed at the very front of the endoscope. Its field of view needs to cover a local area of the target cavity (such as the bone surface and surrounding soft tissue in the joint cavity). When taking pictures, it is necessary to ensure that the marker point at the end of the instrument is within the effective area of the image. The target surface is the anatomical surface that the end of the instrument physically contacts, serving as the physical reference benchmark for subsequent depth correction. The initial image is the input source for depth prediction and needs to contain sufficient scene information to support the geometric inference of the depth model.
[0066] It should be noted that the core of this step lies in acquiring the original image with physical reference anchor points. By contacting the target surface with the end of the instrument, the initial image simultaneously contains the depth scene to be reconstructed and the precisely locatable physical points (the end of the instrument), providing a basis for subsequent comparison of the depth prediction results with the actual physical depth. The minimally invasive nature of the endoscope also ensures that high-quality initial images can still be acquired within the confined target cavity.
[0067] Next, the initial image obtained in this embodiment is input into a pre-trained large depth model to perform monocular depth prediction to generate an initial depth image. The large depth model is a large language model trained based on a deep learning framework (such as a convolutional neural network). The training data includes a large number of endoscopic image pairs labeled with real depth (or synthetic depth data generated through multi-view geometry). The model learns the mapping relationship between image texture, edges, occlusion, and other features and depth, enabling it to infer scene depth from a single two-dimensional image. In addition, this embodiment can also use other models for depth prediction, such as a multi-scale convolutional neural network model, which learns depth information from monocular images through multi-scale feature extraction and disparity regression; or a DepthFormer model based on the Transformer architecture, which uses a self-attention mechanism to capture long-distance dependencies to improve the global consistency of depth prediction; or a semi-supervised model combined with geometric constraints (such as GeoNet), which enhances the scale stability of depth estimation through multi-view geometric priors. These models can all serve as alternative solutions to predict the initial depth image within the target cavity, providing a basic depth reference for subsequent correction processes.
[0068] During monocular depth prediction, the large depth model performs feature extraction and multi-scale feature fusion on the initial image, ultimately outputting an initial depth image with the same resolution as the initial image. This initial depth image uses pixel values to represent the estimated depth (relative depth value) of each pixel relative to the end camera. For example, 16-bit grayscale values are used to encode depth information, with larger values indicating greater distance from the camera.
[0069] Among them, the large depth model is the core algorithm module for realizing depth estimation from two-dimensional to three-dimensional, and its training quality directly affects the accuracy of the initial depth image; monocular depth prediction refers to the technique of inferring depth using only a two-dimensional image, which is suitable for endoscope monocular imaging scenarios; the initial depth image is the direct product of depth prediction, and its pixel-level depth estimation provides basic data for subsequent correction, but due to the complexity of the target cavity environment (such as uneven illumination and lack of texture), there are scale errors and local biases.
[0070] It should be noted that the embodiments of this disclosure can quickly obtain a preliminary estimate of the depth distribution within the target cavity. With its powerful feature learning capability, the large depth model can extract depth clues from images when prior geometric information is lacking, thus meeting the needs of minimally invasive surgery for real-time performance and preliminary depth perception.
[0071] Please see Figure 7 , Figure 7 yes Figure 2 The flowchart further includes step S105. In some embodiments, the process of correcting the initial depth image based on the correction coefficient to obtain the corrected target depth image may further include steps S501 to S503: Step S501: When the instrument end moves to another position on the target surface, update the target pose change between the end camera and the instrument end based on the second pose information of the moved instrument end, and update the correction coefficient based on the updated target pose change. Step S502: Based on the correction coefficients corresponding to different positions of the instrument tip on the target surface, jointly determine the target correction coefficients of the initial depth image; Step S503: Correct the initial depth image based on the target correction coefficient to obtain the corrected target depth image.
[0072] To address the aforementioned steps, this embodiment of the disclosure achieves a comprehensive update of the correction coefficients through multi-position sampling at the instrument tip. Specifically, when the operator moves the surgical instrument tip from one contact point on the target surface to another, such as different regions of the femoral condyle in the knee joint cavity or different anatomical points of the pleura in the thoracic cavity, the pose information of the second positioning mark on the instrument tip can be re-identified to obtain updated second pose information. Subsequently, according to the pose calculation logic in the above embodiment, based on the updated second pose information, the first pose information of the end-effector camera, and the third and fourth pose information of the preset positions of the endoscope and surgical instrument, the target pose change between the end-effector camera and the instrument tip is re-determined. Then, based on the updated target pose change, the actual depth between the end-effector camera and the instrument tip is calculated, and combined with the initial depth of the corresponding position of the instrument tip in the initial depth image at this time, the correction coefficients are updated.
[0073] Among them, the other location refers to the area on the target surface that is independent of the initial contact point, and needs to be distributed in different fields of view of the initial depth image to ensure representativeness; the updated second pose information is the spatial pose data of the instrument end effector at the new location; the update of the target pose change is used to reflect the real-time relative pose relationship between the instrument end effector and the end camera after the instrument end effector moves; the update of the correction coefficient is based on the comparison between the new actual depth and the initial depth to obtain the correction value of the depth prediction error at this location.
[0074] It should be noted that the embodiments of this disclosure achieve the dynamic nature of the correction through multi-location sampling. Since the depth prediction error within the target cavity may have spatial heterogeneity, by collecting correction coefficients at multiple typical locations, a wider range of scene features can be covered, laying the foundation for subsequent fusion to obtain more reliable target correction coefficients and further improving the accuracy of depth image correction.
[0075] Next, this embodiment of the present disclosure obtains the target correction coefficient of the initial depth image by fusing the correction coefficients of the instrument tip at different positions on the target surface. Specifically, it is necessary to collect the correction coefficients calculated at at least two different positions of the instrument tip, and then use a joint determination method, such as arithmetic mean or weighted average, where the weights can be allocated according to the representativeness of each position in the image or the magnitude of the depth prediction error, to calculate a comprehensive target correction coefficient. The target correction coefficient is the unified coefficient finally used for global (or local) correction of the initial depth image, and the linear or nonlinear fusion method can be selected according to the scene characteristics.
[0076] It should be noted that the embodiments disclosed herein can enhance the robustness of the correction. By fusing correction coefficients at multiple locations, the correction deviation caused by the anatomical structure of a single contact point, such as lack of local texture or abnormal light reflection, can be effectively reduced. This makes the target correction coefficients more reflective of the depth prediction error pattern of the entire scene, providing a reliable basis for the accurate correction of subsequent depth images.
[0077] Subsequently, embodiments of this disclosure perform correction processing on the initial depth image based on the target correction coefficient to obtain a corrected target depth image. Specifically, the target correction coefficient is applied to the depth value of each pixel in the initial depth image, and the systematic error of the initial depth image is globally corrected through mathematical operations, such as multiplying the pixel-level depth value by the correction coefficient, or performing nonlinear correction according to the error model. If a regional correction coefficient is used, the corresponding regional coefficient is applied to the pixels in different regions for local correction. The final output image is the corrected target depth image, whose pixel depth values are closer to the true physical depth.
[0078] It should be noted that, in this embodiment, by applying the target correction coefficient fused from multiple locations to the initial depth image, a leap from preliminary depth estimation to precise depth mapping is achieved. The corrected target depth image provides high-precision depth information of intracavitary anatomical structures for minimally invasive surgery, which can significantly improve the accuracy of surgical planning and operation, and provide accurate data for the construction of three-dimensional models.
[0079] Exemplary, the calculations in the embodiments of this disclosure Position components are Its projection onto the z-axis of coordinate system A is the color marker point, which is also the true depth of the second positioning mark on the instrument end relative to the end-effector camera. The true depth value is defined as... Furthermore, the position (ui, vi) of the positioning device's end effector in pixel space can be identified through pixel identification. Using a large depth model (such as Depth-anything), a depth estimate of the entire monocular image can be obtained, with the initial depth defined at pixel position (ui, vi) denoted as d. Therefore, the depth of the previously acquired local region... It can be used to correct d. When the endoscope remains stationary and the instrument contacts different positions within the cavity environment, + can be used to obtain the true depth corresponding to multiple pixel positions. The depth in the calibrated target depth image can be represented as a function , where u, v are pixel positions, d is the depth predicted by the large model, and α is the depth correction coefficient. Therefore, since the correction coefficients corresponding to several pixel positions are obtained, the depth correction coefficients of all pixels in the entire image can be obtained by means of interpolation not limited to polynomial interpolation.
[0080] Please see Figure 8 , Figure 8 yes Figure 7 The flowchart further includes steps S501. In some embodiments, after updating the correction coefficients based on the updated target pose change, the image processing method may further include steps S601 to S602: Step S601: Based on the correction coefficients corresponding to different positions of the instrument tip on the target surface, jointly determine the correction coefficient matrix of the initial depth image; Step S602: Based on the correction coefficient matrix, perform corresponding correction processing on different positions of the initial depth image to obtain the corrected target depth image.
[0081] To address the aforementioned steps, this embodiment of the disclosure constructs a correction coefficient matrix covering the entire initial depth image by fusing correction coefficients from different locations of the instrument tip on the target surface. Specifically, after obtaining the corresponding correction coefficients from anatomical feature points such as the medial and lateral femoral condyles and the tibial plateau in the knee joint cavity at multiple different locations of the instrument tip on the target surface, it is necessary to first determine the pixel coordinates of these locations in the initial depth image. This can be achieved through image feature matching or marker point localization, and then by employing a joint determination method, such as spatial interpolation, region fitting, or anatomically structure-based partitioning mapping, the discrete correction coefficients are expanded into a two-dimensional matrix with the same resolution as the initial depth image, i.e., the correction coefficient matrix.
[0082] The calibration coefficients must be evenly distributed within the field of view of the initial depth image and cover typical anatomical structures within the target cavity, such as bone surfaces and soft tissue interfaces, to ensure spatial representativeness of the calibration. Joint determination is the core process of converting discrete calibration coefficients into a continuous matrix. For example, a bilinear interpolation algorithm can be used to calculate the coefficients of adjacent pixels based on the calibration coefficients of known locations, or to assign corresponding regional calibration coefficients to each region according to anatomical partitions, such as the femoral region and tibia region of the joint cavity. The calibration coefficient matrix is a two-dimensional array with the same size as the initial depth image. Each element corresponds to the depth calibration coefficient of that pixel location in the image, which can accurately reflect the depth prediction error characteristics of different regions, such as smaller errors in bone surface regions and larger errors in soft tissue regions.
[0083] It should be noted that the embodiments of this disclosure can achieve spatial adaptability of correction. The depth prediction error in the target cavity is not uniform globally. For example, uneven illumination and texture differences may cause local error fluctuations. By constructing a correction coefficient matrix, the error patterns at different locations can be captured in a targeted manner, providing a data basis for subsequent pixel-level or region-level accurate correction and avoiding local deviations caused by single coefficient correction.
[0084] Next, based on the obtained correction coefficient matrix, this embodiment performs position-specific correction processing on the initial depth image to generate a corrected target depth image. Specifically, the coefficients at each pixel position in the correction coefficient matrix are matched with the depth values of the corresponding pixels in the initial depth image, such as through pixel-level multiplication, so that the corrected depth = initial depth × the correction coefficient at the corresponding position, or according to the nonlinear correction of the error model; if a region correction coefficient is used, the coefficients of that region are applied to all pixels within the same anatomical region for batch correction. The final output image is the corrected target depth image, in which the depth values of each pixel have been corrected by the coefficients at the corresponding positions, making it closer to the true physical depth.
[0085] Among them, the correction processing corresponding to different positions refers to the process of applying different correction coefficients according to the differences in the spatial position of the image, which reflects the accuracy of correcting the error in a targeted manner. The corrected target depth image is a depth map that has been corrected for spatial heterogeneity error. It retains the scene structure information of the original depth image and eliminates local error deviations through matrix correction, thus having higher spatial consistency and absolute scale accuracy.
[0086] It should be noted that by mapping the correction coefficients one-to-one with the image positions, the differences in local errors caused by the complex environment within the target cavity can be effectively accommodated, such as the difference in error between the edge and center regions of the joint cavity, significantly improving the overall accuracy of the depth image. The corrected target depth image can more realistically reflect the three-dimensional spatial relationships of the anatomical structures within the cavity, providing more reliable depth data support for the subsequent construction of the target cavity's three-dimensional model.
[0087] Furthermore, embodiments of this disclosure can also execute a three-level sampling strategy during arthroscopic dynamic scanning. The first level is anatomical landmark localization to identify key sites such as the intercondylar fossa of the femur and the central region of the tibial plateau. The second level is contact measurement, where the instrument's working end contacts the bone surface with a constant force, and data is recorded after the reading stabilizes. The third level is data fusion, which integrates the measurements from the optical navigation system. With image coordinates Correlate and form calibration data Where C represents the calibration dataset, used to store the correlation data from image coordinates to true depth, providing a basis for subsequent depth image correction. ) is the pixel coordinate of the i-th sampling point in the initial depth image, u is the horizontal pixel index, and v is the vertical pixel index, used to locate the pixel position that needs to be corrected in the image. is the true depth value measured by the optical navigator at the i-th sampling point, reflecting the actual physical distance between the instrument's end effector and the end-effector camera, and serves as the true reference for depth correction. N is the total number of sampling points, i.e., the total number of valid data pairs obtained through the three-level sampling strategy of anatomical landmark localization, contact measurement, and data fusion. i is the index of the sampling point (from 1 to N), used to distinguish data pairs with different image coordinates and true depth.
[0088] Further, please refer to Figure 9 , Figure 9 yes Figure 2 The flowchart further includes steps S105. In some embodiments, after obtaining the corrected target depth image, the image processing method may further include steps S701 to S702: Step S701: When the first pose information represents the movement of the end camera, continuously acquire the corrected target depth image of the end camera at different positions during the movement of the end camera. Step S702: Based on multiple consecutive corrected target depth images, a three-dimensional model of the target cavity is constructed to obtain a three-dimensional intracavity model of the target cavity.
[0089] To address the aforementioned steps, this embodiment tracks the pose changes of the end-effector camera and continuously acquires corrected target depth images during its movement to cover multi-view areas of the target cavity. Specifically, the dynamic change in the first pose information indicates that the end-effector camera is in a moving state. This movement is typically achieved by a physician operating an endoscope to adjust the endoscope's viewing angle and capture images of different areas within the target cavity, such as the anterior and posterior compartments of the knee joint cavity, or the lung lobes and mediastinal space in the thoracic cavity. During movement, the system repeatedly executes the steps described in the above embodiment at a preset frequency to continuously perform image correction, predicting and correcting the depth of the initial images captured at each location, and continuously generating corrected target depth images. These images correspond to the depth observation results of the end-effector camera at different spatial positions.
[0090] Among them, the first pose information characterizes the movement of the end-effector camera by judging whether the camera is in a dynamic adjustment state through the temporal changes of pose data. For example, when the positional deviation of the pose of two consecutive frames exceeds a preset threshold, it is determined to be in a moving state. The corrected target depth image is an accurate depth map after error correction. Each image corresponds to the field of view of the end-effector camera in a specific pose and contains accurate depth information of the target cavity anatomy structure from that perspective. Continuous acquisition is to cover the complete area of the target cavity with multi-view images and avoid the blind spot of observation from a single perspective.
[0091] It should be noted that the core function of this step is to provide a multi-view, high-precision depth data foundation for 3D reconstruction: the target cavity structure is complex and the space is small, and a depth image from a single viewpoint cannot fully reflect its 3D shape. By continuously acquiring calibrated depth images during camera movement, depth observation data from different directions can be obtained, providing sufficient raw materials for subsequent stitching to construct a complete 3D model.
[0092] Next, based on the acquired multiple consecutively corrected target depth images, this embodiment constructs a complete three-dimensional model of the target cavity using a three-dimensional reconstruction algorithm. Specifically, the three-dimensional reconstruction algorithm may include, for each frame of the target depth image, converting the two-dimensional image into a three-dimensional point cloud according to the camera intrinsic parameters and the corrected depth image, with the reference coordinate system of the three-dimensional point cloud being the camera coordinate system. Then, based on the true pose of the camera obtained through the optical navigator, transforming the three-dimensional point cloud in the camera coordinate system to the global coordinate system, merging the point clouds of multiple frames, removing duplicate points or performing filtering, and finally realizing the process from point cloud to reconstructed surface mesh to obtain the required three-dimensional cavity model.
[0093] Among them, multiple consecutive corrected target depth images are the input data for 3D reconstruction. Their continuity ensures that the depth information of adjacent viewpoints overlaps, providing a matching basis for stitching. 3D model construction is the core process of converting 2D depth images into 3D structures. Spatial alignment of depth information from different viewpoints is achieved through pose data, eliminating stitching errors. The 3D intracavitary model is a digital 3D representation of the internal structure of the target cavity, and it can have measurability (such as distance between two points, surface curvature) and visualization characteristics (such as viewing from multiple angles).
[0094] It should be noted that the embodiments of this disclosure, by fusing high-precision depth data from multiple perspectives, construct a three-dimensional intracavitary model that overcomes the perspective limitations of a single depth image. This model can completely and accurately reflect the anatomical structural features of the target cavity, and the entire intracavitary environment can be reconstructed using the actual trajectory of the end-effector and the corrected depth image. This model can provide intuitive three-dimensional navigation for minimally invasive surgery, such as surgical path planning and instrument positioning, significantly improving the precision and safety of the surgery and solving the problem of insufficient visualization of anatomical structures caused by traditional two-dimensional images or low-precision three-dimensional models.
[0095] Please see Figure 10 This disclosure also provides an image processing apparatus that can implement the above-described image processing method. The image processing apparatus includes: The image acquisition module 1001 is used to acquire an initial depth image inside the target cavity and to acquire the initial depth corresponding to the end of the surgical instrument from the initial depth image. The initial depth image is generated by taking an initial image inside the target cavity through the end camera of the endoscope after the end of the instrument moves to the target surface corresponding to the target cavity, and by predicting the depth inside the target cavity based on the initial image. The pose acquisition module 1002 is used to acquire the first pose information of the end-effector camera and the second pose information of the instrument end, and to acquire the third pose information of the first preset position on the endoscope and the fourth pose information of the second preset position on the surgical instrument. The pose relationship determination module 1003 is used to jointly determine the target pose change between the end camera and the instrument end when the third pose information and the fourth pose information belong to the same coordinate system. The correction coefficient determination module 1004 is used to determine the actual depth between the end camera and the instrument end based on the target pose change, and to determine the correction coefficient of the initial depth image based on the actual depth and the initial depth. The image processing module 1005 is used to perform correction processing on the initial depth image based on the correction coefficient to obtain the corrected target depth image.
[0096] In summary, the image processing device, through the image processing method described in the above embodiments, acquires an initial depth image within the target cavity. Then, it utilizes the pose information of the endoscope's end-effector camera and the surgical instrument's end-effector, as well as the pose information of preset positions on the endoscope and surgical instrument, unifying this information into a common coordinate system. The device then jointly calculates the target pose change between the end-effector camera and the instrument's end-effector, thereby deriving the actual depth between them. By comparing the actual depth with the initial depth of the instrument's end-effector in the initial depth image, a correction coefficient is determined. This coefficient reflects the systematic error in depth prediction. Finally, based on the correction coefficient, the initial depth image is corrected to generate an accurate target depth image. This process directly corrects depth errors using high-precision pose data, avoiding reliance on experience-based manual intervention. This not only improves correction efficiency but also significantly enhances the accuracy and reliability of the depth image.
[0097] The specific implementation of this image processing apparatus is basically the same as the specific embodiments of the image processing method described above, and will not be repeated here. While meeting the requirements of the embodiments of this disclosure, the image processing apparatus may also be equipped with other functional modules to implement the image processing method described above.
[0098] This disclosure also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described image processing method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0099] Please see Figure 11 , Figure 11 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 1101 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this disclosure. The memory 1102 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1102 can store operating devices and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1102 and is called and executed by the processor 1101 to execute the image processing method of the embodiments of this disclosure. Input / output interface 1103 is used to implement information input and output; The communication interface 1104 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1105 transmits information between various components of the device (e.g., processor 1101, memory 1102, input / output interface 1103, and communication interface 1104); The processor 1101, memory 1102, input / output interface 1103 and communication interface 1104 are connected to each other within the device via bus 1105.
[0100] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described image processing method.
[0101] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0102] The embodiments described in this disclosure are for the purpose of more clearly illustrating the technical solutions of this disclosure and do not constitute a limitation on the technical solutions provided by this disclosure. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by this disclosure are also applicable to similar technical problems.
[0103] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this disclosure, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0104] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0105] Those skilled in the art will understand that all or some of the steps, apparatuses, or functional modules / units in the methods disclosed above can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0106] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in this disclosure and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, apparatus, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0107] It should be understood that in this disclosure, "at least one item" means one or more, and "more than one" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0108] In the several embodiments provided in this disclosure, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0109] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0110] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0111] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0112] The preferred embodiments of the present disclosure have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present disclosure. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and spirit of the present disclosure shall be within the scope of the claims of the present disclosure.
Claims
1. An image processing method, characterized in that, include: An initial depth image of the target cavity is acquired, and the initial depth corresponding to the end of the surgical instrument is obtained from the initial depth image. The initial depth image is generated by taking an initial image of the target cavity through the end camera of the endoscope after the end of the instrument is moved to the target surface corresponding to the target cavity, and the depth of the target cavity is predicted based on the initial image. Acquire the first pose information of the end-effector camera and the second pose information of the instrument end, and acquire the third pose information of the first preset position on the endoscope and the fourth pose information of the second preset position on the surgical instrument. When the third pose information and the fourth pose information belong to a common coordinate system, the target pose change between the end camera and the end device is jointly determined based on the first pose information, the second pose information, the third pose information and the fourth pose information. The actual depth between the end-effector camera and the instrument end is determined based on the target pose change, and the correction coefficient of the initial depth image is determined based on the actual depth and the initial depth. The initial depth image is corrected based on the correction coefficient to obtain the corrected target depth image.
2. The image processing method according to claim 1, characterized in that, The step of acquiring the first pose information of the end-effector camera and the second pose information of the instrument end, and acquiring the third pose information of the first preset position on the endoscope and the fourth pose information of the second preset position on the surgical instrument, includes: The pose information of the first positioning mark pre-set on the end camera is identified to obtain the first pose information; The pose information of the second positioning mark pre-set on the end of the instrument is identified to obtain the second pose information; The pose information of the first tracker, which is pre-set at the first preset position of the endoscope, is identified to obtain the third pose information; The fourth pose information is obtained by identifying the pose information of the second tracker pre-set at the second preset position of the surgical instrument.
3. The image processing method according to claim 1, characterized in that, The step of jointly determining the target pose change between the end-effector and the instrument end based on the first pose information, the second pose information, the third pose information, and the fourth pose information includes: Based on the difference between the first pose information and the third pose information, the first pose change between the end camera and the first preset position is obtained; Based on the difference between the second pose information and the fourth pose information, the second pose change amount between the instrument end and the second preset position is obtained; Based on the difference between the third pose information and the fourth pose information, the change in the third pose between the first preset position and the second preset position is obtained; The target pose change between the end camera and the instrument end is calculated based on the first pose change, the second pose change, and the third pose change.
4. The image processing method according to claim 1, characterized in that, The acquisition of the initial depth image within the target cavity includes: After the instrument tip is moved to the corresponding target surface inside the target cavity, an initial image is obtained by taking a picture inside the target cavity through the endoscope's end camera; The initial image is input into a pre-trained large depth model for monocular depth prediction, and the initial depth image inside the target cavity is output.
5. The image processing method according to claim 1, characterized in that, The step of correcting the initial depth image based on the correction coefficient to obtain the corrected target depth image includes: When the instrument end moves to another position on the target surface, the target pose change between the end camera and the instrument end is updated again based on the second pose information of the moved instrument end, and the correction coefficient is updated based on the updated target pose change. Based on the correction coefficients corresponding to different positions of the instrument tip on the target surface, the target correction coefficients of the initial depth image are jointly determined. The initial depth image is corrected based on the target correction coefficient to obtain the corrected target depth image.
6. The image processing method according to claim 5, characterized in that, After updating the correction coefficients based on the updated target pose change, the image processing method further includes: Based on the correction coefficients corresponding to different positions of the instrument tip on the target surface, the correction coefficient matrix of the initial depth image is jointly determined; Based on the correction coefficient matrix, corresponding correction processing is performed on different positions of the initial depth image to obtain the corrected target depth image.
7. The image processing method according to claim 1, characterized in that, After obtaining the corrected target depth image, the image processing method further includes: When the first pose information indicates that the end-effector camera is moving, the target depth images after correction of the end-effector camera at different positions are continuously acquired during the movement of the end-effector camera. A three-dimensional model of the target cavity is constructed based on multiple consecutive corrected target depth images to obtain a three-dimensional intracavitary model of the target cavity.
8. An image processing apparatus, characterized in that, include: An image acquisition module is used to acquire an initial depth image within the target cavity and to acquire the initial depth corresponding to the end of the surgical instrument from the initial depth image. The initial depth image is generated by taking an initial image within the target cavity through the end-effector camera after the end of the instrument has moved to the target surface corresponding to the target cavity, and by predicting the depth within the target cavity based on the initial image. The pose acquisition module is used to acquire the first pose information of the end-effector camera and the second pose information of the instrument end, and to acquire the third pose information of the first preset position on the endoscope and the fourth pose information of the second preset position on the surgical instrument. The pose relationship determination module is used to jointly determine the target pose change between the end camera and the instrument end when the third pose information and the fourth pose information belong to a common coordinate system. The correction coefficient determination module is used to determine the actual depth between the end camera and the instrument end based on the target pose change, and to determine the correction coefficient of the initial depth image based on the actual depth and the initial depth. The image processing module is used to perform correction processing on the initial depth image based on the correction coefficient to obtain the corrected target depth image.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the image processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the image processing method according to any one of claims 1 to 7.