Self-navigation overhead support system and method for imaging system
By adopting a multi-degree of freedom overhead support system and a self-navigation positioning system in the X-ray system, combining visual and non-visual sensors to generate three-dimensional mapping maps, the collision problem of the X-ray system when moving in a multi-room environment in the trauma site is solved, and imaging efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202411572396.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-05
- Filing Date
- 2024-11-06
- Publication Date
- 2025-06-06
AI Technical Summary
When existing X-ray systems move in a multi-room environment in the trauma site, it is difficult to effectively avoid collisions with medical equipment and personnel in the room, resulting in inefficient imaging and increased errors.
A multi-degree of freedom overhead support system is adopted, combining vision sensors and non-vision sensors, and a three-dimensional mapping of the environment is generated through self-navigation and positioning systems to navigate the X-ray source in real time to avoid collisions.
It improves the movement efficiency of the X-ray system in a multi-room environment, reduces the risk of collision with objects in the environment, and ensures the stability and accuracy of the imaging process.
Smart Images

Figure CN120093335A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to X-ray systems, and more particularly to X-ray systems that include an overhead support system for moving a portion of the imaging system, such as an X-ray tube, to accommodate various patient positions. Background Art
[0002] A number of X-ray imaging systems of various designs are known and are currently in use. Such systems are generally based on the generation of X-rays directed at a subject of interest. The X-rays traverse the subject and impinge on a detector (e.g., film, an imaging receptor, or a portable cassette). The detector detects the X-rays that are attenuated, scattered, or absorbed by the intervening structures of the subject. For example, in the context of medical imaging, such systems can be used to visualize the internal structures, tissues, and organs of the subject for the purpose of screening or diagnosing disease.
[0003] X-ray systems may be fixed or mobile. Fixed radiography systems typically utilize an X-ray source that is movably mounted to a ceiling in the area where the X-rays are to be obtained. In one prior art configuration, the radiography system is Figure 1 An overhead support X-ray system 100 is shown. The overhead support system 100 (or in an exemplary embodiment, an overhead tube support (OTS) system 100) generally includes a column 105 to which an X-ray source 110 is attached, the column being coupled to an overhead rectangular bridge 115 that travels along a system 120 of rails or tubes oriented perpendicular to the bridge 115. A transport mechanism 125 coupled to the bridge 115 operates to move the column 105 along a longitudinal horizontal axis, while the rail system 120 allows the bridge 115 to travel along a transverse horizontal axis in the same plane. The rail system 120 generally includes a front rail 120a, a rear rail 120b, and a cable drape rail (not shown) mounted to the ceiling of the room or suite housing the fixed radiography system. In some installations, the overhead tube support system 100 may be mounted to a brace system secured to the ceiling to enable the X-ray source 110 to be oriented relative to a fixed table 130 or fixed ledge 135 which holds the detector 140 thereon in order to obtain the desired images of a patient positioned thereon or adjacent thereto.
[0004] One environment in which the OTS system 100 may be employed is in a trauma setting 200, an exemplary embodiment of which is Figure 1, such as an emergency room area within a hospital. The trauma site 200 typically includes a plurality of treatment rooms 202 separated by shortened walls or half walls 204 and each including one or more stretchers or patient tables 130 therein, along with other monitoring equipment and / or medical equipment 208 to facilitate treatment of patients within the rooms 202. In order to accommodate the imaging procedures to be performed in each room 202, the OTS system 100 is configured with rails 120 extending above the shortened walls 204 to allow the carriage or bridge 115 to be moved into a selected room 202 in which the imaging procedure is to be performed. Once positioned within the selected room 202, the OTS system 100 can be operated to position the X-ray source 110 at a location desired for obtaining an image of the patient on the table 130, i.e., allowing the OTS system 100 to move the X-ray source 110 above the table 130 where the patient is located.
[0005] In order to effectively treat patients in each room 202, it is necessary to be able to efficiently move the X-ray source 110 to the desired imaging location within and between each room 202. Therefore, the movement of the OTS system 100 is of primary concern in terms of efficiency because the OTS system 100 is "shared" by several rooms 202 within the trauma site 200. Without smooth and efficient movement to reach the desired target location in the room 202, the efficiency of the imaging process performed by the OTS system 100 and the X-ray source 110 may be less than desired. However, when attempting to standardize the rooms 202 within the trauma site 200 to provide the same level of care in each room 202, due to the emergency nature of the medical problems treated in the rooms 202, the rooms 202 typically contain varying types and quantities of medical equipment 208 therein, with the different positioning of the medical equipment 208 within the respective rooms 202 being used to best perform the necessary treatment for the patients. In this case, the movement of the OTS system 100 and the X-ray source 110 must be customized for each room 202 to accommodate the different locations of the table 130, medical equipment 208, and medical staff within the individual rooms 202, thereby minimizing the risk of collision with the OTS system 100 and the X-ray source 110 when moving within a particular room 202.
[0006] In an attempt to address the problem of minimizing the likelihood of collision of the OTS system 100 and the X-ray source 110 with one or more items in the trauma room 202 while maximizing the efficiency of the movement of the OTS system 100 and the X-ray source 110 within the room 202, two different solutions are currently employed. First, for each room 202, the OTS system 100 is programmed with a predefined route or path between a defined starting or parking / non-use position and a predetermined position of the stretcher / table 130 in each room 202. Because the starting or parking position and the desired position of the X-ray source 110 adjacent to the table 130 are predetermined for a selected imaging procedure, items located within the room 202 (i.e., the table 130, medical equipment 208, and / or medical personnel) can be positioned outside the travel path of the OTS system 100 and the X-ray source 110 to avoid collisions.
[0007] However, a significant disadvantage of this solution is the requirement that all items in the room 202 be positioned outside of the intended path of the OTS system 100 and the X-ray source 110, which may often and easily be overlooked due to the urgent nature of the medical issues being addressed in the room 202 and the resulting changeable positioning of items in the room 202. Specifically, while the intended path of the OTS system 100 and the X-ray source 110 is known to the OTS system 100, it may not be known to the medical personnel in the room 202. Furthermore, while the OTS system 100 may also mitigate situations where a collision does occur by limiting the collision force and / or shutting down the drive mechanism for the OTS system 100, this solution for avoiding collisions between the OTS system 100 and items in the room 202 of the trauma site 200 is suboptimal.
[0008] A second solution to the problem of avoiding collisions between the OTS system 100 and items disposed within the trauma site room 202 is to omit the OTS system 100 entirely and employ one or more mobile X-ray devices within the trauma site 200. However, this solution is also suboptimal because it adds another item that needs to be moved around the table 130, other medical equipment 208, and already existing medical staff within the room 202. In this case, during the clinical imaging process, the radiologist is required to adjust the position of the X-ray device and the wall mount for each patient to achieve the desired orientation for obtaining the necessary images of the patient. However, the manual positioning of the equipment in the room 202 requires a large amount of time and energy, which reduces the imaging efficiency of the X-ray device and prolongs the waiting time for each user for the X-ray device. In addition, the complex arrangement of the X-ray device and other necessary medical equipment used to obtain images often leads to unexpected imaging errors due to incorrect positioning of the imaging device and / or user distraction caused by the extensive equipment positioning process.
[0009] Therefore, it is desirable to develop systems and methods for positioning an OTS system including an X-ray source and optional X-ray detector relative to a patient that overcome these limitations of the current prior art. Summary of the invention
[0010] According to one aspect of an exemplary embodiment of the present disclosure, an imaging system includes: a multi-degree-of-freedom overhead support system, the multi-degree-of-freedom overhead support system is suitable for being mounted to a surface within an environment of the imaging system; an imaging device, the imaging device is mounted to the overhead support system; a visual sensor, the visual sensor is disposed on the imaging device; a non-visual sensor, the non-visual sensor is disposed on the imaging device; a motion controller, the motion controller is operably connected to the overhead support system; a processor, the processor is operably connected to the motion controller and the visual sensor and the non-visual sensor to send control signals to the overhead support system, the visual sensor and the non-visual sensor and to receive signals from the processor. They receive data signals; and a memory operably connected to the processor, the memory storing therein processor-executable instructions for the operation of a self-navigation and positioning system, the self-navigation and positioning system being configured to generate a three-dimensional (3D) map of the environment of the imaging system based on visual data from the visual sensor and non-visual data from the non-visual sensor, wherein the processor-executable instructions, when executed by the processor to operate the self-navigation and positioning system, cause: generating the 3D map of the environment; and navigating the overhead support system from a starting position to an ending position within the environment to avoid collision with one or more objects identified on the 3D map within the environment.
[0011] According to another aspect of the exemplary embodiment of the present disclosure, a method for navigating an overhead support system of an imaging system through an environment includes the following steps: providing an imaging system, the imaging system having: a multi-degree-of-freedom overhead support system, the multi-degree-of-freedom overhead support system being adapted to be mounted to a surface within the environment of the imaging system; an imaging device, the imaging device being mounted to the overhead support system; a visual sensor, the visual sensor being disposed on the imaging device; a non-visual sensor, the non-visual sensor being disposed on the imaging device; a motion controller, the motion controller being operably connected to the overhead support system; a processor, the processor being operably connected to the motion controller and the visual sensor and the non-visual sensor to control signals are sent to the overhead support system, the visual sensor and the non-visual sensor and data signals are received from them; and a memory operably connected to the processor, the memory storing therein processor-executable instructions for the operation of a self-navigation and positioning system, the self-navigation and positioning system being configured to generate a three-dimensional (3D) map of the environment of the imaging system based on visual data from the visual sensor, non-visual data from the non-visual sensor and position data from the motion controller; generate the 3D map of the environment; and navigate the overhead support system from a starting position to an ending position within the environment to avoid collision with one or more objects identified on the 3D map within the environment.
[0012] These and other exemplary aspects, features and advantages of the present invention will become apparent from the following detailed description taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The drawings illustrate the best mode presently contemplated of practicing the invention.
[0014] In the attached picture:
[0015] Figure 1 A diagram of a prior art imaging system utilizing an overhead support system in a trauma field including multiple treatment rooms is shown.
[0016] Figure 2 is an isometric view of a trauma facility having multiple treatment rooms and an imaging system including an overhead support system capable of moving an x-ray source between treatment rooms according to an exemplary embodiment of the present disclosure.
[0017] Figure 3 is an isometric view of a room including an overhead support system with a self-navigating positioning system according to an exemplary embodiment of the present disclosure.
[0018] Figure 4 is with Figure 3Isometric view of an X-ray source utilized together with the overhead support system and universal positioning system.
[0019] Figure 5 is a schematic diagram of the operation of a self-navigation positioning system according to an exemplary embodiment of the present disclosure.
[0020] Figure 6 is a flowchart illustrating an exemplary operation method of a self-navigation positioning system according to an exemplary embodiment of the present disclosure.
[0021] Figure 7 is a schematic diagram of a visual sensor image employed in 2D semantic segmentation performed by a self-navigation and positioning system according to an exemplary embodiment of the present disclosure.
[0022] Figure 8 is a schematic diagram of a non-vision sensor image employed in a homogeneous coordinate transformation performed by a self-navigation positioning system according to an exemplary embodiment of the present disclosure.
[0023] Fig. 9 is a schematic diagram of an alignment image output from a coordinate alignment step performed by a self-navigation positioning system according to an exemplary embodiment of the present disclosure.
[0024] Fig.10 is a schematic diagram of a 3D map output by a self-navigation positioning system according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] One or more specific embodiments will be described below. In order to provide a concise description of these embodiments, all features of an actual implementation may not be described in the specification. It should be understood that, as in any engineering or design project, in the development of any such actual implementation, numerous implementation-specific decisions must be made to achieve the developer's specific goals, such as complying with system-related and business-related constraints that may differ between implementations. In addition, it should be understood that such development efforts may be complex and time-consuming, but are still routine tasks of design, fabrication, and manufacturing for ordinary technicians who benefit from this disclosure.
[0026] When introducing the elements of various embodiments of the present invention, the articles "one", "a", "the" and "said" are intended to mean that there are one or more such elements. The terms "comprising", "including" and "having" are intended to be inclusive and mean that there may be additional elements in addition to the listed elements. In addition, any numerical examples in the following discussion are intended to be non-restrictive, and therefore the additional numerical values, ranges and percentages are within the scope of the disclosed embodiments. As used herein, the terms "substantially", "generally" and "approximately" indicate conditions within reasonably achievable manufacturing and assembly tolerances relative to the ideal desired conditions suitable for achieving the functional purpose of the component or assembly. Moreover, as used herein, "electrically coupled", "electrically connected" and "electrically communicated" mean that the referenced elements are directly or indirectly connected so that current can flow from one to another. The connection may include a direct conductive connection (i.e., without intervening capacitors, inductors or active elements), an inductive connection, a capacitive connection and / or any other suitable electrical connection. There may be intervening components. As used herein, the term "real time" means a level of processing responsiveness that is sensed by the user as being sufficiently immediate or enabling the processor to keep up with the external process.
[0027] Figure 2 and Figure 3 An exemplary embodiment of a trauma site 1000 is shown that includes a plurality of adjacent treatment rooms 300 and includes an imaging system 302 that can be moved between the rooms 300 to obtain desired images of patients located in the rooms 300. The imaging system 302 includes a self-navigation and positioning system 1002 disposed on a workstation 410 for controlling movement and operation of various components of the imaging system 302.
[0028] The imaging system 302 is formed with a first imaging device 306, which may be an X-ray tube or a detector, which is fixed to a portion of the site 1000 or other location where the imaging system 302 is arranged, such as a wall or ceiling 308 of the site 300, by a movable mount 305, wherein the movable mount 305 may be an overhead support system 310 or a robotic arm. The individual rooms 300 each include one or more of a table 312, a wall shelf 314, and a second imaging device 342, which may be the other of the X-ray tube or the detector forming the imaging system 302, and the second imaging device may be incorporated into one or both of the table 312 and the wall shelf 314.
[0029] The overhead support system 310 (which in the exemplary embodiment shown supports an X-ray source, such as a first imaging device 306) provides five (5) separate degrees of freedom / axes of automatic or manual directional movement for the first imaging device 306, and specifically allows lateral and longitudinal movement of a mount 305 along a suspension track 311 of the overhead support system 310, vertical movement via a telescoping column 313 attached to the mount 305 and movable along the suspension track 311 with the mount 305, rotational movement relative to the mount 305 provided by rotation of the column 313, and angular movement provided by a pivot mechanism 315 disposed between the column 313 and the first imaging device 306. The overhead support system 310 also includes one or more suitable position monitors 316 to provide accurate and precise position information regarding the position of the first imaging device 306 disposed on the overhead support system 310 and the individually movable component parts 305, 313, 315 of the overhead support system 310. The movement of each component part 305, 313, 315 of the overhead support system 310 is controlled by a motion controller 317, which is interconnected with the overhead support system 310 and includes one or more motors (not shown) or other power devices capable of operating to independently and selectively move the mounting member 305, the column 313 and the pivot mechanism 315, and an inertial measurement unit (IMU) (not shown) integrated with the motion controller or forming a component part thereof, which can enable the motion controller 317 to provide position data / information about the acceleration and angular velocity of the movement of the overhead support system 310 along each axis of movement of the overhead support system 310 and / or the first imaging device 306 (e.g., horizontal axes - longitudinal (x) axis and lateral (y) axis and / or vertical (z) axis, etc.), as well as position data / information about any rotational movement, optionally in real time.
[0030] It should also be understood that the imaging system 302 may also include other components suitable for implementing the disclosed embodiments. Exemplary imaging procedures that may be performed by the imaging system 302 include radiography procedures such as, but not limited to, computed tomography (CT) procedures, computerized axial tomography (CAT) scanning procedures, and fluoroscopy procedures.
[0031] The overhead support system 310 and the expanded first imaging device 306 are operably connected to a workstation 410, which constitutes at least a portion of the self-navigation and positioning system 1002, which in an exemplary embodiment is located remotely from the imaging system 302, such as at a location outside the room 300 of the wound site 1000. The workstation 410 may include a computer 415, one or more input devices 420 (e.g., a keyboard, mouse, or other suitable input device), and one or more output devices 425 (e.g., a display screen or other device that provides data from the workstation 410). The workstation 410 may receive commands, scanning parameters, and other data from an operator or from a memory 430 and a processor 435 of the computer 415. The computer 415 / processor 435 may use the commands, scanning parameters, and other data to exchange control signals, commands, and data with the overhead support system 310, the first imaging device 306, the table 312, the ledge 314, and the second imaging device 342 through a suitable wired or wireless control interface 440 connected to each of these components of the imaging system 302. For example, the control interface 440 may provide control signals to and receive image, position, or other data signals from one or more of the overhead support system 310, the first imaging device 306, the table 312, the ledge 314, and the second imaging device 342. In addition, the motion controller 317 may be operably connected to the workstation 410 to receive information from the workstation 410 regarding the current and desired positions of the overhead support system 310 and each of the independently movable components of the overhead support system 310 (e.g., the mount 305, the column 313, and the pivot mechanism 315). In a particular exemplary embodiment, motion controller 317 communicates with workstation 410 regarding the current position, angle, and desired position of overhead support system 310 via cables or wireless technology.
[0032] The workstation 410 can control the frequency and amount of radiation generated by the X-ray source 306 or 342, the sensitivity of the detector 306 or 342, and the position of the table 312 and the ledge 314 to facilitate the scanning operation. Signals from the detector 306 or 342 can be sent to the workstation 410 for processing. The workstation 410 may include image processing capabilities for processing the signals from the detector 306 or 342 to produce an output of a real-time 2D or 3D image for display on the one or more output devices 425. In addition, using the five axes of motion provided by each of the overhead support system 310 and the ledge 314, the self-navigation positioning system 1002 enables the imaging system 302 to perform classical table imaging and ledge procedures with only a single detector 306, 342. In addition, the ability of the overhead support system 310, the table 312, and the ledge 314 to operate automatically provides the imaging system 302 utilizing the self-navigation positioning system 1002 with the ability to automatically and / or manually control (such as via the workstation 410) the overhead support system 310.
[0033] Now refer to Figure 4 , the X-ray source 306 disposed on the overhead support system 310 includes a visual sensor 1006 and a non-visual sensor 1008, each of which is disposed on a housing 1010 of the X-ray source 310 and each of which is capable of obtaining three-dimensional (3D) information about the environment surrounding the overhead support system 306 and the X-ray source 306. The visual sensor 1006 may take the form of a camera (such as a depth camera), and the non-visual sensor 1008 may take the form of an infrared sensor, a laser radar or radar sensor, or an ultrasonic sensor, etc. The visual sensor 1006 and the non-visual sensor 1008 are positioned adjacent to each other and oriented in the same direction on the housing 1010 so as to provide both the visual sensor 1006 and the non-visual sensor 1008 with the same or similar field of view of the surrounding environment.
[0034] To control the movement of the overhead support system 310, in an exemplary embodiment, the autonomous navigation and positioning system 1002 is provided on the workstation 410 as processor executable instructions that are stored in the memory 430 and accessible by the computer 415 / processor 435 for operation of a simultaneous localization and mapping (SLAM) algorithm 1012. The SLAM algorithm 1012 enables the use of information from the position sensor 316 and / or motion controller 317 regarding the position of the overhead support system 310 within the room 300 in combination with spatial information about the environment surrounding the overhead support system 310 provided by the vision sensor 1006 and the non-vision sensor 1008 to map the environment surrounding the overhead support system 310 and autonomously guide or navigate the overhead support system 310 through the mapped environment. More specifically, the SLAM algorithm-based self-navigation positioning system 1012 for the overhead support system 310 generates and continuously updates an internal map (i.e., an internal map of the environment (i.e., the trauma site 1000 / room 300) surrounding the overhead support system 310. Fig.10 ) as a reference when moving within the environment. The map provides locations for locating landmarks within the environment (e.g., table 312 and ledge 314) under dynamic conditions with respect to the placement of landmarks and other items within the environment, and enables motion controller 317 to be operated to efficiently move overhead support system 310 to a target location while avoiding obstacles between overhead support system 310 and the target location.
[0035] The self-navigation positioning system 1012 overcomes the problems associated with the prior art single-sensor SLAM algorithm-based systems, in which the reliability of the map generated by the prior art SLAM algorithm is compromised because each of the visual sensor 1006 and the non-visual sensor 1008 has shortcomings in certain circumstances when used alone. For example, the effectiveness of the visual sensor 1006 is highly dependent on the environment, and in particular, even slight lighting changes can have a significant negative impact. In addition, the operation of the non-visual sensor 1008 is significantly degraded by any off-angle reflection of the energy waves used by the non-visual sensor 1008. Therefore, in order to improve the accuracy of sensing and thus enhance the robustness of the 3D map generated by the self-navigation and positioning system 1002 employing the SLAM algorithm 1012, both the visual sensor 1006 and the non-visual sensor 1008 are employed to provide information to the SLAM algorithm 1012, wherein a fusion of the information from each sensor 1006, 1008 is performed by the self-navigation and positioning system 1002 to calibrate the data from these sensors 1006, 1008 in order to address the computational burden of typical camera-based mapping, while mitigating the issues presented by changes in lighting within the environment to provide a more accurate and reliable map of the environment for use in navigating the overhead support system 310 within the environment. In addition, when the non-visual sensor 1008 is operated to scan over a 360° range in 2D, then the mapping process performed by the SLAM algorithm 1012 will not be affected when the overhead support system 310 moves backward or forward.
[0036] Reference now Figure 5 , shows an exemplary method 600 employed by a self-navigating positioning system 1002 using a SLAM algorithm 1012 that employs information from a motion controller 317 in an iterative process (the information including visual data (i.e., 2D images) / first input 606 and non-visual data (i.e., 3D measurements / point clouds) / second input 608 and positional data / motion controller data 614 regarding the position of an overhead support system 310) to generate and continuously update a map 650 of the environment to guide and control the movement of the overhead support system 310 and the connected imaging device 306 within the environment in real time. Figure 6In an exemplary embodiment of method 600 shown in , method 600 includes a first semantic mapping phase or operation 602 and a second map refinement phase or operation 604 to generate (e.g., in real time) a map of the environment (i.e., trauma site 1000 / room 300) through which the overhead support system 310 is moving, so that the self-navigation and positioning system 1002 can move the overhead support system 310 to a desired position along a path within the environment 1000, 300 determined by the self-navigation and positioning system 1002 to obtain an image of the patient while avoiding collisions with other objects and / or individuals present and / or moving within the environment.
[0037] Regarding the first semantic mapping phase or operation 602, initially a first input 606 from a visual sensor 1006 and a second input 608 from a non-visual sensor 1008 are provided to an algorithm 1012 for processing into an environment map 650. In general, using the visual data / first input 606 and the non-visual data / second input 608, the method 600 / algorithm 1012 proceeds to
[0038] a. Synchronizing the 3D measurement / point cloud (eg, second input 608) from the 3D space / non-vision sensor 1006 to a recording timestamp recorded in association with the 2D image (eg, first input 606) obtained by the vision sensor 1008;
[0039] b. Correction / removal of motion artifacts caused by the motion of the overhead support system 310 to ensure 3D measurements
[0040] / matching the point cloud (e.g., second input 608) with a synchronized image plane captured by the vision sensor 1006;
[0041] c. converting the 3D measurement / point cloud (second input 608) into homogeneous coordinates within a homogeneous coordinate system so that the 3D measurement / point cloud can be aligned with the synchronized 2D image frame, optionally taking into account various intrinsic and extrinsic parameters;
[0042] i. Intrinsic parameters include field of view (FOV) and focal length, etc., which are used to eliminate distortion to ensure correct mapping from the sensor plane to the image plane.
[0043] ii. The external parameters are the position correlation between two sensors with different coordinate systems, which include a rotation matrix and a translation vector; - Using the external parameters, synchronization on the sensor plane is generated according to the position relationship;
[0044] iii. Based on the parameters required for alignment, the 3D coordinates obtained from the 3D space / non-vision sensor 1008 will be transformed into a 2D projection;
[0045] d. transforming the homogeneous coordinate system into the Euclidean coordinate system; and
[0046] e. Generate a refined 3D map 650 using the Euclidean coordinate system.
[0047] More specifically, if Figures 6 to 10 As shown, in the illustrated exemplary embodiment of the method 600 of operation of the SLAM algorithm 1012 of the self-navigating positioning system 1002, in the semantic mapping phase or operation of the method 600, initially in step 610, the visual data (i.e., 2D image) / first input 606 and the non-visual data (i.e., 3D measurement / point cloud) / second input 608 are synchronized with each other, such as to temporally match the non-visual / 3D position data from the second input 608 with the image / visual data provided by the first input 606. In an exemplary embodiment, synchronization can be performed by associating a recorded timestamp with the first input 606 to determine the 2D image in the first input 606 obtained during the time used to obtain the 3D measurement / point cloud forming the second input 608. Since the sampling density of the 3D sensor 1008 (e.g., LiDAR) is much less than the sampling density of the 2D sensor 1006 (e.g., camera), a method of associating or registering the LiDAR point cloud to the camera coordinate system based on the timestamps generated during the frame sampling process is used. Thereafter, synchronization in step 610 results in a sparse information mapping result. In addition, synchronization in step 610 may include a third input 612 in the form of position data / motion controller data 614 regarding the position and / or direction and speed of movement of the overhead support system 310 under the guidance of the motion controller 317 at a time associated with the timestamp of the image data in the first input 606.
[0048] In step 616, after synchronization in step 610, the first input 606 data is corrected for any motion artifacts in the 3D measurements / point cloud forming the first input 606, which motion artifacts are caused by the movement of the overhead support system 310 during the time when the 2D images are obtained. Motion artifacts, which generally refer to the inability to obtain sensor odometry information through visual methods when the sensor (e.g., camera) moves rapidly, are primarily related to 2D image data processing. In this case, the pose (e.g., position and orientation) of the visual sensor 1006 is primarily calculated based on consecutive 2D image frames, which is a rough estimate due to the fact that consecutive frames accumulate errors. In particular, when the overhead support system 310 moves or rotates too quickly, the limited frame rate (i.e., the number of images obtained by the visual sensor 1006 in a specified period of time) causes motion blur in the 2D image / first input 606. This artifact correction can be achieved in a known manner by algorithm 1012, for example, where invalid odometry information caused by motion artifacts is replaced by the IMU to ensure correct pose estimation of sensor 1006, and information in a third input 612 from motion controller 317 relating to the position and motion (i.e., speed and direction) of overhead support system 310 during synchronization time can also be used.
[0049] After synchronization and correction, the visual data / first input 606 from the visual sensor 1006 is processed in a 2D semantic segmentation step 618. In this step 618, Figure 6 and Figure 7 As best shown in , the 2D camera image / first input 606 is analyzed by a SLAM algorithm 1012, such as by an artificial intelligence (AI) 1014, such as a deep learning module (e.g., U-Net) or a convolutional neural network (CNN), which forms part of the SLAM algorithm 1012 trained to recognize shapes 607, 609 of various structures within a 2D image of an environment (such as the room 300 of the trauma site 1000). In performing semantic segmentation, the AI 1014 recognizes the shapes 607, 609 in the 2D image / first input 608 and classifies the shapes 607, 609 into various object and surface types present within the 2D image / first input 608, and applies labels 620 (e.g., "table", "ledge", "floor", etc.) to the various objects and / or surfaces identified within the 2D image / first input 606, along with a probability or confidence for each selected label 620, as shown in FIG. Figure 7 as shown in .
[0050] At the same time as the 2D semantic segmentation step, Figure 6 and Figure 8As shown, in step 622, the synchronized and corrected 3D measurement / point cloud of the second input 608 is processed or transformed by the SLAM algorithm 1012 from the homogeneous coordinate system of the 3D measurement / point cloud of the second input 608, such as based on the known position of the 3D sensor 1008 from the motion controller 317 via the data / third input 614 and the position of the 3D measurement / point cloud of the second input 608 relative to the non-vision sensor 1008. More specifically, in step 622, the homogeneous coordinate system of the projected depth information of the 3D measurement / point cloud of the second input 608 is transformed into a globally consistent coordinate system. Each 3D measurement / point cloud is sensed by the 3D sensor 1008 and stored in a homogeneous form, which can be one of a plurality of optional coordinate forms of the 3D measurement / point cloud, and can be stored in the memory 430 to be accessed by the SLAM algorithm 1012 for post-processing purposes defined by the sensor 1008, and then transformed by the algorithm 1012 in step 622 into a consistent (common / standard) coordinate system, which can be easily combined or fused with the odometer information provided by the motion controller 317 and / or the 2D camera image / first input 606 for 3D map generation. Once the 3D measurements / point clouds forming the second input 608 are defined relative to homogeneous coordinates in step 622, the SLAM algorithm 1012 may use the homogeneous coordinates of the 3D measurements / point clouds of each of the individual components of the 3D measurements (i.e., point clouds) of the second input 608 to define / identify one or more volumes 624, wherein each defined volume 624 is transformed into constituent voxels 626 representing the 3D measurements / point clouds of the volume 624 in step 625, as shown in FIG. Figure 8 Because the second input 608 can and often will contain multiple separate and / or overlapping components within the 3D measurement / point cloud, the volume 624 and constituent voxels 626 will represent the various objects and / or surfaces detected by the non-vision sensor 1008 and constituting the 3D measurement / point cloud.
[0051] After the semantic segmentation of the 2D image in step 618 and the formation of the volume 624 in the homogeneous coordinate system in step 622, as Figure 6 and Fig. 9 As shown in the exemplary embodiment of FIG. 6 , the voxelized 3D measurement / point cloud 624 of the second input 608 is aligned with the labeled 2D image / first input 606 to form one or more aligned planes or images 632 in step 628. In this alignment, in addition to introducing the various intrinsic and / or extrinsic parameters 630 described previously, the voxels 626 forming the volume 624 are projected / overlaid onto and aligned with the synchronized and labeled 2D image 606 corresponding to the 3D measurement / point cloud in the second input 608, as shown in FIG. Fig. 9As shown. In an exemplary embodiment of the alignment process of step 628, for each frame (i.e., each view captured by the 2D vision sensor 1006 that can be matched with the 3D sensor data / point cloud from the non-vision sensor 1008, where the number of frames obtained per unit time, i.e., the frame rate described above), the 3D measurement / point cloud of the second input 608 is projected onto the 2D vision plane / image of the first input 606. The projection is based on the determined association (i.e., positional relationship and timestamp) between each 2D vision plane / image of the first input 606 and the associated 3D measurement / point cloud of the second input 608, so as to obtain sparse alignment points with depth values. Thereafter, an upsampling technique is applied to the sparse points to form a dense map corresponding to the 2D image. Since the 2D image with a label as output was previously semantically segmented in step 618, the voxels 626 in the upsampled map are associated with the semantic result of step 618 to assign a label to each voxel 626 containing depth information. This alignment is performed for one or more associated pairs of synchronized 2D images 606 and 3D measurements / point clouds 608 so that each volume 624 in the second input 608 can be aligned with the tagged object present in the synchronized 2D image 606 of the plan view containing the volume 624. In addition, for each 2D image / visual data in the first input 606 obtained by the vision sensor 1006 at a particular position and / or orientation of the vision sensor 1006, alignment of the synchronized 3D spatial measurements / point clouds in the second input 608 can be performed to produce additional aligned images 632 of the same environment 1000, 300 and objects therein, thereby forming a set of aligned images 632. Figure 5 As schematically shown in FIG. 1 , as the visual sensor 1006 and the non-visual sensor 1008 are moved around within the environment 1000 , 300 via the overhead support system 310 , the set of aligned images 632 may be used to define (optionally in real time) a semantic 3D map 633 of the environment 1000 , 300 .
[0052] In conjunction with the projection and alignment of the volume 624 / voxel 626 onto the synchronized 2D image 606, the coordinates of the voxel 626 within the homogeneous coordinate system are also converted or transformed into the Euclidean coordinate system represented within the 2D image 606 during the alignment process performed in step 628. This process is performed using the correspondence between the position of the voxel 626 in the synchronized 2D image 606 and the Euclidean coordinate system defined within the 2D image 606 of the environment 1000, 300 so that the semantic 3D map 640 can conform to the Euclidean coordinate system defined in the environment 1000, 300.
[0053] Once the alignment in step 628 is completed, projection errors are also introduced into the aligned images 632 due to differences in the positioning or alignment of the voxels 626 forming the volume 624 to the representation of the object and / or surface in the 2D image 606. To account for these errors, in step 634, a separate CNN 636 (such as a CNN trained to perform spatial reasoning, e.g., voxel distance clustering) is employed by the algorithm 1012 to decide which voxels 626 correspond to the actual location of the object in the 2D image 606, and which voxels 626 represent projection errors to be removed from the aligned images 632. To assist in the classification of the voxels 626 to be retained and removed, the CNN 636 analyzes the point cloud / volume 624 formed by the voxels 626 in order to identify and provide labels for the objects represented by the point cloud / volume 624 in each aligned image 632. Using the identification / labeling, the CNN 636 can remove voxels 626 that are outside the expected volume of the identified object, i.e., voxels 626 that are identified as being associated with projection errors and not associated with the object.
[0054] In step 638, the SLAM algorithm 1012 may fuse the results of the 2D image caricature segmentation in step 618 into a verification of the recognition provided as an output of the CNN 636, and enable the CNN 636 to define the type and shape of the object represented by the volume 624, so that only voxels 626 that align with the proper position and / or shape of the recognized object may be retained by the CNN 636 within the alignment image 632. This combination of semantic segmentation and recognition using each of the 2D image 606 and the alignment image 632 effectively minimizes errors in accurately locating objects within the environment of the overhead support system 310, and enables the SLAM algorithm 1012 to produce a highly accurate semantic 3D map 640 of the environment 1000, 300.
[0055] In addition, in step 642, the SLAM algorithm 1012 can update the semantic 3D map 640, such as by providing analysis results (e.g., additional alignment images 632 of additional first inputs 606 and second inputs 608, such as first inputs 606 and second inputs 608 acquired or obtained after forming the semantic 3D map). The semantic fusion analysis results from the additional first inputs 606 and second inputs 608 can be used by the SLAM algorithm 1012 to update the semantic 3D map 640 with enhanced probabilities of labeling of different voxels 626 to more accurately define the shape and / or position of objects in the environment 1000, 300, optionally providing the semantic 3D map 640 of the environment 1000, 300 in real time.
[0056] To determine that the semantic 3D map 640 is complete in step 643, in one exemplary embodiment, the mapping process continues with each frame acquired by the vision sensor 1006 to produce an alignment image 632 including the associated 2D semantic information with the 3D projected points from step 618 until all voxels 626 in the 3D map 640 are labeled, although the labels may not be completely accurate. To correct the labels of the voxels 626, the semantic mapping stage or operation 602 is terminated and the semantic 3D map 640 is output to the map refinement stage or operation 604 of the SLAM algorithm 1012. In the refinement stage 604, initially in step 646, the spatial distribution of the labeled / associated objects or volumes 624 in the semantic 3D map 640 is corrected to more clearly identify and segment the space occupied by and disposed between the volumes 624 in the semantic 3D map 640, thereby providing an area through which the overhead support system 310 can be moved from a current position to a desired imaging position. As previously mentioned, some semantic information is incorrect in terms of the labels of the voxels 626, which have a relatively low density and are rarely far from the ground truth about the voxel clusters. By applying the spatial distribution correction in step 646, a final semantic map 650 is generated with fewer errors or mislabeled voxels 626. As a result, better segmentation performance is achieved in each frame of the 3D map.
[0057] Finally, in step 648, voxels 626 previously identified as errors with respect to the position and shape of the labeled volume 624 are removed from the semantic 3D map 640 to provide a 3D map 650, such as Fig.10650 for use by the self-navigation and positioning system 1002 to move the overhead support system 310 from a current and / or parked position 1016 to a desired imaging position 1018 (e.g., adjacent to the table 312) between the locations of the identified and labeled volume 624 on the map 650. The 3D map 650 can be presented in any desired orientation so that even in the top view of the 3D 650, a general representation including the locations and labels of the different objects present in the environment 1000, 300 is provided. Therefore, the 3D map 650 is more reliable for routing use by the self-navigation and positioning system 1002 due to the ability to not only locate but also identify objects in the environment, thereby enabling the system 1002 to distinguish which objects to avoid and which objects are adjacent to the desired location of the overhead support system 310. Furthermore, where the dynamic nature of the environment 1000, 300 can be sensed in real time by the various sensors 1006, 1008 of the system 1002, potential collisions between the overhead support system 310 and objects in the environment 1000, 300 are more detectable and avoidable than prior art collision mitigation systems. Furthermore, while in certain embodiments the 3D map 650 is not presented on the display 425 to a user of the imaging system 302 and is utilized only internally by the self-navigation and positioning system 1002, in other embodiments the 3D map 650 may be provided on the display 425, optionally for further verification by the user of the labels of the objects / volumes 624 presented in the 3D map 650, as well as the starting and / or parked position 660, the ending position 662, and the selected route 664 of the overhead support system 310 that may also be selectively identified within the 3D map 650.
[0058] In addition to the ability of the self-navigation and positioning system 1002 to employ both the visual sensors 1006 and the non-visual sensors 1008 in the fused 3D mapping process performed by the SLAM algorithm 1012 to provide a highly accurate 3D map for identification of and navigation around objects located in the surrounding environment, there are additional benefits to the self-navigation and positioning system 1002 of the present disclosure. More specifically, using input from two different types of sensors (e.g., visual and non-visual) allows the self-navigation and positioning system 1002 to generate a more accurate map that is independent of each of the lighting and scale of the environment in which the system 1002 is operated. In addition, the semantic segmentation information in the 3D map 650 includes system-defined labels 620 for each of the objects or components present in the 3D map 650, allowing a user to immediately identify objects within the 3D map 650.
[0059] In addition to the benefits provided by the self-navigation and positioning system 1002 regarding the information provided to the user and the operation of the motion controller 317 via the 3D map 650, the real-time information obtained by the self-navigation and positioning system 1002 regarding the movement of objects within the environment over time provides new potential for the reconstruction, visualization, and identification of movement patterns in the environment. For example, through the recorded camera images and 3D spatial information, the 3D map generated by the self-navigation and positioning system 1002 over a specific time frame can show the paths taken by objects and people within the environment covered by the 3D map 650, which can be used as clinical evidence to improve the workflow within the room. In a similar manner, the 3D map generated by the self-navigation and positioning system 1002 over a specific time frame can be used to calculate an average of the time spent for each patient based on 4D data consisting of a dynamic 3D map 650 or a movie or video of the 3D map 650 in combination with the associated time stamps.
[0060] Finally, it should also be understood that the self-navigation and positioning system 1002 and / or the SLAM algorithm 1012 may include the necessary computers, electronic devices, software, memory, storage devices, databases, firmware, logic / state machines, microprocessors, communication links, displays or other visual or audio user interfaces, printing devices, and any other input / output interfaces for performing the functions described herein and / or achieving the results described herein. For example, as previously described, the system may include at least one processor / processing unit / computer and system memory / data storage structure, which may include random access memory (RAM) and read-only memory (ROM). At least one processor of the system may include one or more conventional microprocessors and one or more auxiliary coprocessors, such as mathematical coprocessors, etc. The data storage structure discussed herein may include an appropriate combination of magnetic, optical and / or semiconductor memory, and may include, for example, RAM, ROM, flash drives, optical disks such as compact disks, and / or hard disks or drives.
[0061] In addition, the software application / algorithm that adapts the computer / controller to perform the methods disclosed herein can be read from a computer-readable medium into the main memory of at least one processor. As used herein, the term "computer-readable medium" refers to any medium that provides or participates in providing instructions to at least one processor of the system 302, 304 (or any other processor of the device described herein) for execution. Such media can take many forms, including but not limited to non-volatile media and volatile media. Non-volatile media include, for example, optical, magnetic or optical magnetic disks, such as memory. Volatile media include dynamic random access memory (DRAM) that usually constitutes main memory. Common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, tapes, any other magnetic media, CD-ROMs, DVDs, any other optical media, RAMs, PROMs, EPROMs or EEPROMs (electronically erasable programmable read-only memories), FLASH-EEPROMs, any other memory chips or cassettes, or any other media from which a computer can read.
[0062] Although in an embodiment, execution of a sequence of instructions in a software application causes at least one processor to perform the methods / processes described herein, hard-wired circuitry may be used in place of or in combination with software instructions to implement the methods / processes of the present invention. Accordingly, embodiments of the present invention are not limited to any specific combination of hardware and / or software.
[0063] It should be understood that the aforementioned compositions, devices and methods of the present disclosure are not limited to specific embodiments and methods, as these may vary. It should also be understood that the terms used herein are only for the purpose of describing specific exemplary embodiments, and are not intended to limit the scope of the present disclosure, which will only be limited by the appended claims.
Claims
1. An imaging system (302), the imaging system comprising: a. a multi-degree-of-freedom overhead support system (310), the multi-degree-of-freedom overhead support system being adapted to be mounted to a surface (308) within an environment (300) of the imaging system (302); b. an imaging device (306) mounted to the overhead support system (310); c. A visual sensor (1006), the visual sensor being arranged on the imaging device (306) superior; d. a non-visual sensor (1008), the non-visual sensor being disposed on the imaging device (306): e. A motion controller (317) operably connected to the overhead support system (310): f. a processor (435) operably connected to the motion controller (317) and the visual sensor (1006) and the non-visual sensor (1008) to send control signals to the overhead support system (310), the visual sensor (1006) and the non-visual sensor (1008) and receive data signals from them; and g. a memory (430) operably connected to the processor (435) and storing therein processor-executable instructions for operation of a self-navigation and positioning system (1002), the self-navigation and positioning system being configured to generate a three-dimensional (3D) map of the environment (300) of the imaging system (302) using visual data (606) from the visual sensor (1006) and non-visual data (608) from the non-visual sensor (1008), wherein the processor executable instructions, when executed by the processor to operate the self-navigation and positioning system (1002), cause: i. generating the 3D map (650) of the environment (300); and ii. Navigating the overhead support system (310) from a starting position (660) to an ending position (662) within the environment (300) to avoid collision with one or more objects identified on the 3D map (650) within the environment (300).
2. The imaging system (302) of claim 1, wherein the visual sensor (1006) is a camera.
3. The imaging system (302) of claim 1, wherein the non-vision sensor (1008) is a 3D spatial sensor.
4. The imaging system (302) of claim 3, wherein the self-navigation and positioning system (1002) is configured to generate a three-dimensional (3D) map of the environment (300) of the imaging system (302) using visual data (606) from the visual sensor (1006), non-visual data (608) from the non-visual sensor (1008), and position data (614) from the motion controller (317).
5. The imaging system (302) of claim 1, wherein the self-navigation and positioning system (1002) includes a simultaneous positioning and mapping algorithm (1012), and wherein the processor-executable instructions, when executed by the processor (435) to operate the simultaneous positioning and mapping algorithm (1012), cause: a. performing a semantic mapping operation (603) to generate a semantic 3D mapping map (640); and b. Performing a map refinement operation (604) to generate the 3D map (650) from the semantic 3D map (640).
6. The imaging system (302) of claim 5, wherein the processor-executable instructions, when executed by the processor (435) to perform the semantic mapping operation (602), cause: a. synchronizing (610) the visual data (606) with the non-visual data (608); b. converting (622) the non-visual data (608) into homogeneous coordinates; c. aligning (628) the non-visual data (608) with the visual data (606); d. converting the homogeneous coordinates of the non-visual data to Euclidean coordinates; and e. Generating the semantic 3D map (640) based on the non-visual data (608).
7. An imaging system (302) according to claim 6, wherein the processor executable instructions, when executed by the processor (435) to perform the semantic mapping operation (1002), cause motion artifacts (616) in the visual data (606) to be corrected before aligning the non-visual data (608) in homogeneous coordinates with the visual data (606).
8. The imaging system (302) of claim 6, wherein the processor executable instructions, when executed by the processor (435) to perform the semantic mapping operation (602), cause: iii. semantically segmenting (618) the visual data (606) to identify objects represented in the visual data (606) and forming semantic visual data (620) before aligning the non-visual data (608) with the visual data (606); iv. performing feature extraction and classification (634) on the non-visual data (608) to identify objects represented in the non-visual data (608) after aligning the non-visual data (608) with the semantic visual data (620); and v. fusing (638) the semantic visual data (620) and the non-visual data (608) The features are extracted and classified to form a semantic 3D map (640) for creating the 3D map (650).
9. The imaging system (302) of claim 6, wherein the processor-executable instructions, when executed by the processor (435) to align the non-visual data (608) with the visual data (606), cause: a. defining the non-visual data (608) using the homogeneous coordinates into voxels (626) forming one or more volumes (624) in the non-visual data (608); and b. Overlaying the voxels (626) of the non-visual data (608) onto the semantic visual data (620).
10. The imaging system (302) of claim 9, wherein the processor executable instructions, when executed by the processor (435) to convert the homogeneous coordinates of the non-visual data (608) to Euclidean coordinates, cause the homogeneous coordinates of the voxels (626) to be converted to Euclidean coordinates.
11. The imaging system (302) of claim 9, wherein the processor-executable instructions, when executed by the processor (435) to extract and classify features (634) of the non-visual data (608) to identify an object represented in the non-visual data (608), cause: a. marking the one or more volumes (624) within the non-visual data (608) formed by the voxels (626); and b. Remove voxels (626) that are outside the volume (624).
12. A method for navigating an overhead support system (310) of an imaging system (302) through an environment (300), the method comprising the steps of: a. Providing an imaging system (302), the imaging system comprising: i. a multi-degree-of-freedom overhead support system (310) adapted to be mounted to a surface (308) within an environment (300) of the imaging system (302); ii. an imaging device (306) mounted to the overhead support system (310); iii. a visual sensor (1006), wherein the visual sensor is disposed on the imaging device (306); iv. a non-visual sensor (1008), the non-visual sensor being disposed on the imaging device (306); v. a motion controller (317) operably connected to the overhead support system (310); vi. a processor (435) operably connected to the motion controller (317) and the visual sensor (1006) and the non-visual sensor (1008) to send control signals to the overhead support system (310), the visual sensor (1006) and the non-visual sensor (1008) and receive data signals from them; and vii. a memory (430) operably connected to the processor (435) and storing therein processor-executable instructions for operation of a self-navigation and positioning system (1002), the self-navigation and positioning system being configured to generate a three-dimensional (3D) map (650) of the environment (300) of the imaging system (302) using visual data (606) from the visual sensor (1006), non-visual data (608) from the non-visual sensor (1008), and position data (614) from the motion controller (317), b. generating the 3D map (650) of the environment (300); and c. Navigating the overhead support system (310) from a starting position (660) to an ending position (662) within the environment (300) to avoid collision with one or more objects identified on the 3D map (650) within the environment (300).
13. The method according to claim 12, wherein the self-navigation and positioning system (1002) comprises a simultaneous positioning and mapping algorithm (1012), and wherein the method comprises the following steps: a. Performing a semantic mapping operation (602) using the simultaneous localization and mapping algorithm (1012) to generate a semantic 3D map (640); and b. Performing a map refinement operation (604) to generate the 3D map (650) from the semantic 3D map (640).
14. The method of claim 13, wherein the step of performing the semantic mapping operation (602) causes: a. synchronizing (610) the visual data (606) with the non-visual data (608); b. converting (622) the non-visual data (608) into homogeneous coordinates; c. aligning (628) the non-visual data (608) with the visual data (606); d. converting the homogeneous coordinates of the non-visual data (608) to Euclidean coordinates; and e. Generating the semantic 3D map (640) based on the non-visual data (608).
15. The method of claim 14, wherein the step of performing the semantic mapping operation (602) comprises: a. semantically segmenting (618) the visual data (606) to identify objects represented in the visual data (606) before aligning the non-visual data (608) with the visual data (606); b. extracting and feature classifying (634) the non-visual data (608) to identify objects represented in the non-visual data (608) after aligning the non-visual data (608) with the visual data (606); and c. Fusion (638) of the semantic segmentation (618) of the visual data (606) and the feature extraction and classification (634) of the non-visual data (608) to form a semantic 3D map (640) for use in creating the 3D map (650).