Re-envisioning force-torque sensing: simple six-axis sensor with a single camera and fiducial markers

A camera-tracked sensor with mechanical amplification mechanisms addresses the limitations of existing FT sensors by providing high-stiffness and low-cost, reliable six-axis force/torque sensing with minimal environmental drift and noise.

WO2026107484A1PCT designated stage Publication Date: 2026-05-21YALE UNIVERSITY +3
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
YALE UNIVERSITY
Filing Date
2025-11-18
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Existing six-axis force/torque (FT) sensors for robots are expensive, fragile, and prone to drift due to environmental factors, requiring complex assembly and calibration, while vision-based sensors lack stiffness and are limited to specific applications.

Method used

A sensor device using a flexure plate with fiducial markers and mechanical amplification mechanisms, tracked by a single camera, to measure six-axis forces and torques with high stiffness and accuracy, eliminating the need for delicate transducers and precise calibration.

Benefits of technology

The sensor achieves high-stiffness, low-cost, and reliable six-axis force/torque sensing with less than 1.5% relative error, maintaining accuracy over long durations and reducing susceptibility to noise and drift.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025055891_21052026_PF_FP_ABST
    Figure US2025055891_21052026_PF_FP_ABST
Patent Text Reader

Abstract

Described herein is a sensor device comprising a mounting plate, a flexure plate positioned parallel to the mounting plate having an inner surface facing the mounting plate and an outer surface opposite the inner surface, a wall connected to the flexure plate and the mounting plate, forming an enclosure between the flexure plate and the mounting plate, at least one detector disposed on the mounting plate within the enclosure, an end-effector connected to the outer surface of the flexure plate and configured to transfer applied forces to the flexure plate, at least one mirror attached to the inner surface of the flexure plate opposite the at least one detector, and at least one emitter arm located within the enclosure and attached to the inner surface of the flexure plate, positioned to be visible by the at least one detector via the at least one mirror.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No : 047162-5395-00WORE -ENVISIONING FORCE-TORQUE SENSING: SIMPLE SIX-AXIS SENSOR WITH A SINGLE CAMERA AND FIDUCIAL MARKERSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application No. 63 / 721,791 filed on November 18, 2024, incorporated herein by reference in its entirety.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0002] This invention was made with government support under Grant Numbers 1928448, 1900681 and 2132823 awarded by the National Science Foundation. The government has certain rights in the invention.BACKGROUND OF THE INVENTION

[0003] Robots need to be able to sense their environment as they navigate the world and carry out everyday tasks. Six-axis force / torque (FT) sensors are one of the most common ways by which robot manipulators measure contact forces, but commercially available sensors are expensive, fragile, heavy, and prone to drifts from temperature or humidity changes because they measure extremely small deformations in metal structures using delicate sensing elements.

[0004] Developing robot systems that can sense and interact with the world around them requires them to have access to forces and torques from contacts with the environment. Six-axis force / torque (FT) sensors are commonly used to equip robots with the capability of measuring all three force and three moment components and are used in a variety of areas such as manufacturing, medical robots, human-safe systems, manipulation and more (see M. Y. Cao, et al., IEEE Sensors J., 2021). In these applications, six-axis FT sensors allow robots to interactAttorney Docket No : 047162-5395-00WOwith their environment safely and adaptively, such as through impedance or admittance control. Knowing the six-axis force information - compared to measurements along a single or 3 axes -is more effective in reasoning about contacts, because the six-dimensional force / torque vector in the sensor frame can be translated to force / torque vector in a different frame (e.g., at the point of contact) with rigid body transformations (see J. O. Templeman, et al., Sens. Actuators A Phys., 2020). While six-axis FT sensors are useful, they are expensive and susceptible to noise and drift, both stemming from the delicate sensing elements. Most common sensors use transducers such as strain gages (see Y. Sun, et al., Measurement, 2015), capacitors (see U. Kim, et al., IEEE / ASME Trans. Mechatron, 2016), piezoelectric sensors (see Y. J. Li, et al., Meeh. Syst. Signal Process, 2018), or optical techniques (see O. ALMai, et al., IEEE Sensors J., 2018)(see S. Zhong, et al., Opt. Lasers Eng., 2024) that measure small deformations (on the order of 10’9m) in monolithic metal structures. The analog signals from the sensing elements are then amplified and conditioned using high-quality circuitry to be able to derive the corresponding applied force / torque value. The sensing principle in these FT sensors relies on detecting changes in their sensing modality (resistance, capacitance, electrical charge, magnetic field, or optical properties) due to extremely small deformations in the metal structure. As a result, six-axis FT sensor measurements can get noisy or corrupted from electromagnetic interference as well as drift due to changes in environmental factors such as temperature and humidity (see M. Y. Cao, et al., IEEE Sensors J., 2021)(see J. O. Templeman, et al., Sens. Actuators A Phys., 2020). These sensors require extended warm up time, repeated zeroing of the output over long operating durations, or the use of thermally insulating interfaces (see Operational Recommendations for Reducing F / T Sensor Output Drift). On the mechanical side, the desired high stiffness characteristics and requirements for redundant sensing elements to be placed along multiple axes add complexity and cost to the sensing structure. These sensors use intricate, monolithic sensing structures (cross-beams, flexures, and parallel members) that are usually machined out of a single metal alloy block with very tight tolerance fabrication processes to avoid any backlash or slop impacting performance. Sensing transducers are then precisely positioned on the structure in complex assembly processes (like bonding strain gages to metal beams), and subsequent calibration procedures are required for each sensor to capture the large signal variances from differences in transducer properties and placements. All these factors collectively add to the high cost of six-axis FT sensors. While some 3D-printed FT sensors have been proposed as low-costAttorney Docket No : 047162-5395-00WOalternatives with plastic structures (see N. Hendrich, et al., IEEE Access, 2020), their viscoelastic deformation behavior introduces drift and hysteresis, especially over long-term repeated use, making them unsuitable for most traditional applications.

[0005] The promise of using vision for force sensing has been explored in several ways in recent literature (see S. Zhang, et al., IEEE Sensors J., 2022)(see W. Yuan, et al., Sensors, 2017)(see M. Lambeta, et al., IEEE Robot. Autom. Lett., 2020)(see N. Kuppuswamy, et al., IEEE / RSJ Int. Conf. Intell. Robots Syst. (IROS), 2020)(see A. Yamaguchi, et al., IEEE-RAS Int. Conf.Humanoid Robots (Humanoids), 2017), albeit mostly in tactile sensing. Images offer high spatial resolution information that can be parsed to extract useful features with advances in computer vision tools (see A. Yamaguchi, et al., Adv. Robot., 2019). Especially compared to discrete, onedimensional sensing elements, like strain gages, that require complicated electronics, wiring, and assembly, cameras offer a low-cost, non-contact, and reliable approach to force sensing with easily accessible components. However, unlike common sensing transducers that can detect small deformations, vision-based sensing inherently requires structures to be more deformable in order to obtain any useful information from applied forces and moments. As a result, camerabased sensors typically have low stiffness and are limited to use at the ends of kinematic chains (such as fingertips, as opposed to the wrist or joints) where the compliance does not significantly impact the robot’s functionality. These sensors also typically capture images to infer local contact geometry and features, although some works have looked at inferring forces / torques from the sensor images directly (see H. T. Suh, et al., IEEE / RSJ Int. Conf. Intell. Robots Syst. (IROS), 2022)(see C. Zhang, et al., IEEE / RSJ Int. Conf. Intell. Robots Syst. (IROS), 2022)(see A. Yamaguchi, et al., IEEE-RAS Int. Conf. Humanoid Robots (Humanoids), 2016).

[0006] Thus, there is a need in the art for a low cost, high-stiffness, and reliable six-axis FT sensing system.SUMMARY OF THE INVENTION

[0007] Described herein is a sensor device comprising a mounting plate, a flexure plate positioned parallel to the mounting plate having an inner surface facing the mounting plate and an outer surface opposite the inner surface, a wall connected to the flexure plate and theAttorney Docket No : 047162-5395-00WOmounting plate, forming an enclosure between the flexure plate and the mounting plate, at least one detector disposed on the mounting plate within the enclosure, an end-effector connected to the outer surface of the flexure plate and configured to transfer applied forces to the flexure plate, at least one mirror attached to the inner surface of the flexure plate opposite the at least one detector, and at least one emitter arm located within the enclosure and attached to the inner surface of the flexure plate, positioned to be visible by the at least one detector via the at least one mirror. In some embodiments, the at least one detector is a camera. In some embodiments, the at least one emitter arm comprises at least one marker. In some embodiments, the at least one marker comprises a patterned image. In some embodiments, the at least one marker comprises a fiducial marker tag. In some embodiments, the flexure plate comprises a plurality of openings configured to define a plurality of beams within the flexure plate. In some embodiments, the sensor device further comprises a top cover positioned over and secured on the outer surface of the flexure plate.

[0008] In some embodiments, the flexure plate is configured to deform under applied forces and torques. In some embodiments, the flexure plate is configured to allow for lever mechanical amplification. In some embodiments, the at least one emitter arm is configured to magnify a deformation of the flexure plate via lever mechanical amplification. In some embodiments, the at least one mirror is secured in place using a mirror support. In some embodiments, the at least one detector is isolated from the motions or forces on the flexure plate. In some embodiments, the sensor device further comprising at least one Light Emitting Diode (LED) positioned within the enclosure. In some embodiments, the enclosure is attached to the flexure plate and mounting plate using an attachment technique selected from a group comprising adhesive bonding, screwing, soldering, welding, riveting, or clamping. In some embodiments, the wall of the enclosure comprises an opening and a cable grommet positioned within the opening.BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The foregoing purposes and features, as well as other purposes and features, will become apparent with reference to the description and accompanying figures below, which are includedAttorney Docket No : 047162-5395-00WOto provide an understanding of the invention and constitute a part of the specification, in which like numerals represent like elements, and in which:

[0010] FIG. 1A-1B depict perspective views of an example sensor device rendering.

[0011] FIG. 1C depicts an isometric view of an example top cover.

[0012] FIG. 2A-2B depict a design rendering of the components of an example sensor device.

[0013] FIG. 3A depicts example components of a sensor device.

[0014] FIG. 3B is an example sensor device mounted on a UR5 robot arm.

[0015] FIG. 4A-4C depict amplification principles of an example sensor device.

[0016] FIG. 5 depicts response plots obtained during the characterization of an example sensor device.

[0017] FIG. 6 depicts plots illustrating the absolute error in measured forces and moments of a sensor device.

[0018] FIG. 7 depicts an example sensor device and a reference ATI sensor.

[0019] FIG. 8A-8B depicts a sensor device performing example robot arm tasks.

[0020] FIG. 9A-9B depict plots of closed-loop responses of a sensor device during example robot arm tasks.

[0021] FIG. 10 depicts an example experimental setup used for calibrating and characterizing an example sensor device.

[0022] FIG. 11 depicts dimensions of an example sensor device.

[0023] FIG. 12 depicts a plot of mean absolute error on sensor test data as a function of number of calibration samples.Attorney Docket No : 047162-5395-00WO

[0024] FIG. 13 depicts an example Finite Element Analysis (FEA) of the stress distribution on an example sensor device under maximum force / moment in different directions.DETAILED DESCRIPTION

[0025] The following discussion omits or only briefly describes conventional features that are apparent to those skilled in the art. Those of ordinary skill in the pertinent arts may thus recognize that other elements may be desirable and / or necessary to implement the devices, systems, and / or methods described herein. It is noted that various embodiments are described in detail with reference to the drawings. Reference to these various embodiments does not limit the scope of the claims attached hereto. Additionally, any embodiments set forth in this specification are intended to be non-limiting and merely set forth some of the many possible implementations for the appended claims. Further, particular features described herein can be used in combination with other described features in each of the various possible combinations and permutations. As such, it is understood that this detailed description is exemplary and explanatory only and is not restrictive of the broad inventive concepts upon which the embodiments disclosed herein are based.

[0026] Unless otherwise specifically defined herein, all terms are to be given their broadest reasonable interpretation. This includes meanings implied from the specification as well as meanings understood by those skilled in the art and / or as defined in dictionaries, treatises, etc.

[0027] It is noted that, as used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless otherwise specified. The term “includes” and / or “including,” when used in this specification, specify the presence of stated features, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof.

[0028] Relative terms such as “horizontal,” “vertical,” “up,” “down,” “top,” and “bottom” as well as derivatives thereof (e.g., “horizontally,” “downwardly,” “upwardly,” etc.) should be construed to refer to the orientation as then-described or as shown in the drawing figure under discussion. These relative terms are for convenience of description and normally are not intendedAttorney Docket No : 047162-5395-00WOto require a particular orientation in actuality. Terms including “inwardly” versus “outwardly,” “longitudinal” versus “lateral,” and the like are to be interpreted relative to one another or relative to an axis of elongation, or an axis or center of rotation, as appropriate. Terms concerning attachments, coupling and the like, such as “connected” and “interconnected,” refer to a relationship wherein structures are secured or attached to one another either directly or indirectly through intervening structures, as well as both movable or rigid attachments or relationships, unless expressly described otherwise. The phrases “operatively” or “operably connected” indicates such an attachment, coupling, or connection that allows the pertinent structures to operate as intended by virtue of that relationship.

[0029] Reference throughout the specification to “one embodiment,” “an embodiment,” or “some embodiments” means that a particular feature, structure, or characteristic described in connection with at least one example of the subject matter is included in at least one example of the subject matter disclosed. Thus, the appearance of the phrases “in one embodiment,” “in an embodiment,” or “in some embodiments” in various places throughout the specification is not necessarily referring to the same embodiment. Further, the particular features, structures, or characteristics of “one embodiment,” “an embodiment,” or “some embodiments” may be combined in any suitable manner with each other to form additional embodiments of such combinations. It is intended that embodiments of the disclosed subject matter cover modifications and variations thereof. Terms such as “first,” “second,” “third,” etc., merely identify one of a number of portions, components, steps, operations, functions, and / or points of reference as disclosed herein, and likewise to not necessarily limit embodiments of the present disclosure to any particular configuration or orientation.

[0030] Moreover, throughout this disclosure, various aspects can be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the disclosure. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., as well as individual numbers within that range, for example, 1, 2, 2.7, 3, 4, 5, 5.3, 6, and anyAttorney Docket No : 047162-5395-00WOwhole and partial increments therebetween. This applies regardless of the breadth of the range. As used herein, the term “about” in reference to a measurable value, such as an amount, a temporal duration, and the like, is meant to encompass the specified value variations of plus or minus 20%, plus or minus 10%, plus or minus 5%, plus or minus 1%, and plus or minus 0.1% of the specified value, as such variations are appropriate and fit within the confines of a functional system.

[0031] The terms “proximal,” “distal,” “anterior,” “posterior,” “medial,” “lateral,” “superior,” and “inferior” are defined by their standard usage indicating a directional term of reference. For example, “proximal” refers to a position that is situated nearer to the center of a body or point of attachment or interest. In another example, “anterior” refers to the front of a body or structure, while “posterior” refers to the rear of a body or structure, in relation to a relative viewpoint. In another example, “medial” refers to the direction towards the midline of a body or structure, and “lateral” refers to the direction away from the midline of a body or structure. In some embodiments, “lateral” or “laterally” may refer to any sideways direction. In another example, “superior” refers to the top of a body or structure, while “inferior” refers to the bottom of a body or structure. It should be understood, however, that the directional term of reference may be interpreted within the context of a specific body or structure, such that a directional term referring to a location in the context of the reference body or structure may remain consistent as the orientation of the body or structure changes.

[0032] Described herein is a device that measures six-axis forces / torques with a far more robust and ubiquitous sensing element (e.g. a camera) while maintaining a high level of accuracy, stiffness, sensitivity, and range. Described herein is a six-axis Force / Torque sensor, VisualFT, that uses a single camera to track the displacement of fiducial marker tags mounted to a planar flexure plate. The perceived motion of the markers is magnified through lever and mirror amplification mechanisms, so that the sensor retains a highly stiff structure and a large range, while having sufficient sensitivity to applied forces and moments. Because marker motion is mechanically amplified as seen through the camera, it eliminates the need for high-quality signal conditioning circuitry and precise calibration processes that are typically required in strain gage or capacitive force / torque (FT) sensors. Moreover, analog signal-based sensors may be more susceptible to noise from electromagnetic interference compared to digitized camera-basedAttorney Docket No : 047162-5395-00WOsensors. The deformations in the VisualFT sensor structure are localized to a metal flexure plate, not any viscoelastic components, and the resulting linear elastic deformation behavior is simple to model for accurate estimation of applied forces and moments. The performance of the sensor in long duration test runs over multiple days shows less than 1.5% relative error about all six axes compared to a commercial reference FT sensor. This also applies to the sensor in closed-loop force tracking tasks that mimic real-world applications. The components of the sensors are easily fabricated or obtained off-the-shelf.

[0033] The disclosed system combines the advantages of high-stiffness FT sensors with the accessibility and low-cost of modern camera-based sensing. The disclosed sensor uses readily available or easily fabricated components, and the stiffness is achieved by using a flexure plate made of, for example metal, and two mechanical amplification principles - lever and mirror -that augment the perceived motion of the markers without requiring large deformations at the point where the forces / torques are applied. The mechanically amplified motion of the markers may be tracked by a single, off-the-shelf camera module. The sensing principle of the VisualFT sensor does not rely on small magnitude deformations or delicate sensing transducers, which lends several advantages over commercial FT sensors.

[0034] Limited prior works have proposed using fiducial markers to derive proprioceptive feedback and force / torque values (see R. Ouyang, et al., 2020 IEEE Tnt. Conf. Robot. Autom. (ICRA), 2020)(see J. C. S. Ting, et al., Deployable Multimodal Machine Intelligence:Applications in Biomedical Engineering, 2023). These works still use soft, deformable elastic structures with a lot of compliance and low stiffness so that the sensor is sufficiently sensitive to small forces / torques. Instead, the VisualFT sensor uses mechanical amplification that magnifies marker motions outside of the direct force pathway so that the sensor can have high stiffness while still having high sensitivity. This approach also allows for a larger range of applied forces and moments, and the sensor is accurate to within a 1.5% error relative to a commercial FT sensor.

[0035] Referring now to FIG. 1A-1B, a sensor device 100 is shown. In some embodiments, the sensor device 100 may comprise a mounting plate 110, at least one detector 120 disposed on the mounting plate 110, an enclosure 130 having a first end 131 and a second end 132, wherein the first end 131 of the enclosure 130 is secured on the mounting plate 110, and a flexure plate 140Attorney Docket No : 047162-5395-00WOhaving an inner surface 141 (a first side 141), and an outer surface 142 (a second side 142), wherein the first side 141 of the flexure plate 140 is secured to the second end 132 of the enclosure 130. In some embodiments, the device 100 may further comprise an end-effector 280 (see FIG. 3B) disposed on the second side 142 of the flexure plate 140, a plurality of mirrors 150 located within the enclosure 130 and attached to the first side 141 of the flexure plate 140, and at least one emitter arm 170 located within the enclosure 130 and attached to the first side 141 of the flexure plate 140. In some embodiments, the at least one detector 120 is a camera. In some embodiments, the at least one emitter arm 170 comprises a marker holder 271 (see FIG. 2A) and at least one marker 272 (see FIG. 2A). In some embodiments the at least one marker 272 may be a patterned image 272. In some cases, the patterned image 272 of the at least one emitter arm 170 may be a fiducial marker tag 272 (see FIG. 2A) held by the marker holder 271. In some embodiments, the at least one marker 272 comprises a quick-response (QR) code. In some embodiments, the at least one emitter arm 170 may be a plurality of emitter arms, the marker holder 271 may be a plurality of holders 271 and the fiducial marker tag 272 may be a plurality of fiducial marker tags 272. In some embodiments, the plurality of mirrors 150 may be at least one mirror 150. In some embodiments, the fiducial marker tags / marker 272 may be located within the enclosure and attached to the first side 141 of the flexure plate 140, positioned such that the marker 272 is visible to at least one detector via the at least one mirror 150.

[0036] In some embodiments, the device 100 may further comprise atop cover 160 positioned over and secured on the second side 142 of the flexure plate 140. FIG. 1C depicts an isometric view of an example top cover 160. In some embodiments, the top cover 160 may have a raised section 161 located at the bottom and center of the top cover 160. In some embodiments, the raised section 161 creates a gap between a portion of the top cover 160 and the flexure plate 140. The raised section 161 may be configured to prevent or reduce interference with the deflection motions of the plurality of emitter arms 170 on the flexure plate 140. In some embodiments, the top cover 160 may further comprise of a plurality of apertures 162 configured to help the device 100 engage with an end effector 280 (see FIG. 3A). In some embodiments, the flexure plate 140 may comprise a plurality of apertures 146 configured to help the flexure plate 140 engage with the end effector 280. Securing the end effector 280 to the device 100 may be done using any suitable attachment technique. In some embodiments, the top cover 160 is configured to cover the top of the device 100 in order to avoid dust and debris from falling inside the device 100.Attorney Docket No : 047162-5395-00WO

[0037] The flexure plate 140 may be configured to deform under applied forces and torques. In some embodiments, the at least one detector 120 is configured to track the deformation of the flexure plate 140 using the at least one emitter arm 170 and the plurality of mirrors 150. In some embodiments, the flexure plate 140 may comprise a plurality of beams 143. The device 100 may use readily available or easily fabricated components, and the stiffness may be achieved by using the flexure plate 140 and two mechanical amplification principles - lever and mirror. These mechanical amplification principles may augment the perceived motion of the plurality of fiducial marker tags 272 without requiring large deformations at the point where the forces / torques are applied. The mechanically amplified motion of the plurality of fiducial marker tags 272 may be tracked by the at least one detector 120, such as a USB camera module. The flexure plate 140 may be configured to allow for lever mechanical amplification.

[0038] In some embodiments, the sensing principle of the sensor device 100 may not rely on small magnitude deformations or delicate sensing transducers, which lends several advantages over commercial FT sensors: the sensing structure does not need tight tolerance fabrication processes; the assembly process avoids complicated steps like strain gage bonding; the simpler calibration procedure does not have to manage large variances from sensing transducer placements; there may be no signal conditioning or amplifying circuitry; the measurements remain stable over long durations of use or temperature variations; and the digitized camera signals are less prone to noise from electromagnetic interference.

[0039] In some embodiments, the USB camera module 120 may be, for example, a monochrome (black / white) image sensor having a resolution in the range of approximately 0.5 megapixels to 5 megapixels (e.g., 800H x 600V to 2592H x 1944V), and a frame rate in the range of about 30 Hz to 240 Hz. In some embodiments, the camera module 120 may employ a global shutter, which may reduce motion blur relative to typical rolling-shutter cameras. The USB camera module 120 may be implemented using commonly available webcam components and may not require any specialized or proprietary hardware. Such a configuration may provide a cost-effective and easily sourced imaging solution. The ease of use of a single USB cable to connect the force / torque sensor device 100 to a computer is also notably different compared to typical force / torque sensors on the market (e.g. ATI FT sensors) that typically require additional electronics for data acquisition and delicate thin wires originating from the sensing transducers inside the sensor 100.Attorney Docket No : 047162-5395-00WOIn some embodiments, using the USB camera module 120 may result in a straightforward detection method that uses averaged brightness of pixel values instead of marker detection, further simplifying any post-processing on the images obtained, and reduces the computation overload for faster output. This approach may remove the need of each of the plurality of marker tags 272 being well-lit, in focus, and clearly visible in the camera image. Thus, this approach reduces noisy force / torque estimates from the sensor and reduces the amount of data needed to initially calibrate the sensor.

[0040] In some embodiments, the brightness of a set of pixels over each of the plurality of fiducial marker tags 272 are recorded and used to predict the forces / torques. When forces / torques are applied, the black / white plurality of fiducial markers 272 may move, and subsequently the brightness of each pixel changes - getting darker or brighter (depending on whether a white or black portion of at least one fiducial marker 272 covers the pixel). Alternative patterns may be used with this approach, including, but limited to, a simple chessboard pattern. In some embodiments, the pattern does not need to be clearly visible in the image. The plurality of fiducial marker tags 272 may also be used as a backup detector to conduct marker detection. The working principle of the force / torque sensor device 100, which enables the high-stiffness and high-resolution advantages, is the underlying mechanical amplification of lever and the mirror amplification. This sensing structure of the device 100 may be used with any number of detector / emitter arm combinations, not just a camera 120 and a patterned image 272. For instance, an array of analog sensors, such as photo resistors or time-of-flight sensors, may be used as the at least one detector 120. In some embodiments, the at least one emitter arm 170 may comprise an LED which may be used to emit light at the end of the at least one emitter arm 170, and the at least one detector 120 may be a light sensor. In some embodiment, the at least one emitter 170 may be a flat surface and a proximity or distance sensor may be used as the at least one detector 120. In some embodiments, the camera module 120 may be used for the ease-of-use and minimal post-processing requirements. However, the sensor device 100 may easily be adapted to other detector / emitter arms for specific performance criteria, for example, high Hz sampling frequency with analog sensors, or higher sensitivity with a higher resolution camera.

[0041] In some embodiments, the flexure plate 140 may comprise a plurality of openings 144. In some embodiments, the plurality of openings 144 may be configured to define a plurality ofAttorney Docket No : 047162-5395-00WObeams 143 within the flexure plate 140. In some embodiments, the size of the plurality of openings 144 may vary. The plurality of openings 144 may be sized to create a gap between the plurality of beams 143 and the remainder of the flexure plate 140 so as to prevent the plurality of beams 143 from colliding with other portions of the flexure plate 140. In some embodiments, the plurality of openings 144 may be configured to partially see inside the device 100. In some embodiments, the plurality of openings 144 may be reduced in size to prevent inadvertent damage to the camera module 120 from any debris or dirt falling in through the top or second side 142 of the device 100. In some embodiments, the top cover 160 may be used for additional coverage and for more consistent lighting inside the sensor device 100. The thickness of the flexure plate 140 may range from 1 mm - 10 mm. In some embodiments, the thickness of the flexure plate 140 is determined based on the range of forces being measured. A thinner flexure plate 140 may be more sensitive to small forces, whereas a thicker flexure plate 140 may offer increased resistance to deforming or breakage under larger loads. In some embodiments, the thickness of the flexure plate 140 is configured to enhance stiffness and measurement range. The overall height of the sensor device 100 may range from 12 mm - 300 mm. In some embodiments, the height of the device 100 may vary as appropriate, depending at least in part on the physical height of the at least one detector 120 and the minimum detection distance of the at least one detector 120.

[0042] In some embodiments, the at least one emitter arm 170 may be a plurality of emitter arms 170. In some embodiments, the at least one emitter arm 170 is attached to an end 145 of each of the plurality of beams 143. The at least one emitter arm 170 may be configured to magnify the deformation of the flexure plate 140 via lever mechanical amplification. In some embodiments, the plurality of emitter arms 170 comprises the plurality of beams 143. In some embodiments, the plurality of mirrors 150 are positioned at angles ranging from 0-90 degrees. The at least one mirror 150 may be positioned in a flat orientation or in any suitable orientation that allows the at least one mirror 150 to capture the at least one emitter arm 170. In some embodiments, the plurality of mirrors 150 may be, for example, thin plastic adhesive mirror sheets. The plurality of mirrors 150 may be configured to amplify the displacement of the at least one emitter arm 170. In some embodiments, the plurality of mirrors 170 are secured in place using a mirror support 151. In some embodiments, the at least one detector 120 is isolated from the motions or forces on the flexure plate 140. In some embodiments, the flexure plate 140 and mounting plate 110 mayAttorney Docket No : 047162-5395-00WObe attached to the enclosure 130 using any suitable attachment technique, including, but not limited to, adhesive bonding, screwing, soldering, welding, riveting, or clamping. In some embodiments, two or more of the components of the device 100 may be created from one single piece of material. For example, the enclosure 130 and the mounting plate 110 may be machined as one piece on a mill from one piece of metal. In some embodiments, the components of the device 100 may be made as separate parts to facilitate manufacturing, reduce costs, and / or enable replacement or modification of individual components without requiring alteration of the entire device 100. In some embodiments, the enclosure 130 may comprise an opening configured to house a cable grommet 111.

[0043] Referring now to FIG. 2A-2B, another embodiment of a sensing device 200 is shown. FIG. 2A-2B depict a design rendering of the internal components of the device 200. FIG. 2B shows an exploded view of the sensor device 200 laying out the components of the device 200. In some embodiments, the sensor device 200 may comprise a mounting plate 210, at least one detector 220 disposed on the mounting plate 210, an enclosure 230 having a first end 231 and a second end 232, wherein the first end 231 of the enclosure 230 is secured on the mounting plate 210, and a flexure plate 240 having a first side 241 and a second side 242, wherein the first side 241 of the flexure plate 240 is secured to the second end 232 of the enclosure 230. In some embodiments, the device 200 may further comprise an end-effector 280 (see FIG. 3B) disposed on the second side 242 of the flexure plate 240, a plurality of mirrors 250 located within the enclosure 230 and attached to the first side 241 of the flexure plate 240, and at least one emitter arm 270 located within the enclosure 230 and attached to the first side 241 of the flexure plate 240.

[0044] In some embodiments, the at least one detector 220 is a camera. In some embodiments, the at least one emitter arm 270 comprises a marker holder 271 and a patterned image 272. In some cases, the patterned image 272 of the at least one emitter arm 270 may comprise a fiducial marker tag 272 held by the marker holder 271. In some embodiments, the at least one emitter 270 may be a plurality of emitter arms 270, the fiducial marker tag 272 may be a plurality of fiducial marker tags 272 and the marker holder 271 may be a plurality of marker holders 271.Attorney Docket No : 047162-5395-00WO

[0045] In some embodiments, the flexure plate 240 may comprise a plurality of openings 244 and a plurality of beams 243. In some embodiments, the size of the plurality of openings 244 may vary. The plurality of openings 244 may be sized to create a gap between the plurality of beams 243 and the remainder of the flexure plate 240 so as to prevent the plurality of beams 243 from colliding with other portions of the flexure plate 240. In some embodiments, the flexure plate 240 may be configured to deform under applied forces and torques. The flexure plate 240 may be made from any suitable materials, such as metal or fiber-reinforced polymers. In some embodiments, the flexure plate 240 undergoes linear elastic deformations when external force / torques are applied. The deformations may be magnified at the plurality of fiducial marker tags 272 through a lever amplification mechanism. In some embodiments, the at least one detector 220 is configured to track the deformation of the flexure plate 240 using the at least one emitter arm 270 and the plurality of mirrors 250.

[0046] In some embodiments, the deformations are carried to the plurality of fiducial marker tags 272 or fiducial marker clusters through the plurality of beams 243 and / or the plurality of marker holders 271. The flexure plate 240 may be configured to allow for lever mechanical amplification. The device 200 may use readily available or easily fabricated components, and the stiffness of the device 200 may be achieved by utilizing the flexure plate 240 and the two mechanical amplification principles - lever and mirror. In some embodiments, these mechanical amplification principles augment the perceived motion of the plurality of fiducial marker tags 272 without requiring large deformations at the point where the forces / torques are applied. The mechanically amplified motion of the plurality of fiducial marker tags 272 may be tracked by the at least one detector 220 via the mirrors 250. In this configuration, the detector 220 may be a single, off-the-shelf camera module. In some embodiments, the camera module 220 may be placed at the base of the sensor device 200, mechanically isolated from any motions or forces on the flexure plate 240 or the plurality of fiducial marker tags 272. In some embodiments, the camera module 220 may track the reflected image of the plurality of fiducial marker tags or clusters 272 through the plurality of mirrors 250.

[0047] FIG. 2B additionally shows the planar flexure plate 240 that enables the lever mechanism for amplification of the motion of the plurality of fiducial marker tags 272 while maintaining small deformation at the point where the forces and torques are applied. A mirror support 251Attorney Docket No : 047162-5395-00WOmay be positioned underneath the flexure plate 240 and may orient the plurality of mirrors 250, which may be adhesively attached, at a slight angle. In some embodiments, the plurality of mirrors 250 may be made from, for example, thick acrylic mirror sheets, and angled at a range of 0-90 degrees. The angled placement of the plurality of mirrors 250 may further amplify the perceived marker tag 272 displacement.

[0048] The sensing principle of the sensor device 200 may not rely on small magnitude deformations or delicate sensing transducers, which lends several advantages over commercial FT sensors. The sensor device 200 may use the mechanical amplification schemes to convert the small deformations at the point where forces and torques are applied into motions of the plurality of fiducial marker tags 272 that may be discerned as a change in the image signal. Commercial FT sensors use sensing transducers such as strain gages and capacitors that measure a change in signal from small (on the order of 10'9m) deformations in the elastic structure, and the signal is then usually amplified and conditioned to calculate the equivalent force / torque measurement. Relying on small deformations and changes in signal has the challenges of requiring precisely machined, monolithic components and being susceptible to mechanical and electromagnetic interference, or environmental factors such as temperature and humidity. A camera-based sensor would avoid these issues by relying on much larger magnitude changes in the image signal. To obtain a larger signal magnitude for the same magnitude of forces / torques, however, the deformations at the point of force / torque application would also need to be large i.e., the sensor would need to have a much lower overall stiffness which would be highly undesirable. To address this, the device 200 uses amplification mechanisms, for example lever and mirror amplification, to magnify the perceived deformation of the plurality of fiducial markers tags 272 in the view of the camera 220 and requires only small deflections at the point of force / torque application. Thus, the sensor device 200 has both high stiffness and high resolution. The camera module 220 may be coupled to a separate microprocessor, for example, a Raspberry Pi, which is configured to read image data from the camera module at a rate ranging from 30Hz to 120Hz, for example, 56Hz when operating at the full resolution.

[0049] The at least one detector 220 may detect the pose of the plurality of fiducial marker tags 272 in the image and then calculate the applied forces and torques based on the pose of each of the plurality of fiducial marker tags 272. This may require each marker tag 272 to be well-lit, inAttorney Docket No : 047162-5395-00WOfocus, and clearly visible in the camera image. If the plurality of fiducial marker tags 272 are slightly blurry or the lighting conditions aren’t ideal, the resulting detected marker pose may be noisy, which can in turn result in noisy force / torque estimates from the sensor 200. In some embodiments, the device 200 may further comprise at least one light emitting diode (LED) 290 (see FIG. 3A) configured to maintain consistent lighting within the enclosure 230. The at least one LED 290 may be confined within and attached to the inner walls of the enclosure 230. The enclosure 230 covers the internal components of the sensor device 200 and provides structural support to the overall sensor device 200 so that any deformations induced by the applied forces / torques only go through the flexure plate 240 and displace the plurality of marker tags 272. The components used in the sensor device 200 may easily be sourced or manufactured with rapid prototyping methods such as 3D printing and waterjet cutting, and the total cost of the sensor may add up to less than USD 75.

[0050] In some embodiments, the at least one emitter arm 270 is attached to an end 245 of each of the plurality of beams 243. The at least one emitter arm 270 may be configured to magnify the deformation of the flexure plate 240 via lever mechanical amplification. The plurality of mirrors 250 may be angled and configured to amplify the displacement of the at least one emitter arm 270. In some embodiments, the plurality of mirrors 270 are secured in place using the mirror support 251. In some embodiments, the at least one detector 220 is isolated from the motions or forces on the flexure plate 240 and the enclosure 230 may comprise an opening configured to house a cable grommet 111 (see FIG. IB).

[0051] Referring now to FIG. 3A-3B, example components of a sensor device 200 (FIG. 3 A) and an example sensor device 200 mounted on a UR5 robot arm (FIG. 3B) are shown. The components of the device 200 may be put together using any suitable attachment technique, including, but not limited to, adhesive bonding, screwing, soldering, welding, riveting, or clamping. In some embodiments, the sensor 200 is designed with the goal of maintaining a simple structure with easy-to-obtain components, and an accessible fabrication and assembly. In some cases, the device 200 may take under 15 minutes to assemble. As such, most parts for the sensor device 200 may be acquired off-the-shelf and any custom parts may use quick prototyping fabrication methods such as 3D printing or waterjet cutting. The sensor device 200 may use a single camera module to track the plurality of fiducial marker tags 272, and a view of the cameraAttorney Docket No : 047162-5395-00WO220 may be seen in the top right of FIG. 3B. Tn some embodiments, forces / torques may be applied at the black peg / end effector 280 mounted on the top or second side 242 of the flexure plate 240.

[0052] The overall dimensions of the prototype sensor may be designed to be similar to an ATI Gamma sensor. The outer enclosure 230 may have a height ranging from 10 mm - 200 mm and may be made from any suitable material, for example, 6061 Aluminum round tube stock, with a plurality of drilled holes to mount the flexure 240 and mounting 210 plates, and a slot for cables to run through. In some embodiments, the height of the enclosure 230 may take any suitable height, depending at least in part on the physical height of the at least one detector 120 and the minimum detection distance of the at least one detector 120. In some embodiments, the enclosure 230 may be 3D printed for lower-force applications. The height of the overall sensor device 200 may be determined by the minimum focus distance of the camera 220. By using the plurality of mirrors 250 and the reflected image of the plurality of fiducial marker tags 272, the height may be cut by roughly half (since the reflected image is perceived to be behind the mirror plane). In some embodiments, the flexure 240 and mounting 210 plates may be wateijet-cut from materials, such as 6061 Aluminum sheet 353 stock. In some embodiments, components of the device 200 may be 3D printed. Additionally, a Raspberry Pi 5 may be used to detect the plurality of marker tags 272 on images from the camera module 220 and to compute forces, although the images may directly be transmitted from the camera module 220 to the client computer.

[0053] In some embodiments, the plurality of mirrors 250 may be also readily available off-the-shelf and may have an adhesive backing to stick to the mirror support 251 which may orient each of the plurality of mirrors 250 at a slight angle. The angled plurality of mirrors 250 may further extend the viewing field for the camera 220 in order that the plurality of fiducial marker tags 272 may be placed as far away as possible from the end 245 of at least one of the plurality of beams 243 and maximize the lever amplification.

[0054] Referring now to FIG. 4A-4C, amplification principles of the sensor device 200 are shown. The first level of mechanical amplification comes from the lever amplification effect of a cantilevered beam. FIG. 4A shows a simplified cantilever beam model of one of the plurality of emitter arms 270 (one of the three lever arms) on the flexure plate 240. As shown in FIG. 4A, theAttorney Docket No : 047162-5395-00WOpoint where the forces / torques are applied (P, also known as “effort”) is much closer to the base of each deflection beam (O or “fulcrum”) compared to the point where the fiducial markers are attached (Q or “load”). The forces and torques may be transmitted through point P (effort), closer to the base of at least one of the plurality of beams 243 marked by O (fulcrum). In some embodiments, the angled plurality of mirrors 250 are attached underneath a center knob.Subsequently, a portion of the beam (from point P to the base O) may elastically deform under external forces / torques, and as the beam undergoes this deflection, at least one of the plurality of fiducial marker tags 272 experiences a larger deflection (AQ), especially compared to the deflection at point P (AP). The plurality of fiducial marker tags 272 may experience a larger deformation since they are at placed at the end of at least one of the plurality of fiducial marker holds 271 at point Q (load). This class of levers where the effort / input is applied between the fulcrum / base and the load / output is often referred to as speed or distance-multiplier levers, since they have a larger distance advantage than, for example, levers with the fulcrum in the middle. The lever amplification ratio may be characterized through Finite Element Analysis (FEA).

[0055] FIG. 4B shows the deformation of the flexure plate 240 under an example load of 10 N of vertical load in a FEA analysis. The FEA analysis shows the behavior of the plurality of emitter arms 270 and the lever amplification principle in action. The points at the ends of the plurality of holders 271 at Q where the plurality of fiducial marker tags 272 are placed undergo the maximum deformation (URES regions: 0.1244 mm - 0.1748 mm), notably larger than the parts of the flexure plate 240 at P where the load is applied (URES regions: 0.0000 mm - 0.0699 mm).

[0056] The sensor device 200 may not use separate electronic signal amplification unlike commercial sensors. In some embodiments, the cross-beam structure using the plurality of beams 243 may need to be paired with mechanical amplification mechanisms that magnify the motions of the plurality of fiducial marker tags 272. Amplification mechanisms are broadly categorized into two types - lever and triangle (see F. Chen, et al., IEEE Access, 2020). The former uses the mechanical advantage offered by levers, whereas the latter uses the large displacements of a four-bar linkage near its singular configuration. The device 200 may use the lever amplification mechanism for its simplicity, consistent amplification ratio, and predictable stiffness properties. Triangle mechanisms have a more nonlinear amplification ratio and can have some indeterminate motions, especially orthogonal to the output. In some embodiments, among the different classesAttorney Docket No : 047162-5395-00WOof levers, Class 3 lever, also known as displacement or speed-multiplier lever (where the “effort” is in between the “fulcrum” and the “load,” like tweezers, stapler, or broom) may be used. Class 3 lever provides the maximum displacement amplification for a given overall length.

[0057] For each of the plurality of emitter arms 270 in the flexure plate 240, which may behave as a cantilevered beam, the end / base 245 of at least one of the plurality of beams 243 behaves like the “fulcrum” point of the at least one of the plurality of emitter arms 270. Next to the end 245, each of the plurality of beams 243 may have a branch that connects the plurality of beams 243 to the center - this is the point where the external forces / torques are applied (“effort” point). The portion of at least one of the plurality of beams 243 that undergoes stress and elastic deformation may concentrate between the end 245 and the point where the branch extends to connect the at least one beam 243 to the center. The remainder of the at least one of the plurality of beams 243, along with the marker holder 271 that may be attached to the end 245 of the at least one of the plurality of beams 243, may simply serve to extend the at least one emitter arm 270 or ‘lever’ arm, as much as possible to the at least one of the plurality of marker tags 272 (“load” point of the ‘lever’). As a result, the small elastic deformations between the “fulcrum” and the “effort” are amplified in displacement motion to the “load.” In some embodiments, the design of the plurality of emitter arms 270, where forces / torques are applied at a point close to the end 245 or base (for high stiffness) but induced deflections are carried further away (for high resolution), is the basis of the speed-multiplier lever amplification mechanism used in the sensor device 200.

[0058] In some embodiments, to maximize the lever amplification effect, the plurality of markers tags 272 (“load”) may be placed at a point furthest away from the end 245 (“fulcrum”) of each of the plurality of beams 243. For a circular plate, the point may be diametrically opposite to the end 245 of the at least one of the plurality of beams 243. In some embodiments, the described arrangement would not be possible in a plane of the flexure plate 240 for a, for example, three-beam architecture, without the plurality of beams 243 intersecting with each other. In some embodiments, the plurality of marker holders 271 may enable the increased distance between the end 245 and the marker tag 272 by moving the plurality of fiducial marker tags 272 to a different plane below the flexure plate 240. The device 200 may be designed such that the plurality of fiducial marker tags 272 for each of the plurality of beams 243 is as farAttorney Docket No : 047162-5395-00WOdiametrically away as possible from the respective ends 245 each of the plurality of beams 243, while also not obstructing the view of the camera 220 or mechanically interfering with the other plurality of beams 243 and the plurality of marker holders 271. In some embodiments, the plurality of marker holders 271 may have a vertical height, which may add distance between the end 245 of the at least one of the plurality of beams 243 and at least one of the plurality of fiducial maker tags 272, further increasing the amplification effect even more for certain directions.

[0059] The second level of mechanical amplification may come from the deflection on the plurality of mirrors 250 and the resulting change in the reflected image of the plurality fiducial marker tags 272. Some optoelectronic sensors have used mirrors to redirect light beams, but not for leveraging any amplification effect (see Y. Noh, et al., Sensors, 2016)(see G. Palli, et al., Sens. Actuators A Phys., 2014). The device 200 uses the plurality of mirrors 250, for example three mirrors, each angled at, for example, 6 degrees with respect to the central axis so that each of the plurality of mirrors 250 points the camera view to one of the plurality of fiducial marker tags 272. In some embodiments, the marker cluster pattern and calibration process of the plurality of fiducial marker tags 272 are chosen to minimize error. The plurality of mirrors 250 may be attached to the mirror support 251 which orients each of the plurality of mirrors 250 at the prescribed angle. The mirror support 251 may be rigidly attached underneath the flexure plate 240 at the point where forces / torque are applied. Since the camera 220 is always tracking the plurality of fiducial marker tags 272 through the mirrored reflections, even small angular rotations of at least one of the plurality of mirrors 250 may result in large perceived changes in the plurality of fiducial marker tag 272 locations in the reflected image. FIG. 4C shows that the deformations in the flexure plate 240 induce small changes in the mirror position, which further amplifies the perceived motion of the reflected fiducial marker tags 272 in the camera view. Translations of a mirror parallel to its own plane does not result in a change in the reflected image positions. However, the angled plurality of mirrors 250, in some embodiments, may not share any common plane. So, any motion of the mirror support 251 (translation or rotation) may induce a notable change in the reflected marker image, even if the plurality of marker tags 272 don’t move significantly with respect to the camera 220. The plurality of mirrors 250 may also allow the overall height of the sensor 200 design to be reduced by roughly half, because the plurality of mirrors 250 may be placed at half the minimum focus distance of the camera 220.Attorney Docket No : 047162-5395-00WO

[0060] The device 200 may have linear deformations and the advantage of long duration stability. The device 200 may remain stable over long durations of use and maintain calibration over several days attributed to the use of the camera 220 as the sensing module, and the choice of materials like aluminum for the deforming component, the flexure plate 240. In some embodiments, since the camera 220 is mechanically isolated from any moving components, and unlike delicate sensing elements (capacitors, strain gages, optoelectronics and so on) used in commercial FT sensors, it has negligible drift or hysteresis characteristics, especially from any changing temperature or humidity conditions. The digitized camera signal may also be resistant to electromagnetic interference, which affects analog transducers in commercial FT sensors. Moreover, because the deformation signal is mechanically amplified as seen by the camera 220, the sensor 200 does not need any sophisticated amplification or conditioning circuitry. Some low-cost FT sensors use plastic or other viscoelastic materials in their deformation structures (see N. Hendrich, et al., IEEE Access, 2020). However, the flexure plate 240 of the sensor 200 has a linear elastic deformation behavior, which eliminates any noticeable drift and hysteresis in the sensor output from material creep or stress relaxation. The linear elastic mechanics of the flexure plate 240 may be further helped by not attaching adhesive material to the sensor device 200, such as cement-bonded strain gages, that might add any undesirable viscous behavior or loss of deformation information from incorrect or uncured bonding. In some embodiments, the sensor device 200 may have physical overload protection such as hard stops as well as checks in the sensor code to detect wrong initial fiducial marker tag 272 positions and alert the user of any overloading-induced plastic deformation in the flexure plate 240.EXPERIMENTAL EXAMPLES

[0061] The invention is further described in detail by reference to the following experimental examples. These embodiments are provided for purposes of illustration only, and are not intended to be limiting unless otherwise specified. Thus, the invention should in no way be construed as being limited to the following embodiments, but rather, should be construed to encompass any and all variations which become evident as a result of the teaching provided herein.Attorney Docket No : 047162-5395-00WO

[0062] Without further description, it is believed that one of ordinary skill in the art can, using the preceding description and the following illustrative embodiments, make and utilize the system and method of the present invention. The following working embodiments therefore, specifically point out the exemplary embodiments of the present invention, and are not to be construed as limiting in any way the remainder of the disclosure.

[0063] In the subsequent sections, the performance of the sensor device, VisualFT, is experimentally validated through a series of long-duration test runs after calibration over multiple days, as well as through closed loop tasks that use the sensor response for feedback control and mimic real-world applications for a force / torque (FT) sensor.Experimental EvaluationsSensor Response Characterization

[0064] To characterize the performance of the prototype VisualFT sensor, the sensor’s detected marker data (and estimated force / torque values after calibration) and ground truth force / torque data from a six-axis ATI Gamma sensor is collected. The two sensors are mounted rigidly to each other and fixed onto an optical bread board for a characterization test setup. Various combinations of forces and torques are randomly applied to this test setup over the course of roughly 2.5-3 hours spread across 4 days. This corresponds to 25 runs per day, and each run is defined as 5000 data points (-90s) collected at 56 Hz without any interruptions. This sampling rate is determined by the maximum frame rate possible with the camera module at the desired resolution (discussed in more detail in the Methods section). The first day’s 25 runs are used for calibrating the sensor with a linear model. The motive of collecting the test data on days different than the calibration day was to simulate a real-world scenario, where the sensor would be calibrated once but used for a long time after without requiring re-calibration. Collecting the test data across different days also serves a similar purpose of assessing any changes to sensor performance due to any environmental factors or adverse effects applying repeated forces / torques to the sensor on previous days.

[0065] The sensor’s time-domain response for one of the runs from the test set is shown in FIG.5 and the error characterized on the entire data set are shown in FIG. 6 and in Table 1 below. TheAttorney Docket No : 047162-5395-00WOmean absolute error is roughly below 1.5 N for forces and 0.1 Nm for moments compared to the ground truth ATI sensor throughout the test set, which corresponds to less than 1.5% relative error for the ±50-70 N force range and 2-4 Nm moment range (Table 2). And the prototype VisualFT sensortracks the ground truth even with fast, dynamic changes in applied forces / torques as seen in the sample run in FIG. 5. While the characterized error in forces and moments is reported here independently along each axis, the relationship and interactions between different axes is handled by the calibration. The difference in the error metrics along each direction is due to the difference in the magnitude of marker motion perceived by the single camera for different applied forces / torques. Lastly, the performance of the sensor did not exhibit any significant deviation from expected behavior over long durations of use on a single test day, or even across multiple test days after calibration (FIG. 6), highlighting the robustness of the sensor architecture to any changes in environmental conditions or noise.Table 1

[0066] When large errors do arise, they are limited at specific moments in the sensor response, rather than just general noise spread across the entire run. For instance, at the 25-30 seconds and 40-45 seconds marks in FIG. 5, the estimated Fx and Fy values notably stray from ground truth. These happen to be points where the concurrently applied Z-moment is quite high in magnitude. This is also similar to the high error around 700 and 1700 seconds in the Day 2 set. In these cases, the VisualFT sensor might have struggled to decouple each component of the applied force / torque, resulting in high deviation between the estimated and ground truth force / torque values. This could also result from not seeing this specific coupled force / torque combination in the calibration data, discussed in more detail in the subsequent sections.Attorney Docket No : 047162-5395-00WO

[0067] FIG. 5 and FIG. 6 depict the characterization responses of the prototype VisualFT sensor. FIG. 5 shows the sensor response (indicated as a solid line) and the ground truth values obtained from a reference ATI sensor (indicated as a dashed line) along with the mean absolute error for each direction for one of the test runs. Forces are indicated by Fx, Fy, and Fz, and moments by Mx, My, and Mz in the reference frame of the VisualFT sensor. FIG. 6 depicts the absolute error (averaged along the 3 dimensions) in forces and moments for concatenated runs on calibration (day 1) and test (day 2-4) data. The mean noted in the top right and indicated by the dashed line is calculated over all the runs on that day.

[0068] Table 1 above shows the characterization errors in the time-domain test response of the VisualFT sensor. The metrics are reported for the test data from all 3 days after calibration. The magnitude of error is low with differences stemming from the magnitudes of marker motion seen by the camera for different force / torque directions. Table 2 below includes example specifications of a sensor device. The performance and physical attributes of the sensor device are listed alongside a reference ATI sensor (SI-65-5 calibration). The sensor device, VisualFT, (left) and ATI Gamma (right) sensors are also shown side-by-side in FIG. 7. More detailed specifications of the VisualFT sensor device compared against various ATI sensors are shown in Table 3, Table 4, Table 5, and Table 6 below. Table 7 below lists the specifications and performance of another VisualFT device that utilizes a USB camera, thinner mirrors and a top cover (FIG. 1).Attorney Docket No : 047162-5395-00WOTable 2Table 3Attorney Docket No : 047162-5395-00WOTable 4Table 5Table 6Attorney Docket No : 047162-5395-00WOTable 7Closed Loop Control Tasks

[0069] An FT sensor is typically used for automation or manufacturing applications in a feedback loop to maintain or regulate threshold applied forces and moments, typically at the end effector of a robot manipulator. After calibrating the VisualFT sensor from the prior tests, the sensor’s estimated force / torque values are used to carry out two robot manipulation tasks: first, palpating a human dummy model akin to a medical robot scenario, and second, cleaning various kitchen items with a sponge. Both of these tasks require the robot arm to maintain contact at desired force / torque targets without any knowledge of the objects’ surface profiles. To achieve this behavior, a simple impedance controller is run on the robot end-effector position (Xt) that minimizes the error between the sensor estimate (Fest ) and the desired force / torque value (Fd ) with a proportional gain (Kp). This robot arm controller runs at 500 Hz although the VisualFT sensor estimates are only being updated at the maximum possible 56 Hz. And to smooth out any excessive oscillations, especially on more rigid contacts, a low-pass filter (parameterized by a) is run over the new end-effector positions as follows.XX t ~ (X Kp (Fest,t—Fetes) + (1—(Z) XXt-1Attorney Docket No : 047162-5395-00WOXt=Xt-i+ AXtEquation 1

[0070] Equation 1 (see O. Al-Mai, et al., IEEE Sensors J., 2018). For the palpating and cleaning tasks (FIG. 8A-8B), a robot arm (this can be a UR5 robot arm) is teleoperated only in translation directions (no teleoperated change in end-effector orientation) with a 3DConnexion SpaceMouse. However, once the palpating or cleaning task begins (initiated with a key press), only the XY component of the SpaceMouse commands are added to the AXt term in the controller Equation 1 above. This allows the robot arm to be moved over the different parts of the dummy model and kitchen items while the controller maintains contact and the force / torque target using the VisualFT sensor estimates. While the controller could be fine- tuned to achieve better impedance performance, the goal of these two tasks is to show the usability of the VisualFT sensor estimates in the loop and validate that they are well-aligned with the ground truth ATI sensor values (FIG. 9A-9B).

[0071] For the palpating task, the robot end-effector approaches the human dummy model from above with different orientations of the end effector at the start of each run, so that the various combination of forces and moments can be evaluated. The desired force / torque target value is 3 N vertical (Z) force measured in the VisualFT sensor frame, enough to only maintain contact but not create excessive deformations. In the experiment, the human dummy surface is rigid, but it is loosely placed on top of a soft, pliable surface, so any high forces / moments should be visually evident. Moreover, a curved, offset attachment at the tip of the FT sensor is also used in order to more effectively transmit Z-moments (a straight peg would not induce much Z-axis moment about the central axis of the sensor). Demonstrations of the robot performing this palpating task can be found in FIG. 8 A. For the cleaning task, several kitchen items with a variety of profiles and geometries from the Yale-CMU-Berkeley (YCB) Object and Model Set (see B. Calli, et al., Int. I. Robot. Res., 2017) are picked. A sponge is loosely placed (without any adhesives or fixtures) between the end-effector’s tip and the kitchen item surface to be cleaned. Any discontinuities in contact or force target would result in the sponge slipping away. For this task, the desired force target of 10 N vertical (Z) force on the surface measured at the tip of the peg attached to the sensor is required. That is, sensor force / torque estimates are converted toAttorney Docket No : 047162-5395-00WOequivalent force / torque value at a fixed Z distance away (equal to the length of the peg) through a rigid body transformation. The task is conducted for 9 different kitchen item surfaces as shown in FIG. 8B.

[0072] FIG. 8A-8B depict robot arm tasks with the VisualFT sensor estimates in closed-loop feedback. FIG. 8A shows that the robot arm end-effector (oriented at different angles in each trial) is required to maintain a small reference contact force to palpate the human dummy surface. FIG. 8B shows the kitchen items from the YCB object set being cleaned with a sponge (not fixed to the end-effector). The object surfaces are unknown, and the control loop adjusts the pose of the end-effector to maintain the desired normal force on the sponge required for scrubbing the objects. For both of these tasks, only the XY teleoperation commands from the SpaceMouse are added to the arm end-effector pose updates in the impedance control loop.

[0073] FIG. 9A-9B depict a closed-loop response of the VisualFT sensor for dummy palpating and sponge cleaning tasks. Sensor estimates (solid-line) used for the impedance control are overlayed on the ground truth reference values (dashed line). FIG. 9A shows the dummy is palpated in this trial with the end-effector tilted by 30° about X axis and a desired target of 3 N normal Z force (evaluated in the VisualFT frame). FIG. 9B shows the bowl object is cleaned requiring the end-effector to continuously apply 10 N (evaluated in the sponge contact frame) in the normal direction to maintain scrubbing force.DISCUSSION

[0074] The results in this work showed that a VisualFT sensor based on a single, inexpensive camera module can achieve six-axis FT performance within 1.5% relative error of commercially available expensive sensors, while maintaining a similar magnitude of stiffness characteristics. The components of the sensor are easily accessible or use rapid fabrication methods, and its structure is simple enough to be assembled within 15 minutes. In the characterization and real-world task demonstrations above, the VisualFT sensor was able to keep up with fast, dynamic applied forces / torques, and its calibrated performance did not deteriorate during hours of continuous use or any changes in temperature or humidity over multiple days.Long Duration Stability and Linear DeformationsAttorney Docket No : 047162-5395-00WO

[0075] The VisualFT sensor performance remains stable over long durations of use and maintains calibration over several days. This can be attributed to the use of a camera as the sensing module, and the choice of an aluminum flexure plate for the deforming component. The camera is mechanically isolated from any moving components, and unlike delicate sensing elements (capacitors, strain gages, optoelectronics and so on) used in commercial FT sensors, it has negligible drift or hysteresis characteristics of its own, especially from any changing temperature or humidity conditions. The digitized camera signal is also not prone to electromagnetic interference, which affects analog transducers in commercial FT sensors.Moreover, because the deformation signal is mechanically amplified as seen by the camera, the VisualFT sensor does not need any sophisticated amplification or conditioning circuitry. Some low-cost FT sensors use plastic or other viscoelastic materials in their deformation structures (see N. Hendrich, et al., IEEE Access, 2020). However, the metal flexure plate of the VisualFT sensor has a linear elastic deformation behavior, which eliminates any noticeable drift and hysteresis in the sensor output from material creep or stress relaxation. The linear elastic mechanics of the flexure plate are further helped by the fact that no sensing or adhesive material is attached to it, such as cement-bonded strain gages, that might add any undesirable viscous behavior or loss of deformation information from incorrect or uncured bonding. In the case that excessive force / torques are applied to the VisualFT sensor, a part of the flexure plate might undergo permanent plastic deformation and subsequently, the sensor will exhibit some nonlinear behavior and hysteresis effects. The device may add physical overload protection such as hard stops as well as checks in the sensor code to detect wrong initial marker positions and alert the user of any overloading-induced plastic deformation in the flexure plate.Scalability of Sensor Architecture

[0076] Six-axis force / torque sensors are used not just in high-force tasks like manufacturing but also in more delicate applications like medical robots. The VisualFT sensor design can be easily adapted to different ranges and resolutions of forces / torques by simply changing the thickness of the metal flexure plate, which in turn changes the sensor’s stiffness and the magnitude of marker deformations observed by the camera. So, for example, if larger forces / torques are expected in a machining task, the VisualFT sensor for that application would have a thicker flexure plate, notably without any major change to the overall dimensions of the sensor. The higher stiffnessAttorney Docket No : 047162-5395-00WOflexure plate does tradeoff with lower resolution or sensitivity, just like a commercial FT sensor. While commercial FT sensors use expensive monolithic machined parts and complex assembly processes (like strain gage bonding), the simple assembly and choice of components for the VisualFT sensor make it highly scalable to different applications.Limitations and Approximations

[0077] Since fiducial markers were used to track deformations of the flexure mechanisms, the computational overhead of detecting and measuring marker pose was very minimal and not the limiting factor to sampling rate even for the microcomputer used for this work (Raspberry Pi 5). However, cameras almost certainly have a lower sampling frequency than more analog sensing transducers such as resistive strain gages, photoresistors, or capacitors. Some off-the-shelf camera modules can record images at higher frame rates, including the module used (up to 120 frames / sec, although only with a cropped image view). Image sensors in camera modules can often record at even higher speeds, although the limitations lie in the data bandwidth of popular serial interfaces, distortion effects from rolling (versus global) shutter at high frame rates, and thermal limits of the hardware requiring active cooling. Some robot applications might not be dynamic enough to require more than a 100 Hz impedance control loop. But the maximum frame rate of the camera module would always limit the sample frequency of a visual sensor.

[0078] The VisualFT sensor performance was characterized over four days of use, with the first day data being used for calibration. The error characteristics of the sensor performance (FIG. 6) indicate some points where the estimated force / torque error is large. These could indicate gaps in the training data coverage of the calibration dataset. The linear model used to calibrate the marker motions to the ground truth force / torque values may not be flexible enough to estimate combinations of input forces / torques that were not seen during the calibration. While a linear calibration model was chosen as a tradeoff between accuracy and interpretability, a higher order model or more sophisticated methods such as a multilayer perceptron could be better suited to predict force response from images (see O. Al-Mai, et al., IEEE Sensors J., 2022)(see S. Urban, et al., IEEE / RSJ Int. Conf. Intell. Robots Syst. (IROS), 2013). The comparison of some of these other models and the small but notable improvements in the sensor’s performance is shown in Table 8, Table 9, Table 10, and Table 11 below. Table 8 uses a regression model withAttorney Docket No : 047162-5395-00WOinteraction effects - Intercept, linear, and paired products of predictors. Table 9 uses a regression model with quadratic and interaction terms - intercept, linear, squared, and pair products of predictors. Table 10 uses a Gaussian process regression model - a subset of regressors approximation and squared exponential kernel function. Table 11 uses a multilayer perceptron model - ReLU activation and three fully connected layers (size 100 - 50 - 25). There are likely other methods of calibration that can be investigated to further improve the accuracy of the sensor.Table 8Table 9Table 10Attorney Docket No : 047162-5395-00WOTable 11

[0079] For the ground truth force / torque values, the ATI Gamma sensor is used in this work. During these tests, forces / torques were manually applied to the experiment setup in FIG. 10. The manual calibration process may have resulted in inconsistent force application during calibration and test runs, particularly from unintended dynamic effects. Moreover, commercial sensors also have their own error profiles and drift characteristics, and result in deviations from ground truth post-calibration in the VisualFT response. More systematic calibration methods may be used, including custom calibration jigs that load the sensor with known weights and fixed directions to isolate out any unknown errors from the ground truth FT sensor (see D. Chen, et al., Procedia Eng., 2015).

[0080] The VisualFT sensor uses a single camera module to be able to capture high-dimensional data from a single, cheap sensor. Using a single sensor, however, can lead to differences in error magnitudes and sensitivity along different force / torque directions. This is because the amplification effects and effective marker motions perceived in the camera are higher for some directions (such as Mx or My) than others (such as Mz). Subsequently, the stiffness of the flexure plate also needs to be reduced in the low-sensitivity directions (Kz is lOx lower than Kxy). Furthermore, the sensitivity of the sensor is limited by the maximum image resolution of the camera module, and a monocular RGB camera has poorer depth perception. Commercial FT sensors use multiple sensing elements attached along different directions for a high degree of redundancy, which allows for uniform sensitivity and stiffness characteristics in all axes.Multiple cameras to capture marker motions from various angles, or stereo / depth cameras that can provide richer 3D deformation data may also be used. The flexure structure design can also be optimized for more uniform deformation and amplification profile along the six axes.METHODSAttorney Docket No : 047162-5395-00WOMechanical Design and Fabrication

[0081] The VisualFT sensor was designed with the goal of maintaining simple structure, easy-to-obtain components, and accessible fabrication and assembly. As such, most parts for the sensors are acquired off-the-shelf and any custom parts use quick prototyping fabrication methods such as 3D printing or waterjet cutting. The overall dimensions of the prototype sensor were designed to be similar to the reference ATI Gamma sensor. The outer enclosure is a 3.5 in (88.9 mm) 6061 Aluminum 346 round tube stock with 12 drilled holes to mount flexure and base plates and a slot for cables to run through. This is the only machined part in the sensor and can be swapped with a 3D printed version (also available along with the open-sourced design fdes of the sensor) for lower-force applications. The height of the overall sensor is determined by the minimum focus distance of the camera. By using mirrors and using the reflected image of the markers, the distance can be cut by roughly half (since the reflected image is perceived to be behind the mirror plane). The flexure and base plates are wateijet-cut from 6061 Aluminum sheet stock (0.08 in and 0.125 in thick, respectively), and all the remaining components such as holders for the mirrors, camera module, and markers are 3D printed, since they are not in the direct pathway of applied forces / torques. A Raspberry Pi 5 is used to detect markers on images from the camera module and compute forces, although the images could directly be transmitted from the camera module to the client personal computer (PC). FIG. 11 depicts an assembled rendering of the prototype sensor with the overall dimensions noted.

[0082] The acrylic mirrors are also readily available off-the-shelf and have an adhesive backing to stick to the mirror support / holder part that orients each mirror at a slight angle. The angled mirrors further extend the viewing field for the camera so that the markers can be placed as far away as possible from the base of the deflecting beam and maximize the lever amplification. For these marker positions in the prototype sensor and camera position, a mirror angle of 6 degrees was empirically found to keep the markers in the camera view throughout their range of deflection motion under applied forces / torques. The camera module used for the VisualFT sensor needed to have the following performance criteria - sufficiently high image resolution to pick up small marker motions, fast sampling rate at these high resolutions, wide viewing angle to capture multiple distant marker clusters, and short focus distance to keep the overall sensor height short. Based on these specifications, the Raspberry Pi Camera Module 3 Wide was picked for itsAttorney Docket No : 047162-5395-00WOubiquitous, long-term availability and peripheral support compared to most alternatives. To maintain consistent lighting for the camera to detect the fiducial markers, two white 5 V LEDs were attached to the inside walls of the enclosure. The flexure plate on top does let some external light through but can be covered up without disturbing marker detection if the LEDs are powered on.Flexure Plate Architecture

[0083] The VisualFT sensor relies on the compliant mechanism of the flexure plate to transfer the elastic deformations induced by the applied forces / moments to motion at the markers that can be tracked by the camera. The flexure plate was waterjet cut from a stock 0.08 in (2 mm) thick 6061 Aluminum sheet to stay within yield stress for forces up to ~50 N and moments up to ~5 Nm expected from robot manipulation tasks. Although the thickness of the flexure plate could be scaled up (or down) if a larger (or smaller) range of forces / moments is expected by trading off resolution. The design of the flexure plate compliant mechanism incorporates three deflection beams that are cantilevered on the outer edge (where the flexure plate meets the enclosure) and attached to a marker holder on the other end. The three cross-beam architecture is similar to most commercial strain gage FT sensor (such as the reference ATI Gamma used in this work). Two and four-beam versions were evaluated in FEA and it was empirically found that two beams did not capture enough deformation along certain directions (especially in line with the two marker clusters) and a four-beam version was too stiff and required a much thinner flexure members (for the expected 50 N / 5 Nm) that could be hard to consistently manufacture. A thorough design space search comparing different compliant structures (such as cross-beam, parallel, beam type, and so on), and the dimensions of flexure members may be conducted to further optimize performance of the camera-based FT sensor, although the suitability of various compliant structures for six-axis FT sensors is still debated in literature (see M. Y. Cao, et al., IEEE Sensors J., 2021)(see J. O. Templeman, et al., Sens. Actuators A Phys., 2020).Fiducial Marker Placement

[0084] ArUco fiducial markers were used to track the deflection from the beams under external loads and estimate the forces / torques applied. The 5-marker cluster arrangement is borrowedAttorney Docket No : 047162-5395-00WOfrom marker placement strategies that have been shown to improve pose estimation accuracy for ArUco tags (see P. Oscadal, et al., Sensors, 2020)(see G. Cepon, et al., Exp. Tech., 2024).Additionally, fiducial markers may also suffer from higher orientation error compared to translation error, particularly from pose ambiguity (see G. Schweighofer, et al., IEEE Trans. Pattern Anal. Mach. Intell., 2006). So, the cartesian position of the five markers was used in each cluster to obtain information about the 6-DOF pose of the whole cluster. Each cluster of markers is rigidly attached to the marker holder, so there is no relative movement between the markers. Although three markers would be sufficient for this purpose, five are used for redundancy and robustness to error and noise. With three beams and five markers on each beam, 15 markers are detected by the camera, and their cartesian positions (three values per marker) are recorded. So, the calibration method described later in this section regresses over the 45 predictors directly. The markers remain consistently illuminated inside the sensor enclosure because of the LEDs, and fewer than 1 in every 1000 captures failed to detect 1 of the 15 markers. To reduce this drop rate even further, thresholds for marker detection could be lowered, although that would make the detected marker pose values noisier. A library of models for different subsets of detected markers may be used, and since there is a redundant set of 15 markers, the predicted force / torque from fewer than 15 should still be useful.Stiffness and Sensitivity

[0085] The deflection behavior of the flexure plate defines the stiffness and sensitivity of the sensor, because that’s where the deformation under applied load is concentrated in the sensor structure. In this work, the flexure mechanism was designed with the goal of not exceeding yield stresses for forces up to 50 N and moments up to 5 Nm. The actual specifications of range, sensitivity, and stiffness of the final VisualFT sensor prototype were determined as shown in Table 2.

[0086] A FEA model of the flexure plate with attached marker holders was used in the calculations to record the stresses and deformations at various points within the sensor under different loading conditions. As shown in FIG. 4B, the flexure plate is assumed to be rigidly fixtured at the 6 holes where it is fastened to the external enclosure tube. The force and moment loads are applied at the holes in the center where the knob attaches to the flexure plate (in theAttorney Docket No : 047162-5395-00WOsensing frame shown in FIG. 10). The input sensing range is then calculated as the largest force or moment (in either positive or negative directions) that can be applied before the von Mises stress any point on the flexure plate exceeds the yield strength of 6061-T6 Aluminum. The stiffness of the mechanism is determined from the same FEA model by evaluating the deformation of the flexure plate where the knob is attached over the range of applied loads. The forces and moments increase linearly from no load to maximum over 5 seconds, although any inertial or damping effects are ignored in this static analysis. Consequently, the stiffness of the sensor then is the slope of a linear fitted curve on the resultant displacement of the plate under forces in meters (Kx, Ky, and !< / ) and the resultant displacement angle under moments in radians (for Ktx, Kty, and Ktz), respectively.

[0087] The sensitivity of this sensor would be determined by the smallest value of force or moment that would induce a discernable change in the pose of the ArUco markers in the camera view. To calculate this smallest discernable marker movement (in mm), the smallest change in pixel space that the camera would detect needs to be determined. A conservative estimate of this discernable resolution is about 1 / 4 pixel (see R. Ouyang, et al., 2020 IEEE Int. Conf. Robot. Autom. (ICRA), 2020), since ArUco marker detection algorithms like cornerSubPix can find exact corner positions at an accuracy higher than integer pixels.

[0088] For the sensor prototype, the distance between the camera and mirror, and that between the mirror and the marker plane is 15 mm. And the resolution of the camera image is 1296x1296 pixels wide (image is cropped square to the vertical resolution), and field of view is 67° for this module. In the marker space, this image then has the side dimension as calculated below.67°Height of image plane (mm) = 2 x (15 + 15) x tan— = 39.7 mmEquation 2

[0089] Using the image resolution, the height of each pixel can be calculated and subsequently, the smallest marker movement discernable by the camera.Attorney Docket No : 047162-5395-00WOHeight of each pixel (mm) = = 0.031 mmEquation 31 TSmallest discernable marker motion (mm) = - * 0.031 = 7.66* 10 mmEquation 4

[0090] Equations 2 and 3 (see Y. J. Li, et al., Meeh. Syst. Signal Process, 2018). The sensitivity of the sensor is then the minimum force or moment required in each direction to induce this amount of discernable motion at the markers. The force / moment values in each of the cartesian directions is found from the FEA model of the flexure plate and marker holders, and the calculated sensitivity values are shown in Table 2. Note that the mirror amplification effect is not considered here, and the mirrors are assumed to be stationary and flat for a more conservative estimate of the resolution. Also, while only the XY discernable marker motion is calculated above, the smallest motion discernable in Z is of similar magnitude. This is found by dividing the Z distance from mirror to the image plane (15 mm) by the perceived XY width of the marker cluster in the reflection. This perceived width is equal to the width of the marker cluster (3 markers across each 4 mm wide = 12 mm) with an additional factor from the angled mirror (equal to sine of 2* mirror angle = 1.2*) of the perceived XY motion in the reflection for any actual Z motion at the markers.Experiment Setup and Calibration

[0091] The prototype VisualFT sensor is mounted to a reference ATI Gamma through rigid fixture plates and on top of an optical breadboard as shown in FIG. 10. This setup is used for collecting characterization (calibration and test) data to evaluate its response. The fixture plates between the reference and prototype sensor are thick enough to eliminate any relative deformation between the two and the entire setup is fixed to an optical bread board to prevent dynamic effects arising from force / moment-induced vibrations.

[0092] The camera’s sampling rate and resolutions are set to the maximum available 56 Hz and 2304* 1296 (higher sampling rates are possible with a cropped image size and resolution, but notAttorney Docket No : 047162-5395-00WOall markers are visible in the cropped view). The fiducial marker positions from the camera images are detected onboard the Raspberry Pi 5 connected to the camera module. The recorded runtime for marker detection and recording images was noted to be well under 0.01 s per image, and as such, the prototype sensor was able to consistently maintain the 56 Hz sampling frequency. For characterization, the marker positions are time-stamped along with the raw values from the reference ATI sensor, and both these values are synchronously recorded to a client (PC) on the same network. For evaluating the sensor performance in characterization and closed loop tasks, the estimated force and torque values are calculated by the marker positions and calibration matrix onboard the Raspberry Pi.

[0093] Once the 15 markers are detected by the camera, a calibration model can be built to predict forces / torques from the marker positions. In this work, a linear regression is applied with an intercept to predict the 6 force / torque terms (Fx / y / zMx / y / z) from 45 predictors (Rt,xiy / z for the ithmarker). The calibration matrix (A) is thus a 6-by-46 matrix.Equation 5

[0094] Equation 5 (see Y. J. Li, et al., Meeh. Syst. Signal Process, 2018). A linear regression model (with Cauchy weighted robust fitting) was chosen for its interpretability and generalizability. Particularly with the large number of predictors compared to estimates, a more complex model might be more susceptible to overfitting. The ground truth and marker data were collected on 4 different days by running the sensors for 25 runs on a day, with each run having 5000 data points (~90 seconds for a 56 Hz sampling frequency), while different forces and torques are randomly applied at the knob attached to the VisualFT sensor prototype. During thisAttorney Docket No : 047162-5395-00WOprocess, the angle and direction of applied load was varied for obtaining a good coverage for different combinations of forces / torques in the calibration and test data. The linear model is fitted to the marker and ground truth data from the first day of runs. This calibration process may certainly be shortened and optimized in by collecting data more strategically along specific directions. Once the calibration matrix was found from the linear regression, the sensor performance was characterized on the 75 runs as test data collected on the three days (25 runs each) different from the calibration day to evaluate any change in the sensor response over time.

[0095] For the impedance controller used in the palpating and cleaning tasks, the proportional gain values used were KP= [0.002 for force, 1 for moments], and the low-pass filter cutoff factor was set as a = 0.01. While the controller could certainly be fine-tuned further for the desired impedance characteristics, the goal of the task demonstrations was to validate the utility of VisualFT sensor estimates in the loop and assess the alignment with the reference ground truth sensor values. FIG. 10. depicts the experiment setup used for calibrating of the VisualFT sensor protype and characterizing its time-domain response. The prototype sensor and reference ATI sensor are stacked with fixture plates on an optical breadboard. The sensor’s coordinate frame is also shown in FIG. 10.Addition Experimental DiscussionCalibration Sample Size and Models

[0096] In the prior VisualFT sensor’s response characterization experiment, the camera and ground truth data from the first day (1.25E5 samples) was used to calibrate the linear model and validate its performance on the remaining 3 days of test data. One possibly notable factor affecting the sensor and model response is the sample size of the calibration data. Here, the model’s performance (on unseen test data from day 4) is assessed by picking different numbers of calibration runs from days 1 to 3. The mean absolute error (averaged for the forces and moments) on the test data is plotted in FIG. 12 against the corresponding sample size of the calibration data on the horizontal axis. Since the calibration data is collected manually by randomly applying forces / torques on the test setup (FIG. 10), the model might have a different response based on the choice of the calibration runs. So, the calibrated model performance was tested 5 times for every calibration sample size data point, each time randomly selecting theAttorney Docket No : 047162-5395-00WOcalibration sample runs. FIG. 12 depicts the mean absolute error averaged for forces and moments on test data as a function of number of calibration samples. In FIG. 12, the solid line represents the mean over the 5 random selections of calibration runs at each sample size, and the shaded region is the 95% confidence interval. The error response is averaged for forces and moments over the 3 cartesian directions. The sensor FT error seems to stabilize around 1.3-1.5E5 calibration samples, which is roughly equivalent to the calibration sample size used in the characterization tests previously, although the variance is high for smaller sample sizes, indicating that the choice of what runs are picked for calibration (their coverage of different coupled force / moment values) can impact sensor performance. The sensor and linear model could be calibrated with even fewer samples if the collection process was automated to continuously monitor this training data coverage, and sample different combinations of forces / moments more efficiently.

[0097] The VisualFT sensor response was fitted with a linear calibration model that mapped the marker data to the six-dimensional force / torque vector. A simple linear model was chosen for its easy interpretability, generalizability with small amount of calibration samples, and relatively good prediction accuracy on the test data, especially using robust regression methods for less sensitivity to outliers. However, the prediction accuracy might be improved by using higher-order terms in the model or building models with Gaussian Process (GP) regression or neural network architectures. The performance of some of these calibration models is compared in Table 8, Table 9, Table 10, and Table 11. The models are fitted on 2.5*105samples from the first 2 days of sensor and ground truth values collected previously and tested on the remaining 2 days of data. Some models, such as the GP regression, take longer to fit considering the large number of regressors, and others like the multilayer perceptron model could perform even better with more calibration data and hyperparameter optimization. The prediction accuracy is improved (relative errors under 1%) compared to a purely linear model (under 1.3% in Table 2), but the tradeoffs of low interpretability, longer calibration time (sample collection and fitting time) might be acceptable for certain applications.Flexure Plate Stress ProfileAttorney Docket No : 047162-5395-00WO

[0098] One of the main reasons that the VisualFT sensor estimates can be easily predicted with a linear model is because the stress from applied forces / torques in concentrated in the planar flexure plate is made of a single, planar sheet and thus has a linear elastic deformation behavior. This allows the sensor to not suffer from any creep or stress relaxation that would show up as hysteresis and drift in the sensor’s FT estimates. To validate if the induced stress from applied forces and moments is in fact restricted to the metal flexure plate, the stress results are analyzed from the FEA model of the flexure structure of the sensor under the maximum loading conditions within the sensing range of the VisualFT sensor. The von Misses stress distributions for some of these loading conditions are shown in FIG. 13. Any stress in the sensing structure is primarily limited to the flexure plate beams. This stress is highest at the base of each of the 3 cantilevered beams as expected, and some stress is experienced by the branches connecting the 3 beams in the center because of the relative motion between the cantilevered beams. The stress in each cantilevered beam reduces to zero further away from the base and branch, and most notably, no stress is observed by the plastic marker holders that are attached to the flexure plate. This remaining portion of the beams and the marker holders simply serve to extend the elastic deformations and amplify the deflection motion where markers are placed.

[0099] This input sensing range of the sensor was previously noted in Table 2 as the largest force / moment that could be applied before any point on the flexure plate exceeded the yield stress (2.75E8 N / m2) of the plate material (6061 -T6 Aluminum). And as seen in FIG. 13, the von Misses stress for these maximum loading conditions is under the yield stress criterion. The factor of safety for the VisualFT prototype ranges between 1.2-2.5, and physical hard stops could be implemented to prevent accidental overloading of the flexure plate that might induce permanent plastic deformation at some points. However, some applications might necessitate reinforcing the flexure plate at stress points to achieve higher factors of safety or reduce the range of applied forces / moments.

[0100] Six-axis FT sensors may also be used in applications that induce dynamic periodic loads such as grinding or polishing. In these tasks, the resonant frequency of the sensing structure may be important to characterize the loading frequencies that should be avoided. Commercial FT sensors have a solid, monolithic structure because they rely on the sensing transducers picking up small deformations in the material. And as a result, these monolithic structures typically haveAttorney Docket No : 047162-5395-00WOnatural frequencies over 1000 Hz. The VisualFT sensor on the other hand relies on cantilevered beams for mechanical amplification. And a frequency analysis in FEA of the sensing structure showed that the first few resonant modes of the VisualFT sensing structure were at 131 Hz in Fz direction and 148-230 Hz in Tx / Ty directions. These natural frequencies are still high enough for most robotic applications, especially considering robot arms have their resonant frequencies in the tens of Hz (see S. A. Kouritem, et al., Alexandria Eng. J., 2022). But if even higher resonant frequencies are desired in the VisualFT sensor, then the stiffness of the planar spring structure could be increased and paired with a higher resolution camera module so that the image is more sensitive to deflection motions of the markers.

[0101] FIG. 13 depicts the stress distribution under maximum force / moment in the VisualFT sensing range for example loading directions. The von Misses stress results obtained from an FEA model shows that the stress is concentrated in the metal flexure plate, with the highest stress values at the base of each of the three beams and some in the beams extending to the middle where the forces / torques are applied. The FT range of this VisualFT sensor prototype is determined by noting the maximum loads before any point on the flexure plate exceeds the Yield stress of 6061-T6 Aluminum.

[0102] Table 8, Table 9, Table 10, and Table 11 depict the characterization errors in the VisualFT sensor response with other calibration models, such as, a regression model with interaction effects (Table 8), a regression model with quadratic and interaction terms(Table 9), a gaussian process regression model (Table 10) and a multilayer perceptron model (Table 11).

[0103] The disclosures of each and every patent, patent application, and publication cited herein are hereby incorporated herein by reference in their entirety. While this invention has been disclosed with reference to specific embodiments, it is apparent that other embodiments and variations of this invention may be devised by others skilled in the art without departing from the true spirit and scope of the invention. The appended claims are intended to be construed to include all such embodiments and equivalent variations.

Claims

Attorney Docket No : 047162-5395-00WOCLAIMSWhat is claimed is:

1. A sensor device comprising:a mounting plate;a flexure plate positioned parallel to the mounting plate having an inner surface facing the mounting plate and an outer surface opposite the inner surface;a wall connected to the flexure plate and the mounting plate, forming an enclosure between the flexure plate and the mounting plate;at least one detector disposed on the mounting plate within the enclosure;an end-effector connected to the outer surface of the flexure plate and configured to transfer applied forces to the flexure plate;at least one mirror attached to the inner surface of the flexure plate opposite the at least one detector; andat least one emitter arm located within the enclosure and attached to the inner surface of the flexure plate, positioned to be visible by the at least one detector via the at least one mirror.

2. The device of claim 1, wherein the at least one detector is a camera.

3. The device of claim 1, wherein the at least one emitter arm comprises at least one marker.

4. The device of claim 3, wherein the at least one marker comprises a patterned image.

5. The device of claim 4, wherein the at least one marker comprises a fiducial marker tag.

6. The device of claim 1, wherein the flexure plate comprises a plurality of openings configured to define a plurality of beams within the flexure plate.

7. The device of claim 1 further comprises a top cover positioned over and secured on the outer surface of the flexure plate.Attorney Docket No : 047162-5395-00WO8. The device of claim 1 , wherein the flexure plate is configured to deform under applied forces and torques.

9. The device of claim 8, wherein the flexure plate is configured to allow for lever mechanical amplification.

10. The device of claim 9, wherein the at least one emitter arm is configured to magnify a deformation of the flexure plate via lever mechanical amplification.

11. The device of claim 1, wherein the at least one mirror is secured in place using a mirror support.

12. The device of claim 1, wherein the at least one detector is isolated from the motions or forces on the flexure plate.

13. The device of claim 1 further comprising at least one Light Emitting Diode (LED) positioned within the enclosure.

14. The device of claim 1, wherein the enclosure is attached to the flexure plate and mounting plate using an attachment technique selected from a group comprising adhesive bonding, screwing, soldering, welding, riveting, or clamping.

15. The device of claim 1, wherein the wall of the enclosure comprises an opening and a cable grommet positioned within the opening.