Visual and touch integrated manipulator based on sparse visual angle three-dimensional reconstruction and hardness nondestructive testing method

By using a sparse-view 3D reconstruction vision-touch integrated robotic arm, combined with deep learning and force-displacement models, the problems of lighting and environmental interference in fruit non-destructive testing have been solved, achieving high-precision non-destructive testing of fruit hardness, and making it suitable for efficient testing in complex environments.

CN122008301APending Publication Date: 2026-05-12JIANGSU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU UNIV
Filing Date
2026-03-11
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing non-destructive testing methods for fruits mostly rely on machine vision, which is easily affected by lighting conditions, fruit surface stains, and specific varieties, resulting in low detection accuracy. Traditional 3D reconstruction requires expensive multi-camera arrays and is difficult to synchronize, resulting in high costs and difficulty in meeting the needs of industrial production lines. PDMS film deformation is small and nonlinear, and the sparse visual information is limited, making it difficult for traditional geometric reconstruction algorithms to recover high-precision 3D point clouds.

Method used

A vision-touch integrated robotic arm based on sparse viewpoint 3D reconstruction is used to apply pressure to a transparent PDMS film through mechanical grippers. Multi-view images are captured by a camera and a square reflector tube. The 3D shape is reconstructed using a deep learning network, and the hardness is quantified by a force-displacement model to achieve non-destructive testing.

Benefits of technology

It enables accurate, reliable, and non-destructive testing of fruit firmness, overcoming the limitations of traditional testing methods. It is suitable for non-destructive testing of high-value fruits, improves the robustness and accuracy of testing, and is applicable to efficient testing in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122008301A_ABST
    Figure CN122008301A_ABST
Patent Text Reader

Abstract

The invention discloses a visual touch integrated manipulator based on sparse visual angle three-dimensional reconstruction and a hardness nondestructive testing method. Hardware of the manipulator mainly comprises a mechanical clamping jaw, a transparent PDMS film, a square reflector barrel, a polaroid, an annular polarized light source, a camera and a mechanical arm interface. Collected images are transmitted to an external computing device and input to a deep learning network based on a Transform architecture, and a high-precision three-dimensional point cloud model of a deformation area is directly reconstructed. An external computing device accurately calculates the maximum deformation depth by processing the point cloud, and converts the maximum deformation depth into a quantized hardness index in combination with a pre-calibrated force-displacement relation model, so as to classify and discriminate the maturity level of the object. According to the method, the limitation that traditional detection depends on apparent characteristics is broken through, high-precision and high-reliability nondestructive detection is realized by directly quantifying fundamental physical characteristics, and the method has a wide application prospect in the fields of automatic perception and quality control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of agricultural intelligence, specifically involving a vision-touch integrated robotic arm and analysis method based on sparse perspective 3D reconstruction. Background Technology

[0002] The maturity of fruits and vegetables is a key quality indicator determining their harvesting, storage, transportation, and market value. Existing testing methods fall into two main categories: destructive testing and non-destructive maturity analysis. While destructive testing can accurately measure core indicators such as internal components, it suffers from high sample loss, high cost, and long processing time, making it impossible to test all fruits and failing to meet the full inspection needs of the supply chain. In non-destructive testing, manual assessment is highly subjective and inefficient; while instruments based on sound, light, and electricity also have limitations. For example, purely visual methods are easily affected by lighting conditions, fruit surface condition, and varietal differences, assessing only relevant rather than fundamental indicators of maturity. Acoustic and electrical analysis equipment is expensive and has limited applicability. Therefore, the market urgently needs a new technology that can directly, accurately, and non-destructively quantify core physical characteristics such as fruit firmness and enable large-scale testing. Summary of the Invention

[0003] To address the aforementioned problems, the purpose of this invention is to overcome the shortcomings of existing technologies and provide a vision-touch integrated robotic arm based on sparse perspective three-dimensional reconstruction and a non-destructive testing method for fruit hardness. This invention breaks through the limitations of traditional testing methods that rely on indirect related indicators such as color and odor. By directly measuring the three-dimensional deformation of the fruit under controlled pressure, its elastic modulus is quantified, thereby more fundamentally and accurately determining its hardness and ripeness.

[0004] This invention utilizes an external computing device to control a mechanical gripper, performing three-dimensional reconstruction of the deformation of a transparent PDMS film when it comes into contact with an external test object under controlled pressure. Force feedback is obtained by combining this with a pre-calibrated force-displacement model, thereby establishing a correlation model between the elastic modulus and hardness grade of the external test object. This method uses a camera with a square reflector tube to fold the optical path, simultaneously capturing multi-view images of the contact area in a single shot. The external computing device then calls a deep learning network model to reconstruct its three-dimensional morphology and accurately calculates the maximum deformation depth. Based on this maximum deformation depth value and the mechanical property model of the transparent PDMS film, a quantitative hardness index of the external test object is calculated, and this index is used to classify and distinguish the object's hardness grade. This achieves accurate and reliable non-destructive testing of the fundamental physical properties of the external test object, completely overcoming the limitations of traditional testing methods that rely on surface features.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: a vision-touch integrated robotic hand based on sparse perspective three-dimensional reconstruction, the robotic hand structure includes: mechanical gripper (1), transparent PDMS film (2), square reflector tube (3), polarizer (4), ring polarizing light source (5), camera (6), robotic arm interface (7).

[0006] The mechanical gripper (1) is configured to perform a controlled pressure action, causing an external test object to press the transparent PDMS film (2).

[0007] The transparent PDMS film (2) is fixed to the surface of the square reflector tube (3) and serves as a sensing interface that comes into contact with and is subjected to pressure from the external object being tested.

[0008] The square reflector tube (3) is configured as a folded optical path, so that the camera (6) can simultaneously capture the direct center view and multiple virtual indirect view images of the deformed area of ​​the transparent PDMS film (2) in a single shot;

[0009] The annular polarized light source (5) is arranged coaxially with the camera (6), and the polarizer (4) is placed in front of the lens of the camera (6) to jointly suppress the specular reflection on the surface of the transparent PDMS film (2).

[0010] The camera (6) is located at one end of the square reflector tube (3) away from the transparent PDMS film (2), and the light path is directed toward the transparent PDMS film (2).

[0011] The camera (6) is communicatively connected to an external computing device to transmit the acquired multi-view images to the external computing device for three-dimensional reconstruction and quantitative evaluation.

[0012] The robotic arm interface (7) is used to connect to the external support structure.

[0013] The present invention discloses a non-destructive testing method for hardness of a vision-touch integrated robotic arm based on sparse viewpoint 3D reconstruction. The logical computing module of the external computing device includes a system calibration module, a pattern judgment module, a 3D reconstruction module, a point cloud processing module, a hardness assessment module, and a visual classification module. The specific execution includes the following steps:

[0014] System calibration: The system calibration module applies a series of known displacements to the transparent PDMS film (2) through an external precision mechanical testing device and records the force it is subjected to simultaneously. The force-displacement function relationship model of the transparent PDMS film (2) is established through data fitting.

[0015] Detection mode judgment: The mode judgment module obtains the initial image of the external test object captured by the camera (6) and judges whether the external test object is suitable for pure vision detection mode; if it is suitable, the vision detection branch is executed; if it is not suitable, the tactile detection branch is executed.

[0016] Visual inspection branch: When using a pure visual inspection method, the external computing device is configured to run a visual classification module based on the YOLOv13 framework. This module adopts an adaptive sparse neural architecture and can dynamically adjust the network computation volume according to the computing power of the deployed hardware. This module directly quantifies and classifies the external test object into the corresponding hardness level in the pre-established standard hardness database based on its appearance characteristics.

[0017] The tactile detection branch includes the following sub-steps:

[0018] Image acquisition: The external mechanical gripper (1) brings the object under test into contact with the transparent PDMS film (2) and applies a preset pressure under controlled conditions to cause the transparent PDMS film (2) to undergo elastic deformation. Then, the camera (6) is used in conjunction with the square reflector tube (3) to take a multi-view image containing the deformation area of ​​the transparent PDMS film (2) from the direct center view and multiple virtual indirect view.

[0019] Three-dimensional reconstruction: The three-dimensional reconstruction module inputs the acquired multi-view images into a pre-trained end-to-end deep learning network model based on the Transformer architecture. The network model can directly regress a dense three-dimensional coordinate point map from the input multi-view images, thereby generating a three-dimensional point cloud characterizing the deformed surface of the transparent PDMS film (2).

[0020] Target localization: The point cloud processing module is configured to register the output three-dimensional point cloud with coordinate system by random sampling consensus algorithm RANSAC, and to segment the three-dimensional point cloud by combining color space analysis and density clustering algorithm DBSCAN, so as to extract the effective deformation area of ​​the transparent PDMS film (2).

[0021] Hardness assessment and classification: The hardness assessment module calculates the maximum deformation depth value from the processed effective deformation area point cloud, and converts the depth value into the corresponding force value according to the force-displacement relationship model of the mechanical properties of the transparent PDMS film (2) calibrated in step one. The force value is defined as the hardness quantification index of the fruit under the current deformation. Finally, the hardness quantification index is compared with the pre-established standard hardness database to realize the classification and discrimination of the hardness level of the external test object.

[0022] Compared with the prior art, the present invention has the following significant advantages:

[0023] The significant advantages of this invention are: 1. This invention breaks through the limitations of traditional visual inspection and realizes accurate hardness measurement based on physical essence; 2. The sparse-view imaging system composed of a camera and a square reflector tube designed in this invention combines low cost, compact structure and high efficiency; 3. This invention, based on the Transformer-based end-to-end 3D reconstruction algorithm, greatly improves the perception accuracy of small deformations; 4. This invention adopts adaptive feature extraction and multimodal fusion to standardize the detection results and improve the intelligent decision-making level of agricultural equipment.

[0024] This invention assesses maturity based on the fundamental changes in internal tissue structure, rather than solely on superficial features, significantly improving the robustness and reliability of the detection. It eliminates interference from ambient light and unstructured features, making it particularly suitable for detecting high-value fruits with minimal appearance changes.

[0025] This invention greatly shortens the data acquisition time, can acquire multi-view information without mechanical movement and multi-camera synchronous control, and the reflector tube design makes the sensor end lightweight and easy to integrate into the end of the robotic arm.

[0026] The algorithm of this invention has a powerful global context capture capability, which can accurately reconstruct the tiny three-dimensional deformation of the strain film under limited viewing angle, ensuring sub-millimeter accuracy in the calculation of maximum deformation depth, thereby ensuring the accuracy of hardness conversion.

[0027] This invention simulates the human "look first, then touch" sensory perception, endowing harvesting robots with more advanced decision-making capabilities. It can not only determine ripeness but also achieve high-precision positioning even under complex obstructions, achieving an adaptive balance between computational overhead and detection accuracy, thus greatly promoting the intelligent upgrading of agricultural equipment.

[0028] The technical problems solved by this invention are as follows: 1. Existing non-destructive testing of fruits largely relies on machine vision (color, texture), which is easily affected by lighting conditions, surface stains, and specific varieties (such as kiwifruit, whose color change is not obvious when ripe), resulting in low detection accuracy and an inability to reflect the true ripeness of the fruit's interior. 2. Traditional 3D reconstruction usually requires expensive multi-camera arrays, which suffer from synchronization difficulties and high costs. Single-camera multi-angle shooting is time-consuming and has a complex mechanical structure, making it difficult to meet the needs of industrial production lines or real-time robotic operations. 3. The deformation of PDMS film when in contact with fruit is often small and non-linear, and the visual information provided by sparse viewpoints is limited. Traditional geometric reconstruction algorithms are unable to recover high-precision dense 3D point clouds from a limited number of viewpoints. 4. Agricultural scenarios are complex, with fruits exhibiting occlusion, overlap, and subtle differences in ripeness. A single modality cannot simultaneously achieve both detection speed and accuracy.

[0029] Note that the description of these effects does not preclude the existence of other effects. One aspect of the invention does not necessarily have all the aforementioned effects. Effects other than those described above can be readily observed and extracted from the description, drawings, claims, etc. Attached Figure Description

[0030] Figure 1 This is a detailed external structural diagram of a visual-tactile robotic arm according to an embodiment of the present invention.

[0031] Figure 2 This is a detailed structural diagram of the internal structure of a visual-tactile robotic arm according to one embodiment of the present invention.

[0032] Figure 3 This is an overall flowchart of maturity analysis of a visual-tactile robotic hand according to an embodiment of the present invention.

[0033] Figure 4 This is a flowchart of a sparse fruit perspective 3D reconstruction according to an embodiment of the present invention.

[0034] Figure 5 This is a flowchart of fruit positioning according to one embodiment of the present invention.

[0035] Figure 6 This is a flowchart of a method for visually detecting maturity according to an embodiment of the present invention.

[0036] Figure 7 This is a flowchart of one embodiment of the present invention that detects maturity through touch. Detailed Implementation

[0037] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but the scope of protection of the present invention is not limited thereto.

[0038] Figure 1 , 2 The image shows a preferred embodiment of the vision-touch integrated robotic arm. The hardware structure of the device strictly includes a mechanical gripper (1), a transparent PDMS film (2), a square reflector tube (3), a polarizer (4), a ring polarizing light source (5), a camera (6), and a robotic arm interface (7).

[0039] The vision-touch integrated robotic arm is connected to and installed at the end of an external robotic arm via a robotic arm interface (7). The mechanical gripper (1) is located below the system and is used to perform specific gripping and controlled pressure actions under the control of an external computing device.

[0040] The transparent PDMS film (2) is fixed to one end surface of the square reflector tube (3) as a sensing interface that is in direct contact with the external object under pressure.

[0041] In this embodiment, the camera (6) is set at the top of the system and works with the square reflector tube (3) below to fold the light path, so that the direct center view and multiple virtual indirect view images of the contact area of ​​the transparent PDMS film (2) can be captured simultaneously in a single shot.

[0042] The annular polarized light source (5) is preferably arranged coaxially with the camera (6), and a polarizer (4) is installed in front of the lens of the camera (6). The two work together to uniformly illuminate the transparent PDMS film (2) and suppress specular highlights generated on its surface.

[0043] In addition, the system is equipped with an independent external computing device outside the hardware entity. It is connected to the camera (6), the external robotic arm and the mechanical gripper (1) respectively. Its core function is to receive the image data transmitted by the camera (6) and run the three-dimensional reconstruction algorithm based on deep learning, convert the acquired multi-view images into a high-precision three-dimensional point cloud model, and quantify the hardness index of the external test object based on the model.

[0044] Combination Figure 3 As shown, according to this embodiment, preferably, the overall process of the visual-tactile robotic hand performing maturity analysis includes:

[0045] First, the external computing device identifies and locates the image captured by the camera (6) to determine whether the external test object is suitable for pure visual detection or tactile detection. If it is visual detection, the external computing device directly classifies and distinguishes the external test object based on its appearance characteristics such as color under different states. If it is tactile detection, the determination is made by directly measuring the fundamental physical properties (hardness) of the external test object: the external robotic arm and mechanical gripper (1) apply controlled pressure to the external test object and press it onto the transparent PDMS film (2). The camera (6) and the square reflector tube (3) are used to simultaneously capture the deformation image of the transparent PDMS film (2). Then, the external computing device calls the internal Transformer-based deep learning model to process the image and directly reconstruct the three-dimensional point cloud of the deformation area. Finally, the external computing device calculates the maximum deformation depth through the point cloud processing algorithm and converts the depth value into a quantified hardness index according to the pre-calibrated force-displacement model, thereby accurately and reliably quantifying and distinguishing the characteristic level of the external test object.

[0046] According to this embodiment, preferably, the camera model has a resolution of 3840x3104 and captures 15 frames per second.

[0047] According to this embodiment, preferably, the vision-touch integrated robotic arm includes 1-mechanical gripper, 2-transparent PDMS film, 3-square reflector tube, 4-polarizer, 5-ring polarizing light source, 6-camera, and 7-robotic arm interface.

[0048] Before conducting actual testing, the mechanical properties of the transparent PDMS film (2) must first be systematically calibrated. An external standard mechanical testing device (such as a texture analyzer) is used to press the transparent PDMS film (2) at a set speed and depth, while simultaneously recording a series of continuous data points of pressure and displacement. Based on these calibration data, an external computing device uses the least squares method for nonlinear fitting to establish and verify an accurate force-displacement quadratic polynomial model. This model is permanently stored in the external computing device as a conversion benchmark for subsequent hardness calculations.

[0049] Reference Figure 1 When the system enters the tactile detection process, it first places the external test object on the transparent PDMS film (2), which serves as the sensing interface. Subsequently, under the control command of the external computing device, a preset controlled pressure is applied to the external test object through the mechanical gripper (1), causing it to press downwards and resulting in elastic deformation of the transparent PDMS film (2). When the external computing device determines that the maximum deformation state has been reached, it triggers the camera (6) to work with the square reflector tube (3) to capture a high-resolution keyframe image. This composite image fully contains five different perspectives of the deformation area of ​​the transparent PDMS film (2).

[0050] The present invention provides a non-destructive testing method for hardness of a vision-touch integrated robotic arm based on sparse viewpoint 3D reconstruction, comprising the following steps:

[0051] System calibration: A series of known displacements are applied to the elastic sensing film using precision mechanical testing equipment, and the forces acting on it are recorded simultaneously. A force-displacement function relationship model of the film is established through data fitting.

[0052] Image acquisition: The fruit to be tested comes into contact with the elastic sensing film and a preset pressure is applied to deform the film. Then, the cooperative camera (6) and the square reflector tube (3) are used to capture an image containing the direct view of the film deformation area and multiple virtual indirect view.

[0053] 3D Reconstruction: The 3D reconstruction module inputs the acquired multi-view images into a pre-trained end-to-end deep learning network model based on the Transformer architecture. The network model can directly regress a dense 3D coordinate point map from the input sparse view images, thereby generating a 3D point cloud characterizing the deformed surface of the thin film.

[0054] Target localization: The point cloud processing module is configured to register the output three-dimensional point cloud with coordinate system by random sampling consensus algorithm RANSAC, and to segment the three-dimensional point cloud by combining color space analysis and density clustering algorithm DBSCAN, so as to extract the effective deformation area of ​​the transparent PDMS film (2).

[0055] Perform hardness assessment: The hardness assessment module calculates the maximum deformation depth value from the processed effective deformation area point cloud, and converts the depth value into the corresponding force value according to the force-displacement relationship model of the mechanical properties of the transparent PDMS film (2) calibrated in step one. The force value is defined as the hardness quantification index of the fruit under the current deformation.

[0056] Visual inspection of maturity processing: When using a pure visual inspection method, the external computing device is configured to run a visual classification module based on the YOLOv13 framework. This module adopts an adaptive sparse neural architecture, which can dynamically adjust the network computation volume according to the computing power of the deployed hardware. Finally, the hardness index is compared with a pre-established standard hardness database of fruits of different maturity levels, thereby realizing the classification and discrimination of fruit maturity.

[0057] Combination Figure 4 As shown, according to this embodiment, preferably, the sparse viewpoint 3D reconstruction logic steps run by the external computing device are as follows: the 3D reconstruction module inputs the acquired multi-view images into a pre-trained end-to-end deep learning network model based on the Transformer architecture. The network model can directly regress a dense 3D coordinate point map from the input sparse viewpoint images, thereby generating a 3D point cloud characterizing the deformed surface of the thin film; the specific process includes:

[0058] 1. Image acquisition and viewpoint segmentation: The external computing device receives a panoramic composite image acquired by the camera (6) in conjunction with the square reflector tube (3) in a single shot as input; then, the external computing device preprocesses and segments the panoramic image to extract sparse viewpoint sub-images containing the direct central viewpoint and multiple virtual indirect viewpoints.

[0059] 2. Size normalization and feature extraction: The external computing device uses interpolation and other algorithms to normalize the size of the segmented sub-images (e.g., to a uniform resolution of 224*224 and to divide them into blocks). The processed image data is then fed into an encoder based on the Transformer architecture (e.g., the ViT encoder). Through a multi-layer network structure, the external computing device extracts deep visual features from each viewpoint layer by layer and gradually fuses global contextual information.

[0060] 3. Feature fusion and geometric relationship capture: The external computing device inputs the deep visual features output by the encoder into the cross-attention decoder; at this stage, the algorithm realizes cross-view feature depth interaction and fusion through a multi-head cross-attention mechanism, accurately capturing the geometric correspondence between different sparse viewpoints.

[0061] 4. Original point cloud regression and post-processing: Based on the fused features, the external computing device maps them to three-dimensional space through a regression network and directly regresses and outputs the original three-dimensional point cloud data characterizing the deformed surface of the transparent PDMS film (2); finally, the external computing device performs post-processing steps on the output original three-dimensional point cloud, including multi-view global alignment optimization and denoising, to solve the problem of inconsistent scale and pose of local point clouds, thereby outputting a high-precision, globally consistent three-dimensional deformed point cloud model.

[0062] Combination Figure 5 As shown, according to this embodiment, preferably, the target location logic steps executed by the external computing device include:

[0063] 1. Coordinate system registration and attitude correction: The external computing device performs coordinate system registration on the original point cloud data in arbitrary attitudes after 3D reconstruction. The feature plane representing the initial undeformed state of the transparent PDMS film (2) is robustly fitted by the Random Sample Consensus (RANSAC) algorithm (the program directly calls the algorithm module, which is a well-known technology), and its normal vector is calculated; then, through rigid body rotation and translation transformation, the feature plane is precisely aligned to the reference plane of the standard coordinate system to eliminate attitude uncertainty and lay the foundation for subsequent accurate geometric measurement.

[0064] 2. Coarse segmentation of the target region: The external computing device performs strain region segmentation based on spectral features on the registered point cloud. This process uses color space conversion and joint thresholding to perform preliminary segmentation, so as to quickly identify and separate the point cloud representing the actual contact area between the external test object and the transparent PDMS film (2), effectively filtering out most of the background interference and high light reflection noise.

[0065] 3. Refined Localization and Denoising: Based on the coarse segmentation, the external computing device employs an improved DBSCAN density clustering algorithm (the program directly calls this algorithm module, a well-known technique) to perform refined segmentation and localization of the target region. By identifying and retaining the core clusters with the largest cardinality, while completely eliminating all spatially discontinuous discrete noise points, a geometrically complete and spatially accurate point cloud of the contact subject target is obtained.

[0066] Combination Figure 6 As shown, according to this embodiment, preferably, when the system determines to perform a purely visual evaluation, the visual classification processing performed by an external computing device includes the following steps:

[0067] 1. Image Data Acquisition and Augmentation: Using the camera (6), multiple batches of images of external test object samples with different characteristic levels are acquired from multiple angles; and the acquired images are augmented (including random rotation, cropping, color jittering, and brightness adjustment) to expand the dataset size and improve the generalization ability of the target detection model. At the same time, the labeled data is refined and revised, especially the labels of samples in the transition period between adjacent maturity levels, to enhance the model's ability to distinguish subtle appearance differences.

[0068] 2. Adaptive Feature Extraction: The next-generation YOLOv13 object detection framework is adopted, the core of which is an adaptive sparse neural architecture (the program directly calls this algorithm module). This architecture can dynamically select and activate the optimal computation path and network depth during the forward propagation of the network based on the real-time computing power and power consumption constraints of the deployed hardware platform, thereby achieving adaptive optimization of computational load and energy consumption while maintaining extremely high detection accuracy.

[0069] 3. Context-Aware Feature Fusion: In the feature fusion stage of the network, a context-aware cross-scale fusion module is introduced. This module, through a multi-head self-attention mechanism, can more effectively aggregate high-level semantic information from feature maps of different levels with low-level high-resolution detail features, thereby significantly enhancing the model's ability to distinguish subtle differences in illumination changes, partial occlusion, and fruits of similar maturity.

[0070] 4. Target Detection and Maturity Determination: The deep fusion feature map is input into the detection head to accurately locate each external target in the image and output its bounding box coordinates. At the same time, for each successfully detected target, the classifier outputs its maturity level (e.g., immature, mature, overripe) and the corresponding confidence score. It achieves high-precision discrimination, especially in the detection of small targets and overlapping objects in complex backgrounds.

[0071] Combination Figure 7 As shown, according to this embodiment, preferably, the non-destructive testing method for hardness includes the following steps:

[0072] First, the external computing device preprocesses and segments the multi-view composite image acquired by the camera (6) and the square reflector tube (3) to obtain sub-images of the direct central view and virtual indirect view in four directions. Each sub-image is then fed into a pre-defined deep learning network logic for processing. Subsequently, the external computing device runs encoder logic based on the VisionTransformer (ViT) architecture to process the images. Specifically, the external computing device first uses an interpolation algorithm to unify the resolution of each sub-image to 224x224, while preserving the original width and height information of the image. Then, the image with the unified resolution is divided into 14x14 image blocks. For each segmented image block, the external computing device uses the encoder logic to extract the deep visual features of each block, i.e., it uses multiple Transformer blocks for continuous feature extraction calculations. During this process, each Transformer block contains a self-attention mechanism and a feedforward network to gradually fuse the global context information between image blocks from different viewpoints. Finally, the feature tensor and additional positional encoding output by the external computing device will be used as input data for the subsequent decoder logic module. The point cloud reconstruction network employs a dual-branch Transformer decoding architecture, introducing a cross-attention mechanism to achieve deep fusion of cross-frame information. Features extracted from two frames are denoted as F1 and F2, respectively, and fed into two symmetrical cross-attention branches. Each branch consists of multiple stacked decoding blocks. In each decoder layer, not only are its own features decoded, but features from the other frame are also introduced through cross-attention, achieving dynamic interaction and enhancement of cross-frame information. The features output from each decoder layer are then fed into the head module of the DPT structure for fusion processing to generate the final prediction result.

[0073] Each branch ultimately outputs a point cloud map and a confidence map, representing the location distribution of predicted points in the current frame and their corresponding confidence scores, respectively. The cross-attention decoder generates 3D point clouds and confidence maps from the encoded features of the two views. It fuses the features of the two views through a cross-attention mechanism, performing feature interaction to capture the geometric correspondence between viewpoints. Then, a lightweight head maps the decoded features to 3D space, outputting the 3D coordinates of each pixel to achieve point cloud regression, simultaneously generating a confidence map to filter low-quality predicted points. Finally, global optimization unifies multiple pairs of local point clouds to a global coordinate system, resolving inconsistencies in scale, rotation, and translation.

[0074] Although only dual-view input is accepted in a single forward propagation, global 3D scene reconstruction is achieved through the following multi-stage optimization strategy: First, a candidate matching map is constructed based on the texture similarity between images; then, a sliding window sampling strategy is used to sample all dual-view images with overlapping regions. Establish pairwise associations, among which Indicates the index of the two specific images that make up the image pair. This represents the set of all matched image pairs. To ensure the reliability of the matching pairs, the system employs a dual filtering mechanism: firstly, it utilizes a pre-trained dense feature encoder to extract high-dimensional visual features and performs geometric verification using a cosine similarity metric; secondly, it introduces an average confidence threshold in the initial matching stage. To exclude those with confidence levels below a threshold low-quality matching pairs For any pair of valid matching images in the scene. The network will output a set of point clouds in their respective local coordinate systems. and its corresponding confidence mapping ,in Indicates height as Width is The real matrix space. Since these local point clouds have independent coordinate systems and uncertain scales, we define a coordinate system for each matching pair. Introducing rigid transformation parameters and scale factor ,in Represents the rotation matrix in a rigid transformation. Represents the translation vector in a rigid transformation. This represents a special three-dimensional Euclidean group, which is the set of all rigid body rotations and translations in three-dimensional space. Let represent the set of positive real numbers. The goal of global alignment is to find a global point cloud that optimally fuses all local point clouds. The optimization problem is constructed as follows:

[0075] ;

[0076] In the formula It is the confidence level of the network prediction, used to weight the error of each point. Represents a robust kernel function. Represents the set of all matching image pairs. Indicates traversing image pairs The view index in the middle, This represents a pixel index in a point cloud or feature map. The Huber norm is used to resist outlier interference. This indicates the view in the global point cloud. No. Global coordinate estimates of each point. Represents a view The first in the local coordinate system One original observation point. To avoid scale degradation... Additional product normalization constraints This convex relaxation constraint effectively preserves the well-posedness of the optimization problem, while the alternating direction multiplier method (ADMM) enables efficient parameter solving. Finally, a globally consistent 3D model is obtained. And the precise pose and scale of all views. (Refer to...) Figure 5 A series of post-processing operations are performed on the raw point cloud output by the network.

[0077] An external computing device calculates the displacement offset of all points in the final processed, standardized target point cloud, extracting the maximum offset along the Z-axis. This maximum offset corresponds to the maximum pressing depth in the central region of the fruit. Subsequently, this maximum pressing depth is substituted into the force-displacement model established in the initial stage of the system to calculate the corresponding force value, which is defined as the quantitative representation of the fruit's hardness. Finally, the classifier precisely compares this quantified hardness value with a pre-established standard hardness database of fruits at different maturity levels (unripe, ripe, overripe), thereby ultimately determining and outputting the maturity level of the fruit being tested.

[0078] In summary, this invention deeply integrates an innovative visual three-dimensional reconstruction technology with the principle of biomimetic touch, simulating the real tactile perception process of organisms, and uses a transparent PDMS film (2) as the contact sensing interface. Through an external computing device, in conjunction with a camera (6) and a square reflector tube (3), when the external robotic arm and mechanical gripper (1) apply controlled pressure to the external test object, the camera (6) captures multi-view images, and the external computing device performs high-precision reconstruction of the fine three-dimensional deformation generated by the transparent PDMS film (2). Then, combined with a pre-calibrated force-displacement model, its elastic modulus is quantified, thereby achieving direct measurement of the most fundamental physical hardness of the external test object. This method achieves high-precision non-destructive testing of the hardness of the external test object, grasps the core essence of physical characteristic evaluation, and completely breaks through the technical bottleneck of traditional pure visual detection being easily interfered with by external factors such as color, lighting, and appearance differences. In the fields of automated equipment and flexible testing, this technology can endow intelligent equipment with precise "visual and tactile senses," enabling it to make more reliable operation and testing decisions in complex unstructured environments, and significantly improve the success rate of non-destructive testing. In industrial production and quality monitoring applications, by accurately quantifying and classifying the hardness of target objects, production plans can be effectively optimized, defective product losses can be greatly reduced, and high-quality products with highly consistent quality can be provided to end users.

[0079] It should be understood that although this specification describes various embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can be appropriately combined to form other implementation methods that can be understood by those skilled in the art. The series of detailed descriptions listed above are merely specific descriptions of feasible embodiments of the present invention, and they are not intended to limit the scope of protection of the present invention. All equivalent embodiments or modifications made without departing from the spirit of the present invention should be included within the scope of protection of the present invention.

Claims

1. A vision-touch integrated robotic hand based on sparse viewpoint 3D reconstruction, characterized in that, The robotic arm structure includes: mechanical gripper (1), transparent PDMS film (2), square reflector tube (3), polarizer (4), ring polarizing light source (5), camera (6), and robotic arm interface (7). The mechanical gripper (1) is configured to perform a controlled pressure action, causing an external test object to press the transparent PDMS film (2). The transparent PDMS film (2) is fixed to the surface of the square reflector tube (3) and serves as a sensing interface that comes into contact with and is subjected to pressure from the external object being tested. The square reflector tube (3) is configured as a folded optical path, so that the camera (6) can simultaneously capture the direct center view and multiple virtual indirect view images of the deformed area of ​​the transparent PDMS film (2) in a single shot; The annular polarized light source (5) is arranged coaxially with the camera (6), and the polarizer (4) is placed in front of the lens of the camera (6) to jointly suppress the specular reflection on the surface of the transparent PDMS film (2). The camera (6) is located at one end of the square reflector tube (3) away from the transparent PDMS film (2), and the light path is directed toward the transparent PDMS film (2). The camera (6) is communicatively connected to an external computing device to transmit the acquired multi-view images to the external computing device for three-dimensional reconstruction and quantitative evaluation. The robotic arm interface (7) is used to connect to the external support structure.

2. The vision-touch integrated robotic hand according to claim 1, characterized in that, The external computing device is configured to receive the multi-view images acquired by the camera (6) in conjunction with the square reflector tube (3), reconstruct the three-dimensional deformation of the transparent PDMS film (2), calculate the maximum deformation depth based on the three-dimensional deformation, and then combine the pre-calibrated force-displacement model to calculate the hardness index of the external test object, thereby quantitatively evaluating the hardness level of the external test object.

3. A non-destructive testing method for the hardness of a vision-touch integrated robotic arm based on sparse viewpoint 3D reconstruction according to any one of claims 1-2, characterized in that, The logic computing module running on the external computing device includes a system calibration module, a pattern judgment module, a 3D reconstruction module, a point cloud processing module, a hardness assessment module, and a visual classification module. Specific implementation Includes the following steps: System calibration: The system calibration module applies a series of known displacements to the transparent PDMS film (2) through an external precision mechanical testing device and records the force it is subjected to simultaneously. The force-displacement function relationship model of the transparent PDMS film (2) is established through data fitting. Detection mode judgment: The mode judgment module obtains the initial image of the external test object captured by the camera (6) and judges whether the external test object is suitable for pure vision detection mode; If applicable, execute the visual inspection branch; If not applicable, execute the haptic detection branch; Visual inspection branch: When using a pure visual inspection method, the external computing device is configured to run a visual classification module based on the YOLOv13 framework. This module adopts an adaptive sparse neural architecture and can dynamically adjust the network computation volume according to the computing power of the deployed hardware. This module directly quantifies and classifies the external test object into the corresponding hardness level in the pre-established standard hardness database based on its appearance characteristics. The tactile detection branch includes the following sub-steps: Image acquisition: The external mechanical gripper (1) brings the object under test into contact with the transparent PDMS film (2) and applies a preset pressure under controlled conditions to cause the transparent PDMS film (2) to undergo elastic deformation. Then, the camera (6) is used in conjunction with the square reflector tube (3) to take a multi-view image containing the deformation area of ​​the transparent PDMS film (2) from the direct center view and multiple virtual indirect view. Three-dimensional reconstruction: The three-dimensional reconstruction module inputs the acquired multi-view images into a pre-trained end-to-end deep learning network model based on the Transformer architecture. The network model can directly regress a dense three-dimensional coordinate point map from the input multi-view images, thereby generating a three-dimensional point cloud characterizing the deformed surface of the transparent PDMS film (2). Target localization: The point cloud processing module is configured to register the output three-dimensional point cloud with coordinate system by random sampling consensus algorithm RANSAC, and to segment the three-dimensional point cloud by combining color space analysis and density clustering algorithm DBSCAN, so as to extract the effective deformation area of ​​the transparent PDMS film (2). Hardness assessment and classification: The hardness assessment module calculates the maximum deformation depth value from the processed effective deformation area point cloud, and converts the depth value into the corresponding force value according to the force-displacement relationship model of the mechanical properties of the transparent PDMS film (2) calibrated in step one. The force value is defined as the hardness quantification index of the fruit under the current deformation. Finally, the hardness quantification index is compared with the pre-established standard hardness database to realize the classification and discrimination of the hardness level of the external test object.

4. The method according to claim 3, characterized in that, The visual classification module is configured to introduce a context-aware cross-scale fusion mechanism in the feature fusion stage, which aggregates high-level semantics and low-level detail features through an attention mechanism to enhance the ability to distinguish external test objects with changes in illumination, partial occlusion, and similar hardness.

5. The method according to claim 3, characterized in that, The three-dimensional reconstruction module is configured to receive a composite image containing multiple virtual perspective information acquired by the camera (6) in conjunction with the square reflector tube (3), and input it into a deep learning network to directly regress and output a three-dimensional point cloud model representing the deformed surface of the transparent PDMS film (2).

6. The method according to claim 5, characterized in that, The deep learning network adopts a Transformer-based encoder-decoder architecture, including a ViT encoder for extracting features from the multi-view images, and a dual-branch cross-attention decoder for fusing features from different viewpoints through a cross-attention mechanism to regress and output a 3D point cloud. The decoder is configured to fuse feature information between different viewpoints through a cross-attention mechanism to capture geometric correspondences.

7. The method according to claim 3, characterized in that, The visual classification module is configured to calculate the maximum deformation depth value from the processed point cloud, and convert the depth value into a quantified hardness index based on a pre-calibrated force-displacement relationship model that characterizes the mechanical properties of the transparent PDMS film (2). The hardness index is then compared with a pre-established standard hardness database to finally classify and determine the hardness level of the external test object.