A tactile perception data processing method, object cognition method and device

By using tactile coding methods and a generalized contour diffusion model, the problem of efficient utilization in robot tactile perception data processing was solved, achieving high-precision object recognition and generalized contour reconstruction, while reducing data acquisition costs and redundant computation.

CN120816502BActive Publication Date: 2025-11-18TUJIAN TECH (BEIJING) CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511326206.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-11-18
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

Existing robotic tactile perception technologies lack effective model building methods, resulting in high data acquisition costs and inefficient utilization, and an inability to effectively integrate visual and tactile information to achieve high-precision object recognition.

Method used

By using a tactile coding-based approach, a tactile code containing generalized contour information is generated by aligning the constraints of the generalized contour code and the initial tactile code. This is then combined with a generalized contour diffusion model for noise extraction and removal, enabling efficient processing of tactile data and object recognition.

Benefits of technology

This improves the richness and completeness of tactile perception data, providing high-quality input for subsequent high-precision generalized contour reconstruction, reducing redundant calculations in data processing, and improving data processing efficiency and the accuracy of object recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120816502B_ABST
    Figure CN120816502B_ABST
Patent Text Reader

Abstract

The application provides a kind of tactile perception data processing method, object cognitive method and equipment.Application is applied to robot perception technical field.The method comprises: according to the generalized contour data of target object, feature learning is carried out, and generalized contour coding is obtained;According to the tactile data of the target object, deep learning is carried out, and initial tactile coding is obtained;Based on the preset semantic space alignment relationship, the generalized contour coding and the initial tactile coding are aligned by forward propagation, and the tactile coding containing generalized contour information is obtained.The application makes the tactile perception data more rich and complete, and provides better input support for subsequent high-precision generalized contour reconstruction of target object.Meanwhile, through feature compression, redundant calculation is reduced, and data processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of robot perception, in particular to a tactile perception data processing method, an object cognition method and equipment. BACKGROUND

[0002] The object cognition capability of a robot is an important basis for realizing intelligent operation. At present, both vision and touch can provide certain object cognition capability. The vision approach can capture the geometric profile of an object under visible conditions, but cannot provide force-related information (such as inertia, stiffness, friction, etc.); the touch approach can provide force-related information, but currently lacks a corresponding model construction method. The collection of tactile data relies on direct interaction with an object, which can cause damage to the robot, the sensor and the target object, and the cost of data collection is relatively high, so efficient use of data is particularly important. SUMMARY

[0003] Therefore, the present application provides a tactile perception data processing method based on tactile coding, comprising: performing feature learning according to the generalized profile data of a target object to obtain a generalized profile code; performing deep learning according to the tactile data of the target object to obtain an initial tactile code; and performing constraint alignment of the generalized profile code and the initial tactile code based on a preset semantic space alignment relationship through forward propagation to obtain a tactile code containing generalized profile information.

[0004] Optionally, the preset semantic space alignment relationship is established in the following manner: cosine similarity is calculated according to the generalized profile code and the initial tactile code to obtain a cosine similarity matrix; a loss function is calculated according to the cosine similarity matrix to obtain a first cross-entropy loss function; the first cross-entropy loss function is optimized by backpropagation to constrain the alignment of the generalized profile code and the initial tactile code, and the alignment parameters are adjusted through gradient update to obtain the preset semantic space alignment relationship.

[0005] The tactile perception data processing method based on tactile coding provided by the present application further comprises: performing physical feature extraction on the generalized profile data and the tactile data respectively to obtain generalized profile physical features and tactile physical features; calculating inconsistency based on the generalized profile physical features and the tactile physical features to obtain inconsistency information; optimizing the semantic space alignment strategy of the generalized profile code and the initial tactile code based on the inconsistency information; and performing constraint alignment of the generalized profile code and the initial tactile code based on the optimized semantic space alignment strategy to obtain a tactile code containing generalized profile information.

[0006] Optionally, the physical feature extraction is performed according to the generalized contour data, and generalized contour physical features are obtained, including: performing rotation and translation transformation on the generalized contour data and the tactile data, and determining optimal matching points; and extracting physical features in the generalized contour data based on the optimal matching points, to obtain the generalized contour physical features.

[0007] Optionally, the physical feature extraction is performed according to the tactile data, and tactile physical features are obtained, including: dividing the regions with the same time period and the change in the loading state based on the tactile data, to obtain the touch regions; removing the touch regions that do not occur in the contact in the entire time period, to obtain the target touch regions; and performing the physical feature extraction on the target touch regions based on the optimal matching points, to obtain the tactile physical features.

[0008] Optionally, the semantic space alignment strategy of the generalized contour encoding and the initial tactile encoding is optimized based on the inconsistency information, including: calculating the cosine similarity based on the generalized contour encoding and the initial tactile encoding, to obtain a cosine similarity matrix; performing similarity fusion based on the cosine similarity matrix and the inconsistency information, to obtain a fused similarity; and optimizing the semantic space alignment strategy of the generalized contour encoding and the initial tactile encoding based on the fused similarity.

[0009] The second aspect of the present application provides a method for object cognition based on robot tactile, including: obtaining generalized contour data and tactile data of a target object; performing feature learning on the generalized contour data by using a generalized contour encoding model, to obtain generalized contour encoding; performing feature learning on the tactile data by using any one of the tactile perception data processing methods based on tactile encoding, to obtain tactile encoding containing generalized contour information; performing noise extraction on the generalized contour encoding added with random noise and the tactile encoding containing generalized contour information by using a generalized contour diffusion model, to obtain noise, and removing the noise from the generalized contour encoding added with random noise based on the noise, to obtain target generalized contour encoding; and decoding the target generalized contour encoding by using the generalized contour encoding model, to obtain the real generalized contour of the target object.

[0010] Optionally, the generalized contour data of the target object is obtained, including: collecting point clouds of the target object by using a visual collector; sampling based on the point clouds, to obtain second representative points; obtaining displacement information of the target object caused by applying force to the target object at the second representative points; obtaining the proportional relationship information between the force and the displacement based on the applied force and the corresponding displacement information; and obtaining the real generalized contour data of the target object based on the point clouds and the proportional relationship information between the force and the displacement.

[0011] Optionally, the haptic data of the target object is acquired, including: based on the preset position and the preset direction, controlling the robot end effector to move to the target position; controlling the robot end effector to grasp the target object with a preset load; adjusting the position and the load of the robot end effector, and collecting the haptic data of the target object in the adjustment process.

[0012] The third aspect of the present application provides a robot haptic-based object cognition device, comprising: a processor and a memory connected with the processor; wherein the memory stores instructions executable by the processor, and the instructions are executed by the processor to enable the processor to perform the robot haptic-based object cognition method described above.

[0013] The haptic perception data processing method based on haptic coding provided by the present application retains key visual geometric structures by learning general contour data features into general contour coding with fewer points but richer general contour features; retains contact mechanics information by deep learning haptic data into initial haptic coding with fewer points but richer haptic features; and finally fuses the haptic coding with contour information through feature alignment to generate haptic coding containing general contour information, so that the haptic perception data is more rich and complete, providing better input support for subsequent high-precision general contour reconstruction of the target object. At the same time, feature compression reduces redundant calculation and improves data processing efficiency.

[0014] The object cognition method based on robot haptics provided by the present application uses a general contour coding model to learn and compress contour data, retain core key features, learn new fusion features and encode to significantly reduce data volume by taking a small amount of general contour and haptic data of the target object as basic input; then a haptic perception data processing method based on haptic coding encodes the haptic data into haptic coding containing general contour information, not only retains contact mechanics information, but also integrates geometric structure information, and reduces the size of haptic information; then the general contour diffusion model learns the potential offset of the data by adding and removing noise, enhances the adaptability of the model to complex distribution. At the same time, the general contour diffusion model fuses the general contour coding and the haptic coding containing the general contour information, cooperatively maps the geometric structure and contact state information to the general contour space, and finally decodes the data through the general contour encoder to obtain the real general contour and reconstruct the high-precision general contour containing the complete shape of the object and retaining the contact details. Moreover, the complete general contour consistent with the real general contour is reconstructed from a small amount of original data, realizing efficient acquisition of a large number of general contours driven by a small amount of input data, providing a cost-effective general contour input scheme for robot haptic cognition, and improving the efficiency of complete general contour reconstruction of the target object. BRIEF DESCRIPTION OF DRAWINGS

[0015] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0016] Figure 1 This is a flowchart of the first tactile perception data processing method based on tactile coding in an embodiment of the present invention;

[0017] Figure 2 This is a structural diagram of the first tactile coding model in an embodiment of the present invention;

[0018] Figure 3 This is a structural diagram of the encoder and decoder in an embodiment of the present invention;

[0019] Figure 4 This is a flowchart of the second tactile perception data processing method based on tactile coding in an embodiment of the present invention;

[0020] Figure 5 This is a structural diagram of the second tactile coding model in an embodiment of the present invention;

[0021] Figure 6 This is a flowchart of an object recognition method based on robot touch according to an embodiment of the present invention. Detailed Implementation

[0022] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0024] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can also refer to the internal connection of two components; and they can refer to a wireless connection or a wired connection. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0025] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0026] like Figure 1 As shown, this embodiment of the invention provides a tactile perception data processing method based on tactile coding. This method is executed by an electronic device such as a computer or server, and specifically includes:

[0027] S11, perform feature learning based on the generalized contour data of the target object to obtain the generalized contour code.

[0028] In this embodiment, the target object is any object (such as a person or object) for which generalized contour data needs to be acquired. Point cloud data or force response data at point cloud coordinates of the target object are acquired using a force response testing device to obtain generalized contour data, which is used to characterize the geometric contour morphology, mechanical response distribution characteristics, and structural spatial properties of the target object. Feature learning is then performed on the generalized contour data of the target object to ultimately generate a generalized contour code with fewer points but richer generalized contour features.

[0029] S12, perform deep learning based on the tactile data of the target object to obtain the initial tactile code.

[0030] In this embodiment, the tactile data is collected by a robot end effector to assess the tactile characteristics of the target object. This tactile data includes parameters such as the position coordinates of the end effector during adjustment, the magnitude and direction of the applied force, the displacement change at the contact point, pressure distribution characteristics, and deformation response under contact conditions. These parameters characterize the target object's contact mechanical properties (e.g., stiffness, softness), the fine structure of its local geometric contours (e.g., curvature characteristics), and surface pressure distribution patterns. Feature learning is then performed on the target object's tactile data to generate an initial tactile code with fewer points but richer tactile features.

[0031] S13, based on the preset semantic space alignment relationship, the generalized contour code and the initial tactile code are constrained and aligned through forward propagation to obtain a tactile code containing generalized contour information.

[0032] Based on the pre-defined semantic space alignment relationship established during the training phase, the generalized contour code and the initial tactile code are aligned in the feature space through forward propagation. This allows the tactile code to explicitly incorporate key information such as the space and shape of the generalized contour, ultimately generating a tactile code that includes generalized contour information.

[0033] The tactile perception data processing method based on tactile coding provided in this invention is essentially a tactile perception data processing algorithm implemented by a tactile coding model. Besides using common unsupervised dimensionality reduction techniques such as autoencoders or variational autoencoders, the tactile coding model can also utilize... Figure 2 The model shown includes an encoding module 100 (a generalized contour encoder 101 and a tactile encoder 102) and a contrast matching module 103.

[0034] like Figure 3 As shown, the encoding module 100 specifically adopts an encoder-decoder architecture, including a downsampling module 1021 and multiple cross-attention downsampling modules 1022. Figure 2 The number of modules shown is for illustrative purposes only; the actual number can be designed according to specific needs. The upsampling module 1023 and multiple cross-attention downsampling modules 1024 are used to perform deep learning on the generalized contour data (format [B,C,N]) and tactile data (format [B,T,K,(R_1,R_2)+(S_1,S_2)]) of the target object, respectively, to obtain the generalized contour code and the initial tactile code (format [B,H]). The decoder 104 includes an upsampling module and multiple cross-attention upsampling modules, which are used to decode the generalized contour code and the initial tactile code respectively to obtain the generalized contour data of the target object.

[0035] The model is pre-trained using collected generalized contour training data and tactile training data. During training, the generalized contour training data and tactile training data are used as inputs, and tactile encoding is used as the output target. The encoding module 100 extracts key contour features from the generalized contour training data and tactile training data and compresses them into generalized contour encoding and initial tactile encoding, respectively. Then, the decoder of the encoding module 100 decodes the generalized contour encoding and initial tactile encoding to restore the generalized contour and tactile data, respectively. The trained encoding module 100 can achieve feature learning using only the encoder, that is, it can actually perform feature learning based on the obtained generalized contour data and tactile data from the generalized contour encoder 101 and tactile encoder 102. Then, the generalized contour encoding and initial tactile encoding are subjected to constrained alignment training through the comparison and matching module.

[0036] The process of establishing the pre-defined semantic space alignment relationship is as follows:

[0037] When training using the contrast matching module 103, the similarity matrix between the generalized contour code and the initial tactile code is first calculated (cosine similarity can be calculated), and the matrix elements... (B is the batch size) represents the similarity between sample i in the generalized contour encoding batch and sample j in the initial tactile encoding batch. Then, based on this similarity matrix, the cross-entropy loss function is calculated along the rows and columns respectively. The total loss is calculated based on the row cross-entropy loss and the column cross-entropy loss, i.e., the first cross-entropy loss function drives the model optimization. Finally, the total loss is optimized through backpropagation, the alignment parameters are adjusted, and the homologous encodings are forced to align in the feature space, while the non-homologous encodings are separated, thereby establishing the preset semantic space alignment relationship between the "generalized contour encoding and the initial tactile encoding".

[0038] For example, the total loss is calculated as follows:

[0039] ,

[0040] ,

[0041] ,

[0042] ,

[0043] ,

[0044] in, This represents the row probability of sample i and sample j. This represents the row cross-entropy loss between sample i and sample j. Represents the column probabilities of sample i and sample j. This represents the column cross-entropy loss between sample i and sample j. This represents the probability value of sample i corresponding to its own position (the i-th column) in the column-oriented probability distribution. This represents the total loss for sample i and sample j.

[0045] The above training process utilizes a design of "diagonal element matching and off-diagonal element mismatch." Essentially, it aligns contour codes and tactile codes of the same order in the feature space by forcing them to align in the same batch, while distancing code pairs of different orders. This achieves the constrained alignment of the "generalized contour code" and the "initial tactile code," ultimately integrating the tactile code with the structural information associated with the generalized contour.

[0046] The tactile encoding model is continuously trained using generalized contour data and tactile data until the total loss converges (or reaches a preset threshold / no longer decreases significantly). At this point, the model completes the matching and alignment of generalized contours and tactile features. The generalized contour encoding and the initial tactile encoding are aligned in the feature space. The tactile encoding ultimately integrates the structural information associated with the generalized contours.

[0047] This embodiment learns generalized contour data features into a generalized contour code with fewer points but richer generalized contour features, preserving key visual geometric structures; it also uses deep learning to transform tactile data into an initial tactile code with fewer points but richer tactile features, preserving contact mechanics information; finally, feature alignment fuses the tactile code with contour information to generate a tactile code containing generalized contour information, making the tactile perception data richer and more complete, providing better input support for subsequent high-precision generalized contour reconstruction of target objects. Simultaneously, feature compression reduces redundant computation and improves data processing efficiency.

[0048] In some optional embodiments of this example, step S13, based on a preset semantic space alignment relationship, constrains and aligns the generalized contour code and the initial tactile code through forward propagation to obtain a tactile code containing generalized contour information. Specifically, this includes:

[0049] Based on the predefined semantic space alignment relationship, forward feature fusion is performed on the generalized contour code and the initial tactile code to obtain a tactile code containing generalized contour information.

[0050] Using the contrast matching module 103 (with fixed parameters) of the trained tactile coding model, based on the preset semantic space alignment relationship established during the training phase (including fixed feature association parameters such as attention weights and feature mapping matrices), the features of the input generalized contour code and the initial tactile code are associated (e.g., narrowing the feature distance between pairs of samples in the same order and suppressing irrelevant associations between pairs of samples in different orders), so that the initial tactile code is fused with the generalized contour structure information, and finally a tactile code containing generalized contour information is generated.

[0051] This embodiment utilizes the trained model parameters to directly align the features of the generalized contour code and the initial tactile code, enabling the initial tactile code to quickly fuse key information such as the space and shape of the generalized contour, generating a tactile code rich in geometric and mechanical information, providing high-quality input support for subsequent high-precision generalized contour reconstruction.

[0052] like Figure 4 As shown in the figure, the tactile perception data processing method based on tactile coding provided by the embodiment of the present invention further includes:

[0053] S14, Physical features are extracted from the generalized contour data and tactile data respectively to obtain the generalized contour physical features and tactile physical features.

[0054] The generalized contour physical features include the equivalent radius of the contact area, the displacement and force of the first normal force, and the first spatial position and time period, etc. The tactile physical features include the concentration of the normal force along the tangential direction, the displacement and force of the second normal force, and the second spatial position and time period, etc.

[0055] S15, calculate the inconsistency based on the generalized contour physical features and tactile physical features to obtain inconsistency information.

[0056] Calculate the inconsistency between the equivalent radius of the contact area and the concentration of the normal force along the tangential direction, calculate the inconsistency between the displacement and force of the first normal force and the displacement and force of the second normal force, and calculate the inconsistency between the first spatial position and time period and the second spatial position and time period.

[0057] S16, a semantic space alignment strategy for optimizing generalized contour coding and initial tactile coding based on inconsistency information.

[0058] If one or more generalized contour physical features are inconsistent with tactile physical features, it indicates a deviation in feature alignment between the current generalized contour encoding and the initial tactile encoding. These inconsistencies are then synthesized, and the pre-defined semantic space alignment constraints are dynamically adjusted, such as with... Figure 1 The cosine similarity calculated in the middle is used for synthesis.

[0059] S17. Based on the optimized semantic space alignment strategy, the generalized contour code and the initial tactile code are constrained and aligned to obtain a tactile code containing generalized contour information.

[0060] In the comparison matching module, based on the optimized semantic space alignment strategy, the generalized contour code and the initial tactile code are re-aligned to obtain a tactile code containing generalized contour information.

[0061] This embodiment is compared to Figure 1 The method embodiment shown adds physical cognition. By extracting physical features such as the equivalent radius of the contact area and the displacement of the first normal force from the generalized contour data, and physical features such as the concentration of the tangential distribution of the normal force and the displacement of the second normal force from the tactile data, the feature inconsistency between the two is calculated and the semantic space alignment strategy is optimized. This enables the model to more accurately capture the underlying relationship between the geometric structure and the contact mechanics, thereby improving the completeness and accuracy of the subsequent generalized contour reconstruction.

[0062] Another tactile perception data processing method based on tactile coding provided in this embodiment of the invention is essentially also a tactile perception data processing algorithm implemented by a tactile coding model. For example... Figure 5 As shown, the model includes an encoding module 100 (a generalized contour encoder 101 and a tactile encoder 102), a contrast matching module 103, and a physical calculation module 104.

[0063] Among them, the encoding module 100 and the comparison and matching module 103 are... Figure 2 The corresponding modules in the same module have the same structure and function, so they will not be described in detail here.

[0064] The training process of the physical calculation module 104 is as follows: extract the equivalent radius of the contact area, the displacement and force of the first normal force, and the first spatial position and time segment of the input generalized contour training data respectively; extract the concentration of the normal force along the tangential direction, the displacement and force of the second normal force, and the second spatial position and time segment of the input tactile training data respectively.

[0065] Taking the extraction of features from tactile training data as an example, the rules are as follows: For tactile training data, it is considered to include several "contact areas", and it is required that the data is within a time period and the loading state changes (such as the loading state changes significantly between adjacent time points). First, the physical calculation module should calculate whether the end-operation device in the tactile data makes contact with the surface of the target object. If the data of the tactile area that does not make contact throughout the process is not included in the subsequent parameter calculation, the data is removed. Then, if the contact area makes contact, during the contact period, the average spatial position (such as coordinates) and orientation (such as normal vector direction) during that period are calculated. The relationship between the change of the amount of motion in the normal direction (such as displacement) and the sum of the normal force (such as pressure) in the contact area during that period is calculated (for example, "whether the displacement increases linearly as the force increases"). The distribution concentration of the normal force along the array direction of the tactile sensor (such as the row / column direction of the sensor arrangement) during that period is calculated (expressed as the ratio of the standard deviation to the mean on the array, where the array reading is non-negative). Then, the physical calculation module is trained by combining the data of each tactile area calculated above with the corresponding features of the generalized contour data. That is, the inconsistency between the features corresponding to the contact data and the generalized contour data is calculated, and the three inconsistencies are summed to convert into a comprehensive similarity (the smaller the inconsistency, the higher the similarity). The optimal rotation and translation transformation (adjusting the spatial position of the tactile data) is found by calculating the cross-entropy loss function, and the similarity is further optimized (making the tactile data and the generalized contour more spatially matched). The optimized similarity is combined with the cosine similarity of the generalized contour encoding and the initial tactile encoding to obtain the final matching result involving physical cognition, and then the training of the physical calculation module is completed.

[0066] This embodiment extracts the physical features of generalized contour and tactile data respectively, quantifies the feature differences between the two in the dimensions of geometric structure and contact mechanics, and dynamically optimizes the semantic space alignment strategy based on these differences to promote the deep alignment of generalized contour encoding and initial tactile encoding at the level of physical laws, and more accurately integrates geometric structure and contact mechanics information, providing feature support that is more in line with actual physical laws for subsequent complete and accurate generalized contour reconstruction.

[0067] In some optional embodiments of this example, step S14, which involves extracting physical features from the generalized contour data to obtain the generalized contour physical features, specifically includes:

[0068] S141a performs rotation and translation transformations on the generalized contour data and tactile data, and determines the optimal matching point.

[0069] The physical calculation module of the trained tactile coding model is used to perform rotation and translation transformations on the same batch of generalized contour data and tactile data, and to determine the optimal matching points between the two data, such as K optimal matching points.

[0070] S142a, based on the optimal matching point, extract physical features from the generalized contour data to obtain the generalized contour physical features.

[0071] Physical features are extracted at each optimal matching point, such as the equivalent radius of the contact area (format [B,K]), the displacement and force of the first normal force (format [B,K,2]), and the first spatial position and time segment (format [B,K,R_1,R_2], where R_1 and R_2 represent rotation and translation transformation matrices), and other generalized contour physical features.

[0072] This embodiment achieves consistent alignment between generalized contour data and tactile data in the spatial coordinate system by performing rotation and translation transformations on the generalized contour data and determining the optimal matching points. Based on these optimal matching points, the physical features of the generalized contour are extracted, ensuring accurate correspondence between the features and key positions of the geometric structure. This avoids feature errors caused by data position deviations and provides a spatially consistent feature benchmark for subsequent similarity calculation guided by physical cognition and optimization of semantic space alignment strategies, significantly improving the accuracy of the fusion of geometric structure and contact mechanics features.

[0073] In some optional embodiments of this example, step S14, which involves extracting physical features from tactile data to obtain tactile physical features, specifically includes:

[0074] S141b, based on tactile data, divides areas within the same time period and with changing loading states to obtain tactile zones.

[0075] Based on time, tactile data, such as changes in contact force, is divided into multiple local regions, each of which is called a tactile zone.

[0076] S142b removes the tactile areas that have not been touched throughout the entire time period to obtain the target tactile area.

[0077] Some tactile areas may only show a "change in loading state" but no actual effective contact occurred throughout the process (i.e., the contact force was not actually applied to the target object). These areas cannot provide effective contact mechanics information and therefore need to be eliminated. Only tactile areas where actual contact occurred should be filtered and retained to ensure that the objects for subsequent physical feature extraction are effective contact areas.

[0078] S143b extracts physical features of the target tactile area based on the optimal matching point to obtain tactile physical features.

[0079] Based on the K optimal matching points determined after spatial alignment of generalized contour data and tactile data, and using the optimal matching points as spatial references, feature extraction is performed on the tactile data within the target tactile area to obtain tactile physical features such as the concentration of normal force along the tangential direction (format [B,K,1+1], where +1 represents the weight coefficient of the target tactile area filtered in S142b), the displacement and force of the second normal force (format [B,K,2+1], where +1 represents the weight coefficient of the target tactile area filtered in S142b), and the second spatial position and time period (format [B,K,R_1,R_2, +1], where +1 represents the weight coefficient of the target tactile area filtered in S142b).

[0080] This embodiment first segments continuous tactile data into tactile areas that may contain contact events based on the time boundary of the loading state change, then removes invalid areas that were not actually touched throughout the process, and finally accurately locates the target tactile area that was actually touched. By filtering out meaningless data interference, it provides high-quality real contact data to support the subsequent extraction of tactile physical features, and provides accurate data for the accuracy of subsequent fusion of geometric structure and contact mechanical features.

[0081] In some optional implementations of this embodiment, step S16, the process of optimizing the semantic space alignment strategy of the generalized contour encoding and the initial tactile encoding based on inconsistency information, specifically includes:

[0082] S161, calculate the cosine similarity based on the generalized contour coding and the initial tactile coding to obtain the cosine similarity matrix.

[0083] The cosine similarity between the generalized contour code generated by the encoding module and the initial tactile code is calculated using the contrast matching module of the trained tactile encoding model, resulting in a cosine similarity matrix.

[0084] S162, similarity fusion is performed based on the cosine similarity matrix and inconsistency information to obtain the fused similarity.

[0085] The physical calculation module of the trained tactile coding model is used to fuse the inconsistency information calculated in S15 with the cosine similarity matrix.

[0086] S163, a semantic space alignment strategy based on fusion similarity optimization of generalized contour coding and initial tactile coding.

[0087] Aligning generalized contour encoders and initial tactile encoders based on fusion similarity constraints refers to adjusting the parameters of the generalized contour encoder and the initial tactile encoder using optimization algorithms (such as gradient descent) to maximize their fusion similarity. Fusion similarity integrates cosine similarity (basic feature matching degree) with inconsistency information (physical law deviation).

[0088] This embodiment promotes the deep alignment of generalized contour coding and initial tactile coding at the level of physical laws by integrating inconsistencies obtained from cosine similarity and physical cognition guidance. This enables the final tactile coding to retain as much information as possible related to the generalized contour, significantly improving the integrity and accuracy of contact data.

[0089] like Figure 6 As shown, this embodiment of the invention provides a method for object recognition based on robot tactile perception. This method is executed by electronic devices such as computers or servers, and specifically includes:

[0090] S21, Obtain the generalized contour data and tactile data of the target object.

[0091] The process involves collecting generalized contour data of the target object using force response testing equipment and tactile data of the target object using a robot end effector. The generalized contour data includes point cloud data of the target object or force response data at point cloud coordinates, used to characterize the geometric contour morphology, mechanical response distribution characteristics, and structural spatial properties of the target object. The tactile data includes parameters such as the position coordinates of the end effector during adjustment, the magnitude and direction of the applied force, the displacement change at the contact point, pressure distribution characteristics, and deformation response under contact conditions, used to characterize the contact mechanical properties of the target object (e.g., stiffness, flexibility), the fine structure of the local geometric contour (e.g., curvature characteristics), and surface pressure distribution patterns. Data acquisition can be completed with only a small amount of contact information from discrete points or local areas. This step eliminates the need for high-density, full-coverage acquisition methods, significantly reducing data acquisition costs and time.

[0092] S22, use the generalized contour coding model to learn features from the generalized contour data to obtain the generalized contour code.

[0093] Feature learning is achieved through a generalized contour encoder in a generalized contour coding model to obtain the generalized contour code of the target object. Besides common unsupervised dimensionality reduction techniques such as autoencoders or variational autoencoders, generalized contour coding models can also employ encoder-decoder architectures. The generalized contour encoder includes multiple downsampling modules, and the generalized contour decoder includes multiple upsampling modules. Specifically, the input can be generalized contour data of the target object, in the format [B, C, N], where B is the batch size of the input data, C is the number of generalized contour features (including features from the generalized contour data and contact data), and N is the number of points. Then, the downsampling module in the generalized contour encoder is used to transform the generalized contour data, resulting in a data format of [B,C+D_1,N_1]. Here, C represents the number of generalized contour features (inclusive), D_1 is the number of features learned by the model, and N_1 is the number of points after downsampling, which is less than N. Further downsampling modules are then added as needed to continue the transformation, resulting in data formats of [B,C+D_2,N_2], [B,C+D_3,N_3], ..., [B,C+D_m,N_m]. This process continuously increases the number of learned features while decreasing the number of points. Finally, multiple upsampling modules in the generalized contour decoder transform the data in the format [B,C+D_m,N_m] to the format [B,C+D_m-1,N_m-1], ..., until finally transforming it to the format [B, C, N], which is the generalized contour data of the target object.

[0094] S23, using the tactile sensing data processing method based on tactile coding described in any embodiment, feature learning is performed on the tactile data to obtain a tactile code containing generalized contour information.

[0095] By using tactile coding to learn and encode tactile data in the model, a tactile code containing generalized contour information is obtained, reducing the amount of data.

[0096] S24. The generalized contour diffusion model is used to extract noise from the generalized contour code with added random noise and the tactile code containing generalized contour information to obtain noise. Then, the noise is removed from the generalized contour code with added random noise based on the noise to obtain the target generalized contour code.

[0097] Random noise and extracted noise do not refer to physical noise, but rather to random offsets added to the generalized profile diffusion model. This process of adding and removing noise helps the generalized profile diffusion model learn the underlying structure and distribution of the data.

[0098] S25, use the generalized contour coding model to decode the generalized contour of the target object to obtain the true generalized contour of the target object.

[0099] By decoding the generalized contour encoding of the target object using a generalized contour encoding model, the true generalized contour of the target object can be obtained. Specifically, multiple upsampling modules of the generalized contour encoder gradually restore the "compressed information" in the low-dimensional generalized contour encoding to the specific coordinates of the high-dimensional point cloud, and finally reconstruct the generalized contour data that is highly consistent with the true contour of the target object in terms of geometric shape and mechanical response distribution.

[0100] This embodiment uses a small amount of generalized contour and tactile data of target objects as basic input. It then uses a generalized contour encoding model to learn and compress the contour data, retaining core key features and learning new fusion features, which significantly reduces the amount of data. Then, a tactile perception data processing method based on tactile encoding encodes the tactile data into a tactile code containing generalized contour information. This not only retains contact mechanics information but also incorporates geometric structure information and reduces the scale of tactile information. Finally, a generalized contour diffusion model is used to learn potential biases in the data by adding and removing noise, thereby enhancing the model's adaptability to complex distributions. Meanwhile, the generalized contour diffusion model integrates generalized contour coding and tactile coding containing generalized contour information, co-mapping geometric structure and contact state information to the generalized contour space. Finally, the data is decoded by the generalized contour encoder to obtain the real generalized contour, reconstructing a high-precision generalized contour that contains the complete shape of the object and retains contact details. Moreover, it reconstructs a complete generalized contour consistent with the real generalized contour from a small amount of original data, realizing the efficient acquisition of a large number of generalized contours driven by a small amount of input data. This provides a cost-effective generalized contour input scheme for robot tactile cognition and improves the efficiency of reconstructing the complete generalized contour of the target object.

[0101] In some optional embodiments of this example, before performing noise extraction on the generalized contour code with added random noise and the tactile code containing generalized contour information using the generalized contour diffusion model in step S24, the object cognition method based on robot tactile perception further includes:

[0102] Random noise and time steps with the same shape as the generalized contour coding are generated based on the generalized contour coding.

[0103] Based on the shape of the generalized contour encoding, random offsets (such as coordinate offsets and feature offsets) are generated to match it. This process does not depend on actual time, but only controls the offset generation logic through time steps (number of iterations).

[0104] Adding random noise to the generalized contour coding results in a generalized contour coding with added random noise.

[0105] The generated random offsets are progressively superimposed onto the generalized contour code over time steps. For example, a small offset is added at time step t=1, and two offsets are added at t=2 (or the offset strength is increased). Each time step represents a "noise diffusion" operation, which iteratively disrupts the structure of the original code, simulating the cumulative perturbation of data in a complex environment.

[0106] This embodiment adds offsets controlled by time steps, which not only ensures that the subsequent generalized contour diffusion model focuses on core contour information (rather than local perturbations), but also enhances the reconstruction capability of the original data through inverse denoising. Ultimately, the model can decode the target generalized contour more accurately from a small number of generalized contour codes.

[0107] In some optional embodiments of this example, in step S24, a generalized contour diffusion model is used to extract noise from the generalized contour code with added random noise and the tactile code containing generalized contour information to obtain noise, specifically including:

[0108] S241, the generalized contour code with added random noise and the tactile code containing generalized contour information are concatenated to obtain the fused features.

[0109] The generalized contour encoding with added random offsets is concatenated with the tactile encoding containing generalized contour information at the feature level to form a fused feature. This step fuses visual geometric information with tactile mechanical information, allowing the model to utilize global perturbation information of both "shape" and "contact state" simultaneously, providing a more comprehensive input for subsequent offset extraction.

[0110] S242 uses a generalized contour diffusion model based on time steps to iteratively extract noise from the fused features to obtain the noise.

[0111] Using time steps as the control unit (each time step represents one noise diffusion operation), the fused features are processed in multiple iterations. The generalized contour diffusion model progressively separates and extracts the offset components at each time step, ultimately obtaining a clean offset signal.

[0112] This embodiment uses time-step controlled offset iteration extraction. The generalized contour diffusion model simultaneously utilizes global information from both geometry and mechanics, and accurately separates the offset components in stages. The generalized contour diffusion model can more accurately recover the original unoffset generalized contour encoding from the offset generalized contour encoding, providing a more realistic input for subsequent generalized contour reconstruction, and significantly improving the robustness and accuracy of object shape perception in robot tactile cognition.

[0113] In some optional embodiments of this example, obtaining the generalized contour data of the target object in step S21 specifically includes:

[0114] S211a, uses a visual acquisition device to acquire the point cloud of the target object.

[0115] A series of spatial points of the target object are collected using a visual acquisition device (such as a 3D scanner), for example, 2048 points in the form of (x, y, z). The collected point cloud data can also be called contour data. During the acquisition, the target object can be placed in a certain way, such as freely placed on a plane or suspended by an elastic string.

[0116] S212a, based on point cloud sampling, obtains the second representative point.

[0117] A typical way to select the second representative point is to sample from the farthest point to reduce the amount of data.

[0118] S213a, Obtain displacement information of the target object caused by applying force at the second representative point.

[0119] At each second representative point, a force is applied along that point, and the displacement of the target object is measured. The force can be applied along the normal direction at the second representative point. The method for determining the local normal direction is to fit a plane to points "near" the second representative point, and the normal of those points is the normal of the second representative point.

[0120] S214a, based on the applied force and the corresponding displacement information, obtains the proportional relationship information between force and displacement.

[0121] For example, when suspending a target object with an elastic string, the force and displacement should be directly proportional. In this case, this proportionality coefficient is extracted as information about the ratio of force to displacement.

[0122] For example, when a target object is freely placed on a plane, there should be no displacement (or very small displacement) when the force is less than a certain value, and a linear relationship should exist when the force is greater than a certain value. In this case, the limit of the force and the proportional coefficient of the linear relationship are extracted as information on the proportional relationship between force and displacement.

[0123] This step can also involve installing a high-density array sensor at the end of the force-applying device to apply force and measure displacement and force distribution, thereby obtaining information on the ratio of force to displacement.

[0124] S215a, based on point cloud and force-displacement ratio information, obtains the generalized contour data of the target object.

[0125] This embodiment uses a vision-based data acquisition device to collect point cloud data of the target object as basic contour data, further reducing the data volume through sampling at the farthest point. Then, force is applied at a small number of secondary representative points, and displacement is measured. The direction of the force is determined by combining local normal fitting. Mechanical features are extracted by analyzing the relationship between force and displacement. Finally, the point cloud and the proportional relationship between force and displacement are used as generalized contour data. This process requires only a small amount of raw data acquisition to obtain high-information-density generalized contour data, thus reducing data acquisition costs.

[0126] In some optional embodiments of this example, obtaining tactile data of the target object in step S21 specifically includes:

[0127] S211b controls the robot's end effector to move to the target position based on a preset position and preset direction.

[0128] Using a robotic arm or other tactile device, the device moves to a target location near the target object. The preset position and direction can be determined by first finding a general graspable position and direction based on vision using existing grasping decision algorithms, and then randomly fine-tuning it around that position. Alternatively, it can be determined manually.

[0129] S212b controls the robot's end effector to grasp the target object with a preset load.

[0130] The preset load can be a small force to grip the target object to prevent damage to its shape.

[0131] S213b adjusts the position and load of the robot's end effector and collects tactile data of the target object during the adjustment process.

[0132] The robot continuously adjusts the gripping load or the position of the end effector, and collects contact status information of the robot end effector when gripping the target object in real time during the adjustment process, such as contact pressure distribution (force value at each contact point), contact force change rate, contact area ratio, etc.

[0133] This embodiment first controls the robot's end effector to move to a preset position, gently gripping the object with a small load to prevent deformation. Then, by dynamically adjusting the gripping load or the device's position, tactile data of the contact during the adjustment process is collected. This process requires only a few gripping adjustments to obtain key tactile information reflecting the object's contact characteristics, reducing data acquisition complexity and providing core inputs for describing the contact state in subsequent generalized contour coding and tactile coding models. This allows the models to more accurately learn the coupling relationship between geometry and mechanics, ultimately reconstructing extended contour information that includes both shape and contact details.

[0134] In some optional embodiments of this example, the load of the robot end effector is adjusted in step S213b, specifically in two ways:

[0135] The first method is to increase the preset load until it reaches the preset maximum load, and then stop increasing the preset load.

[0136] The second method involves collecting data on the target object's position during the increase of the preset load. It then determines whether the target object's position has changed. If the target object's position has not changed, the preset load is increased further. If the target object's position has changed, and the displacement reaches the preset displacement, the increase of the preset load is stopped.

[0137] Both of the above methods start with a preset load (small load) and gradually increase the load until a certain criterion is met, such as reaching the preset maximum load or the target object's position changing and reaching a certain displacement when the load is increased, at which point the increase in load can be stopped.

[0138] This embodiment achieves dual optimization of accurate tactile data acquisition and object protection through two load adjustment methods: starting with a small load and gradually increasing it avoids deformation or damage to the object caused by sudden force; and triggering a stop by monitoring position changes precisely captures the transition of contact state, ensuring that the acquired tactile data covers the contact characteristics under different loads. These two methods enhance the diversity of tactile data, including contact details under different loads, providing a more comprehensive input for subsequent model learning of the coupling relationship between geometry and mechanics.

[0139] In some optional embodiments of this example, adjusting the position of the robot's end effector in step S213b specifically includes: maintaining a preset load constant while adjusting the position and orientation of the robot's end effector. The adjustment of the position and orientation of the robot's end effector can be random or fixed.

[0140] This embodiment achieves diversified acquisition of tactile data and full coverage of contact scenarios by keeping the preset load of the grip constant and adjusting the position and orientation of the robot's end effector. Random adjustment can cover the contact state of the target object at different positions and angles, avoiding local data deviations caused by fixed positions; fixed adjustment can specifically acquire contact details at key locations (such as edges and recesses). The combination of these two methods makes the tactile data contain richer scene information, reflecting the mechanical properties of the object at different contact positions, and providing a more comprehensive input for subsequent model learning of the coupling relationship between geometry and mechanics.

[0141] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0142] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0143] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0144] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0145] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A tactile perception data processing method based on tactile coding, characterized in that, include: Feature learning is performed based on the generalized contour data of the target object to obtain the generalized contour code; Deep learning is performed on the tactile data of the target object to obtain an initial tactile code; Based on a preset semantic space alignment relationship, the generalized contour code and the initial tactile code are constrained and aligned through forward propagation to obtain a tactile code containing generalized contour information. The preset semantic space alignment relationship is established as follows: cosine similarity is calculated based on the generalized contour code and the initial tactile code to obtain a cosine similarity matrix; a loss function is calculated based on the cosine similarity matrix to obtain a first cross-entropy loss function; the first cross-entropy loss function is optimized using backpropagation to constrain the alignment of the generalized contour code and the initial tactile code; and the alignment parameters are adjusted through gradient updates to obtain the preset semantic space alignment relationship. Physical features are extracted from the generalized contour data and the tactile data respectively to obtain generalized contour physical features and tactile physical features; Inconsistency is calculated based on the generalized contour physical features and the tactile physical features to obtain inconsistency information; Cosine similarity is calculated based on the generalized contour encoding and the initial tactile encoding to obtain a cosine similarity matrix; similarity fusion is performed based on the cosine similarity matrix and the inconsistency information to obtain a fused similarity; and the semantic space alignment strategy of the generalized contour encoding and the initial tactile encoding is optimized based on the fused similarity. The generalized contour code and the initial tactile code are constrained and aligned based on the optimized semantic space alignment strategy to obtain a tactile code containing generalized contour information.

2. The method according to claim 1, characterized in that, Physical features are extracted from the generalized contour data to obtain the generalized contour physical features, including: The generalized contour data and the tactile data are subjected to rotation and translation transformations, and the optimal matching point is determined. Based on the optimal matching point, physical features are extracted from the generalized contour data to obtain the generalized contour physical features.

3. The method according to claim 2, characterized in that, Based on the tactile data, physical features are extracted to obtain tactile physical features, including: Based on the tactile data, regions that are in the same time period and whose loading states change are divided to obtain tactile zones; The target tactile area is obtained by removing tactile areas that have not been touched throughout the entire time period. The physical features of the target tactile area are extracted based on the optimal matching point to obtain the tactile physical features.

4. A method for object cognition based on robot tactile perception, characterized in that, include: Acquire generalized contour data and tactile data of the target object; The generalized contour coding model is used to learn features from the generalized contour data to obtain the generalized contour code; Using the method of any one of claims 1-3, feature learning is performed on the tactile data to obtain a tactile code containing generalized contour information; The noise is extracted from the generalized contour code with added random noise and the tactile code containing generalized contour information using the generalized contour diffusion model to obtain noise. Then, the noise is removed from the generalized contour code with added random noise based on the noise to obtain the target generalized contour code. The target generalized contour is decoded using a generalized contour coding model to obtain the true generalized contour of the target object.

5. The method according to claim 4, characterized in that, Obtain the generalized contour data of the target object, including: The point cloud of the target object is acquired using a visual acquisition device; Based on the point cloud, a second representative point is obtained by sampling. Obtain displacement information when a force is applied to the target object at the second representative point, causing the target object to shift; Based on the applied force and the corresponding displacement information, the proportional relationship between force and displacement is obtained; Based on the point cloud and the proportional relationship between force and displacement, the true generalized contour data of the target object is obtained.

6. The method according to claim 4, characterized in that, Acquire tactile data of the target object, including: Based on a preset position and preset direction, control the robot's end effector to move to the target position; Control the robot's end effector to grasp the target object with a preset load; The position and load of the robot's end effector are adjusted, and tactile data of the target object are collected during the adjustment process.

7. An object recognition device based on robotic tactile perception, characterized in that, include: A processor and a memory connected to the processor; wherein the memory stores instructions executable by the processor, the instructions being executed by the processor to cause the processor to perform the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Robot re-grabbing optimization method based on tactile primitive slip feature feedback

    CN116945166A

  • Tactile perception data processing method and device of intelligent agent, equipment and medium

    CN120122829A