Method, device, equipment and medium for processing tactile perception data of intelligent body

By acquiring low-resolution and high-resolution haptic data, using generative models for training, and generating target data processing models, the existing haptic sensors have solved the problems of high cost and limited resolution in super-resolution perception, and achieved low-cost and efficient haptic perception effects.

CN120122829BActive Publication Date: 2025-08-15BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510592552.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-15
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

When existing haptic sensors realize the super-resolution effect of haptic perception, there are problems of high cost and limited spatial resolution, especially the technical defects of the magnetic haptic sensor and visual haptic sensor cannot achieve high-precision haptic perception efficiently and at low cost.

Method used

By acquiring low-resolution tactile data and high-resolution tactile data, using preset generative models for training, generating target data processing models, realizing super-resolution processing of low-resolution tactile data, including data processing using conditional variational autoencoder models and diffusion models.

Benefits of technology

It realizes efficient and accurate super-resolution tactile perception of low-cost and low-resolution tactile sensors, and can provide high-resolution distributed tactile perception with true value, improving the accuracy and efficiency of tactile perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120122829B_ABST
    Figure CN120122829B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a method for processing tactile perception data for an intelligent agent, which can be applied to the field of artificial intelligence technology. The method includes: acquiring low-resolution tactile data and matching high-resolution tactile data; training a preset generative model based on the low-resolution and high-resolution tactile data to generate a target data processing model; and processing the target low-resolution tactile data using the target data processing model. Embodiments of the present invention also provide an intelligent agent tactile perception data processing device, apparatus, storage medium, and program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, specifically to the field of image processing technology, and more specifically to a method, device, equipment, medium and product for processing tactile perception data of an intelligent body. Background Art

[0002] Artificial Intelligence (AI) is a key driving force behind the new scientific and technological revolution and industrial transformation. It is a new, critical technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. As a key component of intelligence science, AI aims to understand the essence of intelligence and produce new intelligent machines (i.e., agents) that can respond in a manner similar to human intelligence.

[0003] In the field of artificial intelligence, tactile sensing technology focuses on simulating human skin to achieve tactile perception and interaction. Tactile sensing technology relies heavily on tactile sensors to sense and respond to embodied intelligent contact movements. Existing technologies include magnetic-based tactile sensors (MBTS) and vision-based tactile sensors (VBTS). However, these tactile sensors each have corresponding technical deficiencies that need to be addressed. For example, while magnetic tactile sensors combine the advantages of compact design and high-frequency operation, their sparse array of tactile units results in very limited spatial resolution of their tactile data. Meanwhile, vision-based tactile sensors are expensive and large in size, making it difficult to efficiently and cost-effectively achieve super-resolution tactile sensing. Summary of the Invention

[0004] In view of at least one of the above problems, the embodiments of the present invention aim to provide a tactile perception data processing method, apparatus, equipment, medium and product for an intelligent body that can achieve super-resolution perception based on low-resolution tactile perception data, thereby providing a method that can fully utilize high-resolution tactile perception data as a data source for calibrating and improving low-resolution tactile perception data, and ultimately achieve the super-resolution effect of low-resolution tactile perception data, greatly improve the resolution of low-cost low-resolution tactile sensors, and enable the intelligent body to perceive more refined tactile information using only low-cost tactile sensors.

[0005] One aspect of an embodiment of the present invention provides a method for processing tactile perception data of an intelligent body, which includes: obtaining low-resolution tactile data and matching high-resolution tactile data; training a preset generative model based on the low-resolution tactile data and the high-resolution tactile data to generate a target data processing model; and processing the target low-resolution tactile data through the target data processing model to achieve super-resolution tactile perception.

[0006] According to one embodiment of the present invention, in obtaining low-resolution tactile data and matching high-resolution tactile data, it includes: respectively obtaining low-resolution tactile data and high-resolution tactile data based on the same preset contact conditions; wherein the preset contact conditions include: a first contact module of a first tactile sensor that detects corresponding low-resolution tactile data and a second contact module of a second tactile sensor that detects corresponding high-resolution tactile data have a preset mapping relationship, and the contact position and contact force of the first contact module and the second contact module match; wherein the high-resolution tactile data is depth image data of the detection data obtained by the second tactile sensor.

[0007] According to one embodiment of the present invention, before training the preset generative model based on the low-resolution tactile data and the high-resolution tactile data to generate the target data processing model, it also includes: generating the prior distribution data, generated distribution data and variational approximation data corresponding to the low-resolution tactile data and the high-resolution tactile data according to the preset variational autoencoding rules corresponding to the preset generative model; generating the loss function corresponding to the preset generative model based on the prior distribution data, generated distribution data and variational approximation data.

[0008] According to one embodiment of the present invention, training a preset generative model based on low-resolution tactile data and high-resolution tactile data to generate a target data processing model includes: encoding and integrating the low-resolution tactile data and the high-resolution tactile data through the preset generative model; performing decoding and reconstruction processing on the latent variables corresponding to the received encoding and integration processing through the preset generative model; wherein, during the encoding and integration processing and decoding and reconstruction processing, the loss function is optimized until the target data processing model is generated.

[0009] According to one embodiment of the present invention, before training a preset generative model based on low-resolution tactile data and high-resolution tactile data to generate a target data processing model, it also includes: generating noisy tactile data in a forward diffusion process corresponding to the preset generative model by adding preset Gaussian noise to the high-resolution tactile data; performing denoising recovery from the noisy tactile data through the low-resolution tactile data to generate denoised tactile data in a reverse denoising process corresponding to the preset generative model; and generating an objective function of the preset generative model based on the noisy tactile data and the denoised tactile data.

[0010] According to one embodiment of the present invention, training a preset generative model based on low-resolution tactile data and high-resolution tactile data to generate a target data processing model includes: in a forward diffusion process and a reverse denoising process, optimizing the objective function until the target data processing model is generated.

[0011] According to one embodiment of the present invention, the preset generative model includes a conditional variational autoencoder model and / or a diffusion model.

[0012] Another aspect of an embodiment of the present invention provides a tactile perception data processing device for an intelligent agent, comprising a data acquisition module, a model generation module, and a data processing module. The data acquisition module is configured to acquire low-resolution tactile data and matching high-resolution tactile data; the model generation module is configured to train a preset generative model based on the low-resolution and high-resolution tactile data to generate a target data processing model; and the data processing module is configured to process target low-resolution tactile data using the target data processing model to achieve super-resolution tactile perception.

[0013] Another aspect of an embodiment of the present invention provides an electronic device comprising one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above-mentioned method for processing tactile perception data of an intelligent agent.

[0014] Another aspect of an embodiment of the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the above-mentioned method for processing tactile perception data of an intelligent agent.

[0015] Another aspect of an embodiment of the present invention provides a computer program product, including a computer program, which implements the above-mentioned tactile perception data processing method of the intelligent agent when executed by a processor.

[0016] The method for processing tactile perception data of an intelligent agent provided by an embodiment of the present invention can at least partially solve the problem in the related art of being unable to achieve super-resolution of tactile perception in the tactile perception process of an intelligent agent in an efficient and low-cost manner, and thus can achieve at least one of the following technical effects:

[0017] The tactile perception data processing method for an intelligent agent according to an embodiment of the present invention utilizes high-resolution tactile perception data as calibration data during the tactile shape reconstruction process to achieve super-resolution for low-resolution tactile perception data. This enables the intelligent agent to achieve efficient and accurate super-resolution tactile perception using only low-resolution tactile sensors, thereby providing a true, high-resolution distributed tactile perception solution. Specifically, by jointly designing an open-source detection device with the same contact module, a symmetrical calibration setup is used to achieve simultaneous data collection of low-resolution tactile data (such as magnetic signals detected by MBTS) and high-resolution tactile data (such as high-resolution shapes detected by VBTS). Tactile shape reconstruction is then treated as a conditional generative problem, and a conditional generative model is employed to predict super-resolution shapes from low-resolution tactile data input.

[0018] Quantitative experiments have shown that shape reconstruction can be effectively performed while achieving an operating frequency exceeding 250Hz. This cross-modal synergy greatly improves the tactile perception capabilities of low-resolution tactile sensors, including magnetic tactile sensors (MBTS), and can unlock new capabilities for high-precision intelligent robotic tasks.

[0019] It should be understood that the foregoing general description and the following detailed description are merely exemplary and illustrative and are not intended to limit the scope of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:

[0021] Figure 1 A diagram schematically illustrates an application scenario of a method, apparatus, device, medium, and program product for processing tactile perception data of an intelligent agent according to an embodiment of the present invention;

[0022] Figure 2 A flowchart of a method for processing tactile perception data of an intelligent agent according to an embodiment of the present invention is schematically shown;

[0023] Figure 3A A diagram schematically shows a pair of symmetrical structure calibration devices according to a method for processing tactile perception data of an intelligent agent according to an embodiment of the present invention;

[0024] Figure 3B A diagram schematically illustrates a data training scenario of a preset generative model execution method of a tactile perception data processing method for an intelligent agent according to an embodiment of the present invention;

[0025] Figure 3CA schematic diagram of a scene diagram of shape reconstruction of tactile data of a magnetic tactile sensor performed according to a reconstruction network architecture of a target data processing model in a method for processing tactile perception data of an intelligent agent according to an embodiment of the present invention is shown;

[0026] Figure 4A A diagram schematically illustrates a data training scenario for a target data processing model of a Conditional Variational Auto Encoder (CVAE) in a method for processing tactile perception data of an intelligent agent according to an embodiment of the present invention;

[0027] Figure 4B A diagram schematically illustrates a data training scenario for a target data processing model of a diffusion model (DM) in a method for processing tactile perception data of an intelligent agent according to an embodiment of the present invention;

[0028] Figure 5 A block diagram schematically shows a structure of a tactile perception data processing device for an intelligent agent according to an embodiment of the present invention; and

[0029] Figure 6 The block diagram of an electronic device suitable for implementing the tactile perception data processing method of an intelligent agent according to an embodiment of the present invention is schematically shown.

[0030] The above-mentioned drawings are part of the description of the embodiments of the present invention, illustrating exemplary embodiments of the present invention. Together with the description, the drawings are used to illustrate the principles of the embodiments of the present invention. It should be understood that the above general description of the drawings and the following detailed description are merely exemplary and illustrative and are not intended to limit the scope of the present invention. DETAILED DESCRIPTION

[0031] In order to make the objectives, technical solutions and advantages of the embodiments of the present invention more clearly understood, the spirit of the contents disclosed in the present invention will be clearly illustrated with the accompanying drawings and detailed descriptions below. After understanding the embodiments of the contents of the present invention, any technician in the relevant technical field can change and modify the contents of the present invention based on the techniques taught by the contents of the present invention without departing from the spirit and scope of the contents of the present invention.

[0032] The exemplary embodiments of the present invention and their description are used to explain the present invention, but are not intended to limit the present invention. In addition, elements / components with the same or similar reference numerals used in the drawings and embodiments are used to represent the same or similar parts.

[0033] The terms “first,” “second,” etc. used in the present invention do not particularly refer to an order or sequence, nor are they used to limit the present invention. They are only used to distinguish elements or operations described with the same technical terms.

[0034] The directional terms used in the present invention, such as up, down, left, right, front, or back, are only used with reference to the directions in the accompanying drawings. Therefore, the directional terms used are used to illustrate and not to limit the present invention.

[0035] The terms “include,” “including,” “have,” “contain,” etc. used in the present invention are open-ended terms, meaning including but not limited to.

[0036] The term "and / or" used in the present invention includes any or all combinations of the items mentioned.

[0037] Regarding the present invention, "plurality" includes "two" and "more than two"; regarding the present invention, "plurality of groups" includes "two groups" and "more than two groups".

[0038] The terms "substantially" and "approximately" used in this disclosure are intended to modify any quantity or error that may vary slightly, but such variations or errors do not alter the essence of the quantity. Generally speaking, the range of such variations or errors may be 20% in some embodiments, 10% in some embodiments, 5% in some embodiments, or other values. Those skilled in the art will appreciate that the aforementioned values may be adjusted based on actual needs and are not intended to be limiting.

[0039] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0040] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to systems having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, and C, etc.). When expressions such as "at least one of A, B, or C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, or C" should include but is not limited to systems having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, and C, etc.). Those skilled in the art should also understand that any transitional conjunctions and / or phrases indicating two or more optional items, whether in the specification, claims, or drawings, should be understood to provide the possibility of including one, either, or both of these items. For example, the phrase "A or B" should be understood to include the possibilities of "A" or "B", or "A and B".

[0041] In addition to visual tactile sensors, tactile sensors based on other electrical signals generally have limited spatial resolution due to the limited number of measurement units and the difficulty of wiring.

[0042] Taking magnetic tactile sensors as an example, current research on flexible magnetic tactile sensors focuses on the decoupling of three-dimensional force information. Some researchers use theoretical derivation to study how deformation caused by external forces affects the magnetic field distribution on the sensor surface, as well as the relationship between the nonlinear relationship between compressive stress and axial strain and the deformation of the magnetic skin. In addition, a common research method currently uses a data-driven method to establish a mapping relationship between the three-axis magnetic field intensity signal and the three-dimensional force decoupling signal. There are also some studies exploring the super-resolution of magnetic skin, but these methods all use a single-point probe to traverse the entire skin surface and then use a data-driven neural network method to achieve super-resolution. The disadvantage of this is that the model can only achieve single-point super-resolution perception and cannot perceive multi-point contact. The reason is that there is currently a lack of high-resolution distributed tactile perception systems that can provide true values.

[0043] In view of at least one of the above problems, the embodiments of the present invention aim to provide a tactile perception data processing method, apparatus, equipment, medium and product for an intelligent body that can achieve super-resolution perception based on low-resolution tactile perception data, thereby providing a method that can fully utilize high-resolution tactile perception data as a data source for calibrating and improving low-resolution tactile perception data, and ultimately achieve the super-resolution effect of low-resolution tactile perception data, greatly improve the resolution of low-cost low-resolution tactile sensors, and enable the intelligent body to perceive more refined tactile information using only low-cost tactile sensors.

[0044] One aspect of an embodiment of the present invention provides a method for processing tactile perception data of an intelligent body, which includes: obtaining low-resolution tactile data and matching high-resolution tactile data; training a preset generative model based on the low-resolution tactile data and the high-resolution tactile data to generate a target data processing model; and processing the target low-resolution tactile data through the target data processing model.

[0045] Figure 1 The application scenario diagram of the tactile perception data processing method, device, equipment, medium and program product of an intelligent agent according to an embodiment of the present invention is schematically shown.

[0046] like Figure 1 As shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables.

[0047] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0048] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.

[0049] Server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using terminal devices 101, 102, and 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal device.

[0050] It should be noted that the tactile perception data processing method of the intelligent agent provided in the embodiment of the present invention can generally be executed by the server 105. Accordingly, the tactile perception data processing device of the intelligent agent provided in the embodiment of the present invention can generally be set in the server 105. The tactile perception data processing method of the intelligent agent provided in the embodiment of the present invention can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the tactile perception data processing device of the intelligent agent provided in the embodiment of the present invention can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.

[0051] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0052] The following will be based on Figure 1 The scene described by Figures 2 to 4B The tactile perception data processing method of the intelligent agent in the disclosed embodiment is described in detail.

[0053] like Figure 2 As shown, one aspect of an embodiment of the present invention provides a method for processing tactile perception data of an intelligent agent, which includes operations S201 to S203.

[0054] In operation S201 , low-resolution tactile data and high-resolution tactile data matching the low-resolution tactile data are acquired;

[0055] In operation S202 , a preset generative model is trained according to the low-resolution tactile data and the high-resolution tactile data to generate a target data processing model; and

[0056] In operation S203 , the target low-resolution tactile data is processed by the target data processing model to achieve super-resolution tactile perception.

[0057] The intelligent agent can be the executor of the tactile perception data processing method described above in an embodiment of the present invention, or it can be an executor controlled by the tactile perception data processing method. Specifically, it can be a humanoid intelligent robot or other AI device, typically having its own actuators capable of completing specific motion tasks. At least one actuator of the intelligent agent can be equipped with a corresponding tactile sensor, such as a magnetic tactile sensor (MBTS), to enable the intelligent agent to sense tactile objects.

[0058] Low-resolution tactile data can be tactile data acquired by a relatively low-cost tactile sensor, such as sensor reading data acquired by a magnetic tactile sensor (MBTS). Conversely, high-resolution tactile data can be tactile data acquired by a relatively high-cost, larger tactile sensor, such as sensor data acquired by a visual tactile sensor (VBTS), or depth image data corresponding to the sensor data.

[0059] The spatial resolution of the tactile perception image corresponding to the high-resolution tactile data is higher than the spatial resolution of the tactile perception image corresponding to the low-resolution tactile data. In addition, the low-resolution tactile data and the high-resolution tactile data can be acquired simultaneously and separately through matching tactile perception detection conditions to ensure that the two data can match each other. Specifically, Figure 3A shown.

[0060] The preset generative model can be a neural network architecture model that uses a neural network as a function mapped to a probability distribution and then samples from the distribution to obtain the final result. Its output can be a probability distribution in space, such as the CVAE conditional variational autoencoder model and the DM diffusion model. By using the above-mentioned low-resolution tactile data and high-resolution tactile data as training data for the above-mentioned preset generative model, the target data processing model after the final convergence can be generated through the model training process. Therefore, the target data processing model can be a generative model generated after data training with the above-mentioned low-resolution tactile data and high-resolution tactile data.

[0061] After the above-mentioned target data processing model is obtained, the target low-resolution tactile data can be used as input data. Since the target data processing model is a generative model generated after data training using the above-mentioned low-resolution tactile data and high-resolution tactile data, the target super-resolution tactile data can be output.

[0062] Therefore, through the tactile perception data processing method of the intelligent agent in the embodiment of the present invention, the tactile shape reconstruction process can utilize high-resolution tactile perception data as calibration data to achieve super-resolution effects on low-resolution tactile perception data. This enables the intelligent agent to achieve efficient and accurate super-resolution tactile perception using only low-resolution tactile sensors, and can provide a true high-resolution distributed tactile perception solution. Specifically, by jointly designing an open-source detection device with the same contact module, a symmetrical calibration setup is used to achieve simultaneous data collection of low-resolution tactile data (such as magnetic signals detected by MBTS) and high-resolution tactile data (such as high-resolution shapes detected by VBTS). Tactile shape reconstruction is regarded as a conditional generation problem, and a conditional generation model is used to predict super-resolution shapes from the low-resolution tactile data input.

[0063] In order to enable those skilled in the art to have a clearer understanding of the tactile perception data processing method of the above-mentioned intelligent agent in the embodiment of the present invention, the following is further provided: Figures 3A-4B Description.

[0064] like Figure 2-4B As shown, according to an embodiment of the present invention, in operation S201, obtaining low-resolution tactile data and high-resolution tactile data matching therewith includes:

[0065] respectively acquiring low-resolution tactile data and high-resolution tactile data based on the same preset contact conditions;

[0066] Among them, the preset contact conditions include: there is a preset mapping relationship between the first contact module of the first tactile sensor that detects corresponding low-resolution tactile data and the second contact module of the second tactile sensor that detects corresponding high-resolution tactile data, and the contact position and contact force of the first contact module and the second contact module match; wherein the high-resolution tactile data is depth image data of the detection data obtained by the second tactile sensor.

[0067] To achieve the initial super-resolution setup of a tactile sensor (such as a magnetic tactile sensor (MBTS)) corresponding to low-resolution tactile data guided by a tactile sensor (such as a visual tactile sensor (VBTS)) corresponding to high-resolution tactile data, it is necessary to ensure that both sensors operate under the same contact conditions. These preset contact conditions involve two key aspects:

[0068] 1) The contact modules of the two sensors must have a preset mapping relationship, where the preset mapping relationship can be, for example, the same material and size, or a mapping relationship established between the materials and sizes to ensure that the difference in contact deformation is minimal;

[0069] 2) During data acquisition, the contact position and contact force detected by the two sensors must match, for example, the midlines of the contact positions must be consistent and the tactilely sensed contact forces must be the same.

[0070] like Figure 3A The calibration setup, shown in Figure 1, uses a visual tactile sensor (VBTS) and a magnetic tactile sensor (MBTS) as examples. The two different tactile sensors can be mounted on either side of the gripper, with their contact modules aligned along the centerline. The tactile images from the visual tactile sensor (VBTS) are calibrated, mirrored, and cropped to match the dimensions of the magnetic tactile sensor (MBTS). The original sampled data (325×325 pixels per frame) is downsampled to 200×200 pixels for use as a sampling dataset, which is then used as training data.

[0071] Among them, Figure 3A As shown in Figure 2, a 3D printed hexagonal wrench and the letter "R" can be used as calibration objects during the training data collection process. The gripper can be closed at a pressure threshold of 10N, as shown in Figure 2. Figure 3B The final dataset includes 2025 pairs of tactile readings of hexagonal wrenches and 2025 pairs of tactile readings of the letter "R".

[0072] Among them, the visual tactile sensor VBTS can achieve ultra-high spatial resolution tactile perception by using a camera to capture the deformation of flexible materials. The deformation information of the surface depth can be obtained through sensor calibration, so it can be used as a high-resolution tactile system that provides the true value of deformation.

[0073] Therefore, by using high-resolution tactile data from high-resolution tactile sensors such as visual tactile sensors as the true depth value, and by Figure 3A The modular, detachable, and reusable calibration plate system shown can realize a super-resolution system for multi-point contact perception. At the same time, the calibration plate system also significantly improves the efficiency of calibration.

[0074] Compared to the low-resolution tactile data x of the magnetic tactile sensor MBTS (such as MBTS sensor readings), the reconstruction of the high-resolution tactile data of the visual tactile sensor VBTS is usually based on the depth value I in the multimodal data (such as visual information, contact force information, and surface texture information) contained in the tactile data of the visual tactile sensor VBTS. Specifically, due to changes in surface texture, lighting conditions, or sensor noise, multiple valid visual results may correspond to similar magnetic measurement values. Figure 3CAs shown in Figure 3, the generative model provides a probabilistic reconstruction network architecture (Recon Network) that is well-suited to handling this uncertainty. Unlike deterministic models, which map low-resolution tactile data x to a single estimate of high-resolution tactile data I, the generative model models the distribution of possible outputs given the low-resolution tactile data x. This capability enables the generative model to capture complex relationships and variations in tactile data, thereby providing richer predictions.

[0075] like Figure 2-4B As shown, according to one embodiment of the present invention, the preset generative model includes a conditional variational autoencoder model and / or a diffusion model. Figure 4A As shown in , the conditional variational autoencoder model can be a deep learning model, which mainly includes two parts: encoder and decoder. The encoder can compress the input data into low-dimensional latent variables, while the decoder can reconstruct or generate data based on these latent variables and conditional input. Figure 4B As shown in the figure, the diffusion model is a type of generative model, which usually gradually adds noise to the input data to turn it into pure random noise data, and then uses the neural network architecture to gradually remove the noise in reverse and finally restore the input data.

[0076] Among them, the following text takes "the data collected by the magnetic tactile sensor MBTS corresponds to low-resolution tactile data, and the depth image data of the data collected by the visual tactile sensor VBTS corresponds to high-resolution tactile data" and "the preset generative models are the conditional variational autoencoder model CVAE and the diffusion model DM" as examples to explain in detail the data training process for the preset generative model combining low-resolution tactile data and high-resolution tactile data.

[0077] like Figure 2-4B As shown, according to an embodiment of the present invention, before operation S202 of training a preset generative model according to the low-resolution tactile data and the high-resolution tactile data to generate a target data processing model, the method further includes:

[0078] Generate prior distribution data, generative distribution data, and variational approximation data corresponding to the low-resolution tactile data and the high-resolution tactile data according to a preset variational autoencoding rule corresponding to the preset generative model;

[0079] Generate the loss function corresponding to the preset generative model based on the prior distribution data, generated distribution data and variational approximation data.

[0080] The preset variational autoencoder rule can be a data training rule of a preset generative model defined by a conditional variational autoencoder model CVAE. The magnetic tactile sensor MBTS reading can be regarded as low-resolution tactile data and used as the conditional variable x of the preset generative model, which is typically a vector from a magnetometer (such as Figure 3B and Figure 3C Sensor readings are shown. To achieve the learning goal of mapping from low-resolution tactile data x to high-resolution tactile data I, a neural network architecture based on high-resolution tactile data can be constructed, where the high-resolution tactile data I can represent the corresponding depth image data generated by the visual tactile sensor VBTS.

[0081] In the conditional variational autoencoder model CVAE, according to its preset variational autoencoder rules, a latent variable z can be used for modeling to define the conditional distribution data of the preset generative model. , as shown in the following formula 1:

[0082]

[0083] in, represents the prior distribution data (such as can be fixed to standard Gaussian distribution data), and It is to generate distribution data. Since the true posterior distribution is usually impossible to calculate, it is necessary to introduce approximate posterior distribution data (i.e., variational approximation data), where and They are the learnable parameters of the decoder and encoder networks defined by the conditional variational autoencoder model CVAE.

[0084] To train the conditional variational autoencoder model CVAE, we can maximize the evidence lower bound (ELBO) of the log-likelihood of the high-resolution tactile data I given the low-resolution tactile data x, which is expressed in the form of a loss function as follows:

[0085] (2)

[0086] in, is the Kullback–Leibler divergence (KL divergence), which is used to measure the variational approximation data Data with prior distribution The first term in Equation 2 above forces the decoder to define the generated distribution data It can accurately recover the high-resolution tactile data I from the latent variable z and the low-resolution tactile data x; and the second constraint encoder , making it close to the prior distribution data , thereby preventing overfitting and enhancing generalization ability, as shown in Figure 4A shown.

[0087] During backpropagation, a reparameterization trick can be used so that gradient optimization can still be performed during the sampling of the latent variable z.

[0088] like Figure 2-4B As shown, according to one embodiment of the present invention, in operation S202, a preset generative model is trained according to the low-resolution tactile data and the high-resolution tactile data to generate a target data processing model, including:

[0089] Encoding and integrating low-resolution tactile data and high-resolution tactile data through a preset generative model;

[0090] Performing decoding and reconstruction processing on the received latent variables corresponding to the encoding and integration processing through a preset generative model;

[0091] In the process of encoding integration and decoding reconstruction, the loss function is optimized until the target data processing model is generated.

[0092] like Figure 4A As shown, the network architecture of the conditional variational autoencoder model CVAE is shown. Encoder network and decoder network By parameters and express.

[0093] The encoder processes high-resolution tactile data I (such as depth image data corresponding to the visual tactile sensor (VBTS)) through convolutional (Conv) and fully connected (FC) layers. Its first layer uses residual connections to generate parameters that approximate the posterior distribution. The conditional input (i.e., readings from the magnetic tactile sensor (MBTS) as low-resolution tactile data x) is embedded through a separate fully connected layer and incorporated into the encoder as a multiplicative condition in two stages to enhance memory retention and prevent catastrophic forgetting.

[0094] The decoder can receive the latent variable z sampled using the reparameterization technique and generate a reconstructed image through convolutional layers and fully connected layers. Similar to the encoder, the low-resolution tactile data x of the magnetic tactile sensor MBTS can also be used as a multiplicative condition in the decoder. Modeled as a Gaussian distribution whose mean and variance are learned by the decoder to generate the final reconstructed image .

[0095] During the training process, we can use mini-batch stochastic gradient descent (SGD) to optimize the negative evidence lower bound (Negative ELBO, i.e., the negative loss function) until convergence. In order to make the training more stable, we can use the reparameterization technique to sample from the standard Gaussian distribution output by the encoder. , and calculate ,in and are the mean and standard deviation of the distribution.

[0096] In addition, the robustness and generalization ability of the network can be improved through techniques such as batch normalization, group normalization, and max / average pooling. At the same time, the selection of the number of latent dimensions is based on the balance between reconstruction accuracy and generation diversity.

[0097] Therefore, after the training is completed, the corresponding target data processing model can be generated. Through the target data processing model, for a new reading x′ of the magnetic tactile sensor MBTS, the prior distribution Sample the latent variable z and generate multiple possible prediction results ′, to capture the uncertainty and variability in tactile perception.

[0098] It should be noted that in actual application scenarios, compared to the Generative Adversarial Network (GAN) or the Diffusion Model (DM), the Conditional Variational Autoencoder (CVAE) model typically has a simpler network structure and inference process, and has higher computational efficiency. This computational efficiency makes it more suitable for high-frequency operation in the target data processing model training scenario of tactile perception data in the embodiments of the present invention, where fast inference of low-resolution tactile data (such as magnetic tactile sensor (MBTS) data) is required.

[0099] In addition, if Figure 3B As shown in Figure 3, by conditionally constraining the low-resolution tactile data x during the generation process, the conditional generative model can also ensure that the latent representation can be aligned with the specific features of the magnetic measurement data, making the generation process more interpretable and controllable than the standard (unconditional) generative model.

[0100] like Figure 2-4B As shown, according to an embodiment of the present invention, before operation S202 of training a preset generative model according to the low-resolution tactile data and the high-resolution tactile data to generate a target data processing model, the method further includes:

[0101] By adding preset Gaussian noise to the high-resolution tactile data, noisy tactile data in a forward diffusion process corresponding to a preset generative model is generated;

[0102] Performing denoising recovery from the noisy tactile data using the low-resolution tactile data to generate denoised tactile data in a reverse denoising process corresponding to a preset generative model;

[0103] An objective function of a preset generative model is generated based on the noisy tactile data and the denoised tactile data.

[0104] As mentioned above, the diffusion model (DM) is a generative model that simulates physical diffusion processes (such as gas diffusion) to gradually denoise Gaussian noise and generate data from a specific distribution. This process can be divided into two stages: forward diffusion and backward denoising.

[0105] (1) Forward diffusion process: gradually adding noise to the data to transform the complex data distribution into a simple distribution (such as Gaussian distribution);

[0106] (2) Reverse denoising process: learning to gradually restore the original data distribution from the noise distribution to generate high-quality samples.

[0107] Similar to the above-mentioned conditional variational autoencoder model CVAE, the tactile readings of a magnetic tactile sensor such as MBTS can be regarded as low-resolution tactile data x to be used as a conditioning variable, which is usually a vector from a magnetometer, such as Figure 3B and Figure 3C shown.

[0108] In order to learn the mapping from low-resolution tactile data x to high-resolution tactile data I, in the modeling process based on the Denoising Diffusion Probabilistic Models (DDPM), we can learn the conditional distribution The model is built through a step-by-step denoising process, wherein the high-resolution tactile data I may represent corresponding depth image data generated from a visual tactile sensor (VBTS).

[0109] Specifically, in the forward diffusion process, Gaussian noise can be gradually added to the high-resolution tactile data I to generate a series of noisy data I1, I2, ..., I T , where I T Close to pure noise. The added Gaussian noise can be preset according to the actual application scenario requirements, that is, preset Gaussian noise; the series of noise-containing data generated above can be the noise tactile data.

[0110] The process of gradually adding Gaussian noise to generate a series of noise data can be expressed as the following formula 3:

[0111] (3)

[0112] in, is a noise scheduling parameter to control the noise intensity at each step, which can usually be a manually set hyperparameter.

[0113] In addition, in the reverse denoising process, a denoising model can be learned , the model is derived from noisy tactile data , supplemented by the low-resolution tactile data x of the sensor data as a condition, and gradually restore the original data Among them, the original data The tactile data can be denoised.

[0114] The above reverse denoising process can be expressed as the following formula 4:

[0115] (4)

[0116] in and are the mean and variance parameterized by the denoising diffusion probability model DDPM neural network.

[0117] The goal of training the denoising diffusion probability model DDPM is to minimize the KL divergence between the forward process and the posterior denoising process. Specifically, the objective function of the following formula 5 can be defined and optimized accordingly:

[0118] (5)

[0119] in, is the preset Gaussian noise added during the forward diffusion process, is the noise predicted by the above denoising model.

[0120] Similar to the above-mentioned conditional variational autoencoder model CVAE, the denoising diffusion probability model DDPM also relies on the reparameterization technique to perform gradient optimization on noise sampling during the backpropagation process.

[0121] Specifically, the noise tactile data It can be expressed as the following formula 6:

[0122] (6)

[0123] in, , is the standard Gaussian noise added above.

[0124] like Figure 2-4B As shown, according to one embodiment of the present invention, in operation S202, a preset generative model is trained according to the low-resolution tactile data and the high-resolution tactile data to generate a target data processing model, including:

[0125] In the forward diffusion process and the reverse denoising process, the objective function is optimized until the target data processing model is generated.

[0126] like Figure 4B As shown in the figure, through this modeling scheme of the diffusion model DM based on the denoising diffusion probability model DDPM, the depth image corresponding to the high-resolution tactile data I can be gradually recovered from the noisy data, and the low-resolution tactile data x is used as a conditional variable to guide the generation process, thereby achieving more accurate mapping learning.

[0127] Denoising model based on denoising diffusion probability model DDPM , we can use the UNet network structure, such as Figure 4B shown.

[0128] The downsampling (encoding) process consists of five parts: one residual block, three convolutional blocks, and one average pooling block (AvgPool). The residual block contains two convolutional layers (Conv) and one batch normalization layer (BatchNorm), using GELU as the activation function. The convolutional and residual blocks are essentially the same, with the exception of an additional max pooling layer (MaxPool).

[0129] During the downsampling process, the conditions introduced by the sensor data are embedded twice in a multiplicative manner. Meanwhile, the upsampling (decoding) process consists of one deconvolution block (ConvT), three convolution blocks, and one output block. The deconvolution block uses group normalization (GroupNorm) and ReLU as activation functions; the convolution block consists of one deconvolution layer and two small convolution blocks. The small convolution blocks are similar in structure to the convolution blocks used in downsampling (excluding MaxPool). The output block has a similar structure to the deconvolution block, except that it contains two deconvolution layers at the beginning and end.

[0130] Similar to the downsampling process, the introduction of conditions in the upsampling process also uses a fully connected layer (FC) to perform two multiplicative embeddings. In addition, the time step t is also embedded twice in the upsampling process using a fully connected layer (FC). At the same time, the three intermediate results of the downsampling are directly passed to the upsampling / encoder using skip connections. In terms of hyperparameter selection, linear noise The range is 1e-4 to 0.02, 400 denoising steps, fixed learning rate 1e-4, optimizer is AdamW, and learning is done for about 400 epochs.

[0131] Therefore, by repeatedly optimizing the objective function in the above-mentioned forward diffusion process and reverse denoising process until the preset generative model converges, the corresponding target data processing model can be obtained.

[0132] In order to further verify the processing capability of the target data processing model, reflect the technical effect of the tactile perception data processing method of the intelligent agent in the embodiment of the present invention, and also reflect the step of "processing the target low-resolution tactile data through the target data processing model" defined in operation S203 in the embodiment of the present invention, the following model verification experiment is used to verify and evaluate the model processing effect of the shape reconstruction quality on a series of different visible and invisible (Unseen) partial and complete objects.

[0133] Specifically, from the collected dataset, we selected the Allen wrench and the letter "R" for training due to their unique shape features. The high-resolution tactile data from the visual tactile sensor (VBTS) was denoised and downsampled from 325×325 pixels to 200×200 pixels to balance image quality and training efficiency. The low-resolution tactile data from the magnetic tactile sensor (MBTS) was limited to a range of ±500 and normalized to the interval [-1, 1] to ensure training stability.

[0134] The network architecture of the lightweight conditional variational autoencoder (CVAE) model contains 47.2 million parameters. The encoder consists of six CNN blocks, including two max-pooling layers and one average-pooling layer, followed by two fully-connected layers (latent dimension 512). All encoder blocks use BatchNorm and GELU activation functions. The decoder consists of four deconvolution blocks and eight intermediate CNN blocks. The deconvolution blocks (ConvT) use GroupNorm and ReLU, and the CNN blocks use BatchNorm and GELU activation functions. Conditions are embedded in both the encoder and decoder via two fully-connected layers (with GELU activation). Training was performed on an NVIDIA RTX 3090 GPU for approximately 350 epochs. The training set contains 1822 Allen wrenches and 1822C Rs (a total of 3645 examples). The AdamW optimizer was used with a fixed learning rate of 1e-4. On the same hardware, the inference time for shape reconstruction from a single reading of the magnetic tactile sensor (MBTS) is within 25 milliseconds.

[0135] The tactile shape reconstruction capabilities of the super-resolution solution implemented by the tactile perception data processing method according to an embodiment of the present invention (such as SuperMag described below) were evaluated against those of three baseline methods:

[0136] (1) Perform bilinear / bicubic interpolation on the data of the z-axis magnetic tactile sensor MBTS as a simple spatial upsampling baseline;

[0137] (2) SuperMag, a super-resolution method based on z-axis data, uses only z-axis magnetic data for reconstruction;

[0138] (3) SuperMag, a super-resolution method based on AK data, is trained only on the Allen wrench data (half of the training set) to analyze the limitations of generalization ability.

[0139] (4) The standard super-resolution method uses the complete three-axis magnetotactile sensor MBTS data and multi-object training (hexagonal wrench + letter "R").

[0140] Among them, all methods are tested on unseen objects to evaluate their generalization ability.

[0141] In order to evaluate the performance of the proposed super-resolution method on tactile shape reconstruction, we can use it on a set of unseen objects (e.g. Figure 3C Figure 3 compares different models and versions of super-resolution methods using three metrics: peak signal-to-noise ratio, structural similarity, and Fréchet inception distance:

[0142] (1) Peak signal-to-noise ratio (PSNR) can measure the pixel-level accuracy of each image. A PSNR value below 20dB indicates serious defects and is unacceptable.

[0143] (2) Structural Similarity (SSIM): This evaluates the visual quality of each image, taking into account brightness, contrast, and texture. The SSIM score ranges from -1 to 1, with negative values being unacceptable and higher values indicating better quality.

[0144] (3) Fréchet Initial Distance (FID): Evaluates the similarity between the reconstructed image and the real visual image. Lower values indicate better alignment.

[0145] The specific comparison information of the validation data of SuperMag and the baseline method (Baselines) for peak signal-to-noise ratio (PSNR), structural similarity (SSIM), and Fréchet inception distance (FID) is shown in Table 1 below.

[0146]

[0147] As shown in Table 1, SuperMag outperforms all baseline methods across all metrics, demonstrating its high efficiency for tactile shape reconstruction. Bilinear and bicubic interpolation methods perform poorly, confirming that simple upsampling cannot recover high-resolution tactile details. Super-resolution methods using only single-axis magnetic data (z-axis) achieve modest improvements but are still limited by incomplete deformation information. SuperMag (AK), a super-resolution method trained on only a single object (an Allen wrench), further reduces the Fréchet Inception Distance (FID) and improves the Peak Signal-to-Noise Ratio (PSNR), highlighting the advantages of multi-axis magnetic conditioning. The full SuperMag achieves the best performance by leveraging multi-axis data and multi-object training. This emphasizes the critical role of training diversity and cross-modal supervision. High structural similarity (SSIM) values indicate high perceptual fidelity of the reconstructed shape, while the Peak Signal-to-Noise Ratio (PSNR) verifies pixel-level accuracy.

[0148] The results show that combining the multi-axis magnetic tactile sensor MBTS data with the supervision guided by the visual tactile sensor VBTS effectively solves the trade-off between resolution and frequency, and realizes real-time high-resolution tactile perception for the robot.

[0149] Achieving efficient and accurate super-resolution tactile perception through low-resolution tactile sensors can provide a true high-resolution distributed tactile sensing solution. Specifically, by jointly designing an open-source detection device with the same contact module, a symmetrical calibration setup is used to achieve simultaneous data collection of low-resolution tactile data (such as magnetic signals detected by MBTS) and high-resolution tactile data (such as high-resolution shape detected by VBTS). Tactile shape reconstruction is formulated as a conditional generative problem, and a conditional generative model is used to predict super-resolution shapes from low-resolution tactile data input.

[0150] Quantitative experiments have shown that shape reconstruction can be effectively performed while achieving an operating frequency exceeding 250Hz. This cross-modal synergy greatly improves the tactile perception capabilities of low-resolution tactile sensors, including magnetic tactile sensors (MBTS), and can unlock new capabilities for high-precision intelligent robotic tasks.

[0151] Based on the above-mentioned method for processing tactile perception data of an intelligent agent, the present invention also provides a device for processing tactile perception data of an intelligent agent. Figure 5 The device is described in detail.

[0152] Figure 5 The structure block diagram of the tactile perception data processing device 500 of an intelligent agent according to an embodiment of the present invention is schematically shown.

[0153] like Figure 5As shown, the tactile perception data processing device 500 of the intelligent agent in this embodiment includes a data acquisition module 510, a model generation module 520 and a data processing module 530.

[0154] The data acquisition module 510 is used to acquire low-resolution tactile data and high-resolution tactile data that matches the low-resolution tactile data. In one embodiment, the data acquisition module 510 can be used to perform the operation S201 described above, which will not be described in detail here.

[0155] The model generation module 520 is used to train a preset generative model based on the low-resolution tactile data and the high-resolution tactile data to generate a target data processing model. In one embodiment, the model generation module 520 can be used to perform the operation S202 described above, which will not be repeated here.

[0156] The data processing module 530 is used to process the target low-resolution tactile data using the target data processing model to achieve super-resolution tactile perception. In one embodiment, the data processing module 530 can be used to perform the operation S203 described above, which will not be repeated here.

[0157] According to embodiments of the present invention, any multiple modules among the data acquisition module 510, the model generation module 520, and the data processing module 530 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present invention, at least one of the data acquisition module 510, the model generation module 520, and the data processing module 530 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or any suitable combination of these. Alternatively, at least one of the data acquisition module 510, the model generation module 520, and the data processing module 530 may be at least partially implemented as a computer program module that, when executed, performs the corresponding functionality.

[0158] Figure 6 The block diagram of an electronic device suitable for implementing the tactile perception data processing method of an intelligent agent according to an embodiment of the present invention is schematically shown.

[0159] The above-mentioned electronic device provided by an embodiment of the present invention includes one or more processors and a memory, and the memory is used to store one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors execute the above-mentioned intelligent body's tactile perception data processing method.

[0160] like Figure 6 As shown, an electronic device 600 according to an embodiment of the present invention includes a processor 601, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 602 or programs loaded from a storage unit 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or related chipsets and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0161] Various programs and data required for the operation of the electronic device 600 are stored in the RAM 603. The processor 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The processor 601 executes the programs in the ROM 602 and / or RAM 603 to perform various operations according to the method flow of the embodiment of the present invention. It should be noted that the programs may also be stored in one or more memories other than the ROM 602 and RAM 603. The processor 601 may also execute the programs stored in the one or more memories to perform various operations according to the method flow of the embodiment of the present invention.

[0162] According to an embodiment of the present invention, electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to bus 604. Electronic device 600 may also include one or more of the following components connected to I / O interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 608 including a hard disk; and a communication section 609 including a network interface card such as a LAN card or modem. Communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to I / O interface 605 as needed. Removable media 611, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 610 as needed, so that computer programs read from the removable media can be installed into storage section 608 as needed.

[0163] The present invention also provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the above-mentioned method for processing tactile perception data of an intelligent agent.

[0164] The computer-readable storage medium may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiment of the present invention.

[0165] According to an embodiment of the present invention, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, a computer-readable storage medium may include the ROM 602 and / or RAM 603 described above, and / or one or more memories other than ROM 602 and RAM 603.

[0166] An embodiment of the present invention also includes a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the above-mentioned tactile perception data processing method of the intelligent agent.

[0167] The computer program includes program codes for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program codes are used to enable the computer system to implement the method provided by the embodiment of the present invention.

[0168] The computer program executes the above functions defined in the system / device of the embodiment of the present invention when the computer program is executed by the processor 601. According to the embodiment of the present invention, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0169] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 609, and / or installed from a removable medium 611. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0170] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609 and / or installed from a removable medium 611. When the computer program is executed by the processor 601, the above-described functions defined in the system of the embodiment of the present invention are performed. According to the embodiment of the present invention, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.

[0171] According to an embodiment of the present invention, the program code for executing the computer program provided by the embodiment of the present invention can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0172] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0173] In addition, all actions of acquiring information, signals or data in the present invention are carried out in compliance with the relevant data protection laws, regulations and policies of the country where they are located, and with the authorization given by the owner of the corresponding device.

[0174] Those skilled in the art will appreciate that various combinations and / or combinations of features described in the various embodiments and / or claims of the present invention may be made, even if such combinations and / or combinations are not explicitly described in the present invention. In particular, various combinations and / or combinations of features described in the various embodiments and / or claims of the present invention may be made, without departing from the spirit and teachings of the present invention. All such combinations and / or combinations fall within the scope of the present invention.

[0175] The above describes embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. The scope of the present invention is defined by the appended claims and their equivalents. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which are intended to fall within the scope of the present invention.

Claims

1. A method for processing tactile perception data of an intelligent agent, characterized in that: include: Low-resolution tactile data and matching high-resolution tactile data are respectively acquired based on the same preset contact conditions; wherein the preset contact conditions include a preset mapping relationship based on the same material and size between a first contact module of the magnetic tactile sensor that detects the low-resolution tactile data and a second contact module of the visual tactile sensor that detects the high-resolution tactile data, and wherein the contact positions of the first contact module and the second contact module are aligned and have the same contact force; and wherein the high-resolution tactile data is depth image data of the detection data acquired by the visual tactile sensor; Training a preset generative model based on the low-resolution tactile data and the high-resolution tactile data to generate a target data processing model; and The target data processing model is used to process the target low-resolution tactile data to achieve super-resolution tactile perception.

2. The method according to claim 1, characterized in that Before training the preset generative model according to the low-resolution tactile data and the high-resolution tactile data to generate the target data processing model, the method further includes: generating prior distribution data, generative distribution data, and variational approximation data corresponding to the low-resolution tactile data and the high-resolution tactile data according to a preset variational autoencoding rule corresponding to a preset generative model; A loss function corresponding to the preset generative model is generated according to the prior distribution data, the generated distribution data and the variational approximation data.

3. The method according to claim 2, characterized in that The step of training a preset generative model according to the low-resolution tactile data and the high-resolution tactile data to generate a target data processing model includes: Performing encoding and integration processing of the low-resolution tactile data and the high-resolution tactile data through a preset generative model; Performing a decoding and reconstruction process on the received latent variables corresponding to the encoding and incorporation process by using a preset generative model; In the process of encoding integration processing and decoding reconstruction processing, the loss function is optimized until a target data processing model is generated.

4. The method according to claim 1, wherein Before training the preset generative model according to the low-resolution tactile data and the high-resolution tactile data to generate the target data processing model, the method further includes: generating noise tactile data in a forward diffusion process corresponding to the preset generative model by adding preset Gaussian noise to the high-resolution tactile data; performing denoising recovery on the noisy tactile data using the low-resolution tactile data to generate denoised tactile data in a reverse denoising process corresponding to the preset generative model; An objective function of the preset generative model is generated according to the noisy tactile data and the denoised tactile data.

5. The method according to claim 4, characterized in that The step of training a preset generative model according to the low-resolution tactile data and the high-resolution tactile data to generate a target data processing model includes: In the forward diffusion process and the reverse denoising process, the objective function is optimized until a target data processing model is generated.

6. A tactile perception data processing device for an intelligent agent, characterized in that: include: a data acquisition module configured to acquire low-resolution tactile data and matching high-resolution tactile data based on identical preset contact conditions; wherein the preset contact conditions include a preset mapping relationship between a first contact module of a magnetic tactile sensor that detects the low-resolution tactile data and a second contact module of a visual tactile sensor that detects the high-resolution tactile data, based on identical materials and dimensions, and wherein the midlines of the contact positions of the first contact module and the second contact module are aligned and the contact forces are identical; and wherein the high-resolution tactile data is depth image data of the detection data acquired by the visual tactile sensor; a model generation module, configured to train a preset generative model based on the low-resolution tactile data and the high-resolution tactile data to generate a target data processing model; and The data processing module is used to process the target low-resolution tactile data through the target data processing model to achieve super-resolution tactile perception.

7. An electronic device comprising: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to execute the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 5.

9. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Tactile glove array signal super-resolution reconstruction method based on diffusion model

    CN119579410A