Tactile-visual collaborative focusing method based on multi-modal deep learning

Through multimodal deep learning combined with visual and tactile data, the focal length of the photography equipment is dynamically adjusted, which solves the problem of insufficient focus accuracy in the existing technology, and achieves a high-precision focus effect, which is suitable for macro photography and industrial robot vision.

CN120434503APending Publication Date: 2025-08-05SHENZHEN APICAL TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510566174.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing automatic focus technology is difficult to achieve high-precision focus in macro photography and industrial robot vision, especially in complex and changing environments, and it is difficult to maintain stable focus performance.

Method used

The haptic-visual collaborative focus method based on multimodal deep learning is adopted to obtain visual data and tactile data through photography equipment, extract multimodal features, and focus prediction is used to dynamically adjust the focal length of the photography equipment.

Benefits of technology

It improves the accuracy of focus and scene adaptability, and is suitable for shooting needs in macro photography and industrial robot vision, significantly improving shooting effects and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120434503A_ABST
    Figure CN120434503A_ABST
Patent Text Reader

Abstract

The invention provides a touch-vision collaborative focusing method based on multi-modal deep learning. The method comprises the following steps: acquiring multi-modal data of a target object; the method comprises the following steps: shooting a target object by using a photographic device to obtain multi-modal data, the multi-modal data comprising visual data and tactile data; multi-modal features corresponding to the multi-modal data are extracted; wherein the multi-modal features comprise visual features, tactile features and distance change features; inputting the multi-modal feature into a pre-trained multi-modal deep learning model to obtain a focus prediction result of the target object; and adjusting the focal length of the photographic equipment according to the focus prediction result. Through the above method, the focusing accuracy and scene adaptability are improved, so as to meet the shooting requirements in macro photography and industrial robot vision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of photography technology, and in particular to a tactile-visual collaborative focusing method based on multimodal deep learning. Background Art

[0002] In modern industrial automation and intelligent manufacturing, robotic vision and industrial photography technologies have become core means for achieving high-precision inspection, identification, positioning, and manipulation. As industrial production demands ever-increasing product quality, production efficiency, and intelligence, robotic vision systems and industrial photography equipment must quickly and accurately capture clear images of target objects in complex and ever-changing scenarios. Autofocus technology, a critical component for ensuring image quality, directly impacts the reliability and stability of the entire vision system. Robots in robotic vision applications must complete tasks such as part assembly, quality inspection, and material handling, requiring the vision system to accurately capture the target object's features in real time. In industrial inspection, precise measurement and assessment of surface defects and dimensional accuracy of mechanical parts require high-resolution, high-contrast images. While existing autofocus technologies have met the application requirements of robotic vision and industrial photography to a certain extent, in certain scenarios, insufficient focusing accuracy makes it difficult to determine object distance, resulting in large focusing errors or prone to focus failures, making it difficult to maintain stable focusing performance in diverse environments. Summary of the Invention

[0003] The main technical problem solved by the present invention is to improve the focusing accuracy and scene adaptability to meet the shooting requirements of macro photography and industrial robot vision.

[0004] According to the first aspect, an embodiment provides a tactile-visual collaborative focusing method based on multimodal deep learning, comprising:

[0005] Acquiring multimodal data of a target object; wherein the target object is photographed using a photographic device to acquire the multimodal data, wherein the multimodal data includes visual data and tactile data;

[0006] Extracting multimodal features corresponding to the multimodal data; wherein the multimodal features include visual features, tactile features, and distance change features;

[0007] Inputting the multimodal features into a pre-trained multimodal deep learning model to obtain a focus prediction result of the target object;

[0008] The focal length of the photographic device is adjusted according to the focus prediction result.

[0009] In some embodiments, the pre-trained multimodal deep learning model is trained in the following manner:

[0010] Acquire a training data set; wherein the training data set includes training tactile data under different environmental conditions, training image sequences corresponding to the training tactile data, and training distance change data, wherein the environmental conditions include the distance between the photographic device and the target object, the texture of the target object, and the lighting conditions of the target object;

[0011] Extracting a training feature set corresponding to the training data set; wherein the training feature set includes training tactile features, training image features corresponding to the training tactile features, and training distance change features;

[0012] Establishing a mapping relationship among the training tactile feature, the training image feature corresponding to the training tactile feature, the training distance change feature, and a preset focus reference result through supervised learning;

[0013] The multimodal deep learning model to be trained is trained based on the mapping relationship to obtain a trained multimodal deep learning model.

[0014] In some embodiments, the training of the multimodal deep learning model to be trained based on the mapping relationship to obtain a trained multimodal deep learning model includes:

[0015] Inputting the training tactile features in the mapping relationship, the training image features corresponding to the training tactile features, and the training distance change features into a fusion network in the multimodal deep learning model to be trained to obtain a focus prediction result; wherein the fusion network includes a Transformer-based cross-modal attention mechanism;

[0016] Evaluating the focus prediction result based on a preset focus reference position and a preset evaluation index in the mapping relationship to obtain an evaluation result;

[0017] The model parameters of the multimodal deep learning model to be trained are optimized according to the evaluation results until the multimodal deep learning model to be trained reaches convergence or a preset number of training rounds, thereby obtaining a trained multimodal deep learning model.

[0018] In some embodiments, inputting the training tactile features in the mapping relationship, the training image features corresponding to the training tactile features, and the training distance change features into a fusion network in the multimodal deep learning model to be trained to obtain a focus prediction result includes:

[0019] assigning corresponding weights to the training tactile feature, the training image feature corresponding to the training tactile feature, and the training distance change feature respectively;

[0020] performing a weighted summation of the training tactile features, the training image features corresponding to the training tactile features, and the training distance change features based on the fusion network and the assigned weights to obtain a comprehensive feature representation;

[0021] The comprehensive feature representation is decoded into a corresponding focus prediction result.

[0022] In some embodiments, after acquiring multimodal data of the target object, the tactile-visual collaborative focusing method further includes:

[0023] Determining a scene type of a scene in which the target object is located based on the visual data and the tactile data; wherein the scene type includes a macro scene, a dynamic contact scene, and a complex lighting scene;

[0024] Before inputting the multimodal features into a pre-trained multimodal deep learning model, the tactile-visual collaborative focusing method further includes:

[0025] Adjust the model parameters of the pre-trained multimodal deep learning model according to the scenario type.

[0026] In some embodiments, the photographic device includes a tactile sensor and a camera, and the visual data of the target object is obtained by using the visual sensor, and the tactile data of the target object is obtained by using the tactile sensor; the visual features include the image clarity, edge features and position of the target object of the visual image, and the tactile features include the object distance between the target object and the tactile sensor, the pressure generated when the tactile sensor contacts the target object, and the texture information of the target object, and the distance change feature is determined by fusing the depth estimation information in the visual feature and the pressure information in the tactile data.

[0027] In some embodiments, the photographic device includes a single camera or multiple cameras; and acquiring multimodal data of the target object includes:

[0028] Using a photographic device including a single camera to photograph the target object to obtain the multimodal data; or;

[0029] A photographic device comprising multiple cameras is used to obtain multimodal data of the target object at different viewing angles.

[0030] According to the second aspect, an embodiment provides a tactile-visual collaborative focusing device based on multimodal deep learning, comprising:

[0031] A data acquisition module, configured to acquire multimodal data of a target object; wherein the target object is photographed using a photographic device to acquire the multimodal data, the multimodal data including visual data and tactile data;

[0032] A feature extraction module, configured to extract multimodal features corresponding to the multimodal data; wherein the multimodal features include visual features, tactile features, and distance change features;

[0033] A focus prediction module, configured to input the multimodal features into a pre-trained multimodal deep learning model to obtain a focus prediction result of the target object;

[0034] A focal length adjustment module is used to adjust the focal length of the photographic device according to the focus prediction result.

[0035] According to a third aspect, an embodiment provides a tactile-visual collaborative focusing device based on multimodal deep learning, comprising:

[0036] Memory, used to store programs;

[0037] A processor is configured to implement the tactile-visual collaborative focusing method by executing the program stored in the memory.

[0038] According to a fourth aspect, an embodiment provides a computer program product, comprising a computer program and / or instructions, which implement the tactile-visual collaborative focusing method when executed by a processor.

[0039] According to the multimodal deep learning-based tactile-visual collaborative focusing method, apparatus, device, and computer program product of the above-described embodiments, by utilizing a pre-trained multimodal deep learning model to fuse the visual and tactile data of a target object, the multimodal deep learning model predicts the focus of the target object and adjusts the focal length of the photographic device based on this focus prediction, thereby improving focusing accuracy. Compared to using visual data alone, combining the visual and tactile data of the target object to predict focus results is suitable for macro photography and industrial robot vision, and improves the scene applicability of focusing. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 This is a flowchart of a tactile-visual collaborative focusing method based on multimodal deep learning according to an embodiment of the present application;

[0041] Figure 2 A flowchart of a method for training a pre-trained multimodal deep learning model according to an embodiment;

[0042] Figure 3 A flowchart of an embodiment of training a multimodal deep learning model to be trained based on a mapping relationship to obtain a trained multimodal deep learning model;

[0043] Figure 4 A flowchart of an embodiment of inputting training tactile features in a mapping relationship, training image features corresponding to the training tactile features, and training distance change features into a fusion network in a multimodal deep learning model to be trained to obtain a focus prediction result;

[0044] Figure 5 Schematic diagram of the structure of a tactile-visual collaborative focusing device based on multimodal deep learning in one embodiment. DETAILED DESCRIPTION

[0045] The present invention will be further described in detail below by means of specific embodiments in conjunction with the accompanying drawings. Similar elements in different embodiments are numbered with associated similar elements. In the following embodiments, many detailed descriptions are provided to enable the present application to be better understood. However, those skilled in the art will readily appreciate that some of the features may be omitted in different circumstances, or may be replaced by other elements, materials, or methods. In some cases, some operations related to the present application are not shown or described in the specification. This is to avoid the core portion of the present application being overwhelmed by excessive descriptions, and for those skilled in the art, it is not necessary to describe these related operations in detail. They will fully understand the related operations based on the description in the specification and the general technical knowledge in the art.

[0046] In addition, the features, operations, or characteristics described in the specification may be combined in any appropriate manner to form various embodiments. Furthermore, the steps or actions in the method description may be reordered or adjusted in a manner readily apparent to those skilled in the art. Therefore, the various sequences in the specification and drawings are provided solely for the purpose of clearly describing a particular embodiment and are not intended to be mandatory, unless otherwise specified.

[0047] The serial numbers assigned to components herein, such as "first," "second," etc., are used solely to distinguish the objects being described and do not convey any sequential or technical meaning. References to "connection" and "coupling" herein, unless otherwise specified, include both direct and indirect connections (couplings).

[0048] Current mainstream autofocus technologies primarily adjust focus based on single visual data (such as image clarity and contrast). However, when faced with macro photography, complex contact scenarios, and the need for multi-dimensional information fusion, traditional focusing methods struggle to accurately determine object distance, resulting in large focusing errors. This is particularly true in robotic vision or industrial photography, where traditional methods lack multi-dimensional information support. For example, in industrial environments, the position and posture of target objects can change at any time, and lighting conditions can be unstable, making it difficult to achieve high-precision focus control based solely on visual data.

[0049] In order to improve the accuracy and scene adaptability of focusing and to be suitable for the shooting requirements in macro photography and industrial robot vision, an embodiment of the present application provides a tactile-visual collaborative focusing method based on multimodal deep learning. In this visual-tactile collaborative focusing method, multimodal data of the target object is obtained; wherein, the target object is photographed by a photographic device to obtain multimodal data, and the multimodal data includes visual data and tactile data; multimodal features corresponding to the multimodal data are extracted; wherein the multimodal features include visual features, tactile features and distance change features; the multimodal features are input into a pre-trained multimodal deep learning model to obtain the focus prediction result of the target object; and the focal length of the photographic device is adjusted according to the focus prediction result.

[0050] The following describes the tactile-visual collaborative focusing method based on multimodal deep learning provided by the embodiments of the present application in conjunction with the accompanying drawings.

[0051] Figure 1 A flowchart of a tactile-visual collaborative focusing method based on multimodal deep learning provided by an embodiment of the present application is shown, and is described in detail below:

[0052] Step S10: Acquire multimodal data of the target object.

[0053] Specifically, a target object is photographed using a photographic device to obtain multimodal data, wherein the multimodal data includes visual data and tactile data. The photographic device includes a tactile sensor and a camera, and the visual data of the target object is obtained using the visual sensor, and the tactile data of the target object is obtained using the tactile sensor.

[0054] For example, in macro photography of flowers, when capturing details of flowers, tactile sensors sense the distance between petals, while cameras capture visual data of the flowers. In industrial robot parts inspection scenarios, when inspecting small parts, tactile sensors detect tactile data from the parts to determine their position, while cameras use visual data to optimize focus, ensuring clear edges and textures. In dynamic contact robotic photography, during robotic grasping tasks, tactile and visual sensors detect tactile data and visual data of the target object, collaboratively analyzing the distance and shape of the target object.

[0055] In the embodiments of the present application, compared with the mainstream automatic focusing technology that mainly adjusts the focus based on a single visual data (image clarity or contrast), the visual data and tactile data of the target object are obtained, so that when facing macro, complex contact scenes and multi-dimensional information fusion needs, a rich data basis can be provided and the problem of low focusing accuracy can be solved based on this data basis.

[0056] Step S20: extracting multimodal features corresponding to the multimodal data.

[0057] Specifically, multimodal features corresponding to the multimodal data are extracted, where multimodal features include visual features, tactile features, and distance change features. Visual feature extraction primarily relies on computer vision technology, processing image or video data to obtain representative features, while tactile feature extraction primarily relies on tactile sensor technology, processing tactile data detected by tactile sensor technology to obtain information related to the object's surface characteristics. Distance change features are extracted by combining tactile and visual data to analyze the dynamic distance between the target object and the camera of the photographic device.

[0058] For example, visual feature extraction can involve extracting color features from visual data, reflecting the overall color features of the visual data by statistically analyzing the distribution of different colors in the visual data. Alternatively, it can involve extracting shape features from visual data, extracting shape features by detecting edges in the visual data. Tactile feature extraction can involve extracting pressure distribution features, using a tactile sensor to obtain pressure values at each point on a physical surface, forming a pressure matrix that reflects the pressure distribution. Alternatively, it can involve detecting vibration signals generated by contact with an object's surface through a tactile sensor, analyzing the vibration frequency and amplitude, and extracting texture and roughness features. Distance change features are determined by fusing pressure information from tactile data with depth estimation information from visual data, where the fusion process is achieved through a multimodal feature extraction algorithm.

[0059] In the embodiments of the present application, visual data and tactile data usually contain a large amount of redundant information. By extracting features, key information can be retained and the computational complexity can be reduced. At the same time, in the robot grasping task, after extracting visual features (object contour features) and tactile features (hardness features), the grasping strategy can be quickly matched to reduce the real-time calculation time. Therefore, by extracting multimodal features corresponding to multimodal data, redundant information is reduced, the algorithm operation is accelerated, and at the same time, robustness and anti-interference ability can be enhanced. The distance change feature can be coordinated with tactile features and visual features to achieve more sophisticated physical interaction.

[0060] Step S30: Input the multimodal features into a pre-trained multimodal deep learning model to obtain a focus prediction result of the target object.

[0061] Specifically, a multimodal deep learning model aims to process and fuse data from multiple modalities (e.g., vision and touch) to achieve more comprehensive perception and understanding. The multimodal deep learning model can be a multimodal model based on the Transformer architecture.

[0062] For example, the multimodal deep learning model is a multimodal pre-training model (Contrastive Language–Image Pre-training, CLIP), which maps images and texts into the same feature space through contrastive learning to achieve image-text matching.

[0063] In this embodiment, multimodal features are input into a pre-trained multimodal deep learning model. Visual features provide spatial information, while tactile features supplement material details, reducing the errors introduced by a single modality to improve the accuracy of focus prediction results for the target object. Furthermore, visual features are sensitive to changes in illumination, while tactile features are tolerant to occlusion or stains. Combining these two improves system stability and robustness. The pre-trained multimodal deep learning model can process and fuse multimodal features to produce accurate focus prediction results.

[0064] Step S40: adjusting the focal length of the photographic device according to the focus prediction result.

[0065] Specifically, the focus prediction result can be a focus position or a focus parameter. The focus position refers to the coordinates or area range of the area in the picture that needs to be clearly imaged, while the focus parameter is a parameter used to describe the imaging characteristics of the focus area and is used to optimize the focusing effect.

[0066] In the embodiments of this application, dynamically adjusting the focal length of a photographic device based on focus prediction results significantly improves shooting quality and efficiency, enabling precise focusing. Real-time focus adjustment based on focus prediction results eliminates the need for manual adjustment, making it suitable for fast-moving scenes. Furthermore, in the specialized context of macro photography, the camera automatically switches to a macro focus distance when a tiny subject is predicted, eliminating the need for manual switching.

[0067] In some embodiments, please refer to Figure 2 The pre-trained multimodal deep learning model is trained by following steps S31 to S34:

[0068] Step S31: Obtain a training data set.

[0069] Specifically, the training dataset includes training tactile data under different environmental conditions, training image sequences corresponding to the training tactile data, and training distance change data. Environmental conditions include the distance between the camera and the target object, the texture of the target object, and the lighting conditions of the target object.

[0070] For example, the distance between the camera and the target object can be very close (1-10 cm) for macro photography of insects or plant details, moderate (1-3 meters) for standard photography of portraits or everyday scenery, or long (10-100 meters) for long-range photography of landscapes or wildlife. The target object can be a high-texture object with rich surface details, such as fabric, wood grain, or brick walls, a low-texture object with a smooth surface and less detail, such as glass, metal, or plastic, or a mixed-texture object with a surface containing multiple textures, such as the casing of an electronic device. The lighting conditions of the target object can be strong light environments, such as direct sunlight or high-intensity light sources, weak light environments, such as low-light or low-light scenes, or uniform lighting environments, such as scenes with evenly distributed optical fibers and no obvious shadows.

[0071] In the embodiment of the present application, by obtaining a diverse training data set to train the model, the generalization ability and robustness of the model can be improved.

[0072] Step S32: extracting a training feature set corresponding to the training data set.

[0073] Specifically, the training feature set includes training tactile features, training image features corresponding to the training tactile features, and training distance change features.

[0074] Step S33: establishing a mapping relationship among the training tactile features, the training image features corresponding to the training tactile features, the training distance change features, and the preset focus reference results through supervised learning.

[0075] Specifically, a mapping relationship is established between the training tactile features, the training image features corresponding to the training tactile features, and the training distance change features, and the corresponding focus reference results are marked, so as to establish a mapping relationship between the training tactile features, the training image features corresponding to the training tactile features, the training distance change features, and the preset focus reference results.

[0076] In this embodiment, a mapping relationship between training tactile features, training image features corresponding to the training tactile features, training distance change features, and preset focus reference results is established through supervised learning, which can integrate multiple modal information, enhance robustness, achieve feature complementarity and knowledge transfer, and reduce training costs.

[0077] Step S34: Train the multimodal deep learning model to be trained based on the mapping relationship to obtain a trained multimodal deep learning model.

[0078] Specifically, the training tactile features in the mapping relationship, the training image features corresponding to the training tactile features, and the training distance change features are used as the input of the multimodal deep learning model to be trained, and the preset focus reference results in the mapping relationship are used as the output labels. The cross entropy loss function or the mean square error loss function is used to optimize the model to obtain a trained multimodal deep learning model.

[0079] In the embodiments of the present application, a multimodal deep learning model to be trained is trained based on the mapping relationship, so that the trained multimodal deep learning model can cope with complex environments and enhance its environmental adaptability. The multimodal deep learning model is optimized through the training data set, covering macro objects, industrial parts and dynamic contact scenarios.

[0080] In some embodiments, please refer to Figure 3 , step S34: train the multimodal deep learning model to be trained based on the mapping relationship to obtain a trained multimodal deep learning model, including steps S341 to S343, which are described in detail below.

[0081] Step S341: input the training tactile features in the mapping relationship, the training image features corresponding to the training tactile features, and the training distance change features into the fusion network in the multimodal deep learning model to be trained to obtain the focus prediction result.

[0082] Specifically, the fusion network includes a Transformer-based cross-modal attention mechanism.

[0083] In an embodiment of the present application, the focus prediction result is obtained by adopting a fusion network based on the Transformer cross-modal attention mechanism, which can achieve alignment and fusion of cross-modal features.

[0084] Step S342: Evaluate the focus prediction result based on the preset focus reference position and the preset evaluation index in the mapping relationship to obtain an evaluation result.

[0085] Specifically, the preset evaluation indicators include focus prediction error and focusing success rate. Among them, the focus prediction error includes but is not limited to two-dimensional focus prediction error, three-dimensional focus prediction error and normalized focus error. Two-dimensional focus prediction error refers to the deviation in two-dimensional space between the predicted focus position in the focus prediction result and the real focus position in the preset focus reference position, and the Euclidean distance between the two can be calculated. Three-dimensional focus prediction error refers to the deviation in three-dimensional space between the predicted focus position in the focus prediction result and the real focus position in the preset focus reference position. Normalized focus error refers to scaling the error between the predicted focus position in the focus prediction result and the real focus position in the preset focus reference position to between 0 and 1 or a specific range to eliminate the dimensional effect. The focusing success rate is used to evaluate the proportion of successfully predicted real focus positions in the focus prediction results.

[0086] In an embodiment of the present application, the evaluation result is used to measure the error between the focus reference position and the focus prediction position to facilitate subsequent optimization and training of the model.

[0087] Step S343: Optimize the model parameters of the multimodal deep learning model to be trained according to the evaluation results until the multimodal deep learning model to be trained reaches convergence or a preset number of training rounds to obtain a trained multimodal deep learning model.

[0088] In the examples of this application, please refer to Figure 4 , step S341: input the training tactile features in the mapping relationship, the training image features corresponding to the training tactile features, and the training distance change features into the fusion network in the multimodal deep learning model to be trained to obtain the focus prediction results, including steps S341a to S341c, which are described in detail below.

[0089] Step S341a: assign corresponding weights to the training tactile features, the training image features corresponding to the training tactile features, and the training distance change features respectively.

[0090] Specifically, the assigned weights can be dynamically adjusted, and the weights of the training tactile features, the training image features corresponding to the training tactile features, and the training distance change features can be adjusted according to scene changes.

[0091] For example, in low-light conditions, the weight of trained tactile features is automatically increased to reduce reliance on training image features and achieve adaptive decision-making.

[0092] Step S341b: performing weighted summation on the training tactile features, the training image features corresponding to the training tactile features, and the training distance change features based on the fusion network and the assigned weights to obtain a comprehensive feature representation.

[0093] Step S341c: Decode the comprehensive feature representation into the corresponding focus prediction result.

[0094] Specifically, the comprehensive feature representation is decoded and mapped into the corresponding focus prediction result.

[0095] In some embodiments, the scene type of the target object is determined based on visual and tactile data. Deep learning techniques are used to analyze scene characteristics (e.g., the target object's type and contact state) in the visual and tactile data to determine the scene type. Focus strategies are then dynamically optimized based on the scene type. Scene types include macro scenes, dynamic contact scenes, and complex lighting scenes.

[0096] For example, in macro scenarios, the object surface is determined based on tactile perception, and in industrial scenarios, the focus depth is adjusted according to the part texture and distance.

[0097] In the embodiment of the present application, the scene type obtained through scene analysis provides multi-dimensional support for subsequent focus prediction, thereby improving the focusing effect.

[0098] In some embodiments, before inputting the multimodal features into a pre-trained multimodal deep learning model, the tactile-visual collaborative focusing method further includes:

[0099] Adjust the model parameters of the pre-trained multimodal deep learning model according to the scenario type.

[0100] For example, in macro scenes, the focus is prioritized on the tactile surface of the object, and in industrial scenes, the focus depth is adjusted according to the part texture and distance.

[0101] In an embodiment of the present application, the model parameters are adaptively adjusted based on the scene type, so that the focus prediction results output by the multimodal deep learning model after the adjustment of the model parameters are more scene-adaptive, and are particularly suitable for shooting requirements in macro photography and industrial robot vision.

[0102] In some embodiments, the focus can be optimized according to the shooting task type, for example, according to preset parameters (macro priority, industrial priority), or the focus area can be adjusted according to the contact pressure or distance change of the detected target object. The position offset of the target object can also be judged in combination with the image clarity for focusing.

[0103] In some embodiments, the photographic device includes a tactile sensor and a camera, and uses a visual sensor to obtain visual data of the target object, and uses a tactile sensor to obtain tactile data of the target object. Multimodal data of the target object is obtained, and multimodal features corresponding to the multimodal data are extracted, wherein the multimodal features include visual features, tactile features, and distance change features. The visual features include image clarity, edge features, and the position of the target object in the visual image. The tactile features include the object distance between the target object and the tactile sensor, the pressure generated when the tactile sensor contacts the target object, and the texture information of the target object. The distance change feature is determined by fusing the depth estimation information in the visual features with the pressure information in the tactile data.

[0104] Specifically, the depth estimation information in the visual feature can be the depth change rate. The camera obtains continuous depth estimates of the target object and calculates the depth change rate. The pressure information in the tactile data can be the tactile distance change. The contact pressure is obtained through the tactile sensor and combined with a physical model (such as the elastic deformation equation) to infer the contact distance change. The visual depth change rate and the tactile distance change are weighted and fused to generate the final distance change feature.

[0105] In the embodiment of the present application, visual features, tactile features and distance change features work together to monitor the state changes of the target object in real time and respond to environmental changes more accurately and efficiently.

[0106] In some embodiments, the photographic device may include a single camera or multiple cameras. When using multiple cameras to photograph a target object, visual and tactile data of the target object at different viewing angles can be collected synchronously, or visual and tactile data of the target object at different viewing angles can be collected asynchronously to obtain more comprehensive visual data. Furthermore, it is worth noting that when using multiple cameras to photograph a target object, the cameras can be divided into a primary camera and a secondary camera. The primary camera can be focused in combination with tactile data, while the secondary camera can provide perspective supplementation and distance correction. This is used in industrial robot parts inspection scenarios to capture high-definition images of tiny parts.

[0107] In some embodiments, the conditions for acquiring multimodal data of a target object include: tactile sensor sensitivity, image resolution, and the ambient lighting of the target object. All three of these conditions must meet a preset minimum threshold to ensure effective focusing. The minimum threshold is set based on the specific application scenario and actual conditions.

[0108] In the embodiments of the present application, when the tactile sensor's sensitivity meets a minimum threshold, it can perceive subtle distance changes. When the image resolution meets a minimum threshold, the accuracy of visual feature extraction can be ensured. When the ambient lighting is maintained within a set range, visual data distortion can be avoided. When these three conditions meet the minimum threshold, focusing accuracy can be guaranteed.

[0109] In some embodiments, the tactile-visual collaborative focusing method based on multimodal deep learning provided in the embodiments of the present application is suitable for robotic photography equipment, macro cameras and industrial vision systems. It can be integrated by embedding the multimodal deep learning model into the device chip and combining it with the real-time calculation of the tactile sensor and the camera, thereby improving the focusing accuracy by more than 50%, significantly improving the shooting quality in macro and industrial scenarios.

[0110] In some embodiments, the tactile-visual collaborative focusing method based on multimodal deep learning provided in the embodiments of the present application can be applied in the following scenarios. Scenario 1: Flower photography for macro photography. When shooting flower details, the distance of the petals is perceived by the tactile sensor, and the focus is adjusted to the stamen area in combination with the visual data to generate a high-definition macro image. Scenario 2: Parts detection for industrial robots. When detecting small parts, tactile data is used to determine the position of the parts, and visual data is used to optimize the focus to ensure that the edges and textures are clearly visible. Scenario 3: Robot photography for dynamic contact scenes. In robot grasping tasks, the distance and shape of the target object are analyzed through tactile-visual collaboration, and the focus is dynamically adjusted to the contact point to provide real-time shooting support.

[0111] Corresponding to the tactile-visual collaborative focusing method based on multimodal deep learning in the above embodiment, Figure 5 A structural schematic diagram of a tactile-visual collaborative focusing device based on multimodal deep learning provided in an embodiment of the present application is shown. For the sake of convenience, only the parts related to the embodiment of the present application are shown.

[0112] Please refer to Figure 5 The tactile-visual collaborative focusing device based on multimodal deep learning includes a data acquisition module 10, a feature extraction module 20, a focus prediction module 30 and a focal length adjustment module 40, which are described in detail below.

[0113] The data acquisition module 10 is used to acquire multimodal data of the target object. The target object is photographed by a photographic device to acquire the multimodal data, and the multimodal data includes visual data and tactile data.

[0114] The feature extraction module 20 is used to extract multimodal features corresponding to the multimodal data; wherein the multimodal features include visual features, tactile features and distance change features.

[0115] The focus prediction module 30 is used to input the multimodal features into a pre-trained multimodal deep learning model to obtain the focus prediction result of the target object.

[0116] The focal length adjustment module 40 is used to adjust the focal length of the photographic device according to the focus prediction result.

[0117] In some embodiments, the focus prediction module 30 is further used to train a pre-trained multimodal deep learning model:

[0118] Obtain a training data set; wherein the training data set includes training tactile data under different environmental conditions, training image sequences corresponding to the training tactile data, and training distance change data. The environmental conditions include the distance conditions between the photographic device and the target object, the texture conditions of the target object, and the lighting conditions of the target object.

[0119] A training feature set corresponding to the training data set is extracted; wherein the training feature set includes training tactile features, training image features corresponding to the training tactile features, and training distance change features.

[0120] Through supervised learning, a mapping relationship is established among the training tactile features, the training image features corresponding to the training tactile features, the training distance change features and the preset focus reference results.

[0121] The multimodal deep learning model to be trained is trained based on the mapping relationship to obtain a trained multimodal deep learning model.

[0122] In some embodiments, the focus prediction module 30 is further configured to train the multimodal deep learning model to be trained based on the mapping relationship to obtain a trained multimodal deep learning model, including:

[0123] The training tactile features in the mapping relationship, the training image features corresponding to the training tactile features, and the training distance change features are input into the fusion network in the multimodal deep learning model to be trained to obtain the focus prediction results; wherein, the fusion network includes a cross-modal attention mechanism based on Transformer.

[0124] The focus prediction result is evaluated based on the preset focus reference position and the preset evaluation index in the mapping relationship to obtain an evaluation result.

[0125] The model parameters of the multimodal deep learning model to be trained are optimized according to the evaluation results until the multimodal deep learning model to be trained reaches convergence or a preset number of training rounds to obtain a trained multimodal deep learning model.

[0126] In some embodiments, the focus prediction module 30 is further configured to input the training tactile features in the mapping relationship, the training image features corresponding to the training tactile features, and the training distance change features into a fusion network in the multimodal deep learning model to be trained, to obtain a focus prediction result, including:

[0127] Corresponding weights are assigned to the training tactile features, the training image features corresponding to the training tactile features, and the training distance change features.

[0128] Based on the fusion network and the assigned weights, the training tactile features, the training image features corresponding to the training tactile features, and the training distance change features are weightedly summed to obtain a comprehensive feature representation.

[0129] Decode the comprehensive feature representation into the corresponding focus prediction result.

[0130] In some embodiments, after acquiring the multimodal data of the target object, the data acquisition module 10 is further configured to:

[0131] The scene type of the scene in which the target object is located is determined based on the visual data and the tactile data; wherein the scene types include macro scenes, dynamic contact scenes and complex lighting scenes.

[0132] In some embodiments, before inputting the multimodal features into a pre-trained multimodal deep learning model, the focus prediction module 30 is further configured to:

[0133] Adjust the model parameters of the pre-trained multimodal deep learning model according to the scenario type.

[0134] In some embodiments, the photographic device includes a tactile sensor and a camera, and uses a visual sensor to obtain visual data of the target object, and uses a tactile sensor to obtain tactile data of the target object; the visual features include image clarity, edge features and the position of the target object of the visual image, and the tactile features include the object distance between the target object and the tactile sensor, the pressure generated when the tactile sensor contacts the target object, and the texture information of the target object. The distance change feature is determined by fusing the depth estimation information in the visual features and the pressure information in the tactile data.

[0135] In some embodiments, the photographic device includes a single camera or multiple cameras; the data acquisition module 10 is further configured to acquire multimodal data of the target object:

[0136] photographing a target object using a photographic device including a single camera to obtain multimodal data; or

[0137] A photographic device comprising multiple cameras is used to obtain multimodal data of a target object at different viewing angles.

[0138] In some embodiments, the present application further provides a tactile-visual collaborative focusing device based on multimodal deep learning, comprising:

[0139] Memory, used to store programs;

[0140] The processor is configured to implement the tactile-visual collaborative focusing method by executing a program stored in the memory.

[0141] In some embodiments, the present application also provides a computer program product, including a computer program and / or instructions, which implements a tactile-visual collaborative focusing method when the computer program and / or instructions are executed by a processor.

[0142] Those skilled in the art will appreciate that all or part of the functions of the various methods in the above embodiments can be implemented by hardware or by computer program. When all or part of the functions in the above embodiments are implemented by computer program, the program can be stored in a computer-readable storage medium, and the storage medium can include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to implement the above functions. For example, the program is stored in the memory of the device, and when the program in the memory is executed by the processor, all or part of the above functions can be implemented. In addition, when all or part of the functions in the above embodiments are implemented by computer program, the program can also be stored in a storage medium such as a server, another computer, disk, optical disk, flash disk or mobile hard disk, and saved in the memory of the local device by downloading or copying, or the system of the local device is updated. When the program in the memory is executed by the processor, all or part of the functions in the above embodiments can be implemented.

[0143] The above examples are used to illustrate the present invention, which are only used to help understand the present invention and are not intended to limit the present invention. Those skilled in the art can make several simple deductions, modifications or substitutions based on the concept of the present invention.

Claims

1. A tactile-visual collaborative focusing method based on multimodal deep learning, characterized in that: include: Acquiring multimodal data of a target object; wherein the target object is photographed using a photographic device to acquire the multimodal data, wherein the multimodal data includes visual data and tactile data; Extracting multimodal features corresponding to the multimodal data; wherein the multimodal features include visual features, tactile features, and distance change features; Inputting the multimodal features into a pre-trained multimodal deep learning model to obtain a focus prediction result of the target object; The focal length of the photographic device is adjusted according to the focus prediction result.

2. The tactile-visual collaborative focusing method according to claim 1, wherein: The pre-trained multimodal deep learning model is trained in the following way: Acquire a training data set; wherein the training data set includes training tactile data under different environmental conditions, training image sequences corresponding to the training tactile data, and training distance change data, wherein the environmental conditions include the distance between the photographic device and the target object, the texture of the target object, and the lighting conditions of the target object; Extracting a training feature set corresponding to the training data set; wherein the training feature set includes training tactile features, training image features corresponding to the training tactile features, and training distance change features; Establishing a mapping relationship among the training tactile feature, the training image feature corresponding to the training tactile feature, the training distance change feature, and a preset focus reference result through supervised learning; The multimodal deep learning model to be trained is trained based on the mapping relationship to obtain a trained multimodal deep learning model.

3. The tactile-visual collaborative focusing method according to claim 2, wherein: The method of training the multimodal deep learning model to be trained based on the mapping relationship to obtain a trained multimodal deep learning model includes: Inputting the training tactile features in the mapping relationship, the training image features corresponding to the training tactile features, and the training distance change features into a fusion network in the multimodal deep learning model to be trained to obtain a focus prediction result; wherein the fusion network includes a Transformer-based cross-modal attention mechanism; Evaluating the focus prediction result based on a preset focus reference position and a preset evaluation index in the mapping relationship to obtain an evaluation result; The model parameters of the multimodal deep learning model to be trained are optimized according to the evaluation results until the multimodal deep learning model to be trained reaches convergence or a preset number of training rounds, thereby obtaining a trained multimodal deep learning model.

4. The tactile-visual collaborative focusing method according to claim 3, wherein: The step of inputting the training tactile features in the mapping relationship, the training image features corresponding to the training tactile features, and the training distance change features into a fusion network in the multimodal deep learning model to be trained to obtain a focus prediction result includes: assigning corresponding weights to the training tactile feature, the training image feature corresponding to the training tactile feature, and the training distance change feature respectively; performing a weighted summation of the training tactile features, the training image features corresponding to the training tactile features, and the training distance change features based on the fusion network and the assigned weights to obtain a comprehensive feature representation; The comprehensive feature representation is decoded into a corresponding focus prediction result.

5. The tactile-visual collaborative focusing method according to claim 1, wherein: After acquiring the multimodal data of the target object, the tactile-visual collaborative focusing method further includes: Determining a scene type of a scene in which the target object is located based on the visual data and the tactile data; wherein the scene type includes a macro scene, a dynamic contact scene, and a complex lighting scene; Before inputting the multimodal features into a pre-trained multimodal deep learning model, the tactile-visual collaborative focusing method further includes: Adjust the model parameters of the pre-trained multimodal deep learning model according to the scenario type.

6. The tactile-visual collaborative focusing method according to claim 1, wherein: The photographic device includes a tactile sensor and a camera, and uses the visual sensor to obtain visual data of the target object, and uses the tactile sensor to obtain tactile data of the target object; the visual features include image clarity, edge features and the position of the target object of the visual image, and the tactile features include the object distance between the target object and the tactile sensor, the pressure generated when the tactile sensor contacts the target object, and the texture information of the target object. The distance change feature is determined by fusing the depth estimation information in the visual feature and the pressure information in the tactile data.

7. The tactile-visual collaborative focusing method according to claim 1, wherein: The photographic device includes a single camera or multiple cameras; the acquiring of multimodal data of the target object includes: Using a photographic device including a single camera to photograph the target object to obtain the multimodal data; or; A photographic device comprising multiple cameras is used to obtain multimodal data of the target object at different viewing angles.

8. A tactile-visual collaborative focusing device based on multimodal deep learning, characterized in that: include: A data acquisition module, configured to acquire multimodal data of a target object; wherein the target object is photographed using a photographic device to acquire the multimodal data, the multimodal data including visual data and tactile data; A feature extraction module, configured to extract multimodal features corresponding to the multimodal data; wherein the multimodal features include visual features, tactile features, and distance change features; A focus prediction module, configured to input the multimodal features into a pre-trained multimodal deep learning model to obtain a focus prediction result of the target object; A focal length adjustment module is used to adjust the focal length of the photographic device according to the focus prediction result.

9. A tactile-visual collaborative focusing device based on multimodal deep learning, characterized in that: include: Memory, used to store programs; A processor, configured to implement the tactile-visual collaborative focusing method according to any one of claims 1 to 7 by executing the program stored in the memory.

10. A computer program product comprising a computer program and / or instructions, characterized in that When the computer program and / or instructions are executed by a processor, the tactile-visual collaborative focusing method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Focusing control method and related equipment

    CN121037681A