Humanoid robot multi-modal data processing method, system and equipment and medium
By collecting multimodal data and processing methods, the robot's data processing capability is improved, and the task processing capability in industrial tasks is adapted. The robot's data processing capability is improved, and the task processing capability in industrial tasks is adapted. The robot's data processing capability is improved, and the task processing capability in industrial tasks is adapted. The robot's data processing capability is improved, and the task processing capability in industrial tasks is adapted. The robot's data processing capability is improved, and the robot's data processing efficiency is improved.
Patent Information
- Application Number
- CN202511051131.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-29
AI Technical Summary
In existing technologies, humanoid robots lack generalization capabilities and execution accuracy in industrial tasks and lack optimization for the specificity of industrial scenarios, resulting in insufficient visual recognition accuracy and weak tactile-visual cross-modal fusion capabilities, making them difficult to adapt to precision operations.
By collecting multimodal data sets, preprocessing and generating multi-layer prompt information, and using pre-trained multimodal models for collaborative reasoning processing, data processing capabilities and task execution accuracy can be improved.
It realizes the effective fusion of multimodal information of humanoid robots, such as vision and touch, to adapt to the precision operation requirements in industrial scenarios, improves the robot's data processing capabilities, adapts to the task processing capabilities in industrial scenarios, adapts to the accuracy of industrial tasks, and improves the accuracy of the robot's execution of industrial tasks.
Smart Images

Figure CN120687743A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent robot technology, and in particular to a method, system, device and medium for multimodal data processing of a humanoid robot. Background Art
[0002] In industrial manufacturing, humanoid robots must process multimodal information, including vision, touch, and speech, to complete complex tasks like assembly and quality inspection. While existing technologies handle general data processing, these methods are not optimized for the specific needs of industrial scenarios, resulting in insufficient generalization and precision for humanoid robots in industrial tasks.
[0003] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to propose a multimodal data processing method, system, device and medium for a humanoid robot, which can improve the data processing efficiency of the humanoid robot.
[0005] To achieve the above objectives, one aspect of an embodiment of the present application provides a method for processing multimodal data of a humanoid robot, the method comprising:
[0006] A multimodal dataset is obtained through humanoid robot collection;
[0007] Preprocessing the multimodal dataset to obtain a preprocessed dataset;
[0008] generating multi-layer prompt information according to the industrial scenario of the humanoid robot;
[0009] The preprocessed data set and the multi-layer prompt information are input into a pre-trained multimodal model for collaborative reasoning processing, and an industrial task processing result is output.
[0010] In some embodiments, the multimodal dataset includes visual images, tactile signals, and operation instruction text, and preprocessing the multimodal dataset to obtain a preprocessed dataset includes:
[0011] Performing contour feature extraction processing on the visual image to obtain contour features;
[0012] Performing defect detection and vector generation processing on the visual image according to the contour features to obtain visual features;
[0013] Performing time-frequency domain transformation and frequency band feature extraction processing on the tactile signal to obtain tactile features;
[0014] Performing semantic enhancement processing on the operation instruction text according to a preset industrial field vocabulary to obtain a semantically enhanced text;
[0015] Performing semantic vector conversion and constraint feature addition processing on the semantically enhanced text to obtain text features;
[0016] The preprocessing data set is obtained according to the visual features, the tactile features and the text features.
[0017] In some embodiments, generating multi-layer prompt information according to the industrial scenario of the humanoid robot includes:
[0018] Performing prompt generation processing on the role attributes of the humanoid robot according to the industrial scenario to obtain basic layer prompt information;
[0019] Performing prompt generation processing on the industrial task of the humanoid robot according to the industrial scenario to obtain task-layer prompt information;
[0020] Performing prompt generation processing on the task constraints of the humanoid robot according to the industrial scenario to obtain constraint layer prompt information;
[0021] The multi-layer prompt information is obtained according to the base layer prompt information, the task layer prompt information and the constraint layer prompt information.
[0022] In some embodiments, inputting the preprocessed dataset and the multi-layer prompt information into a pre-trained multimodal model for collaborative reasoning processing, and outputting an industrial task processing result, includes:
[0023] Performing feature extraction and knowledge graph embedding processing on the preprocessed data set to obtain visual recognition features;
[0024] Performing temporal dependency capture processing on the preprocessed data set to obtain tactile recognition features;
[0025] Encoding the multi-layer prompt information to obtain a prompt feature vector;
[0026] The visual recognition feature and the tactile recognition feature are cross-modally fused according to the prompt feature vector to obtain the industrial task processing result.
[0027] In some embodiments, performing cross-modal fusion processing on the visual recognition feature and the tactile recognition feature according to the prompt feature vector to obtain the industrial task processing result includes:
[0028] Performing attention weight calculation processing on the visual recognition feature according to the prompt feature vector to obtain visual task attention;
[0029] Performing attention weight calculation processing on the tactile recognition feature according to the prompt feature vector to obtain tactile task attention;
[0030] performing feature fusion processing on the visual recognition feature and the tactile recognition feature according to the visual task attention and the tactile task attention to obtain a fused feature;
[0031] Task output prediction processing is performed based on the fusion features to obtain the industrial task processing result.
[0032] In some embodiments, before inputting the pre-processed dataset and the multi-layer prompt information into a pre-trained multimodal model for collaborative reasoning, the method further includes training the multimodal model, including:
[0033] Get the training data set and training prompt set;
[0034] Performing sample construction processing on the training data set according to the training prompt set to obtain a sample data set; the sample data set includes positive sample pairs and negative sample pairs;
[0035] Inputting the sample data set into the multimodal model according to a progressive training strategy to obtain a prediction result;
[0036] A model loss is calculated based on the prediction result, and parameters of the multimodal model are adjusted based on the model loss.
[0037] In some embodiments, inputting the sample dataset into the multimodal model according to the progressive training strategy to obtain a prediction result includes:
[0038] Performing unimodal prompt processing on the sample data set according to the training prompt set to obtain a first training set;
[0039] performing multimodal joint prompt processing on the sample data set according to the training prompt set to obtain a second training set;
[0040] Performing industrial scene prompt processing on the sample data set according to the training prompt set to obtain a third training set;
[0041] According to different training stages, the first training set, the second training set and the third training set are respectively input into the multimodal model for task prediction processing to obtain the prediction result.
[0042] To achieve the above objectives, another aspect of the present application provides a multimodal data processing system for a humanoid robot, the system comprising:
[0043] A data acquisition module is used to obtain a multimodal data set through a humanoid robot;
[0044] A preprocessing module, configured to preprocess the multimodal dataset to obtain a preprocessed dataset;
[0045] a prompt generation module, configured to generate multi-layer prompt information according to the industrial scenario of the humanoid robot;
[0046] The collaborative reasoning module is used to input the preprocessed data set and the multi-layer prompt information into a pre-trained multimodal model for collaborative reasoning processing, and output an industrial task processing result.
[0047] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned method when executing the computer program.
[0048] To achieve the above objectives, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0049] To achieve the above object, another aspect of the present application provides a computer program product, including a computer program, which implements the above method when executed by a processor.
[0050] The embodiments of the present application include at least the following beneficial effects: The present application provides a method, system, device, and medium for multimodal data processing of a humanoid robot. The solution obtains a multimodal data set through the collection of the humanoid robot, and can process the multimodal data, thereby improving the humanoid robot's data processing capabilities. In addition, the solution generates multiple layers of prompt information based on the industrial scenario of the humanoid robot, and can generate different prompt information based on specific industrial scenarios to guide the generation of model processing. It can flexibly adjust the processing strategy according to the characteristics of different industrial tasks, thereby improving the task processing capabilities in complex industrial scenarios. In addition, the solution inputs the preprocessed data set and the multiple layers of prompt information into a pre-trained multimodal model for collaborative reasoning processing, and outputs the industrial task processing results. This enables the humanoid robot to effectively integrate multimodal information such as vision and touch, adapt to the precision operation requirements in industrial scenarios, and improve the accuracy of the robot in performing industrial tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a flowchart of a multimodal data processing method for a humanoid robot provided in an embodiment of the present application;
[0052] Figure 2 This is a schematic diagram of the structure of a multimodal model provided in an embodiment of the present application;
[0053] Figure 3 1 is a schematic structural diagram of a multimodal data processing system for a humanoid robot provided in an embodiment of the present application;
[0054] Figure 4 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are merely examples of systems and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.
[0056] It will be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0057] The terms "at least one", "plurality", "each", "any", etc. used in this application include "at least one", "two" or more, "plurality" or "each", "any" or "any one", "each" or "any one" as used herein.
[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0059] In industrial manufacturing, humanoid robots must process multimodal information, including vision, touch, and voice, to complete complex tasks like assembly and quality inspection. In related technologies, general-purpose humanoid robots lack infusion of industrial domain knowledge, resulting in insufficient visual recognition accuracy for industrial parts, weak cross-modal tactile-visual fusion capabilities, and difficulty adapting to precision operations. Furthermore, their semantic understanding of industrial process instructions exhibits significant deviations. The data processing of related humanoid robots is not optimized for the specificities of industrial scenarios, such as high precision requirements and strict process constraints. This results in insufficient generalization and execution accuracy for humanoid robots in industrial tasks.
[0060] In view of this, the embodiments of the present application provide a method, system, device, and medium for multimodal data processing of a humanoid robot. This solution obtains a multimodal data set through the collection of a humanoid robot, and can process the multimodal data, thereby improving the humanoid robot's data processing capabilities. In addition, this solution generates multiple layers of prompt information based on the industrial scenario of the humanoid robot, and can generate different prompt information based on specific industrial scenarios to guide the generation of model processing. It can flexibly adjust the processing strategy according to the characteristics of different industrial tasks, thereby improving the task processing capabilities in complex industrial scenarios. In addition, this solution inputs the preprocessed data set and the multiple layers of prompt information into a pre-trained multimodal model for collaborative reasoning processing, and outputs the industrial task processing results. This enables the humanoid robot to effectively integrate multimodal information such as vision and touch, adapt to the precision operation requirements in industrial scenarios, and improve the accuracy of the robot in performing industrial tasks.
[0061] The embodiment of the present application provides a method for processing multimodal data of a humanoid robot, which relates to the field of intelligent robot technology. The method for processing multimodal data of a humanoid robot provided in the embodiment of the present application can be applied to a humanoid robot, can also be applied to a server, and can also be software running in a humanoid robot or a server. In some embodiments, the server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements the method for processing multimodal data of a humanoid robot, etc., but is not limited to the above forms.
[0062] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0063] Figure 1 This is an optional flowchart of a multimodal data processing method for a humanoid robot provided in an embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S104.
[0064] Step S101, obtaining a multimodal dataset through collection by a humanoid robot;
[0065] Step S102, preprocessing the multimodal dataset to obtain a preprocessed dataset;
[0066] Step S103, generating multi-layer prompt information according to the industrial scene of the humanoid robot;
[0067] Step S104: input the preprocessed data set and the multi-layer prompt information into a pre-trained multimodal model for collaborative reasoning processing, and output an industrial task processing result.
[0068] In steps S101 to S104 shown in the embodiment of the present application, a multimodal dataset is obtained by collecting data through a humanoid robot. The multimodal dataset may include visual images of parts collected by a camera, pressure signals collected by a tactile sensor, and operation instruction texts in process documents. The multimodal dataset is then preprocessed, specifically by performing convolution feature extraction on the visual image and combining it with an industrial defect detection algorithm for preprocessing, performing time-frequency domain transformation on the tactile signal to extract features, and performing semantic enhancement on the operation instruction text through an industrial domain vocabulary to obtain a preprocessed dataset. The embodiment of the present application also generates multi-layer prompt information based on the industrial scenario of the humanoid robot, wherein the prompt information may include priority weights based on the characteristics of the industrial task. The preprocessed multimodal data and the layered prompt information are input into a multimodal model including a cross-modal fusion layer for collaborative reasoning processing, and the industrial task processing result is output.
[0069] The embodiment of the present application obtains a multimodal training set from the industrial manufacturing field for preprocessing, and generates multi-layer target prompt information including a base layer, a task layer, and a constraint layer based on the industrial scenario. The preprocessed data and prompt information are input into a multimodal encoder, and a target multimodal model is obtained through comparative learning and progressive strategy training. Through customized multimodal data processing and a layered prompt mechanism for the industrial field, the embodiment of the present application can improve the model's environmental perception, motion control, and task planning capabilities in industrial scenarios, and is suitable for industrial tasks such as humanoid robot assembly and quality inspection.
[0070] In some embodiments, the multimodal dataset includes visual images, tactile signals, and operation instruction text, and preprocessing the multimodal dataset to obtain a preprocessed dataset includes:
[0071] Performing contour feature extraction processing on the visual image to obtain contour features;
[0072] Performing defect detection and vector generation processing on the visual image according to the contour features to obtain visual features;
[0073] Performing time-frequency domain transformation and frequency band feature extraction processing on the tactile signal to obtain tactile features;
[0074] Performing semantic enhancement processing on the operation instruction text according to a preset industrial field vocabulary to obtain a semantically enhanced text;
[0075] Performing semantic vector conversion and constraint feature addition processing on the semantically enhanced text to obtain text features;
[0076] The preprocessing data set is obtained according to the visual features, the tactile features and the text features.
[0077] In an embodiment of the present application, a visual image of a component is acquired by camera capture, and convolution feature extraction is performed on the visual image and preprocessed in combination with an industrial defect detection algorithm to extract key features in the image and identify potential defects. Visual features are obtained by vector generation of defect features and fusion of contour features. In an embodiment of the present application, a pressure signal is acquired by a tactile sensor, and the pressure signal is used as a tactile signal, thereby performing a time-frequency domain transformation on the tactile signal to extract impact features, and the time domain signal is converted into a time-frequency feature that is easy to analyze to obtain tactile features. In an embodiment of the present application, the operation instruction text in the process document is obtained, the operation instruction text is semantically enhanced through an industrial domain vocabulary, and constraint features are added to the enhanced features to obtain text features, which can improve the model's ability to understand industrial professional terms and instructions.
[0078] In one feasible embodiment, a camera captures multi-angle images of components with a resolution of at least 1280×720 pixels to ensure discernible component details. The preprocessing phase first extracts contour features using a convolutional network consisting of at least 10 convolutional layers, employing batch normalization and ReLU activation functions to enhance feature extraction. Image preprocessing is combined with industrial defect detection algorithms: edge detection operators are used to identify component contours, and morphological operations are used to optimize contour continuity. A feature extraction network is then used to generate multidimensional feature vectors for potential defect areas, which are then used for subsequent defect classification. For example, in motor housing assembly scenarios, preprocessed visual features can accurately distinguish bolt hole locations from surface scratches. A multidimensional tactile sensor on the dexterous hand collects pressure signals during component grasping. In one embodiment, the sampling frequency is 1kHz to capture instantaneous contact force changes. The tactile signals are transformed into the time-frequency domain using a wavelet transform to decompose the time-domain signals into characteristic components in different frequency bands. High-frequency impact features (such as sudden force changes at the moment of grasping) and low-frequency trend features (such as continuous changes in grasping force) are specifically extracted. Taking gear grabbing as an example, the features after time-frequency domain transformation can reflect the force distribution law when the gears are engaged, providing data support for subsequent action control. Extract operation instruction texts from process documents, such as "tighten the bolts to the specified torque" and "locate the bearing mounting hole". Semantic enhancement is performed through an industrial vocabulary. One embodiment of the vocabulary contains no less than 1,000 industrial professional terms, covering component names, process parameters, operating specifications, etc. Word embedding technology is used to convert text into semantic vectors, and process constraints (such as torque range, accuracy requirements, etc.) are introduced as additional feature dimensions. For example, the instruction "tighten the cylinder head bolts" is converted into a semantic vector containing keywords such as "cylinder head", "bolts", and "tighten", and the constraint feature of "torque 20-25N·m" is added.
[0079] In some embodiments, generating multi-layer prompt information according to the industrial scenario of the humanoid robot includes:
[0080] Performing prompt generation processing on the role attributes of the humanoid robot according to the industrial scenario to obtain basic layer prompt information;
[0081] Performing prompt generation processing on the industrial task of the humanoid robot according to the industrial scenario to obtain task-layer prompt information;
[0082] Performing prompt generation processing on the task constraints of the humanoid robot according to the industrial scenario to obtain constraint layer prompt information;
[0083] The multi-layer prompt information is obtained according to the base layer prompt information, the task layer prompt information and the constraint layer prompt information.
[0084] In an embodiment of the present application, prompts are generated for the role attributes of a humanoid robot based on an industrial scenario. The industrial scenario can be determined by classifying visual images or by user input of specific scenario data. This embodiment defines the role attributes of the humanoid robot, and then generates base-level prompts by defining the robot's role and task scope. For example, the robot's role is defined as "automotive engine assembly robot," and its task scope is defined as "identifying cylinder block components." The base-level prompts clarify the robot's role attributes in the industrial scenario, such as "acting as an automotive engine assembly robot" or "acting as an electronic component quality inspection robot." Task scopes are also defined, such as "identifying engine block components" or "detecting circuit board solder joint defects." This prompt layer utilizes a templated design, including scenario keywords (such as "engine block" and "circuit board") and role keywords (such as "assembly" and "quality inspection"), ensuring the domain knowledge foundation for model building. This embodiment also breaks down the specific execution steps of the industrial tasks that the humanoid robot needs to perform to generate task-level prompts. For example, prompts are generated for the logical sequence of tasks in the industrial task, such as "locate bolt hole → tighten." This embodiment of the application also generates constraint layer prompts based on process standards and precision requirements, such as using the mandatory technical parameter "torque error ≤ 1%" as a constraint layer prompt. In a parts assembly scenario, the base layer prompt might be "Acting as an automotive assembly robot," the task layer prompt might be "Locate bolt holes and perform tightening operations," and the constraint layer prompt might be "Torque error ≤ 1%, tightening sequence: top left → bottom right → top right → bottom left."
[0085] In some embodiments, inputting the preprocessed dataset and the multi-layer prompt information into a pre-trained multimodal model for collaborative reasoning processing, and outputting an industrial task processing result, includes:
[0086] Performing feature extraction and knowledge graph embedding processing on the preprocessed data set to obtain visual recognition features;
[0087] Performing temporal dependency capture processing on the preprocessed data set to obtain tactile recognition features;
[0088] Encoding the multi-layer prompt information to obtain a prompt feature vector;
[0089] The visual recognition feature and the tactile recognition feature are cross-modally fused according to the prompt feature vector to obtain the industrial task processing result.
[0090] In the embodiment of the present application, the pre-processed data set and the multi-layer prompt information are input into the pre-trained multimodal model, such as Figure 2As shown, the multimodal model includes a visual branch, a tactile branch, and a cross-modal fusion layer. The visual branch extracts features from a preprocessed dataset and embeds them into a knowledge graph to obtain visual recognition features. Specifically, this is done by using a 3D convolutional network to process the visual data in the preprocessed dataset. In one feasible embodiment, the visual branch's network depth is no less than 15 layers, including a feature extraction module and a knowledge graph embedding module. The component knowledge graph stores no fewer than 100,000 relationship triples, such as "bolt-nut-connection" and "bearing-shaft-support." These are encoded via a graph neural network and then injected into the visual feature space, enhancing the model's understanding of component assembly relationships. In this embodiment, the tactile branch captures temporal dependencies on the preprocessed dataset to obtain tactile recognition features. The pressure signal sequence is processed via a temporal neural network. The tactile branch's network includes a bidirectional LSTM layer and an attention mechanism, capable of capturing the temporal dependencies of force signals. This embodiment can also establish a tactile feature library based on the grasping force-displacement curve. This feature library stores force feedback patterns for more than 500 industrial operations, such as "smooth grasping," "thread engagement," and "rigid collision."
[0091] The embodiment of the present application obtains a prompt feature vector by encoding multiple layers of prompt information, and specifically generates operation instruction features through semantic enhancement through language branches, including base layer prompts (roles), task layer prompts (goals), and constraint layer prompts (process parameters). Finally, the dynamic weighted fusion of visual and tactile features is achieved through the gating unit. The fusion weight is automatically adjusted by the type of industrial task. For example, in assembly tasks, the weight of visual features accounts for 60%, and the weight of tactile features accounts for 40%; the weight of visual features in quality inspection tasks is increased to 80%. The gating unit calculates the weight coefficient through nonlinear transformation to ensure that key modal features dominate the fusion result.
[0092] In a feasible embodiment, the visual branch prioritizes the identification of cylinder bolt holes under the prompt of the base layer, the task layer prompts to strengthen the extraction of hole edge features, and the constraint layer prompts to increase the loss weight of the hole coordinate error by 50%, and the final positioning error is controlled at 0.05mm; the tactile branch matches the "tightening" force signal template according to the prompt of the task layer, and the constraint layer prompts to set the loss weight of the torque prediction to 2 times that of the ordinary feature, and the torque deviation is controlled at 0.7%; during cross-modal fusion, the constraint layer prompts to increase the tactile weight from 0.4 to 0.6, ensuring that the tightening process is dominated by tactile feedback, and the final assembly accuracy meets the ISO 9283 industrial standard. The embodiment of the present application fully demonstrates the guiding mechanism of multi-layer prompt information in the multimodal model through the logical chain of "fusion input → branch interaction → cross-modal modulation → progressive optimization".
[0093] In some embodiments, performing cross-modal fusion processing on the visual recognition feature and the tactile recognition feature according to the prompt feature vector to obtain the industrial task processing result includes:
[0094] Performing attention weight calculation processing on the visual recognition feature according to the prompt feature vector to obtain visual task attention;
[0095] Performing attention weight calculation processing on the tactile recognition feature according to the prompt feature vector to obtain tactile task attention;
[0096] performing feature fusion processing on the visual recognition feature and the tactile recognition feature according to the visual task attention and the tactile task attention to obtain a fused feature;
[0097] Task output prediction processing is performed based on the fusion features to obtain the industrial task processing result.
[0098] In an embodiment of the present application, the cross-modal fusion layer calculates the fusion weight of the visual-tactile features based on the prompt information, wherein the visual task attention is obtained by the interaction of three layers of prompt information and visual features. For example, the base layer prompt (such as "automobile engine assembly robot") is fused with the parts relationship in the knowledge graph (such as the "bolt-cylinder" connection relationship) to generate domain prior features and inject them into the fully connected layer of the 3D convolutional network to improve the accuracy of parts recognition. The task layer prompt (such as "positioning bolt hole") is used as a visual attention mask to enhance the feature activation value of the bolt hole area in the image and suppress background noise. The constraint layer prompt (such as "installation error ≤ 0.1mm") is converted into the weight coefficient of the spatial coordinate loss function, so that the hole coordinates output by the model are closer to the industrial standard requirements. Among them, the base layer prompt vector is point-multiplied with the parts knowledge graph vector of the visual branch to strengthen the domain prior knowledge; the task layer prompt vector is used as the query vector of the attention mechanism to guide the model to focus on the modal features related to the task; the constraint layer prompt vector is converted into the regularization term of the loss function to increase the penalty weight for the output that does not meet the process requirements. In the embodiment of the present application, the linear transformation of the visual recognition feature and the prompt feature vector is calculated by the activation function to obtain the visual task attention. Similarly, the attention weight of the tactile recognition feature can be calculated by the prompt feature vector to obtain the tactile task attention. The tactile task attention can also be directly calculated by 1-visual task attention. Finally, the visual recognition feature and the tactile recognition feature are subjected to feature fusion processing according to the visual task attention and the tactile task attention, that is, the fusion feature is calculated by the formula: fusion feature = visual weight × visual feature + tactile weight × tactile feature. The fused feature is used to predict the task output, for example, to predict the control instruction of the humanoid robot arm, and the control instruction is used as the result of industrial task processing.
[0099] In a feasible embodiment, different modal features are mapped to a unified feature space through linear transformation, and the task-level prompt (such as "locate the bolt hole and tighten it") is used as the query vector Q, and the visual feature V' and the tactile feature T' are used as the key K respectively. v ,K t Sum V v ,V t , the constraint layer prompt (such as "torque error ≤ 1%") is used as the weight bias term b constraint , adjust the attention weight. The correlation between visual features and task prompts is calculated according to the visual task attention calculation formula. The visual task attention calculation formula is shown as follows:
[0100]
[0101] Where, Attn v represents the visual task attention, D k is the key vector dimension. This visual task attention reflects the degree to which the task cue focuses on visual features, such as focusing on the visual features of the bolt hole during bolt assembly. Similarly, tactile task attention is calculated by calculating the correlation between tactile features and task cues. The tactile task attention calculation formula is as follows:
[0102]
[0103] Where, Attn t Indicates attention to tactile tasks. In this embodiment, a gating unit dynamically fuses visual, tactile, and language features, where the gating signal is automatically adjusted based on the type of industrial task, for example, favoring visual features in assembly tasks and tactile features in quality inspection tasks.
[0104] In some embodiments, before inputting the pre-processed dataset and the multi-layer prompt information into a pre-trained multimodal model for collaborative reasoning, the method further includes training the multimodal model, including:
[0105] Get the training data set and training prompt set;
[0106] Performing sample construction processing on the training data set according to the training prompt set to obtain a sample data set; the sample data set includes positive sample pairs and negative sample pairs;
[0107] Inputting the sample data set into the multimodal model according to a progressive training strategy to obtain a prediction result;
[0108] A model loss is calculated based on the prediction result, and parameters of the multimodal model are adjusted based on the model loss.
[0109] In an embodiment of the present application, a training data set and a training prompt set are obtained, wherein the training data set may include a motion control training set, which may include control instructions (such as "lift the robotic arm to a specified position") and action execution information (such as "joint angle change" and "end pose coordinates"). During training, the mapping relationship between control instructions and action execution is learned through a temporal neural network in combination with kinematic constraints (such as joint angle limits and speed limits). For example, in the robotic arm painting task, the model masters the correspondence between the "swing amplitude" instruction and the paint trajectory through training, while meeting the process constraint of "spraying speed 10-15mm / s". In an embodiment of the present application, the training data set may also include an object operation training set and a voice interaction training set, wherein the object operation training set includes operation requests (such as "grab the gear and install it to the shaft end") and operation knowledge information (such as "gear weight 5kg" and "friction coefficient 0.3"). During training, the operation request is converted into a semantic vector, combined with the physical property constraints of the object, and the operation strategy is generated through the cross-modal fusion layer. For example, when grasping fragile parts, the model will automatically lower the grasping force threshold and adjust the grasping posture to reduce impact. The voice interaction training set contains industrial scenario voice requests (such as "query the current workstation task" and "report equipment failure") and interactive response information (such as "current task: assemble motor housing" and "fault code: E001, conveyor belt jam"). During training, the model is combined with industrial scenario speech templates (such as fault reporting process and task query format) to enhance the model's understanding of professional terms through the attention mechanism. For example, when receiving a voice request to "detect abnormal noise in the gearbox", the model can associate it with industrial terms such as "gearbox" and "abnormal noise detection" and generate corresponding detection process responses.
[0110] In a feasible embodiment, the training data set is subjected to sample construction processing through a training prompt set to obtain a sample data set, and the sample data set includes positive sample pairs and negative sample pairs. Specifically, the "semantic-visual" positive samples can be defined through the base layer prompts, such as the "bolt" text and the bolt image, and the "action-tactile" positive samples can be defined through the task layer prompts, such as the "tighten" instruction and the torque force signal, and the "parameter-feature" positive samples can be defined through the constraint layer prompts, such as "30N·m" and the force signal peak. According to the progressive training strategy, the sample data set is input into the multimodal model for prediction to obtain the prediction result. The model loss can be calculated through the loss function, and the parameters of the multimodal model are adjusted according to the model loss. The embodiment of the present application improves the accuracy of cross-modal semantic alignment by comparing the learning loss function to force the distance between similar samples in the feature space to be smaller than that of heterogeneous samples.
[0111] In some embodiments, inputting the sample dataset into the multimodal model according to the progressive training strategy to obtain a prediction result includes:
[0112] Performing unimodal prompt processing on the sample data set according to the training prompt set to obtain a first training set;
[0113] performing multimodal joint prompt processing on the sample data set according to the training prompt set to obtain a second training set;
[0114] Performing industrial scene prompt processing on the sample data set according to the training prompt set to obtain a third training set;
[0115] According to different training stages, the first training set, the second training set and the third training set are respectively input into the multimodal model for task prediction processing to obtain the prediction result.
[0116] In an embodiment of the present application, a sample data set is processed with single-modal prompts according to a training prompt set, that is, only basic layer prompts are used to establish a basic representation of the domain, and a first training set is obtained. A sample data set is processed with multimodal joint prompts according to the training prompt set to obtain a second training set, that is, task layer prompts are superimposed to guide multimodal collaborative reasoning. A sample data set is processed with industrial scenario prompts according to the training prompt set to obtain a third training set, that is, all three layers of prompts are enabled, and the model output is dynamically adjusted in combination with the real-time constraints of the production line. The first training set, the second training set, and the third training set are respectively input into the multimodal model for task prediction processing at different training stages, and prediction results at different training stages can be obtained.
[0117] The following describes the solution of the embodiment of the present invention in detail with reference to specific application examples:
[0118] The embodiment of the present application is applied to the field of intelligent robot technology, and is suitable for industrial tasks such as humanoid robot assembly and quality inspection. Specifically, by obtaining a multimodal training set in the field of industrial manufacturing, including visual images, tactile signals, and operation instruction texts; preprocessing the multimodal data, which may include convolutional feature extraction, wavelet transform, semantic enhancement, etc.; and generating hierarchical target prompt information including a base layer, a task layer, and a constraint layer; inputting the preprocessed data and the prompt information into a multimodal encoder, and obtaining a target multimodal model through comparative learning and progressive strategy training, and obtaining the industrial task processing result through the multimodal model output. In the actual industrial production line, the model is fine-tuned online through reinforcement learning. The reward function is designed based on task execution accuracy (such as assembly completion rate) and efficiency (such as operation time), so that the model can adapt to the constraints of a specific industrial environment.
[0119] See also Figure 3 The present application also provides a humanoid robot multimodal data processing system that can implement the above method. The system includes:
[0120] The data acquisition module 301 is used to obtain a multimodal data set through a humanoid robot;
[0121] A preprocessing module 302 is used to preprocess the multimodal dataset to obtain a preprocessed dataset;
[0122] a prompt generating module 303, configured to generate multi-layer prompt information according to the industrial scenario of the humanoid robot;
[0123] The collaborative reasoning module 304 is used to input the pre-processed data set and the multi-layer prompt information into a pre-trained multimodal model for collaborative reasoning processing, and output an industrial task processing result.
[0124] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0125] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above method when executing the computer program. The electronic device can be any smart terminal including a tablet computer, an in-vehicle computer, or the like.
[0126] It can be understood that the contents of the above method embodiments are applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0127] See also Figure 4 , Figure 4 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0128] The processor 401 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0129] The memory 402 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 402 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 402 and is called by the processor 401 to execute the above-mentioned methods of the embodiments of this application.
[0130] Input / output interface 403, used to implement information input and output;
[0131] Communication interface 404, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0132] Bus 405 , which transmits information between various components of the device (e.g., processor 401 , memory 402 , input / output interface 403 , and communication interface 404 );
[0133] The processor 401 , the memory 402 , the input / output interface 403 and the communication interface 404 are connected to each other in communication within the device via a bus 405 .
[0134] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above method is implemented.
[0135] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0136] An embodiment of the present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.
[0137] It is understandable that the contents of the above method embodiments are all applicable to the present program product embodiments, the functions specifically implemented by the present program product embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0138] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0139] The embodiments of the present application provide a method, system, device, and medium for processing multimodal data for a humanoid robot. This solution collects a multimodal data set through the humanoid robot and can process the multimodal data, thereby improving the humanoid robot's data processing capabilities. Furthermore, this solution generates multiple layers of prompt information based on the industrial scenario of the humanoid robot, can generate different prompt information based on specific industrial scenarios to guide the generation of model processing, can flexibly adjust processing strategies based on the characteristics of different industrial tasks, and improve task processing capabilities in complex industrial scenarios. Furthermore, this solution inputs the preprocessed data set and multiple layers of prompt information into a pre-trained multimodal model for collaborative reasoning processing, and outputs the industrial task processing results. This enables the humanoid robot to effectively integrate multimodal information such as vision and touch, adapt to the precision operation requirements in industrial scenarios, and improve the accuracy of the robot in performing industrial tasks.
[0140] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0141] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0142] The system embodiment described above is merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0143] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0144] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0145] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0146] In the several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the above units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of systems or units, which can be electrical, mechanical or other forms.
[0147] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0148] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0149] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0150] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A method for processing multimodal data of a humanoid robot, characterized in that: The method comprises the following steps: A multimodal dataset is obtained through humanoid robot collection; Preprocessing the multimodal dataset to obtain a preprocessed dataset; generating multi-layer prompt information according to the industrial scenario of the humanoid robot; The preprocessed data set and the multi-layer prompt information are input into a pre-trained multimodal model for collaborative reasoning processing, and an industrial task processing result is output.
2. The method according to claim 1, characterized in that The multimodal dataset includes visual images, tactile signals, and operation instruction texts. The preprocessing of the multimodal dataset to obtain a preprocessed dataset includes: Performing contour feature extraction processing on the visual image to obtain contour features; Performing defect detection and vector generation processing on the visual image according to the contour features to obtain visual features; Performing time-frequency domain transformation and frequency band feature extraction processing on the tactile signal to obtain tactile features; Performing semantic enhancement processing on the operation instruction text according to a preset industrial field vocabulary to obtain a semantically enhanced text; Performing semantic vector conversion and constraint feature addition processing on the semantically enhanced text to obtain text features; The preprocessing data set is obtained according to the visual features, the tactile features and the text features.
3. The method according to claim 1, characterized in that The generating of multi-layer prompt information according to the industrial scene of the humanoid robot includes: Performing prompt generation processing on the role attributes of the humanoid robot according to the industrial scenario to obtain basic layer prompt information; Performing prompt generation processing on the industrial task of the humanoid robot according to the industrial scenario to obtain task-layer prompt information; Performing prompt generation processing on the task constraints of the humanoid robot according to the industrial scenario to obtain constraint layer prompt information; The multi-layer prompt information is obtained according to the base layer prompt information, the task layer prompt information and the constraint layer prompt information.
4. The method according to claim 1, wherein The step of inputting the pre-processed data set and the multi-layer prompt information into a pre-trained multimodal model for collaborative reasoning processing and outputting an industrial task processing result includes: Performing feature extraction and knowledge graph embedding processing on the preprocessed data set to obtain visual recognition features; Performing temporal dependency capture processing on the preprocessed data set to obtain tactile recognition features; Encoding the multi-layer prompt information to obtain a prompt feature vector; The visual recognition feature and the tactile recognition feature are cross-modally fused according to the prompt feature vector to obtain the industrial task processing result.
5. The method according to claim 4, characterized in that The cross-modal fusion processing of the visual recognition feature and the tactile recognition feature according to the prompt feature vector to obtain the industrial task processing result includes: Performing attention weight calculation processing on the visual recognition feature according to the prompt feature vector to obtain visual task attention; Performing attention weight calculation processing on the tactile recognition feature according to the prompt feature vector to obtain tactile task attention; performing feature fusion processing on the visual recognition feature and the tactile recognition feature according to the visual task attention and the tactile task attention to obtain a fused feature; Task output prediction processing is performed based on the fusion features to obtain the industrial task processing result.
6. The method according to claim 1, characterized in that Before inputting the pre-processed data set and the multi-layer prompt information into a pre-trained multimodal model for collaborative reasoning, the method further includes training the multimodal model, including: Get the training data set and training prompt set; Performing sample construction processing on the training data set according to the training prompt set to obtain a sample data set; the sample data set includes positive sample pairs and negative sample pairs; Inputting the sample data set into the multimodal model according to a progressive training strategy to obtain a prediction result; A model loss is calculated based on the prediction result, and parameters of the multimodal model are adjusted based on the model loss.
7. The method according to claim 6, characterized in that The step of inputting the sample data set into the multimodal model according to the progressive training strategy to obtain a prediction result includes: Performing unimodal prompt processing on the sample data set according to the training prompt set to obtain a first training set; performing multimodal joint prompt processing on the sample data set according to the training prompt set to obtain a second training set; Performing industrial scene prompt processing on the sample data set according to the training prompt set to obtain a third training set; According to different training stages, the first training set, the second training set and the third training set are respectively input into the multimodal model for task prediction processing to obtain the prediction result.
8. A multimodal data processing system for a humanoid robot, characterized in that: The system comprises: A data acquisition module is used to obtain a multimodal data set through a humanoid robot; A preprocessing module, configured to preprocess the multimodal dataset to obtain a preprocessed dataset; a prompt generation module, configured to generate multi-layer prompt information according to the industrial scenario of the humanoid robot; The collaborative reasoning module is used to input the preprocessed data set and the multi-layer prompt information into a pre-trained multimodal model for collaborative reasoning processing, and output an industrial task processing result.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-modal data processing method applied to robot interaction
CN113894779A
Man-machine co-fusion mechanical arm self-adaptive grabbing method and system based on multi-modal large model
CN118789551A
Robot control method and device based on multi-modal data fusion
CN119260752A
Smart home scene understanding and interaction method and system based on multi-modal fusion
CN119398159A
Multi-modal intelligent interaction optimization method and system
CN119937787A
Cited By
Training method of humanoid robot task planning model and task planning method
CN121340271A
Behavior labeling method based on end-to-end humanoid robot large model and related equipment
CN121705817A