Operation control method and device based on multi-modal information fusion, equipment and medium
Through the operation control method of multimodal information fusion, visual and language features are obtained, action sequences are generated and actuators are controlled, which solves the adaptability problems of existing technologies in dynamic environments and diversified tasks and realizes more precise and intelligent operation control.
Patent Information
- Application Number
- CN202510918112.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies lack intelligent adaptability when faced with dynamic environments and diverse tasks, and are unable to adjust decision-making and operational strategies in real time, resulting in limited application of robotic arms, medical robots and financial technology systems in complex environments.
By acquiring image data and task instructions of the operating environment, extracting visual and language features, fusing them to generate fusion features, and inputting them into the action generation model to generate action sequences, control the actuator to perform operations, and adjust the model parameters based on feedback signals.
It improves the adaptability and multi-tasking processing capabilities of the operation control system in complex environments, and significantly enhances the system's generalization ability and execution efficiency.
Smart Images

Figure CN120791752A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to an operation control method and device based on multi-modal information fusion, equipment and a storage medium. BACKGROUND
[0002] In the field of embodied intelligence, traditional robotic arm grasping systems usually rely on pre-set fixed programs and simple visual feedback mechanisms, which makes these systems unable to effectively cope with dynamically changing environments and diverse objects. In existing systems, visual recognition technology often cannot cope with changes in object shape, color, texture and other characteristics, especially when the object placement angle, lighting conditions change or the object category changes, the recognition accuracy of traditional visual recognition systems decreases sharply. In addition, traditional decision planning methods mostly rely on simple rules or limited experience data, which makes the system lack the ability to cope with unexpected situations in complex environments. When the environment changes, the task requirements change, the system usually cannot generate effective grasping strategies in time, resulting in the robotic arm unable to correctly perform the task. Therefore, the application of existing systems in complex environments is limited, and it is urgent to introduce more intelligent perception and decision systems to improve their adaptability and generalization ability.
[0003] In the field of medical health, intelligent robots need to perform tasks in complex and delicate environments, such as precise grasping of surgical tools or processing of pathological specimens. Traditional medical robot systems usually rely on fixed visual recognition and grasping strategies, which makes the system unable to adapt to the diverse needs in medical scenarios. For example, when the types and shapes of instruments used in surgery change, traditional systems may not be able to accurately identify and grasp the instruments, thereby affecting the smooth progress of the surgery. At the same time, existing medical robot systems usually lack deep understanding and reasoning ability for complex tasks, resulting in poor adaptability of the system in dynamic environments, and the system cannot adjust the strategy according to real-time feedback. Therefore, improving the intelligent level of medical robot systems, especially the coordination between visual perception, decision planning and execution control modules, is a key problem that needs to be solved in current technology.
[0004] In the field of financial technology, intelligent robots need to perform tasks in a dynamic financial business scenario and handle various financial products. Existing automated systems usually plan tasks based on fixed rules, which are difficult to adapt to the rapid changes in the types of financial products or the complexity of business scenarios. For example, when new product models are introduced or the structure of existing products is modified in a bank business, traditional systems may not be able to adapt to these changes in a timely manner, resulting in task execution failures. In addition, existing financial technology systems exhibit poor adaptability and accuracy in responding to dynamic environments and complex business scenarios, failing to meet the actual needs of customer service, product processing, and other multi-task scenarios. Therefore, developing intelligent robot systems that can learn and adjust in real time to adapt to the rapid changes and complex environments of financial scenarios has become a pressing problem in the field of financial technology. SUMMARY
[0005] The main purpose of the present application is to provide a multi-modal information fusion-based operation control method, device, equipment and storage medium, aiming to solve the technical problems that the existing technology lacks intelligent adaptability to dynamic environmental changes and diversified tasks, and cannot adjust decision and operation strategy in real time.
[0006] To achieve the above-mentioned purpose, the present application provides a multi-modal information fusion-based operation control method, comprising:
[0007] obtaining image data of an operation environment and task instructions for describing an operation task;
[0008] extracting visual features from the image data and language features from the task instructions;
[0009] fusing the visual features and the language features to generate fusion features;
[0010] inputting the fusion features into an action generation model to generate an action sequence for controlling an actuator;
[0011] controlling the actuator to perform an operation in the operation environment according to the action sequence;
[0012] obtaining the execution result of the operation to generate a feedback signal, and adjusting the parameters of the action generation model based on the feedback signal.
[0013] Further, to achieve the above-mentioned purpose, the present application provides a multi-modal information fusion-based operation control device, comprising:
[0014] a data acquisition module for obtaining image data of an operation environment and task instructions for describing an operation task;
[0015] a feature extraction module for extracting visual features from the image data and language features from the task instructions;
[0016] a feature fusion module configured to fuse the visual features and the language features to generate fused features;
[0017] an action generation module configured to input the fused features into an action generation model to generate an action sequence for controlling an effector;
[0018] an effector control module configured to control the effector to perform an operation in the operation environment according to the action sequence;
[0019] a feedback adjustment module configured to obtain an execution result of the operation to generate a feedback signal, and adjust parameters of the action generation model based on the feedback signal.
[0020] Further, to achieve the above object, the present application also provides a computer device, which comprises a memory, a processor, and a multi-modal information fusion based operation control program stored in the memory and executable on the processor, and the multi-modal information fusion based operation control program, when executed by the processor, implements the steps of the multi-modal information fusion based operation control method.
[0021] Further, to achieve the above object, the present application also provides a computer readable storage medium, which stores a multi-modal information fusion based operation control program, and the multi-modal information fusion based operation control program, when executed by a processor, implements the steps of the multi-modal information fusion based operation control method.
[0022] Beneficial effects: The present application relates to the field of artificial intelligence technology, and can be applied to embodied intelligence, financial technology, medical health and other business scenarios, and discloses a multi-modal information fusion based operation control method, device, equipment and medium, which comprises the following steps: obtaining image data of an operation environment and a task instruction, extracting visual features in the image data and language features in the task instruction, fusing the visual features and the language features to generate fused features, inputting the fused features into an action generation model to generate an action sequence, controlling an effector to perform an operation in the operation environment according to the action sequence, obtaining an execution result to generate a feedback signal, and adjusting parameters of the action generation model based on the feedback signal. The present application fuses visual and language features, improves the adaptability of an operation control system to complex environmental changes and the multi-task processing capability, and thus can realize more accurate and intelligent operation control, and significantly improves the generalization capability and execution efficiency of the system, especially when facing diversified tasks and dynamic environments. BRIEF DESCRIPTION OF DRAWINGS
[0023] The present application will be further described below with reference to the accompanying drawings and embodiments, in which:
[0024] Figure 1 An application environment schematic diagram of the operation control method based on multi-modal information fusion in an embodiment of the present application;
[0025] Figure 2 A flowchart of the operation control method based on multi-modal information fusion in an embodiment of the present application;
[0026] Figure 3 A functional module schematic diagram of the operation control device based on multi-modal information fusion in a preferred embodiment of the present application;
[0027] Figure 4 A structure schematic diagram of a computer device in an embodiment of the present application;
[0028] Figure 5 Another structure schematic diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION
[0029] It should be understood that the specific embodiments described herein are merely exemplary and not intended to limit the present application.
[0030] The operation control method based on multi-modal information fusion provided by the embodiments of the present application can be applied in an application environment as shown in Figure 1 , wherein a user end communicates with a service end through a network. The service end can obtain image data and task instructions of an operation environment through the user end, extract visual features in the image data and language features in the task instructions, generate fusion features by fusing the visual features and the language features, input the fusion features into an action generation model to generate an action sequence, control an executor to perform operations in the operation environment according to the action sequence, obtain an execution result to generate a feedback signal, and adjust parameters of the action generation model based on the feedback signal. The present application improves the adaptability to complex environment changes and the multi-task processing capability of the operation control system by fusing visual and language features, so as to realize more accurate and intelligent operation control, and significantly improves the generalization capability and execution efficiency of the system, especially when facing diversified tasks and dynamic environments. The user end can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The service end can be realized by an independent server or a server cluster composed of multiple servers. The present application will be described in detail through specific embodiments.
[0031] Please refer to Figure 2 , Figure 2 A flowchart of the operation control method based on multi-modal information fusion provided by the present application in an embodiment. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that herein.
[0032] As shown in Figure 2 The multi-modal information fusion-based operation control method includes the following steps:
[0033] S10, acquiring image data of an operation environment and task instructions for describing an operation task;
[0034] In this embodiment, in the intelligent device control, the acquisition of image data of the operation environment and the task instructions is the basis for effective control. The image data of the operation environment is usually collected by image sensors (such as cameras, lidar, depth cameras, etc.), and through these image data, the device can perceive and understand its surrounding environment. The data collected by these sensors includes visual information of the environment, such as the position, pose, shape, color and texture of the object, and various features.
[0035] The task instructions are instructions for describing the operation target, which are usually manually input or automatically generated by other systems. These instructions include the target of the task, the execution requirements of the task, the time limit, etc. For example, when a mechanical arm performs a grasping operation, the task instructions may include the name of the grasped object, the target position, and whether to avoid obstacles, etc. By combining image data and task instructions, comprehensive decision support can be provided for subsequent operations.
[0036] In actual implementation, the image sensor collects image data according to the actual environmental situation, and transmits the data to the processing unit through a hardware interface. The task instructions are usually received through an interface (such as a speech recognition system or text input), and are parsed into operable instructions in the processing unit. This process can be optimized through pre-set formats and protocols to ensure efficient and accurate data transmission. In more complex environments, multiple sensors may need to work in parallel to obtain information from different angles and depths in real time, ensuring that the control system can fully understand the operation environment. The acquisition of image data and task instructions is further improved through sensor fusion technology to improve the accuracy and reliability of the data.
[0037] In the implementation of image data acquisition, the use of multiple image sensors and their cooperation can obtain multi-angle and depth information of the environment. For example, in a complex environment, combining RGB images, depth images and point cloud data can provide rich environmental information. The task instructions can be automatically converted into standard control commands through natural language processing technology, or directly input by humans. In order to enhance the accuracy and real-time performance of task execution, a speech recognition module can be integrated into the device to control the operation of the device through voice input instructions. In addition, the fusion of image data and task instructions usually requires the timeliness of data processing, and sensor data should be processed and transmitted within a short time to ensure that the system can respond to environmental changes in real time.
[0038] In implementation, the image data of the operating environment is first processed by the sensor of the image acquisition module to generate raw images or point cloud data, and then the noise is removed and the data accuracy is improved through the denoising and preprocessing module. After the task instruction is input, the natural language processing module is used for syntax analysis, key information extraction, and conversion to structured data, which is transmitted to the control system for subsequent operation. Multi-modal sensor fusion technology can be used to further enhance the accuracy and processing efficiency of information. In this environment, image and instruction data are dynamically alternating, and the system must be able to efficiently and stably process a large amount of real-time information.
[0039] Example: In embodied intelligence applications, intelligent devices need to be able to perceive their operating environment and perform tasks. For example, when a robot performs a cargo handling task in a warehouse, image sensors collect object information in the environment, and through task instructions, it learns that the items need to be moved. At this time, the robot generates control signals based on the analysis results of the image data (such as the position, shape, and posture of the items) and the specific requirements in the task instructions (such as the target to be moved and path selection), guiding the robot to complete the moving task. As the items change, the robot continuously acquires new image data and task instructions to ensure accurate operation in a dynamic environment.
[0040] In the field of medical health, when intelligent devices perform surgical assistance tasks, the image data obtained is analyzed by deep learning algorithms to understand the situation of the surgical area, and the task instructions provide the required medical operations for the device, such as positioning, grabbing, or cutting. Through the close combination of image data and task instructions, surgical robots can accurately perform operations and have high precision and real-time reaction capability when dealing with complex surgeries, ensuring patient safety and surgical effectiveness.
[0041] In the field of financial technology, intelligent devices need to quickly identify and process different types of financial products and data. By obtaining market data (such as real-time images or documents) and task instructions (such as user operation requirements), intelligent devices can efficiently perform automated processing. For example, in customer query services, image data can be analyzed by OCR technology to parse contracts or financial statements, and task instructions guide the device to perform corresponding operations, such as querying, generating reports, or handling other financial services, significantly improving the efficiency and accuracy of customer service.
[0042] This embodiment combines image data and task instructions to enable devices to more accurately perceive and understand the environment, and make intelligent decisions based on task goals. Image sensors can provide accurate environmental perception, and task instructions can guide the direction of device action. This combination provides the system with strong adaptability and flexibility, especially in dynamic environments, devices can quickly respond to various situations to achieve more efficient and accurate operations.
[0043] S20, extracting visual features from the image data and extracting language features from the task instruction;
[0044] In this embodiment, the extraction of visual features from image data and the extraction of language features from task instructions are key steps to provide accurate data for subsequent operations. Image data contains rich environmental information such as the shape, color, texture, position and pose of objects. Task instructions usually include operation requirements, target positions, obstacle avoidance rules and other content. By extracting these features, the device can obtain information from both visual and semantic aspects, realizing the fusion of multi-modal data.
[0045] Visual feature extraction is usually completed by deep neural networks such as convolutional neural networks (CNN). First, image data is obtained from sensors and preprocessed (such as denoising, scaling, etc.), and then input into the visual feature extraction network. CNN will extract low-level features (such as edges, corners, etc.) in the image, and abstract high-level semantic information (such as the type, position and pose of objects) in the image through multiple layers of network. These visual features will be used as input data for subsequent decision-making. In implementation, image data is collected by cameras or other visual sensors, and the output data of the sensor is transmitted to the image processing unit through the hardware interface. In the processing unit, convolutional neural networks or other deep learning methods are used for feature extraction. The data after feature extraction is normalized and dimensionally reduced to adapt to the subsequent processing flow. In addition, in order to ensure real-time performance, the computing power of the image processing and feature extraction module needs to match the processing power of the hardware device. An edge computing platform can be used to reduce data transmission delay and improve processing efficiency.
[0046] The extraction of language features in task instructions mainly relies on natural language processing (NLP) technology. First, the user's instruction or the system-generated instruction is received through the text input interface. After receiving the task instruction, the system will perform syntax analysis, entity recognition and semantic understanding. Through these processes, the system can identify key task information in the instruction, such as target objects, operation requirements, execution time limits, etc. NLP models such as BERT or GPT series models can be used to process task instructions to extract specific semantic information, ensuring that the system can understand and execute user or predetermined tasks. In implementation, task instructions can be received through a speech recognition system, a text input module or other interfaces. After these text information is received, it will be processed by the text analysis module, first performing word segmentation and part-of-speech tagging on the sentence, and then using a pre-trained natural language processing model for semantic understanding. Through these steps, the system can understand the intent of the task instruction and then extract key information related to task execution. In order to improve processing efficiency, parallel computing architecture can be used to ensure the real-time performance and accuracy of language feature extraction.
[0047] In specific implementation, image data is processed in real time through deep learning models. For example, convolutional neural networks can be used to process images, and existing deep learning frameworks such as TensorFlow and PyTorch can be used to build models. For language feature extraction, natural language processing models such as BERT can be used. These models have strong generalization ability in processing complex instructions through extensive pre-training and fine-tuning. To ensure system stability and efficiency, image data and task instruction processing modules can be designed as independent subsystems, using multi-threading or distributed computing architecture for data processing. Image processing and language feature extraction modules can synchronize data through intermediate caching mechanisms to ensure real-time updating and transmission of data. In different application scenarios, different image feature extraction algorithms can be selected as needed, such as Mask R-CNN for object detection or YOLO for real-time processing. For language feature extraction, in addition to text-based task instructions, voice recognition systems can also be used to receive voice instructions, which are then processed after being converted from voice to text.
[0048] Example: In the field of embodied intelligence, for example, when a robot performs a cargo handling task in a warehouse, it first obtains image data of the environment through a camera, extracts visual features such as the location and shape of the items. Through a voice recognition system, it receives task instructions such as "move the red box to the designated location", and the language features in the task instructions provide the robot with a clear task target. After fusing visual features and language features, the robot can accurately identify the target object based on image data and calculate the optimal grabbing path based on task instructions to complete the task.
[0049] In the field of medical health, surgical robots need to perform surgical operations based on real-time visual data (such as images of surgical instruments) and task instructions (such as grabbing a specific surgical tool). By extracting visual features from image data, the robot can identify the location, shape, and posture of the surgical tool. Task instructions provide the robot with specific operation targets and requirements, and the robot can accurately perform tasks and adjust operation strategies in real time by fusing these two types of information.
[0050] In the field of financial technology, intelligent devices can automatically process tasks such as customer service and financial statement analysis by combining real-time image data (such as scanned images of contracts or forms) and task instructions (such as customer requests). Visual features in image data can help the device understand the location of contract terms, and language features provide the device with specific requirements for processing tasks. By fusing these information, the device can quickly and accurately complete customer requests and improve business processing efficiency.
[0051] By combining the extraction of visual features and language features, the intelligent device can have strong environmental perception and task understanding capabilities, enabling it to complete precise tasks in complex and variable environments. Visual features provide real-time perception of the environment, while language features give the device the ability to understand and execute tasks. In this way, the device can not only handle dynamic changes in different environments, but also flexibly respond to diverse task requirements according to user instructions.
[0052] S30, fusing the visual features and the language features to generate fused features;
[0053] In this embodiment, the fusion of visual features and language features to generate fused features is a key step in realizing multi-modal data collaborative processing. Visual features and language features usually represent perception of the environment and understanding of the task, respectively, and their fusion can provide the intelligent system with more comprehensive perception capabilities and decision support. In this embodiment, visual feature extraction obtains environmental information through image data, while language feature extraction understands the operation target through task instructions. The fusion of these two types of features enables the device to generate more accurate and effective operation commands based on understanding of the environment and the task.
[0054] Visual features include low-level and high-level features extracted from image data, such as object shape, color, texture, position, pose, etc. These features are usually extracted from images through deep convolutional neural networks (CNN) or other visual processing networks (such as ResNet, YOLO, etc.). After processing by convolutional layers and pooling layers, CNN can extract geometric information and spatial layout of objects, and then generate high-level semantic features.
[0055] Language features are semantic information extracted from task instructions, which are usually input through text or speech and describe specific tasks that the intelligent device needs to complete. Common techniques for language feature extraction include natural language processing (NLP) methods such as Word2Vec, BERT, etc. Through these methods, the system can understand key information in the task instruction, such as target objects, operation types, execution time limits, etc.
[0056] The process of fusing visual features and language features is usually carried out through multi-modal learning methods. In this process, techniques such as feature-level fusion, decision-level fusion, or model-level fusion can be used. Feature-level fusion is to directly concatenate or weight combine the two features after feature extraction to form a unified feature representation; decision-level fusion is to combine the information of the two features through the decision layer of the model after separately extracting and processing the visual features and language features; model-level fusion is to design a multi-modal network architecture in the multi-layer structure of the neural network, so that visual information and language information can be learned together in the model.
[0057] A common implementation is to combine visual features and language features through concatenation to form a joint feature vector. This joint feature vector can help the intelligent device consider both the perception information of the environment and the semantic information of the task during further model training and inference, providing more accurate action sequence generation. For different tasks and environments, weighted fusion or more refined fusion operations through attention mechanisms can also be selected according to actual needs.
[0058] In implementation, first, the visual features and language features are standardized to ensure they have similar scales and ranges. Then, a fusion module (such as a fully connected layer, weighted summation, attention mechanism, etc.) is used to combine the two features. The fused features are used as input in subsequent models to drive the intelligent device to generate action sequences.
[0059] In specific implementation, visual feature extraction can use a convolutional neural network (CNN) architecture such as ResNet, VGG16, etc. The processed visual features can be low-level information such as object edges, colors, depths, or high-level information such as object categories, positions, sizes, etc. After training, the visual network model can extract object features from input image data as input for subsequent steps. For task instruction processing, an NLP model (such as BERT, GPT, etc.) can extract language features related to the task, such as target objects, action types, operation time limits, etc., by inputting the task instruction into text data. Through model training, the semantics of the task instruction can be effectively parsed and understood. The feature fusion process can be accelerated through parallel computing architecture, with visual features and language features processed separately and optimized through weighted concatenation or attention mechanisms in the fusion stage. The fused features are input into the action generation model to further generate action sequences for controlling the actuators. To adapt to different scenarios, the weight distribution of visual features and language features can be adjusted, and the fusion method can be optimized through experiments to improve the accuracy and adaptability of task execution.
[0060] Example: In embodied intelligence, robots need to perform precise grasping and carrying tasks based on image data in the environment (such as object position, shape, etc.) and task instructions (such as carrying a certain object to a specific location). By extracting visual features (such as object position, pose, etc.) from image data and language features (such as task target, location requirements, etc.) from task instructions, the system can fuse these two types of information into a unified feature vector, which can guide the robot to generate action sequences and achieve precise grasping and carrying.
[0061] In the medical health field, surgical robots need to perform tasks based on real-time image data (such as images of surgical instruments) and task instructions (such as grabbing a specific surgical tool). By extracting visual features from images and language features from task instructions, robots can integrate both pieces of information to accurately identify surgical tools and perform precise grabbing operations, ensuring the smooth progress of the surgery.
[0062] In the financial technology field, intelligent devices can process financial documents or customer requests. When processing contract scan image data, the device integrates key information in the image (such as contract terms, amounts, etc.) and language features in the task instructions (such as customer requirements for operations) to help the device make accurate decisions, quickly process customer requests, and improve work efficiency.
[0063] This embodiment can make intelligent devices simultaneously obtain environmental information and task information by integrating visual features and language features, providing more accurate and comprehensive decision support for subsequent action generation. This multi-modal integration method significantly improves the adaptability of devices in complex dynamic environments and enhances their precision and generalization ability in task execution. The integrated features can help devices accurately identify objects and understand task requirements in a changing environment, thereby better completing tasks.
[0064] S40, input the integrated features into an action generation model to generate an action sequence for controlling the actuator;
[0065] In this embodiment, inputting the integrated features into the action generation model is an important step to achieve automated decision-making and execution. The integrated features are usually a combination of visual information and task instructions extracted from the features, which are used to guide the system to generate action sequences that can adapt to different environments and task requirements. The core role of the action generation model is to generate an action sequence for controlling the actuator based on the input integrated features, ensuring that the actuator can perform the task according to the design goal.
[0066] The action generation model usually adopts a deep learning-based method, especially reinforcement learning or supervised learning models such as deep neural networks (DNN), convolutional neural networks (CNN), or recurrent neural networks (RNN). In this process, the integrated features are input into the model as input data, and the model processes the input data through internal network layers to generate the corresponding action sequence. The action sequence is a specific operation instruction for the actuator, which may include joint angles, velocities, accelerations, and other physical motion parameters.
[0067] The integrated features include visual information (such as object pose, position, shape, etc.) and task instruction information (such as task goals, operation steps, time limits, etc.). Through the combination of these features, the action generation model can better understand the environment and task requirements, and then make reasonable action choices.
[0068] In a specific implementation, the fused features are input to the first layer (e.g., fully connected layer or convolutional layer) of the action generation model, and then important information in the input features is extracted through layer-by-layer processing of the network. The action generation model generates a predicted action sequence based on historical data and the current input features. For example, the action sequence of a robot grasping operation can include moving the robot arm from the current pose to the target position, performing the grasping action, adjusting the grasping force, etc. The output of the model is the action sequence, which will be passed to the actuator as control instructions to perform the actual physical operation.
[0069] In a specific implementation, the fused features can be processed in various forms. A common way is to combine the visual features and task instructions through concatenation or weighted sum. The action generation model after that is usually a deep neural network, which can be a convolutional neural network (CNN) for processing visual data or a recurrent neural network (RNN) for processing time series data. After the fused features are input, the model generates an action sequence through a series of nonlinear transformations (e.g., ReLU activation, fully connected layer, etc.). Finally, the action sequence output by the model is further converted into specific instructions to control the actuator. For example, in a robot arm grasping task, the action generation model receives the object position information extracted from the image and the target object information from the task instructions, and generates a series of motion instructions for the robot arm joints through internal calculations to ensure the successful completion of the grasping task.
[0070] Example: In the field of embodied intelligence, for example, on an industrial production line, the robot arm obtains image data of the operating environment through a camera and obtains task instructions through a sensor. The system extracts visual features such as the shape, color, and position of the object from the image data, and extracts description information of the target object from the task instructions. After fusing these two types of features, the action generation model generates an action sequence that includes specific grasping actions that the robot arm needs to perform, such as moving, grasping, and fixing.
[0071] In the field of medical health, a surgical robot needs to perform precise operations according to surgical images and doctors' instructions. Visual features such as the position and pose of surgical instruments are extracted from image data, and task target information (e.g., grasping a specific tool) is extracted from the doctor's instructions. After fusing these two types of features, the system generates an action sequence to control the surgical robot, including operations such as instrument grasping, moving, and adjusting the position to ensure the smooth progress of the surgery.
[0072] In the field of financial technology, intelligent devices can perform processing tasks on financial documents based on image data and instructions. By extracting key content from the document, such as contract terms, amounts, and other information, and combining it with specific task requirements in the instructions (such as reviewing contracts, extracting information), the system generates a sequence of actions that the device controls to perform, completing tasks such as document processing, information extraction, and review.
[0073] This embodiment generates precise action sequences based on real-time environmental and task requirements by inputting fused features into the action generation model, greatly improving the efficiency and adaptability of intelligent devices in complex tasks. By fusing visual information and task instruction information, the system's perception and decision-making capabilities are enhanced, allowing the device to efficiently perform tasks in a changing environment. Ultimately, the system not only handles complex scenarios that traditional methods struggle with, but also improves the accuracy and generalization of task completion.
[0074] S50, controlling the actuator to perform operations in the operating environment according to the action sequence;
[0075] In this embodiment, the actuator is controlled by the action sequence to perform operations in the environment. The actuator usually refers to devices such as robotic arms, robots, etc., and the action sequence serves as a control signal to drive the actuator to perform specific physical operations. These operations include but are not limited to object grasping, carrying, placing, assembling, etc. According to the generated action sequence, the actuator can complete the target task in a changing operating environment.
[0076] The action sequence is usually output by the action generation model and contains various action instructions required to control the actuator to perform tasks, such as joint angles, velocities, accelerations, etc. of the robotic arm. These parameters are usually transmitted through the driver and servo system connected to the control system. Specifically, the action sequence can be generated in the following ways:
[0077] Joint angle and position control: According to task requirements, the action sequence indicates the target angle and position of each joint to ensure that the actuator moves along the predetermined trajectory.
[0078] Speed and acceleration control: Control the speed and acceleration of the actuator during movement to ensure that the task is completed without damaging objects or the environment.
[0079] Force and feedback control: During execution, the action sequence may need to be combined with an external force feedback system to ensure the stability and reliability of the actuator in complex tasks such as precise grasping.
[0080] In actual applications, when the action sequence is transmitted to the actuator, it usually goes through the following steps:
[0081] Action Sequence Parsing: The action sequence is first transmitted to the control system of the actuator, which parses each action instruction (such as joint angle, speed requirement, etc.) in it.
[0082] Motion Planning and Execution: Based on the parsed action instructions, the control system calculates the motion trajectory of each joint of the robot arm or other actuators. This calculation process usually relies on inverse kinematics (IK) algorithms to determine the target position of each joint, ensuring that the end effector can reach the target position.
[0083] Action Execution and Monitoring: The actuator moves according to the planned trajectory. During the process, the sensors of the actuator will monitor the execution effect in real time, ensuring that each joint moves according to the predetermined trajectory and can make timely adjustments in the event of collision or error.
[0084] In implementation, the action sequence is usually sent to the drive module of the actuator through the embedded control system, and these drive modules convert the action sequence into corresponding control instructions such as current and signals to drive the motor of the actuator to perform the corresponding action. The control system also optimizes according to the complexity of the task, such as through multiple iterations to ensure the successful completion of the task. In addition, the system can also introduce sensors and feedback mechanisms to ensure the accuracy and stability of the execution process.
[0085] In specific application scenarios, such as robot arm grabbing and handling tasks, the actuator not only executes the predetermined trajectory, but also needs to handle dynamic changes in the environment, such as obstacle avoidance and adjustment of grabbing force. These changes are usually captured by sensors and fed back to the control system, which then updates the action sequence or adjusts the operation.
[0086] Example Explanation: In the field of embodied intelligence, for example in manufacturing, the robot arm needs to perform grabbing tasks according to task instructions and environmental data. By inputting the generated action sequence, the robot arm can flexibly adjust its actions, such as grabbing different types of objects, avoiding obstacles, and adjusting the grabbing force. Each action sequence indicates how the robot arm moves each joint to ensure accurate grabbing and handling of objects.
[0087] In the field of medical health, surgical robots need to perform precise grabbing and operation tasks in complex surgical environments. By inputting the action sequence, the surgical robot can automatically perform various operations such as instrument grabbing, position adjustment, tool exchange, etc. according to image data and task instructions. Each action sequence includes precise motion parameters of the joints to ensure smooth operation of the surgery.
[0088] In the field of financial technology, intelligent robots need to process and identify relevant data for various financial products. When performing data processing tasks, the robot generates a sequence of actions to control the actuators by obtaining image data and task instructions. For example, the robot grabs specific files according to the task instructions, analyzes the contract content and organizes it, and generates an operation sequence for each step to ensure accurate task execution.
[0089] This embodiment can automatically execute tasks in complex operating environments by controlling actuators according to the action sequence, reducing manual intervention and improving the efficiency and accuracy of task execution. Especially in the face of complex and variable environments, the actuators can flexibly respond and adjust the operation strategy in a timely manner to ensure the smooth completion of the task. This approach can effectively improve the generalization ability and adaptability of the robot and enhance the execution effect of the system in various dynamic environments.
[0090] S60, obtaining the execution result of the operation to generate a feedback signal, and adjusting the parameters of the action generation model based on the feedback signal.
[0091] In this embodiment, the action generation model is optimized by monitoring feedback information in real time during operation execution. The generation and use of feedback signals are the core mechanisms of reinforcement learning systems, and the quality of feedback signals directly determines the efficiency and accuracy of model updates. Feedback signals usually come from the actual execution results of the actuators and the perception data of the sensors, aiming to evaluate the effectiveness of the task completion of the actuators in the operating environment and respond to problems that occur during the execution process.
[0092] The execution result of the operation is collected in real time by the sensor system and other monitoring devices of the actuator. These sensors may include position sensors, force sensors, vision sensors, etc., which can provide accurate task execution information. Specifically, the execution result can be the following types of data:
[0093] Position and attitude feedback: confirms whether the actuator successfully reaches the target position and whether it follows the predetermined trajectory.
[0094] Action success flag: for example, whether the object is successfully grabbed, whether it is successfully placed, or whether the task such as carrying in an industrial robot is successfully completed.
[0095] Execution quality evaluation: evaluates the quality of task completion, such as whether the grabbing force is appropriate, whether the object is damaged, etc.
[0096] The actuators obtain data through these sensors and transmit the operation results to the control system. The control system generates a feedback signal for the execution result based on these data, and the strength and type of the signal reflect the quality of task completion.
[0097] Based on the execution results of the operations, the system needs to generate feedback signals. These feedback signals are usually quantitative or qualitative evaluations of the operation results. For example:
[0098] Success and failure flags: Mark whether the task is completed (e.g., whether the grabbing is successful).
[0099] Reward or penalty signals: Assign a reward value or a penalty value to the execution result according to the quality of task completion. For example, the success of grabbing an object gives a positive reward, while failure may give a negative reward.
[0100] Precision feedback: If the task involves precision requirements (such as precise grabbing), the system will adjust the value of the feedback signal according to the degree of deviation.
[0101] The feedback signal can be a simple binary signal (success / failure) or a continuous value (such as a score of task precision). The generation of feedback signals is based on the evaluation of the control system on the execution result and the predetermined evaluation criteria.
[0102] The generated feedback signal is used to update the parameters of the action generation model. The action generation model is usually based on the principle of reinforcement learning, and by optimizing the parameters of the model, the accuracy and efficiency of subsequent operations are improved. Specifically:
[0103] State value function update: Calculate the state value function (such as Q value) according to the feedback signal, which is used to measure the expected return of taking a certain action in a certain state.
[0104] Gradient calculation and update: By calculating the gradient of the feedback signal, the system can use gradient descent or other optimization algorithms to adjust the parameters of the action generation model to improve the prediction accuracy and generalization ability of the model.
[0105] Incremental learning: In order to adapt to new tasks or environmental changes, the model can use incremental learning methods to gradually update its parameters without retraining the entire model.
[0106] This process is achieved through reinforcement learning or deep reinforcement learning methods. The model adjusts its parameters based on the continuously updated feedback signals, so that it can gradually improve its decision-making ability during execution, achieving higher execution accuracy and efficiency.
[0107] During implementation, the system first obtains the results of the operation through the actuator's sensors and transmits these results to the central control system. The control system evaluates the execution results using pre-defined rules and generates corresponding feedback signals. The feedback signal can be a numerical value representing the degree of success of the task execution or a binary value indicating whether the task was successful. Based on these feedback signals, the system uses optimization algorithms (such as Q-learning or other reinforcement learning algorithms) to calculate updates to the model parameters, thereby updating the parameters of the action generation model.
[0108] In its implementation, the system calculates the gradient of the value function based on the feedback signal from the current task and uses this gradient to adjust the parameters of the action generation model. This way, as more tasks are executed and feedback accumulates, the model can continuously learn and adapt, optimizing its decision-making process and ultimately achieving more precise and efficient control.
[0109] Example: In the field of embodied intelligence, robots may need to perform a variety of tasks in complex environments, such as grasping, carrying, and stacking. By obtaining feedback signals from task execution, the system can adjust its motion generation model and optimize grasping strategies to ensure efficient and accurate operation in different environments. For example, if the robot fails to grasp the target object in a task, the feedback signal may indicate failure. The system will use this feedback to adjust its strategy to improve the success rate for the next execution.
[0110] In the healthcare field, surgical robots must perform tasks in dynamically changing environments, such as precisely grasping surgical instruments or pathological specimens. By acquiring real-time feedback signals from task execution, the system can adjust its operating strategy to ensure accurate grasping and positioning of instruments. For example, if the surgical environment changes (such as the type or shape of the surgical instrument), the system can automatically adjust its grasping mode based on the feedback signals to ensure a smooth operation.
[0111] In the fintech sector, intelligent robots may be tasked with processing complex financial documents or providing customer service. By acquiring real-time feedback signals from task execution, robots can continuously optimize their decision-making models, improving processing efficiency and accuracy. For example, when handling a financial product crawling task, if the system fails to correctly identify or classify a product, a feedback signal will indicate task failure. The system will then adjust its recognition strategy based on this signal to improve accuracy and responsiveness in subsequent tasks.
[0112] The embodiment obtains the execution result of the operation and adjusts the parameters of the action generation model based on the feedback signal, so that the system can be self-optimized and gradually improve the success rate and precision of task execution. This adaptive mechanism enables the system to cope with tasks in different environments and improves its adaptability and generalization ability in dynamic environments. By continuously adjusting the model parameters, the system can learn more effective action strategies and significantly improve the task completion quality of the actuator.
[0113] The application relates to the technical field of artificial intelligence, can be applied to business scenarios such as embodied intelligence, financial technology and medical health, and discloses an operation control method and device based on multi-modal information fusion, equipment and a medium, which comprises the following steps: acquiring image data of an operation environment and a task instruction, extracting visual features in the image data and language features in the task instruction, fusing the visual features and the language features to generate fusion features, inputting the fusion features into an action generation model to generate an action sequence, controlling an actuator to execute an operation in the operation environment according to the action sequence, acquiring an execution result to generate a feedback signal, and adjusting parameters of the action generation model based on the feedback signal. The application fuses visual and language features, improves the adaptability of an operation control system to complex environmental changes and the multi-task processing capability, so that more accurate and intelligent operation control can be realized, and the generalization ability and execution efficiency of the system are significantly improved, especially when facing diversified tasks and dynamic environments.
[0114] In one embodiment, the above step S20 comprises:
[0115] S201, acquiring an original image of an operation environment through an image sensor, and generating a denoised image by denoising the original image;
[0116] S202, performing geometric correction on the denoised image to generate a corrected image, and extracting image data of the operation environment from the corrected image;
[0117] S203, receiving an operation instruction in the form of text through an instruction input interface;
[0118] S204, performing syntax analysis on the operation instruction in the form of text to generate an analyzed instruction, and extracting a task instruction describing an operation task from the analyzed instruction.
[0119] In this embodiment, obtaining image data of the operating environment and task instructions is the first step to achieve intelligent device control, involving image data collection, processing, and task instruction analysis. These steps provide necessary information input for subsequent operation execution. Image sensors are responsible for acquiring raw image data in the operating environment. Raw images are often affected by noise, blur, or environmental lighting, resulting in poor image quality. Denoising processing is used to reduce noise interference in images. Common methods include median filtering, mean filtering, bilateral filtering, and other techniques. These methods can effectively remove small noise in images while preserving image edges and texture features as much as possible, generating clear and accurate denoised images.
[0120] After denoising processing, image data usually needs to be geometrically corrected to address image distortion or distortion problems. Geometric correction is achieved by performing perspective transformation or affine transformation on images to ensure that objects in the image are consistent with their actual positions in the real environment. Through this step, the shape and position of objects in the image are corrected to ensure their correspondence with the actual operating environment, providing an accurate basis for subsequent image feature extraction. The corrected image is used to extract key visual features in the operating environment, including object position, size, shape, and other important information, which are crucial for subsequent action generation and task execution.
[0121] Complementary to image data extraction is the process of receiving and analyzing task instructions. Task instructions are usually input in text form and received through an instruction input interface. These instructions contain descriptions of the operating task, such as task objectives, execution methods, etc. To convert task instructions into a machine-understandable format, the system needs to parse the text. The goal of syntax parsing is to identify key elements in the text and separate the various components of the task, such as target objects, action types, and action sequences. Through this process, text-based task instructions are converted into standardized parsed instructions, providing clear operation objectives and steps for subsequent control systems.
[0122] After extracting and analyzing task instructions, the system can accurately understand and execute user-set operating tasks. The environmental information obtained from image data and the operating task information extracted from task instructions are combined through fusion technology, providing comprehensive perception and decision support for subsequent action generation and execution. This series of processing steps ensures that the system can complete tasks in a variable environment, enhancing the robustness and flexibility of the system.
[0123] The embodiment significantly improves the environmental perception ability of the system and the accuracy of task execution by detailed processing and extraction of image data and task instructions of the operating environment. The denoising processing and geometric correction preserve the quality and accuracy of the image data, providing a reliable data foundation for subsequent visual feature extraction. Through accurate task instruction analysis, the system can accurately understand the user's operation intention and convert it into executable control signals. The system can effectively cope with dynamically changing environments, improving the reliability and flexibility of task execution, thereby enhancing the adaptability and generalization ability of intelligent devices in complex environments.
[0124] In one embodiment, the above step S30 comprises:
[0125] S301, input the image data into a visual feature extraction network to generate initial visual features;
[0126] S302, perform spatial pooling processing on the initial visual features to generate pooled visual features;
[0127] S303, perform dimension reduction processing on the pooled visual features to generate reduced dimension visual features, and extract visual features from the reduced dimension visual features;
[0128] S304, input the task instruction into a language feature extraction network to generate initial language features;
[0129] S305, perform context encoding on the initial language features to generate context language features;
[0130] S306, perform normalization processing on the context language features to generate normalized language features, and extract language features from the normalized language features.
[0131] In the embodiment, first, the image data is input into a visual feature extraction network to generate initial visual features. The visual feature extraction network generally uses a convolutional neural network (CNN), which can extract low-level features of an image from the original image data, such as edges, textures, colors, etc. In this stage, the high-dimensional data of the image is abstracted by the neural network to generate initial visual features, which are the basis for subsequent analysis.
[0132] Next, the generated initial visual features are subjected to spatial pooling processing. Pooling operation reduces the dimension of feature map through down-sampling, reducing the spatial size of features, thereby improving the calculation efficiency, while preserving important spatial information in the image. In this stage, the pooling operation usually uses the method of maximum pooling or average pooling to extract the most representative features from multiple local regions. The pooled feature map is called pooled visual features.
[0133] Then, the pooled visual features are dimensionally reduced to generate reduced visual features. Dimensional reduction generally uses techniques such as principal component analysis (PCA), t-SNE, etc. to reduce the dimensionality of the features, avoid the computational complexity caused by high-dimensional data, and remove redundant information, so that the subsequent features are more concise and easy to process. Finally, the system extracts meaningful visual features from the reduced visual features, which contain the main information of the image, such as the shape, color, and position of the object, and can be effectively applied in subsequent tasks.
[0134] For language feature extraction in task instructions, first input the text form of the task instruction into the language feature extraction network. The language feature extraction network is usually based on a natural language processing (NLP) model, such as a long short-term memory network (LSTM) or a transformer (Transformer), which aims to extract grammatical and semantic information from the instruction to generate initial language features. These initial language features embody the basic meaning of the task instruction.
[0135] Next, the initial language features are subjected to context encoding processing. Context encoding is mainly to understand the relative relationship between words in language, and through the encoder to analyze the context of each word, generate context language features that can better express the connotation of language. This step can make the system understand the long-range dependency and context information in language, and improve the accuracy of language understanding.
[0136] Finally, normalization processing standardizes the context language features to ensure that the numerical range of the features is consistent, avoiding instability in the training process due to numerical differences. After normalization, the system can extract the final language features from the normalized language features for subsequent task decision-making and execution.
[0137] This embodiment extracts visual features from image data and language features from task instructions, so that the system can better understand the requirements of the task and the changes in the environment. The fusion of visual features and language features enables the intelligent device to fully consider the relationship between environmental information and task instructions when executing tasks, improving the accuracy and flexibility of task execution. This process optimizes the perception ability of the system, enabling it to cope with dynamic changes in the environment and complex task requirements, thereby improving the generalization ability and adaptability of the system, especially in complex scenarios, showing higher robustness.
[0138] In one embodiment, the above step S40 comprises:
[0139] S401, input the fusion features into the encoding network of the action generation model to generate encoding features;
[0140] S402, a diffusion process is performed on the encoded feature to generate a noisy feature by gradually adding Gaussian noise;
[0141] S403, an inverse diffusion process is performed on the noisy feature by a denoising network to generate a denoised feature;
[0142] S404, an action sequence decoding operation is performed based on the denoised feature to generate a candidate action sequence;
[0143] S405, an action sequence with the highest prediction success rate is selected from the candidate action sequences as the action sequence for controlling the actuator.
[0144] In this embodiment, first, the fusion feature is input to the encoding network of the action generation model, and the purpose of this stage is to convert the image and language features into encoded features that can be processed by the action generation model. The encoding network usually adopts a deep neural network architecture, such as a convolutional neural network (CNN) or a transformer network (Transformer), which maps the input fusion feature to a low-dimensional feature space, preserving the key features of the input information. The core of this process is to capture the common information of image data and language instructions to provide a unified representation for subsequent action generation.
[0145] Next, the encoded feature will perform a diffusion process, the core of which is to generate a noisy feature by gradually adding Gaussian noise to the encoded feature. The purpose of the diffusion process is to expand the distribution range of the original data by adding noise, so that the subsequent denoising process can more effectively recover useful information. The addition of Gaussian noise simulates the uncertainty or disturbance of the data, thereby increasing the robustness and adaptability of the model to environmental changes.
[0146] After generating the noisy feature, the system will perform an inverse diffusion process on these noisy features through a denoising network to generate a denoised feature. The purpose of the inverse diffusion process is to recover the true information in the data by removing noise. The denoising network is usually based on a deep convolutional neural network or a generative adversarial network (GAN), which aims to train the generation network to gradually remove noise and restore clearer and more stable feature information. Through the inverse diffusion process, the network can eliminate external disturbances or information loss, improving the accuracy of the subsequent generation results.
[0147] The generated denoised feature will be used to perform an action sequence decoding operation. The purpose of the decoding process is to generate a candidate action sequence based on the denoised feature. The action sequence decoder usually uses a recurrent neural network (RNN), a long short-term memory network (LSTM), or a transformer network (Transformer) to generate an action sequence. The decoding network infers the possible action sequence based on the input denoised feature, and these candidate action sequences can correspond to a series of actions performed by the actuator in the operating environment.
[0148] Finally, the system selects a candidate action sequence with the highest predicted success rate from the generated sequences to control the specific actions of the executor. The process of selecting an action sequence relies on a success rate prediction mechanism that selects the most likely successful action by calculating the success probability of each candidate action sequence under specific environmental conditions. The prediction of success rate is usually achieved through reinforcement learning algorithms, which estimate the success probability of each action based on historical execution results, thereby achieving dynamic optimization.
[0149] This embodiment effectively enhances the adaptability of the action generation model in complex environments through diffusion and inverse diffusion processes. By combining visual and language features, the system can more comprehensively understand the operation task and environmental information, generating high-quality action sequences. This process not only improves the accuracy of action generation, but also has strong generalization ability in diverse operating environments. By selecting the action sequence with the highest predicted success rate, the system ensures high efficiency and success rate in executing tasks, thereby improving the execution ability of intelligent devices in complex scenarios.
[0150] In one embodiment, the above step S50 includes:
[0151] S501, parsing the action sequence to obtain joint angle parameters, joint velocity parameters, and end effector pose parameters of the executor;
[0152] S502, determining joint motion trajectories based on the joint angle parameters and joint velocity parameters;
[0153] S503, generating end effector path planning of the executor based on the end effector pose parameters of the executor;
[0154] S504, converting the joint motion trajectories and the end effector path planning of the executor into motor control instructions;
[0155] S505, sending the motor control instructions to the control interface of the executor driver;
[0156] S506, driving each joint motor of the executor synchronously through the driver control signal;
[0157] S507, controlling the end effector of the executor to perform operations in the operating environment along the end effector path planning.
[0158] In this embodiment, first, the step of parsing the action sequence provides key data for subsequent control by obtaining joint angle parameters, joint velocity parameters, and end effector pose parameters of the end effector in the action sequence. The joint angle parameters indicate the target angular position of each joint of the robot arm, the joint velocity parameters specify the desired motion speed of each joint, and the pose parameters of the end effector provide the position and orientation of the end effector in space. These data can be extracted through a sensor input interface or from the action sequence generated in the previous link, ensuring that the control system can understand and parse the motion target.
[0159] Based on the acquisition of joint angles and joint velocities, the system uses these parameters to determine the joint motion trajectory. The joint motion trajectory refers to the motion path of each joint from the current state to the target state, which is usually calculated by a motion planning algorithm (such as the shortest path algorithm or quintic polynomial interpolation), ensuring that the joint reaches the specified position within the specified time and meets the dynamics constraints of the robot arm. Motion trajectory planning not only includes joint position, but also needs to ensure that the velocity and acceleration of the joint meet the physical constraints of the robot arm, avoiding fast movement or non-smooth motion.
[0160] Next, based on the pose parameters of the end effector, the system generates the end effector path planning of the actuator. The purpose of end effector path planning is to ensure that the end effector of the robot arm (such as gripper, welding gun, etc.) can accurately move along the predetermined path in space to the target position. In this process, the system uses path planning algorithms (such as A* algorithm or spline curve) to calculate the path of the end effector according to the operation task requirements and current environmental conditions, and optimizes the adjustment according to dynamic environmental data. This step ensures that the end effector can avoid obstacles and smoothly reach the target when performing tasks.
[0161] After the joint motion trajectory and end effector path planning are generated, the system converts these information into motor control instructions. Motor control instructions are signals that drive the motors of each joint of the actuator, which usually include control of motor rotation angle, speed, acceleration, etc. Control instructions will combine inverse kinematics solution and path planning results, taking into account the dynamics and constraints of each joint, to accurately control the motion of each joint of the robot arm. These instructions can be transmitted to the driver interface of the actuator through common industrial protocols such as PWM (Pulse Width Modulation) signal or CANopen protocol.
[0162] After generating the motor control instructions, the system sends them to the control interface of the actuator driver. The main function of the driver is to receive instructions from the control system and convert them into signals suitable for motor driving. The design of the control interface ensures the accuracy and timeliness of instruction transmission, ensuring that the motor can respond and adjust in real time according to the instructions.
[0163] Next, the actuator driver drives the motors of each joint in synchronization according to the received control signals. Synchronization refers to the simultaneous activation of multiple joint motors in a predetermined time and sequence to ensure the coordination of the robot arm's movements. The driver control signal needs to have sufficient precision to ensure that the movement of each joint is smoothly executed within the given trajectory, thereby ensuring the coordination and stability of the overall movement.
[0164] Finally, the end effector of the actuator performs operations along the predetermined path in the operating environment. The movement of the end effector follows the pre-generated path planning and joint motion trajectory, and can complete tasks such as object grasping, welding, assembly, etc. according to task requirements. This process needs to ensure that the end effector not only accurately performs tasks, but also flexibly adapts to unexpected situations in dynamic environments, avoiding interference caused by obstacles or changes.
[0165] This embodiment ensures efficient coordinated movement and task execution of the robot arm in complex operating environments through accurate path planning, joint motion trajectory generation, and effective conversion of motor control instructions. Through meticulous joint and end effector control, the adaptability of traditional systems in complex environments can be effectively addressed, significantly improving the movement precision and operation reliability of the robot arm. At the same time, the robot arm can accurately perform various tasks in a variable environment, improving the success rate of task completion and enhancing the intelligence level and applicability of the robot arm.
[0166] In one embodiment, the above step S60 comprises:
[0167] S601, monitoring the execution result of the operation to obtain execution state data;
[0168] S602, identifying the operation success flag and abnormal event flag in the execution state data;
[0169] S603, matching a preset reward strategy based on the operation success flag and the abnormal event flag to generate a quantitative value;
[0170] S604, converting the quantitative value into a feedback signal;
[0171] S605, determining the state value function gradient of the action generation model based on the feedback signal;
[0172] S606, adjusting the network weight parameters of the action generation model using the state value function gradient.
[0173] In this embodiment, first, the monitoring operation obtains execution state data. This step relies on sensors and feedback mechanisms to collect state information in real time when the robot arm performs tasks. These state data can include the pose, speed, acceleration, etc. of the robot arm, and can also include the success or failure flag of the operation. For example, in a grasping task, the sensor can feedback whether the object is successfully grasped or whether the actuator reaches the predetermined position. Through these execution state data, the operation of the robot arm can be comprehensively monitored and analyzed to ensure that the execution state of the task is timely grasped.
[0174] Next, identify the operation success flag and abnormal event flag in the execution state data. By analyzing the monitored execution state data, the system can identify whether there is an operation success or abnormal situation. The operation success flag indicates that the task is completed successfully as expected, while the abnormal event flag represents that the system encounters problems during task execution, such as grasping failure, object loss, path blockage, etc. For each task, the system will judge whether there is an abnormal event according to the set standard or tolerance, and generate the corresponding flag. This analysis step is critical because it provides the basis for further decision-making of the system.
[0175] After identifying the operation success flag and abnormal event flag, the next step is to match the pre-set reward strategy based on these flags to generate a quantitative value. The reward strategy is a set of pre-defined rules for measuring the success of each executed task. For example, if the operation is successful, the system may generate a positive quantitative value (e.g. +1); if an abnormal event occurs, the system may generate a negative quantitative value (e.g. -1). This quantitative value is an important basis for the system to evaluate the operation result, and is used in the subsequent learning and optimization process.
[0176] Then, convert the quantitative value into a feedback signal. After this conversion, the quantitative value becomes a feedback signal, which will be used as input for model optimization. Through the feedback signal, the system can evaluate the pros and cons of task execution, thereby providing decision-making basis. The feedback signal not only helps the system understand the past operation performance, but also provides learning signals for optimizing the model, promoting the continuous improvement of the system.
[0177] Based on the feedback signal, generate the state value function gradient of the action generation model. In this process, the feedback signal will be used to calculate the gradient of the model, which reflects the difference in the performance of the model in the current state. By calculating the gradient of the state value function, the system can measure the gap between the current strategy and the optimal strategy, and provide necessary information for the subsequent optimization process. The calculation of the state value function gradient relies on the Bellman equation in reinforcement learning, which can provide the learning direction and intensity of the model.
[0178] Finally, the network weight parameters of the action generation model are adjusted using the state value function gradient. By applying gradient information, the system can update the parameters of the action generation model. The updated model can better adapt to environmental changes, improving the accuracy and stability of task execution. This process usually uses optimization algorithms such as gradient descent to adjust network weights, allowing the system to gradually optimize performance through continuous learning, ultimately achieving a more intelligent and efficient decision-making process.
[0179] This embodiment accurately evaluates the results of operation tasks and optimizes in real time according to feedback information by implementing a state monitoring and feedback mechanism. By generating feedback signals and adjusting the parameters of the action generation model, continuous improvement of robotic arm operation can be achieved, thereby improving the success rate and adaptability of task execution. Effectively solves the problem that traditional systems cannot adjust strategies in real time according to environmental changes, has stronger intelligent decision-making and learning ability, and helps to improve the performance and reliability of automated operation systems in complex environments.
[0180] In one embodiment, the above step S60 further comprises:
[0181] S701, monitoring sensor data of the operating environment to obtain the current environmental state;
[0182] S702, comparing the current environmental state with the historical environmental state to generate environmental change features;
[0183] S703, judging environmental change events based on the environmental change features;
[0184] S704, in response to the environmental change events, collecting new visual information and new task instructions;
[0185] S705, processing the new visual information and new task instructions to generate model update data;
[0186] S706, applying the model update data to adjust the network weight parameters of the action generation model.
[0187] In this embodiment, first, the sensor data of the operating environment is monitored to obtain the current environmental state. This step involves collecting real-time sensor data from the operating environment to accurately describe the current state of the environment. Sensors may include cameras, lidar, ultrasonic sensors, etc., which provide multi-dimensional information about the environment. For example, image sensors provide visual data of the environment, distance sensors provide changes in object distance, temperature and humidity sensors provide temperature and humidity information of the environment, etc. By integrating the data from these sensors, the system can comprehensively understand the current state of the operating environment.
[0188] The environment change features are generated by comparing the current environment state with the historical environment state. This step compares the current environment state with the historical records to identify what changes have occurred in the environment. By calculating the differences between the current environment state and the historical environment state, the system can derive "environment change features" that reflect changes in object positions, addition or disappearance of obstacles, changes in environmental conditions, etc. The extraction of differentiated features is crucial for understanding the dynamic changes of the environment and responding to them.
[0189] Based on the environment change features, the system determines whether an environment change event has occurred. By analyzing and judging the environment change features, the system can identify whether a significant environment change event has occurred. For example, the environment change features may indicate that an object has been moved, a target has been lost, or a new obstacle has appeared in the work area. Based on these change features, the system determines whether a corresponding environment change event has been triggered, and then decides whether to adjust the decision-making process of the action generation model.
[0190] In response to the environment change event, new visual information and new task instructions are collected. In response to the detected environment change event, the system needs to take appropriate measures to adapt to the changes. Specifically, the system will re-collect new visual data (such as images of new target objects) in response to these change events, and update the task instructions (such as adjusting the grasping strategy, path planning, etc.) according to the changes. The collection of new visual information and task instructions is crucial to ensure that the system can adapt to the dynamic changing environment.
[0191] The new visual information and task instructions are processed to generate model update data. This step involves processing the newly collected visual information and task instructions to generate model update data that adapts to the current environmental changes. The processing of new visual information usually includes feature extraction, image analysis, target recognition, etc., while the processing of task instructions includes parsing and adjusting task objectives, etc. Through these processes, the system can generate new update data to provide input to the action generation model.
[0192] The model update data is applied to adjust the network weight parameters of the action generation model. Finally, by inputting the generated model update data into the action generation model, the system can adjust its network weight parameters. This process is usually optimized through a backpropagation algorithm, which updates the decision parameters of the model based on the newly obtained data. The updated model can better adapt to the current environmental changes, improving the success rate and precision of task execution.
[0193] Example Explanation: In the field of embodied intelligence, intelligent devices need to perform tasks in complex and dynamic operating environments. Traditional fixed programs and visual feedback mechanisms often cannot adapt to the dynamic changes of the environment and the diversity of task requirements. By fusing visual perception and language instructions, a deep learning-based action generation model enables intelligent devices to acquire and process data in real-time environments, automatically adjust their execution strategies, and respond quickly to sudden changes.
[0194] First, the system collects raw image data of the operating environment through image sensors. This image data comes from high-precision cameras, depth cameras, or other visual devices, which can capture key objects and targets in the environment. Then, the image data is denoised to remove background noise and interference, generating a denoised image. This step is crucial to ensure the accuracy of subsequent processing.
[0195] Next, the denoised image data is adjusted through geometric correction. The purpose of geometric correction is to correct the perspective of the image, so that the position of the objects in the image conforms to the position relationship in the actual physical space. After the corrected image is generated, the system extracts the image data of the operating environment from the image, which will provide the basis for subsequent visual feature extraction.
[0196] At the same time, the task instructions are received through the instruction input interface, usually in the form of text instructions. The instruction input interface can be a speech recognition system, manual input, or instructions generated by an automated system. After receiving the task instructions, the system analyzes the text instructions through syntax parsing, extracting key information from them to generate parsed instructions. The parsed instructions can accurately describe the target and method of executing the task, providing clear guidance for task planning and execution.
[0197] The visual features extracted from the above image data and the language features extracted from the task instructions are input into the encoding network of the action generation model. The action generation model generates fusion features by fusing visual features and language features, which serve as the basis for generating action sequences. Visual features include object shape, position, texture, etc., while language features include task description, target instruction, etc. After fusing these features, the system can generate more accurate and reasonable action sequences to control the actuators to complete the task in the operating environment.
[0198] Based on the generated action sequence, the actuator begins to execute the task in the operating environment. The actuator adjusts its joint angles and speed according to the action sequence to complete the predetermined operation. During this process, the motion parameters and path planning data of the actuator are accurately calculated to ensure that each step in the execution process is optimized based on real-time feedback.
[0199] During task execution, the system monitors the operation results in real-time and obtains execution state data. Feedback information is obtained through sensors and other devices to analyze whether the operation is proceeding as expected. If the operation is successful, the system marks it as a success flag; if an abnormality occurs, such as a failed grasp or the object not moving to the predetermined position, it is marked as an abnormal event. At this time, the system matches the pre-set reward strategy according to the success flag and the abnormal event flag to generate a quantitative evaluation value. This evaluation value reflects the quality and success rate of the current operation, and the generated feedback signal will be used for subsequent model adjustment.
[0200] Finally, based on the generated feedback signal, the network weight parameters of the action generation model are adjusted. At this time, through the incremental learning algorithm and gradient descent method, the model fine-tunes the parameters to adapt to the changes in the current operation environment. With the continuous optimization of the model, the system can perform better in new tasks and environments, thereby improving the efficiency and success rate of task execution.
[0201] In the field of medical health, the application of intelligent device control methods plays a crucial role, especially in automated surgical systems, pathological specimen processing, and precise operation of medical devices. Traditional medical robots often rely on pre-set programs and rule systems, which cannot quickly adapt to dynamic patient conditions and environmental changes in complex surgical environments. In this system, intelligent devices can automatically generate operation actions based on real-time image data and task instructions, and optimize the action generation model based on feedback signals, thereby improving the precision and flexibility of medical operations.
[0202] First, the original image data of the operation environment is collected through image sensors, including key information such as the patient's physical condition and the surgical area. After noise removal, geometric correction, and other processing, clear and accurate patient images can be generated. Through the task input interface, the doctor's operation instructions (such as needing to remove a certain part or perform a certain operation) are received. These task instructions will be parsed into operation tasks that the machine can understand, providing a basis for subsequent action generation.
[0203] Next, the visual features (such as the shape, structure, and depth of the surgical area) and language features (such as surgical instructions and target positions) are fused through the action generation model of the intelligent device. The fused feature data will be used to generate an action sequence for controlling the robot to perform precise actions. This action sequence includes multiple action steps such as robot joint movement and gripper operation, and by controlling the robot to perform these actions, the doctor's set tasks can be accurately completed.
[0204] When the robot operates according to the generated action sequence, it is very important to monitor and obtain the execution results of the operation in real time. The intelligent device monitors the state of the robot operation through sensors, obtains the execution state data in real time, and judges whether the operation is successful. If the operation is successful, the system adjusts the parameters of the action generation model to the optimal state through the generated feedback signal, thereby ensuring the accuracy and efficiency of future operations. If there is an operation deviation (such as position deviation, improper clamping, etc.), the system will automatically generate an adjustment scheme and adjust the parameters of the action generation model through incremental learning, so that the robot can avoid similar problems in subsequent operations.
[0205] For example, during surgery, the robot needs to accurately grasp surgical tools and perform operations. By obtaining image data of the surgical area through image sensors and combining with the operation instructions input by the doctor, the robot can automatically perform complex surgical operations. If a sudden situation occurs during the operation, such as a shift in the target position, the system will adjust the model parameters in time to avoid operation errors and ensure the smooth progress of the operation. In addition, the system can also optimize based on historical operation data, so that the robot has stronger adaptability when facing different patients and operation situations. This can greatly improve the automation and accuracy of medical operations, reduce human operation errors, and especially in high-risk and high-complexity surgical procedures, can effectively ensure the safety of patients.
[0206] In the field of financial technology, the application of intelligent device control methods is also of great significance, especially in complex business scenarios such as financial transactions, automated customer service, and financial product processing. Traditional financial automation systems often rely on fixed rules and pre-set processes, which makes them lack flexibility and adaptability when facing changing financial products and complex business scenarios. This system obtains environmental information in real time, automatically generates operation actions, and optimizes the action generation model through the feedback mechanism, so that the financial technology system can quickly respond to changes in a dynamic environment and improve processing efficiency and accuracy.
[0207] First, image data of the operation environment is obtained through image sensors to capture key visual information related to financial transactions. For example, in an automated customer service system, the system may need to extract important information (such as identification, contracts, transaction credentials, etc.) from the customer's submitted documents. At the same time, the system receives task descriptions from customers or system administrators through task instructions, which may include operations such as processing applications for certain financial products, fund transfers, etc. Task instructions are accessed to the system through the instruction input interface, and the system performs syntax analysis on these instructions to extract task information related to financial business (such as transaction type, amount, transaction counterpart, etc.).
[0208] Next, the image data and task instructions are input into the action generation model of the intelligent device for processing. The model fuses image features and language features. Through this fused feature, the system can generate a corresponding action sequence for controlling the system to perform financial tasks. These actions include automatically filling out forms, initiating transaction requests, performing verifications, etc. The intelligent device controls the system to perform tasks related to financial operations according to the generated action sequence. For example, the system can automatically identify the file information uploaded by the customer and automatically complete operations such as fund transfer and account authentication based on the task instructions.
[0209] During execution, the system continuously monitors the execution results of the operations to obtain execution state data. This includes information such as whether the transaction is successful, whether the file processing is completed, etc. The system generates feedback signals based on these execution state data. If deviations are found during execution (such as unsuccessful fund transfer, account information mismatch, etc.), the system will automatically adjust the model parameters to ensure that subsequent operations are more accurate and efficient. This mechanism of adjusting model parameters through feedback signals enables the financial system to continuously optimize its operation performance in a complex and dynamic environment, improving the accuracy and reliability of the system.
[0210] For example, in a cross-border payment scenario, the system needs to identify payment information in different currencies and automatically calculate based on exchange rates. If the exchange rate changes, the system will adjust its calculation model through the feedback mechanism to ensure the accuracy of all calculations and conversions in the cross-border payment process. The generation and adjustment mechanism of feedback signals helps the system adapt to changes in the financial market and provide more accurate services.
[0211] This embodiment improves the adaptability and robustness of the system by dynamically monitoring changes in the operating environment and responding and adjusting the task execution strategy in a timely manner. It can automatically adjust the model according to real-time environmental changes to ensure efficient and accurate task execution in complex environments. By introducing environmental change detection in the feedback mechanism, the system can more flexibly cope with unknown and uncertain situations, enhancing the autonomous decision-making ability of the intelligent device in dynamic environments.
[0212] In an embodiment, a multi-modal information fusion-based operation control device is provided, which corresponds one-to-one to the multi-modal information fusion-based operation control method in the above embodiment. Referring to Figure 3 , Figure 3 A functional module diagram of a preferred embodiment of the multi-modal information fusion-based operation control device of the present application. Data acquisition module 10, feature extraction module 20, feature fusion module 30, action generation module 40, actuator control module 50, and feedback adjustment module 60. The detailed description of each functional module is as follows:
[0213] The data collection module 10 is configured to acquire image data of an operation environment and task instructions describing an operation task.
[0214] The feature extraction module 20 is configured to extract visual features from the image data and language features from the task instructions.
[0215] The feature fusion module 30 is configured to fuse the visual features and the language features to generate fused features.
[0216] The action generation module 40 is configured to input the fused features into an action generation model to generate an action sequence for controlling an effector.
[0217] The effector control module 50 is configured to control the effector to perform an operation in the operation environment according to the action sequence.
[0218] The feedback adjustment module 60 is configured to acquire an execution result of the operation to generate a feedback signal, and adjust parameters of the action generation model based on the feedback signal.
[0219] In an embodiment, the feature extraction module 20 is specifically configured to:
[0220] acquire a raw image of the operation environment through an image sensor, and generate a denoised image by performing denoising processing on the raw image;
[0221] perform geometric correction on the denoised image to generate a corrected image, and extract image data of the operation environment from the corrected image;
[0222] receive an operation instruction in a text form through an instruction input interface;
[0223] perform syntax analysis on the operation instruction in the text form to generate an analyzed instruction, and extract task instructions describing the operation task from the analyzed instruction.
[0224] In an embodiment, the feature fusion module 30 is specifically configured to:
[0225] input the image data into a visual feature extraction network to generate initial visual features;
[0226] perform spatial pooling processing on the initial visual features to generate pooled visual features;
[0227] perform dimension reduction processing on the pooled visual features to generate reduced-dimension visual features, and extract visual features from the reduced-dimension visual features;
[0228] input the task instructions into a language feature extraction network to generate initial language features;
[0229] contextually encode the initial language feature to generate a context language feature;
[0230] normalize the context language feature to generate a normalized language feature, and extract a language feature from the normalized language feature.
[0231] In an embodiment, the action generation module 40 is specifically configured to:
[0232] input the fusion feature into an encoding network of an action generation model to generate an encoded feature;
[0233] perform a diffusion process on the encoded feature to generate a noise feature by gradually adding Gaussian noise;
[0234] perform an inverse diffusion process on the noise feature through a denoising network to generate a denoised feature;
[0235] perform an action sequence decoding operation based on the denoised feature to generate a candidate action sequence;
[0236] select an action sequence with the highest prediction success rate from the candidate action sequence as an action sequence for controlling the effector.
[0237] In an embodiment, the effector control module 50 is specifically configured to:
[0238] parse the action sequence to obtain joint angle parameters, joint velocity parameters, and end effector pose parameters of the effector;
[0239] determine a joint motion trajectory according to the joint angle parameters and the joint velocity parameters;
[0240] generate an end effector path plan of the effector based on the end effector pose parameters of the effector;
[0241] convert the joint motion trajectory and the end effector path plan of the effector into motor control instructions;
[0242] send the motor control instructions to a control interface of an effector driver;
[0243] drive each joint motor of the effector to act synchronously through the driver control signal;
[0244] control the end effector of the effector to perform an operation in the operating environment along the end effector path plan.
[0245] In an embodiment, the feedback adjustment module 60 is specifically configured to:
[0246] monitor an execution result of the operation to obtain execution state data;
[0247] identifying an operation success flag and an abnormal event flag in the execution state data;
[0248] matching a preset reward strategy based on the operation success flag and the abnormal event flag, and generating a quantization value;
[0249] converting the quantization value into a feedback signal;
[0250] determining a state value function gradient of the action generation model based on the feedback signal;
[0251] adjusting network weight parameters of the action generation model using the state value function gradient.
[0252] In an embodiment, the feedback adjustment module 60 is specifically configured to:
[0253] monitoring sensor data of the operation environment to obtain a current environment state;
[0254] comparing the current environment state with a historical environment state to generate an environment change feature;
[0255] judging an environment change event based on the environment change feature;
[0256] in response to the environment change event, collecting new visual information and new task instructions;
[0257] processing the new visual information and the new task instructions to generate model update data;
[0258] applying the model update data to adjust network weight parameters of the action generation model.
[0259] In one embodiment, a computer device is provided, which can be a server, and an internal structure diagram thereof can be as shown in Figure 4 The computer device includes a processor, a memory, a network interface and a database connected through a system bus. The processor of the computer device is configured to provide determination and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium, an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the computer device is configured to communicate with an external user terminal through a network connection. The computer program is executed by the processor to implement the functions or steps of the server side of the operation control method based on multi-modal information fusion.
[0260] In one embodiment, a computer device is provided, which can be a user terminal, and an internal structure diagram thereof can be as shown in Figure 5As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of an operation control method based on multimodal information fusion.
[0261] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0262] acquiring image data of an operating environment and a task instruction for describing an operating task;
[0263] extracting visual features from the image data and extracting language features from the task instructions;
[0264] fusing the visual features and the language features to generate a fused feature;
[0265] Inputting the fused features into an action generation model to generate an action sequence for controlling an actuator;
[0266] controlling the actuator to perform operations in the operating environment according to the action sequence;
[0267] An execution result of the operation is obtained to generate a feedback signal, and parameters of the action generation model are adjusted based on the feedback signal.
[0268] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0269] acquiring image data of an operating environment and a task instruction for describing an operating task;
[0270] extracting visual features from the image data and extracting language features from the task instructions;
[0271] fusing the visual features and the language features to generate a fused feature;
[0272] Inputting the fused features into an action generation model to generate an action sequence for controlling an actuator;
[0273] controlling the effector to perform an operation in the operation environment according to the action sequence;
[0274] obtaining an execution result of the operation to generate a feedback signal, and adjusting a parameter of the action generation model based on the feedback signal.
[0275] It should be noted that the functions or steps described above with respect to the computer-readable storage medium or the computer device can correspond to the related descriptions of the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0276] A person of ordinary skill in the art can understand that all or part of the processes in the foregoing method embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a nonvolatile computer readable storage medium. When the computer program is executed, the computer program can include the processes of the foregoing method embodiments. Any reference to a memory, storage, database or other medium used in the embodiments provided in the present application can include nonvolatile and / or volatile memory. The nonvolatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. The volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0277] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is exemplified. In actual applications, the above functions can be completed by different functional units or modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0278] It should be explained that if the software tools or components of other companies appear in the embodiments of the present application, they are only used for example introduction and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present application, but not limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, it should be understood by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. An operation control method based on multimodal information fusion, characterized in that: The following steps are involved: acquiring image data of an operating environment and a task instruction for describing an operating task; extracting visual features from the image data and extracting language features from the task instructions; fusing the visual features and the language features to generate a fused feature; Inputting the fused features into an action generation model to generate an action sequence for controlling an actuator; controlling the actuator to perform operations in the operating environment according to the action sequence; An execution result of the operation is obtained to generate a feedback signal, and parameters of the action generation model are adjusted based on the feedback signal.
2. The operation control method based on multimodal information fusion according to claim 1, characterized in that: Acquire image data of the operating environment and task instructions for describing the operating task, including: Collecting an original image of the operating environment through an image sensor, and performing denoising processing on the original image to generate a denoised image; Performing geometric correction on the denoised image to generate a corrected image, and extracting image data of the operating environment from the corrected image; Receive operation instructions in text form through the instruction input interface; The text-based operation instruction is parsed to generate parsed instructions, and task instructions describing the operation task are extracted from the parsed instructions.
3. The operation control method based on multimodal information fusion according to claim 1, characterized in that: Extracting visual features from the image data and extracting language features from the task instructions includes: Inputting the image data into a visual feature extraction network to generate initial visual features; Performing spatial pooling processing on the initial visual features to generate pooled visual features; Performing dimensionality reduction processing on the pooled visual features to generate reduced-dimensionality visual features, and extracting visual features from the reduced-dimensionality visual features; Inputting the task instruction into a language feature extraction network to generate initial language features; Performing context encoding on the initial language features to generate contextual language features; Normalization processing is performed on the contextual language features to generate normalized language features, and language features are extracted from the normalized language features.
4. The operation control method based on multimodal information fusion according to claim 1, characterized in that: The fused features are input into the action generation model to generate an action sequence for controlling the actuator, including: Inputting the fused features into an encoding network of an action generation model to generate encoding features; performing a diffusion process on the coding feature to generate a noise feature by gradually adding Gaussian noise; Performing an inverse diffusion process on the noise feature through a denoising network to generate a denoised feature; Performing an action sequence decoding operation based on the denoising features to generate a candidate action sequence; An action sequence with the highest prediction success rate is selected from the candidate action sequences as the action sequence for controlling the actuator.
5. The operation control method based on multimodal information fusion according to claim 1, characterized in that: Controlling the actuator to perform an operation in the operating environment according to the action sequence includes: Analyzing the motion sequence to obtain joint angle parameters, joint velocity parameters, and end effector posture parameters of the actuator; determining a joint motion trajectory according to the joint angle parameter and the joint velocity parameter; generating an end-effector path plan for the actuator based on the end-effector pose parameters of the actuator; Converting the joint motion trajectory and the end effector path planning of the actuator into motor control instructions; Sending the motor control instruction to the control interface of the actuator driver; Drive the joint motors of the actuator to move synchronously through the driver control signal; An end effector of the effector is controlled to perform an operation in the operation environment along the end effector path plan.
6. The operation control method based on multimodal information fusion according to claim 1, characterized in that: Obtaining an execution result of the operation to generate a feedback signal, and adjusting parameters of the action generation model based on the feedback signal, including: Monitoring the execution result of the operation to obtain execution status data; Identifying an operation success flag and an abnormal event flag in the execution status data; Generate a quantitative value based on matching a preset reward strategy with the operation success flag and the abnormal event flag; converting the quantized value into a feedback signal; determining a state-value function gradient of the action generation model based on the feedback signal; The state value function gradient is used to adjust the network weight parameters of the action generation model.
7. The operation control method based on multimodal information fusion according to claim 1, characterized in that: After obtaining the execution result of the operation to generate a feedback signal, and adjusting the parameters of the action generation model based on the feedback signal, the method further includes: Monitoring sensor data of the operating environment to obtain a current environmental state; Comparing the current environmental state with the historical environmental state to generate an environmental change signature; determining an environmental change event based on the environmental change characteristics; In response to the environmental change event, new visual information and new task instructions are collected; Processing the new visual information and the new task instructions to generate model update data; The model update data is applied to adjust the network weight parameters of the action generation model.
8. An operation control device based on multimodal information fusion, characterized in that: The operation control device based on multimodal information fusion includes: A data acquisition module, used to obtain image data of the operating environment and task instructions for describing the operating task; a feature extraction module, configured to extract visual features from the image data and language features from the task instructions; A feature fusion module, configured to fuse the visual features with the language features to generate fused features; an action generation module, configured to input the fused features into an action generation model to generate an action sequence for controlling an actuator; An actuator control module, configured to control the actuator to perform operations in the operating environment according to the action sequence; A feedback adjustment module is used to obtain the execution result of the operation to generate a feedback signal, and adjust the parameters of the action generation model based on the feedback signal.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and an operation control program based on multimodal information fusion stored in the memory and capable of running on the processor. When the operation control program based on multimodal information fusion is executed by the processor, the steps of the operation control method based on multimodal information fusion as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The storage medium stores an operation control program based on multimodal information fusion, and when the operation control program based on multimodal information fusion is executed by the processor, the steps of the operation control method based on multimodal information fusion according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Power robot control method and system based on visual voice action model
CN121043156A
Robot control method, system, device, equipment and medium
CN121290446A
Robot control method and device, computer equipment and storage medium
CN121468574A
Open-ground amphibious robot cross-domain autonomous positioning method and device based on view cone transformation
CN122041849A
Cross-domain autonomous positioning method and device for air-ground amphibious robot based on frustum transformation
CN122041849B