Seabed operation robot advancing planning method and system based on visual language model
Through visual language model combined with track/sled chassis data, the seabed operation robot travel planning method is solved, and the traditional method is unstable in the seabed sparse soft soil environment is achieved, stable and efficient seabed operation is achieved, and data acquisition costs are reduced.
Patent Information
- Application Number
- CN202510564484.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-08
AI Technical Summary
Traditional robotic maneuver planning methods are difficult to adapt to complex and changeable subsea sparse soft soil environments, especially ignoring the numerical data of the track/sled chassis, resulting in unstable travel. The existing visual navigation methods are insufficient in multimodal data fusion and complex environmental decision-making, and data acquisition is expensive and risky.
The seabed operation robot travel planning method based on visual language model is adopted, and image information is obtained through the robot's front RGB camera, combined with the numerical data of the crawler/sled chassis, and the instruction fine-tuning is used to form closed-loop control to achieve stable travel.
It improves the adaptability and travel stability of the robot on the surface of the seabed thin soft soil, reduces the cost and risk of data acquisition, improves the efficiency of model training, and ensures the stable operation of the robot in complex environments.
Smart Images

Figure CN120449476A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of marine robotics technology, and more particularly to a method and system for seabed operation robot movement planning based on a visual language model. Background Art
[0002] With the continuous development of marine resources, the application of submarine operation robots in marine engineering is becoming increasingly widespread. However, the complex and changing seabed environment, especially the soft soil surface of the seabed, whose topography, water flow, marine life, and other factors pose severe challenges to the robot's movement stability. Traditional robot movement planning methods are usually based on preset environmental models and rules, which are difficult to adapt to the complex and changing seabed environment. Although existing visual navigation methods can use image information for environmental perception, they still have shortcomings when dealing with multimodal data fusion and complex environmental decision-making. For example, these methods often ignore numerical data of the robot's crawler / skid chassis, such as pressure data, speed, seabed depth, etc., which are critical for the robot's stable movement. In addition, traditional data acquisition methods are costly and risky in real seabed environments, which limits the efficiency and quality of model training. Summary of the Invention
[0003] In response to the technical problems existing in the prior art, the present invention provides a seabed operation robot movement planning method and system based on a visual language model, aiming to solve the problem of stable movement of the robot when operating on the rare soft soil surface of the seabed, ensure the stable movement of the robot, and improve the robot's adaptability to the rare soft soil surface of the seabed.
[0004] According to a first aspect of the present invention, a method for seabed operation robot movement planning based on a visual language model is provided, comprising: Use the robot's front-facing RGB camera to acquire image information, use a visual encoder to encode the RGB information, and extract key features from the image; Combined with the numerical data of the robot's crawler / skid chassis, the pressure data, speed, and seabed depth data are organized into prompt words to provide the vision-language model with the robot's current operating status and environmental information; The vision-language model uses the encoded RGB image and organized prompt words to determine the robot's best action plan in the current environment based on the combined information of the image and numerical data, and make step-by-step decisions; After making a step decision, the command is fine-tuned using LoRA. During the execution process, the robot's status and environmental changes are monitored in real time, and the feedback information is used as the input for the next round of reasoning to form a closed-loop control.
[0005] On the basis of the above technical solution, the present invention can also make the following improvements.
[0006] Optionally, the key features in the image include seabed topography and obstacle locations; the steps of acquiring image information using a front-mounted RGB camera of the robot, encoding the RGB information using a visual encoder, and extracting the key features in the image include: The robot acquires image information through its front-facing RGB camera, and uses a visual encoder to encode the RGB information to obtain word units containing the key features of the image, corresponding to the key features in the image. The seabed slope, whether there are obstacles, and whether there are depressions in the image are transmitted to the encoder in the form of an image. The encoder extracts potential features through various layers of neural networks for use by the vision-language model.
[0007] Optionally, after the robot front RGB camera is used to acquire image information, the acquired image needs to be preprocessed; the preprocessing includes: image enhancement, denoising, and sharpening operations.
[0008] Optionally, the numerical data of the combined robot crawler / skid chassis includes: A seabed mechanics simulation model is constructed using simulation software. Based on the numerical model of the seabed sediment and the robot crawler / skid chassis, the stress state and deformation law of the crawler / skid in the seabed environment are simulated to generate training data. The simulation software includes EDEM and ABAQUS, which are used to simulate the interaction between the seabed mechanical environment and the robot crawler / skid.
[0009] Optionally, organizing the pressure data, speed, and seabed depth data into prompt words to provide the vision-language model with the robot's current operating state and environment information includes: Collect numerical data of the crawler / skid chassis in real time, filter and calibrate the collected numerical data, and extract key numerical features through data processing algorithms, including: average pressure, pressure change rate, speed stability, and depth change trend; The collected numerical features are delayed aligned. Once consistent data is obtained, template filling or natural language generation is used to organize the processed numerical features into prompt words to provide the visual-language model with the robot's current operating status and environmental information.
[0010] Optionally, the vision-language model uses RGB images and organized prompt words to make step decisions, including forward, left turn, right turn and stop actions. The vision-language model judges the best action plan of the robot in the current environment based on the comprehensive information of the image and numerical data, ensuring the stable movement of the robot on the rare soft soil surface of the seabed.
[0011] Optionally, judging the best action plan of the robot in the current environment based on the comprehensive information of the image and numerical data and making a step decision includes: Construct a seabed mechanics simulation model. Set the model parameters based on the physical properties of seabed sediments and the robot's track / skid chassis structure. Simulate the stress state and deformation behavior of the track / skid in the seabed environment to generate training data. Design a variety of submarine operation scenarios to simulate different submarine terrains, environmental conditions, and robot operation tasks. Run the robot model in the simulation environment, collect training data, and generate a diverse dataset through simulation to cover scenarios and situations that are difficult to collect in the actual submarine environment. The training data includes image data, numerical data, and corresponding step-by-step decision labels. The collected simulation dataset is preprocessed and the processed data is used for training and fine-tuning the vision-language model.
[0012] Optionally, the method of using LoRA to fine-tune instructions and monitor the robot's status and environmental changes in real time, and using the feedback information as the next round of reasoning input to form a closed-loop control includes: The system acquires RGB images and numerical data of the track / skid chassis in real time. After image encoding and prompt word generation, it obtains visual feature vectors and prompt words. These visual feature vectors and prompt words are input into a fine-tuned vision-language model. The model performs reasoning based on the input information and outputs step-by-step decision results.
[0013] Optionally, the step decision results include forward, left turn, right turn and stop actions and corresponding confidence scores. According to the step decision results output by the vision-language model, the robot control system executes corresponding actions, and through continuous iteration, ensures the stable movement of the robot on the surface of rare soft soil on the seabed.
[0014] According to a second aspect of the present invention, a seabed operation robot travel planning system based on a visual language model is provided, comprising: The image information acquisition and encoding module is used to acquire image information using the robot's front-mounted RGB camera, encode the RGB information using a visual encoder, and extract key features from the image; The prompt word generation module is used to combine the numerical data of the robot crawler / skid chassis, pressure data, speed, and seabed depth data into prompt words, and provide the vision-language model with the robot's current operating status and environmental information; The step decision reasoning module is used by the vision-language model to use the encoded RGB image and the organized prompt words to determine the robot's optimal action plan in the current environment based on the combined information of the image and numerical data, and make step decisions; The instruction fine-tuning module is used to fine-tune the instructions using LoRA after making a step decision. During the execution process, it monitors the robot's status and environmental changes in real time, and uses the feedback information as the input for the next round of reasoning to form a closed-loop control.
[0015] Technical effects and advantages of the present invention: This paper proposes a method and system for seabed operation robot movement planning based on a visual language model, aiming to improve the robot's adaptability and stability on soft seabed surfaces. Leveraging the reasoning capabilities of a large multimodal language model, combined with image information and numerical data, this method achieves precise movement planning for the robot through a combination of visual encoders, prompt word generation, and model fine-tuning. Furthermore, the method utilizes simulation software to generate training data, reducing the cost and risk of data acquisition in real-world seabed environments and improving the efficiency of model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A diagram showing the steps of a seabed operation robot movement planning method based on a visual language model provided by an embodiment of the present invention; Figure 2 A flowchart of data collection and preparation for a seabed operation robot motion planning method based on a visual language model provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0018] It is understandable that, based on the defects in the background technology, the embodiment of the present invention proposes a seabed operation robot travel planning method based on a visual language model. The method is suitable for stable travel planning of a seabed robot when operating on the surface of rare soft soil on the seabed, and improves the robot's adaptability and travel stability to the surface of rare soft soil on the seabed. Specifically, Figure 1 Said method comprises the following steps: Step S1: Use the robot's front RGB camera to acquire image information, use a visual encoder to encode the RGB information, and extract key features in the image; In this implementation step, key features in the image, such as seabed topography, obstacle locations, etc., will serve as input to the vision-language model, providing a basis for environmental perception.
[0019] Front-facing RGB camera installation and commissioning: A high-resolution RGB camera is installed on the front of the seabed robot, ensuring that the camera covers the robot's field of view in its direction of travel. The camera should be installed in a location that avoids obstructions and interference to ensure clear and continuous image acquisition. Camera parameters, such as focal length, aperture, and frame rate, are adjusted to adapt to the complex optical environment of low light and high scattering on the seabed.
[0020] As the robot navigates the seabed, its RGB camera collects image data in real time. These images may contain noise, blur, and color distortion. Therefore, preprocessing is required when acquiring image information using the robot's front-mounted RGB camera. This preprocessing includes image enhancement, denoising, and sharpening to improve image quality and enhance the discernibility of key features.
[0021] The encoding of RGB information using a visual encoder includes: A visual encoder is selected to extract features from the preprocessed RGB image, and the RGB image is encoded using CLIP's visual encoder. These models can automatically extract high-level semantic features from the image, such as seabed topography, obstacles, and sediment types. The extracted feature vectors are used as input to the vision-language model, providing the visual information foundation for subsequent movement planning.
[0022] Step S2: combining the numerical data of the robot crawler and the skid chassis, and organizing the pressure data, speed, and seabed depth data into prompt words; In this embodiment, numerical data from the robot's tracks / skid chassis, including but not limited to pressure, speed, and seabed depth, is combined and organized into prompt words. These prompt words provide the vision-language model with information about the robot's current operating state and environment, helping the model make more accurate decisions.
[0023] Track / Skid Chassis Sensor Configuration: Various sensors, including pressure sensors, speed sensors, and depth sensors, are installed on the robot's tracks or skid chassis. The pressure sensor is used to monitor the pressure distribution between the track or skid and the seabed in real time. The speed sensor is used to measure the robot's travel speed. The depth sensor is used to obtain the robot's vertical position information on the seabed.
[0024] Numerical data acquisition and processing: The sensor collects numerical data from the crawler / skid chassis in real time. This data may be affected by noise and needs to be filtered and calibrated. Through data processing algorithms, key numerical features such as average pressure, pressure change rate, speed stability, depth change trend, etc. are extracted. The data preparation process is as follows: Figure 2As shown in the figure, in the simulation environment, the camera data and the seabed data come from different sensors, and the acquisition delays need to be aligned. After alignment, with consistent data, prompt words can be further prepared to prepare for fine-tuning.
[0025] Prompt word generation strategy: The processed numerical data is organized into prompt words, which are used to provide the vision-language model with information about the robot's current operating status and environment. Prompt word generation can be achieved through template filling or natural language generation. For example, a prompt might be something like "The current track / skid chassis pressure is XX Pa, the speed is XX m / s, and the depth is XX m. The next step should be forward." These prompt words are concise and clear, accurately reflecting the key information in the numerical data and making them easier for the vision-language model to understand and process.
[0026] Step S3: The vision-language model uses the encoded RGB image and the organized prompt words to determine the best action plan for the robot in the current environment based on the comprehensive information of the image and numerical data, and makes a step decision; In this example, a vision-language model uses encoded RGB images and organized prompts to make step decisions, including forward, left, right, and stop. Based on the combined information from the image and numerical data, the model determines the robot's optimal action plan in the current environment, ensuring stable movement on the soft seabed surface.
[0027] Determining the best action plan for the robot in the current environment and making step-by-step decisions include: Simulation Software Selection and Model Construction: Select appropriate simulation software, such as EDEM or ABAQUS, to build a seabed mechanics simulation model. Based on the physical properties of the seabed sediments and the robot's track / skid chassis structure, set the simulation model parameters, including soil density, internal friction angle, cohesion, and compression modulus.
[0028] Training Data Collection and Simulation: To efficiently collect training data, we use simulation software (such as EDEM and ABAQUS) to build a seabed mechanics model. This simulates the stress state and deformation patterns of the crawler / skid in the seabed environment, generating a large amount of training data. This data will be used for model training and optimization, improving the model's generalization capabilities and decision-making accuracy.
[0029] Simulation Scenario Design and Data Collection: We design various subsea operation scenarios, simulating diverse seafloor topography, environmental conditions, and robotic tasks. We run the robot model in the simulation environment and collect a large amount of training data, including image data, numerical data, and corresponding step-by-step decision labels. Through simulation, we can efficiently generate diverse datasets covering scenarios and situations that are difficult to capture in real-world subsea environments.
[0030] Data Processing and Model Training: Collected simulation data is preprocessed, including image enhancement, data cleaning, and feature extraction. The processed data is used to train and fine-tune the vision-language model, improving its generalization and decision-making accuracy. By leveraging simulation data, the cost and risk of data collection in real-world submarine environments are reduced, improving the efficiency and quality of model training.
[0031] Step S4: After making a step decision, LoRA is used to fine-tune the instructions. During the execution process, the robot's status and environmental changes are monitored in real time, and the feedback information is used as the input for the next round of reasoning to form a closed-loop control.
[0032] In this example, a multimodal large language model, such as Qwen2-VL, is selected as the base model. These models have strong cross-modal understanding capabilities and can integrate visual features and language information to perform complex reasoning and decision-making.
[0033] Application of the LoRA fine-tuning method: The LoRA (Low-Rank Adaptation) method is used to fine-tune the vision-language model, enabling it to understand the seabed environment and make stable stepping direction decisions. The LoRA method inserts low-rank matrices into key layers of the model to efficiently adjust model parameters while maintaining model stability and generalization. During the fine-tuning process, the model is trained with images and prompt word data from seabed operation scenes, enabling it to understand the unique characteristics of the seabed environment and the robot's movement requirements. The LoRA fine-tuning method optimizes the task while maintaining efficient model parameters, improving the model's adaptability and decision-making accuracy.
[0034] Fine-tuning data preparation and training: Prepare a training dataset containing seafloor images and corresponding prompt words. The dataset should cover a variety of seafloor topography, environmental conditions, and robot states. Through transfer learning and fine-tuning training, enable the model to make accurate step-by-step decisions based on the images and prompt words. During training, use appropriate optimization algorithms and loss functions, such as the Adam optimizer and the cross-entropy loss function. Monitor loss changes and model performance during training to ensure that the model converges to the desired state.
[0035] Reasoning input preparation: During the robot's movement, RGB images and numerical data of the track / skid chassis are acquired in real time. After image encoding and prompt word generation steps, visual feature vectors and prompt words are obtained.
[0036] Vision-Language Model Inference: The visual feature vector and prompt word are input into the fine-tuned vision-language model. The model performs inference based on the input information and outputs step decision results. Decision results include actions such as forward, turn left, turn right, and stop, as well as corresponding confidence scores.
[0037] Decision execution and feedback: Based on the step-by-step decision output from the model, the robot control system executes the corresponding action. During execution, the robot's status and environmental changes are monitored in real time, and feedback is used as input for the next round of reasoning, forming a closed-loop control. Through continuous iteration, the robot's stable movement on the soft seabed surface is ensured.
[0038] In summary, the method of the present invention, by leveraging the reasoning capabilities of a multimodal large language model, can adapt to complex submarine rare-soft soil environments, improve the robot's ability to navigate these environments, and ensure the robot can successfully complete submarine operations. By combining image information and numerical data, the model can make more accurate stepping decisions, preventing unstable or dangerous situations during submarine operations and improving operational safety and reliability. Using simulation software to generate training data reduces the cost and risk of data acquisition in actual submarine environments, improves the efficiency of model training, and provides strong support for model optimization and improvement.
[0039] According to a second aspect of the present invention, a seabed operation robot travel planning system based on a visual language model is provided, comprising: The image information acquisition and encoding module is used to acquire image information using the robot's front-mounted RGB camera, encode the RGB information using a visual encoder, and extract key features from the image; The prompt word generation module is used to combine the numerical data of the robot crawler / skid chassis, pressure data, speed, and seabed depth data into prompt words, and provide the vision-language model with the robot's current operating status and environmental information; The step decision reasoning module is used by the vision-language model to use the encoded RGB image and the organized prompt words to determine the robot's optimal action plan in the current environment based on the combined information of the image and numerical data, and make step decisions; The instruction fine-tuning module is used to fine-tune the instructions using LoRA after making a step decision. During the execution process, it monitors the robot's status and environmental changes in real time, and uses the feedback information as the input for the next round of reasoning to form a closed-loop control.
[0040] It can be understood that the seabed operation robot travel planning system based on the visual language model provided by the present invention corresponds to the seabed operation robot travel planning method based on the visual language model provided in the aforementioned embodiments. The relevant technical features of the seabed operation robot travel planning system based on the visual language model can refer to the relevant technical features of the seabed operation robot travel planning method based on the visual language model, which will not be repeated here.
[0041] To address these issues, the present invention proposes a method and system for seabed operation robot movement planning based on a visual language model, designed to improve the robot's adaptability and stability on soft seabed surfaces. By leveraging the reasoning capabilities of a large multimodal language model, combined with image information and numerical data, and employing techniques such as visual encoders, prompt word generation, and model fine-tuning, precise robot movement planning is achieved. Furthermore, the present invention utilizes simulation software to generate training data, reducing the cost and risk of data acquisition in actual seabed environments and improving the efficiency of model training.
[0042] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0043] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0044] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0045] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
[0046] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A seabed operation robot movement planning method based on a visual language model, characterized in that: The following steps are involved: Use the robot's front-facing RGB camera to acquire image information, use a visual encoder to encode the RGB information, and extract key features from the image; Combined with the numerical data of the robot's crawler / skid chassis, the pressure data, speed, and seabed depth data are organized into prompt words to provide the vision-language model with the robot's current operating status and environmental information; The vision-language model uses the encoded RGB image and organized prompt words to determine the robot's best action plan in the current environment based on the combined information of the image and numerical data, and make step-by-step decisions; After making a step decision, the command is fine-tuned using LoRA, and the robot's status and environmental changes are monitored in real time. The feedback information is used as the input for the next round of reasoning to form a closed-loop control.
2. The seabed operation robot movement planning method based on the visual language model according to claim 1 is characterized in that: Key features in said imagery include seafloor topography and obstacle locations; The robot's front-facing RGB camera is used to acquire image information, and a visual encoder is used to encode the RGB information. The key features extracted from the image include: The robot acquires image information through its front-facing RGB camera, and uses a visual encoder to encode the RGB information to obtain word units containing the key features of the image, corresponding to the key features in the image. The seabed slope, whether there are obstacles, and whether there are depressions in the image are transmitted to the encoder in the form of an image. The encoder extracts potential features through various layers of neural networks for use by the vision-language model.
3. The seabed operation robot movement planning method based on the visual language model according to claim 1 is characterized in that: After the robot's front RGB camera is used to obtain image information, the collected image needs to be preprocessed; Preprocessing includes: image enhancement, denoising, and sharpening operations.
4. The seabed operation robot movement planning method based on the visual language model according to claim 1 is characterized in that: The numerical data of the combined robot crawler / skid chassis include: A seabed mechanics simulation model is constructed using simulation software. Based on the numerical model of the seabed sediment and the robot crawler / skid chassis, the stress state and deformation law of the crawler / skid in the seabed environment are simulated to generate training data. The simulation software includes EDEM and ABAQUS, which are used to simulate the interaction between the seabed mechanical environment and the robot crawler / skid.
5. The seabed operation robot movement planning method based on the visual language model according to claim 1 is characterized in that: The pressure data, speed, and seabed depth data are organized into prompt words to provide the vision-language model with the robot's current operating status and environmental information, including: Collect numerical data of the crawler / skid chassis in real time, filter and calibrate the collected numerical data, and extract key numerical features through data processing algorithms, including: average pressure, pressure change rate, speed stability, and depth change trend; The collected numerical features are delayed aligned. Once consistent data is obtained, template filling or natural language generation is used to organize the processed numerical features into prompt words to provide the visual-language model with the robot's current operating status and environmental information.
6. The seabed operation robot movement planning method based on the visual language model according to claim 5 is characterized in that: The vision-language model uses RGB images and organized prompt words to make step decisions, including forward, left, right, and stop actions. Based on the combined information of images and numerical data, the vision-language model determines the robot's optimal action plan in the current environment, ensuring stable movement of the robot on the rare soft soil surface of the seabed.
7. The seabed operation robot movement planning method based on the visual language model according to claim 6 is characterized in that: The process of determining the best action plan for the robot in the current environment based on the comprehensive information of the image and numerical data and making a step-by-step decision includes: Construct a seabed mechanics simulation model. Set the model parameters based on the physical properties of seabed sediments and the robot's track / skid chassis structure. Simulate the stress state and deformation behavior of the track / skid in the seabed environment to generate training data. Design a variety of submarine operation scenarios to simulate different submarine terrains, environmental conditions, and robot operation tasks. Run the robot model in the simulation environment, collect training data, and generate a diverse dataset through simulation to cover scenarios and situations that are difficult to collect in the actual submarine environment. The training data includes image data, numerical data, and corresponding step-by-step decision labels. The collected simulation dataset is preprocessed and the processed data is used for training and fine-tuning the vision-language model.
8. The seabed operation robot movement planning method based on the visual language model according to claim 1 is characterized in that: The use of LoRA to fine-tune instructions and monitor the robot's state and environmental changes in real time, using feedback information as input for the next round of reasoning, forming a closed-loop control, including: The system acquires RGB images and numerical data of the track / skid chassis in real time. After image encoding and prompt word generation, it obtains visual feature vectors and prompt words. These visual feature vectors and prompt words are input into a fine-tuned vision-language model. The model performs reasoning based on the input information and outputs step-by-step decision results.
9. The seabed operation robot movement planning method based on the visual language model according to claim 8 is characterized in that: The step decision results include forward, left turn, right turn, and stop actions, as well as corresponding confidence scores. Based on the step decision results output by the vision-language model, the robot control system executes corresponding actions. Through continuous iteration, the robot ensures stable movement on the surface of rare soft soil on the seabed.
10. A seabed operation robot movement planning system based on a visual language model, characterized by: include: The image information acquisition and encoding module is used to acquire image information using the robot's front-mounted RGB camera, encode the RGB information using a visual encoder, and extract key features from the image; The prompt word generation module is used to combine the numerical data of the robot crawler / skid chassis, pressure data, speed, and seabed depth data into prompt words, and provide the vision-language model with the robot's current operating status and environmental information; The step decision reasoning module is used by the vision-language model to use the encoded RGB image and the organized prompt words to determine the robot's optimal action plan in the current environment based on the combined information of the image and numerical data, and make step decisions; The instruction fine-tuning module is used to fine-tune the instructions using LoRA after making a step decision. During the execution process, it monitors the robot's status and environmental changes in real time, and uses the feedback information as the input for the next round of reasoning to form a closed-loop control.