A multi-point robot navigation system with fine-tuned adaptive small language model

By fine-tuning and iteratively optimizing small language models on the robot, the problem of local deployment of language models in robot navigation is solved, and efficient and stable robot navigation and human-computer interaction are achieved.

CN118862954BActive Publication Date: 2025-05-13SHANGHAI JIAOTONG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410833513.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-25
Publication Date
2025-05-13
Estimated Expiration
2044-06-25

AI Technical Summary

Technical Problem

The prior art has difficulty in deploying language models locally on robots, resulting in poor performance of robot navigation in complex environments and natural language instruction processing.

Method used

Through fine-tuning and teacher-student iteration methods, the small language model is optimized so that it can run on the robot side with limited resources and enable local deployment. The model fine-tuning module uses the human cycle generation method to generate data sets and fine-tune parameters of small language models; the teacher-student iteration module guides the student model through the teacher model to improve its performance in navigation tasks.

Benefits of technology

It significantly reduces the demand for hardware resources, avoids dependence on remote servers, improves the autonomy and stability of the system, and enhances the robot's human-computer interaction and complex environment understanding capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118862954B_ABST
    Figure CN118862954B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-point robot navigation system of fine-tuning an adaptive small language model, which relates to the field of robot navigation, fine-tunes a pre-trained small language model, obtains a usable model through a teacher-student iteration process, and the model is directly connected to the navigation system, and a specified target point is given according to a user instruction to guide the robot to navigate; including: a model fine-tuning module fine-tunes the parameters of the small language model, and the output of the fine-tuned small language model can be used in the context of robot navigation; a teacher-student iteration module implements the teacher-student iteration process; and a robot navigation controller module is responsible for converting the output of the small language model into actual navigation actions to realize multi-point navigation tasks. The present invention adopts fine-tuning, iteration and other methods to significantly increase the ability to deploy low-cost small language models, reduce the cost of local deployment of the model on the robot side, reduce dependence on the network, and ensure that data is processed locally, so as to protect user privacy well.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of robot navigation, and in particular to a multi-point robot navigation system with a fine-tuned adaptive small language model. Background Art

[0002] Driven by the rapid development of robotics technology, robots have been widely used in various fields such as medical, household and industry. However, traditional robot control methods are limited to simple tasks and are not sufficient to handle complex environments and tasks described in natural language. Large language models (LLMs), such as ChatGPT and Llama, have demonstrated excellent natural language processing and logical reasoning capabilities, highlighting their potential in enhancing robot navigation.

[0003] In the field of robot navigation combined with large models, a lot of excellent work has emerged in the academic community. SayCan demonstrates a method that can perform temporary extension tasks without pre-training, which has strong interpretability. SayTap directly generates the motion pattern of a quadruped robot and can respond to various types of instructions. However, the design of a random pattern generator and the training of a large number of gaits involve a trade-off between sample balance and data efficiency. VLMaps proposes a method for spatial target navigation in zero-sample conditions, allowing obstacle maps to be shared among multiple robots. VLNBERT reaches the SOTA level and simplifies the model architecture, but relies on a large amount of pre-training data, and its performance needs to be improved when processing long instructions. All of the above methods have well integrated large models into robot navigation, but most of the use of large models adopts remote calls, and the problem of local deployment of models has never been solved.

[0004] In terms of locally deployable small language models, although small language models (SLMs) have been very versatile, their performance in specific tasks is not outstanding, and untrained SLMs also struggle to perform robot control tasks. In order to improve the capabilities of small language models, a series of techniques can be used. Parameter Efficient Fine-Tuning (PEFT) is a common and effective method that improves the performance of small language models by training them with a small test set, thereby significantly improving their capabilities in specific tasks. Knowledge distillation is another common method to improve the performance of SLMs by leveraging the stored knowledge of large language models. Through step-by-step distillation, SLMs can achieve performance close to LLMs in specific tasks even with small datasets.

[0005] Therefore, technicians in this field are committed to developing a multi-point robot navigation system with a fine-tuned adaptive small language model. Summary of the invention

[0006] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is to meet the challenge of locally deploying a language model on a robot, to realize the local deployment of SLM for robot navigation in a real environment, and to enable the robot to interact with the user accurately and instantly.

[0007] To achieve the above-mentioned object, the present invention provides a multi-point robot navigation system with a fine-tuned adaptive small language model, characterized in that the navigation system fine-tunes the pre-trained small language model, obtains a model that can be used by the navigation system through a teacher-student iterative process, and the model is directly connected to the navigation system, and a specified target point is given according to a user instruction to guide the robot to navigate; wherein the navigation system includes:

[0008] A model fine-tuning module generates a test data set by a human-in-the-loop generation method, and uses the data set to fine-tune the parameters of the small language model. The fine-tuned small language model outputs a context that can be used for robot navigation;

[0009] A teacher-student iteration module implements the teacher-student iteration process, including a teacher model and a student model, wherein the student model is the small language model fine-tuned by the model fine-tuning module, and the teacher model is used to guide the student model; the teacher model generates prompts according to task requirements and environmental information, and adjusts the prompts according to the performance feedback of the student model; the student model attempts to complete the task, and the results of the student model's completion are used as the basis for the next round of prompts;

[0010] The robot navigation controller module is responsible for converting the output of the small language model into the actual navigation action of the multi-point robot to realize the multi-point navigation task.

[0011] Furthermore, during the fine-tuning process of the model fine-tuning module, the context output by the small language model adopts a fixed JSON format, and the context includes explanation and location information, wherein:

[0012] The explanation is set as a string, showing the analysis and reasoning of the task by the small language model, and the reason for the success or failure of the task;

[0013] The position is set to a list of number pairs containing the coordinates of the target point;

[0014] The navigation system uses the position to navigate the robot and displays the robot's thought process through the interpretation.

[0015] Furthermore, the human in-loop generation method comprises the following steps when generating the data set:

[0016] S1.1: Write a first task based on a real-life scenario, and combine the first task and environmental information into a prompt;

[0017] S1.2: the large language model generates a second task according to the prompt;

[0018] S1.3: The human evaluator scores the second task, and the second task and the corresponding score are used as data in the dataset;

[0019] S1.4: The second tasks and the scores are brought back to the large language model, and the large language model generates a new set of the second tasks according to the scores;

[0020] S1.5: Repeat S1.3 to S1.4 until enough high-quality data sets are obtained.

[0021] Furthermore, during the fine-tuning process, the model fine-tuning module uses the PEFT library of Huggingface to quickly implement parameter fine-tuning, and adjusts the learning rate and the number of training rounds to obtain the accuracy of the small language model on the data set under different parameters.

[0022] Furthermore, the model fine-tuning module uses the following steps to fine-tune the parameters of the small language model:

[0023] S2.1: Generate the data set by the human-in-the-loop generation method;

[0024] S2.2: Using the data set, using the LoRA method to perform low-rank adjustment on the model parameters of the small language model;

[0025] S2.3: testing the adjusted small language model by inputting the navigation task test set to obtain an output result of the small language model;

[0026] S2.4: extracting the target point in the output result, and calculating the accuracy of the small language model on the test set;

[0027] S2.5: Determine whether the format of the output result and the accuracy meet the requirements. If not, repeat S2.2 to S2.5;

[0028] S2.6: Output the fine-tuned small language model.

[0029] Furthermore, in the teacher-student iteration module, the teacher model is set as a knowledgeable large language model, and the teacher model acts as a prompt engineer, generates appropriate prompt content according to the current task and map environment information, obtains feedback from the results of the previous iteration, and adjusts the prompt content in real time according to the feedback; the student model receives the prompt content and generates corresponding output according to the current task, and the output is recorded and used to provide feedback information to the teacher model in subsequent iterations.

[0030] Furthermore, the teacher-student iterative process includes the following steps:

[0031] S3.1: Using the small language model fine-tuned by the model fine-tuning module as the student model;

[0032] S3.2: inputting the navigation task test set into the teacher model and the student model;

[0033] S3.3: extracting the output results of the student model and performing simulation experiments to obtain performance data of the student model;

[0034] S3.4: Feeding back the performance data to the teacher model, and the teacher model providing teaching guidance to the student model;

[0035] S3.5: Repeat S3.2 to S3.4 for multiple rounds of iterations to obtain the small language model that can finally be used.

[0036] Furthermore, the robot navigation controller module extracts a list of target points from the output of the small language model, guides the robot toward the target points, and updates the planned path in real time to help the robot navigate from the initial position to the expected points in sequence.

[0037] Furthermore, the small language model makes navigation decisions based on a pre-built map. After receiving a natural language command, the small language model uses the map to determine a series of target waypoints, which accurately reflect the intent of the natural language instruction and comply with the constraints of the map.

[0038] Furthermore, the natural language instruction is composed of a series of words, and the small language model maps the natural language instruction and the map information to the target waypoint sequence through the navigation decision, and the navigation decision process is formalized by the function f:

[0039] f(W,M)={p1,p2,…,p n},

[0040] Where W is a natural language command, which consists of a series of words w1, w2…, composition,

[0041]

[0042] M is the map, p i is the coordinate or identifier of a location on the map M, i = 1, 2…, n.

[0043] In a preferred embodiment of the present invention, compared with the prior art, the present invention has the following beneficial effects:

[0044] 1. The present invention significantly reduces the demand for hardware resources by adopting fine-tuning and teacher-student iteration methods, so that high-performance language models can run on the robot side with limited resources, avoiding dependence on remote servers, and improving the autonomy and stability of the system. It is particularly outstanding in applications with poor network conditions or requiring high real-time performance. At the same time, the data is processed locally, which well protects the privacy and data security of users.

[0045] 2. The present invention uses a different idea from the mainstream compression method. It does not compress the large model to obtain the small model, but starts directly from the pre-trained small model. Through fine-tuning and teacher-student iteration, it learns the knowledge of the large model to solve specific tasks, providing a new idea for the compression and local deployment of large models, promoting the promotion and application of large models, and promoting the construction of low-cost and high-performance large models.

[0046] 3. Innovatively integrate low-cost, high-performance small language models into the robot's multi-point real-time navigation, promote the development of large-model-based robot navigation technology that does not rely on the network, and significantly enhance the robot's human-computer interaction and understanding of complex dynamic environments.

[0047] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is an overall schematic diagram of the robot navigation process of an embodiment of the present invention;

[0049] Figure 2 is a schematic diagram of a small language model parameter adjustment process for robot navigation according to an embodiment of the present invention;

[0050] Figure 3 is a schematic diagram of a small language model parameter fine-tuning process according to an embodiment of the present invention;

[0051] Figure 4 is a schematic diagram of a small language model teacher-student iteration process of an embodiment of the present invention;

[0052] Figure 5It is a schematic diagram of the algorithm implementation of the small language model teacher-student iterative process in an embodiment of the present invention. DETAILED DESCRIPTION

[0053] The following describes several preferred embodiments of the present invention with reference to the drawings in the specification, so that the technical content is clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.

[0054] In the drawings, components with the same structure are indicated by the same numerical reference numerals, and components with similar structures or functions are indicated by similar numerical reference numerals. The size and thickness of each component shown in the drawings are arbitrarily shown, and the present invention does not limit the size and thickness of each component. In order to make the illustration clearer, the thickness of the components is appropriately exaggerated in some places in the drawings.

[0055] Traditional robot navigation methods find it difficult to handle and predict dynamic changes in complex environments, as well as to achieve better human-computer interaction. Currently, large models have too many parameters to be deployed locally on the robot side. In order to meet the challenge of locally deploying language models on the robot, the present invention proposes an innovative, adaptive, task-specific model compression method that significantly reduces the cost of local model deployment and can achieve local deployment of small language models for robot navigation in real environments, enabling the robot to interact with users accurately and instantly.

[0056] like Figure 1 , Figure 2 As shown, an embodiment of the present invention provides a multi-point robot navigation system FastNav with a fine-tuned adaptive small language model. The pre-trained small language model is fine-tuned, and a model that can be used by the navigation system FastNav is obtained through a teacher-student iterative process. The model is directly connected to the navigation system, and specifies the target point according to the user's instructions to guide the robot to navigate.

[0057] In this embodiment, the navigation system FastNav mainly includes the following modules: a model fine-tuning module, a teacher-student iteration module and a robot navigation controller module, wherein:

[0058] The model fine-tuning module generates a test data set through a human-in-the-loop generation method, and uses the data set to fine-tune the parameters of the small language model. The output of the fine-tuned small language model can be used in the context of robot navigation;

[0059] The teacher-student iteration module implements the teacher-student iteration process. This module includes a teacher model and a student model. The student model is a small language model that has been fine-tuned by the model fine-tuning module. The teacher model is used to guide the student model. The teacher model generates prompts based on task requirements and environmental information, and adjusts the prompts based on the performance feedback of the student model. The student model tries to complete the task, and the results of the student model are used as the basis for the next round of prompts.

[0060] The robot navigation controller module is responsible for converting the output of the small language model into the actual navigation actions of the multi-point robot to achieve the multi-point navigation task.

[0061] Through the FastNav system, the adjusted small language model SLM can achieve performance close to or even better than large models in the navigation field while maintaining its lightweight. Experimental results show that the SLM processed by the FastNav system has demonstrated high success rate, low navigation error and high implementation efficiency in both simulated environments and real robots.

[0062] The following is a detailed description of each module of the FastNav system in this embodiment.

[0063] 1. Model fine-tuning module

[0064] In order to adapt the pre-trained SLM to the specific field of robot navigation and avoid defects in handling specific navigation tasks, the SLM needs to be fine-tuned first. The model fine-tuning module in the FastNav system is responsible for completing the parameter fine-tuning of the SLM.

[0065] like Figure 3 As shown, during the fine-tuning process of the model fine-tuning module, the context output by the small language model SLM adopts a fixed JSON format. The context includes explanation and location information, where:

[0066] The explanation is a string that shows the small language model’s analysis and reasoning of the task, as well as the reasons for the success or failure of the task, making it easier to understand the thinking process of the SLM model;

[0067] position is a list of pairs containing the x and y coordinates of the target point.

[0068] With this design, it is hoped that the SLM will be able to output in this format and use the positions to navigate the robot, showing the robot's thought process through interpretation.

[0069] Before fine-tuning parameters, human-in-the-loop generation is used to generate datasets for parameter fine-tuning. First, some tasks are written using real-life examples. These tasks and some environmental information form a prompt. Then, the large language model LLM generates some tasks based on the prompt. Human evaluators join this process and score the tasks. The tasks and scores are brought back to the LLM, and then it generates a new set of tasks based on human feedback. We will repeat this cycle until we have enough high-quality data. After multiple cycles, a high-quality dataset that meets human requirements is obtained.

[0070] In this embodiment, the detailed data generation process of the human in loop generation method includes:

[0071] Step 1: Write the first task according to the real-life scenario and combine the first task and environmental information into a prompt.

[0072] Step 2: The large language model LLM generates the second task according to the prompt. In this embodiment, the LLM can use a large language model such as ChatGPT or Llama.

[0073] Step 3: The human evaluator scores the second task, and the second task and the corresponding score are used as data in the dataset;

[0074] Step 4: The second task and the score are brought back to the large language model, which generates a new set of second tasks based on the score.

[0075] Step 5: Repeat steps 3 and 4 until enough high-quality data sets are obtained.

[0076] After obtaining enough data sets, the SLM is fine-tuned using the data sets so that it outputs the correct JSON text as expected. This embodiment uses the PEFT library of Huggingface to quickly implement this process. During the fine-tuning process, parameters such as the learning rate and the number of training rounds can be adjusted to obtain the accuracy of the SLM on the test set under different parameters, and the accuracy is compared.

[0077] For the relevant fine-tuning process, please refer to the attached Figure 3In this embodiment, the pre-trained language model is selected for adjustment. First, the LoRA method is selected to perform low-rank adjustment on the model parameters using the data set generated above. The adjusted model is input into the artificially defined navigation task test set to obtain the output result of the model. Then, it is determined whether the result is in the required JSON format, and then the output target point is extracted, and the accuracy of the model on the test set is calculated. The relevant parameters are adjusted according to whether the JSON format is correct and the accuracy is high or low. This process is repeated until the best model is obtained, and the best model is selected for subsequent processes.

[0078] In this embodiment, the model fine-tuning module uses the following steps to fine-tune the parameters of the small language model SLM:

[0079] Step 1: Generate a dataset using the human-in-the-loop generation method described above;

[0080] Step 2: Using this dataset, use the LoRA method to perform low-rank adjustment on the model parameters of the small language model;

[0081] Step 3: Test the adjusted small language model SLM by inputting the navigation task test set to obtain the output result of the small language model SLM;

[0082] Step 4: Extract the target points in the output results and calculate the accuracy of the small language model SLM on the test set;

[0083] Step 5: Determine whether the format and accuracy of the output result meet the requirements. If not, repeat steps 2 to 5.

[0084] Step 6: After parameter adjustment, the format and accuracy of the SLM model output result meet the design requirements, and the fine-tuned small language model SLM is output.

[0085] 2. Teacher-student iteration module

[0086] The teacher-student model is a paradigm that has been proven to be effective in language model training and knowledge transfer. The teacher-student iteration module in this example adopts the idea of ​​a multi-agent system, in which multiple agents collaborate and learn from each other to achieve a common goal. This approach consists of a more knowledgeable teacher model guiding a less knowledgeable student model, with the goal of improving the student's performance through iterative learning and feedback.

[0087] like Figure 4As shown, in this embodiment, a strong language model (such as GPT-4) is designated as the teacher, and a weaker SLM acts as the student. In each iteration, the teacher acts as a prompt engineer, responsible for generating appropriate prompts for the student based on the current task and map environment information. The teacher also receives feedback from the results of the previous iteration and adjusts the prompt content in real time based on this feedback. This ensures that the student makes correct judgments and corrects incorrect judgments as much as possible. The student is usually a fine-tuned SLM designed to receive prompts from the teacher and generate corresponding outputs based on the current task. In each iteration, the student's output is recorded and used to provide feedback to the teacher in subsequent iterations.

[0088] In this embodiment, the teacher model selects the knowledgeable large language model GPT-4 to act as a prompt engineer, generates appropriate prompt content according to the current task and map environment information, obtains feedback from the results of the previous iteration, and adjusts the prompt content in real time according to the feedback; the student model uses the SLM model obtained by the model fine-tuning module, the student model receives the prompt content, and generates corresponding output according to the current task. The output is recorded and used to provide feedback information to the teacher model in subsequent iterations. A manually defined navigation data set is selected and input to GPT-4 and SLM respectively. By extracting the results of SLM output, a simulation experiment is performed in Gazebo to obtain the performance of SLM without iteration. Then, the performance of SLM on each test set is fed back to GPT-4, and GPT-4 guides SLM according to the relevant situation of each test set. The model obtained by guidance repeats the above process, and after multiple rounds of iterations, a model that can be used is finally obtained.

[0089] like Figure 4 As shown, the teacher-student iteration process in this embodiment includes the following steps:

[0090] Step 1: Use the small language model SLM fine-tuned by the model fine-tuning module as the student model;

[0091] Step 2: Input the navigation task test set into the teacher model GPT-4 and the student model SLM;

[0092] Step 3: Extract the output results of the student model SLM and conduct simulation experiments to obtain the performance data of the student model SLM;

[0093] Step 4: Feedback the performance data to the teacher model GPT-4, which then provides guidance to the student model SLM.

[0094] Step 5: Repeat steps 2 to 4 for multiple rounds of iterations to obtain the final usable small language model SLM.

[0095] The algorithm of the teacher-student iteration process in this embodiment is as follows: Figure 5 shown.

[0096] 3. Robot navigation controller module

[0097] The robot navigation controller module is a key component in the FastNav system, which is responsible for converting the target points generated by the language model into the actual navigation path of the robot.

[0098] In this embodiment, the "robot language navigation problem" is first defined: given a natural language command, which consists of a series of words. The task of the small language model SLM is to make navigation decisions based on a pre-built map. After receiving the command, the SLM model uses the map to determine a series of target waypoints, which determine the navigation path.

[0099] After obtaining the target point list from the small language model SLM, a robot navigation controller is needed to guide the robot toward the target point and update the planned path in real time. In this embodiment, the modular navigation framework Navigation2 is selected. The target point list is extracted from the output of the small language model SLM and placed in Navigation2. Then, the Navigation2 controller will help the robot navigate from the initial position to the expected point in sequence according to the list.

[0100] Specifically, the natural language instruction W consists of a sequence of words w. The small language model SLM maps the natural language instructions and map information to a target waypoint sequence through navigation decisions. The navigation decision process is formalized by the function f, which maps the language commands and map information to a waypoint sequence. The main goal of the small language model SLM is to optimize the waypoint sequence so that it accurately reflects the intention of the language instruction while complying with the constraints and practicality of the map.

[0101] The function f is specifically:

[0102] f(W,M)={p1,p2,…,p n},

[0103] Where W is a natural language command, which consists of a series of words w1, w2…, composition:

[0104]

[0105] M is the map, p i is the coordinate or identifier of a location on the map M, i = 1, 2…, n.

[0106] In this embodiment, the robot navigation controller module ensures that FastNav can effectively convert the output of the language model into the actual navigation actions of the robot, thereby achieving efficient and accurate multi-point navigation tasks.

[0107] Compared with the prior art, the multi-point robot navigation system with fine-tuned adaptive small language model provided in this embodiment has the following advantages:

[0108] 1. In response to the problem that the parameters of the currently used large models are too large and difficult to deploy locally on the robot side, the present invention proposes an innovative adaptive model compression method for specific tasks, which significantly reduces the cost of local model deployment. By adopting methods such as fine-tuning and iteration, the ability of small language models with low deployment costs in specific tasks is significantly increased, and the cost of local deployment of models on the robot side is significantly reduced, which promotes the application and development of embodied intelligence. While reducing dependence on the network, it ensures that data is processed locally, protects user privacy well, and at the same time meets the demand for low-cost deployment of large language models on the edge, promoting the promotion and application of large language models.

[0109] 2. In view of the problem that most of the current model compression methods have too much computational complexity and the compression effect is not good enough, the present invention uses a different idea from the mainstream compression method. Instead of compressing the large model to obtain the small model, it starts directly from the pre-trained small model and learns the knowledge of the large model to solve specific tasks through fine-tuning and teacher-student iteration. The teacher-student iteration includes a knowledgeable teacher model, such as GPT-4, guiding a student model with less knowledge, that is, a small language model. The teacher generates prompts based on task requirements and environmental information, and repeatedly adjusts these prompts based on student performance feedback. Then, the student tries to complete the task, and the result is used as the basis for the next round of prompts, which provides a new idea for the compression and local deployment of large models, promotes the promotion and application of large models, and promotes the construction of low-cost and high-performance large models.

[0110] 3. In view of the fact that traditional robot navigation methods are difficult to handle and predict dynamic changes in complex environments, and to achieve better human-computer interaction, the present invention innovatively integrates low-cost and high-performance small language models into the robot's multi-point real-time navigation, promotes the development of large-model-based robot navigation technology that does not rely on the network, and significantly enhances the robot's human-computer interaction and understanding of complex dynamic environments.

[0111] Different embodiments are provided for testing and verification of the multi-point robot navigation system FastNav with a fine-tuned adaptive small language model provided by the present invention.

[0112] Example 1: FastNav deployed on a real robot

[0113] 1.1 Environment Setup

[0114] Direct Drive Tech’s DIABLO robot was selected as the experimental platform and the FastNav system was deployed in a laboratory corridor environment.

[0115] 1.2 Model Selection and Deployment

[0116] TinyLlama-1.1B is selected as the SLM for the experiment, and the FastNav system is deployed on the NVIDIA Jetson Orin NX edge device.

[0117] 1.3 Experimental process

[0118] 1) Model fine-tuning:

[0119] TinyLlama-1.1B is fine-tuned using a navigation task dataset for a lab environment.

[0120] 2) Teacher-student iteration:

[0121] Perform multiple rounds of teacher-student iterations in a simulated environment to optimize model performance.

[0122] Keep detailed records of prompts, model outputs, and feedback adjustments for each iteration.

[0123] 3) Real environment testing:

[0124] Deploy the FastNav-processed model to a real robot to perform navigation tasks.

[0125] 1.4 Experimental Results

[0126] The experimental results show that TinyLlama-1.1B processed by FastNav has a higher success rate and lower navigation error in real environments, verifying the effectiveness of FastNav in practical applications. The unadjusted TinyLlama language model has a lower success rate and a larger navigation error. After 2 iterations and 6 iterations, the accuracy of the model gradually increased to a higher level, the navigation error gradually decreased, and the usability in the real environment steadily improved. Based on a series of data, the language model fully adjusted by FastNav also showed excellent performance in the real environment.

[0127] Example 2: FastNav does not include fine-tuning process

[0128] 2.1 Experimental Purpose

[0129] The purpose of this example is to verify the role of the fine-tuning module in FastNav, that is, to demonstrate its importance to the output format and success rate by omitting the fine-tuning step.

[0130] 2.2 System Settings

[0131] Choose a pre-trained SLM without fine-tuning as the base model.

[0132] 2.3 Omit the fine-tuning step

[0133] The pre-trained SLM is directly used for teacher-student iterations without any additional fine-tuning training.

[0134] 2.4 Teacher-Student Iteration Module

[0135] Teacher Model: Choose a knowledgeable large language model as the teacher model.

[0136] Student model: Use the unfine-tuned SLM as the student model.

[0137] Iterative process: the teacher model generates prompts, the student model tries to generate navigation target points based on the prompts, and the teacher model provides feedback based on the output of the student model.

[0138] 2.5 Navigation Controller

[0139] Use a standard robot navigation controller to receive the output of the student model and attempt to perform the navigation task.

[0140] 2.6 Experimental Results

[0141] Expected Result: Due to the lack of fine-tuning, SLM may have difficulty in generating output formats suitable for specific navigation tasks, resulting in low success rates.

[0142] 2.7 Conclusion

[0143] Through this experiment, we can conclude that fine-tuning is a key step to adapt SLM to domain-specific tasks, which significantly constrains the output format of SLM and improves the success rate of the task.

[0144] Example 3: FastNav does not include the teacher-student iteration process

[0145] 3.1 Experimental Purpose

[0146] The purpose of this example is to verify the role of the teacher-student iteration module in FastNav, that is, to demonstrate its importance in improving model performance by omitting the iteration process.

[0147] 3.2 System Settings

[0148] Choose an SLM that has been fine-tuned.

[0149] 3.3 Omitting the teacher-student iteration step

[0150] The fine-tuned SLM is used directly for the navigation task without involving teacher-student iterations.

[0151] 3.4 Navigation Controller

[0152] A standard robotic navigation controller is used to receive the output of the fine-tuned SLM and perform navigation tasks.

[0153] 3.5 Experimental Results

[0154] Expected results: Although the model has been fine-tuned, due to the lack of teacher-student iteration process, the performance of the model may not be further improved, resulting in limited increase in success rate.

[0155] 3.6 Conclusion

[0156] Through this experiment, it can be concluded that teacher-student iteration is an effective means to improve the performance of a fine-tuned SLM, which can significantly increase the success rate of the model by up to 30%-40%.

[0157] Through the test verification of the above-mentioned different embodiments, the technical solution of the present invention has significant technical advantages and broad application prospects. Through innovative model compression and deployment methods, it provides new directions and solutions for the development of robotics technology and has high practicality and industrialization potential.

[0158] In summary, the multi-point robot navigation system FastNav with adaptive small language model provided by the embodiment of the present invention is specifically designed to meet the needs of deploying a large language model locally on the robot side. Compared with the prior art, the present invention has significant features, which are specifically reflected in the following aspects:

[0159] 1. Technical advantages

[0160] 1) Low-cost and efficient deployment

[0161] By adopting the method of fine-tuning and teacher-student iteration, the present invention significantly reduces the demand for hardware resources, allowing high-performance language models to run on the resource-limited robot side. It avoids dependence on remote servers, improves the autonomy and stability of the system, and performs outstandingly in applications with poor network conditions or requiring high real-time performance. At the same time, the data is processed locally, which well protects the privacy and data security of users.

[0162] 2) High-performance task-specific processing capabilities

[0163] The adaptive compression method enhances the performance of the small language model in specific tasks, enabling it to efficiently handle complex tasks and dynamic environments. Through the teacher-student iterative process, the knowledge of the large model is effectively transferred to the small model, improving the latter's task performance.

[0164] 3) Enhanced knowledge transfer and learning capabilities

[0165] The performance of the model in practical applications can be optimized through continuous iteration to further improve its performance and adaptability. The teacher model generates prompts based on task requirements and environmental information, and repeatedly adjusts these prompts based on the performance feedback of the student model to ensure the learning effect of the student model.

[0166] 2. Performance indicators

[0167] 1) Compression rate and performance improvement

[0168] The model compression method of the present invention can enable a small language model to achieve performance comparable to that of a large model with 10 times the number of parameters on a specific robot navigation task while maintaining or approaching the performance of a large model.

[0169] 2) Enhanced edge computing capabilities

[0170] By deploying a small language model locally, the robot can complete tasks autonomously in an offline state, improving the capability and reliability of edge computing, reducing the delay in data transmission and processing, and enhancing the real-time response capability of the system.

[0171] 3. Production implementation

[0172] 1) Easy to integrate

[0173] The technical implementation threshold is low. It is based on the existing language model architecture and can be implemented through fine-tuning and training. It is easy to integrate into the existing robot system. Due to the simplicity of technical implementation and short deployment cycle, it can be quickly applied in the existing system. It reduces the demand for high-performance hardware and reduces production and maintenance costs. By significantly reducing the deployment cost, it improves the cost-effectiveness and is suitable for large-scale promotion and application.

[0174] 2) Adapt to various application scenarios

[0175] It is not only suitable for industrial robots, but also can be applied to robot systems in many fields such as medical and household, and has broad application prospects. The technical solution has strong adaptability and flexibility, and can be customized according to specific application requirements.

[0176] Industrial application prospects The technical advantages will significantly reduce the hardware cost and deployment difficulty of the robot system, improve the processing ability and adaptability of the model in specific tasks, and enable the robot to work efficiently in more practical application scenarios. These technical advantages will help promote the widespread application of large language models in the field of robotics and promote the development of embodied intelligence, with a wide range of application scenarios and market potential.

[0177] 3) Social and economic benefits

[0178] This technical solution is applicable to multiple fields such as industry, medical treatment, and household, and has broad application prospects. With the maturity and promotion of technology, it is expected that market demand will continue to grow, and it has huge market potential. By improving the application effect of robots in the fields of medical treatment, household, etc., the present invention will bring more convenience and benefits to society. In addition, reducing costs and improving efficiency will bring significant economic benefits to enterprises and promote the rapid development of related industries.

[0179] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. A multi-point robot navigation system with a fine-tuned adaptive small language model, characterized in that: The navigation system fine-tunes the pre-trained small language model, and obtains a model that can be used by the navigation system through a teacher-student iterative process. The model is directly connected to the navigation system, and a specified target point is given according to the user's instructions to guide the robot to navigate; wherein the navigation system includes: A model fine-tuning module generates a test data set by a human-in-the-loop generation method, and uses the data set to fine-tune the parameters of the small language model. The fine-tuned small language model outputs a context that can be used for robot navigation; A teacher-student iteration module implements the teacher-student iteration process, including a teacher model and a student model, wherein the teacher model is a large language model, and the student model is the small language model fine-tuned by the model fine-tuning module, and the teacher model is used to guide the student model; the teacher model generates prompts according to task requirements and environmental information, and adjusts the prompts according to the performance feedback of the student model; the student model attempts to complete the task, and the results of the student model's completion are used as the basis for the next round of prompts; A robot navigation controller module is responsible for converting the output of the small language model into actual navigation actions of the multi-point robot to achieve the multi-point navigation task; The model fine-tuning module uses the following steps to fine-tune the parameters of the small language model: S2.1: Generate the data set by the human-in-the-loop generation method; S2.2: Using the data set, using the LoRA method to perform low-rank adjustment on the model parameters of the small language model; S2.3: testing the adjusted small language model by inputting the navigation task test set to obtain an output result of the small language model; S2.4: extracting the target point in the output result, and calculating the accuracy of the small language model on the test set; S2.5: Determine whether the format of the output result and the accuracy meet the requirements. If not, repeat S2.2 to S2.5; S2.6: Output the fine-tuned small language model.

2. The navigation system according to claim 1, characterized in that During the fine-tuning process of the model fine-tuning module, the context output by the small language model adopts a fixed JSON format, and the context includes explanation and location information, wherein: The explanation is set as a string, showing the analysis and reasoning of the task by the small language model, and the reason for the success or failure of the task; The position is set to a list of number pairs containing the coordinates of the target point; The navigation system uses the position to navigate the robot and displays the robot's thought process through the interpretation.

3. The navigation system according to claim 2, characterized in that: The human in loop generation method comprises the following steps when generating the data set: S1.1: Write a first task based on a real-life scenario, and combine the first task and environmental information into a prompt; S1.2: The large language model generates a second task based on the prompt; S1.3: The human evaluator scores the second task, and the second task and the corresponding score are used as data in the dataset; S1.4: The second tasks and the scores are brought back to the large language model, and the large language model generates a new set of the second tasks according to the scores; S1.5: Repeat S1.3 to S1.4 until enough high-quality data sets are obtained.

4. The navigation system according to claim 3, characterized in that: During the fine-tuning process, the model fine-tuning module uses the PEFT library of Huggingface to quickly implement parameter fine-tuning, and adjusts the learning rate and the number of training rounds to obtain the accuracy of the small language model on the data set under different parameters.

5. The navigation system according to claim 4, characterized in that: In the teacher-student iteration module, the teacher model is set as a knowledgeable large language model. The teacher model acts as a prompt engineer, generates appropriate prompt content according to the current task and map environment information, obtains feedback from the results of the previous iteration, and adjusts the prompt content in real time according to the feedback; the student model receives the prompt content and generates corresponding output according to the current task. The output is recorded and used to provide feedback information to the teacher model in subsequent iterations.

6. The navigation system according to claim 5, characterized in that: The teacher-student iteration process includes the following steps: S3.1: Using the small language model fine-tuned by the model fine-tuning module as the student model; S3.2: inputting the navigation task test set into the teacher model and the student model; S3.3: extracting the output results of the student model and performing simulation experiments to obtain performance data of the student model; S3.4: Feeding back the performance data to the teacher model, and the teacher model providing teaching guidance to the student model; S3.5: Repeat S3.2 to S3.4 for multiple rounds of iterations to obtain the small language model that can finally be used.

7. The navigation system according to claim 6, characterized in that: The robot navigation controller module extracts a list of target points from the output of the small language model, guides the robot toward the target points, and updates the planned path in real time to help the robot navigate from the initial position to the expected points in sequence.

8. The navigation system according to claim 7, characterized in that: The small language model makes navigation decisions based on a pre-built map. After receiving a natural language instruction, the small language model uses the map to determine a series of target waypoints, which accurately reflect the intent of the natural language instruction and comply with the constraints of the map.

9. The navigation system according to claim 8, characterized in that: The natural language instruction consists of a series of words. The small language model maps the natural language instruction and the map information to the target waypoint sequence through the navigation decision. The navigation decision process is formalized by the function f: f(W,M)={p1,p2,…,p n }, Among them, W is a natural language instruction, which consists of a series of words composition, M is the map, p i is the coordinate or identifier of a location on the map M, i = 1, 2…, n.

Citation Information

Patent Citations

  • Small language model paraphrase generation method and device oriented to government affair information

    CN118798213A

  • Machine learning method for adjusting a preform

    WO2024142021A1

  • Natural language training and / or augmentation with large language models

    WO2024216304A1