Large language model-based ship autonomous navigation lightweight network training method
By constructing a lightweight network training method based on a large language model, using a multi-dimensional evaluation model to screen high-quality trajectory sets and dynamically optimize the evaluation model weights, the shortcomings of existing ship autonomous navigation systems in terms of generalization and real-time performance, data dependence and dynamic evaluation are solved, and highly reliable and low-latency autonomous decision support is achieved.
Patent Information
- Application Number
- CN202510762989.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-10-14
AI Technical Summary
Existing ship autonomous navigation systems have shortcomings in terms of the contradiction between generalization and real-time performance, data dependence and lack of dynamic evaluation, making it difficult to provide highly reliable and low-latency autonomous decision-making support in complex maritime environments.
By constructing a lightweight network training method based on a large language model, using the large language model to generate a set of candidate trajectories, combining safety, compliance and efficiency evaluation indicators to screen training trajectories, building a multi-dimensional quantitative evaluation model, and verifying the lightweight network output through actual navigation scenarios, the evaluation model weights are dynamically optimized.
It achieves highly reliable and low-latency autonomous decision support in complex maritime environments, taking into account the generalization ability of large language models and the real-time performance of lightweight networks, ensuring the accuracy and adaptability of decision instructions.
Smart Images

Figure CN120781075A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence and autonomous ship navigation, and in particular to a lightweight network training method for autonomous ship navigation based on a large language model. Background Art
[0002] Ship navigation decision-making systems are a core component for ensuring safe and efficient ship navigation. Their core task is to generate reasonable navigation instructions based on environmental perception, dynamic obstacle information, and navigation rules. Traditional systems primarily make decisions based on fixed rules or pre-built mathematical models, such as the rule-based encoding of the International Regulations for Preventing Collisions at Sea (COLREGs) or path planning algorithms based on kinematic models. These systems have the advantage of determinism in structured scenarios, but face complex and changing maritime environments (such as dense ship encounters, sudden weather changes, and equipment failures), lacking decision-making flexibility and difficulty adapting to the real-time needs of dynamic scenarios. In recent years, the development of deep neural networks and large language models (LLMs) has provided new ideas for intelligent decision-making. Their powerful reasoning and generalization capabilities are expected to address the limitations of traditional methods.
[0003] Currently, intelligent ship navigation systems fall into two main categories: one is based on rule engines, implementing decision logic through manually designed rule bases. This leads to high maintenance costs and a lack of coverage for long-tail scenarios. The other utilizes traditional reinforcement learning (RL) or supervised learning models, relying on manually labeled data or training in simulation environments. This suffers from high data acquisition costs and poor generalization. In existing deployments, large language models are being explored for decision generation due to their natural language understanding and logical reasoning capabilities. However, direct application faces bottlenecks such as large parameter counts and high inference latency, making it difficult to meet the real-time requirements of embedded devices and lacking closed-loop optimization capabilities in dynamic environments. Furthermore, system evaluations are often based on offline simulations or static rule verification, which prevents real-time correction of decision biases and leads to error accumulation.
[0004] Existing technologies have the following defects: (1) Conflict between generalization and real-time performance: rule engines and lightweight models have limited generalization capabilities, while large language models cannot meet low-latency requirements; (2) Data dependence: supervised learning requires a large amount of labeled data, and there are differences between simulated data and real-world scenarios, making it difficult to cope with sudden or rare events; (3) Lack of dynamic evaluation: existing evaluation mechanisms lack online self-verification capabilities, and small errors continue to amplify in the decision-execution cycle, resulting in reduced system reliability.
[0005] Therefore, there is an urgent need for a lightweight network training method for autonomous ship navigation based on a large language model. Summary of the Invention
[0006] In view of this, the present invention provides a lightweight network training method for autonomous navigation of ships based on a large language model, which limits the large language model to the non-real-time data generation link. Its output decision instructions are not directly applied to ship navigation control, but are used as candidate data for lightweight network training. By constructing a dynamic evaluation model, a training trajectory set is screened out to train the lightweight network, so that it has the dual advantages of knowledge generalization and fast reasoning. Combined with the closed-loop verification mechanism of actual navigation scenarios, the evaluation model is dynamically optimized to provide highly reliable and low-latency autonomous decision-making support for ship navigation.
[0007] To this end, the present invention provides the following technical solutions:
[0008] A lightweight network training method for autonomous ship navigation based on a large language model, comprising:
[0009] Collect ship navigation scene information to construct structured text and time series feature vectors;
[0010] The large language model outputs ship navigation decision instructions based on the structured text at preset periodic intervals; the scenario information contained in all structured texts and the output decision instructions within the training round are arranged in chronological order as a trajectory;
[0011] Construct a candidate trajectory set based on the output trajectory of each round of the large language model;
[0012] Calculate multiple evaluation indicators of candidate trajectories, assign weights to each evaluation indicator, and calculate the comprehensive score of each candidate trajectory through weighted summation;
[0013] Based on the comprehensive score, candidate trajectories that meet the conditions are selected from the candidate trajectory set as training trajectories to construct a training trajectory set;
[0014] The lightweight network is trained using the training trajectory set, and the difference between the actual trajectory score probability distribution and the training trajectory score probability distribution during the training process is characterized by discreteness; and the weight of each evaluation indicator is updated based on the discreteness.
[0015] Furthermore, the ship navigation scene information includes:
[0016] Ship status, obstacle information, environmental parameters and navigation tasks.
[0017] Furthermore, the evaluation indicators include:
[0018] Safety assessment indicators, compliance assessment indicators, and efficiency assessment indicators.
[0019] Furthermore, the termination conditions for the large language model to output a decision instruction include: the ship reaches the destination or the large language model reaches the maximum number of decision cycles.
[0020] Furthermore, the candidate trajectory set is constructed based on the output trajectories of each round of the large language model, which also includes:
[0021] The trajectories that complete the current round of navigation mission within the maximum number of decision cycles are selected as candidate trajectories, and a candidate trajectory set is constructed.
[0022] Furthermore, based on the comprehensive score, candidate trajectories that meet the conditions are screened out from the candidate trajectory set as training trajectories, including:
[0023] Define the screening threshold by the comprehensive score of historical trajectories;
[0024] Candidate trajectories with comprehensive scores greater than the screening threshold and safety assessment indicators that meet the standards are selected as training trajectories.
[0025] Furthermore, the weight of each evaluation indicator is updated based on the dispersion, including:
[0026] A discreteness threshold is defined, and if the discreteness is less than the discreteness threshold, the weight of each evaluation indicator is not updated;
[0027] If the discreteness is greater than the discreteness threshold, the weight of each indicator is updated according to the evaluation result of the actual trajectory.
[0028] Furthermore, the calculation formula of the screening threshold is:
[0029] θ(t)=μ(t)-k·σ(t)
[0030] Where θ(t) is the screening threshold, μ(t) is the mean of the comprehensive score in the sliding window, σ(t) is the standard deviation of the comprehensive score in the sliding window, k is the sensitivity coefficient, and the window size is W, which means W data items. For each new data item, the window slides forward once and the oldest data item is discarded.
[0031] Furthermore, the updating of the weight of each indicator according to the evaluation trajectory of the actual trajectory includes:
[0032]
[0033] Among them, α, β and γ represent weight factors, α″, β″ and γ″ are updated weight factors, Score the safety of the actual trajectory, Score the compliance of the actual trajectory, is the efficiency score of the actual trajectory, and λ is the learning rate.
[0034] Advantages and positive effects of the present invention:
[0035] This method leverages the large language model's deep understanding of structured scene information to generate candidate trajectory sets, breaking through the limitations of traditional rule bases. A multi-dimensional quantitative evaluation model is constructed to screen high-security, high-compliance, and high-efficiency trajectories to construct training trajectory sets, preventing low-quality data from contaminating the model. A lightweight network uses the large language model to output high-quality datasets for training, preserving the large language model's generalization reasoning capabilities while also meeting real-time requirements. The lightweight network output is verified through actual navigation scenarios, comparing the consistency of the evaluation model with the real-world results, and dynamically correcting the evaluation indicator weights to adapt to the evolution of complex scenarios and ensure the accurate generation of decision instructions. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0037] Figure 1 A flowchart of the steps for lightweight network training method for autonomous ship navigation based on a large language model;
[0038] Figure 2 This is a lightweight network decision flow chart for autonomous navigation of ships based on a large language model in an embodiment. DETAILED DESCRIPTION
[0039] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0040] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0041] The present invention provides a lightweight network training method for autonomous navigation of ships based on a large language model. The method converts the generalized reasoning ability of the large language model into a training trajectory set, establishes a multidimensional evaluation model and a dynamic screening mechanism, and performs weighted scoring on the trajectory set based on the three-dimensional indicators of safety, compliance, and efficiency, and screens out high-quality data as the training set for the lightweight network. At the same time, a closed-loop feedback and parameter adaptation method are designed, and the lightweight network output is verified through actual navigation scenarios. The consistency of the evaluation model and the real-scene results is compared, and the evaluation indicator weights are dynamically corrected.
[0042] Combine Figure 1 The steps of the lightweight network training method for ship autonomous navigation based on a large language model include:
[0043] A. Process ship navigation scene information.
[0044] A1. Construct structured text and time series feature vectors based on ship navigation scene information.
[0045] 1) Structured text, including:
[0046] ① Ship status: position coordinates (x, y), current speed v, current heading angle θ;
[0047] ②Obstacle information: obstacle type (ship, buoy, reef, etc.), obstacle location coordinates (x i ,y i ), relative distance d obs , relative speed v obs , the static obstacle speed is set to 0, the relative heading angle θ obs ;
[0048] ③Environmental parameters: ocean current velocity c v , wave height w h , visibility d vis ;
[0049] ④ Navigation mission: starting point coordinates (x0, y0), end point coordinates (x dst ,y dst ).
[0050] 2) Construct the time series feature vector as:
[0051] s=[x,y,v,θ,d obs ,v obs ,θ obs ,c v ,w h ,d vis ,x0,y0,x dst ,y dst ]∈Rn
[0052] Structured text is used as input to the large language model, and time series feature vectors are used as input to the lightweight network.
[0053] A2. Determine the ship's navigation decision.
[0054] Define the ship navigation decision cycle T c As the benchmark, each period is marked as k, and the time range is recorded as [t k ,t k +T c ), where t k =k×T c The starting moment of the cycle is , and in each decision cycle, the ship navigation decision instruction needs to be given according to the current scene information until the navigation mission is completed or the maximum number of decision cycles is exceeded.
[0055] B. Obtain a training trajectory set through a large language model.
[0056] The Large Language Model (LLM) uses a pre-trained language model based on the Transformer architecture. It fine-tunes the model through massive navigation logs and rule texts to adapt it to the decision-making reasoning of maritime navigation.
[0057] The input is structured text;
[0058] Output decision instruction a = (Δv, Δθ), including speed adjustment Δv and heading adjustment Δθ;
[0059] Training objective: Minimize binary cross entropy loss:
[0060]
[0061] Where N is the number of samples in the batch, y i ∈{0,1} represents the rationality label of the manual annotation instruction, p i Predict probabilities for the model.
[0062] The cyclic decision-making process of the large language model in the simulation environment is:
[0063] t k At this moment, enter structured text;
[0064] t k +T c At this moment, output decision instruction a t =(Δv t ,Δθ t );
[0065] The ship executes the command and moves to the state at time t+1;
[0066] t k+1 Moments, similarly, input structured text;
[0067] t k+1 +T c At this moment, output decision instruction a t+1 =(Δv t+1 ,Δθ t+1 );
[0068] Termination condition: When the ship completes the navigation mission and reaches the destination (x dst ,y dst ) or exceeds the maximum number of decision cycles T mac , ending this round of tasks.
[0069] The trajectory of the task completion round is taken as the candidate trajectory set: τ cand ={τ1,τ2,…,τ M}
[0070] Among them, τ k ={(s0,a0),(s1,a1),…,(s N ,a N )}, indicating that the trajectory information of a round contains N sets of corresponding scene-command pairs. The number of scene-command pairs corresponds to the number of decision cycles required to complete the task. For rounds that exceed the maximum number of decision cycles, the task is considered failed and the corresponding trajectory data will be directly filtered out.
[0071] C. Establish a dynamic evaluation model for autonomous navigation performance, quantitatively score each trajectory in the candidate trajectory set, and select high-quality data to train a lightweight network.
[0072] C1. Determine the evaluation indicators of the dynamic evaluation model for autonomous navigation performance.
[0073] In this embodiment, they are: safety evaluation indicators, compliance evaluation indicators, and efficiency evaluation indicators.
[0074] 1) Safety assessment indicators:
[0075] Establish a safety evaluation index S based on collision risk probability safe :Based on the dynamic relationship between the ship and the obstacle, the minimum approach distance (DCPA) between the ship and the obstacle and the shortest approach time (TCPA) for the ship and the obstacle to reach the minimum distance are calculated. The calculation formula is as follows:
[0076]
[0077] Where, is the relative position vector between the ship and the obstacle, is the obstacle relative velocity vector,
[0078] Safety evaluation index S safe for:
[0079]
[0080] In the formula, λ and μ are weight coefficients, σ d and σ t is an empirical constant.
[0081] 2) Compliance assessment indicators:
[0082] Based on the International Regulations for Preventing Collisions at Sea (COLREGs) and local waterway regulations, a rule matching scoring model evaluation index (S rule ):
[0083]
[0084] Where N is the total number of COLREGs rules, C i is the compliance judgment of the i-th rule, ω i is the rule weight coefficient, the default value is Σω i =1.
[0085] 3) Efficiency evaluation indicators:
[0086] The efficiency evaluation index (S) is calculated based on the path deviation loss and fuel consumption model. eff ):
[0087] Path deviation loss (E path ):
[0088]
[0089] Where, d i is the actual navigation distance of the i-th segment of the command path, d ideal is the theoretical shortest distance of the i-th segment, n is the number of path segments (usually divisible by time or distance), and its physical meaning is the degree of deviation between the actual path and the ideal path. The greater the deviation, the higher the loss.
[0090] Fuel consumption model (E fuel ):
[0091] Estimation of fuel consumption based on ship dynamics model:
[0092]
[0093] Where, E fuel is the fuel consumption of the current instruction, E max It is the historical maximum fuel consumption value. The physical meaning is that the closer the fuel consumption is to the historical optimal value, the higher the score.
[0094] Efficiency evaluation index (S eff ):
[0095] S eff =ρ·(100-E path )+(1-ρ)·S fuel
[0096] Where S fuel Indicates fuel consumption.
[0097] C2. Calculate a comprehensive score based on security assessment indicators, compliance assessment indicators, and efficiency assessment indicators.
[0098] The comprehensive score is calculated using a weighted summation with adjustable weights:
[0099] S totle =α·S safe +β·S rule +γ·S eff
[0100] Here, α, β, and γ represent weight factors, satisfying α+β+γ=1.
[0101] C3. Dynamically screen the candidate trajectory set based on the comprehensive score and construct the training trajectory set.
[0102] Define screening thresholds based on the scoring distribution of historical traces to adapt to data change trends:
[0103] θ(t)=μ(t)-k·σ(t)
[0104] Where μ(t) is the mean of the comprehensive score in the sliding window, σ(t) is the standard deviation of the comprehensive score in the sliding window, k is the sensitivity coefficient, and the window size is W, which means W pieces of data. For each new piece of data, the window slides forward once and the oldest data is discarded.
[0105] The set of successful turn candidate trajectories τ obtained from the large language model cand In the training set, select round trajectories with high comprehensive scores and safety standards to build a training trajectory set:
[0106] D train ={τ k |S totle (τ k )≥θ(t),S safe (τ k )≥S min}
[0107] Among them, τ k represents the training trajectory, S min To set the minimum security assessment indicator threshold.
[0108] D. Conduct lightweight network training for autonomous navigation based on the training trajectory set.
[0109] Lightweight Neural Networks (LNN) uses multi-layer neural network nesting and uses the training trajectory set for supervised training of the network model. The fully connected layer is used as the output layer, and the output instructions include speed adjustment and heading adjustment. LNN =(Δv LNN ,Δθ LNN );
[0110] The lightweight network uses mean square error to calculate the loss and introduces an optimization weight coefficient. If λ = 1, it means that the optimization weights of heading and speed are the same. If λ < 1, more attention is paid to heading accuracy.
[0111]
[0112] Where N is the number of samples in the batch, Δθ LNN is the heading adjustment value predicted by LNN for the i-th sample, Δθ train is the heading angle adjustment of the high-quality training set of the i-th sample, Δv LNN is the speed change predicted by LNN for the i-th sample, Δv train is the speed adjustment of the high-quality training set of the i-th sample, and λ is the weight coefficient of the speed term.
[0113] The trained lightweight network will be deployed in the ship's navigation decision-making system to generate decision instructions based on actual navigation scenario information and control the ship's autonomous navigation.
[0114] E. Closed-loop verification and optimization of autonomous navigation performance evaluation model.
[0115] By quantifying the difference between the actual trajectory output by the lightweight network and the trajectory evaluation results of the large language model training set, we ensure that the decision logic of the two remains consistent in dynamic scenarios, and thereby optimize the evaluation model parameters.
[0116] E1. Calculate the probability distribution of scores.
[0117] 1) Probability distribution of training set scores:
[0118] For the training trajectory set D in the same task train , for each trajectory τ k Construct its rating vector as Normalized training set scoring probability distribution:
[0119]
[0120] Where i is the scoring dimension index (security / compliance / efficiency), j is the sum index used for normalization, |D train | is the number of trajectories in the training set.
[0121] 2) Actual trajectory score probability distribution:
[0122] The lightweight network outputs actual trajectory information τ for the same task in actual deployment and operation real ={(s0,a0),(s1,a1),…,(s N ,a N )}, evaluate and score it to get S total (τ real ), and construct the scoring vector Normalized actual trajectory evaluation probability distribution:
[0123]
[0124] Where i is the scoring dimension index (security / compliance / efficiency), and j is the summation index used for normalization.
[0125] E2. Calculate the KL divergence to measure the difference between the actual trajectory score probability distribution and the training set score probability distribution.
[0126] The difference is calculated based on KL divergence (Kullback-Leibler Divergence):
[0127]
[0128] Where i is the scoring dimension index (security / compliance / efficiency), P real (i) is the probability distribution value of the score of each dimension of the actual trajectory output by the lightweight network, P train (i) is the probability distribution value of the score of each dimension of the training set trajectory, KL divergence D KL The larger the value, the more significant the difference between the two.
[0129] Set threshold D th , when D KL ≥D th When the decision logic of the training set generated by the lightweight network and the large prediction model deviates too much, it is considered that the weight adjustment needs to be triggered.
[0130] E3. Dynamic update of weights
[0131] Through feedback data from actual scenarios, the weights α, β, and γ of each indicator in the evaluation model are dynamically modified to optimize the screening criteria.
[0132] Collect the actual trajectory τ after completing the task based on the lightweight network realThe actual score is calculated based on the data results of the evaluation model: Safety score Compliance Score Efficiency score
[0133] Adjust each weight in turn:
[0134]
[0135] Normalization processing:
[0136]
[0137] Make α″+β″+γ″=1, α, β, γ are the original weights, α″, β″, γ″ are the new weights after adjustment, and λ is the learning rate, which controls the adjustment amplitude.
[0138] If the actual score is low, it means that the current weight does not pay enough attention to security and the weight value needs to be increased; if the score is high, the weight should be maintained or fine-tuned.
[0139] At this point, a round of closed-loop feedback is completed, the evaluation model weights are optimized and adjusted, and the next round of iteration, re-screening and training can be carried out based on the new parameter values.
[0140] Specific application examples
[0141] Combine Figure 2 The lightweight network training process for ship autonomous navigation based on a large language model includes:
[0142] A. Extract and process scene information to build input for large language models and lightweight networks.
[0143] In this example, the structured text reads: Current ship position (22.5°N, 118.0°E), speed 10 knots, heading angle 90°; Obstacle information: Relative distance 1.2 nautical miles, speed 8 knots, relative heading angle 75°; Environmental characteristics: Current speed 0.5 knots, wave height 2.5 meters, visibility 3 nautical miles; Navigation mission: Starting point (22.5°N, 118.0°E), destination (23.0°N, 118.5°E). Please provide the ship's decision-making instructions.
[0144] In this embodiment, the time series feature vector:
[0145] s=[22.5,118.0,10,90,1.2,8,75,0,5,2.5,3,22.5,118.0,23.5,118.5].
[0146] B. Generate candidate training datasets through large language models.
[0147] Input: structured text;
[0148] Model architecture: Deepseek_R1 model based on Transformer is used as the large language model base. A large number of group navigation scene Prompt-Response decision pairs are prepared for pre-training, covering collision avoidance, channel planning, rule application, etc. A low-rank matrix is constructed to fine-tune the model without modifying the original model parameters. In the simulation environment, the information of each task training is saved.
[0149] In this embodiment, the trajectory after 5 decision cycles is:
[0150]
[0151] The corresponding information is shown in Table 1.
[0152] Table 1
[0153]
[0154] The scene-instruction pair form of trajectories τ2 and τ3 is consistent with the above table. Different number of trajectories will be generated in different navigation tasks, among which there will be abnormal trajectories that make the task unable to be completed normally. For the trajectory τ3 of the reasoning timeout task failure, it is directly filtered out and not entered into the candidate training data set. Therefore, the output candidate training data set is: cand {τ1,τ2}.
[0155] C, Establish a dynamic evaluation model of autonomous navigation performance.
[0156] C1, Evaluate and calculate the candidate trajectory set according to ship dynamics
[0157] 1) Safety evaluation index: S safe (τ1) = 0.92, S safe (τ2) = 0.88;
[0158] 2) Compliance evaluation index:
[0159] τ1 meets the COLREGs rules in 5 decision cycles, so
[0160]
[0161] τ2 gives the instruction at t = 1 decision cycle violates the right turn limited rule, corresponding to the deduction, trajectory 2 executes 5 decision cycles, so
[0162]
[0163] 3) Efficiency evaluation index:
[0164] Efficiency evaluation indicators are calculated based on path deviation and fuel consumption:
[0165] S eff (τ1)=0.85,S eff (τ2)=0.70.
[0166] C2. Calculate the comprehensive score.
[0167] Set the initial weights α = 0.6, β = 0.3, γ = 0.1:
[0168] S total (τ1)=0.6×0.92+0.3×1.0+0.1×0.85=0.937
[0169] S total (τ2)=0.6×0.88+0.3×0.95+0.1×0.70=0.883
[0170] C3. Based on the comprehensive score, dynamically screen the candidate trajectory set and construct the training trajectory set.
[0171] In this embodiment, the sliding window W=50, the sensitivity coefficient k=1.2, the window mean μ(t)=0.85 and the standard deviation σ(t)=0.1 are calculated based on the recent instruction scoring results, and the screening threshold is calculated as:
[0172] θ(t)=μ(t)-k·σ(t)=0.91-1.2×0.05=0.85
[0173] Training trajectory set D train ={τ1,τ2}.
[0174] D. Training an autonomous navigation lightweight network based on the training trajectory set.
[0175] Input: Time series feature vector.
[0176] Lightweight network architecture: A recurrent neural network (RNN) is used as the lightweight network model, with a single-layer gated recurrent unit as the hidden layer to solve the vanishing gradient problem of traditional RNNs.
[0177] The fully connected layer serves as the output layer and outputs instruction a during the decision cycle. LNN =(Δv LNN ,Δθ LNN ), until the task is completed and the corresponding trajectory information τ is generated LNN .
[0178] The trained lightweight network will be deployed in the ship's navigation decision-making system to generate decision instructions based on actual navigation scenario information and control the ship's autonomous navigation.
[0179] E. Closed-loop verification and optimization of autonomous navigation performance evaluation model.
[0180] E1. Constructing probability distribution
[0181] Probability distribution of training trajectory set scores:
[0182] Based on the training trajectory set and the corresponding scoring results, construct the scoring vector:
[0183]
[0184] Calculate the probability distribution of the evaluation index:
[0185] ΣS safe =0.92+0.88=1.80
[0186] ΣS rule =1.0+0.95=1.95
[0187] ∑S effe =0.85+0.70=1.55
[0188]
[0189] Actual trajectory score probability distribution:
[0190] The lightweight network is used in actual deployment and operation for different scenario information of the same task real Verify that its output actual trajectory information τ real Compliance with requirements:
[0191] Structured text: Current ship position (22.5°N, 118.0°E), speed 10 knots, heading 90°; Obstacle information: Relative distance 1.5 nautical miles, speed 10 knots, relative heading 80°; Environmental characteristics: Current speed 0.5 knots, wave height 2.5 meters, visibility 3 nautical miles; Navigation mission: Starting point (22.5°N, 118.0°E), end point (23.0°N, 118.5°E). Please provide the ship's decision-making instructions.
[0192] Time series feature vector:
[0193] s=[22.5,118.0,10,90,1.5,10,80,0,5,2.5,3,22.5,118.0,23.5,118.5];
[0194] Actual trajectory τ real The scene-command information is shown in Table 2.
[0195] Table 2
[0196]
[0197] Evaluate and score the actual trajectory to get S total (τ real ), construct the rating vector:
[0198]
[0199] Normalized actual trajectory evaluation probability distribution:
[0200]
[0201]
[0202] E2. Calculate KL divergence:
[0203] D KL (P real ||P reain )≈0.192
[0204] In this implementation, the initial threshold D th =0.15, then D KL ≥D th , then the decision logic of the lightweight network and the large language model deviates too much, and weight adjustment is required.
[0205] E3. Dynamically update weights.
[0206] The lightweight network outputs the actual trajectory score vector: S real = [0.85, 0.90, 0.75], calculate the average rating vector of the training set trajectory based on τ1 and τ2: S real =[0.90,0.98,0.78]
[0207] Each weight is updated in turn, where the learning rate λ = 0.1:
[0208]
[0209] The normalization process obtains the final optimized weight of this round:
[0210]
[0211] At this point, a round of closed-loop feedback is completed, and the evaluation model weights are optimized and adjusted. The next round of iteration can be carried out based on the new parameter values, and re-screening and training can be carried out to ensure that the output of decision instructions is stable and reliable.
[0212] In summary, this method ensures high efficiency and real-time performance while enabling lightweight network output instructions to have the generalization ability and high reliability of large language models, which can effectively improve the reliability and adaptability of ship autonomous decision-making.
[0213] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A lightweight network training method for ship autonomous navigation based on a large language model, characterized by: include: Collect ship navigation scene information to construct structured text and time series feature vectors; The large language model outputs ship navigation decision instructions at preset periodic intervals based on structured text; Arrange all the scene information and decision instructions contained in the structured text in the training round in chronological order as a trajectory; construct a candidate trajectory set based on the output trajectory of each round of the large language model; Calculate multiple evaluation indicators of candidate trajectories, assign weights to each evaluation indicator, and calculate the comprehensive score of each candidate trajectory through weighted summation; Based on the comprehensive score, candidate trajectories that meet the conditions are selected from the candidate trajectory set as training trajectories to construct a training trajectory set; Using the training trajectory set to train a lightweight network, using discreteness to characterize the difference between the actual trajectory score probability distribution and the training trajectory score probability distribution during the training process; The weight of each evaluation indicator is updated based on the dispersion.
2. The method according to claim 1, characterized in that The ship navigation scene information includes: Ship status, obstacle information, environmental parameters and navigation tasks.
3. The method according to claim 1, characterized in that The evaluation indicators include: Safety assessment indicators, compliance assessment indicators, and efficiency assessment indicators.
4. The method according to claim 1, wherein The termination conditions for the large language model to output decision instructions include: the ship reaches the destination or the large language model reaches the maximum number of decision cycles.
5. The method according to claim 1, wherein The candidate trajectory set is constructed based on the output trajectory of each round of the large language model, which also includes: The trajectories that complete the current round of navigation mission within the maximum number of decision cycles are selected as candidate trajectories, and a candidate trajectory set is constructed.
6. The method according to claim 1, characterized in that Based on the comprehensive score, candidate trajectories that meet the conditions are selected from the candidate trajectory set as training trajectories, including: Define the screening threshold by the comprehensive score of historical trajectories; Candidate trajectories with comprehensive scores greater than the screening threshold and safety assessment indicators that meet the standards are selected as training trajectories.
7. The method according to claim 1, characterized in that The weight of each evaluation indicator is updated based on the dispersion, including: A discreteness threshold is defined, and if the discreteness is less than the discreteness threshold, the weight of each evaluation indicator is not updated; If the discreteness is greater than the discreteness threshold, the weight of each indicator is updated according to the evaluation result of the actual trajectory.
8. The method according to claim 6, characterized in that The calculation formula of the screening threshold is: θ(t)=μ(t)-k·σ(t) Where θ(t) is the screening threshold, μ(t) is the mean of the comprehensive score in the sliding window, σ(t) is the standard deviation of the comprehensive score in the sliding window, k is the sensitivity coefficient, and the window size is W, which means W data items. For each new data item, the window slides forward once and the oldest data item is discarded.
9. The method according to claim 7, characterized in that The weight of each indicator is updated according to the evaluation trajectory of the actual trajectory, including: Among them, α, β and γ represent weight factors, α″, β″ and γ″ are updated weight factors, Score the safety of the actual trajectory, Score the compliance of the actual trajectory, is the efficiency score of the actual trajectory, and λ is the learning rate.
Citation Information
Cited By
Large model training data measuring and screening method giving consideration to universality and safety performance
CN121303389A
Self-vehicle driving control method and device, electronic equipment and storage medium
CN121341228A