Underwater vehicle intelligent navigation large model pre-training method, system and medium
By constructing a large-scale intelligent navigation model for underwater vehicles based on the Transformer architecture, the problems of insufficient robustness and data mining in UV automatic control were solved, achieving more efficient multi-task processing and autonomous control capabilities, and reducing maintenance costs.
Patent Information
- Application Number
- CN202511416136.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2025-12-26
AI Technical Summary
Existing UV automatic control technology is based on classical control theory and suffers from problems such as insufficient robustness to complex environments, lack of learning and adaptive capabilities, difficulty in parameter adjustment, and insufficient data value mining.
By adopting a decoder-only Transformer architecture model and collecting data from multiple sources and unifying the formats, a large-scale intelligent navigation model for underwater vehicles is constructed. This model enables deep feature extraction and pre-training of state and control information. Furthermore, it combines a causal language model (CLM) for multi-task learning, thereby improving the model's generalization performance and multi-task processing capabilities.
It significantly enhances the model's generalization performance and multi-task processing capabilities, solves the error propagation problem in traditional multi-stage model pipelines, reduces maintenance costs, and improves the model's robustness and autonomous control capabilities for intelligent navigation.
Smart Images

Figure CN121209554A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent navigation of underwater vehicles, and in particular to an intelligent navigation large model pre-training method and system for underwater vehicles and a medium. BACKGROUND
[0002] Intelligent navigation technology for underwater vehicles (UV) is an important development direction in the field of ocean technology, and its development is driven by the growing demand for human development of marine resources, scientific exploration and national security. Traditional manned submarines and remote-controlled submarines (ROVs) have significant limitations in efficiency, safety and cost, especially in deep sea, complex terrain and high-risk areas, and traditional technology cannot meet the task requirements. Therefore, underwater vehicles with autonomous or semi-autonomous capabilities have become a key technology to solve the above problems. Intelligent navigation technology, as the core support of UV, significantly improves the autonomy, adaptability and task execution capability of the submarine through the integration of artificial intelligence (AI), sensor technology, automatic control and communication technology.
[0003] The development history of UV intelligent navigation technology can be divided into three stages: 1. Early stage (mid-20th century): UV mainly relies on pre-programmed paths and simple sensors, with limited functionality to basic data collection and limited area exploration, with low autonomy and environmental adaptability.
[0004] 2. Development stage (late 20th century to early 21st century): With the advancement of sensor technology, computer technology and communication technology, UV began to integrate multiple types of sensors (such as sonar, camera, etc.), and introduced preliminary path planning and obstacle avoidance algorithms, achieving semi-autonomous navigation capabilities.
[0005] 3. Intelligent stage (21st century to present): The breakthrough of artificial intelligence technology has promoted the intelligent development of UV, and the application of deep learning, machine learning and computer vision technologies has enabled the submarine to have the ability of environmental perception, autonomous decision-making, task execution and learning optimization, and intelligent navigation technology has gradually matured.
[0006] The implementation of intelligent navigation technology relies on the coordinated action of multiple core technologies: 1. Environmental perception technology: through sonar, camera, depth sensor, inertial navigation system and other devices, real-time acquisition of underwater environmental physical, chemical and biological information.
[0007] 2. Data processing and artificial intelligence algorithms: use deep learning, computer vision and other technologies to process sensor data to achieve obstacle recognition, terrain mapping and target detection.
[0008] 3. Path planning and decision-making technology: Based on mission objectives and environmental information, generate optimal navigation paths and dynamically adjust strategies to respond to environmental changes.
[0009] 4. Automation control technology: Through the control of actuators such as propellers and rudders, achieve stable navigation, precise steering and depth adjustment of the underwater vehicle.
[0010] 5. Communication and collaboration technology: Use acoustic communication, optical communication or underwater acoustic communication technology to realize information exchange between underwater vehicles, surface ships, other underwater vehicles or control centers, and support multi-underwater vehicle collaborative operation.
[0011] 6. Energy management and endurance optimization technology: According to task requirements and navigation conditions, optimize energy distribution strategies to extend the endurance of underwater vehicles.
[0012] 7. Learning and adaptive technology: Through machine learning algorithms and task experience accumulation, continuously optimize navigation strategies to improve task efficiency and success rate.
[0013] UV intelligent navigation technology has great application potential in many fields: 1. Ocean science research: Used for seabed topography mapping, marine biology investigation, hydrological data collection, etc., providing data support for marine ecosystem research and resource development.
[0014] 2. Resource exploration: Improve the efficiency and accuracy of oil, gas and mineral resource exploration and evaluation.
[0015] 3. Military and defense: Applied to underwater reconnaissance, anti-submarine warfare, mine detection and clearance, etc., to enhance national defense security capabilities.
[0016] 4. Engineering and maintenance: Used for inspection and maintenance of submarine pipelines and cables, reducing the risk and cost of manual operation.
[0017] 5. Environmental monitoring: Provides real-time data and long-term monitoring capabilities for marine pollution monitoring, coral reef protection and ecological assessment.
[0018] Although intelligent navigation technology has made significant progress, it still faces many challenges: 1. Complex environment adaptability: The complexity of underwater environments (such as ocean currents, terrain, temperature changes) puts higher requirements on the perception and decision-making capabilities of underwater vehicles.
[0019] 2. Communication limitations: Underwater communication is limited by the speed of sound propagation and signal attenuation, and efficient communication technology still needs to be broken through.
[0020] 3. Energy management: The energy supply and endurance of underwater vehicles directly affect the execution of long-term tasks.
[0021] 4. Multi-AUV Cooperation: Issues such as task allocation, communication synchronization, and conflict avoidance need to be addressed.
[0022] In the future, intelligent navigation technology will develop in the following directions: 1. Higher autonomy: Through reinforcement learning and adaptive algorithms, improve the autonomous decision-making ability of underwater vehicles in complex environments.
[0023] 2. Intelligent cooperation: Develop multi-AUV cooperative networks to achieve efficient execution of large-scale tasks.
[0024] 3. New energy technology: Explore new energy sources (such as fuel cells, wave energy) and energy optimization strategies to extend the endurance of underwater vehicles.
[0025] 4. Integration and modularity: Through modular design, improve the functional expandability and maintenance convenience of underwater vehicles.
[0026] In summary, underwater vehicle intelligent navigation technology is an important direction for the development of marine technology, and its core lies in the integration of artificial intelligence, sensor technology, automatic control, and communication technology to improve the autonomy, adaptability, and task execution capability of underwater vehicles. Despite the many challenges, intelligent navigation technology has broad application prospects in the fields of marine scientific research, resource exploration, military defense, and will continue to drive the development of cutting-edge marine technology, providing strong technical support for human exploration and development of the ocean.
[0027] Existing technical problems: Existing UV automatic control technology solutions are mostly based on classical control theory, using PID (Proportional-Integral-Derivative) control, fuzzy control, sliding mode control, and combinations and variants of these control methods to achieve precise control of UV's attitude, depth, speed, and heading. However, UV automatic control based on classical control theory has the following technical limitations: 1. Lack of robustness to complex environments: Underwater environments are complex and variable, with factors such as ocean currents, turbulence, and temperature changes. Classical control methods are difficult to dynamically adjust to respond to these changes.
[0028] 2. Lack of learning and adaptive ability: Classical control methods are usually fixed and cannot learn or optimize online based on system state or environmental changes.
[0029] 3. Parameter adjustment difficulty: Classical control methods (such as PID control) require manual adjustment of control parameters, and the parameter adjustment process is tedious and relies on expert experience.
[0030] 4. Lack of data value mining: There is a clear lack of data value mining, which cannot fully utilize the potential of data-driven optimization of control performance. SUMMARY
[0031] In view of the above problems, the present application provides an underwater vehicle intelligent navigation large model pre-training method, system and medium, which can not only significantly enhance the generalization performance and multi-task processing capability of the model, provide more reliable technical support for UV autonomous control, but also effectively solve the error propagation problem existing in the traditional multi-stage model pipeline, reduce the maintenance cost, and improve the robustness of the model.
[0032] To achieve the above object and other related objects, the technical solutions provided by the present application are as follows: An underwater vehicle intelligent navigation large model pre-training method, the method comprising: M1. Multi-source data collection and format unification: collecting data information of the control of unmanned underwater vehicles, remote control underwater vehicles and manned underwater vehicles in normal navigation state, the control in abnormal navigation state, multi-device collaborative navigation control, control in different navigation environments, control in different navigation depths and control in different navigation time periods, and performing preprocessing; M2. Construction of pre-training task: dividing each sample point at each time according to the execution process into two categories of state class information and control class information; M3. Construction of underwater vehicle intelligent navigation model: based on the current position and all historical information, using a Decoder-only Transformer structure model to predict the vector representation of the next position.
[0033] Further, the preprocessing includes feature alignment and unit unification of data from different sources, the feature alignment is matching data features collected by different devices, and the unit unification is converting data in different units into a unified unit.
[0034] Further, in step M2, the state information function State i and the control information Control i of the state model at the i-th time of each sample point at each time according to the execution process are divided into two categories. , , Wherein, F1 is the calculation formula of the state model, F2 is the calculation formula of the control model, Target i is the target information at the i-th time, Control i-1 is the control information at the i-1-th time, History j is the historical information at the j-th time, and n is the sample capacity.
[0035] Further, the state sequence prediction task, State0, State1,..., State n , aims to learn the changing trend of state information sequences under different devices and task types, and model the time-dependent relationship of state information. n The control sequence prediction task, Control0, Control1,..., Control n , aims to learn the changing trend of control information sequences under different devices and task types, and model the time-dependent relationship of control information.
[0036] Further, the state and control cross sequence prediction task, State0, Control0, State1, Control1,..., State n , Control n , aims to learn the alternating changing trend of state information and control information in time sequence, and model the dynamic interaction between the two. The state and control joint sequence prediction task, [State0, Control0], [State1, Control1],..., [State n , Control n ], aims to learn the collaborative changing law of state information and control information in the joint sequence, and model the overall correlation between the two. The state sequence classification task, [State0, State1,..., State i ] → Class i , where Class n is the i-th classification of the state sequence, aims to predict the device type or task type to which the given state information sequence belongs, and realize high-level semantic understanding of the sequence. The control sequence classification task, [Control0, Control1,..., Control i ] → Classp i , where Classp n is the i-th classification of the control sequence, aims to predict the device type or task type to which the given control information sequence belongs, and realize semantic mapping of the control information. The state and control joint sequence classification task, [State0, Control0, State1, Control1,..., State n , Control i ] → LClass iFor the i-th classification of the joint control sequence, the goal is to predict the device type or task type it belongs to based on the given state and control joint information sequence, and realize the semantic recognition of the joint sequence.
[0037] Further, the formula of the Decoder-only Transformer structure model is P(x1, x2,..., x T )=P(x1)·P(x2|x1)·P(x3|x1,x2)·...·P(x T |x1,x2,...x T-1 ), wherein x T is the historical information at time T, and T is a positive integer.
[0038] Further, the construction of the underwater vehicle intelligent navigation model includes model input dimension unification, and the model input dimension unification is to uniformly map the L-dimensional state information vector, the M-dimensional control information vector and the N-dimensional target information vector to an L+M+N-dimensional space.
[0039] In order to achieve the above object and other related objects, the present application further provides an underwater vehicle intelligent navigation large model pre-training system, comprising a computer device programmed or configured to perform the steps of the underwater vehicle intelligent navigation large model pre-training method.
[0040] In order to achieve the above object and other related objects, the present application further provides a computer readable storage medium having stored thereon a computer program programmed or configured to perform the underwater vehicle intelligent navigation large model pre-training method.
[0041] The present application has the following positive effects: 1. The present application realizes the deep feature extraction and pre-training of the multi-source historical operation data accumulated in the UV running process by constructing the pre-training data set, designing the large-scale model architecture and formulating the pre-training strategy, and finally obtains a high-performance UV intelligent navigation large-scale pre-training model. In the model architecture design, the Decoder-only Transformer structure is adopted, combined with the input dimension unification scheme, which effectively processes the multi-dimensional and multi-task data, and the complex relationship between the modeling state, control and target information. At the same time, through the multi-task learning strategy and the modeling method of the causal language model (CLM), the generalization ability and multi-task processing ability of the model are significantly improved. Compared with the method based on traditional control theory and supervised learning, the pre-training model constructed according to the scaling law (Scaling Law) theory can significantly enhance the generalization performance and multi-task processing ability of the model in theory, and provides more reliable technical support for UV autonomous control. 2. This invention significantly enhances the model's intelligent execution capability for complex tasks by leveraging the high-dimensional feature extraction capabilities of a large-scale pre-trained model and a collaborative optimization mechanism for multi-task learning. It supports mixed-task training and improves the model's generalization performance and prediction accuracy. Furthermore, based on dynamic interaction and overall correlation modeling of state and control information, it achieves accurate prediction and optimization of state and control information during UV navigation. Through a unified input dimension scheme and a Decoder-only Transformer structure design, it effectively solves the error propagation problem present in traditional multi-stage model pipelines, reduces maintenance costs, and improves model robustness. Attached Figure Description
[0042] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is a schematic diagram of the data sample of the present invention; Figure 3 This is a schematic diagram of the pre-training task of the present invention; Figure 4 This is a schematic diagram of the unified vector of the present invention; Figure 5 This is a schematic diagram of the intelligent navigation model of the underwater vehicle of the present invention. Detailed Implementation
[0043] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0044] Example 1: As Figure 1 As shown, a pre-training method for a large-scale intelligent navigation model of an underwater vehicle is provided, the method comprising: M1. Multi-source data collection and format unification: Collects data on the control of unmanned underwater vehicles, remotely operated underwater vehicles and manned underwater vehicles under normal navigation conditions, control under abnormal navigation conditions, multi-device collaborative navigation control, control under different navigation environments, control at different navigation depths and control within different navigation time periods, and performs preprocessing. M2. Construction of pre-training tasks: The sample points at each time step are divided into two categories based on the execution process: state information and control information. M3. Construction of intelligent navigation model of underwater submersible: based on the current position and all historical information, the vector representation of the next position is predicted by using the Decoder-only Transformer structure model.
[0045] In the embodiment, the preprocessing includes feature alignment and unit unification of data from different sources. The feature alignment is matching data features collected by different devices, and the unit unification is converting data in different units into a uniform unit.
[0046] In the embodiment, in step M2, the state information function State i and the control information Control i of the i-th moment of the state model of each time point of the sample point are divided into two categories according to the execution process. , , wherein F1 is the calculation formula of the state model, F2 is the calculation formula of the control model, Target i is the target information of the i-th moment, Control i-1 is the control information of the i-1-th moment, History j is the historical information of the j-th moment, and n is the sample capacity.
[0047] In the embodiment, the state sequence prediction task, State0, State1,..., State n , aims to learn the change trend of the state information sequence under different devices and task types, and model the time sequence dependence of the state information; the control sequence prediction task, Control0, Control1,..., Control n , aims to learn the change trend of the control information sequence under different devices and task types, and model the time sequence dependence of the control information.
[0048] In the embodiment, the state and control cross sequence prediction task, State0, Control0, State1, Control1,..., State n , Control n , aims to learn the alternating change trend of the state information and the control information in time sequence, and model the dynamic interaction relationship between the two; the state and control joint sequence prediction task, [State0, Control0], [State1, Control1],..., [State n , Controln The target is to learn the cooperative variation law of state information and control information in the joint sequence, and model the overall correlation between the two. State sequence classification task, [State0, State1,..., State n ]→Class i , wherein Class i is the i-th classification of the state sequence, and the target is to predict the device type or task type to which the given state information sequence belongs based on the given state information sequence, so as to realize high-level semantic understanding of the sequence. Control sequence classification task, [Control0, Control1,..., Control n ]→Classp i , wherein Classp i is the i-th classification of the control sequence, and the target is to predict the device type or task type to which the given control information sequence belongs based on the given control information sequence, so as to realize semantic mapping of the control information. State and control joint sequence classification task, [State0, Control0, State1, Control1,..., State n , Control n ]→LClass i , wherein LClass i is the i-th classification of the joint control sequence, and the target is to predict the device type or task type to which the given state and control joint information sequence belongs based on the given state and control joint information sequence, so as to realize semantic recognition of the joint sequence.
[0049] In this embodiment, the formula of the Decoder-only Transformer structure model is P(x1, x2,..., x T )=P(x1)·P(x2|x1)·P(x3|x1,x2)·...·P(x T |x1,x2,...x T-1 ), wherein x T is the historical information at time T, and T is a positive integer.
[0050] In this embodiment, the construction of the underwater vehicle intelligent navigation model includes model input dimension unification, which unifies the L-dimensional state information vector, the M-dimensional control information vector and the N-dimensional target information vector into an L+M+N-dimensional space.
[0051] Embodiment 2: Based on the underwater vehicle intelligent navigation large model pre-training method in embodiment 1, the present application is further described and explained.
[0052] AsFigure 1 As shown in a method for pre-training an intelligent navigation large model of an underwater vehicle, the method comprises: M1. Multi-source data collection and format unification: Collect data information of the control of unmanned underwater vehicles, remotely operated vehicles and manned underwater vehicles in normal navigation state, the control in abnormal navigation state, multi-device collaborative navigation control, control in different navigation environments, control in different navigation depths and control in different navigation time periods, and perform preprocessing; M2. Construction of pre-training task: According to the execution process, each time sample point is divided into state class information and control class information; M3. Construction of intelligent navigation model of underwater vehicle: Based on the current position and all historical information, a Decoder-only Transformer structure model is used to predict the vector representation of the next position.
[0053] I. Construction of pre-training data set Multi-source data collection and format unification: The pre-training large model needs to rely on large-scale data set for training, and the data set should have good multi-task coverage and multi-type coverage. In order to achieve this goal, data sets related to unmanned underwater vehicle (UV) control need to be collected from multiple sources. Data sources can include unmanned underwater vehicles (UUVs), remotely operated vehicles (ROVs), manned underwater vehicles (HOVs) and other devices. Task types should include but are not limited to: control of various devices in normal navigation state, control in abnormal navigation state, multi-device collaborative navigation control (including homogeneous devices and heterogeneous devices), control in different navigation environments, control in different navigation depths and control in different navigation time periods, etc.
[0054] After the collection of multi-source data set is completed, the data from different sources needs to be preprocessed, including feature alignment and unit unification. Feature alignment refers to matching the data features collected by different devices, for example, the control data of a UV with a perception device contains perception-related features, while the control data of a UV without a perception device lacks such features, and corresponding alignment processing is needed. Unit unification refers to converting data in different units to a unified unit, for example, angle data may be expressed in radians or degrees, and needs to be unified to one unit according to needs.
[0055] After feature alignment and unit unification, the data can be organized into time series data according to the device and task execution period, and stored in the form of CSV or Excel table. For example, the preprocessed feature vector can be represented as an N-dimensional vector, including X-axis linear velocity, Y-axis linear velocity, Z-axis linear velocity, X-axis angular velocity, Y-axis angular velocity, Z-axis angular velocity, X-axis coordinate, Y-axis coordinate, Z-axis coordinate, X-axis rotation angle, Y-axis rotation angle, Z-axis rotation angle, left sonar perception information, right sonar perception information, etc. The feature vector is filled in time dimension T in turn to form a two-dimensional table of T x N, that is, an operation data sample. The operation data does not need to be a complete task data set (i.e. containing all data samples from the beginning to the end of the task), and incomplete data (such as stopping halfway through the task) can also be used. In addition, there is an important information that the target information corresponding to the navigation task (such as the depth, heading angle, attitude, etc. to be reached in this navigation period).
[0056] As shown in Figure 2 , the first row is the table header composed of feature vectors, the first column is the time dimension, and the blue data area represents the specific values of the state, position, perception, control, etc. of the underwater vehicle changing with time.
[0057] Pre-training task design: As shown in Figure 3 , in the navigation data, each sample point at time t can be divided into two categories: state information and control information according to the execution process. Specifically, the state information at time t is determined by the state information, control information, target information and historical information at time t-1; and the control information at time t is determined by the state information, target information and historical information at time t.
[0058] The state information function State i of the i-th moment of the state model and the control information Control i of the i-th moment of each sample point at each moment according to the execution process are divided into two categories: , , where F1 is the calculation formula of the state model, F2 is the calculation formula of the control model, Target i is the target information at the i-th moment, Control i-1 is the control information at the i-1-th moment, History j is the historical information at the j-th moment, and n is the sample capacity.
[0059] In this embodiment, the state sequence prediction task is defined as State0, State1, ..., State2. n The goal is to learn the changing trends of state information sequences under different devices and task types, and to model the temporal dependencies of state information; to control sequence prediction tasks, Control0, Control1, ..., Control n The goal is to learn the changing trends of control information sequences under different devices and task types, and to model the temporal dependencies of control information.
[0060] In this embodiment, the state and control cross-sequence prediction task, State0,Control0,State1,Control1,...,State n Control n The goal is to learn the alternating trends of state information and control information over time and model the dynamic interaction between the two. Joint sequence prediction task for state and control, [State0,Control0], [State1,Control1], ..., [State0,Control0] n Control n The goal is to learn the coordinated change patterns of state information and control information in a joint sequence and to model the overall correlation between the two. State sequence classification task, [State0, State1, ..., State...] n ]→Class i , where Class i For the i-th classification of the state sequence, the goal is to predict the device type or task type to which it belongs based on the given state information sequence, thereby achieving a high-level semantic understanding of the sequence. Control sequence classification task, [Control0, Control1, ..., Control n ]→Classp i , among which, Classp i For the i-th classification of the control sequence, the goal is to predict the device type or task type to which it belongs based on a given control information sequence, thereby achieving semantic mapping of control information; State and control joint sequence classification task, [State0,Control0,State1,Control1,...,State] n Control n →LClass i Among them, LClass iFor the i-th classification of the joint control sequence, the goal is to predict the device type or task type it belongs to based on the given state and control joint information sequence, achieving semantic recognition of the joint sequence.
[0061] II. UV intelligent navigation model structure design Model input dimension unification: Because the dimensions of the state information vector, control information vector, and target information vector are different (L, M, N respectively), the data of different tasks cannot be directly mixed for training. To solve this problem, an input information dimension unification scheme is designed to map different information vectors to the same dimension space (L+M+N). The specific implementation method is as follows: For each vector, only fill the valid information in its corresponding dimension, and fill the rest of the dimensions with all-zero placeholders.
[0062] For example, the state information vector fills valid values in L dimensions, and zeros in M and N dimensions; the control information vector fills valid values in M dimensions, and zeros in L and N dimensions; the target information vector fills valid values in N dimensions, and zeros in L and M dimensions.
[0063] As shown in Figure 4 , through this design, the data of different tasks can be represented in a unified dimension space, supporting mixed task training.
[0064] As shown in Figure 5 , the formula of the Decoder-only Transformer structure model is, P(x1,x2,...,x T )=P(x1)·P(x2|x1)·P(x3|x1,x2)·...·P(x T |x1,x2,...x T-1 ), where x T is the historical information at time T, and T is a positive integer.
[0065] The input and output are all state, control, or joint vectors after dimension unification, with a dimension of L+M+N. The input vector first passes through the feature extraction module to obtain the corresponding feature vector. On the basis of the feature vector, position encoding and task type encoding are added as the input of the Decoder module. The decoding result vector passes through the output mapping module to obtain the unified vector representation of the next position.
[0066] The modeling method adopts the causal language model (CLM) modeling method, which predicts the vector representation at the next time based on historical information. The loss function can use mean square error (MSE) or other custom loss functions, and the specific choice depends on the task requirements. The optimization method uses batch gradient descent (Batch Gradient Descent) optimization method to improve training efficiency and model convergence.
[0067] The core design idea is to use the input dimension unification scheme and the Decoder-only Transformer structure design, so that the intelligent navigation pre-training large model can effectively process multi-dimensional and multi-task data, and model the complex relationship between state, control and target information. This design not only supports mixed task training, but also improves the generalization ability and prediction accuracy of the model, providing a solid technical foundation for the application of intelligent navigation.
[0068] In the implementation scheme of UV intelligent navigation pre-training large model, the core innovation points mainly lie in the following aspects: 1. Multi-source data collection and format unification: The data sources include unmanned underwater vehicles (UUV), remotely operated vehicles (ROV), manned underwater vehicles (HOV), and other devices. The task types cover normal navigation, abnormal navigation, multi-device cooperative navigation, different navigation environments, depth, and time period control. The feature alignment and unit unification are completed, and the data is organized into a time series (T×N two-dimensional table) and stored in CSV or Excel format. The navigation task target information (such as depth, heading angle, attitude, etc.) needs to be included in the data set.
[0069] 2. Pre-training task design: State sequence prediction task: modeling the time sequence dependence of state information. Control sequence prediction task: modeling the time sequence dependence of control information. State and control cross sequence prediction task: modeling the dynamic interaction between state and control information. State and control joint sequence prediction task: modeling the overall correlation between state and control information. State sequence classification task: predicting device or task type based on state sequence. Control sequence classification task: predicting device or task type based on control sequence. State and control joint sequence classification task: predicting device or task type based on joint sequence. Extensible goal-oriented sequence generation task to enhance model generalization ability.
[0070] 3. Model input dimension unification: The state information vector (L dimension), the control information vector (M dimension) and the target information vector (N dimension) are uniformly mapped to an L+M+N dimensional space. Each vector is only filled with valid information in its corresponding dimension, and the rest of the dimensions are filled with zero placeholders. Mixed training of different task data is supported.
[0071] 4. Intelligent navigation model structure design: The Decoder-only Transformer structure is adopted to predict the vector representation of the next position based on historical information. The input and output are all state, control or joint vectors after uniform dimension (L+M+N dimension). The input vector is input to the Decoder after feature extraction, position encoding and task type encoding. The decoding result vector is input to the output mapping module to obtain the uniform vector representation of the next position. The model is built in a causal language model (CLM) manner, and the loss function is mean square error (MSE) or other custom loss function. The optimization method is batch gradient descent.
[0072] In this embodiment, the present application provides an underwater vehicle intelligent navigation large model pre-training system, which comprises a computer device programmed or configured to perform the steps of the underwater vehicle intelligent navigation large model pre-training method.
[0073] In this embodiment, the present application provides a computer readable storage medium having a computer program stored thereon, which is programmed or configured to perform the underwater vehicle intelligent navigation large model pre-training method.
[0074] Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM) and the like.
[0075] In summary, the application can not only significantly enhance the generalization performance and multi-task processing capability of the model, but also provide more reliable technical support for UV autonomous control, effectively solve the error transmission problem existing in the traditional multi-stage model pipeline, reduce the maintenance cost, and improve the robustness of the model.
[0076] The specific embodiments described above do not constitute a limitation of the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall fall within the protection scope of the present disclosure.
Claims
1. A method for pre-training a large model for intelligent navigation of an underwater vehicle, characterized in that, The method comprises: M1. Multi-source data collection and format unification: Collecting data information of the operation of the unmanned underwater vehicle, remote control underwater vehicle and manned underwater vehicle in normal navigation state, abnormal navigation state, multi-device cooperative navigation operation, different navigation environments, different navigation depths and different navigation time periods, and preprocessing; M2. Construction of pre-training task: Dividing each time sample point into state class information and control class information according to the execution process; M3. Construction of intelligent navigation model of underwater vehicle: Based on the current position and all historical information, using a Decoder-only Transformer structure model to predict the vector representation of the next position. 2.The underwater vehicle intelligent navigation large model pre-training method according to claim 1, wherein: The preprocessing includes feature alignment and unit unification of different sources of data, the feature alignment is to match the data features collected by different devices, and the unit unification is to convert data of different units into a unified unit.
3. The underwater vehicle intelligent navigation large model pre-training method according to claim 2, characterized in that: After the feature alignment and unit unification of the different sources of data, the data can be organized into time series data according to the device and task execution period, and stored in the form of CSV or Excel table.
4. The underwater vehicle intelligent navigation large model pre-training method according to claim 1, wherein, In step M2, the sample point of each time is divided into the state information function State of the state model of the i-th time in two categories of state class information and control class information according to the execution process i and the control information Control of the i-th time i , , , Wherein, F1 is the calculation formula of state model, F2 is the control model calculation formula, Target i is the target information at the i-th moment, Control i-1 is the control information at the i-1-th moment, History j is the history information at the j-th moment, and n is the sample size.
5. The underwater vehicle intelligent navigation large model pre-training method according to claim 4, characterized in that: State sequence prediction task, State0, State1,..., State n , the goal is to learn the trend of state information sequence under different devices and task types, and model the time-dependent relationship of state information; control sequence prediction task, Control0, Control1,..., Control n , the goal is to learn the trend of control information sequence under different devices and task types, and model the time-dependent relationship of control information.
6. The underwater vehicle intelligent navigation large model pre-training method according to claim 5, characterized in that: State and Control Cross-Sequence Prediction Task, State0, Control0, State1, Control1,..., State n , n , the goal is to learn the alternating trend of state information and control information in time series, and model the dynamic interaction relationship between the two; State and control joint sequence prediction task, [State0, Control0], [State1, Control1],..., [State n , Control n ], the goal is to learn the collaborative change rule of state information and control information in the joint sequence, and model the overall correlation between the two; State Sequence Classification Task, [State0, State1,..., State n ] → Class i where Class i is the i-th classification of the state sequence, the goal is to predict the device type or task type it belongs to based on the given state information sequence, and to achieve high-level semantic understanding of the sequence; Control sequence classification task, [Control0, Control1,..., Control n ]→Classp i where Classp i is the i-th class of control sequence, the goal is to predict the device type or task type it belongs to based on the given control information sequence, and realize the semantic mapping of control information; State and control joint sequence classification task, [State0,Control0,State1,Control1,...,State] n Control n →LClass i Among them, LClass i For the i-th classification of the joint control sequence, the goal is to predict the device type or task type to which it belongs based on a given state and control joint information sequence, thereby achieving semantic recognition of the joint sequence.
7. The method of claim 1, wherein: The formula for the Decoder-only Transformer structure model is P(x1,x2,...,x...). T )=P(x1)·P(x2|x1)·P(x3|x1,x2)·...·P(x T |x1,x2,...x T-1 ), where x T This represents the historical information at time T, where T is a positive integer.
8. The intelligent navigation large model pre-training method for an underwater vehicle according to claim 1, characterized in that: The construction of the intelligent navigation model of the underwater vehicle includes model input dimension unification, which unifies the L-dimensional state information vector, M-dimensional control information vector and N-dimensional target information vector into L+M+N-dimensional space.
9. An underwater vehicle intelligent navigation large model pre-training system comprising a computer device, characterized in that, The computer device is programmed or configured to perform the steps of the underwater vehicle intelligent navigation large model pre-training method of any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program programmed or configured to perform the underwater vehicle intelligent navigation large model pre-training method of any one of claims 1-8.