Unmanned ship parameter identification method and device based on multi-modal deep learning
Through multimodal deep learning technology, multi-source ocean data of unmanned boats in extreme sea conditions are processed, and feature information is extracted and integrated, which solves the problem of low sea conditions identification and performance indicator calculation efficiency of unmanned boats in extreme sea conditions, achieving high-precision safe navigation design and rapid verification.
Patent Information
- Application Number
- CN202411919859.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to effectively identify and process multi-source marine data in extreme marine environments, resulting in low efficiency of safe navigation control of unmanned boats in extreme marine conditions, and it is difficult to achieve high-precision sea condition identification and performance indicator calculation.
The multimodal deep learning method is adopted to process multimodal unmanned boat sea condition parameters through a convolutional neural network intelligent characterization learning model, extract multimodal feature information, and fusion of information through explicit definition and feature weighting algorithm to generate fused multi-source data to improve the accuracy and efficiency of sea condition recognition and performance indicator calculation.
The design criteria for safe navigation of unmanned boats in extreme sea conditions have been formed, the overall performance indicator demonstration and design efficiency of designers in extreme sea conditions have been improved, and the rapid verification and iteration capabilities of unmanned boats in extreme sea conditions have been enhanced.
Smart Images

Figure CN120045891A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous and safe navigation of unmanned boats, and in particular relates to a method and device for identifying parameters of unmanned boats based on multimodal deep learning. Background Art
[0002] Unmanned intelligent technology has become a representative technology of the world's cutting-edge military technological revolution. At present, the safe navigation of medium and large unmanned boats in extreme marine environments is an important bottleneck restricting the development of their combat capabilities.
[0003] The testing and evaluation of unmanned boats has become a hot topic of common concern internationally. Currently, many organizations have included this topic in their discussions, such as the Euron Good Experimental Methodology (GEM) Special Interest Group, the IEEE Robotics and Automation TCPEBRAS16 Technical Committee, the IEEE International Conference on Robotics and Automation Reproducible Experiment Research Section, and the International Conference on Intelligent Robots and Systems.
[0004] As for the safe navigation control of unmanned boats, foreign researcher Froude was one of the early pioneers in studying the rolling motion of ships. His work led to the invention of the bilge keel for anti-rolling. As early as 1874, Froude tried to use water tanks to reduce rolling motion. However, using anti-rolling water tanks to stabilize the ship's attitude is passive and uncontrolled, that is, it uses the physical properties of the facility itself to change the roll response characteristics of the ship, and does not exert active control.
[0005] Regarding the ship attitude stabilization control, foreign researcher Froude was one of the early pioneers in studying the ship's rolling motion. His work led to the invention of the bilge keel for anti-rolling. As early as 1874, Froude tried to use water tanks to reduce rolling motion. However, the use of anti-rolling water tanks to stabilize the ship's attitude is passive and uncontrolled, that is, the physical properties of the facility itself are used to change the ship's rolling response characteristics, and no active control is applied. The current ship attitude stabilization control technology is mainly based on anti-rolling control. At present, the sea state recognition of multi-source ocean data is troubled by data heterogeneity, environmental complexity, and many interference factors. It is manifested in the large feature differences caused by the diverse sources of multi-source ocean environment data, high fusion difficulty, strong nonlinearity and non-stationarity of the extreme sea state time series caused by the changes in the ocean environment, and the limited parameter recognition efficiency of the traditional sea state prediction model caused by the interference of many external factors, resulting in the lack of efficiency and reliability of the existing methods in practical applications. Therefore, how to provide an unmanned boat parameter recognition method based on multimodal deep learning has become a technical problem that needs to be solved in this field. Summary of the invention
[0006] The purpose of the present invention is to provide a method and device for unmanned boat parameter identification based on multimodal deep learning.
[0007] According to a first aspect of the present invention, a method for identifying parameters of an unmanned boat based on multimodal deep learning is provided, comprising:
[0008] Obtain multi-modal sea state parameters of unmanned boats;
[0009] Using an intelligent representation learning model based on a convolutional neural network, the multimodal unmanned boat sea state parameters are processed to obtain multimodal feature information corresponding to the multimodal unmanned boat sea state parameters;
[0010] Determine the contribution information and complementary information of the multi-modal unmanned vehicle sea state parameters in different tasks by using explicit definition and the multi-modal characteristic information;
[0011] A feature weighting algorithm is used to fuse the features with high information gain in the contribution information and the complementary information to obtain fused multi-source data corresponding to the multi-modal unmanned boat sea condition parameters. Optionally, the method further includes:
[0012] Determining, based on the multimodal unmanned boat sea state parameters, uncertainty estimates corresponding to the multimodal unmanned boat sea state parameters;
[0013] Calculate sample activation vectors of multimodal sample unmanned boat sea condition parameters for different sea conditions based on the uncertainty estimation and deep feature learning model;
[0014] Different sample activation vectors are processed to obtain the extreme value distribution and confidence information of sea conditions;
[0015] The bounds and uncertainties of the activation vectors are estimated using extreme value and evidence theory to generate a probability distribution model for each sea state;
[0016] The matching degree between the sample activation vector and the probability distribution model of each sea condition is calculated to obtain the probability that the sea condition parameters of the sample unmanned boat belong to different sea conditions.
[0017] Optionally, the convolutional neural network-based intelligent representation learning model is obtained in the following manner:
[0018] In the forward propagation stage, the training samples are taken out and input into the network to represent the ideal output. The corresponding actual output is calculated using the forward propagation rules. The information is transformed step by step through the layered structure and transmitted to the output layer.
[0019] In the back propagation stage, the back propagation of the convolutional neural network performs error calculation from the last output layer to the previous layer, and infers the previous pooling layer from the convolution layer or the previous convolution layer from the pooling layer. The input and output of each layer is a three-dimensional tensor.
[0020] Optionally, the method further comprises:
[0021] Calculate the error between the ideal output and the actual output;
[0022] The error of each layer is obtained by back propagation method;
[0023] Use the error of each layer to get the weight change and update the weight.
[0024] Optionally, the use of explicit definition and the multimodal feature information to determine contribution information and complementary information of the multimodal unmanned vehicle sea state parameters in different tasks includes:
[0025] After completing the multimodal feature extraction, use information gain or entropy information theory to quantify the information content of the feature;
[0026] The contribution of each feature to the fusion task is calculated through feature contribution evaluation. Information gain is used as a metric to evaluate the contribution of all features, and features with high information gain are selected as candidate features.
[0027] Analyze the complementarity between different features through mutual information;
[0028] According to the information gain and mutual information results, weights are assigned to each feature to achieve weighted selection of features.
[0029] According to a second aspect of the present invention, there is provided an unmanned boat parameter identification device based on multimodal deep learning, the device comprising:
[0030] An acquisition module is used to obtain multi-modal sea state parameters of unmanned boats;
[0031] A processing module, used to process the multimodal unmanned boat sea state parameters using an intelligent representation learning model based on a convolutional neural network to obtain multimodal feature information corresponding to the multimodal unmanned boat sea state parameters;
[0032] A contribution module, used to determine the contribution information and complementary information of the multi-modal unmanned vehicle sea state parameters in different tasks by using explicit definitions and the multi-modal feature information;
[0033] The weighting module is used to adopt a feature weighting algorithm to fuse the features with high information gain in the contribution information and the complementary information to obtain fused multi-source data corresponding to the multi-modal unmanned boat sea condition parameters.
[0034] Optionally, the weighting module is used to:
[0035] Determining, based on the multimodal unmanned boat sea state parameters, uncertainty estimates corresponding to the multimodal unmanned boat sea state parameters;
[0036] Calculate sample activation vectors of multimodal sample unmanned boat sea condition parameters for different sea conditions based on the uncertainty estimation and deep feature learning model;
[0037] Different sample activation vectors are processed to obtain the extreme value distribution and confidence information of sea conditions;
[0038] The bounds and uncertainties of the activation vectors are estimated using extreme value and evidence theory to generate a probability distribution model for each sea state;
[0039] The matching degree between the sample activation vector and the probability distribution model of each sea condition is calculated to obtain the probability that the sea condition parameters of the sample unmanned boat belong to different sea conditions.
[0040] Optionally, the acquisition module is further used for:
[0041] In the forward propagation stage, the training samples are taken out and input into the network to represent the ideal output. The corresponding actual output is calculated using the forward propagation rules. The information is transformed step by step through the layered structure and transmitted to the output layer.
[0042] In the back propagation stage, the back propagation of the convolutional neural network performs error calculation from the last output layer to the previous layer, and infers the previous pooling layer from the convolution layer or the previous convolution layer from the pooling layer. The input and output of each layer is a three-dimensional tensor.
[0043] Optionally, the weighting module is used to:
[0044] Calculate the error between the ideal output and the actual output;
[0045] The error of each layer is obtained by back propagation method;
[0046] Use the error of each layer to get the weight change and update the weight.
[0047] Optionally, the contribution module is used to:
[0048] After completing the multimodal feature extraction, use information gain or entropy information theory to quantify the information content of the feature;
[0049] The contribution of each feature to the fusion task is calculated through feature contribution evaluation. Information gain is used as a metric to evaluate the contribution of all features, and features with high information gain are selected as candidate features.
[0050] Analyze the complementarity between different features through mutual information;
[0051] According to the information gain and mutual information results, weights are assigned to each feature to achieve weighted selection of features.
[0052] In a third aspect, the present application shows an electronic device, which includes: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the method described in any of the above aspects.
[0053] In a fourth aspect, the present application illustrates a non-temporary computer-readable storage medium, which, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute a method as described in any of the above aspects.
[0054] In a fifth aspect, the present application illustrates a computer program product. When instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to execute the method described in any of the above aspects.
[0055] The beneficial effects brought by the present invention are as follows:
[0056] It can be seen from the above scheme that the embodiment of the present invention provides an unmanned boat parameter identification method and device based on multimodal deep learning, by acquiring multimodal unmanned boat sea state parameters; using an intelligent representation learning model based on a convolutional neural network to process the multimodal unmanned boat sea state parameters, and obtain multimodal feature information corresponding to the multimodal unmanned boat sea state parameters; using explicit definition and the multimodal feature information to determine the contribution information and complementary information of the multimodal unmanned boat sea state parameters in different tasks; using a feature weighting algorithm to perform feature weighting on the high information gain features in the contribution information and complementary information. Through fusion processing, fused multi-source data corresponding to the multi-modal unmanned boat sea condition parameters can be obtained, which can form design criteria for safe navigation of large and medium-sized unmanned boats under extreme sea conditions, effectively guide designers to scientifically carry out the demonstration and design of the overall performance indicators of unmanned boats under extreme sea conditions, and realize the rapid verification and iteration of the design of the safe navigation function and the calculation of performance indicators of unmanned boats under extreme sea conditions, providing strong technical support for solving and breaking through the design difficulties of key capabilities such as autonomous and safe navigation of large and medium-sized unmanned boats in extreme marine environments and breakthroughs in key technologies such as the calculation of the overall performance indicators of unmanned boats under extreme sea conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 A schematic diagram of a flow chart of an unmanned boat parameter identification method based on multimodal deep learning provided according to an embodiment;
[0058] Figure 2 A schematic diagram of a convolutional neural network model provided according to an embodiment;
[0059] Figure 3A schematic diagram of a pooling process provided according to an embodiment;
[0060] Figure 4 It is a structural block diagram of an unmanned boat parameter identification device based on multimodal deep learning in this application;
[0061] Figure 5 is a block diagram of an electronic device of the present application;
[0062] Figure 6 It is a block diagram of a computer-readable storage medium of the present application. DETAILED DESCRIPTION
[0063] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0064] Reference Figure 1 , shows a flowchart of the steps of a method for identifying unmanned boat parameters based on multimodal deep learning of the present application, which can be applied to electronic devices, wherein the method can specifically include the following steps:
[0065] S101, obtaining multi-modal sea condition parameters of the unmanned boat;
[0066] S102, using an intelligent representation learning model based on a convolutional neural network to process the multimodal unmanned boat sea state parameters to obtain multimodal feature information corresponding to the multimodal unmanned boat sea state parameters;
[0067] S103, using explicit definition and multi-modal feature information to determine the contribution information and complementary information of multi-modal unmanned boat sea state parameters in different tasks;
[0068] S104, using a feature weighting algorithm to fuse the features with high information gain in the contribution information and the complementary information, to obtain fused multi-source data corresponding to the multi-modal sea state parameters of the unmanned boat.
[0069] Specifically, the terminal device of the embodiment of the present application obtains multimodal sea state parameters of the unmanned boat, and while optimizing the fusion representation of multi-source sea state data, introduces an efficient deep feature learning method to optimize the application efficiency of multimodal data in deep feature extraction. The specific research method is: to study the efficient fusion method of multimodal data, by explicitly defining the amount of information of multi-source features to evaluate the contribution of different features to the task and their complementarity, and select feature combinations with high information gain in a feature weighted manner to improve the model's processing ability for heterogeneous data; to study the cross-modal uncertainty estimation of the representation, and combine it with deep feature learning technology, to perform extreme value distribution and confidence modeling on the activation vectors of different sea conditions, and to use extreme value / evidence theory to estimate the boundaries and uncertainties of the activation vectors, and to generate a probability distribution model for each sea condition, so as to calculate the matching degree between the sample activation vector and the extreme value distribution of each sea condition in the test phase, and obtain the probability that the sample belongs to different sea conditions. Through the above research method, a sea condition parameter identification model with efficient deep feature extraction capability is constructed to ensure accurate identification of sea condition parameters in extreme marine environments.
[0070] Another embodiment of the present application further supplements the unmanned boat parameter identification method based on multimodal deep learning provided in the above embodiment.
[0071] The research on the construction technology of multimodal sea state parameter identification model for the field of marine environmental monitoring evaluates the complementarity of multi-source features through information calculation, realizes the effective fusion and efficient representation of multimodal data, and combines the network structure and deep feature learning strategy research for cross-modal uncertainty estimation to build a multimodal sea state parameter identification model with high recognition accuracy and strong robustness in extreme marine environments. This part of the research achieves more accurate and robust sea state identification based on the evaluation of the complementarity of multimodal data, providing an environmental detection model for the project.
[0072] Optionally, the method further comprises:
[0073] According to the multi-modal sea state parameters of the unmanned boat, uncertainty estimates corresponding to the multi-modal sea state parameters of the unmanned boat are determined;
[0074] Based on uncertainty estimation and deep feature learning models, the sample activation vectors of multi-modal sample unmanned boat sea state parameters of different sea conditions are calculated;
[0075] Different sample activation vectors are processed to obtain the extreme value distribution and confidence information of sea conditions;
[0076] The bounds and uncertainties of the activation vectors are estimated using extreme value and evidence theory to generate a probability distribution model for each sea state;
[0077] The matching degree between the sample activation vector and the probability distribution model of each sea condition is calculated to obtain the probability that the sea condition parameters of the sample unmanned boat belong to different sea conditions.
[0078] Optionally, the intelligent representation learning model based on convolutional neural network is obtained by:
[0079] In the forward propagation stage, the training samples are taken out and input into the network to represent the ideal output. The corresponding actual output is calculated using the forward propagation rules. The information is transformed step by step through the layered structure and transmitted to the output layer.
[0080] In the back propagation stage, the back propagation of the convolutional neural network performs error calculation from the last output layer to the previous layer, and infers the previous pooling layer from the convolution layer or the previous convolution layer from the pooling layer. The input and output of each layer is a three-dimensional tensor.
[0081] Optionally, the method further comprises:
[0082] Calculate the error between the ideal output and the actual output;
[0083] The error of each layer is obtained by back propagation method;
[0084] Use the error of each layer to get the weight change and update the weight.
[0085] Optionally, explicit definition and multi-modal feature information are used to determine the contribution information and complementary information of multi-modal unmanned vehicle sea state parameters in different tasks, including:
[0086] After completing the multimodal feature extraction, use information gain or entropy information theory to quantify the information content of the feature;
[0087] The contribution of each feature to the fusion task is calculated through feature contribution evaluation. Information gain is used as a metric to evaluate the contribution of all features, and features with high information gain are selected as candidate features.
[0088] Analyze the complementarity between different features through mutual information;
[0089] According to the information gain and mutual information results, weights are assigned to each feature to achieve weighted selection of features.
[0090] The complementarity between the features of the current multi-source sea condition data cannot be effectively utilized, resulting in low recognition accuracy. Especially in the complex and changeable marine environment, the features of a single data source are difficult to provide enough information to achieve high-precision target recognition. Firstly, the multimodal features are extracted using an intelligent representation learning model based on convolutional neural networks. Then, the information content of multi-source features is explicitly defined and calculated, and their contribution and complementarity in different tasks are evaluated. Finally, feature weighting is used to select feature combinations with high information gain, and multi-source data is fused to achieve more accurate and efficient representation and recognition.
[0091] Similar to traditional neural networks, general deep neural networks are also composed of input layers, hidden layers, and output layers, but they differ greatly in many aspects such as structural principles. Convolutional neural networks mainly rely on convolutional layers and pooling layers for feature extraction and dimensionality reduction. The feature extraction of the entire image is achieved by sliding the convolution kernel to obtain a feature map. Each layer of nodes is partially connected to the previous layer, rather than the full connection of traditional neural networks. The pooling layer reduces the dimensionality of the feature map and makes the network have better displacement invariance. Its network model is as follows: Figure 2 As shown in Figure 2. Convolutional neural networks can directly input images into the network, avoiding the complex feature extraction process. Convolutional layers and pooling layers are important parts of feature extraction and dimensionality reduction and are the most representative structures.
[0092] like Figure 2 As shown in the figure, the convolution layer extracts features through convolution kernels and suppresses other components. The specific implementation process is: each convolution kernel is connected to the local receptive field of the previous feature map, that is, the convolution process is performed, and the bias is added and the activation function is passed to obtain the local features. The convolution kernel is moved on the image to extract each local feature, and the feature map of the layer is obtained by combining them. Different convolution kernels are used as different feature detectors to extract feature information and obtain different feature maps of the layer. In the convolution operation, the convolution kernel moves in the entire area with a certain step size, that is, the feature map obtained by the convolution process shares weights. This method provides an effective learning method for processing large images, which greatly reduces the number of model parameters.
[0093] Direct detection of the large amount of feature data extracted after convolution will lead to overfitting. In order to reduce training parameters, the pooling layer uses the local correlation of the image to perform subsampling while retaining useful information, reducing the amount of data and enhancing the network's adaptability to image displacement. Pooling usually uses the maximum value method and the mean value method. The network has good adaptability to the translation of the pooled area. Figure 3 It is a schematic diagram of the pooling process.
[0094] The training process of convolutional neural network includes the process of signal forward propagation, error back propagation and weight update, which are repeated continuously to achieve the purpose of weight optimization. The training algorithm mainly includes two stages and five steps: the first stage, the forward propagation stage: first take out the training sample input into the network to represent the ideal output; then use the forward propagation rule to calculate the corresponding actual output. In the forward propagation stage, the information is transformed layer by layer and transmitted to the output layer. The second stage, the back propagation stage: the back propagation of the convolutional neural network is only the same as the traditional neural network in the error calculation method from the last output layer to the previous layer, but the calculation method of inferring the previous pooling layer from the convolution layer or from the pooling layer to the previous convolution layer is very different from the traditional neural network. And the input and output of each layer is not a vector like in the neural network, but a three-dimensional tensor, that is, several matrices. First, calculate the error, that is, the difference between the ideal output and the actual output; then use the back propagation method to obtain the error of each layer; finally, use the error of each layer to obtain the weight change and update the weight.
[0095] After completing the multimodal feature extraction, information theory methods such as information gain or entropy are used to quantify the information content of the feature. Information entropy measures the uncertainty of random variables. The entropy of feature X is defined as follows:
[0096]
[0097] Where P(xi) is the probability of the i-th value in feature X. Information gain measures the amount by which the uncertainty of the target variable Y is reduced when feature X is known. It is defined as follows:
[0098] IG(Y,X)=H(Y)-H(Y|X)
[0099] Where H(X) is the entropy of the target variable Y, and H(Y|X) is the conditional entropy of Y given feature X. Based on the above definition, the contribution of each feature to the fusion task is calculated through feature contribution evaluation. Information gain is used as a metric to evaluate the contribution of all features, and features with high information gain are selected as candidate features. The complementarity between different features is then analyzed through mutual information. Mutual information measures the mutual dependence between two variables. The mutual information of features X and Z is defined as follows:
[0100]
[0101] Where P(x,z) is the joint probability distribution of X and Z, and P(x) and P(z) are the marginal probability distributions of X and Z, respectively.
[0102] According to the information gain and mutual information results, weights are assigned to each feature to achieve weighted selection of features. For feature Xi, its weight assignment parameter w iIt can be calculated based on information gain and complementarity:
[0103]
[0104] Where a and β are weight adjustment coefficients. The above method evaluates the contribution and complementarity of features by defining and calculating the amount of information, selects feature combinations with high information gain and strong complementarity, and realizes weighted selection and fusion of multi-source features. Based on the fused features, through network training and optimization, the recognition accuracy and robustness of the model in complex marine environments are improved, the efficient fusion and representation of multimodal data are realized, and the recognition performance of the model is enhanced.
[0105] The embodiments of the present application can form the design criteria for safe navigation of large and medium-sized unmanned boats under extreme sea conditions, effectively guide the designers to scientifically carry out the demonstration and design of the overall performance indicators of unmanned boats under extreme sea conditions, and realize the rapid verification and iteration of the design of the safe navigation function of unmanned boats under extreme sea conditions and the calculation of performance indicators, so as to solve and break through the design difficulties of key capabilities such as autonomous and safe navigation of large and medium-sized unmanned boats in extreme marine environments and the breakthrough of key technologies such as the calculation of the overall performance indicators of unmanned boats under extreme sea conditions. At the same time, the simulation design platform for safe navigation of unmanned boats under extreme sea conditions developed by this project provides convenient research means for designers, further improves the design efficiency and accuracy of designers, and shortens the research and development cycle of the overall performance indicator calculation, scheme design, and navigation decision-making and control technology related to safe navigation of unmanned boats in extreme sea conditions. Therefore, the research of this project has good military and social benefits. At the same time, the research results of this project can also be promoted and applied in the industry to the design of the safe navigation function of small surface unmanned boats in extreme sea conditions, and the potential application units are the Navy, Equipment Development, Military Science and Technology Commission, and national defense scientific research units of the Ship Group.
[0106] It should be noted that each implementable method in this embodiment may be implemented separately, or may be implemented in combination in any manner without conflict, and this application is not limited thereto.
[0107] The embodiment of the present invention provides an unmanned boat parameter identification method and device based on multimodal deep learning, which obtains multimodal unmanned boat sea state parameters; uses an intelligent representation learning model based on a convolutional neural network to process the multimodal unmanned boat sea state parameters to obtain multimodal feature information corresponding to the multimodal unmanned boat sea state parameters; uses explicit definition and multimodal feature information to determine the contribution information and complementary information of the multimodal unmanned boat sea state parameters in different tasks; uses a feature weighting algorithm to fuse the features with high information gain in the contribution information and complementary information to obtain fused multi-source data corresponding to the multimodal unmanned boat sea state parameters, which can form a design criterion for safe navigation of large and medium-sized unmanned boats under extreme sea conditions, effectively guide designers to scientifically carry out the demonstration and design of the overall performance indicators of unmanned boats under extreme sea conditions, and realize the rapid verification and iteration of the design of the safe navigation function of the unmanned boat under extreme sea conditions and the calculation of the performance indicators, so as to provide strong technical support for solving and breaking through the design difficulties of key capabilities such as autonomous and safe navigation of large and medium-sized unmanned boats in extreme marine environments and the breakthrough of key technologies such as the calculation of the overall performance indicators of unmanned boats under extreme sea conditions.
[0108] Another embodiment of the present application provides an unmanned boat parameter identification device based on multimodal deep learning, which is used to execute the unmanned boat parameter identification method based on multimodal deep learning provided in the above embodiment.
[0109] like Figure 4 , which is a schematic diagram of the structure of the unmanned boat parameter identification device based on multimodal deep learning provided in an embodiment of the present application. The unmanned boat parameter identification device based on multimodal deep learning includes an acquisition module 401, a processing module 402, a contribution module 403 and a weighting module 404, wherein:
[0110] The acquisition module 401 is used to obtain multi-modal sea condition parameters of the unmanned boat;
[0111] The processing module 402 is used to process the multimodal unmanned boat sea state parameters using an intelligent representation learning model based on a convolutional neural network to obtain multimodal feature information corresponding to the multimodal unmanned boat sea state parameters;
[0112] The contribution module 403 is used to determine the contribution information and complementary information of the multi-modal unmanned vehicle sea state parameters in different tasks by using explicit definition and multi-modal feature information;
[0113] The weighting module 404 is used to adopt a feature weighting algorithm to fuse the features with high information gain in the contribution information and the complementary information to obtain fused multi-source data corresponding to the multi-modal sea state parameters of the unmanned boat.
[0114] Regarding the device in this embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0115] Another embodiment of the present application further supplements the unmanned boat parameter identification device based on multimodal deep learning provided in the above embodiment.
[0116] Optionally, a weighting module is used to:
[0117] According to the multi-modal sea state parameters of the unmanned boat, uncertainty estimates corresponding to the multi-modal sea state parameters of the unmanned boat are determined;
[0118] Based on uncertainty estimation and deep feature learning models, the sample activation vectors of multi-modal sample unmanned boat sea state parameters of different sea conditions are calculated;
[0119] Different sample activation vectors are processed to obtain the extreme value distribution and confidence information of sea conditions;
[0120] The bounds and uncertainties of the activation vectors are estimated using extreme value and evidence theory to generate a probability distribution model for each sea state;
[0121] The matching degree between the sample activation vector and the probability distribution model of each sea condition is calculated to obtain the probability that the sea condition parameters of the sample unmanned boat belong to different sea conditions.
[0122] Optionally, the acquisition module is also used to:
[0123] In the forward propagation stage, the training samples are taken out and input into the network to represent the ideal output. The corresponding actual output is calculated using the forward propagation rules. The information is transformed step by step through the layered structure and transmitted to the output layer.
[0124] In the back propagation stage, the back propagation of the convolutional neural network performs error calculation from the last output layer to the previous layer, and infers the previous pooling layer from the convolution layer or the previous convolution layer from the pooling layer. The input and output of each layer is a three-dimensional tensor.
[0125] Optionally, a weighting module is used to:
[0126] Calculate the error between the ideal output and the actual output;
[0127] The error of each layer is obtained by back propagation method;
[0128] Use the error of each layer to get the weight change and update the weight.
[0129] Optionally, contribute modules to:
[0130] After completing the multimodal feature extraction, use information gain or entropy information theory to quantify the information content of the feature;
[0131] The contribution of each feature to the fusion task is calculated through feature contribution evaluation. Information gain is used as a metric to evaluate the contribution of all features, and features with high information gain are selected as candidate features.
[0132] Analyze the complementarity between different features through mutual information;
[0133] According to the information gain and mutual information results, weights are assigned to each feature to achieve weighted selection of features.
[0134] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0135] The embodiment of the present invention provides an unmanned boat parameter identification method and device based on multimodal deep learning, which obtains multimodal unmanned boat sea state parameters; uses an intelligent representation learning model based on a convolutional neural network to process the multimodal unmanned boat sea state parameters to obtain multimodal feature information corresponding to the multimodal unmanned boat sea state parameters; uses explicit definition and multimodal feature information to determine the contribution information and complementary information of the multimodal unmanned boat sea state parameters in different tasks; uses a feature weighting algorithm to fuse the features with high information gain in the contribution information and complementary information to obtain fused multi-source data corresponding to the multimodal unmanned boat sea state parameters, which can form a design criterion for safe navigation of large and medium-sized unmanned boats under extreme sea conditions, effectively guide designers to scientifically carry out the demonstration and design of the overall performance indicators of unmanned boats under extreme sea conditions, and realize the rapid verification and iteration of the design of the safe navigation function of the unmanned boat under extreme sea conditions and the calculation of the performance indicators, so as to provide strong technical support for solving and breaking through the design difficulties of key capabilities such as autonomous and safe navigation of large and medium-sized unmanned boats in extreme marine environments and the breakthrough of key technologies such as the calculation of the overall performance indicators of unmanned boats under extreme sea conditions.
[0136] Optionally, an embodiment of the present application further provides an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the various processes of the above-mentioned method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0137] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, each process of the above method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0138] Figure 5 800 is a block diagram of an electronic device 800 shown in the present application. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0139] Reference Figure 5 , the electronic device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .
[0140] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0141] The memory 804 is configured to store various types of data to support operations on the device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, images, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0142] The power supply component 806 provides power to the various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 800.
[0143] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and the rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.
[0144] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.
[0145] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: home button, volume button, start button, and lock button.
[0146] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the device 800, the relative positioning of components, such as the display and keypad of the electronic device 800, and the sensor assembly 814 can also detect the position change of the electronic device 800 or a component of the electronic device 800, the presence or absence of contact between the user and the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0147] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast operation information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0148] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.
[0149] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by a processor 820 of an electronic device 800 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0150] Figure 619 is a block diagram of a computer-readable storage medium 1900 shown in the present application. For example, the computer-readable storage medium 1900 may be provided as a server.
[0151] Reference Figure 6 , the computer-readable storage medium 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions, such as an application, that can be executed by the processing component 1922. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0152] The computer readable storage medium 1900 may also include a power supply component 1926 configured to perform power management of the computer readable storage medium 1900, a wired or wireless network interface 1950 configured to connect the computer readable storage medium 1900 to a network, and an input / output (I / O) interface 1958. The computer readable storage medium 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™ or the like.
[0153] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0154] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0155] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
[0156] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the present application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0157] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0158] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0159] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0160] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0161] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard drives, ROM, RAM, magnetic disks, or optical disks.
[0162] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0163] The above are preferred embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A method for identifying unmanned boat parameters based on multimodal deep learning, characterized in that: include: Obtain multi-modal sea state parameters of unmanned boats; Using an intelligent representation learning model based on a convolutional neural network, the multimodal unmanned boat sea state parameters are processed to obtain multimodal feature information corresponding to the multimodal unmanned boat sea state parameters; Determine the contribution information and complementary information of the multi-modal unmanned vehicle sea state parameters in different tasks by using explicit definition and the multi-modal characteristic information; A feature weighting algorithm is used to fuse the features with high information gain in the contribution information and complementary information to obtain fused multi-source data corresponding to the multi-modal unmanned boat sea condition parameters.
2. The unmanned boat parameter identification method based on multimodal deep learning according to claim 1 is characterized in that: The method further comprises: Determining, based on the multimodal unmanned boat sea state parameters, uncertainty estimates corresponding to the multimodal unmanned boat sea state parameters; Calculate sample activation vectors of multimodal sample unmanned boat sea condition parameters for different sea conditions based on the uncertainty estimation and deep feature learning model; Different sample activation vectors are processed to obtain the extreme value distribution and confidence information of sea conditions; The bounds and uncertainties of the activation vectors are estimated using extreme value and evidence theory to generate a probability distribution model for each sea state; The matching degree between the sample activation vector and the probability distribution model of each sea condition is calculated to obtain the probability that the sea condition parameters of the sample unmanned boat belong to different sea conditions.
3. The unmanned boat parameter identification method based on multimodal deep learning according to claim 1 is characterized in that: The convolutional neural network-based intelligent representation learning model is obtained in the following way: In the forward propagation stage, the training samples are taken out and input into the network to represent the ideal output. The corresponding actual output is calculated using the forward propagation rules. The information is transformed step by step through the layered structure and transmitted to the output layer. In the back propagation stage, the back propagation of the convolutional neural network performs error calculation from the last output layer to the previous layer, and infers the previous pooling layer from the convolution layer or the previous convolution layer from the pooling layer. The input and output of each layer is a three-dimensional tensor.
4. The unmanned boat parameter identification method based on multimodal deep learning according to claim 3 is characterized in that: The method further comprises: Calculate the error between the ideal output and the actual output; The error of each layer is obtained by back propagation method; Use the error of each layer to get the weight change and update the weight.
5. The unmanned boat parameter identification method based on multimodal deep learning according to claim 4 is characterized in that: The explicit definition and the multi-modal feature information are used to determine the contribution information and complementary information of the multi-modal unmanned vehicle sea state parameters in different tasks, including: After completing the multimodal feature extraction, use information gain or entropy information theory to quantify the information content of the feature; The contribution of each feature to the fusion task is calculated through feature contribution evaluation. Information gain is used as a metric to evaluate the contribution of all features, and features with high information gain are selected as candidate features. Analyze the complementarity between different features through mutual information; According to the information gain and mutual information results, weights are assigned to each feature to achieve weighted selection of features.
6. An unmanned boat parameter identification device based on multimodal deep learning, characterized in that: The device comprises: An acquisition module is used to obtain multi-modal sea state parameters of unmanned boats; A processing module, used to process the multimodal unmanned boat sea state parameters using an intelligent representation learning model based on a convolutional neural network to obtain multimodal feature information corresponding to the multimodal unmanned boat sea state parameters; A contribution module, used to determine the contribution information and complementary information of the multi-modal unmanned vehicle sea state parameters in different tasks by using explicit definitions and the multi-modal feature information; The weighting module is used to adopt a feature weighting algorithm to fuse the features with high information gain in the contribution information and the complementary information to obtain fused multi-source data corresponding to the multi-modal unmanned boat sea condition parameters.
7. The unmanned boat parameter identification device based on multimodal deep learning according to claim 6, characterized in that: The weighting module is used for: Determining, based on the multimodal unmanned boat sea state parameters, uncertainty estimates corresponding to the multimodal unmanned boat sea state parameters; Calculate sample activation vectors of multimodal sample unmanned boat sea condition parameters for different sea conditions based on the uncertainty estimation and deep feature learning model; Different sample activation vectors are processed to obtain the extreme value distribution and confidence information of sea conditions; The bounds and uncertainties of the activation vectors are estimated using extreme value and evidence theory to generate a probability distribution model for each sea state; The matching degree between the sample activation vector and the probability distribution model of each sea condition is calculated to obtain the probability that the sea condition parameters of the sample unmanned boat belong to different sea conditions.
8. The unmanned boat parameter identification device based on multimodal deep learning according to claim 6, characterized in that: The acquisition module is also used for: In the forward propagation stage, the training samples are taken out and input into the network to represent the ideal output. The corresponding actual output is calculated using the forward propagation rules. The information is transformed step by step through the layered structure and transmitted to the output layer. In the back propagation stage, the back propagation of the convolutional neural network performs error calculation from the last output layer to the previous layer, and infers the previous pooling layer from the convolution layer or the previous convolution layer from the pooling layer. The input and output of each layer is a three-dimensional tensor.
9. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method according to any one of claims 1 to 5 when executed by the processor.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Cited By
Lightweight unmanned aerial vehicle high-voltage line inhabitation point dynamic selection method and system, medium, equipment and product
CN121095808A