Time series multimodal ship vibration identification method and device

By converting ship vibration data into image data and using a deep learning network model for identification, the problems of time series correlation loss and data redundancy in existing technologies are solved, achieving higher accuracy in identifying ship vibration characteristics.

CN120431438BActive Publication Date: 2025-09-26NAVAL UNIV OF ENG PLA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510933387.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-09-26
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

Existing vibration information analysis methods lose time-series correlation information, resulting in inaccurate identification of ship vibration characteristics and data redundancy.

Method used

The ship vibration data is converted into image data and recognized through a deep learning network model. The recognition accuracy is improved by using local feature extraction, multi-branch extraction and feature fusion units. The deep learning network model includes a local feature extraction unit, a multi-branch extraction unit, a feature fusion unit and a classifier unit.

Benefits of technology

Through improvements in graphical expression and deep learning network models, data redundancy can be reduced, the characteristics of ship vibration data can be fully extracted, and the accuracy of ship status identification can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431438B_ABST
    Figure CN120431438B_ABST
Patent Text Reader

Abstract

This application discloses a method and device for identifying time-series multimodal ship vibrations. The method includes: acquiring ship vibration data; converting the ship vibration data into image data; inputting the image data into a trained deep learning network model for identification, thereby obtaining a recognition result for the ship vibration data; wherein the deep learning network model includes a local feature extraction unit, a multi-branch extraction unit, a feature fusion unit, and a classifier unit; obtaining the recognition result for the ship vibration data includes: inputting the image data into the local feature extraction unit to obtain a first local feature map; inputting the local feature map into the multi-branch extraction unit to obtain multiple feature maps; inputting the multiple feature maps into the feature fusion unit to obtain a fused feature map; and inputting the fused feature map into the classifier unit to obtain a recognition result for the ship vibration data. This application can improve the recognition results of ship vibration data and increase the accuracy of ship status identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of vibration data recognition, and in particular to a method, device, storage medium and electronic device for time-series multimodal ship vibration recognition. Background Art

[0002] Vibration information analysis involves processing and analyzing vibration signals. It helps identify various components within the signal, including key parameters such as frequency, amplitude, and phase. These parameters provide valuable information about the vibration source, propagation path, and impact on the system. Vibration information analysis provides a deeper understanding of the physical nature and dynamic characteristics of vibration phenomena. This is crucial for solving practical engineering problems, improving product design, and optimizing production processes.

[0003] Current vibration information analysis methods typically use time-domain or frequency-domain analysis of vibration data. This approach suffers from three main drawbacks: 1) it loses the temporal correlation between data from different channels; and 2) it suffers from severe data redundancy, leading to inaccurate identification of a ship's vibration characteristics. Summary of the Invention

[0004] The present application provides a time-series multimodal ship vibration identification method, device, storage medium and electronic equipment, which can improve the recognition results of ship vibration data and improve the accuracy of ship status identification.

[0005] This application provides a time-series multimodal ship vibration identification method, comprising:

[0006] Obtain ship vibration data;

[0007] converting the ship vibration data into image data according to a mapping relationship between the ship vibration data and RGB channels of the image;

[0008] Inputting the image data into a trained deep learning network model for recognition to obtain a recognition result of the ship vibration data;

[0009] The deep learning network model includes a local feature extraction unit, a multi-branch extraction unit, a feature fusion unit, and a classifier unit; the image data is input into the trained deep learning network model for recognition to obtain the recognition result of the ship vibration data, including:

[0010] The image data is input into the local feature extraction unit to obtain a first local feature map; the local feature map is input into the multi-branch extraction unit to obtain multiple feature maps; the multiple feature maps are input into the feature fusion unit to obtain a fused feature map; the fused feature map is input into the classifier unit to obtain a recognition result of the ship vibration data.

[0011] Furthermore, in the above-mentioned time-series multimodal ship vibration recognition method, the step of converting the ship vibration data into image data according to a mapping relationship between the ship vibration data and RGB channels of an image comprises:

[0012] The ship vibration data includes ship vibration data in the x-direction, ship vibration data in the y-direction, and ship vibration data in the z-direction, and the ship vibration data in the x-direction is mapped to the R channel, the ship vibration data in the y-direction is mapped to the G channel, and the ship vibration data in the z-direction is mapped to the B channel to obtain a plurality of single-row color images;

[0013] The image data is obtained by folding the multiple single-row color images.

[0014] Furthermore, in the above-mentioned time series multimodal ship vibration identification method, the multi-branch extraction unit includes a first branch, a second branch, and a third branch, and the classifier unit includes a first classifier, a second classifier, and a third classifier; the first local feature map is input into the multi-branch extraction unit to obtain multiple feature maps, including:

[0015] Inputting the first local feature map into the first branch to obtain a second local feature map;

[0016] Inputting the first local feature map into the second branch and the third branch respectively to obtain a first global feature map and a second global feature map;

[0017] Inputting the fused feature map into the classifier unit to obtain the recognition result of the ship vibration data includes:

[0018] Inputting the fused feature map into the first classifier to obtain a first classification result;

[0019] Inputting the first global feature map into the second classifier to obtain a second classification result, and inputting the second global feature map into the third classifier to obtain a third classification result;

[0020] An identification result of the ship vibration data is obtained based on the first classification result, the second classification result, and the third classification result.

[0021] Furthermore, in the above-mentioned time series multimodal ship vibration recognition method, the first branch includes several feature extraction layers, each feature extraction layer includes multiple convolution blocks, each convolution block includes multiple bottleneck layers, and the bottleneck layer includes a convolution layer, a spatial convolution layer, and a residual connection block;

[0022] In the bottleneck layer, the input feature map is sequentially passed through a 1×1 convolution layer, a 3×3 spatial convolution layer, and a 1×1 convolution layer to obtain an output feature map.

[0023] Furthermore, in the above-mentioned time series multimodal ship vibration recognition method, the second branch includes a linear projection layer and multiple Transformer blocks, each of the Transformer blocks includes a multi-head attention module, a normalization layer, and a multi-layer perceptron connected in sequence, and the multi-head attention module and the multi-layer perceptron are further connected via a residual connection;

[0024] Inputting the first local feature map into the second branch to obtain a first global feature map, comprising:

[0025] Inputting the first local feature map into the linear projection layer to obtain a plurality of image blocks;

[0026] Pass the plurality of image blocks through the plurality of Transformer blocks to obtain the first global feature map.

[0027] Furthermore, in the above-mentioned time series multimodal ship vibration identification method, the third branch includes a plurality of hybrid modules, and the hybrid modules include two different types of multilayer perceptrons;

[0028] In the first type of multilayer perceptron, the input feature map is passed through a linear projection layer to obtain several image blocks. After the columns of each image block are transposed, the output feature map is obtained through a shared multilayer perceptron.

[0029] In the second type of multilayer perceptron, the input feature map is passed through a linear projection layer to obtain several image blocks. After the rows of each image block are transposed, the output feature map is obtained through a shared multilayer perceptron.

[0030] Furthermore, in the above-mentioned time series multimodal ship vibration recognition method, the feature fusion unit includes a 1×1 convolution layer and a downsampling module;

[0031] Inputting the plurality of feature maps into the feature fusion unit to obtain a fused feature map includes:

[0032] Inputting the first local feature map into the 1×1 convolutional layer and the downsampling module in sequence to obtain a first local feature map with adjusted channel number and spatial dimension;

[0033] The first local feature map after the number of channels and spatial dimensions are adjusted and the first global feature map and the second global feature map after upsampling are added together and then regularized to obtain the fused feature map.

[0034] The present application also provides a time-series multimodal ship vibration identification device, comprising:

[0035] An acquisition module is used to acquire ship vibration data;

[0036] a data conversion module, configured to convert the ship vibration data into image data according to a mapping relationship between the ship vibration data and the RGB channels of the image;

[0037] A recognition module is used to input the image data into a trained deep learning network model for recognition, thereby obtaining a recognition result of the ship vibration data;

[0038] The deep learning network model includes a local feature extraction unit, a multi-branch extraction unit, a feature fusion unit, and a classifier unit; the image data is input into the trained deep learning network model for recognition to obtain the recognition result of the ship vibration data, including:

[0039] The image data is input into the local feature extraction unit to obtain a first local feature map; the local feature map is input into the multi-branch extraction unit to obtain multiple feature maps; the multiple feature maps are input into the feature fusion unit to obtain a fused feature map; the fused feature map is input into the classifier unit to obtain a recognition result of the ship vibration data.

[0040] The present application also provides a computer-readable storage medium, in which a plurality of instructions are stored. The instructions are suitable for being loaded by a processor to execute any of the above-mentioned time-series multi-modal ship vibration identification methods.

[0041] The present application also provides an electronic device, including a processor and a memory, wherein the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used to perform the steps in any of the above-mentioned time-series multimodal ship vibration identification methods.

[0042] The present application provides a method, device, storage medium, and electronic device for identifying time-series multimodal ship vibrations. The present application converts ship vibration data into image data, and inputs the image data into a deep learning network model for identification to obtain an identification result. The present application graphically represents the data and then identifies the graphical data, which can reduce data redundancy and facilitate the extraction of ship vibration data features. Furthermore, the present application improves the structure of the deep learning network model, extracting local and global features through multiple branches and then fusing them, fully extracting the features of the image data, which can improve the recognition results of ship vibration data and improve the accuracy of ship status identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The following detailed description of the specific embodiments of the present application in conjunction with the accompanying drawings will make the technical solutions and other beneficial effects of the present application apparent.

[0044] Figure 1 This is a flowchart of the time-series multimodal ship vibration identification method provided in an embodiment of the present application.

[0045] Figure 2 This is the converted image data provided in the embodiment of the present application.

[0046] Figure 3 A schematic diagram of the structure of the deep learning network model provided in the embodiments of the present application.

[0047] Figure 4 This is a flowchart of the feature fusion unit processing provided in an embodiment of the present application.

[0048] Figure 5 This is a schematic diagram of the structure of the time-series multimodal ship vibration identification device provided in an embodiment of the present application.

[0049] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0050] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0051] The present invention provides a method, device, storage medium, and electronic device for identifying time-series multimodal ship vibrations. The device can be integrated into an electronic device, such as a terminal or server. The terminal can include a tablet computer, laptop computer, personal computer (PC), microprocessor, or other device.

[0052] See also Figure 1 , Figure 1 This is a flowchart of a method for identifying time-series multimodal ship vibrations provided in an embodiment of the present application, which is applied to an electronic device. The method for identifying time-series multimodal ship vibrations includes the following steps:

[0053] S1, obtain ship vibration data.

[0054] Specifically, the ship vibration data is acquired through an acceleration sensor. The ship vibration data is time series data and includes time series data in three directions: x, y, and z.

[0055] S2, converting the ship vibration data into image data according to the mapping relationship between the ship vibration data and the RGB channels of the image.

[0056] In one embodiment, step S2 includes the following steps:

[0057] S21, the ship vibration data includes the ship vibration data in the x direction, the ship vibration data in the y direction and the ship vibration data in the z direction. The ship vibration data in the x direction is mapped to the R channel, the ship vibration data in the y direction is mapped to the G channel, and the ship vibration data in the z direction is mapped to the B channel to obtain multiple single-row color images.

[0058] Specifically, the ship vibration data in the x, y, and z directions correspond to ax, ay, and az, respectively. There is a one-to-one mapping relationship between these data and the RGB channels of the image. The mapping relationship is as follows:

[0059] R channel:

[0060] G channel:

[0061] B channel:

[0062] S22, folding multiple single-row color images to obtain image data.

[0063] Figure 2 The converted image data provided in the embodiment of the present application is as follows: Figure 2As shown, the four small images on the left are images generated from ship vibration data while sailing in ice-free waters, with a predominantly green tint. The four small images on the right are images generated from ship vibration data while sailing in ice-covered waters. The vibrations generated by the ship's collision with sea ice are primarily parallel to the sea surface and are primarily low-frequency signals. This increases the Z-axis value of the acquired vibration data, which in turn appears purple after conversion to image data.

[0064] S3, input the image data into the trained deep learning network model for recognition, and obtain the recognition result of the ship vibration data.

[0065] The deep learning network model includes a local feature extraction unit, a multi-branch extraction unit, a feature fusion unit, and a classifier unit. The image data is input into the trained deep learning network model for recognition, and the recognition results of the ship vibration data are obtained, including:

[0066] The image data is input into the local feature extraction unit to obtain a first local feature map; the local feature map is input into the multi-branch extraction unit to obtain multiple feature maps; the multiple feature maps are input into the feature fusion unit to obtain a fused feature map; the fused feature map is input into the classifier unit to obtain the recognition result of the ship vibration data.

[0067] Figure 3 A schematic diagram of the structure of the deep learning network model provided in the embodiment of the present application is shown in FIG. Figure 3 As shown, the multi-branch extraction unit includes a first branch, a second branch, and a third branch, and the classifier unit includes a first classifier, a second classifier, and a third classifier; the image data is input into the trained deep learning network model for recognition, and the recognition results of the ship vibration data are obtained, including:

[0068] Inputting the first local feature map into the first branch to obtain a second local feature map;

[0069] Input the first local feature map into the second branch and the third branch respectively to obtain the first global feature map and the second global feature map;

[0070] The fused feature map is input into the classifier unit to obtain the recognition results of the ship vibration data, including:

[0071] Input the fused feature map into the first classifier to obtain the first classification result;

[0072] Inputting the first global feature map into the second classifier to obtain a second classification result, and inputting the second global feature map into the third classifier to obtain a third classification result;

[0073] The recognition result of the ship vibration data is obtained based on the first classification result, the second classification result and the third classification result.

[0074] From the above steps, it can be seen that step S3 includes the following steps:

[0075] S31, inputting the image data into a local feature extraction unit to obtain a first local feature map.

[0076] Specifically, the local feature extraction unit includes a 7×7 convolution layer (with a stride of 2) and a 3×3 maximum pooling layer (with a stride of 2). The task of the local feature extraction unit is to extract initial local features, such as edge and texture information, in preparation for subsequent processing.

[0077] S32: Input the first local feature map into the first branch to obtain a second local feature map.

[0078] Among them, the first branch includes several feature extraction layers, each feature extraction layer includes multiple convolution blocks, each convolution block includes multiple bottleneck layers, and the bottleneck layer includes a convolution layer, a spatial convolution layer and a residual connection block.

[0079] In the bottleneck layer, the input feature map is sequentially passed through a 1×1 convolution layer, a 3×3 spatial convolution layer, and a 1×1 convolution layer to obtain the output feature map.

[0080] Specifically, the first branch can be a CNN (Convolutional Neural Network) branch. The CNN branch adopts a feature pyramid structure, which means that as the network depth increases, the resolution of the feature map gradually decreases, but the number of channels increases. The entire CNN branch is divided into four feature extraction layers, each of which includes multiple convolutional blocks, and each convolutional block contains n bottleneck layers. The bottleneck layer includes a 1×1 convolutional layer (for reducing the number of channels), a 3×3 spatial convolutional layer, a 1×1 convolutional layer (for restoring the number of channels), and a residual connection between the input and output. In the specific example, n is set to 1 for the first convolutional block, and n is set to ≥ 2 in the subsequent N-1 convolutional blocks.

[0081] S33: Input the first local feature map into the second branch to obtain a first global feature map.

[0082] The second branch includes a linear projection layer and multiple Transformer blocks. Each Transformer block includes a multi-head attention module, a normalization layer, and a multi-layer perceptron connected in sequence. The multi-head attention module and the multi-layer perceptron are also connected through residual connections.

[0083] Specifically, the second branch can be a Transformer branch, which includes N repeated Transformer blocks. Each Transformer block consists of a multi-head self-attention module and a multi-layer perceptron (MLP block). Before each layer, a normalization layer (layer norm) is applied, and residual connections are used in both the self-attention layer and the MLP block to ensure effective information transfer and gradient flow.

[0084] Step S33 includes:

[0085] S331: Input the first local feature map into a linear projection layer to obtain several image blocks.

[0086] To process the first local feature map from the local feature extraction unit, a linear projection layer compresses it into 14×14 patch embeddings. The linear projection layer is a 4×4 convolution with a stride of 4. Given that the CNN branch (3×3 convolution) already encodes local features and spatial position information, additional position embeddings are no longer required. This design helps improve image resolution for downstream vision tasks while still allowing the Transformer branch to focus on higher-level global feature modeling. This meticulous structural design contributes to better performance and higher image quality.

[0087] S332: Pass the plurality of image blocks through a plurality of Transformer blocks to obtain a first global feature map.

[0088] S34: Input the first local feature map into the third branch to obtain a second global feature map.

[0089] The third branch contains multiple hybrid modules, which include two different types of multilayer perceptrons.

[0090] In the first type of multilayer perceptron, the input feature map is passed through a linear projection layer to obtain several image blocks. After the columns of each image block are transposed, the output feature map is obtained through a shared multilayer perceptron.

[0091] In the second type of multilayer perceptron, the input feature map is passed through a linear projection layer to obtain several image blocks. After the rows of each image block are transposed, the output feature map is obtained through a shared multilayer perceptron.

[0092] Specifically, the third branch can be an MLP branch, and the mixing module is a Mixer architecture. The Mixer architecture employs two different types of MLP layers: the token-mixing MLP and the channel-mixing MLP. The token-mixing MLP allows information exchange between different spatial locations (tokens). It is an MLP applied across patches, meaning it helps "mix" spatial information. The channel-mixing MLP, on the other hand, allows information exchange between different channels. It is an MLP applied independently to image patches, meaning it helps "mix" features at each location. Specifically, the token-mixing MLP block operates on the columns of each patch, first transposing the patches, then applying a shared MLP1 to compute the output, and finally transposing again to obtain the final result. The channel-mixing MLP block operates on the rows of each patch, with all rows sharing a common MLP for computation. These two types of MLP layers are executed alternately to promote information exchange between different dimensions, thereby helping the model better understand and utilize feature information in the image. This design helps improve the performance of the Mixer architecture, making it suitable for various visual tasks.

[0093] S35, inputting multiple feature maps into the feature fusion unit to obtain a fused feature map.

[0094] Specifically, the feature fusion unit (FCU) consists of a 1×1 convolutional layer and a downsampling module. Figure 4 The flowchart of the feature fusion unit processing provided in the embodiment of the present application is as follows: Figure 4 As shown in Figure 1, the FCU acts as a bridging module, fusing local features from the CNN branch with the global representations from the Transformer and MLP branches. The FCU is applied starting from the second block, as the initialization features for all three branches are identical. Throughout the branch structure, the FCU interactively and progressively fuses feature maps and patch embeddings, ensuring an organic and synergistic integration of information. This module helps the model comprehensively consider local and global information at different levels, improving its performance.

[0095] In one embodiment, step S35 includes the following steps:

[0096] S351, input the first local feature map into the 1×1 convolution layer and the downsampling module in sequence to obtain the first local feature map after the number of channels and spatial dimension are adjusted;

[0097] S352: The first local feature map after the number of channels and spatial dimensions are adjusted and the first global feature map and the second global feature map after upsampling are added together and regularized to obtain a fused feature map.

[0098] In order to solve the misalignment problem between the second local feature map output by the CNN branch and the image patches in the Transformer branch and the MLP branch, FCU is proposed, which continuously couples local features and global representations in an interactive manner.

[0099] First, it's important to consider the mismatch in feature dimensions between CNNs, Transformers, and MLPs. CNN feature maps have dimensions of C×H×W (where C, H, and W represent channels, height, and width, respectively), while patch embeddings have a shape of (K+1)×E, where K, 1, and E represent the number of image patches, class tokens, and embedding dimensions, respectively. To align these dimensions, when the feature maps are fed into the Transformer and MLP branches, they first pass through a 1×1 convolutional layer to adjust the number of channels. Then, a downsampling module aligns the spatial dimensions. Finally, the feature maps are summed with the patch embeddings. When fed back from the Transformer and MLP branches to the CNN branch, the patch embeddings are upsampled to align the spatial scales. Then, a 1×1 convolutional layer aligns the channel dimension with the CNN feature map dimension, and the two are summed. Layer Norm and Batch Norm modules are also used to normalize the features to ensure stability and consistency.

[0100] On the other hand, there is a significant semantic gap between the second local feature map and the patch embeddings (the first and second global feature maps). The second local feature map is extracted from local convolution operators, while the patch embeddings are derived through global self-aggregation and attention mechanisms. Therefore, FCU is applied in each block (except the first) to gradually fill the semantic gap and ensure that the model can effectively fuse local and global information. This design helps maintain feature consistency and information coherence, thereby improving the performance and expressiveness of deep learning models.

[0101] S36: Input the fused feature map into the first classifier to obtain a first classification result.

[0102] S37: Input the first global feature map into the second classifier to obtain a second classification result, and input the second global feature map into the third classifier to obtain a third classification result.

[0103] Specifically, for the Transformer branch and the MLP branch, class tokens are extracted from the first global feature map and the second global feature map, and then fed into the second classifier and the third classifier respectively.

[0104] S38, obtaining a recognition result of the ship vibration data based on the first classification result, the second classification result, and the third classification result.

[0105] In step S3, the global context information from the Transformer branch is first imported into the convolutional feature map to enhance the global perception capabilities of the CNN branch. This allows for a better understanding of the overall image context, thereby improving the model's global perception and comprehension capabilities. Secondly, local features from the CNN branch are gradually fed back into the patterning, which helps enrich the local detail information of the Transformer and MLP branches. This allows the model to better capture and utilize local features in the image, improving its sensitivity to detail. Finally, the patterning of the MLP and Transformer branches is fused together, further enhancing the MLP branch's capabilities in spatial interaction. This process helps the model better understand the relationships between different regions, thereby improving the effect of spatial interaction. Through the complementary nature of these three, more powerful representation learning will be achieved, improving the performance and expressiveness of the model.

[0106] This application converts ship vibration data into image data and inputs the image data into a deep learning network model for recognition to obtain a recognition result. This application presents the data in an image format and then recognizes the image data, which can reduce data redundancy and facilitate the extraction of ship vibration data features. Furthermore, this application improves the structure of the deep learning network model, extracting local and global features through multiple branches and then fusing them, fully extracting the features of the image data, which can improve the recognition results of ship vibration data and increase the accuracy of ship status recognition.

[0107] According to the method described in the above embodiment, this embodiment will be further described from the perspective of a time-series multimodal ship vibration identification device. The time-series multimodal ship vibration identification device can be implemented as an independent entity or integrated into an electronic device. The electronic device can be a terminal, a server, etc., wherein the terminal can include a tablet computer, a laptop computer, a personal computer (PC), a micro processing box, or other devices.

[0108] See also Figure 5 , Figure 5The present invention specifically describes a time-series multimodal ship vibration recognition device provided in an embodiment of the present invention, which is applied to an electronic device. The time-series multimodal ship vibration recognition device may include:

[0109] An acquisition module is used to acquire ship vibration data;

[0110] A data conversion module is used to convert the ship vibration data into image data according to the mapping relationship between the ship vibration data and the RGB channels of the image;

[0111] The recognition module is used to input the image data into the trained deep learning network model for recognition and obtain the recognition results of the ship vibration data;

[0112] The deep learning network model includes a local feature extraction unit, a multi-branch extraction unit, a feature fusion unit, and a classifier unit. The image data is input into the trained deep learning network model for recognition, and the recognition results of the ship vibration data are obtained, including:

[0113] The image data is input into the local feature extraction unit to obtain a first local feature map; the local feature map is input into the multi-branch extraction unit to obtain multiple feature maps; the multiple feature maps are input into the feature fusion unit to obtain a fused feature map; the fused feature map is input into the classifier unit to obtain the recognition result of the ship vibration data.

[0114] During specific implementation, the above modules and / or units can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above modules and / or units can refer to the previous method embodiments. The specific beneficial effects that can be achieved can also be found in the beneficial effects in the previous method embodiments, which will not be repeated here.

[0115] In addition, embodiments of the present application further provide an electronic device, which may be a computer, tablet computer, or other device. This electronic device can implement the steps of any of the embodiments of the time-series multimodal ship vibration identification method provided in the embodiments of the present application, and thus can achieve the beneficial effects achieved by any of the time-series multimodal ship vibration identification methods provided in the embodiments of the present application. For details, please refer to the previous embodiments and will not be repeated here.

[0116] Figure 6 The following figure shows a block diagram of the specific structure of an electronic device provided in an embodiment of the present invention. This electronic device can be used to implement the time-series multimodal ship vibration identification method provided in the above embodiments. The electronic device 500 can be a terminal, server, or other device. The terminal can include a tablet computer, laptop computer, personal computer (PC), microprocessor box, or other device.

[0117] RF circuit 510 is used to receive and transmit electromagnetic waves, converting them into electrical signals, thereby enabling communication with a communications network or other devices. RF circuit 510 may include various existing circuit components for performing these functions, such as an antenna, a radio frequency transceiver, a digital signal processor, an encryption / decryption chip, a subscriber identity module (SIM) card, memory, and the like. RF circuit 510 can communicate with various networks, such as the Internet, an intranet, or a wireless network, or with other devices via a wireless network. These wireless networks may include cellular telephone networks, wireless local area networks, or metropolitan area networks. The wireless networks may utilize various communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as Institute of Electrical and Electronics Engineers standards IEEE 802.11a, IEEE 802.11b, IEEE802.11g, and / or IEEE802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging, and short messaging, and any other suitable communication protocols, including those currently undeveloped.

[0118] The memory 520 can be used to store software programs and modules, such as the corresponding program instructions / modules in the above-mentioned embodiments. The processor 580 executes various functional applications and data processing by running the software programs and modules stored in the memory 520, that is, realizing functions such as taking pictures with the front camera, processing the captured images, and switching the display color of the displayed content on the display screen. The memory 520 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 520 may further include a memory remotely located relative to the processor 580, and these remote memories may be connected to the electronic device 500 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0119] The input unit 530 may be used to receive input digital or character information, and generate a keyboard and a mouse related to user settings and function control.

[0120] The display unit 540 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces. These graphical user interfaces can be composed of graphics, text, icons, videos, or any combination thereof. The display unit 540 may include a display panel 541. Optionally, the display panel 541 can be configured in the form of an LCD (Liquid Crystal Display), an OLED (Organic Light-Emitting Diode), or the like.

[0121] Audio circuit 560, speaker 561, and microphone 562 provide an audio interface between the user and electronic device 500. Audio circuit 560 converts received audio data into electrical signals and transmits them to speaker 561, which then converts them into sound signals for output. Microphone 562, on the other hand, converts collected sound signals into electrical signals, which are then received by audio circuit 560 and converted into audio data. The audio data is then processed by output processor 580 and transmitted via RF circuit 510 to, for example, another terminal. Alternatively, the audio data may be output to memory 520 for further processing. Audio circuit 560 may also include an earphone jack to allow communication between external headphones and electronic device 500.

[0122] Electronic device 500, through a transmission module 570 (e.g., a Wi-Fi module), can help users receive requests, send information, and so on, providing users with wireless broadband Internet access. Although the figure shows transmission module 570, it is understood that it is not a required component of electronic device 500 and can be omitted as needed without changing the essence of the invention.

[0123] Processor 580 is the control center of electronic device 500. It connects all components of the phone using various interfaces and circuits. By running or executing software programs and / or modules stored in memory 520 and accessing data stored in memory 520, it executes various functions of electronic device 500 and processes data, thereby providing overall monitoring of the electronic device. Optionally, processor 580 may include one or more processing cores. In some embodiments, processor 580 may integrate an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 580.

[0124] Electronic device 500 also includes a power supply 590 (e.g., a battery) for powering various components. In some embodiments, the power supply can be logically connected to processor 580 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. Power supply 590 can also include any components, such as one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0125] Although not shown, the electronic device 500 also includes a camera (such as a front camera and a rear camera), a Bluetooth module, etc., which will not be described in detail here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal also includes a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include instructions for performing the following operations:

[0126] Obtain ship vibration data;

[0127] According to the mapping relationship between the ship vibration data and the RGB channels of the image, the ship vibration data is converted into image data;

[0128] The image data is input into the trained deep learning network model for recognition, and the recognition results of the ship vibration data are obtained;

[0129] The deep learning network model includes a local feature extraction unit, a multi-branch extraction unit, a feature fusion unit, and a classifier unit. The image data is input into the trained deep learning network model for recognition, and the recognition results of the ship vibration data are obtained, including:

[0130] The image data is input into the local feature extraction unit to obtain a first local feature map; the local feature map is input into the multi-branch extraction unit to obtain multiple feature maps; the multiple feature maps are input into the feature fusion unit to obtain a fused feature map; the fused feature map is input into the classifier unit to obtain the recognition result of the ship vibration data.

[0131] During specific implementation, the above modules can be implemented as independent entities, or can be arbitrarily combined and implemented as the same or several entities. The specific implementation of the above modules can be found in the previous method embodiments and will not be repeated here.

[0132] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be accomplished via instructions, or by controlling related hardware via instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. To this end, an embodiment of the present invention provides a storage medium storing a plurality of instructions that can be loaded by a processor to execute the steps of any of the embodiments of the time-series multimodal ship vibration identification method provided in the embodiments of the present invention.

[0133] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0134] Since the instructions stored in the storage medium can execute the steps in any embodiment of the time-series multimodal ship vibration identification method provided in the embodiments of the present invention, the beneficial effects that can be achieved by any time-series multimodal ship vibration identification method provided in the embodiments of the present invention can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0135] The above is a detailed introduction to a time-series multimodal ship vibration identification method, device, storage medium and electronic device provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present application.

Claims

1. A method for identifying time series multimodal ship vibration, characterized in that: include: Obtain ship vibration data; converting the ship vibration data into image data according to a mapping relationship between the ship vibration data and RGB channels of the image; Inputting the image data into a trained deep learning network model for recognition to obtain a recognition result of the ship vibration data; The deep learning network model includes a local feature extraction unit, a multi-branch extraction unit, a feature fusion unit, and a classifier unit; the image data is input into the trained deep learning network model for recognition to obtain the recognition result of the ship vibration data, including: Inputting the image data into the local feature extraction unit to obtain a first local feature map; inputting the local feature map into the multi-branch extraction unit to obtain multiple feature maps; inputting the multiple feature maps into the feature fusion unit to obtain a fused feature map; inputting the fused feature map into the classifier unit to obtain a recognition result of the ship vibration data; Wherein, the multi-branch extraction unit includes a first branch, a second branch and a third branch, and the classifier unit includes a first classifier, a second classifier and a third classifier; Inputting the first local feature map into the multi-branch extraction unit to obtain multiple feature maps, including: inputting the first local feature map into the first branch to obtain a second local feature map; inputting the first local feature map into the second branch and the third branch respectively to obtain a first global feature map and a second global feature map; Inputting the fused feature map into the classifier unit to obtain a recognition result of the ship vibration data includes: inputting the fused feature map into a first classifier to obtain a first classification result; inputting the first global feature map into a second classifier to obtain a second classification result; inputting the second global feature map into a third classifier to obtain a third classification result; and obtaining a recognition result of the ship vibration data based on the first classification result, the second classification result, and the third classification result.

2. The time series multimodal ship vibration identification method according to claim 1, characterized in that: The converting of the ship vibration data into image data according to the mapping relationship between the ship vibration data and the RGB channels of the image includes: The ship vibration data includes ship vibration data in the x-direction, ship vibration data in the y-direction, and ship vibration data in the z-direction, and the ship vibration data in the x-direction is mapped to the R channel, the ship vibration data in the y-direction is mapped to the G channel, and the ship vibration data in the z-direction is mapped to the B channel to obtain a plurality of single-row color images; The image data is obtained by folding the multiple single-row color images.

3. The method for identifying time series multimodal ship vibration according to claim 1, characterized in that: The first branch includes several feature extraction layers, each of which includes multiple convolution blocks, each of which includes multiple bottleneck layers, and the bottleneck layers include convolution layers, spatial convolution layers, and residual connection blocks; In the bottleneck layer, the input feature map is sequentially passed through a 1×1 convolution layer, a 3×3 spatial convolution layer, and a 1×1 convolution layer to obtain an output feature map.

4. The method for identifying time series multimodal ship vibration according to claim 1, wherein: The second branch includes a linear projection layer and multiple Transformer blocks, each of the Transformer blocks includes a multi-head attention module, a normalization layer, and a multi-layer perceptron connected in sequence, and the multi-head attention module and the multi-layer perceptron are further connected via a residual connection; Inputting the first local feature map into the second branch to obtain a first global feature map, comprising: Inputting the first local feature map into the linear projection layer to obtain a plurality of image blocks; Pass the plurality of image blocks through the plurality of Transformer blocks to obtain the first global feature map.

5. The method for identifying time series multimodal ship vibration according to claim 1, characterized in that: The third branch includes a plurality of hybrid modules, each of which includes two different types of multilayer perceptrons; In the first type of multilayer perceptron, the input feature map is passed through a linear projection layer to obtain several image blocks. After the columns of each image block are transposed, the output feature map is obtained through a shared multilayer perceptron. In the second type of multilayer perceptron, the input feature map is passed through a linear projection layer to obtain several image blocks. After the rows of each image block are transposed, the output feature map is obtained through a shared multilayer perceptron.

6. The method for identifying time series multimodal ship vibration according to claim 1, characterized in that: The feature fusion unit includes a 1×1 convolution layer and a downsampling module; Inputting the plurality of feature maps into the feature fusion unit to obtain a fused feature map includes: Inputting the first local feature map into the 1×1 convolutional layer and the downsampling module in sequence to obtain a first local feature map with adjusted channel number and spatial dimension; The first local feature map after the number of channels and spatial dimensions are adjusted and the first global feature map and the second global feature map after upsampling are added together and then regularized to obtain the fused feature map.

7. A time-series multimodal ship vibration identification device, the time-series multimodal ship vibration identification device is used to implement the time-series multimodal ship vibration identification method according to claim 1, characterized in that: include: An acquisition module is used to acquire ship vibration data; a data conversion module, configured to convert the ship vibration data into image data according to a mapping relationship between the ship vibration data and the RGB channels of the image; A recognition module is used to input the image data into a trained deep learning network model for recognition, thereby obtaining a recognition result of the ship vibration data; The deep learning network model includes a local feature extraction unit, a multi-branch extraction unit, a feature fusion unit, and a classifier unit; the image data is input into the trained deep learning network model for recognition to obtain the recognition result of the ship vibration data, including: The image data is input into the local feature extraction unit to obtain a first local feature map; the local feature map is input into the multi-branch extraction unit to obtain multiple feature maps; the multiple feature maps are input into the feature fusion unit to obtain a fused feature map; the fused feature map is input into the classifier unit to obtain a recognition result of the ship vibration data.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, which are suitable for being loaded by a processor to execute the time-series multimodal ship vibration identification method according to any one of claims 1 to 6.

9. An electronic device, characterized in that: The invention comprises a processor and a memory, wherein the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used to execute the steps in the time series multimodal ship vibration identification method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Ship bearing fault diagnosis method based on parallel scaling Gramb angle field and CNN (Convolutional Neural Network)

    CN119779682A