A method and system for ship water gauge reading based on hyperspectral multi-domain fusion
Through the ship's water ruler reading method based on hyperspectral multi-domain fusion, the hyperspectral image feature extraction network and multi-scale feature fusion module are used to solve the accuracy and safety problems of traditional methods in complex environments, and high-precision water ruler readings are achieved.
Patent Information
- Application Number
- CN202411185032.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-27
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-08-27
AI Technical Summary
Traditional ship water ruler reading methods are susceptible to environmental factors such as water surface waves, obstacles, water traces and inclinations, resulting in reduced observation accuracy and reliability and pose a threat to the safety of observers.
The ship's water ruler reading method based on hyperspectral multi-domain fusion is adopted. By acquiring the hyperspectral water ruler images, the hyperspectral image feature extraction network, multi-scale feature fusion module, character recognition network and waterline positioning network are used to automatically extract and calculate the water ruler reading to reduce manual intervention.
It improves the accuracy and reliability of water ruler readings, reduces the threat to observer safety, and can effectively identify water ruler characters and position water lines in complex environments to achieve accurate water ruler readings.
Smart Images

Figure CN119131808B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ship water gauge reading, and in particular to a ship water gauge reading method and system based on hyperspectral multi-domain fusion. Background Art
[0002] The ship's draft gauge reading is an important part of the draft gauge weighing method. The traditional method mainly relies on manual observation to determine the draft gauge reading. However, manual observation is easily affected by complex environmental factors such as surface waves, surface obstacles, ship water tracks, and tilt. These environmental factors not only affect the accuracy and reliability of manual observation, but also pose a certain threat to the personal safety of observers. Summary of the invention
[0003] The present invention provides a method and system for reading a ship water gauge based on hyperspectral multi-domain fusion to solve the problems existing in the above-mentioned prior art. The technical solution is as follows:
[0004] On the one hand, a method for reading ship water gauge based on hyperspectral multi-domain fusion is provided, including:
[0005] S1, obtaining a hyperspectral water gauge image taken on site;
[0006] S2, inputting the hyperspectral water gauge image into the trained hyperspectral image feature extraction network, outputting feature maps of three scales, the hyperspectral image feature extraction network including three cascaded hyperspectral image feature extraction sub-networks, downsampling the hyperspectral water gauge image, inputting the first hyperspectral image feature extraction sub-network, and outputting large-scale feature F 1 , the F 1 After downsampling, the second hyperspectral image feature extraction subnetwork is input to output the mid-scale feature F 2 , the F 2 After downsampling, the third hyperspectral image feature extraction subnetwork is input to output the small-scale feature F 3 ;
[0007] S3, the feature maps F of the three scales 1 、F 2 、F 3 Input the multi-scale feature fusion module, perform feature fusion, and output fusion features;
[0008] S4, inputting the fusion features into the trained character recognition network and waterline positioning network respectively, and outputting the character recognition results and the waterline positioning results respectively;
[0009] S5. Input the character recognition result and the waterline positioning result into a water gauge reading unit, and output the water gauge reading of the hyperspectral water gauge image.
[0010] Optionally, the hyperspectral image feature extraction subnetwork is composed of a spectral information perception module, a spatial semantic information perception module, a learnable frequency encoder and a cross-domain feature fusion module connected in sequence;
[0011] The spectral information perception module uses spectral information to fully reflect the physical structure and chemical composition differences inside an object, and obtains spectral feature information of different characters and waterlines based on the differences in reflectivity of different bands of the hyperspectral image;
[0012] The spatial semantic information perception module is composed of a spatial feature decomposition unit and a spatial feature reconstruction unit. The spatial feature decomposition unit separates features with rich spatial information from features with less spatial information. The spatial feature reconstruction unit fully combines the two weighted different information features to strengthen the flow between feature information, obtain a feature map after spatial refinement, and retain detail features.
[0013] The learnable frequency encoder captures the minute differences in the spectra of different substances in the frequency domain and extracts the detailed features in the frequency domain according to the absorption characteristics of different substances in the image to different wavelengths;
[0014] The cross-domain feature fusion module automatically obtains the importance of each feature channel by learning the relationship between different spectral channels, and then assigns different weight coefficients to each channel to strengthen important features and suppress unimportant features. It also utilizes the complementary relationship between spatial domain information and frequency domain information to enhance the expression and description of image features and improve the performance of the model.
[0015] Optionally, the spectral information perception module is composed of a spectral feature segmentation unit and a spectral feature integration unit;
[0016] Among them, the spectral feature segmentation unit divides the input feature I obtained after downsampling the hyperspectral image into two parts according to the channel, namely αC and (1-α)C, 0≤α≤1 is the segmentation ratio, C is the number of channels of the feature, and then respectively increases the dimension through 1×1 convolution and k×k convolution, using the spectral information to fully reflect the characteristics of the physical structure and chemical composition differences inside the object. According to the differences in reflectivity of different bands of the hyperspectral image, the spectral feature information of different characters and waterlines is obtained to obtain the corresponding feature graphs Y 1 and Y 2 ;
[0017] The spectral feature integration unit reduces the redundancy of the refined feature map on the channel. First, a global average pooling layer is used to collect global spatial information S 1 and S 2 , and then the obtained S 1 and S2 Superimposed together, and the channel feature importance vector δ is generated through the soft attention operation 1 and δ 2 , the soft attention mechanism assigns weights to each element in the sequence, rather than just selecting one element, and weights different parts of the input data through the soft function SoftMax. The formula is as follows:
[0018]
[0019] Finally, Y 1 and Y 2 The channel feature importance vector δ corresponding to them 1 and δ 2 The multiplication results are concatenated by channel to obtain spectral feature I c .
[0020] Optionally, the spatial semantic information perception module is composed of the spatial feature decomposition unit and the spatial feature reconstruction unit;
[0021] The spatial feature decomposition unit uses the scaling factor in the group normalization layer to evaluate the information content of different feature maps and converts the spectral feature I c , perform group normalization, and obtain the group normalized feature map I out , the formula is as follows:
[0022]
[0023] Where GN is group normalization, μ and σ are I c The mean and standard deviation of ε are added to stabilize the division. γ and β are trainable parameters. The parameter γ is used as a method to measure the spatial pixel variance of each channel. A larger γ of a channel means that the spatial information contained in this channel is richer. The normalized weight W is obtained by the following formula γ :
[0024]
[0025] Where i, j = 1, 2, 3, ..., C, C represents the number of channels of the feature map, and then W γ And the feature map I after group normalization out Multiply them together to get the reweighted feature I W , then I W The information weight W is obtained by mapping it to the range (0, 1) through the Sigmoid function and gating it by setting the threshold. The weights above the threshold are set to 1 and the rest are set to 0. 1 , set the weights below the threshold to 1, and the rest to 0, and get the non-information weight W 2Finally, the spectral feature I c Respectively with the W 1 and W 2 Multiply to get features with high information content and features with low information content
[0026] The spatial feature reshaping unit uses a cross-reconstruction operation to fully combine the two weighted different information features and strengthen the flow between feature information. First, the feature Divide and feature Divide and Afterwards, and and Add together to get the reconstructed feature I 1 and I 2 Finally, I 1 and I 2 Splice to obtain the spatial feature F after spatial refinement S .
[0027] Optionally, the learnable frequency encoder is composed of layer normalization, two-dimensional discrete Fourier transform, a learnable global filter, a two-dimensional inverse discrete Fourier transform, a feedforward network and a first residual connection;
[0028] Among them, the two-dimensional discrete Fourier transform transforms the features from the spatial domain to the frequency domain, and can learn the global filter to filter out unimportant frequency domain features and further strengthen the acquired frequency domain features. The two-dimensional discrete Fourier inverse transform transforms the frequency domain features back to the spatial domain. The feedforward network captures the complex relationship between the input features and expands the input representation space. The specific implementation process is as follows:
[0029] For the spatial feature F S ∈R H×W×D , after normalizing its layers, use the two-dimensional discrete Fourier transform to transform the image features of each channel f[x,y[, 0≤x≤H-1,0≤y≤W-1, into the frequency domain, and concatenate them according to the channel dimension to obtain the frequency domain features X∈R H×W×D , is a complex tensor and represents the frequency spectrum of the input features;
[0030] Then, a learnable global filter K∈R is added H×W×D Multiply it by X to modulate the spectrum and obtain enhanced frequency domain features
[0031] Then, the enhanced frequency domain features are transformed into Transform channel by channel back to the spatial domain and update f[x,y], concatenate the transformed f[x,y] according to the channel dimension, and get
[0032] Finally Input a feedforward network, the feedforward network consists of a layer normalization, a multilayer perceptron and a second residual connection, and the output of the feedforward network, the spatial domain feature F of the first residual connection S , the second residual connection The three are added together to finally obtain the frequency domain refinement enhancement feature F P .
[0033] Optionally, the cross-domain feature fusion module obtains representative overall features and significant features in different channels of spatial domain features by using average pooling and maximum pooling, and uses them as weight coefficients to measure the importance of different channels. The weight coefficients are multiplied by the frequency domain refinement enhancement features after the convolution operation to further strengthen the important frequency domain features, while suppressing the unimportant frequency domain features, and utilizing the complementary relationship between spatial domain information and frequency domain information to strengthen the expression and description of image features and improve the performance of the model. The specific implementation process is as follows:
[0034] The spatial feature F S , after the average pooling layer and the maximum pooling layer, two one-dimensional feature vectors are obtained. The average pooling layer uses the average value in the selected area as the value after pooling of this area to retain the overall characteristics of the feature map. The maximum pooling layer uses the maximum value in the selected area as the value after pooling of this area to retain the most significant features of the feature map. The two one-dimensional feature vectors obtained after the pooling operation are reduced in dimension through a fully connected layer with shared weights, and then added together, and then a fully connected layer is used to increase the dimension.
[0035] The frequency domain refinement enhancement feature F P , after passing through two convolutional layers, multiplying the one-dimensional vector after dimensionality increase, and combining the obtained result with the spatial domain feature F S Add them together to get the fused cross-domain fusion feature F.
[0036] Optionally, the multi-scale feature fusion module uses a feature pyramid algorithm to fuse the three-scale feature maps obtained by the hyperspectral image extraction network, and upsamples the small-scale feature map by 1 times 2 to make its size consistent with the large-scale feature map. Figure 1 The middle-scale feature map is upsampled by 1 times to make its size equal to the large-scale feature map, and then the three feature maps are superimposed together to form the fused feature, which contains richer semantic information.
[0037] Optionally, the character recognition network is composed of a convolutional layer and a fully connected layer, and outputs the detection box attributes of the target, the category and confidence of the target respectively according to the fusion features;
[0038] The waterline positioning network is composed of a global pooling layer and a fully connected layer. The fused features are passed through the global pooling layer to obtain a one-dimensional tensor, and the one-dimensional tensor is passed through a fully connected layer to fit the coordinates of the final waterline.
[0039] Optionally, the water gauge reading unit determines the two characters T closest to the water surface according to the character recognition result. 1 and T 2 , and according to T 1 , T 2 and the waterline coordinates, the distance between the three is used to calculate the specific scale where the water surface is located, including:
[0040] Calculate the waterline to the nearest character T respectively 1 The distance L 1 , character T 1 to T 2 The distance L 2 , T 1 and T 2 The scale difference V, and according to T 1 The scale value R is used to determine the water gauge reading L. The specific calculation formula is as follows:
[0041]
[0042] On the other hand, a ship water gauge reading system based on hyperspectral multi-domain fusion is provided, the system comprising:
[0043] An acquisition module is used to acquire the hyperspectral water gauge image taken on site;
[0044] The feature extraction module is used to input the hyperspectral water gauge image into the trained hyperspectral image feature extraction network, and output feature maps of three scales. The hyperspectral image feature extraction network includes three cascaded hyperspectral image feature extraction sub-networks. After the hyperspectral water gauge image is downsampled, it is input into the first hyperspectral image feature extraction sub-network to output the large-scale feature F 1 , the F 1 After downsampling, the second hyperspectral image feature extraction subnetwork is input to output the mid-scale feature F 2 , the F 2 After downsampling, the third hyperspectral image feature extraction subnetwork is input to output the small-scale feature F 3 ;
[0045] The feature fusion module is used to combine the feature maps F of the three scales1 、F 2 、F 3 Input the multi-scale feature fusion module, perform feature fusion, and output fusion features;
[0046] A recognition and positioning module, used to input the fusion features into the trained character recognition network and waterline positioning network respectively, and output the character recognition results and the waterline positioning results respectively;
[0047] The water gauge reading module is used to input the character recognition result and the waterline positioning result into the water gauge reading unit, and output the water gauge reading of the hyperspectral water gauge image.
[0048] On the other hand, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned ship water level reading method based on hyperspectral multi-domain fusion.
[0049] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement the above-mentioned ship water level reading method based on hyperspectral multi-domain fusion.
[0050] The beneficial effects brought about by the technical solution provided by the present invention include at least:
[0051] The present invention proposes a method and system for reading ship water gauges based on hyperspectral multi-domain fusion. The method uses hyperspectral images to replace traditional RGB images to make up for the missing channel information of RGB images, and designs a spectral information perception module, which can use the differences in reflectivity of different bands of hyperspectral images to solve the problems of false detection, missed detection, and wrong detection of water gauge characters in scenes with water surface reflections, water rust, and insufficient lighting conditions; a spatial semantic information perception module is designed to enable the model to learn sufficient detail features and enhance the model's ability to recognize target characters with inconsistent sizes, shapes, and shooting angles on water gauge images; a learnable frequency encoder is used to obtain frequency domain refinement enhancement feature information of hyperspectral images, and a cross-domain feature fusion module is designed to obtain hyperspectral water gauge image features that simultaneously fuse spatial and frequency domain information, making up for the limitations of using only spatial or frequency domain features; a character recognition network and a waterline positioning network are built, and the water gauge reading is determined through a water gauge reading unit to meet the task requirements of accurately reading water gauges using hyperspectral images. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0053] Figure 1 It is a flow chart of a method for reading a ship water gauge based on hyperspectral multi-domain fusion provided by an embodiment of the present invention;
[0054] Figure 2 This is a general flow chart of a method for reading a ship water gauge using hyperspectral multi-domain fusion provided by an embodiment of the present invention;
[0055] Figure 3 It is a schematic diagram of the structure of a spectral information perception module provided by an embodiment of the present invention;
[0056] Figure 4 is a schematic diagram of the structure of a spatial semantic information perception module provided by an embodiment of the present invention;
[0057] Figure 5 is a schematic diagram of the structure of a learnable frequency encoder provided by an embodiment of the present invention;
[0058] Figure 6 It is a schematic diagram of the structure of a cross-domain feature fusion module provided by an embodiment of the present invention;
[0059] Figure 7 It is a schematic diagram of a water gauge reading unit provided in an embodiment of the present invention reading a water gauge reading;
[0060] Figure 8 This is a block diagram of a ship water gauge reading system based on hyperspectral multi-domain fusion provided by an embodiment of the present invention;
[0061] Fig. 9 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0062] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0063] The embodiment of the present invention provides a method for reading a ship water gauge based on hyperspectral multi-domain fusion, which can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The flowchart of a method for reading a ship water gauge based on hyperspectral multi-domain fusion is shown in FIG. The processing flow of the method may include the following steps:
[0064] S1, obtaining a hyperspectral water gauge image taken on site;
[0065] The embodiment of the present invention uses a hyperspectral camera to obtain a hyperspectral water gauge image taken on-site at a port or dock.
[0066] S2, input the hyperspectral water gauge image into the trained hyperspectral image feature extraction network, and output feature maps of three scales. The hyperspectral image feature extraction network includes three cascaded hyperspectral image feature extraction sub-networks. After downsampling the hyperspectral water gauge image (using convolution with a step size of 2), the hyperspectral water gauge image is input into the first hyperspectral image feature extraction sub-network to output large-scale features F 1 , the F 1 After downsampling, the second hyperspectral image feature extraction subnetwork is input to output the mid-scale feature F 2 , the F 2 After downsampling, the third hyperspectral image feature extraction subnetwork is input to output the small-scale feature F 3 ;
[0067] The training of the hyperspectral image feature extraction network in the embodiment of the present invention is performed using the image data in the training set, including:
[0068] A hyperspectral camera is used to obtain water gauge photos of multiple ships at different positions in the port or dock. The water gauge characters appearing in the image are annotated using target detection and annotation software to obtain a water gauge scale character dataset. The waterline position in the image is annotated using annotation software to obtain the waterline coordinates and obtain a waterline positioning dataset. The obtained water gauge scale character dataset and waterline positioning dataset are divided into a training set and a test set, respectively.
[0069] Alternatively, if Figure 2 As shown, the hyperspectral image feature extraction subnetwork is composed of a spectral information perception module, a spatial semantic information perception module, a learnable frequency encoder and a cross-domain feature fusion module connected in sequence;
[0070] The spectral information perception module uses spectral information to fully reflect the physical structure and chemical composition differences inside an object, and obtains spectral feature information of different characters and waterlines based on the differences in reflectivity of different bands of the hyperspectral image;
[0071] The spectral information perception module can solve the problems of false detection, missed detection, and wrong detection of water ruler characters in scenes with water surface reflections, water rust, and insufficient lighting conditions.
[0072] The spatial semantic information perception module is composed of a spatial feature decomposition unit and a spatial feature reconstruction unit. The spatial feature decomposition unit separates features with rich spatial information from features with less spatial information. The spatial feature reconstruction unit fully combines the two weighted different information features to strengthen the flow between feature information, obtain a feature map after spatial refinement, and retain detail features.
[0073] The spatial semantic information perception module can solve the problem of poor recognition effect of target characters with inconsistent sizes, shapes and shooting angles on the water gauge image.
[0074] The features corresponding to different materials in the hyperspectral water level image have small differences in real numbers, which limits the feature extraction ability of the model. The learnable frequency encoder captures the slight differences in the spectra of different materials in the frequency domain and extracts the detailed features in the frequency domain according to the absorption characteristics of different materials (hull, characters, water surface, etc.) in the image at different wavelengths.
[0075] Since the spectral information of hyperspectral images is richer, in order to make full use of the characteristics of the spectral dimension and realize further enhancement and fusion of features, the cross-domain feature fusion module automatically obtains the importance of each feature channel by learning the relationship between different spectral channels, and then assigns different weight coefficients to each channel to enhance important features and suppress unimportant features. It also utilizes the complementary relationship between spatial domain information and frequency domain information to enhance the expression and description of image features and improve the performance of the model.
[0076] Alternatively, if Figure 3 As shown, the spectral information perception module is composed of a spectral feature segmentation unit and a spectral feature integration unit;
[0077] Among them, the spectral feature segmentation unit divides the input feature I obtained after downsampling the hyperspectral image into two parts according to the channel, namely αC and (1-α)C, 0≤α≤1 is the segmentation ratio, C is the number of channels of the feature, and then respectively increases the dimension through 1×1 convolution and k×k convolution, using the spectral information to fully reflect the characteristics of the physical structure and chemical composition differences inside the object. According to the differences in reflectivity of different bands of the hyperspectral image, the spectral feature information of different characters and waterlines is obtained to obtain the corresponding feature graphs Y 1 and Y 2 ;
[0078] The spectral feature integration unit reduces the redundancy of the refined feature map on the channel. First, a global average pooling layer is used to collect global spatial information S 1 and S 2 , and then the obtained S 1 and S 2Superimposed together, and the channel feature importance vector δ is generated through the soft attention operation 1 and δ 2 , the soft attention mechanism assigns weights to each element in the sequence, rather than just selecting one element, and weights different parts of the input data through the soft function SoftMax. The formula is as follows:
[0079]
[0080] Finally, Y 1 and Y 2 The channel feature importance vector δ corresponding to them 1 and δ 2 The multiplication results are concatenated by channel to obtain spectral feature I c .
[0081] Alternatively, if Figure 4 As shown, the spatial semantic information perception module is composed of the spatial feature decomposition unit and the spatial feature reconstruction unit;
[0082] The spatial feature decomposition unit uses the scaling factor in the group normalization (GN) layer to evaluate the information content of different feature maps and converts the spectral feature I c , perform group normalization, and obtain the group normalized feature map I out , the formula is as follows:
[0083]
[0084] Where GN is group normalization, μ and σ are I c The mean and standard deviation of ε are added to stabilize the division. γ and β are trainable parameters. The parameter γ is used as a method to measure the spatial pixel variance of each channel. A larger γ of a channel means that the spatial information contained in this channel is richer. The normalized weight W is obtained by the following formula γ :
[0085]
[0086] Where i, j = 1, 2, 3, ..., C, C represents the number of channels of the feature map, and then W γ And the feature map I after group normalization out Multiply them together to get the reweighted feature I W , then I W The information weight W is obtained by mapping it to the range (0, 1) through the Sigmoid function and gating it by setting the threshold. The weights above the threshold are set to 1 and the rest are set to 0. 1, set the weights below the threshold to 1, and the rest to 0, and get the non-information weight W 2 Finally, the spectral feature I c Respectively with the W 1 and W 2 Multiply to get features with high information content and features with low information content
[0087] The spatial feature reshaping unit uses a cross-reconstruction operation to fully combine the two weighted different information features and strengthen the flow between feature information. First, the feature Divide and feature Divide and Afterwards, and and Add together to get the reconstructed feature I 1 and I 2 Finally, I 1 and I 2 Splice to obtain the spatial feature F after spatial refinement S .
[0088] Alternatively, if Figure 5 As shown, the learnable frequency encoder is composed of layer normalization, two-dimensional discrete Fourier transform, a learnable global filter, a two-dimensional inverse discrete Fourier transform, a feed forward network (FFN) and a first residual connection;
[0089] Among them, the two-dimensional discrete Fourier transform transforms the features from the spatial domain to the frequency domain, and can learn the global filter to filter out unimportant frequency domain features and further strengthen the acquired frequency domain features. The two-dimensional discrete Fourier inverse transform transforms the frequency domain features back to the spatial domain. The feedforward network captures the complex relationship between the input features and expands the input representation space. The specific implementation process is as follows:
[0090] For the spatial feature F S ∈R H×W×D , after normalizing its layers, use the two-dimensional discrete Fourier transform to transform the image features f[x,y] of each channel, 0≤x≤H-1,0≤y≤W-1, to the frequency domain, and concatenate them according to the channel dimension to obtain the frequency domain features X∈R H×W×D , is a complex tensor and represents the frequency spectrum of the input features;
[0091] Then, a learnable global filter K∈R is added H×W×DMultiply it by X to modulate the spectrum and obtain enhanced frequency domain features
[0092] Then, the enhanced frequency domain features are transformed into Transform channel by channel back to the spatial domain and update f[x,y], concatenate the transformed f[x,y] according to the channel dimension, and get
[0093] Finally Input a feedforward network, the feedforward network consists of a layer normalization, a multi-layer perceptron (MLP) and a second residual connection, and the output of the feedforward network, the spatial domain feature F of the first residual connection S , the second residual connection The three are added together to finally obtain the frequency domain refinement enhancement feature F P .
[0094] Alternatively, if Figure 6 As shown in the figure, the cross-domain feature fusion module obtains representative overall features and significant features in different channels of spatial domain features by using average pooling and maximum pooling, and uses them as weight coefficients to measure the importance of different channels. The weight coefficients are multiplied by the frequency domain refinement enhancement features after the convolution operation to further strengthen the important frequency domain features, while suppressing the unimportant frequency domain features. The complementary relationship between spatial domain information and frequency domain information is used to strengthen the expression and description of image features and improve the performance of the model. The specific implementation process is as follows:
[0095] The spatial feature F S , after the average pooling layer and the maximum pooling layer, two one-dimensional feature vectors are obtained. The average pooling layer uses the average value in the selected area as the value after pooling of this area to retain the overall characteristics of the feature map. The maximum pooling layer uses the maximum value in the selected area as the value after pooling of this area to retain the most significant features of the feature map. The two one-dimensional feature vectors obtained after the pooling operation are reduced in dimension through a fully connected layer with shared weights, and then added together, and then a fully connected layer is used to increase the dimension.
[0096] The frequency domain refinement enhancement feature F P , after passing through two convolutional layers, multiplying the one-dimensional vector after dimensionality increase, and combining the obtained result with the spatial domain feature F S Add them together to get the fused cross-domain fusion feature F.
[0097] S3, the feature maps F of the three scales 1 、F 2 、F 3 Input the multi-scale feature fusion module, perform feature fusion, and output fusion features;
[0098] Optionally, the multi-scale feature fusion module uses a feature pyramid algorithm to fuse the three-scale feature maps obtained by the hyperspectral image extraction network, and upsamples the small-scale feature map by 1 times 2 to make its size consistent with the large-scale feature map. Figure 1 The middle-scale feature map is upsampled by 1 times to make its size equal to the large-scale feature map, and then the three feature maps are superimposed together to form the fused feature, which contains richer semantic information.
[0099] S4, inputting the fusion features into the trained character recognition network and waterline positioning network respectively, and outputting the character recognition results and the waterline positioning results respectively;
[0100] Optionally, the character recognition network is composed of a convolutional layer and a fully connected layer, and outputs the detection box attributes (x-coordinate, y-coordinate, box width, box height) of the target, the category and confidence of the target according to the fusion features;
[0101] The waterline positioning network is composed of a global pooling layer and a fully connected layer. The fused features are passed through the global pooling layer to obtain a one-dimensional tensor, and the one-dimensional tensor is passed through a fully connected layer to fit the coordinates of the final waterline.
[0102] S5. Input the character recognition result and the waterline positioning result into a water gauge reading unit, and output the water gauge reading of the hyperspectral water gauge image.
[0103] Alternatively, if Figure 7 As shown, the water gauge reading unit determines the two characters T closest to the water surface according to the character recognition result. 1 and T 2 , and according to T 1 , T 2 and the waterline coordinates, the distance between the three is used to calculate the specific scale where the water surface is located, including:
[0104] Calculate the waterline to the nearest character T respectively 1 The distance L 1 , character T 1 to T 2 The distance L 2 , T 1 and T 2 The scale difference V, and according to T 1 The scale value R is used to determine the water gauge reading L. The specific calculation formula is as follows:
[0105]
[0106] by Figure 7 For example, the character T 1 and T2 The scale difference V is 2, T 1 The scale value R is 4, then the water gauge reading is calculated as follows:
[0107] like Figure 8 As shown, an embodiment of the present invention further provides a ship water gauge reading system based on hyperspectral multi-domain fusion, the system comprising:
[0108] An acquisition module 810 is used to acquire a hyperspectral water gauge image taken on site;
[0109] The feature extraction module 820 is used to input the hyperspectral water gauge image into the trained hyperspectral image feature extraction network, and output feature maps of three scales. The hyperspectral image feature extraction network includes three cascaded hyperspectral image feature extraction sub-networks. After the hyperspectral water gauge image is downsampled, it is input into the first hyperspectral image feature extraction sub-network to output large-scale features F 1 , the F 1 After downsampling, the second hyperspectral image feature extraction subnetwork is input to output the mid-scale feature F 2 , the F 2 After downsampling, the third hyperspectral image feature extraction subnetwork is input to output the small-scale feature F 3 ;
[0110] The feature fusion module 830 is used to combine the feature maps F of the three scales 1 、F 2 、F 3 Input the multi-scale feature fusion module, perform feature fusion, and output fusion features;
[0111] The recognition and positioning module 840 is used to input the fusion features into the trained character recognition network and waterline positioning network, and output the character recognition results and waterline positioning results respectively;
[0112] The water gauge reading module 850 is used to input the character recognition result and the waterline positioning result into a water gauge reading unit, and output the water gauge reading of the hyperspectral water gauge image.
[0113] An embodiment of the present invention provides a ship water gauge reading system based on hyperspectral multi-domain fusion, and its functional structure corresponds to a ship water gauge reading method based on hyperspectral multi-domain fusion provided by an embodiment of the present invention, which will not be repeated here.
[0114] Fig. 9It is a structural schematic diagram of an electronic device 900 provided in an embodiment of the present invention. The electronic device 900 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 901 and one or more memories 902, wherein at least one instruction is stored in the memory 902, and the at least one instruction is loaded and executed by the processor 901 to implement the steps of the above-mentioned ship water level reading method based on hyperspectral multi-domain fusion.
[0115] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions, which can be executed by a processor in a terminal to complete the above-mentioned ship water gauge reading method based on hyperspectral multi-domain fusion. For example, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0116] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0117] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for reading ship water gauge based on hyperspectral multi-domain fusion, characterized in that: The method comprises: S1, obtaining a hyperspectral water gauge image taken on site; S2, inputting the hyperspectral water gauge image into the trained hyperspectral image feature extraction network, outputting feature maps of three scales, the hyperspectral image feature extraction network including three cascaded hyperspectral image feature extraction sub-networks, downsampling the hyperspectral water gauge image, inputting the first hyperspectral image feature extraction sub-network, and outputting large-scale feature F 1 , the F 1 After downsampling, the second hyperspectral image feature extraction subnetwork is input to output the mid-scale feature F 2 , the F 2 After downsampling, the third hyperspectral image feature extraction subnetwork is input to output the small-scale feature F 3 ; S3, the feature maps F of the three scales 1 、F 2 、F 3 Input the multi-scale feature fusion module, perform feature fusion, and output fusion features; S4, inputting the fusion features into the trained character recognition network and waterline positioning network respectively, and outputting the character recognition results and the waterline positioning results respectively; S5, inputting the character recognition result and the waterline positioning result into a water gauge reading unit, and outputting the water gauge reading of the hyperspectral water gauge image; The hyperspectral image feature extraction subnetwork is composed of a spectral information perception module, a spatial semantic information perception module, a learnable frequency encoder and a cross-domain feature fusion module connected in sequence; The spectral information perception module uses spectral information to fully reflect the physical structure and chemical composition differences inside an object, and obtains spectral feature information of different characters and waterlines based on the differences in reflectivity of different bands of the hyperspectral image; The spatial semantic information perception module is composed of a spatial feature decomposition unit and a spatial feature reconstruction unit. The spatial feature decomposition unit separates features with different spatial information, and the spatial feature reconstruction unit fully combines two weighted different information features to strengthen the flow between feature information, thereby obtaining a feature map after spatial refinement and retaining detail features. The learnable frequency encoder captures the minute differences in the spectra of different substances in the frequency domain and extracts the detailed features in the frequency domain according to the absorption characteristics of different substances in the image to different wavelengths; The cross-domain feature fusion module automatically obtains the importance of each feature channel by learning the relationship between different spectral channels, and then assigns different weight coefficients to each channel to strengthen important features and suppress unimportant features. It also utilizes the complementary relationship between spatial domain information and frequency domain information to enhance the expression and description of image features and improve the performance of the model.
2. The method according to claim 1, characterized in that The spectral information perception module is composed of a spectral feature segmentation unit and a spectral feature integration unit; The spectral feature segmentation unit divides the input feature I obtained after downsampling the hyperspectral image into two parts according to the channel, namely αC and (1-α)C, 0≤α≤1 is the segmentation ratio, C is the number of channels of the feature, and then respectively increases the dimension through 1×1 convolution and k×k convolution, using the spectral information to fully reflect the physical structure and chemical composition differences inside the object. According to the difference in reflectance of different bands of the hyperspectral image, the spectral feature information of different characters and waterlines is obtained to obtain the corresponding feature maps Y1 and Y2; The spectral feature integration unit reduces the redundancy of the refined feature map on the channel. First, the global average pooling layer is used to collect the global spatial information S1 and S2, and then the obtained S1 and S2 are superimposed together, and the channel feature importance vectors δ1 and δ2 are generated through the soft attention operation. The soft attention mechanism assigns a weight to each element in the sequence instead of just selecting one element. The soft function SoftMax is used to weight different parts of the input data. The formula is as follows: Finally, the results of multiplying Y1 and Y2 by their corresponding channel feature importance vectors δ1 and δ2 are concatenated by channel to obtain the spectral feature I c .
3. The method according to claim 2, characterized in that The spatial semantic information perception module is composed of the spatial feature decomposition unit and the spatial feature reconstruction unit; The spatial feature decomposition unit uses the scaling factor in the group normalization layer to evaluate the information content of different feature maps and converts the spectral feature I c , perform group normalization, and obtain the group normalized feature map I out , the formula is as follows: Where GN is group normalization, μ and σ are I c The mean and standard deviation of ε are added to stabilize the division. γ and β are trainable parameters. The parameter γ is used as a method to measure the spatial pixel variance of each channel. A larger γ of a channel means that the spatial information contained in this channel is richer. The normalized weight W is obtained by the following formula γ : Where i, j = 1, 2, 3, ..., C, C represents the number of channels of the feature map, and then W γ And the feature map I after group normalization out Multiply them together to get the reweighted feature I W , then I W The sigmoid function is used to map the data to the range (0, 1), and the threshold is set for gating. The weights above the threshold are set to 1, and the rest are set to 0 to obtain the information weight W1. The weights below the threshold are set to 1, and the rest are set to 0 to obtain the non-information weight W2. Finally, the spectral feature I c Multiply them by W1 and W2 respectively to get features with high information content and features with low information content The spatial feature reshaping unit uses a cross-reconstruction operation to fully combine the two weighted different information features and strengthen the flow between feature information. First, the feature Divide and feature Divide and Afterwards, and and Add them together to get the reconstructed features I1 and I2. Finally, I1 and I2 are concatenated to get the spatial domain feature F after spatial refinement. S .
4. The method according to claim 3, characterized in that The learnable frequency encoder is composed of layer normalization, two-dimensional discrete Fourier transform, a learnable global filter, a two-dimensional inverse discrete Fourier transform, a feedforward network and a first residual connection; Among them, the two-dimensional discrete Fourier transform transforms the features from the spatial domain to the frequency domain, and can learn the global filter to filter out unimportant frequency domain features and further strengthen the acquired frequency domain features. The two-dimensional discrete Fourier inverse transform transforms the frequency domain features back to the spatial domain. The feedforward network captures the complex relationship between the input features and expands the input representation space. The specific implementation process is as follows: For the spatial feature F S ∈R H×W×D , after normalizing its layers, use the two-dimensional discrete Fourier transform to transform the image features f[x,y] of each channel, 0≤x≤H-1,0≤y≤W-1, to the frequency domain, and concatenate them according to the channel dimension to obtain the frequency domain features X∈R H×W×D , is a complex tensor and represents the frequency spectrum of the input features; Then, a learnable global filter K∈R is added H×W×D Multiply it by X to modulate the spectrum and obtain enhanced frequency domain features Then, the enhanced frequency domain features are transformed into Transform channel by channel back to the spatial domain and update f[x,y], concatenate the transformed f[x,y] according to the channel dimension, and get Finally Input a feedforward network, the feedforward network consists of a layer normalization, a multilayer perceptron and a second residual connection, and the output of the feedforward network, the spatial domain feature F of the first residual connection S , the second residual connection The three are added together to finally obtain the frequency domain refinement enhancement feature F P .
5. The method according to claim 4, characterized in that The cross-domain feature fusion module obtains representative overall features and significant features in different channels of spatial domain features by using average pooling and maximum pooling, and uses them as weight coefficients to measure the importance of different channels. The weight coefficients are multiplied by the frequency domain refinement enhancement features after the convolution operation to further strengthen the important frequency domain features, while suppressing the unimportant frequency domain features. The complementary relationship between spatial domain information and frequency domain information is used to strengthen the expression and description of image features and improve the performance of the model. The specific implementation process is as follows: The spatial feature F S , after the average pooling layer and the maximum pooling layer, two one-dimensional feature vectors are obtained. The average pooling layer uses the average value in the selected area as the value after pooling of this area to retain the overall characteristics of the feature map. The maximum pooling layer uses the maximum value in the selected area as the value after pooling of this area to retain the most significant features of the feature map. The two one-dimensional feature vectors obtained after the pooling operation are reduced in dimension through a fully connected layer with shared weights, and then added together, and then a fully connected layer is used to increase the dimension. The frequency domain refinement enhancement feature F P , after passing through two convolutional layers, multiplying the one-dimensional vector after dimensionality increase, and combining the obtained result with the spatial domain feature F S Add them together to get the fused cross-domain fusion feature F.
6. The method according to claim 1, characterized in that The multi-scale feature fusion module uses a feature pyramid algorithm to fuse the feature maps of three scales obtained by the hyperspectral image extraction network, upsamples the small-scale feature map by 1 times 2 to make its size consistent with the large-scale feature map, upsamples the medium-scale feature map by 1 times 1 time to also make its size equal to the large-scale feature map, and then superimposes the three feature maps to form the fused feature, which contains richer semantic information.
7. The method according to claim 1, characterized in that The character recognition network is composed of a convolutional layer and a fully connected layer, and outputs the detection box attributes, the category and the confidence of the target according to the fusion features; The waterline positioning network is composed of a global pooling layer and a fully connected layer. The fused features are passed through the global pooling layer to obtain a one-dimensional tensor, and the one-dimensional tensor is passed through a fully connected layer to fit the coordinates of the final waterline.
8. The method according to claim 7, characterized in that The water gauge reading unit determines the two characters T1 and T2 closest to the water surface according to the character recognition result, and calculates the specific scale of the water surface according to the distance between T1, T2 and the waterline coordinates, including: Calculate the distance L1 from the waterline to the nearest character T1, the distance L2 from character T1 to T2, and the scale difference V between T1 and T2 respectively, and determine the water gauge reading L according to the scale value R of T1. The specific calculation formula is as follows:
9. A ship water gauge reading system based on hyperspectral multi-domain fusion, characterized in that: The system comprises: An acquisition module is used to acquire the hyperspectral water gauge image taken on site; The feature extraction module is used to input the hyperspectral water gauge image into the trained hyperspectral image feature extraction network, and output feature maps of three scales. The hyperspectral image feature extraction network includes three cascaded hyperspectral image feature extraction sub-networks. After the hyperspectral water gauge image is downsampled, it is input into the first hyperspectral image feature extraction sub-network to output the large-scale feature F 1 , the F 1 After downsampling, the second hyperspectral image feature extraction subnetwork is input to output the mid-scale feature F 2 , the F 2 After downsampling, the third hyperspectral image feature extraction subnetwork is input to output the small-scale feature F 3 ; The feature fusion module is used to combine the feature maps F of the three scales 1 、F 2 、F 3 Input the multi-scale feature fusion module, perform feature fusion, and output fusion features; A recognition and positioning module, used to input the fusion features into the trained character recognition network and waterline positioning network respectively, and output the character recognition results and the waterline positioning results respectively; A water gauge reading module, used for inputting the character recognition result and the waterline positioning result into a water gauge reading unit, and outputting the water gauge reading of the hyperspectral water gauge image; The hyperspectral image feature extraction subnetwork is composed of a spectral information perception module, a spatial semantic information perception module, a learnable frequency encoder and a cross-domain feature fusion module connected in sequence; The spectral information perception module uses spectral information to fully reflect the physical structure and chemical composition differences inside an object, and obtains spectral feature information of different characters and waterlines based on the differences in reflectivity of different bands of the hyperspectral image; The spatial semantic information perception module is composed of a spatial feature decomposition unit and a spatial feature reconstruction unit. The spatial feature decomposition unit separates features with different spatial information, and the spatial feature reconstruction unit fully combines two weighted different information features to strengthen the flow between feature information, thereby obtaining a feature map after spatial refinement and retaining detail features. The learnable frequency encoder captures the minute differences in the spectra of different substances in the frequency domain and extracts the detailed features in the frequency domain according to the absorption characteristics of different substances in the image to different wavelengths; The cross-domain feature fusion module automatically obtains the importance of each feature channel by learning the relationship between different spectral channels, and then assigns different weight coefficients to each channel to strengthen important features and suppress unimportant features. It also utilizes the complementary relationship between spatial domain information and frequency domain information to enhance the expression and description of image features and improve the performance of the model.