Active Recognition Method for Ship Names in Inland River Maritime Video Surveillance

Through the active recognition method of ship names based on computer vision, semantic part detection, OCR text detection and text filtering algorithms, the problem of time-consuming and labor-consuming ship identity confirmation in the existing technology is solved, and the rapid and accurate ship identity recognition is achieved, and maritime supervision efficiency is improved.

CN114445776BActive Publication Date: 2025-05-30WUHAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210054488.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-18
Publication Date
2025-05-30
Estimated Expiration
2042-01-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and confirm the identity of the ship, especially when illegal ships close or tamper with AIS equipment, which makes the maritime regulatory authorities time-consuming and labor-intensive investigations.

Method used

The ship's identity is identified and confirmed by using computer vision-based ship names by detecting ships in videos, using semantic partial detection, OCR text detection and text filtering algorithms based on editing distance.

Benefits of technology

The ship's identity can be identified as soon as it is discovered, effectively filtering non-ship name text interference, improving identification accuracy, and improving the reconnaissance efficiency of the maritime department on illegal ship behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114445776B_ABST
    Figure CN114445776B_ABST
Patent Text Reader

Abstract

The present invention relates to an active ship name recognition method for inland river maritime video surveillance, comprising the following steps: S1, detecting ships in the video and determining the ship types: when the ships are container ships, bulk carriers and oil tankers, semantic parts are used for detection, and while detecting the ships, the parts of the ships where ship names exist are detected; when the ships are passenger ships and law enforcement ships, semantic part detection is not performed; S2, using an OCR text detection and recognition framework to detect the ships or ship parts described in step S1 and recognize the existing text information; S3, using a designed text filtering algorithm based on the edit distance rule to determine the true ship name. The present invention can effectively improve the reconnaissance efficiency of law enforcement departments for the acts of turning off AIS or using false license plates of illegal ships, improve the work efficiency of maritime departments, and reduce labor and time costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of inland waterway video surveillance, and more specifically, to an active ship name recognition method for inland maritime video surveillance. Background Art

[0002] Currently, the main method for the maritime department to confirm the identity of ships is to use the Automatic Identification System (AIS) for ships. The on-board AIS device periodically transmits radio signals containing ship information, and the shore receiver remains in a waiting state to receive the signals. When the ship-shore devices are within the effective communication range, the shore device will receive the signals. Due to this passive communication method, when an illegal ship turns off the on-board AIS device or tampers with and uses the information of other ships, the shore device responsible for signal reception does not have the ability to identify such behaviors, resulting in very time-consuming and laborious investigations by the maritime supervision department in similar cases.

[0003] Most of the existing ship name recognition methods based on computer vision are still in the theoretical research stage, and for many problems existing on-site, such as the ship name being too small in the field of view, insufficient visual information of the ship name caused by rust, occlusion, etc., and interference from other non-ship name texts in the field of view, the existing methods cannot solve them. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide an active ship name recognition method for inland maritime video surveillance, which is an active ship name recognition method based on computer vision, makes up for the deficiencies of AIS, overcomes the interference of non-ship name texts in the field of view, and can identify the identity of a ship at the first time when the ship is discovered.

[0005] The technical solution adopted by the present invention to solve its technical problems is: to construct an active ship name recognition method for inland maritime video surveillance, including the following steps:

[0006] S1. Detect ships in the video and determine the ship type: When the ship is a container ship, a bulk carrier, or an oil tanker, use semantic parts for detection. While detecting the ship, detect the part of the ship where the ship name exists. When the ship is a passenger ship or a law enforcement ship, semantic part detection is not performed;

[0007] S2. Use the OCR text detection and recognition framework to detect the ship or the ship part described in step S1 and identify the existing text information;

[0008] S3. Use the designed text filtering algorithm based on the edit distance rule to determine the real ship name.

[0009] According to the above solution, step S1 includes the following steps:

[0010] S101. Input the ship instances and hull part instances generated by the ship region proposal network and the hull part region proposal network respectively. The instances include the bounding box coordinates bbox, that is, Bounding box, and the corresponding feature sequences. The bounding box coordinates bbox include the upper left corner point coordinates and the width and height, which are x, y, w, and h respectively.

[0011] S102. Calculate the correlation between the ship instances and the hull part instances. After calculating the correlation coefficients between the instances to obtain the correlation matrix, sum along the coordinate axes and multiply it to the corresponding instance feature sequences to complete the calculation of the correlation between the ship instances and the hull part instances.

[0012] S103. Send the instances in the region weighted by the correlation coefficients into the final classification and coordinate regression network to obtain the coordinate frames and classification results of the ship and the hull part.

[0013] According to the above solution, the method of inputting the ship instances and hull part instances generated by the ship region proposal network and the hull part region proposal network respectively in step S101 is as follows: The images of the ship instances and hull part instances are input and pass through the convolutional feature extraction network to obtain the feature map. The feature map is sent into the Region Proposal Network, that is, the RPN network to obtain the detection results. The duplicate results are filtered through Non Maximum Suppression, that is, NMS. According to the coordinate information of the results, the feature segments at the corresponding positions on the feature map are intercepted. After passing through the Region of Interest, that is, the RoI Pooling layer, the feature segments are processed into new feature maps of the same size.

[0014] According to the above solution, the method of obtaining the correlation matrix in step S102 is as follows: The images of the ship instances and hull part instances are input and pass through the networks for ship detection, classification, and hull part detection, and then the results of ship and hull detection are output, including the target location coordinates: x, y, w, h, and the feature maps of the unified size after RoI Pooling. Calculate the IoU between the ship instance coordinate frame and the hull part instance coordinate frame in turn, and finally obtain the correlation matrix of the ship and the hull part. The method of completing the calculation of the correlation between the ship instances and the hull part instances is as follows: Sum the correlation matrix along the rows or columns, and then multiply it by the feature map of the ship or the hull part to complete the weighting of the correlation coefficients.

[0015] According to the above solution, the method of using the OCR text detection and recognition framework in step S2 is as follows: The original image size of the ship instance and the hull part instance is (W, H), and the image size input into the ship detection, classification, and hull part detection network after scaling is (Wr, Hr). If the bbox in the output result is (x, y, w, h), then the position coordinates bboxo(x o 、y o 、w o 、h o ) in the original image are calculated by equations (11) and (12) as follows:

[0016]

[0017]

[0018] The above calculation method restores the position bboxo of the target in the original image, and the corresponding part is intercepted in the original image according to the bboxo coordinates as the original output of the second step; text detection is performed on the input picture, and the corresponding area is intercepted on the input image according to the text detection result for text recognition.

[0019] According to the above solution, the text filtering algorithm based on the edit distance rule in step S3 is as follows:

[0020] a. When there are multiple recognition results, calculate the edit distance of each recognition result in turn: Calculate the edit distance score of each recognition result according to the set edit distance threshold. When the edit distance is lower than the threshold for similar results, it is the ship name with recognition errors. When the edit distance is higher than the threshold for similar results, it is the one closest to the true ship name;

[0021] b. When the edit distance of two results is 0, that is, the two recognition results are exactly the same, then the ship name is determined.

[0022] Implementing the active ship name recognition method for inland river maritime video surveillance of the present invention has the following beneficial effects:

[0023] The present invention effectively filters non-ship name texts, such as container trademarks, warning slogans, etc. by adding semantic part detection technology, and at the same time ensures the clarity of the OCR input, indirectly ensuring the accuracy of the recognition result; as a supplementary method to the AIS system, the present invention can confirm the identity of the ship at the first time when the ship is discovered, effectively improving the reconnaissance efficiency of law enforcement departments for the behaviors of illegal ships turning off the AIS or using fake license plates, improving the work efficiency of the maritime department, and reducing labor and time costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:

[0025] Figure 1 It is the flow chart of the method for actively identifying ship names in inland river maritime video surveillance according to the present invention;

[0026] Figure 2 It is the network structure diagram of ship detection, classification and hull part detection according to the present invention;

[0027] Figure 3 It is the detection network structure based on region proposals according to the present invention;

[0028] Figure 4 It is the RPN structure in the region proposal neural network according to the present invention;

[0029] Figure 5 It is the schematic diagram of the principle of RoIPooling according to the present invention;

[0030] Figure 6 It is the schematic diagram of the detection target coordinate regression and classification network structure according to the present invention;

[0031] Figure 7 It is the schematic diagram of the ship name potential area text detection network structure according to the present invention;

[0032] Figure 8 It is the schematic diagram of the principle of text area contraction according to the present invention;

[0033] Figure 9 It is the schematic diagram of the text recognition network structure according to the present invention;

[0034] Figure 10 It is the schematic diagram of the text filtering algorithm based on edit distance according to the present invention. Detailed implementation manners

[0035] For a clearer understanding of the technical features, objectives and effects of the present invention, the detailed implementation manners of the present invention will now be described in detail with reference to the accompanying drawings.

[0036] As Figure 1 shown, the method for actively identifying ship names in inland river maritime video surveillance according to the present invention includes the following steps:

[0037] S1. Detect ships in the video and determine the ship types: When the ships are container ships, bulk carriers and oil tankers, semantic parts are used for detection. While detecting the ships, the parts of the ships where ship names exist are detected. When the ships are passenger ships and law enforcement ships, semantic part detection is not performed;

[0038] The specific content is as follows: S101. Input the ship instances and hull part instances generated by the ship area proposal network and the hull part area proposal network respectively. The instances include the bounding box coordinates bbox, that is, Bounding box, and the corresponding feature sequences. The bbox includes the coordinates of the upper left corner point and the width and height, that is, x, y, w, h;

[0039] The specific content is as follows: Input the image to obtain a feature map through a convolutional feature extraction network, send the feature map into the Region Proposal Network, that is, the RPN network to obtain detection results, filter duplicate results through Non Maximum Suppression, that is, NMS, and intercept the feature fragments at the corresponding positions on the feature map according to the coordinate information of the results. After passing through the Region of Interest, that is, RoI Pooling layer, the feature fragments can be processed into new feature maps of the same size.

[0040] S102. Calculate the correlation between the ship instance and the hull part instance. After the correlation coefficients between the instances are calculated to obtain a correlation matrix, sum along the coordinate axis directions and multiply to the corresponding instance feature sequences to complete the correlation calculation between the ship instance and the hull part instance;

[0041] The specific content is as follows: After the input image passes through the networks for ship detection, classification, and hull part detection, the results of ship and hull detection are output, including the coordinates of the target location: x, y, w, h, and the feature maps of the unified size after RoI Pooling. Calculate the IoU between the ship instance coordinate box and the hull part instance coordinate box in turn, and finally obtain the correlation matrix between the ship and the hull part. Sum the correlation matrix row by row or column by column, and then multiply it by the feature map of the ship or hull part to complete the correlation coefficient weighting.

[0042] S103. Send the instances in the area weighted by the correlation coefficients to the final classification and coordinate regression network to obtain the coordinate boxes and classification results of the ship and the hull part.

[0043] S2. Use the OCR text detection and recognition framework to detect the ship or ship part in step S1 and recognize the existing text information;

[0044] The specific content is as follows: The original size of the image is (W, H), the size of the image input to the ship detection, classification, and hull part detection network after scaling is (Wr, Hr), and the bbox in the output result is (x, y, w, h). Then the position coordinates bboxo(x o , y o , w o , h o) The calculation method is as shown in equations (11) and (12):

[0045]

[0046]

[0047] The above calculation method restores the position bboxo of the target in the original image, and intercepts the corresponding part in the original image according to the bboxo coordinates as the original output of the second step; performs text detection on the input picture, and intercepts the corresponding area on the input image according to the text detection result for text recognition.

[0048] S3. Use the designed text filtering algorithm based on the edit distance rule to determine the real ship name. The specific content is as follows:

[0049] a. When there are multiple recognition results, calculate the edit distance of each recognition result in turn: calculate the edit distance score of each recognition result according to the set edit distance threshold. When the edit distance is lower than the threshold for similar results, it is the ship name with recognition errors. When the edit distance is higher than the threshold for similar results, it is the one closest to the real ship name.

[0050] b. When the edit distance of two results is 0, that is, the two recognition results are exactly the same, then determine the ship name.

[0051] The preferred embodiment of the present invention includes the following three steps:

[0052] S1. Detect the ships in the video and determine the ship types. If the ships are container ships, bulk carriers, oil tankers, etc., use semantic part detection technology to detect the parts of the ships where the ship names may exist while detecting the ships. If the ships are passenger ships, law enforcement ships, etc., semantic part detection is not performed;

[0053] S2. Use the OCR text detection and recognition framework to detect the ships or ship parts detected in the first step and recognize the text information therein;

[0054] S3. Use the designed text filtering algorithm based on the edit distance rule to determine the real ship name

[0055] Such as Figure 2The network structure diagram for ship detection, classification, and hull part detection is shown. The first step in the overall structure of the present invention is as follows: First, several ship instances and hull part instances are generated through the ship region proposal network and the hull part region proposal network respectively. Each instance contains the bounding box coordinates bbox (Bounding box) and the corresponding feature sequence. The bbox contains (x, y, w, h), that is, the coordinates of the upper left corner point of the bbox and the width and height. Subsequently, the correlation between each ship instance and the hull part instance is calculated. After the correlation coefficients between all instances are calculated to obtain the correlation matrix, the sum is calculated in the coordinate axis direction and multiplied to the corresponding instance feature sequence to complete the correlation calculation between each ship instance and the hull part instance. Finally, the region proposal instances weighted by the correlation coefficients are sent into the final classification and coordinate regression network to obtain the final ship and hull part coordinate frames and classification results.

[0056] As Figure 3 shown, the ship and hull part region proposal neural networks adopt the same structure, which consists of a feature extraction network, an RPN (Region Proposal Network), and an RoI (Region of Interest) Pooling. The input image passes through the convolutional feature extraction network to obtain the feature map. The feature map is sent into the RPN network to obtain the detection results. The duplicate results are filtered through NMS (Non-Maximum Suppression). According to the coordinate information of the filtered detection results, the feature segments corresponding to the positions on the feature map are intercepted. After passing through the RoI Pooling layer, these feature segments can be processed into a new feature map of the same size.

[0057] 1) Feature extraction

[0058] The structural parameters of the feature extraction network in the region proposal neural network are shown in Table 1, where conv represents the convolutional layer, and the two parameters in the parentheses represent the convolutional kernel size and the number of channels respectively. maxpool represents the max pooling layer. There is a ReLU layer after each convolutional layer. Since it is an element-wise operation, it is ignored in the table. The length and width of the feature map obtained through the feature extraction network are both 1 / 16 of the input image.

[0059] Table 1. Feature extraction network parameter table

[0060] Network layer name Structure parameter Layer1 conv(3*3,64)conv(3*3,64)maxpool(2*2) Layer2 conv(3*3,128)conv(3*3,128)maxpool(2*2) Layer3 conv(3*3,256)conv(3*3,256)conv(1*1,256)maxpool(2*2) Layer4 conv(3*3,512)conv(3*3,512)conv(1*1,512)maxpool(2*2) Layer5 conv(3*3,512)conv(3*3,512)conv(1*1,512)

[0061] 2) RPN network

[0062] As Figure 4The proposed RPN structure in the neural network suggests that before the feature map is fed into the first 3x3 convolutional layer, nine bounding boxes (bboxes) of different sizes are preset for each point. The unit lengths are {128, 256, 512}, and the aspect ratios are {1:1, 1:2, 2:1}, which are combined pairwise. After passing through the first 3x3 convolutional layer, the length and width of the feature map remain unchanged, and the number of channels becomes 256. Subsequently, the two branches have different computational tasks. The upper 1x1 convolutional layer with Softmax is responsible for calculating the probability that the object or background is contained within each rectangular box, and the lower 1x1 convolutional layer is responsible for calculating the bbox coordinates (x, y, w, h) of the target, from which the true position of the object can be obtained.

[0063] 3) NMS algorithm

[0064] The principle of NMS in the region proposal neural network is as shown in Algorithm 1:

[0065]

[0066]

[0067] The calculation process of IoU (Intersection of Union) is as follows: For the given bbox A (x a , y a , w a , h a ), bbox B (x b , y b , w b , h b ), first calculate the coordinates of the upper left and lower right corners of the overlapping region of the two bboxes (x 1 , y 1 ), (x 2 , y 2 ), as shown in equations (1) and (2):

[0068] x 1 = max(x a , x b )), y 1 = max(y a , y b ) (1)

[0069] x 2 = min(x a + w a , x b + w b ), y 2 = min(y a + h a, y b +h b ) (2)

[0070] The coordinates of the two corner points of the overlapping region can be obtained to calculate the area S of the overlapping region i , as shown in Equation (3):

[0071] S i = max(x 2 - x 1 , 0) * max(y 2 - y 1 , 0) (3)

[0072] Next, the area S of the union region of the two bounding boxes u , that is, the bounding box A , the bounding box B minus S i , as shown in Equation (4), then IoU is the ratio of Si to Su, as shown in Equation (5):

[0073] S u = w a h a + w b h b - S i (4)

[0074]

[0075] The significance of NMS is to filter out redundant results with relatively poor confidence and position, reduce the workload of subsequent RoIPooling, and avoid a large number of repeated calculations in the subsequent precise calculations of object classification and localization.

[0076] 4) RoIPooling

[0077] As Figure 5 shown, the principle of RoIPooling in the region proposal neural network is to divide feature maps of different sizes into smaller segments according to the set output size. In the figure, the output is taken as 2*2 as an example, and the actual output size set in the network is 7*7. When the corresponding dimension cannot be divided evenly, the extra row (column) is assigned to the first segment of the corresponding dimension. As shown by the black border in the figure, the divided feature segments are obtained, and then the maximum value is calculated for each black border to obtain the value of the corresponding output position.

[0078] The advantage of RoIPooling is that it transforms the previous feature maps of different sizes into a unified size through calculation, which is convenient for the parameter setting of the subsequent neural network responsible for object correlation calculation and final classification and regression, and improves the calculation speed of the network.

[0079] Calculation of the correlation between the ship and the hull part instances. The steps for calculating the correlation between the ship and the hull part detection instances are as follows:

[0080] After the input image passes through the networks for ship detection, classification, and hull part detection, the results of ship and hull detection will be output, including the coordinates (x, y, w, h) of the target location and the feature map of a unified size after RoIPooling. Calculate the IoU of each ship instance coordinate box and the hull part instance coordinate box in turn. In the experiment, the IoU threshold γ is set to 0.9. If the IoU is less than γ, the correlation coefficient CC ij of the two instances corresponding to the bboxes will be set to zero. Otherwise, the feature maps corresponding to the two instances will be unfolded into one-dimensional sequences and concatenated, and the concatenated feature sequence will be fed into a fully connected layer neural network to obtain the correlation coefficient CC ij value. Finally, the correlation matrix of all ships and hull parts is obtained. The size of the correlation matrix is related to the number of detection results of the networks for ship detection, classification, and hull part detection. For example, if there are n ship detection results and m hull detection results, then the dimension of the correlation matrix is m*n.

[0081] Sum the correlation matrix by rows or columns, and then multiply it by the corresponding feature map of the ship or hull part, and the correlation coefficient weighting is completed. After weighting, when performing the final classification and bbox regression calculations, the ship instance also contains the information of the hull part, and the hull part also contains the information of the ship instance, realizing the information interaction between the two branches.

[0082] Ship and hull part coordinate regression and classification. The ship and hull part coordinate regression and classification networks adopt the same structure, as Figure 6 shown. fc represents the fully connected layer. Among them, the first two common fc layers are followed by ReLU layers. The top fc layer at the end is responsible for the precise regression of the target coordinates, and the bottom fc layer and softmax layer are responsible for classifying the target. Among them, ship classification includes 6 categories: container ship, bulk carrier, oil tanker, car carrier, passenger ship, and law enforcement ship. Hull classification includes bow and cabin.

[0083] Calculation of the loss of ship and hull part detection. The loss of ship and hull part detection is divided into two parts. One is the loss of the region proposal output of the ship and hull part, and the other is the loss of the final bbox regression and classification.

[0084] 1) Loss of ship and hull part region proposal output

[0085] The loss of the ship and hull part region proposal output includes the loss of the region proposal for foreground and background classification and the loss of calculating the offset of the region proposal coordinate box, as shown in Equation (6):

[0086]

[0087]

[0088] Among them C i is the foreground and background classification probability of the i-th region proposal output, C i * represents the supervision value of classification, and B i , B i * The processing of is relatively complicated. In order to make the model sensitive to the subtle differences of bbox when calculating the target position and improve the detection accuracy, the second set of outputs and corresponding labels of the region proposal network are processed, as shown in formula (7), where (x, y, w, h) is the bbox coordinate of the output of the region proposal network, (x a ,y a ,w a ,h a ) represents the width and height of the initial anchor of the position coordinate corresponding to the current output, (x * ,y * ,w * ,h * ) represents the true value of the target position used for supervision; Nc, N b Respectively indicate compliance with C i ≥0.7 or C i ≤0.3 This condition C i The number and C i ≥0.7 or C i ≤0.3 and B is 1 i In order to balance the two losses during training, the hyperparameter μ=10 is set.

[0089] In formula (6), L c and L b The specific calculation method of is as follows:

[0090]

[0091]

[0092] 2) Final output loss

[0093] The final output loss of the model includes the target classification loss and the target coordinate regression loss, as shown in formula (10), where v is a hyperparameter to ensure the balance of the two losses and is set to 5 in the experiment.

[0094]

[0095] The data collection for the ship and hull part detection dataset was carried out in the Wuhan section of the Yangtze River. All the collection and annotation work was completed from May 2020 to July 2020. The dataset consists of 221 images, which were annotated using labelme and contain a total of 235 ships, 174 bow parts, and 202 stern parts.

[0096] To ensure the robustness of the model, the model was first pre-trained using the Pascal-Part dataset. Subsequently, based on the training parameters of the Pascal-Part dataset, it was trained on the ship and hull part detection dataset. Due to the limited number of images in the dataset, all the data was used for training. The hyperparameters for both trainings are shown in Table 2:

[0097] Table 2 Ship and Hull Part Detection Training Parameter Settings

[0098] epoch 150 optimizer SGD learning rate 1.00E-03 momentum 0.9 Weightdecay 1.00E-06

[0099] After training, the AP (Average Precision) of the model is shown in Table 3:

[0100] Table 3 Ship and Hull Part Detection Average Precision

[0101] Detection target type IoU = 0.5 IoU = 0.75 Ship 99.4% 94.9% Hull part 99.1% 75.3%

[0102] In addition, to test the filtering effect of hull part detection on non-ship name text interference, 20 unannotated images were used for testing. The number of non-ship name texts that could be filtered out when the hull part area was used as the input for text detection was counted, as shown in Table 4:

[0103] Table 4. Statistical Results of the Filtering Effect of Hull Part Detection on Non-Ship Name Texts

[0104] Number of text instances Number of non-ship-name texts filtered Number of non-ship-name texts not filtered out Number of ship names 193 153 5 35

[0105] Before the algorithm process officially enters the second step, it is necessary to process the detection results of the ship detection, classification, and hull part detection network in the first step. The reason is that in order to improve the overall efficiency of the first step, the size of the input image is scaled proportionally. Therefore, the results output in the first step are naturally the positions of the targets on the scaled image. In the second step, in order to ensure the clarity of the text area in the image and improve the recognition accuracy, it is planned to intercept the relevant target parts from the original high-resolution image. Therefore, it is necessary to restore the size of the bbox in the detection results of the first step. The specific steps are as follows: Assume that the size of the original image is (W, H), the size of the image input into the ship detection, classification, and hull part detection network after scaling is (Wr, Hr), and the bbox in the output result is (x, y, w, h). Then the calculation method of the position coordinates bboxo(xo, yo, wo, ho) of the target in the original image is as shown in equations (11) and (12):

[0106]

[0107]

[0108] Restore the position bbox of the target in the original image according to the above method o After that, according to the bbox o coordinates, intercept the corresponding part in the original image as the original output of the second step. The method in the second stage is a standard two-stage method, that is, perform text detection on the input image, and intercept the corresponding area on the input image according to the text detection results and then perform text recognition.

[0109] The text detection of the potential ship name area is carried out through the following steps:

[0110] 1) Overall network structure

[0111] The text detection network is an image segmentation network constructed based on the Unet idea. As Figure 7 shown in the overall structure, after the input image enters the network, it passes through layer0 to layer4 in sequence. The width and height of the output feature map of each layer are half of the previous level.

[0112] Then, starting from the feature map of layer4 at the bottommost layer, perform upsampling successively through the ConvTp layer, and then perform element-wise addition with the output features of layer3, from bottom to top until layer0. The structure of the ConvTp layer is Conv(1*1), BN(), ReLU(), ConvTranspose(), BN(). Note that the output channels of Conv(1*1) are half of the input channels, so as to ensure that the output of the ConvTp layer can perform element-wise addition with the output of the downsampling layer. Next, the output of each stage of upsampling is extended to the width and height dimensions of the output of layer0 using the nearest-neighbor interpolation method, and then concatenated from the channel dimension and fed into the final classifier for pixel classification at different scales.

[0113] Finally, based on the segmentation results of three different scales, use the pixel set expansion algorithm to obtain the final segmentation result, and the bbox is the minimum bounding rectangle generated according to the pixel set of each text region.

[0114] 2) Downsampling network structure

[0115] The structures of each downsampling layer of the ship name potential area text detection network are shown in Table 5. The first two parameters of the Conv convolutional layer represent the number of input and output channels respectively, k represents the kernel size, s represents the convolution stride, and p represents the padding dimension at the edge; in particular, the feature map passes through the layers in red font to achieve downsampling, and the number on the rightmost side represents the number of repeated superpositions of the corresponding part, which is folded for easy reading.

[0116] Table 5 Structures of each downsampling layer of the ship name potential area text detection network

[0117]

[0118]

[0119] 3) Multi-scale classifier and supervision method

[0120] To accurately detect the boundaries of the text region, a multi-scale progressive pixel classifier that gradually expands outward from the center region of the text is designed. As Figure 7In the Multi-scale classifier, after the output results of each stage are interpolated, expanded, and concatenated, they pass through a Conv(3*3,256) convolutional layer (there are also BN() and ReLU() after the convolutional layer, which are omitted in the figure for simplicity of expression), and then enter three branches respectively. The three branches have the same structure. After a 1*1 convolution, the Sigmoid function is used for classification, and the results P1, P2, and P3 are output respectively. Among them, P1 is responsible for distinguishing text and non-text regions, and P2 and P3 are responsible for distinguishing the text and text regions after boundary contraction, so as to generate accurate text region boundaries in the subsequent process.

[0121] The three segmentation branches go from bottom to top, and the supervision labels are heatmaps generated by shrinking the original labels by 40%, 20%, and 0% with the center of the original text region as the bbox.

[0122] As Figure 8 shown, the calculation method for shrinking the text region of the supervision heatmap is assumed that the black bbox in the figure is the original text region in the annotation, and the four vertex coordinates in the coordinate system are the black (x 1 ,y 1 ), (x 2 ,y 2 ), (x 3 ,y 3 ), (x 4 ,y 4 ) in the figure.

[0123] First, the center point coordinates (x c ,y c ) of the bbox need to be obtained. The calculation formula for the center point coordinates is as shown in equations (13) and (14):

[0124]

[0125]

[0126] Secondly, the distances between the center point and the four vertices in the X and Y directions need to be obtained respectively. Taking (x 1 ,y 1 ) as an example, as shown in equations (15) and (16):

[0127] d x1 =x c -x 1 (15)

[0128] d y1 =y c -y 1 (16)

[0129] After obtaining d x1 、d y1After that, the coordinates of the first vertex of the bbox after contracting by a certain proportion centered can be obtained as (x 1shrink , y 1shrink ), as shown in Equations (17) and (18):

[0130] x 1shrink = x c - (1 - r)d x1 (17)

[0131] y 1shrink = y c - (1 - r)d y1 (18)

[0132] Among them, the letter r represents the contraction rate of the bbox. Modify the value of r to calculate the bbox for supervising classifiers at different scales, and finally generate the corresponding heatmap according to the bbox coordinates.

[0133] Combined with the development of natural scene text inspection algorithms in recent years, DiceLoss is selected as the loss function, as shown in Equation (19):

[0134]

[0135] Among them, P i,x,y represents the value at the coordinate point (x, y) of the output P i of the i-th order, that is, the predicted value, and G i,x,y represents the value at the coordinate point (x, y) of the supervised heatmap G i of the i-th order, that is, the true value.

[0136] The first-order output P 1 is responsible for distinguishing text and non-text regions. Therefore, the loss calculation strategy for the first order starts from the global perspective to evaluate the classification ability of P 1 ; moreover, in order to improve the calculation efficiency, the OHEM (Online Hard Example Mining) algorithm is added to process during the first-order loss calculation, and the positive and negative sample ratio is set to 1:3, as shown in Equation (20):

[0137] L c = 1 - D(P 1 * M, G 1 * M) (20)

[0138] Among them, M represents the heatmap generated after adding OHEM.

[0139] The outputs P 2 and P 3Responsible for distinguishing the text after boundary contraction and the text region. The focus is on the part of the original text region that was originally positive and becomes negative after contraction. Therefore, when calculating the loss of these two-stage outputs, it is set to ignore the part judged as non-text region in the first-stage output, as shown in Equation (21):

[0140]

[0141] Where W is the heat map screened according to the output of P 1 The selection principle is as shown in Equation (22):

[0142]

[0143] In summary, the loss function Loss for text detection in the potential ship name region is as shown in Equation (23), where λ is a hyperparameter used to balance each loss and is set to 0.7 in the experiment.

[0144] Loss = λL c +(1 - λ)L s (23)

[0145] 4) Pixel set expansion algorithm

[0146] In order to obtain an accurate segmentation result based on the multi-scale output results, the pixel set expansion algorithm merges the results of the multi-scale classifiers step by step to obtain the final text region segmentation result, and calculates the minimum bounding rectangle of each text region according to the segmentation heat map. The process of the pixel set expansion algorithm is as shown in Algorithm 2.

[0147]

[0148]

[0149] 5) Data collection and annotation

[0150] The data for text detection in the potential ship name region was taken from multiple shots in the Wuhan section of the Yangtze River since 2019. Currently, there are 1,220 images in total, and 2,542 text detection instances have been annotated. The division ratio of the training set to the test set is 4:1.

[0151] The annotation tool used is labelme, and the four-point method is used to frame the text region. The annotation information is the coordinates of the four vertices of the text region in the image.

[0152] 6) Experiment

[0153] First, to ensure accuracy and robustness, the ICDAR2017 MLT text detection dataset was used to pre-train the model. The specific experimental settings are shown in Table 6:

[0154] Table 6 Text detection pre-training parameter settings

[0155] epoch 600 optimizer SGD learning rate 1.00E-03 momentum 0.99 weight decay 5.00E-04 lr step 0.1(200,400)

[0156] Subsequently, based on the pre-trained model parameters, training was carried out on the text detection dataset in the potential area of ship names. The specific experimental settings are shown in Table 7:

[0157] Table 7 Training parameter settings for text detection in the potential area of ship names

[0158]

[0159]

[0160] In addition, the performances of several excellent natural scene text detection methods were tested and compared with the present invention. As shown in Table 8, to ensure fairness, for all methods, the training was carried out on the text detection dataset in the potential area of ship names based on the parameters trained on the public dataset.

[0161] Table 8 Experimental accuracy (IoU = 0.5)

[0162] Method Recall Precision H-mean DBNet 0.776 0.946 0.847 DRRG 0.781 0.877 0.826 FCENet 0.773 0.822 0.797 Mask RCNN 0.89 0.836 0.862 PANet 0.628 0.869 0.729 Our method 0.861 0.868 0.864

[0163] As can be seen from the above table, although the present invention does not reach the optimal level, it is not inferior to the existing text detection algorithms at all.

[0164] The steps for recognizing the text detection results are as follows:

[0165] 1) Structural principle

[0166] As Figure 9 shown, for the overall structure of the text recognition network, the input is a segment intercepted from the original image according to the output of the text detection model. After feature extraction by ResNet31, the feature map is separated in the channel dimension and fed into the LSTM encoder in the form of a sequence. The decoder combines the output of the encoder and the information calculated by the attention mechanism to sequentially output the recognition results.

[0167] Among them, both the encoder and the decoder are LSTM recurrent network structures with two layers of 512 hidden states. After separating the feature map along the channel dimension, vertical max-pooling is first performed to compress the dimension, and then it is fed into the encoder for encoding. During training, the input of the decoder is the character sequence of the label; during inference, the input of the decoder is the output of the previous time step.

[0168] To improve the accuracy, a two-dimensional attention mechanism is introduced, which calculates the attention map by combining the hidden state of the decoder and the feature map extracted by the CNN. The feature map is first convolved with a 3x3 kernel, and the hidden state of the decoder at the current time step is convolved with a 1x1 kernel. Then, the two results are added along the channel dimension. After that, the two-dimensional attention feature map is obtained through attention calculation. The two-dimensional attention map and the CNN feature map are weighted and summed along the channel dimension. Finally, the weighted feature map is summed and reduced in dimension along the channel dimension. The reduced weighted feature is concatenated with the hidden state of the decoder at the current time step and then passed through a fully connected layer to be transformed into the dimension of the label dictionary, and the output at the current time step is determined by the Softmax layer for classification.

[0169] 2) Data Generation

[0170] To improve the accuracy, an artificial data generation method is adopted to expand the ship name recognition dataset. First, 2754 Chinese ship names and 3226 English and pinyin ship names are collected through web crawling. Then, more than 30,000 training images and 5980 test images are artificially generated using image synthesis technology, and the label dictionary of SynthText is used.

[0171] 3) Experiments

[0172] To ensure the robustness of the model, the model is first trained on the SynthText dataset, and the specific parameters are shown in Table 9:

[0173] Table 9 Settings of Recognition Training Parameters for SynthText

[0174]

[0175] Subsequently, based on the training parameters of SynthText, ship name recognition experiments are carried out, and the specific parameters are shown in Table 10:

[0176] Table 10 Settings of Recognition Training Parameters for Ship Names

[0177]

[0178] After training, the recognition accuracy of the model is shown in Table 11:

[0179] Table 11 Ship Name Recognition Accuracy

[0180] Ship name recognition accuracy Character recognition accuracy 95.42% 99.19%

[0181] The specific steps of the text filtering algorithm are as follows:

[0182] As Figure 10As shown in the flowchart of the text filtering algorithm, if there are multiple recognition results, the edit distance of each recognition result is calculated in turn. If there are two results whose edit distance is 0, that is, the two recognition results are exactly the same, then it is determined that this is the ship name to be found. Otherwise, the edit distance score of each recognition result is calculated according to the set edit distance threshold. The edit distance threshold is set here mainly to ensure the robustness of the entire method. It is hoped that when similar situations occur, such as errors in the recognition results that may be caused by the unsatisfactory quality of the actual scene image, it can also provide reasonable reference results to the downstream. Similar results with edit distances lower than the threshold are identified as ship names with recognition errors. At this time, it can be determined that the result with a higher recognition confidence is the closest to the real ship name.

[0183] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation modes, which are merely illustrative rather than restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are within the protection of the present invention.

Claims

1. An active recognition method for ship names for inland river maritime video surveillance, characterized in that, it includes the following steps: S1. Detect ships in the video and determine the ship types: When the ships are container ships, bulk carriers and oil tankers, use the semantic part for detection. While detecting the ships, detect the ship parts with ship names. When the ships are passenger ships and law enforcement ships, semantic part detection is not performed; The step S1 includes the following steps: S101. Input the ship instances and hull part instances generated by the ship region proposal network and the hull part region proposal network respectively. The instances contain the bounding box coordinates bbox, that is, Bounding box, and the corresponding feature sequences. The bounding box coordinates bbox contain the upper left corner point coordinates and the width and height, which are: x, y, w, h respectively; S102. Calculate the correlation between the ship instance and the hull part instance. After the correlation coefficients between the instances are calculated to obtain the correlation matrix, sum along the coordinate axes and multiply to the corresponding instance feature sequences to complete the calculation of the correlation between the ship instance and the hull part instance; S103. Send the instances in the region weighted by the correlation coefficients to the final classification and coordinate regression network to obtain the coordinate frames and classification results of the ships and hull parts; S2. Use the OCR text detection and recognition framework to detect the ships or ship parts in step S1 and recognize the existing text information; The method of detecting using the OCR text detection and recognition framework in step S2 is as follows: The original image sizes of the ship instance and the hull part instance are ( W, H ), and the image sizes of the ship detection, classification, and hull part detection network after scaling are ( Wr , Hr ). If the bbox in the output result is ( x, y, w, h ), then the position coordinates bboxo ( x o 、y o 、w o 、h o ) of the target in the original image are calculated by equations (11) and (12) as follows: (11) (12) The calculation method restores the position bboxo of the target in the original image, and intercepts the corresponding part in the original image according to the bboxo coordinates as the original output of the second step; perform text detection on the input picture, and intercept the corresponding area on the input image according to the text detection result for text recognition; S3. Use a text filtering algorithm based on the edit distance rule to determine the real ship name; The text filtering algorithm based on the edit distance rule in the step S3 is: a. When there are multiple recognition results, calculate the edit distances of each recognition result in turn: Calculate the edit distance scores of each recognition result according to the set edit distance threshold. When the edit distance is lower than the threshold for similar results, it is a ship name with recognition errors. When the edit distance is higher than the threshold for similar results, it is the one closest to the real ship name; b. When the edit distances of two results are 0, that is, the two recognition results are exactly the same, determine the ship name.

2. The active recognition method for ship names for inland river maritime video surveillance according to claim 1, characterized in that, The method of inputting the ship instances and hull part instances generated by the ship area proposal network and the hull part area proposal network respectively in step S101 is as follows: The images of the ship instances and hull part instances are input into a convolutional feature extraction network to obtain a feature map. The feature map is sent to the Region Proposal Network, that is, the RPN network, to obtain detection results. The duplicate results are filtered through Non Maximum Suppression, that is, NMS. According to the coordinate information of the results, the feature segments corresponding to the positions on the feature map are intercepted. After passing through the Region of Interest, that is, the RoI Pooling layer, the feature segments are processed into new feature maps of the same size.

3. The active ship name recognition method for inland river maritime video surveillance according to claim 2, characterized in that the method of obtaining the correlation matrix in step S102 is as follows: The images of the ship instances and hull part instances are input. After passing through the network for ship detection, classification, and hull part detection, the results of ship and hull detection are output, including the coordinates of the target location: x, y, w, h, and the feature map of the unified size after RoI Pooling. The IoU between the coordinate frames of the ship instances and the coordinate frames of the hull part instances is calculated in sequence, and finally the correlation matrix of the ship and the hull part is obtained; the method of completing the correlation calculation between the ship instance and the hull part instance is: summing the correlation matrix row by row or column by column, and then multiplying it by the feature map of the ship or the hull part to complete the weighting of the correlation coefficient.

Citation Information

Patent Citations

  • Intelligent ship identity recognition method and system based on twin network

    CN112232269A

  • System and Method for Extremely Efficient Image and Pattern Recognition and Artificial Intelligence Platform

    US20200184278A1