A method for identifying a ship name of an inland river ship by combining autoregressive and non-autoregressive decoding

CN119206695BActive Publication Date: 2026-09-25HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411193552.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2026-09-25
Estimated Expiration
2044-08-28

AI Technical Summary

Technical Problem

[0012]本发明的目的在于提供一种组合自回归与非自回归解码的内河船舶船名识别方法,改善了以往视觉-语言多模态船名识别模型繁琐的训练过程和对外部语言模型的依赖,这些以往方法由于利用外部语料训练语言模型,产生了词库依赖问题,对视觉模型预测正确的结果出现了矫枉过正的现像

Benefits of technology

[0052]本发明针对现有船舶船名识别方法在实际应用中效率的不足,通过结合AR和NAR解码机制构建一种简单高效的船舶船名识别的集成架构。融合字符表示动态整合视觉以及语言线索,以自适应地调整它们在识别语义相关性较弱的船舶名称方面的贡献,以及使用实例级和基于字符级置信度的重新加权机制来明确挖掘困难数据,促使模型从困难船名文本图像中学习更多。大幅度得解决了实际应用中拍摄不清、模糊、光照不均等情况的船名图像的识别问题,构建了更准确、更实用的船舶船名识别模型。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119206695B_ABST
    Figure CN119206695B_ABST
Patent Text Reader

Abstract

The application discloses a kind of methods for identifying ship name of inland river ship by combining autoregressive and non-autoregressive decoding.The application is as follows:1, data set construction;2, feature extractor pre-training;3, build ship name recognition model: the ship name image features extracted by feature extractor are sent into autoregressive and non-autoregressive decoder branches respectively, to obtain two types of character representations;4, adaptively fuse two types of character representations through gating mechanism;5, take the character representations extracted by autoregressive decoder branch and non-autoregressive decoder branch and the fused character representations as inputs respectively, build three linear classifiers, to obtain three groups of predictions;6, by fusing the losses of autoregressive decoder branch and non-autoregressive decoder branch, respectively build character-level and instance-level re-weighted weights, dynamically adjust the contribution of batch samples to the total loss of the current batch during backpropagation.The application eliminates the cumbersome training process and dependence on external language models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a visual-linguistic multimodal ship name recognition method that dynamically fuses autoregressive and non-autoregressive character representations. Specifically, it is a method for recognizing inland waterway ship names by combining autoregressive and non-autoregressive decoding. Background Technology

[0002] Inland waterway transport, characterized by low pollution, large capacity, and low energy consumption, has profoundly influenced economic development and cultural dissemination throughout human history. However, with the increase in the number of inland waterway vessels and the volume of inland waterway traffic, regulatory pressure has become increasingly severe. This is mainly manifested in the fact that some inland waterway transport vessels attempt to evade supervision by disabling their Automatic Identification Systems (AIS). This places a tremendous burden on maintaining order in inland waterway shipping.

[0003] Meanwhile, there is also a waste in the current application of information resources in the shipping system. Specifically, although numerous high-definition cameras have been installed at docks, ports, or along canals, these devices currently mostly serve only to provide enforcement evidence for regulatory authorities. This image or video data contains a wealth of information useful for ship identification, such as ship names. This patent proposes an efficient ship name recognition algorithm.

[0004] Current research on ship name recognition technology is limited, and the challenges in terms of image quantity, image quality, and language semantics have not yet been fully resolved. Specifically, these challenges are reflected in the following aspects:

[0005] 1. The complexity of the training process:

[0006] Current ship license plate recognition systems rely on complex training procedures and additional corpora to improve recognition performance. This reliance not only increases the difficulty of system implementation but also limits their applicability and flexibility in diverse real-world application scenarios.

[0007] 2. The issue of dependence on external language models:

[0008] Existing methods typically integrate pre-trained language models to correct errors made by the visual model during recognition. However, this reliance on external language models can lead to the recognizer's performance becoming overly dependent on the accuracy of the language model, which is inconsistent with the weak semantic relevance of ship license plate text images, potentially affecting the overall performance and generalization ability of the recognizer.

[0009] 3. The issue of efficiency adaptability to real-world application scenarios:

[0010] Because the semantic relevance of character sequences in ship license plate text images is weak, existing language model-based recognition methods have limitations in terms of processing efficiency and adaptability to practical applications. In particular, existing technologies may fail to efficiently and accurately complete the recognition task when faced with irregular layouts and variations in appearance.

[0011] To address the aforementioned challenges, overcome the difficulties in ship name recognition, and provide solid technical support for the construction of an intelligent integrated management system for inland waterway transportation, this invention proposes a novel ship name recognizer. By adaptively merging dual-branch character representations, it avoids the cumbersome training process and eliminates dependence on external language models. Summary of the Invention

[0012] The purpose of this invention is to provide a method for recognizing inland waterway vessel names by combining autoregressive and non-autoregressive decoding. This method improves upon the cumbersome training process and dependence on external language models of previous visual-linguistic multimodal vessel name recognition models. These previous methods, due to their reliance on external corpora for language model training, introduced a lexicon dependency problem, resulting in overcorrection of the visual model's correct predictions. The technology proposed in this patent adaptively fuses representations from the non-autoregressive visual model branch and the autoregressive language model branch at the character-level representation level, eliminating the cumbersome training process and the dependence on external language models. Furthermore, this invention introduces a confidence-based dynamic sample loss reweighting strategy, including two confidence-based hard example mining losses at the character and instance levels. This allows the model training process to benefit more from difficult or challenging samples. Simultaneously, considering the scale variability of vessel name text images and the problem of ambiguous samples, this invention constructs a pre-training process for a visual feature extractor based on masked image modeling, used to initialize the feature extraction module of the vessel name recognizer.

[0013] To achieve the above objectives, the technical solution of the present invention mainly includes the following steps:

[0014] Step 1: Ship Name Recognition Dataset Construction: Images containing ship names are captured using a camera, the ship name text lines are cropped, and the data is labeled using LabelMe annotation software. The text content is used as the annotation information to create a recognition training set.

[0015] Step 2, Visual Encoding Model Pre-training: The network is modeled based on masked images, and the feature extractor is pre-trained using self-supervised learning.

[0016] Step 3: Construct the ship name recognition model: Initialize the pre-trained feature extractor from Step 2. The extracted ship name image features are fed into the autoregressive decoder branch and the non-autoregressive decoder branch respectively to obtain two types of character representations.

[0017] Step 4: Adaptively fuse the representations of the autoregressive decoder branch and the non-autoregressive decoder branch through a gating mechanism.

[0018] Step 5: Model Training: Using the character representations extracted by the autoregressive decoder branch and the non-autoregressive decoder branch, and the fused character representation as inputs, three linear classifiers are constructed to obtain three sets of predictions. The model is then trained under the supervision of cross-entropy loss.

[0019] Step 6: Implement hard sample mining. By fusing the losses of the autoregressive decoder branch and the non-autoregressive decoder branch, construct character-level and instance-level reweighted weights respectively, and dynamically adjust the contribution of batch samples to the total loss of the current batch during backpropagation.

[0020] The specific steps of step 1 are as follows:

[0021] Step 1-1: On the banks of the Qiantang River, take images with ships and their names using a camera. By adjusting the focal length and shooting angle, try to capture images that cover as many different aspects as possible, such as scale, background, lighting conditions, and shooting location.

[0022] Steps 1-2: Based on the perspective transformation of the four corner coordinates of the ship name text line, perform coarse correction on the text line area to obtain a small-scale ship name text line image.

[0023] Steps 1-3: While maintaining the aspect ratio of the images, a series of data augmentation techniques were employed, including but not limited to random rotation, random contrast adjustment, random scale adjustment, random resolution adjustment, and random blurring, to process the small-scale image dataset. Subsequently, the processed images were precisely placed on a 100×32 grayscale template. Data augmentation was used to improve the generalization ability of the optimized model during subsequent training.

[0024] Steps 1-4: Ship Name Text Recognition Model Annotation. The cropped and expanded ship name text images are saved as .txt files with "filename annotation" format, serving as annotation information. This completes the creation of the ship name text recognition model training dataset.

[0025] Step 2, specifically, is as follows:

[0026] Step 2-1: Select the VIT network and modify it by deleting the class tokens used for classification and retaining only the embedding vectors corresponding to the image patches as the extracted image features.

[0027] Steps 2-3: By randomly selecting and masking a certain proportion of the ship name image, the network is prompted to learn to predict the features of the masked region from the remaining region. Mean Squared Error (MSE) or contrast loss is used to measure the difference between the predicted features and the true features, ensuring that the network can effectively recover the masked region.

[0028] Steps 2-4: Configure the training process. Select Adam as the optimizer, set the initial learning rate, and dynamically adjust the learning rate according to the training progress. Introduce L2 regularization and Dropout techniques to reduce the risk of model overfitting and enhance the model's generalization ability to new data.

[0029] Steps 2-5: Periodically evaluate the model's performance during training. Evaluate the model's performance by calculating metrics such as the loss function value on the validation set, and fine-tune it as needed to achieve satisfactory performance on the self-supervised learning task.

[0030] Step 3, the specific steps are as follows:

[0031] Step 3-1: Input the ship's name and image from the input terminal. The ship name image features are obtained by outputting the feature extractor trained in step 2. Let C represent the set of real numbers, C represent the number of output image channels, and H×W represent the height and width of the output image.

[0032] Step 3-2: Construct an Autoregressive Decoder (AR). The autoregressive decoder acts as a language model to capture the semantic relationships between sequences of ship name characters. Therefore, a Transformer decoder is used as the autoregressive decoder, which includes three key modules: a masked multi-head self-attention module, a multi-head cross-attention module, and a feed-forward network (FFN) module. The autoregressive decoder branches use right-shifted text annotations. (for known parameters) (y) T The input consists of the character representing time step T and the ship name image feature X obtained in step 3-1. The input is then processed layer by layer through a masked multi-head self-attention module, a multi-head cross-attention module, and a feedforward network module to finally extract context-aware autoregressive character representation features. The predicted probability of a character at time step t is expressed as: That is, the probability of predicting time step t. The input needs to be the output within the time interval from 1 to t-1. And the ship name image feature X, where the softmax function is used to convert a set of real numbers into a probability distribution, such that the output value is in the range [0,1] and the sum of all output values ​​is 1.θ Represents a linear classifier, l t Let t be the character representation feature at time step t.

[0033] Step 3-3: Construct a Non-Autoregressive Decoder (NAR). The NAR decoder employs a multi-attention dynamic fusion mechanism to refine the decoder's input features. The ship name image features X obtained in Step 3-1 are enhanced using three independent components: channel attention, spatial attention, and self-attention. Then, the enhanced features are further enhanced through a gated full fusion mechanism. Dynamic fusion yields improved features Finally, attention is obtained through character order query, which will improve the features. Convert to non-autoregressive branch character representation feature F NAR =[v1,...,v T ]. At the same time through Generate the complete target transcribed text, where the argmax function is used to find the value of the independent variable when the function reaches its maximum value in a given domain. The characters predicted by the non-autoregressive decoder.

[0034] Step 4, specifically, is as follows:

[0035] Step 4-1: Apply a gating mechanism to the outputs of Steps 3-2 and 3-3 for adaptive character representation fusion to obtain the fusion feature F, as shown in the following formula:

[0036] G=σ([F AR ,F NAR ]W)

[0037]

[0038] in, Let σ represent the learnable weights, and let σ represent the non-linear activation function sigmoid. This represents element-wise multiplication. It is F AR and F NAR The adaptive weight matrix.

[0039] The final output fused feature F is determined using the following formula:

[0040]

[0041] in The character representing the prediction at time step t. g represents the predicted character within time steps 1-t. t L represents the adaptive weights at time step t.t v represents the autoregressive character representation feature at time step t. t The non-autoregressive character representation feature is used to represent the time step t.

[0042] Step 5, the specific steps are as follows:

[0043] Step 5-1: The character representation F extracted from the autoregressive decoder branch and the non-autoregressive decoder branch in steps 3-3 and 3-4. AR and F NAR Using the fused features F from step 4-1 as input, three linear classifiers c are constructed. θ Three sets of predictions were obtained. Among them, c θ Its main function is to divide the input space into different regions, each region mapping to a category. It learns the weights of features through training data to determine which features are more important for classification decisions. Specifically, it takes the form f(x) = w·x + b, where x is the input feature vector; w is the weight vector, representing the direction in the feature space and determining the classification decision boundary; and b is the bias term, controlling the position of the decision boundary.

[0044] Step 6, the specific steps are as follows:

[0045] Step 6-1: During training, predictions for ship license plate text images with incomplete visual information (such as occlusion or blur) show low confidence scores. Therefore, adding these challenging ship license plate text images can enhance the robustness of the model's recognition. The loss contribution of difficult, misidentified samples is inversely proportional to the prediction probability, i.e. To refine the learnable weights W to better handle data information generated when learning from difficult data, two weights are designed: instance-level weights w. i and character-level weight w i This allows for dynamic adjustment of the contribution of batch samples to the total loss of the current batch during backpropagation.

[0046] Step 6-2: Design instance-level weights w i Instance-level weights It is generated by multiplying the confidence scores over T time steps, and then reweighting the recognition loss at the granularity of text images. The specific implementation formula is as follows: The sigmoid function compresses any real value into the range of 0 to 1, which is equivalent to performing an activation function operation.

[0047] Step 6-3: Design character-level weights w i Character-level weighting Character level This represents assigning different weights to the recognition loss at time step t for the i-th character, and the specific implementation formula is as follows: To balance learning from all data and learning from difficult data, a warm-up mechanism is employed, defining the reweighted working period as t*E, where E represents a predefined period. Instance-level weights w i Or character-level weight w i Based on fusion loss L Fused Generated.

[0048] Step 6-4: Combining steps 6-1 to 6-3, design the following loss function:

[0049]

[0050] The model is trained under the supervision of this cross-entropy loss, where B represents the number of training batches, and the CrossEntropy function represents the cross-entropy loss function, which measures the difference between the probability distribution predicted by the model and the probability distribution of the true labels. The branch loss L of the autoregressive decoder is obtained by solving the cross-entropy loss function. AR Non-autoregressive decoder branch loss L NAR and fusion loss L Fused The final overall optimization objective is determined to be L = αL. AR +βL NAR +γL Fused The loss reweighting method, where α, β, and γ are used as balancing factors, aims to improve the model's robustness to difficult-to-handle samples, thereby improving the recognition performance of ship name texts.

[0051] The beneficial effects of this invention are as follows:

[0052] This invention addresses the inefficiencies of existing ship name recognition methods in practical applications by constructing a simple and efficient integrated architecture for ship name recognition by combining AR and NAR decoding mechanisms. It dynamically integrates visual and linguistic cues through character representation to adaptively adjust their contributions to recognizing ship names with weak semantic relevance. Furthermore, it employs instance-level and character-level confidence-based reweighting mechanisms to explicitly mine challenging data, enabling the model to learn more from difficult ship name text images. This significantly solves the problem of recognizing ship name images that are unclear, blurry, or unevenly lit in practical applications, resulting in a more accurate and practical ship name recognition model. Attached Figure Description

[0053] Figure 1 A detailed flowchart of the present invention. Detailed Implementation

[0054] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0055] Step 1: Ship Name Recognition Dataset Construction: Images containing ship names are captured using a camera, the ship name text lines are cropped, and the data is labeled using LabelMe annotation software. The text content is used as the annotation information to create a recognition training set.

[0056] Step 2, Visual Encoding Model Pre-training: The network is modeled based on masked images, and the feature extractor is pre-trained using self-supervised learning.

[0057] Step 3: Construct the ship name recognition model: Initialize the pre-trained feature extractor from Step 2. The extracted ship name image features are fed into the autoregressive decoder branch and the non-autoregressive decoder branch respectively to obtain two types of character representations.

[0058] Step 4: Adaptively fuse the representations of the autoregressive decoder branch and the non-autoregressive decoder branch through a gating mechanism.

[0059] Step 5: Model Training: Using the character representations extracted by the autoregressive decoder branch and the non-autoregressive decoder branch, and the fused character representation as inputs, three linear classifiers are constructed to obtain three sets of predictions. The model is then trained under the supervision of cross-entropy loss.

[0060] Step 6: Implement hard sample mining. By fusing the losses of the autoregressive decoder branch and the non-autoregressive decoder branch, construct character-level and instance-level reweighted weights respectively, and dynamically adjust the contribution of batch samples to the total loss of the current batch during backpropagation.

[0061] The specific steps of step 1 are as follows:

[0062] Step 1-1: On the banks of the Qiantang River, take images with ships and their names using a camera. By adjusting the focal length and shooting angle, try to capture images that cover as many different aspects as possible, such as scale, background, lighting conditions, and shooting location.

[0063] Steps 1-2: Based on the perspective transformation of the four corner coordinates of the ship name text line, perform coarse correction on the text line area to obtain a small-scale ship name text line image.

[0064] Steps 1-3: While maintaining the aspect ratio of the images, a series of data augmentation techniques were employed, including but not limited to random rotation, random contrast adjustment, random scale adjustment, random resolution adjustment, and random blurring, to process the small-scale image dataset. Subsequently, the processed images were precisely placed on a 100×32 grayscale template. Data augmentation was used to improve the generalization ability of the optimized model during subsequent training.

[0065] Steps 1-4: Ship Name Text Recognition Model Annotation. The cropped and expanded ship name text images are saved as .txt files with "filename annotation" format, serving as annotation information. This completes the creation of the ship name text recognition model training dataset.

[0066] Step 2, specifically, is as follows:

[0067] Step 2-1: Select the VIT network and modify it by removing the class tokens used for classification and retaining only the embedding vectors corresponding to the image patches as extracted image features, such as... Figure 1 The “Specific Structure of VIT Encoder” section is shown below.

[0068] Steps 2-3, such as Figure 1 As shown in the "Pre-training of Ship Name Feature Extractor Based on Masked Image Modeling" section, by randomly selecting and masking a certain proportion of the ship name image, the network is prompted to learn to predict the features of the masked region from the remaining region. Mean Squared Error (MSE) or contrast loss is used to measure the difference between the predicted features and the real features, ensuring that the network can effectively recover the masked region.

[0069] Steps 2-4: Configure the training process. Select Adam as the optimizer, set the initial learning rate, and dynamically adjust the learning rate according to the training progress. Introduce L2 regularization and Dropout techniques to reduce the risk of model overfitting and enhance the model's generalization ability to new data.

[0070] Steps 2-5: Periodically evaluate the model's performance during training. Evaluate the model's performance by calculating metrics such as the loss function value on the validation set, and fine-tune it as needed to achieve satisfactory performance on the self-supervised learning task.

[0071] Step 3, the specific steps are as follows:

[0072] Step 3-1: Input the ship's name and image from the input terminal. The ship name image features are obtained by outputting the feature extractor trained in step 2. Let C represent the set of real numbers, C represent the number of output image channels, and H×W represent the height and width of the output image. A structure is constructed as follows: Figure 1 The model structure shown is "Construction of a ship name recognizer based on combined autoregressive and non-autoregressive decoding".

[0073] Step 3-2: Construct an Autoregressive Decoder (AR). The autoregressive decoder acts as a language model to capture the semantic relationships between sequences of ship name characters. Therefore, a Transformer decoder is used as the autoregressive decoder, which includes three key modules: a masked multi-head self-attention module, a multi-head cross-attention module, and a feed-forward network (FFN) module. The autoregressive decoder branches use right-shifted text annotations. (for known parameters) (y) T The input consists of the character representing time step T and the ship name image feature X obtained in step 3-1. The input is then processed layer by layer through a masked multi-head self-attention module, a multi-head cross-attention module, and a feedforward network module to finally extract context-aware autoregressive character representation features. The predicted probability of a character at time step t is expressed as: That is, the probability of predicting time step t. The input needs to be the output within the time interval from 1 to t-1. And the ship name image feature X, where the softmax function is used to convert a set of real numbers into a probability distribution, such that the output value is in the range [0,1] and the sum of all output values ​​is 1. θ Represents a linear classifier, l t Let t be the character representation feature at time step t.

[0074] Step 3-3: Construct a Non-Autoregressive Decoder (NAR). The NAR decoder employs a multi-attention dynamic fusion mechanism to refine the decoder's input features. The ship name image features X obtained in Step 3-1 are enhanced using three independent components: channel attention, spatial attention, and self-attention. Then, the enhanced features are further enhanced through a gated full fusion mechanism. Dynamic fusion yields improved features Finally, attention is obtained through character order query, which will improve the features. Convert to non-autoregressive branch character representation feature F NAR =[v1,...,v T ]. At the same time through Generate the complete target transcribed text, where the argmax function is used to find the value of the independent variable when the function reaches its maximum value in a given domain. The characters predicted by the non-autoregressive decoder.

[0075] Step 4, the specific steps are as follows:

[0076] Step 4-1: Apply a gating mechanism to the outputs of Steps 3-2 and 3-3 for adaptive character representation fusion to obtain the fusion feature F, as shown in the following formula:

[0077] G=σ([F AR ,F NAR ]W)

[0078]

[0079] in, Let σ represent the learnable weights, and let σ represent the non-linear activation function sigmoid. This represents element-wise multiplication. It is F AR and F NAR The adaptive weight matrix.

[0080] The final output fused feature F is determined using the following formula:

[0081]

[0082] in This represents the predicted character at time step t. This represents the predicted character within time steps 1-t, g t L represents the adaptive weights at time step t. t v represents the autoregressive character representation feature at time step t. t The non-autoregressive character representation features at time step t.

[0083] Step 5, the specific steps are as follows:

[0084] Step 5-1: The character representation F extracted from the autoregressive decoder branch and the non-autoregressive decoder branch in steps 3-3 and 3-4. AR and F NAR Using the fused features F from step 4-1 as input, three linear classifiers c are constructed. θ Three sets of predictions were obtained. Among them, c θ Its main function is to divide the input space into different regions, each region mapping to a category. It learns the weights of features through training data to determine which features are more important for classification decisions. Specifically, it takes the form f(x) = w·x + b, where x is the input feature vector; w is the weight vector, representing the direction in the feature space and determining the classification decision boundary; and b is the bias term, controlling the position of the decision boundary.

[0085] Step 6, the specific steps are as follows:

[0086] Step 6-1: During training, predictions for ship license plate text images with incomplete visual information (such as occlusion or blur) show low confidence scores. Therefore, adding these challenging ship license plate text images can enhance the robustness of the model's recognition. The loss contribution of difficult, misidentified samples is inversely proportional to the prediction probability, i.e. To refine the learnable weights W to better handle data information generated when learning from difficult data, two weights are designed: instance-level weights w. i and character-level weight w i This allows for dynamic adjustment of the contribution of batch samples to the total loss of the current batch during backpropagation.

[0087] Step 6-2: Design instance-level weights w i Instance-level weights It is generated by multiplying the confidence scores over T time steps, and then reweighting the recognition loss at the granularity of text images. The specific implementation formula is as follows: The sigmoid function compresses any real value into the range of 0 to 1, which is equivalent to performing an activation function operation.

[0088] Step 6-3: Design character-level weights w i Character-level weighting Character level This represents assigning different weights to the recognition loss at time step t for the i-th character, and the specific implementation formula is as follows: To balance learning from all data and learning from difficult data, a warm-up mechanism is employed, defining the reweighted working period as t*E, where E represents a predefined period. Instance-level weights w i Or character-level weight w i Based on fusion loss L Fused Generated.

[0089] Step 6-4: Combining steps 6-1 to 6-3, design the following loss function:

[0090]

[0091] The model is trained under the supervision of this cross-entropy loss, where B represents the number of training batches, and the CrossEntropy function represents the cross-entropy loss function, which measures the difference between the probability distribution predicted by the model and the probability distribution of the true labels. The branch loss L of the autoregressive decoder is obtained by solving the cross-entropy loss function. AR Non-autoregressive decoder branch loss L NAR and fusion loss L Fused The final overall optimization objective is determined to be L = αL. AR +βLNAR +γL Fused The loss reweighting method, where α, β, and γ are used as balancing factors, aims to improve the model's robustness to difficult-to-handle samples, thereby improving the recognition performance of ship name texts.

[0092] Table 1: Ablation experiments of the boat license plate text recognition algorithm based on representation fusion multimodal learning

[0093]

[0094] Experiments showed that this invention achieved a recognition accuracy of 92.74% on the SLPR-R dataset, as shown in Table 1. On the SLPR-P dataset, it achieved a recognition accuracy of 89.44%, as shown in Table 1 below. Therefore, the ship name text recognition model proposed in this invention, with its integrated architecture of autoregressive and non-autoregressive decoding branches, utilizes visual and linguistic modalities, interacts with them through a multi-attention mechanism, and leverages the complementarity between autoregressive and non-autoregressive decoding to further improve the recognition accuracy of ship name text images with distorted appearances.

Claims

1. A method for identifying inland waterway vessel names by combining autoregressive and non-autoregressive decoding, characterized in that, At the character-level representation level, autoregressive and non-autoregressive decoder branches are adaptively fused, and a confidence-based dynamic sample loss reweighting strategy is introduced, including two confidence-based hard example mining losses at the character level and instance level. The specific steps are as follows: Step 1: Ship Name Recognition Dataset Construction: Images containing ship names are captured using a camera, the ship name text lines are cropped, and the data is labeled using Labelme annotation software. The text content is used as the annotation information to create a recognition training set. Step 2, Visual coding model pre-training: The network is modeled based on masked images, and the feature extractor is pre-trained in a self-supervised learning manner; Step 3: Construct the ship name recognition model: Initialize the pre-trained feature extractor from Step 2. The extracted ship name image features are fed into the autoregressive decoder branch and the non-autoregressive decoder branch respectively, resulting in two types of character representations. The specific steps are as follows: Step 3-1: Input the ship's name and image from the input terminal. The ship name image features are obtained by outputting the feature extractor trained in step 2. , Represents the set of real numbers. Indicates the number of output image channels. Indicates the height and width of the output image; Step 3-2: Construct an autoregressive decoder; The Transformer decoder is used as the autoregressive decoder, which includes three key modules: a masked multi-head self-attention module, a multi-head cross-attention module, and a feedforward network module; the autoregressive decoder branches use right-shifted text annotation. and the ship name image features obtained in step 3-1 As input, the data is processed layer by layer through a masked multi-head self-attention module, a multi-head cross-attention module, and a feedforward network module to finally extract context-aware autoregressive character representation features. , The character representing time step T; The predicted probability of a character in time step t is expressed as: Predict the probability of time step t The input needs to be the output within the time interval from 1 to t-1. and ship name image features ,in The function is used to convert a set of real numbers into a probability distribution such that the output value is in the range [0,1] and the sum of all output values ​​is 1; Represents a linear classifier. The character representation feature at time step t; Step 3-3: Construct a non-autoregressive decoder; using a NAR decoder, first perform a multi-attention dynamic fusion mechanism to refine the decoder input features; The ship name image features obtained in step 3-1 Enhanced features are obtained through three independent components: channel attention, spatial attention, and self-attention. , , Then, the enhanced features are further enhanced through a gated full fusion mechanism. , , Dynamic fusion yields improved features In the middle; finally, attention is obtained by querying character order to improve features. Convert to non-autoregressive branch character representation features ; at the same time through Generate the complete target transcribed text, where The function is used to find the value of the independent variable when the function reaches its maximum value over a given domain; The characters predicted by the non-autoregressive decoder; Step 4: Adaptively fuse the representations of the autoregressive decoder branch and the non-autoregressive decoder branch through a gating mechanism; Step 5, Model Training: Using the character representations extracted from the autoregressive decoder branch and the non-autoregressive decoder branch, and the fused character representations as inputs, three linear classifiers are constructed to obtain three sets of predictions; The model is trained under the supervision of cross-entropy loss; Step 6: Implement hard sample mining. By fusing the losses of the autoregressive decoder branch and the non-autoregressive decoder branch, construct character-level and instance-level reweighted weights respectively, and dynamically adjust the contribution of batch samples to the total loss of the current batch during backpropagation.

2. The method for identifying inland waterway vessel names by combining autoregressive and non-autoregressive decoding according to claim 1, characterized in that, The specific steps of step 1 are as follows: Step 1-1: Take images with the ship and its name using a camera; adjust the focal length and shooting angle to capture images that cover a wider range of characteristics. Step 1-2: Perform coarse correction on the text line area based on the perspective transformation of the four corner coordinates of the ship name text line to obtain a small-scale ship name text line image. Steps 1-3: While maintaining the aspect ratio of the image, data augmentation techniques are used to process the image dataset, including but not limited to random rotation, random contrast adjustment, random scale adjustment, random resolution adjustment, and random blurring; the processed image is precisely placed on a gray template with a size of 100×32. Steps 1-4: The ship name text recognition model saves the cropped and expanded ship name text image as a .txt file with "filename annotation" as annotation information; This completes the creation of the training dataset for the ship name character recognition model.

3. The method for identifying inland waterway vessel names by combining autoregressive and non-autoregressive decoding according to claim 1, characterized in that, The specific steps of step 2 are as follows: Step 2-1: Select the VIT network and modify it. Delete the category labels used for classification in the VIT network and only keep the embedding vectors corresponding to the image patches as the extracted image features. Steps 2-3: By randomly selecting and masking a certain proportion of the ship name image, the network is prompted to learn to predict the features of the masked region from the remaining region. Mean squared error or contrast loss is used to measure the difference between the predicted features and the real features, ensuring that the network can effectively recover the masked region. Steps 2-4: Configure the training process; select Adam as the optimizer, set the initial learning rate, and dynamically adjust the learning rate according to the training progress; introduce... Regularization and Dropout techniques are used to reduce the risk of model overfitting and enhance the model's ability to generalize to new data. Steps 2-5: Periodically evaluate the model's performance during training; evaluate the model's performance by calculating the loss function value on the validation set and fine-tune it as needed.

4. The method for identifying inland waterway vessel names by combining autoregressive and non-autoregressive decoding according to claim 3, characterized in that, The specific steps of step 4 are as follows: Step 4-1: Adaptively fuse the character representations of the outputs from Steps 3-2 and 3-3 using a gating mechanism to obtain the fused features. The formula is as follows: ; in, Represents the learnable weights. The sigmoid function represents a non-linear activation function. This represents element-wise multiplication. yes and The adaptive weight matrix; Final output fusion features The decision is made using the following formula: ; in The character representing the prediction at time step t. This indicates the predicted character within time steps 1-t. This represents the adaptive weights at time step t. This represents the autoregressive character representation feature at time step t. The non-autoregressive character representation feature is used to represent the time step t.

5. The method for identifying inland waterway vessel names by combining autoregressive and non-autoregressive decoding according to claim 4, characterized in that, The specific steps of step 5 are as follows: Step 5-1: Represent the characters extracted from the autoregressive decoder branch and the non-autoregressive decoder branch in steps 3-3 and 3-4. and and the fusion features after fusion in step 4-1 For the input, construct three linear classifiers. Three sets of predictions were obtained; among them The main function is to divide the input space into different regions, each region mapping to a category, and learn the weights of features through training data to determine which features are more important for classification decisions; specifically, it takes the form of... , It is the input feature vector. It is a weight vector. It is a bias term.

6. The method for identifying inland waterway vessel names by combining autoregressive and non-autoregressive decoding according to claim 5, characterized in that, The specific steps of step 6 are as follows: Step 6-1: To refine the learnable weights To better handle the data information generated when learning from difficult data, two weights are designed, one for instance-level weights and the other for other instance-level weights. and character-level weights This allows for dynamic adjustment of the contribution of batch samples to the total loss of the current batch during backpropagation; Step 6-2: Design instance-level weights Instance-level weights It is generated by multiplying the confidence scores over T time steps, and then reweighting the recognition loss at the granularity of text images. The specific implementation formula is as follows: ,in The function compresses any real value to the range of 0 to 1, which is equivalent to performing an activation function operation; Step 6-3: Design character-level weights Character-level weighting , including character level This represents assigning different weights to the recognition loss at time step t for the i-th character, and the specific implementation formula is as follows: To balance learning from all data and learning from difficult data, a warm-up mechanism is employed, defining the reweighted working period as... ,here Indicates a predefined period; instance-level weights Or character-level weights Based on fusion loss Generated; Step 6-4: Combining steps 6-1 to 6-3, design the following loss function: ; The model is trained under the supervision of this cross-entropy loss, where Indicates the number of training batches. The function represents the cross-entropy loss function, which measures the difference between the probability distribution predicted by the model and the probability distribution of the true labels; the branch loss of the autoregressive decoder is obtained by solving the cross-entropy loss function. Non-autoregressive decoder branch loss and fusion loss The final overall optimization objective was determined as follows: α, β, and γ serve as balancing factors.