ViT-based port entering and leaving fishing boat name area extraction method, medium and system

Through ViT-based three-level pyramid feature extraction and multi-feature fusion, combined with the self-attention mechanism and MarineIDNet neural network, the problem of accurate extraction of fishing vessel names in complex marine environments is solved, and high-precision and robust recognition is achieved.

CN120635674AActive Publication Date: 2025-09-12BEIHAI FORECASTING CENT OF STATE OCEANIC ADMINISTRATION ((QINGDAO MARINE FORECASTING STATION OF STATE OCEANIC ADMINISTRATION) (QINGDAO MARINE ENVIRONMENT MONITORING CENT OF STATE OCEANIC ADMINISTRATION))
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510955791.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-12
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

Existing technologies make it difficult to extract the names of fishing vessels with high precision in complex marine environments, especially under different observation distances, lighting conditions and changes in hull posture, where the recognition accuracy is low.

Method used

A three-level pyramid feature extraction structure based on ViT is adopted, combined with the self-attention mechanism and multi-feature fusion function. The recognition difficulty is evaluated through the distortion matrix and the bias matrix. An adaptive region enhancement algorithm is designed, and the MarineIDNet neural network is used for precise positioning and verification.

Benefits of technology

The extraction accuracy and recognition robustness of the fishing vessel name area in complex marine environments have been improved, adapting to different observation distances and lighting conditions to ensure the accuracy of text recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635674A_ABST
    Figure CN120635674A_ABST
Patent Text Reader

Abstract

The invention provides a ViT-based port entering and leaving fishing boat name area extraction method, medium and system, belongs to the technical field of machine vision, and realizes progressive positioning from the whole boat body to the accurate position of the boat name through a three-stage pyramid feature extraction structure. The method comprises the following steps: firstly, determining a first-level region of interest including a ship outline by utilizing a self-attention mechanism, then extracting a second-level region of interest where a ship name number is located through position coding and multi-head attention operation, and finally, accurately positioning a ship name number region by applying a fine-grained feature extraction network. According to the method, an optical imaging degradation equation and a hull attitude light reflection model are introduced, a distortion matrix and a deviation matrix are constructed, a recognizable matrix is generated through a multi-feature fusion function, and targeted enhancement processing on a low-recognizable-degree area is guided. And finally, verifying an extraction result by adopting a Mar ineI DNet model of a double-branch structure to realize high-precision extraction of the ship name number area in the complex marine environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of machine vision technology, and in particular relates to a method, medium and system for extracting the name area of ​​fishing vessels entering and leaving a port based on ViT. Background Art

[0002] Fishing vessel names are key information for vessel identification and are crucial for fishery resource management, port security monitoring, and the regulation of illegal fishing. Traditional fishing vessel name recognition techniques rely primarily on convolutional neural network (CNN)-based object detection algorithms, such as Faster R-CNN and YOLOv4, which treat the vessel name area as a common target for detection and location. These methods achieve reasonable recognition results under ideal conditions and are widely used in scenarios such as port monitoring, maritime law enforcement, and fishery resource management.

[0003] However, traditional CNN-based techniques for extracting ship name and number regions face numerous challenges in real-world marine environments. First, fishing vessels' constantly changing postures at sea can cause significant perspective distortion in the ship name and number regions. Second, image resolution varies significantly at different observation distances, blurring details in the ship name and number regions when photographed from long distances. Furthermore, complex sea surface lighting conditions, weather factors, and the reflective properties of the ship's hull material further complicate recognition. These factors contribute to the generally low accuracy of traditional methods in practical applications, especially in adverse weather conditions.

[0004] Existing technologies struggle to accurately extract fishing vessel names in complex and changing environments. In particular, they lack the ability to adaptively process changes in observation distance, lighting conditions, and vessel posture. This results in low extraction accuracy for the vessel name region, limiting subsequent recognition accuracy and failing to meet the practical needs of maritime law enforcement and fisheries management. In other words, existing technologies suffer from low accuracy in automatically extracting fishing vessel names in complex marine environments. Summary of the Invention

[0005] In view of this, the present invention provides a ViT-based method, medium and system for extracting the name area of ​​fishing vessels entering and leaving the port, which can solve the technical problem in the existing technology of low accuracy in automatic extraction of the name area of ​​fishing vessels in complex marine environments.

[0006] The present invention is implemented as follows: In a first aspect, the present invention provides a method for extracting the name area of ​​fishing vessels entering and leaving a port based on ViT, comprising:

[0007] Based on the ViT (Vision Transformer model), a three-level pyramid feature extraction structure is designed for fishing vessel image analysis; the self-attention mechanism is used to determine the first-level area of ​​interest; the second-level area of ​​interest is extracted based on the first-level area of ​​interest; a fine-grained feature extraction network is applied to determine the third-level area of ​​interest; a distortion matrix is ​​constructed for the third-level area of ​​interest; the deviation matrix of the third-level area of ​​interest is calculated; a recognizable matrix is ​​generated based on the comprehensive calculation of the distortion matrix and the deviation matrix; an adaptive region enhancement algorithm is designed, and targeted image enhancement processing is performed on low-recognizable areas according to the guidance of the recognizable matrix; the final optimized ship name area image is output, and the coordinate information of the ship name area is provided for subsequent text recognition and verification systems.

[0008] Among them, the first-level area of ​​interest refers to the large-scale area containing the main structure of the hull determined after the initial analysis of the overall image of the fishing vessel by ViT. The first-level area of ​​interest covers the main outline of the fishing vessel and the surrounding area; the second-level area of ​​interest refers to the hull structure area determined after further narrowing the scope on the basis of the first-level area of ​​interest. The second-level area of ​​interest includes the key parts of the ship's side and the outer wall of the cab where the ship's name is marked; the third-level area of ​​interest refers to the precise location area of ​​the ship's name determined after fine processing. The third-level area of ​​interest only includes the ship's name text and its immediate background.

[0009] Among them, the distortion matrix is ​​a mathematical model that quantitatively represents the impact of different observation distances on image clarity. The distortion matrix is ​​constructed by analyzing the degree of attenuation of high-frequency information in the image and is used to evaluate the recognition difficulty caused by the distance factor; the deviation matrix is ​​a mathematical model that characterizes the degree of geometric deformation of the ship name area caused by changes in the ship's posture angle. The deviation matrix mainly considers the difficulties brought by perspective projection transformation to text recognition; the recognizable matrix is ​​an evaluation index generated by combining the distortion matrix and the deviation matrix. The recognizable matrix uses numerical values ​​to quantify the difficulty of recognizing the ship name area under current observation conditions.

[0010] This also includes applying the optical imaging degradation equation to simulate the change pattern of image quality of the ship name area at different observation distances. The input of the optical imaging degradation equation includes the observation distance obtained from the image metadata, the atmospheric visibility obtained from the meteorological database, the camera optical parameters obtained from the camera parameter file, the actual size of the ship name area measured from image analysis, and the imaging time obtained from the image metadata. The output of the optical imaging degradation equation is the distance-dependent point spread function and the signal-to-noise ratio attenuation matrix.

[0011] This also includes applying the hull posture light reflection model to analyze the impact of the light reflection characteristics of the hull surface on the recognition of the ship name. The input of the hull posture light reflection model includes the hull surface normal vector calculated from image analysis, the incident light angle obtained from the illumination estimation algorithm, the hull material reflection coefficient obtained from the material recognition network, the hull surface roughness obtained from texture analysis, and the coating aging degree obtained from the image degradation assessment. The output of the hull posture light reflection model is the effective contrast of the ship name area and the text edge sharpness evaluation matrix.

[0012] Among them, it also includes applying a multi-feature fusion function to comprehensively consider multiple factors to determine the recognizability of the ship name area. The input of the multi-feature fusion function includes the distortion matrix, the deviation matrix, the ambient brightness matrix obtained from image analysis, the hull texture complexity obtained from texture analysis, and the ship name font feature vector obtained from the font recognition network. The output of the multi-feature fusion function is the area recognizability score and optimization strategy recommendation.

[0013] The method also includes using the MarineIDNet neural network to verify the ship name area and fine-tune the boundaries. The specific structure of the MarineIDNet model is a dual-branch network architecture, which includes a feature extraction branch and a region verification branch. The feature extraction branch adopts an improved ViT structure, and the region verification branch adopts a four-layer convolutional neural network and two fully connected layers.

[0014] The steps for establishing a training dataset during the MarineIDNet model pre-training process specifically include collecting high-definition images of different types of fishing vessels under various lighting conditions, weather environments, and observation distances, using semi-automatic annotation tools to annotate the vessel name area with rectangular frames and transcribe the text content, and performing stratified sampling and balancing processing based on vessel size, observation distance, and weather conditions. Finally, a pre-training dataset containing 50,000 images is constructed.

[0015] A second aspect of the present invention provides a computer-readable storage medium having program instructions stored therein. When the program instructions are run in a computer, the method for extracting the name area of ​​fishing vessels entering and leaving the port based on ViT is executed.

[0016] The third aspect of the present invention provides a ViT-based system for extracting the name and number areas of fishing vessels entering and leaving the port, which includes the above-mentioned computer-readable storage medium. The system is any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is arranged in the system, and the system is provided with a microprocessor for executing the program instructions stored in the computer-readable storage medium.

[0017] This method combines a distortion matrix and a bias matrix to construct an identifiable matrix evaluation mechanism, achieving high-precision extraction of fishing vessel names and numbers. This method gradually pinpoints the location of vessel names and numbers through multiple levels of regions of interest, effectively overcoming the limitations of traditional CNN methods, which struggle to adapt to complex marine environments.

[0018] By incorporating optical imaging degradation equations and a ship attitude light reflection model, this method accurately quantifies the impact of varying observation distances and ship attitude changes on the recognition of ship name and number regions. Furthermore, adaptive region enhancement processing guided by an identifiable matrix significantly improves the quality of ship name and number region extraction under adverse conditions. In particular, the application of the MarineIDNet neural network, through a dual-branch network architecture combined with a local receptive field constraint mechanism, further ensures the precise location of the ship name and number region boundaries.

[0019] The present invention solves the technical problem of low automatic extraction accuracy of the fishing vessel name area in a complex marine environment, and improves the recognition robustness under different observation distances, lighting conditions and hull postures. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a flow chart of the method of the present invention. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0022] like Figure 1 FIG. 1 is a flowchart of a method for extracting the name region of fishing vessels entering and leaving a port based on ViT according to the first aspect of the present invention. The method comprises the following steps:

[0023] S01. Design a three-level pyramid feature extraction structure based on ViT for fishing boat image analysis and construct multiple feature representations at different scales;

[0024] S02. Use the self-attention mechanism to determine the first-level region of interest in the input image, focusing on the area in the image that has the hull contour features and occupies a large visual area;

[0025] S03. Based on the first-level ROI, the second-level ROI is extracted through position coding enhancement and multi-head attention operation to accurately locate the upper and lower structures of the hull and the area where the ship's name is located;

[0026] S04. Applying a fine-grained feature extraction network to determine a third-level region of interest, wherein the third-level region of interest precisely corresponds to the location of the ship's name, and achieving precise boundary positioning through Token segmentation technology;

[0027] S05. Constructing a distortion matrix for the third-level region of interest, using an optical imaging degradation equation to quantify changes in image clarity at different observation distances, and establishing a mapping function between distance and image quality degradation;

[0028] S06, calculating the deviation matrix of the third-level region of interest, applying the hull attitude light reflection model to analyze the degree of perspective deformation of the ship name area caused by the change of the ship attitude, and establishing a function between attitude angle and text deformation degree;

[0029] S07. Generate a recognizable matrix based on the comprehensive calculation of the distortion matrix and the deviation matrix, and quantitatively evaluate the difficulty coefficient of ship name area recognition under current conditions through a multi-feature fusion function;

[0030] S08. Designing an adaptive region enhancement algorithm to perform targeted image enhancement processing on low recognizability regions according to the recognizability matrix;

[0031] S09. Output the final optimized ship name area image and provide the ship name area coordinate information for subsequent text recognition and verification systems. Use the MarineIDNet neural network to verify the ship name area and fine-tune the boundaries.

[0032] Among them, the first-level area of ​​interest refers to the large-scale area containing the main structure of the hull determined after the ViT preliminary analysis of the overall image of the fishing vessel. The first-level area of ​​interest covers the main outline of the fishing vessel and the surrounding area.

[0033] Among them, the second-level area of ​​interest refers to the hull structure area determined after further narrowing the scope on the basis of the first-level area of ​​interest. The second-level area of ​​interest usually includes the key parts of the ship's side and the outer wall of the wheelhouse where the ship's name is marked.

[0034] The third-level area of ​​interest refers to the precise location area of ​​the ship's name that is finally determined after fine processing, and the third-level area of ​​interest only includes the ship's name and its adjacent background.

[0035] The distortion matrix is ​​a mathematical model that quantifies the effect of different observation distances on image clarity. The distortion matrix is ​​constructed by analyzing the attenuation degree of high-frequency information in the image and is used to evaluate the recognition difficulty caused by distance factors.

[0036] The deviation matrix is ​​a mathematical model that characterizes the degree of geometric deformation of the ship name area caused by changes in the ship's posture angle. The deviation matrix mainly considers the difficulties brought by perspective projection transformation to text recognition.

[0037] Among them, the identifiable matrix is ​​an evaluation index generated by combining the distortion matrix and the deviation matrix. The identifiable matrix uses numerical values ​​to quantify the difficulty of identifying the ship name area under current observation conditions, providing an optimization basis for subsequent processing.

[0038] Among them, the optical imaging degradation equation is used to simulate the change law of image quality of the ship name area at different observation distances. The input of the optical imaging degradation equation includes the observation distance obtained from the image metadata, the atmospheric visibility obtained from the meteorological database, the camera optical parameters obtained from the camera parameter file, the actual size of the ship name area measured from image analysis, and the imaging time obtained from the image metadata. The output of the optical imaging degradation equation is the distance-dependent point spread function and the signal-to-noise ratio attenuation matrix. The point spread function is used for subsequent image sharpening processing, and the signal-to-noise ratio attenuation matrix is ​​used to evaluate the difficulty of identifying different areas.

[0039] Among them, the hull posture light reflection model is used to analyze the influence of the light reflection characteristics of the hull surface on the recognition of the ship name and number. The input of the hull posture light reflection model includes the hull surface normal vector calculated from the image analysis, the incident light angle obtained from the illumination estimation algorithm, the hull material reflection coefficient obtained from the material recognition network, the hull surface roughness obtained from the texture analysis, and the coating aging degree obtained from the image degradation assessment. The output of the hull posture light reflection model is the effective contrast of the ship name and number area and the text edge sharpness evaluation matrix. The effective contrast of the ship name and number area is used for subsequent image enhancement processing, and the text edge sharpness evaluation matrix is ​​used to determine the edge enhancement strategy.

[0040] Among them, the multi-feature fusion function is used to comprehensively consider multiple factors to determine the recognizability of the ship name area. The input of the multi-feature fusion function includes the distortion matrix generated in the above steps, the deviation matrix calculated in the above steps, the ambient brightness matrix obtained from image analysis, the hull texture complexity obtained from texture analysis, and the ship name font feature vector obtained from the font recognition network. The output of the multi-feature fusion function is the regional recognizability score and optimization strategy recommendation. The regional recognizability score is used to determine whether enhancement processing is required, and the optimization strategy recommendation is used to guide the selection of the enhancement algorithm.

[0041] The specific structure of the MarineIDNet model is a dual-branch network architecture, which includes a feature extraction branch and a region verification branch. The feature extraction branch adopts an improved ViT structure, divides the input image into a sequence of image blocks of 16×16 size as the basic processing unit, has an embedding dimension of 768, includes 12 Transformer encoder layers, and the number of self-attention heads in each layer is 12. The hidden layer dimension of the feedforward neural network is 3072. At the same time, a local receptive field constraint mechanism is introduced to enhance the sensitivity to small-sized text areas, so that the attention mechanism pays more attention to the edge detail features of the ship name area; the region verification branch is composed of a four-layer convolutional neural network and two fully connected layers. The convolution kernel sizes are 7×7, 5×5, 3×3, and 3×3, respectively, and the number of channels are 64, 128, 256, and 256, respectively, to verify whether the candidate area contains valid ship name information.

[0042] The steps of establishing a training dataset in the pre-training process of the MarineIDNet model specifically include collecting high-definition images of different types of fishing vessels under various lighting conditions, weather environments and observation distances, using semi-automatic annotation tools to annotate the ship name area with rectangular frames and transcribe the text content, and performing stratified sampling and balanced processing based on ship size, observation distance, and weather conditions. Finally, a pre-training dataset containing 50,000 images is constructed, of which 80% is used for training, 10% is used for verification, and 10% is used for testing. At the same time, data enhancement technology is used to expand data diversity, and key sampling is performed on difficult-to-identify examples in complex environments to improve the robustness of the model under harsh conditions.

[0043] The pre-training steps of the MarineIDNet model specifically include first learning basic visual features on the ImageNet dataset, then performing transfer learning on a general marine vessel dataset to obtain the ability to represent hull structure-related features, and finally fine-tuning on a fishing vessel name recognition dataset. The AdamW optimizer is used for training, the initial learning rate is set to 0.0001, the weight decay is 0.05, the cosine annealing learning rate scheduling strategy is adopted, the batch size is 32, the number of training rounds is 200, and an early stopping mechanism is introduced during training to avoid overfitting. In the key parameter adaptive adjustment strategy, the number of attention mechanism heads will be automatically adjusted according to the size of the ship name area in the input image. When the area size is less than 1 / 20 of the input image, the number of heads is increased to 16 to improve the ability to perceive details. At the same time, the attention weight attenuation coefficient is dynamically adjusted according to the image contrast, and a smaller attenuation coefficient is applied to low-contrast images to enhance the ability to extract weak features.

[0044] The specific implementation of the above steps is described in detail below.

[0045] Step S01 is implemented by constructing a three-level pyramid feature extraction structure based on ViT. This structure employs a multi-scale feature representation strategy, capturing hierarchical information in the fishing boat image through feature maps at different resolutions. First, the input image is resized to 224×224 pixels and segmented into 16×16 image blocks, generating a total of 196 basic image blocks. Each image block is then linearly projected into a 768-dimensional feature space, and positional encoding is added to preserve spatial location information. These feature representations are then processed through three Transformer encoders of varying depths: the shallow encoder contains four attention layers to capture local texture features, the middle encoder contains eight attention layers to learn structural features of moderate complexity, and the deep encoder contains 12 attention layers to understand global semantic information. Finally, the feature representations from these three different levels are fused through a feature pyramid network structure to generate a multi-scale feature map. This step aims to construct a comprehensive representation of the fishing boat image, providing rich visual feature information for subsequent region-of-interest localization. In particular, the multi-scale representation improves adaptability to regions with different boat names and sizes.

[0046] The specific implementation of step S02 is to use the self-attention mechanism to determine the first-level regions of interest. This step first calculates the attention weight matrix for each token in the input image and normalizes the weight distribution using the Softmax function. It then analyzes the attention heatmap and identifies regions with weight values ​​above a threshold of 0.65 as candidate regions of interest. A connected region analysis algorithm is then applied to merge adjacent high-attention regions into a single region. The visual area percentage of each region is then calculated, retaining regions with an area percentage exceeding 15% of the total image area. Finally, a geometric shape analysis algorithm is applied to calculate features such as the perimeter-to-area ratio and contour complexity of the region, screening for regions with hull characteristics. These regions have a perimeter-to-area ratio between 0.08 and 0.25 and a contour complexity index below 1.8. This step aims to quickly locate the primary region containing the hull within the global image, reducing the search space for subsequent processing and improving computational efficiency. This eliminates background interference and ensures that attention is focused on the hull structure.

[0047] The specific implementation of step S03 involves determining a second-level ROI based on the first-level ROI. This step first applies enhanced position encoding to the first-level ROI, combining sine-cosine position encoding with learnable position encoding to enhance spatial position perception. A multi-head attention mechanism then performs fine-grained analysis of the region, using eight attention heads, each with an attention matrix dimension of 64, to capture features from different semantic subspaces. An attention response map is then calculated and threshold segmentation is applied, extracting regions with an attention response value greater than 0.78 as candidate regions. Prior knowledge constraints are then introduced, limiting the search to the forward area of ​​the hull and the sidewalls based on statistical information from the ship type database. Finally, morphological operations are applied to optimize the region boundaries, using a combination of opening and closing operations to eliminate noise and smooth edges. The structural element size is set to 3×3 pixels. The purpose of this step is to further narrow the recognition range and accurately locate key areas on the hull that may contain the ship's name, such as the sidewalls and the bridge surface, laying the foundation for subsequent fine-grained recognition.

[0048] Step S04 specifically involves applying a fine-grained feature extraction network to determine the third-level region of interest. This step first applies a high-resolution feature extraction module to the second-level region of interest, segmenting the region into smaller 4×4 pixel blocks for detailed analysis. Feature enhancement is then performed using a fine-grained attention network, consisting of four attention heads, each focusing on character-level visual features. A text region detection algorithm is then applied to identify potential text regions based on the stroke density map and text direction field, with a text region probability threshold set to 0.82. Token segmentation is then used to precisely locate the boundaries of the text regions. An adaptive thresholding method is employed to separate foreground text from background. The threshold calculation is dynamically adjusted based on local region contrast, with a contrast weighting factor of 0.7. Finally, a boundary optimization algorithm is applied to fine-tune the detected text regions, including minimum bounding rectangle calculation, boundary extension and reservation, and boundary regularization. The boundary extension ratio is set to 10% of the text region width. The goal of this step is to precisely locate the specific region containing the ship's name, extract the ship's name region with pixel-level accuracy, and ensure that subsequent processing focuses on the minimum necessary area containing the ship's name.

[0049] The specific implementation of step S05 involves constructing a distortion matrix for the third-level region of interest. This step first extracts key parameters, such as observation distance from image metadata, atmospheric visibility from a meteorological database, optical parameters from a camera parameter file, the actual size of the ship name area measured through image analysis, and imaging time from metadata. A mathematical model is then established based on the optical imaging degradation equation, which accounts for the combined effects of atmospheric scattering, the optical system's diffraction limit, and focus shift on image quality. The point spread function is then calculated at different observation distances to quantify the degree of image blur. When the observation distance exceeds 50 meters, the increase in the point spread function radius should not be less than 0.5 pixels per 10 meters. A signal-to-noise ratio attenuation matrix is ​​then generated, analyzing the signal intensity to noise level ratio of each subregion. Areas with a value below 6 decibels are marked as high-distortion risk areas. Finally, a normalization process is performed to generate a standardized distortion matrix, with values ​​ranging from 0 to 1, where 0 represents no distortion and 1 represents complete distortion. The purpose of this step is to quantitatively assess the impact of observation distance on the image quality of the ship name area, providing a scientific basis for subsequent image enhancement processing and ensuring targeted improvement of low-quality areas.

[0050] The specific implementation of step S06 is to calculate the third-level region of interest deviation matrix. This step first applies the hull posture light reflection model to analyze the interaction between light and the hull surface. The hull surface normal vector is extracted from the image, the incident light angle is obtained through the illumination estimation algorithm, the hull material reflectance coefficient is obtained using the material recognition network, the surface roughness is obtained through texture analysis, and the coating aging degree is obtained based on image degradation assessment. Then, a model is established to determine the relationship between hull posture and perspective deformation. The tilt angle of the hull surface relative to the camera is calculated. When the tilt angle exceeds 60 degrees, the text deformation coefficient should be greater than 1.5. The degree of perspective deformation of the text area under different posture angles is then analyzed and a quantitative deformation index is established. For every 10-degree increase in the perspective deformation angle, the text recognition difficulty coefficient increases by 0.2. Then, an effective contrast matrix is ​​generated for the ship name area to evaluate the changes in the contrast between the text and the background under different lighting conditions. Areas with an effective contrast ratio below 0.3 are marked as high-risk areas. Finally, a text edge sharpness evaluation matrix is ​​calculated to analyze the clarity of the text outline. Areas with edge sharpness below 0.4 require special enhancement processing. The purpose of this step is to quantitatively evaluate the degree of geometric deformation of the ship name area caused by the change of the hull posture, and provide accurate parameters for subsequent perspective correction and image enhancement.

[0051] The specific implementation of step S07 is to generate a recognition matrix based on the comprehensive calculation of the distortion matrix and the deviation matrix. This step first applies a multi-feature fusion function to integrate various influencing factors, including the distortion matrix generated in the previous step, the calculated deviation matrix, the ambient brightness matrix obtained from image analysis, the hull texture complexity obtained from texture analysis, and the ship name font feature vector obtained from the font recognition network. Then, a comprehensive score is calculated by weighted summation, with the weights of each factor being 0.35 for the distortion factor, 0.3 for the deviation factor, 0.15 for the brightness factor, 0.1 for the texture complexity, and 0.1 for the font feature. Then, a nonlinear mapping function is applied to convert the original score into a standardized recognition difficulty coefficient, using a Sigmoid function for mapping, with the output ranging from 0 to 1, where 0 indicates easy recognition and 1 indicates extremely difficult recognition. Then, a recognition heat map is generated based on the recognition difficulty coefficient, which intuitively displays the recognition difficulty distribution of each part of the ship name area. Areas with a recognition difficulty coefficient higher than 0.7 are marked as red high-risk areas. Finally, optimization strategy recommendations are determined based on the recognition score, including recommended image processing methods and parameter settings. When the overall recognition is lower than 0.5, adaptive enhancement processing is recommended. The purpose of this step is to comprehensively evaluate the recognition difficulty of the ship name area under current conditions and provide a decision basis for subsequent targeted enhancement processing.

[0052] The specific implementation of step S08 is to design an adaptive region enhancement algorithm. This step first identifies regions requiring enhancement based on the recognizability matrix, with regions with a recognizability coefficient greater than 0.6 being the focus. Differentiated enhancement strategies are then designed for different problems. For distortion-dominated problems, a deconvolution recovery algorithm is applied with a convolution kernel size of 5×5 and a regularization parameter of 0.02. For deviation-dominated problems, a perspective correction algorithm is applied, performing a perspective transformation based on four-point correspondences, with control points located at the four corners of the text. For insufficient contrast, local contrast enhancement is applied, with the enhancement coefficient adaptively adjusted based on local brightness, with the maximum gain not exceeding 200% of the original contrast. Edge enhancement is then applied to improve text edge clarity, using a nonlinear gradient enhancement method with the enhancement strength inversely proportional to the edge sharpness matrix. Finally, a noise suppression filter is applied, using a guided filter algorithm to protect text edges while suppressing noise. The filter radius is set to 3 pixels and the regularization coefficient is 0.1. The purpose of this step is to perform targeted enhancement of low-recognition regions based on the aforementioned analysis results, improving the overall recognizability of the ship name region and providing high-quality input for the subsequent text recognition system.

[0053] The specific implementation of step S09 involves outputting the optimized image and coordinate information of the ship name area. This step first integrates the aforementioned processing results to generate a high-quality image of the ship name area. It then calculates the precise coordinate information of the ship name area, including the coordinates of the rectangular bounding box, rotation angle, and confidence score. The processed image and coordinate information of the ship name area are then transmitted to the verification system. The ship name area is then verified using the MarineIDNet neural network, which employs a dual-branch structure consisting of feature extraction and region verification branches. The region validity is verified through multi-layer feature extraction and comparative analysis. Finally, the boundaries of the verified areas are fine-tuned using a precise text edge positioning algorithm, with the adjustment accuracy required to be within 2 pixels of the original boundary. This step aims to provide the final optimized image and precise location information of the ship name area, providing standardized input for the subsequent text recognition system. Furthermore, the verification process ensures the accuracy and reliability of the extracted results.

[0054] Furthermore, the detailed structure of the MarineIDNet model is a dual-branch network architecture, consisting of a feature extraction branch and a region verification branch. The feature extraction branch uses an improved ViT structure, which divides the input image into a sequence of 16×16 image blocks. Each image block is converted into a 768-dimensional feature vector through a linear projection layer, and a learnable positional encoding is added. This branch contains 12 Transformer encoder layers, each consisting of a multi-head self-attention module and a feedforward neural network. The multi-head self-attention module contains 12 attention heads, each with a query, key, and value matrix of 64 dimensions. The feedforward neural network consists of two fully connected layers with a hidden layer dimension of 3072. This branch also introduces a local receptive field constraint mechanism, which improves sensitivity to small-sized text by limiting the scope of attention calculation to a local area (with a radius of 3 tokens). Finally, the features are converted into feature vectors by the classification head for downstream tasks. The area verification branch is composed of a four-layer convolutional neural network and two fully connected layers. The convolution kernel size of the first layer is 7×7, and the number of output channels is 64. The convolution kernel size of the second layer is 5×5, and the number of output channels is 128. The convolution kernel size of the third and fourth layers is 3×3, and the number of output channels is 256. Each convolution layer is followed by batch normalization and ReLU activation function. The second and fourth layers are followed by a maximum pooling layer with a pooling kernel size of 2×2. The output dimensions of the fully connected layers are 512 and 2, respectively. Finally, the Softmax function is used to output the probability that the area contains a valid ship name.

[0055] Furthermore, the steps for establishing the pre-training dataset of the MarineIDNet model specifically include the following processes: first, collect high-definition images of different types of fishing boats under various environmental conditions, including samples under various lighting conditions (sunny, cloudy, dusk, night, etc.), weather environments (sunny, cloudy, rainy, foggy, etc.) and different observation distances (close distance within 5 meters, medium distance 5-30 meters, long distance above 30 meters); then use a semi-automatic annotation tool to mark the ship name area in each image with a rectangular frame, and record the text content as the true value label. The annotation accuracy is required to be controlled at the pixel level; then, according to the size of the ship (small ships are less than 12 meters long, medium-sized ships are 12-24 meters long, and large ships are greater than 24 meters long), observation distance, weather conditions, etc. Stratified sampling is performed to ensure a balanced number of samples under various conditions; then a pre-training dataset of 50,000 images is constructed, of which 80% (40,000) are used for training, 10% (5,000) are used for validation, and 10% (5,000) are used for testing; finally, data diversity is expanded through data augmentation techniques, including random rotation (±15 degrees), scaling (0.8 to 1.2 times), brightness adjustment (±20%), contrast change (±15%), noise addition (Gaussian noise standard deviation 0.01 to 0.05), and simulation of severe weather conditions (rain, fog, uneven lighting, etc.). Focus on sampling and enhancement of difficult-to-recognize examples in complex environments, improve the coverage of samples in edge cases, and thus enhance the robustness of the model under harsh conditions.

[0056] Furthermore, the pre-training process of the model specifically includes first learning basic visual features on the ImageNet dataset, then performing transfer learning on the general marine vessel dataset to obtain the ability to represent hull structure-related features, and finally fine-tuning on the fishing vessel name recognition dataset. The AdamW optimizer is used for training, the initial learning rate is set to 0.0001, the weight decay is 0.05, the cosine annealing learning rate scheduling strategy is adopted, the batch size is 32, the number of training rounds is 200, and the early stopping mechanism is introduced during the training process to avoid overfitting. In the key parameter adaptive adjustment strategy, the number of attention mechanism heads will be automatically adjusted according to the size of the ship name area in the input image. When the area size is less than 1 / 20 of the input image, the number of heads will be increased to 16 to improve the ability to perceive details. At the same time, the attention weight attenuation coefficient is dynamically adjusted according to the image contrast. A smaller attenuation coefficient is applied to low-contrast images to enhance the ability to extract weak features.

[0057] A second aspect of the present invention provides a computer-readable storage medium having program instructions stored therein. When the program instructions are run in a computer, the program instructions are used to execute the above-mentioned ViT-based method for extracting the name area of ​​fishing vessels entering and leaving the port.

[0058] The third aspect of the present invention provides a ViT-based system for extracting the name and number areas of fishing vessels entering and leaving the port, which includes the above-mentioned computer-readable storage medium. The system is any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is arranged in the system, and the system is provided with a microprocessor for executing the program instructions stored in the computer-readable storage medium.

[0059] The mathematical model or calculation process involved in the present invention is described in detail below.

[0060] The three-level pyramid feature extraction structure in step S01 involves multiple calculation processes, which are specifically represented as follows:

[0061] The process of segmenting the input image into image blocks can be expressed as:

[0062] P={p i , j |i=1,2,...,N h ; j = 1, 2, ..., N w};

[0063] Where P is the set of image blocks; p i,j represents the image block at position (i, j); N h and N w Respectively represent the number of blocks in the height and width directions of the image. For a 224×224 image, it is divided into 16×16 image blocks. N h =N w =14.

[0064] The linear projection transformation process of the image block is expressed as:

[0065] z0=[x class ;p1E;p2E;...;p N E]+E pos ;

[0066] Where z0 represents the initial feature representation of the image block sequence; x class is a special mark used for classification; p n represents the nth image block; E is the linear projection matrix, which maps the image block to the D-dimensional feature space, D = 768; E pos is the position encoding matrix, with the same dimension as z0. E is obtained through pre-training, and E pos Calculated by the following formula:

[0067] E pos (pos, 2i) = sin(pos / 10000 2i / D );

[0068] E pos(pos, 2i+1) = cos(pos / 10000) 2i / D );

[0069] Where pos represents the position in the sequence; i represents the feature dimension index, ranging from 0≤i <D / 2。

[0070] The multi-head self-attention calculation process is expressed as:

[0071]

[0072] Where Q, K, and V represent query, key, and value matrices, respectively; k Indicates the dimension of the key vector, used to scale the dot product result to prevent the gradient from disappearing, d k =D / h=64, where h=12 is the number of attention heads.

[0073] The output feature fusion process of encoders at different levels can be expressed as:

[0074] F multi =α1F shallow +α2F middle +α3F deep ;

[0075] Where, F multi Represents the fused multi-scale features; F shallow 、F middle 、F deep They represent the output features of the shallow, middle and deep encoders respectively; α1, α2 and α3 are fusion weight coefficients obtained by back-propagation optimization, and the initial values ​​are generally set to α1 = 0.2, α2 = 0.3 and α3 = 0.5.

[0076] The self-attention mechanism in step S02 determines the first-level interest area, which involves the following calculation process:

[0077] The calculation of the attention heat map is expressed as:

[0078]

[0079] Where A map represents the average attention heat map; A i represents the attention matrix of the i-th attention head; h represents the number of attention heads, h = 12.

[0080] The recognition process of high attention areas is expressed as:

[0081] M roi1 ={(x, y)|A map (x, y) > τ1};

[0082] Where Mroi1 represents the first-level interest region mask; (x, y) represents the pixel coordinates in the image; τ1 represents the attention threshold, τ1 = 0.65.

[0083] The regional area ratio is calculated as:

[0084]

[0085] Where R area Indicates the area ratio of the region; Area(M roi1 ) represents the area of ​​interest; Area(I total ) represents the total area of ​​the image. area >0.15 area.

[0086] The contour complexity calculation is expressed as:

[0087]

[0088] Where C complexity represents the contour complexity index; P represents the region perimeter; A represents the region area. And C complexity <1.8 area.

[0089] Determining the second-level ROI in step S03 involves the following calculation process:

[0090] The enhanced position encoding calculation is expressed as:

[0091]

[0092] Where, represents enhanced position encoding; Represents sine-cosine position encoding; represents the learnable position encoding; λ is the fusion coefficient, λ = 0.6.

[0093] The multi-head attention response calculation is expressed as:

[0094]

[0095] Where A response represents the attention response map; w i Represents the weight of the i-th attention head; Attention i represents the output of the i-th attention head. i Obtained through training, the initial value is set to equal weight

[0096] The second level area of ​​interest is determined as:

[0097] Mroi2 ={(x, y)|A response (x, y)>τ2}∩R prior ;

[0098] Where M roi2 represents the second-level interest region mask; τ2 represents the attention response threshold, τ2 = 0.78; R prior Represents the constraint area based on prior knowledge and determined according to the ship type database.

[0099] Determining the third-level ROI in step S04 involves the following calculation process:

[0100] Fine-grained feature extraction is expressed as:

[0101]

[0102] Where, F fine Represents fine-grained features; Conv represents a convolution operation; Resize represents a resampling operation, which reduces the second-level ROI to a quarter of its original size to obtain a finer feature representation.

[0103] The text area detection calculation is expressed as:

[0104] P text =σ(W2·ReLU(W1·F fine +b1)+b2);

[0105] Where R text Represents the text area probability map; W1, W2, b1, b2 are model parameters; σ represents the Sigmoid activation function; ReLU represents the rectified linear unit activation function.

[0106] The adaptive threshold calculation is expressed as:

[0107]

[0108] Where, T adaptive represents the adaptive threshold; μ(x, y) represents the local area mean; σ(x, y) represents the local area standard deviation; γ represents the contrast weight coefficient, γ = 0.7.

[0109] The third level area of ​​interest is determined as:

[0110] M roi3 ={(x, y)|P text (x, y)>τ3};

[0111] Where M roi3 represents the third-level interest area mask; τ3 represents the text area probability threshold, τ3 = 0.82.

[0112] The boundary extension process is expressed as:

[0113] B expanded =Expand(B original , r expand );

[0114] Where B expanded Indicates the expanded boundary; B original represents the original boundary; r expand Indicates the expansion ratio, r expand =0.1, indicating that the expanded width is 10% of the original text area width.

[0115] Constructing the distortion matrix in step S05 involves the following calculation process:

[0116] The optical imaging degradation equation is expressed as:

[0117] g(x, y)=h(x, y)*f(x, y)+n(x, y);

[0118] Where g(x, y) represents the observed image; h(x, y) represents the point spread function; f(x, y) represents the ideal undistorted image; n(x, y) represents additive noise; * represents the convolution operation.

[0119] The point spread function calculation is expressed as:

[0120] h(x, y) = h system (x, y)*h atmospheric (x, y)*h defocus (x, y);

[0121] Where h system represents the point spread function of the optical system; h atmospheric represents the atmospheric scattering point spread function; h defocus represents the out-of-focus spread function.

[0122] The point spread function of the optical system is expressed as:

[0123]

[0124] Where J1 represents the first-order Bessel function; r represents the radial distance; r0 represents the radius of the Airy disk. Where λ is the wavelength of light, f is the focal length, and D is the aperture diameter.

[0125] The atmospheric scattering point spread function is expressed as:

[0126]

[0127] Where r represents the radial distance; r0 represents the atmospheric coherence length, Where L represents the observation distance, C n Represents the atmospheric refractive index structure constant, which is related to meteorological conditions.

[0128] The out-of-focus spread function is expressed as:

[0129] $h_{defocus}(r)=

[0130] $;

[0131] Where R defocus represents the defocus radius, Where s represents the actual object distance.

[0132] The distance-dependent point spread function radius increase is calculated as:

[0133]

[0134] Where, ΔR PSF represents the increase in the radius of the point spread function; d represents the observation distance (meters); d0 represents the reference distance, d0 = 50 meters; α PSF represents the radius increase coefficient, α PSF =0.5 pixels / 10 meters.

[0135] The signal-to-noise ratio attenuation matrix calculation is expressed as:

[0136]

[0137] Where SNR(x, y) represents the signal-to-noise ratio (dB) at position (x, y); S(x, y) represents the signal strength;

[0138] N(x, y) represents the noise intensity. Areas with a signal-to-noise ratio below 6 dB are marked as high-distortion risk areas.

[0139] The normalized distortion matrix calculation is expressed as:

[0140]

[0141] Where D norm represents the normalized distortion matrix; SNR threshold Indicates the signal-to-noise ratio threshold, SNR threshold =20dB. D norm The value range is 0 to 1, where 0 means no distortion and 1 means full distortion.

[0142] The calculation of the deviation matrix in step S06 involves the following calculation process:

[0143] The calculation of the hull surface normal vector is expressed as:

[0144]

[0145] Where, represents the surface normal vector; Represents image gradient; represents the gradient modulus.

[0146] The incident light angle estimation is expressed as:

[0147]

[0148] Where, Represents the normalized incident light direction vector; S represents the pixel set in the highlight area; Represents the half vector of the i-th pixel; Represents the surface normal vector of the i-th pixel.

[0149] The relationship model between the hull posture and perspective deformation is expressed as:

[0150] D perspective (θ)=1+β·max(0, θ-θ0);

[0151] Where D perspective Indicates the text deformation coefficient; θ indicates the angle between the hull surface and the camera plane (degrees); θ0 indicates the critical angle, θ0 = 30 degrees; β indicates the deformation coefficient, β = 0.02 / degree, that is, for every 1 degree increase in the angle, the deformation coefficient increases by 0.02. When θ exceeds 60 degrees, D perspective Greater than 1+0.02·(60-30)=1.6, meeting the requirements.

[0152] The effect of perspective deformation angle on recognition difficulty is expressed as:

[0153]

[0154] Where D difficulty Indicates the recognition difficulty coefficient; θ indicates the perspective deformation angle (degrees). The recognition difficulty coefficient increases with each 10-degree increase in the perspective deformation angle.

[0155] The effective contrast ratio is calculated as:

[0156]

[0157] Where C effective Indicates effective contrast; I text Indicates the brightness of the text area; I background Indicates the brightness of the background area; I max and I min Represent the maximum and minimum brightness of the image respectively; γ represents the perspective influence coefficient, γ=0.3; D perspective(θ) represents the text deformation coefficient. Areas with an effective contrast lower than 0.3 are marked as high-risk areas.

[0158] The text edge sharpness calculation is expressed as:

[0159]

[0160] Where S edge Indicates edge sharpness; Represents the gradient magnitude at position (x, y); represents the maximum gradient amplitude; δ represents the perspective influence coefficient, δ = 0.4; D perspective (θ) represents the text deformation coefficient. Areas with edge sharpness lower than 0.4 require special enhancement processing.

[0161] The deviation matrix calculation is expressed as:

[0162] B(x, y) = 1-min(C effective (x, y), S edge (x, y));

[0163] Where B(x, y) represents the deviation matrix; C effective Indicates effective contrast; S edge Indicates edge sharpness. The value of B(x, y) ranges from 0 to 1, where 0 indicates no deviation and 1 indicates maximum deviation.

[0164] The comprehensive calculation to generate the identifiable matrix in step S07 involves the following calculation process:

[0165] The multi-feature fusion function is expressed as:

[0166] R(x, y) = w1·D norm (x, y)+w2·B(x,y)+w3·L(x,y)+w4·T(x,y)+w5·F(x,y);

[0167] Where R(x, y) represents the original comprehensive score; D norm represents the normalized distortion matrix; B represents the bias matrix; L represents the ambient brightness matrix; T represents the complexity of the hull texture; F represents the font feature vector of the ship name; w1, w2, w3, w4, and w5 represent the weights of each factor, w1 = 0.35, w2 = 0.3, w3 = 0.15, w4 = 0.1, and w5 = 0.1, respectively.

[0168] The ambient brightness matrix calculation is expressed as:

[0169]

[0170] Where L(x, y) represents the ambient brightness deviation; I(x, y) represents the brightness value at position (x, y); I max Indicates the maximum brightness value. The closer the L(x, y) value is to 0, the moderate brightness is achieved, while the closer it is to 0.5, the higher or lower the brightness is achieved.

[0171] The texture complexity calculation is expressed as:

[0172]

[0173] Where T(x, y) represents the texture complexity; N represents the number of pixels in the local area; Indicates position (x i ,y i ) is the gradient magnitude at .

[0174] The calculation of the standardized recognition difficulty coefficient is expressed as:

[0175]

[0176] Where R norm R represents the standardized recognition difficulty coefficient; R represents the original comprehensive score; k represents the Sigmoid function steepness parameter, k = 5; R0 represents the center point parameter, R0 = 0.5. norm The value range is 0 to 1, where 0 is easy to identify and 1 is extremely difficult to identify.

[0177] The identifiable matrix generation is expressed as:

[0178] $RM(x,y)=

[0179] $;

[0180] Where RM represents the identifiable matrix; R norm represents the standardized recognition difficulty coefficient; τ4 represents the recognition difficulty threshold, τ4 = 0.7. Areas with a recognition difficulty coefficient higher than 0.7 are marked as high-risk areas.

[0181] Designing the adaptive region enhancement algorithm in step S08 involves the following calculation process:

[0182] The deconvolution restoration algorithm is expressed as:

[0183]

[0184] Where, represents the restored image; h represents the point spread function; g represents the observed image; D represents the difference operator; θ represents the regularization parameter, θ = 0.02.

[0185] The perspective correction algorithm is expressed as:

[0186]

[0187] Where H represents the perspective transformation matrix; A1 represents the coordinate matrix of the four corner points of the source quadrilateral; A2 represents the coordinate matrix of the four corner points of the target rectangle. The transformed image is calculated using the following formula:

[0188]

[0189] Where, I corrected Represents the corrected image; I original represents the original image; h ij Represents the elements of the perspective transformation matrix H.

[0190] The local contrast enhancement algorithm is expressed as:

[0191] I enhanced (x,y)=I(x,y)+α(x,y)·(I(x,y)-I blur (x, y));

[0192] Where, I enhanced represents the enhanced image; I represents the original image; I blur represents the image after Gaussian blur; α represents the enhancement coefficient, Make sure the maximum gain does not exceed 200% of the original contrast.

[0193] The edge enhancement process is expressed as:

[0194] I edge (x,y)=I(x,y)+β(x,y)·(I(x,y)-I blur (x, y) 2 ·sign(I(x,y)-I blur (x, y));

[0195] Where, I edge represents the edge enhanced image; I represents the original image; I blur represents the image after Gaussian blur; β represents the nonlinear enhancement coefficient, Where β0 represents the basic enhancement coefficient, β0=10, ∈ represents a small constant to prevent division by zero, ∈=0.01.

[0196] The guided filtering algorithm is expressed as:

[0197] I filtered (x, y)=a(x, y)·G(x, y)+b(x, y);

[0198] Where, I filteredrepresents the filtered image; G represents the guide image, which is generally the original image; a and b are the coefficients of the local linear model, which is solved by minimizing the cost function:

[0199]

[0200] Where, ω (x,y) represents a window centered at (x, y); ∈ represents the regularization coefficient, ∈ = 0.1. The filter radius is set to 3 pixels.

[0201] The principles and significance of the equations involved in the above formula are explained as follows: The optical imaging degradation equation simulates the physical phenomena in the real imaging process, including the effects of factors such as the diffraction limit of the optical system, atmospheric scattering, and defocus on image quality. The ideal image is converted to the actual observed image using the point spread function, while additive noise simulates sensor noise and quantization error. This equation, which incorporates the principles of physical optics, accurately describes how image quality varies at different observation distances.

[0202] The calculation of the point spread function (PSF) takes into account three key factors: the diffraction limit of the optical system, atmospheric scattering, and defocus effects. The optical system PSF uses the Airy disk model, based on wave optics theory; the atmospheric scattering PSF considers the effects of atmospheric turbulence on light wave propagation, based on Kolmogorov turbulence theory; and the defocus point spread function is based on the principles of geometric optics. This comprehensive consideration of these three factors enables the PSF to comprehensively reflect changes in image quality under varying observation conditions.

[0203] The ship posture and perspective deformation relationship model establishes a quantitative relationship between the tilt angle of the ship surface and the degree of text deformation. A linear growth function is used to describe the effect of increasing angle on deformation. The critical angle is set to take into account the fact that small angle changes have little impact on recognition in practical applications. This model simplifies complex perspective transformation calculations and provides a method for quickly evaluating the impact of perspective deformation.

[0204] The effective contrast calculation comprehensively considers the brightness difference between the text and the background, as well as the effects of perspective distortion. Normalization is used to keep the contrast value between 0 and 1 for ease of subsequent processing. The introduction of the perspective influence factor accounts for the fact that perspective distortion reduces text edge clarity, thereby reducing effective contrast. This design aligns with human visual perception.

[0205] The multi-feature fusion function comprehensively considers various factors affecting the difficulty of ship name recognition through a weighted summation method. The weight distribution reflects the varying degrees of influence of each factor on the recognition difficulty. The distortion matrix and the bias matrix have higher weights, indicating that image quality and geometric deformation are the primary factors affecting recognition. This multi-feature fusion method makes the recognition difficulty assessment more comprehensive and accurate.

[0206] The standardized recognition difficulty coefficient uses a Sigmoid function for nonlinear mapping, converting the raw composite score into a recognition difficulty coefficient ranging from 0 to 1. The Sigmoid function was chosen for its smoothness and sensitivity to intermediate values, making it suitable for representing continuously varying difficulty levels. The parameter k controls the function's steepness, while R0 controls the center point's location. These two parameters adjust the sensitivity of the mapping.

[0207] Adaptive enhancement algorithms select different processing strategies based on image characteristics. Deconvolution restoration algorithms, based on regularization theory, restore degraded images by minimizing a cost function. Perspective correction algorithms, based on projective geometry, achieve geometric correction of images using a perspective transformation matrix. Local contrast enhancement and edge enhancement, based on nonlinear image processing techniques, adjust the enhancement strength through adaptive parameter adjustments. Guided filtering, based on edge-preserving filtering theory, suppresses noise while preserving edge details. Combining these algorithms can provide effective solutions to different types of image degradation problems.

[0208] The hull attitude light reflection model is used to analyze the influence of the hull surface light reflection characteristics on the ship name recognition. Its calculation process is expressed as follows:

[0209] Optionally, the hull surface reflection model is expressed as:

[0210]

[0211] Where, I reflect Indicates the intensity of reflected light; k d Represents the diffuse reflection coefficient, obtained by the material recognition network; represents the surface normal vector, which is calculated from image analysis; Indicates the direction of incident light, obtained from the illumination estimation algorithm; k s Represents the specular reflection coefficient, obtained by the material recognition network; Indicates the direction of reflected light; Indicates the viewing direction; n represents the roughness index, which is obtained from texture analysis and generally ranges from 1 to 200. The larger the value, the smoother the surface.

[0212] Optionally, the surface roughness calculation is expressed as:

[0213]

[0214] Where R roughness Represents surface roughness; σ gradient represents the local gradient standard deviation; σ gradient,max Indicates the global maximum gradient standard deviation. Roughness values ​​range from 0 to 1, with larger values ​​indicating smoother surfaces.

[0215] Optionally, the coating aging degree is calculated as:

[0216]

[0217] Where A coating Indicates the degree of coating aging; CV local Represents the local color variation coefficient; CV reference Indicates the reference color variation coefficient. A larger aging degree value indicates more serious coating aging.

[0218] Optionally, the effective contrast matrix of the ship name area is calculated as:

[0219]

[0220] Where C effective represents the effective contrast matrix; I text Indicates the brightness of the text area; I background Indicates the brightness of the background area; A coating Indicates the degree of coating aging; α indicates the incidence angle influence coefficient, α = 0.6.

[0221] Optionally, the text edge sharpness evaluation matrix is ​​calculated as:

[0222]

[0223]

[0224] Where S edge represents the edge sharpness evaluation matrix; represents the gradient amplitude; R roughness Indicates surface roughness; A coating Indicates the degree of coating aging; β is the roughness influence coefficient, β = 0.4; γ is the aging influence coefficient, γ = 0.5.

[0225] Optionally, outputting the optimized ship name area image and coordinate information in step S09 involves the following calculation process:

[0226] Optionally, the coordinate information of the ship name area is calculated as follows:

[0227] B final ={x min ,y min , x max ,y max ,θ,cnof};

[0228] Where B final Represents the final bounding box information; x min 、y min 、x max 、ymax They represent the coordinates of the upper left corner and lower right corner of the bounding box respectively; θ represents the rotation angle; conf represents the confidence score.

[0229] Optionally, the confidence score calculation is expressed as:

[0230]

[0231] Where, conf represents the confidence score; |M roi3 | represents the number of pixels in the third-level interest area; P text represents the text region probability; RM represents the recognition matrix.

[0232] Optionally, the boundary fine-tuning algorithm is expressed as:

[0233] B adjusted =B final +ΔB;

[0234] Where B adjusted Indicates the adjusted boundary; B final represents the initial boundary; ΔB represents the boundary adjustment amount, ΔB=f adjustment (G text ), where f adjustment represents the boundary adjustment function, G text Represents the text edge gradient map. Adjustment accuracy is controlled within 2 pixels of the original boundary.

[0235] It's important to note that token segmentation is a key method for accurately locating the boundaries of fishing vessel names and numbers in ViT-based image processing. Based on the characteristics of the Vision Transformer, it decomposes the image into a sequence of tokens and then performs refined processing to accurately segment the target area.

[0236] The core of token segmentation technology lies in applying the basic unit (token) processed by the Vision Transformer to image segmentation tasks. In traditional ViT models, the input image is segmented into fixed-size image blocks (typically 16×16 pixels). Each image block is converted into a token vector, which is then processed by the Transformer encoder to form a feature representation. Token segmentation technology innovates on this foundation by treating each token as a classifiable semantic unit and analyzing the correlation and semantic similarity between tokens to determine whether they belong to the same target region.

[0237] Specifically, in the process of determining the third-level interest area, the implementation steps of the token segmentation technology are as follows:

[0238] High-resolution Token mapping: Divide the image of the second-level region of interest into smaller image blocks (such as 8×8 or 4×4 pixels) to generate a high-density token sequence to improve spatial resolution.

[0239] Token semantic feature extraction: For each token, deep feature representation is extracted through position encoding enhancement and multi-head self-attention mechanism to capture local texture, edge and text features.

[0240] Token clustering and grouping: Based on the similarity calculation of token features, an adaptive threshold clustering algorithm is used to group tokens with similar semantics and determine the token set that may contain the ship name.

[0241] Boundary precision processing: Use the correlation strength between tokens to build a connectivity graph, apply graph segmentation algorithms (such as normalized cut) to determine the precise boundary of the ship name area, and eliminate background interference.

[0242] Multi-scale Token consistency verification: By comparing the consistency of token segmentation results at different scales, the stability and accuracy of boundary positioning are improved and the impact of noise is reduced.

[0243] In this way, Token segmentation technology can break through the limitations of rectangular bounding boxes in traditional target detection and achieve accurate description of the irregular shape of the ship name area, which is particularly suitable for processing text areas on the curved surface of the hull.

[0244] Token segmentation technology plays an important role in accurate boundary positioning, mainly in the following aspects:

[0245] First, Token segmentation leverages the Transformer model's global contextual awareness. In ship image processing, the ship's name and number region often exhibit subtle visual differences from the surrounding background and may have an irregular shape due to the ship's curved surface. Traditional CNN-based bounding box detection methods struggle to accurately capture these features. However, Token segmentation leverages a self-attention mechanism to establish associations between distant pixels, analyze global contextual information, and more accurately identify the true boundaries of the ship's name and number region.

[0246] Secondly, token segmentation technology enables fine-grained pixel-level classification. During the third-level region of interest (ROI) determination process, each token actually represents a small region in the image. By classifying these tokens (as belonging to the ship's name or background) and combining their location information, a precise mask of the ship's name region can be constructed. This pixel-level classification is more precise than traditional bounding boxes and can accommodate deformation and perspective effects of the ship's name on the curved surface of the ship.

[0247] Token segmentation technology also boasts robustness. By deeply extracting and analyzing the features of each token, it maintains high boundary location accuracy even in poor lighting conditions, damaged or partially obscured hulls. Especially when processing low-contrast or blurred images, token segmentation can integrate contextual information, compensate for deficiencies in local features, and improve boundary location reliability.

[0248] Finally, token segmentation technology can be tightly integrated with subsequent recognition modules (such as MarineIDNet). By providing a precise mask for the ship name region, background interference can be eliminated, improving text recognition accuracy. Token segmentation results also serve as an important input for recognizability matrix calculation, helping the system assess recognition difficulty and guiding subsequent enhancement processing.

[0249] Through this token-based fine segmentation method, the system can accurately locate the ship name area in complex marine environments, providing a solid foundation for subsequent ship identification and verification.

[0250] Specifically, the principle of the present invention is: the technical principle of the present invention is based on the organic combination of ViT's global feature modeling capability and the multi-level interest region refinement extraction strategy, and by introducing a quantitative evaluation mechanism, it can achieve accurate extraction of the fishing vessel name area in a complex environment.

[0251] First, the ViT-based three-level pyramid feature extraction architecture leverages its powerful global contextual awareness to overcome the limitations of traditional CNNs, which rely on local receptive fields. By segmenting the input image into sequential image blocks for processing, ViT is able to capture the overall structural features of the ship's hull and effectively identify key areas within the hull's outline through a self-attention mechanism. The strategy of gradually refining the three-level region of interest (ROI) allows for progressive localization, from the overall hull to the precise location of the ship's name. This avoids the difficulty of searching for small targets directly within the entire image, improving localization accuracy and computational efficiency.

[0252] Secondly, the introduction of the distortion matrix and the bias matrix is ​​a key technical innovation of this invention. The distortion matrix, established using the optical imaging degradation equation, quantitatively describes the impact of observation distance on image quality. The bias matrix, constructed based on the ship's attitude and light reflection model, accurately depicts the perspective deformation of the text area caused by changes in the ship's attitude. The resulting discernible matrix provides a scientific basis for subsequent enhancement processing, enabling adaptive optimization of the ship's name area under different observation conditions, addressing the inability of traditional methods to cope with complex environmental changes.

[0253] The design of the MarineIDNet neural network fully considers the unique needs of ship name recognition. The feature extraction branch of the dual-branch architecture inherits the global modeling advantages of ViT while introducing local receptive field constraints to enhance sensitivity to small text regions. The region verification branch ensures that the extracted region contains valid ship name information to prevent false detections. Stratified sampling and key sampling techniques are used during pre-training to specifically enhance the model's robustness under harsh conditions, addressing the low recognition accuracy of existing technologies in complex environments.

[0254] In summary, the present invention solves the problem of accurate extraction of fishing vessel name areas in complex environments through ViT system design combined with multi-level interest region strategy, quantitative evaluation mechanism and adaptive enhancement processing.

[0255] A specific embodiment 1 of the present invention is provided below. The specific implementation of each step in this embodiment 1 is described in detail as follows.

[0256] The specific implementation of step S01 is to build a three-level pyramid feature extraction structure based on ViT. The structure adopts a multi-scale feature representation strategy to capture the hierarchical information in the fishing boat image through feature maps at different resolutions. First, the input image is resized to 224×224 pixels and divided into image blocks of 16×16 size. The image block set is represented as P={p i,j |i=1,2,...,N h ; j = 1, 2, ..., N w}, where P is the set of image blocks, p i,j represents the image block at position (i, j), N h and N w Respectively represent the number of blocks in the height and width directions of the image. For a 224×224 image, it is divided into 16×16 image blocks. N h =N w =14. Then each image block is linearly projected into a 768-dimensional feature space, and position encoding is added to retain the spatial position information. This process is expressed as z0=[x class ;p1E;p2E;...;p N E]+E pos , where z0 represents the initial feature representation of the image block sequence, x class is a special mark used for classification, p n represents the nth image block, E is the linear projection matrix, which maps the image block to the D-dimensional feature space, D = 768, E pos Is the position encoding matrix, the dimension is the same as z0, the position encoding is through E pos (pos, 2i) = sin(pos / 10000 2i / D ) and Epos (pos, 2i + 1) = cos(pos / 10000 2i / D ) Calculate, where pos represents the position in the sequence, and i represents the feature dimension index, with the range 0 ≤ i < D / 2. Then, these feature representations are processed by three Transformer encoders with different depths. The shallow encoder contains 4 attention layers for capturing local texture features, the middle encoder contains 8 attention layers for learning medium-complexity structural features, and the deep encoder contains 12 attention layers for understanding global semantic information. The multi-head self-attention calculation process is expressed as where Q, K, and V represent the query, key, and value matrices respectively, and d k represents the dimension of the key vector, which is used to scale the dot product result to prevent gradient vanishing, and d k = D / h = 64, where h = 12 is the number of attention heads. Finally, the feature pyramid network structure is used to fuse the feature representations at three different levels to generate a multi-scale feature map. The fusion process is expressed as F multi = α1F shallow + α2F middle + α3F deep , where F multi represents the fused multi-scale feature, and F shallow , F middle , F deep represent the output features of the shallow, middle, and deep encoders respectively. α1, α2, and α3 are fusion weight coefficients, which are optimized through backpropagation and are generally set to α1 = 0.2, α2 = 0.3, and α3 = 0.5 initially. The purpose of this step is to construct a comprehensive representation of the fishing boat image and provide rich visual feature information for subsequent region of interest localization, especially to improve the adaptability to boat name areas of different sizes through multi-scale representation.

[0257] The specific implementation of step S02 is to use the self-attention mechanism to determine the first-level region of interest. This step first calculates the attention weight matrix for each Token in the input image and normalizes the weight distribution through the Softmax function; then analyzes the attention heat map, and the calculation of the attention heat map is expressed as where A map represents the average attention heat map, A i represents the attention matrix of the i-th attention head, and h represents the number of attention heads, h = 12. The regions with weight values higher than the 0.65 threshold are identified as candidate regions of interest. The identification process of high-attention regions is expressed as M roi1 = {(x, y)|A map (x, y) > τ1}, where M roi1represents the first-level interest region mask, (x, y) represents the pixel coordinates in the image, τ1 represents the attention threshold, τ1 = 0.65. Then, the connected component analysis algorithm is applied to merge the adjacent high-attention regions into the overall region; then the visual area ratio of each region is calculated, and the regional area ratio calculation is expressed as where R area Indicates the area ratio of the region, Area(M roi1 ) represents the area of ​​the region of interest, Area(I total ) represents the total area of ​​the image, and the area that accounts for more than 15% of the total area of ​​the image is retained. Finally, the geometric shape analysis algorithm is applied to calculate the perimeter and area ratio of the region, contour complexity and other features. The contour complexity calculation is expressed as Among them C complexity The contour complexity index (P) represents the region perimeter, and A represents the region area. Regions with a perimeter-to-area ratio between 0.08 and 0.25 and a contour complexity index below 1.8 are retained. This step aims to quickly locate the primary region containing the ship hull in the global image, reducing the search space for subsequent processing and improving computational efficiency. It also eliminates background interference and ensures that the focus is on the ship structure.

[0258] The specific implementation of step S03 is to determine the second level interest area based on the first level interest area. This step first applies enhanced position coding to the first level interest area, using a combination of sine-cosine position coding and learnable position coding. The enhanced position coding calculation is expressed as in represents enhanced position encoding, represents the sine-cosine position encoding, Represents the learnable position encoding, λ is the fusion coefficient, λ = 0.6. Then the region is fine-grained analyzed by the multi-head attention mechanism, 8 attention heads are set, and the dimension of the attention matrix of each head is 64 to capture the features of different semantic subspaces. The multi-head attention response calculation is expressed as Among them A response represents the attention response map, w i Represents the weight of the i-th attention head, Attention i represents the output of the i-th attention head, w i Obtained through training, the initial value is set to equal weight Then the attention response map is calculated and threshold segmentation is applied to extract the area with an attention response value greater than 0.78 as the candidate area. The second-level interest area is determined and represented as M roi2 ={(x, y)|A response (x, y)>τ2}∩R prior , where M roi2represents the second-level interest region mask, τ2 represents the attention response threshold, τ2 = 0.78, R prior The ' ' represents a constrained region based on prior knowledge, determined from a ship database. Prior knowledge constraints are then introduced, and based on statistical information from the ship database, the search range is limited to the front of the hull and the sidewalls. Finally, morphological operations are applied to optimize the region boundaries, using a combination of opening and closing operations to remove noise and smooth edges. The structural element size is set to 3×3 pixels. This step aims to further narrow the recognition range and accurately locate key areas on the hull that may contain the ship's name, such as the sidewalls and the wheelhouse surface, laying the foundation for subsequent fine-grained recognition.

[0259] The specific implementation of step S04 is to apply the fine-grained feature extraction network to determine the third-level interest region. This step first applies the high-resolution feature extraction module to the second-level interest region, dividing the region into smaller 4×4 pixel blocks for fine analysis. The fine-grained feature extraction is expressed as Among them F fine Represents fine-grained features, Conv represents convolution operation, and Resize represents resampling operation. The second-level interest area is reduced to one-fourth of its original size to obtain a finer feature representation. Feature enhancement is then performed through a fine-grained attention network, which contains four attention heads, each of which focuses on character-level visual features. Then, a text region detection algorithm is applied to identify potential text regions based on the stroke density map and text direction field. The text region detection calculation is represented by P text =σ(W2·ReLU(W1·F fine +b1)+b2), where P text represents the text area probability map, W1, W2, b1, b2 are model parameters, σ represents the Sigmoid activation function, ReLU represents the rectified linear unit activation function, and the text area probability threshold is set to 0.82. Then, the text area boundary is accurately located by Token segmentation technology, and the adaptive threshold method is used to separate the foreground text and background. The adaptive threshold calculation is expressed as Where T adaptive represents the adaptive threshold, μ(x, y) represents the local area mean, σ(x, y) represents the local area standard deviation, γ represents the contrast weight coefficient, γ = 0.7. The third level interest area is determined as M roi3 ={(x, y)|P text (x, y)>τ3}, where M roi3 Denotes the third-level interest area mask, τ3 denotes the text area probability threshold, τ3 = 0.82. Finally, the boundary optimization algorithm is applied to fine-tune the detected text area, including minimum bounding rectangle calculation, boundary extension reservation and boundary regularization processing. The boundary extension processing is represented by B expanded=Expand(B original , r expand ), where B expanded represents the expanded boundary, B original represents the original boundary, r expand Indicates the expansion ratio, r expsand =0.1, indicating that the expansion width is 10% of the original text area width. The purpose of this step is to accurately locate the specific area where the ship name is located, to achieve pixel-level precision in the extraction of the ship name area, and to ensure that subsequent processing can focus on the minimum necessary area containing the ship name.

[0260] The specific implementation of step S05 is to construct a distortion matrix for the third-level area of ​​interest. This step first extracts the observation distance from the image metadata, obtains the atmospheric visibility from the meteorological database, obtains the optical parameters from the camera parameter file, measures the actual size of the ship name area from the image analysis, and obtains the imaging time from the metadata. Then, a mathematical model is established based on the optical imaging degradation equation. The model considers the comprehensive impact of factors such as atmospheric scattering, the diffraction limit of the optical system, and the focus offset on the image quality. The optical imaging degradation equation is expressed as g(x, y) = h(x, y)*

[0261] f(x, y) + n(x, y), where g(x, y) represents the observed image, h(x, y) represents the point spread function, f(x, y) represents the ideal undistorted image, n(x, y) represents additive noise, and * represents the convolution operation. The point spread function calculation is expressed as h(x, y) = h system (x, y)*h atmospheric (x, y)*h defocus (x, y), where h system represents the point spread function of the optical system, h atmospheric represents the atmospheric scattering point spread function, h defocus Specifically, the point spread function of the optical system is expressed as Where J1 represents the first-order Bessel function, r represents the radial distance, and r0 represents the radius of the Airy disk. Where λ represents the wavelength of light, f represents the focal length, and D represents the aperture diameter; the atmospheric scattering point spread function is expressed as Where r0 represents the atmospheric coherence length, Where L represents the observation distance, C n represents the atmospheric refractive index structure constant, which is related to meteorological conditions; the out-of-focus spread function is expressed as where R defocus represents the defocus radius, Where s represents the actual object distance. Then the point spread function at different observation distances is calculated to quantify the degree of image blur. The distance-dependent point spread function radius increase is calculated as where ΔR PSF represents the increase in the point spread function radius, d represents the observation distance (meters), d0 represents the reference distance, d0 = 50 meters, α PSF represents the radius increase coefficient, α PSF = 0.5 pixels / 10 meters. When the observation distance exceeds 50 meters, the point spread function radius increase should not be less than 0.5 pixels / 10 meters. Then generate the signal-to-noise ratio attenuation matrix and analyze the ratio of signal intensity to noise level in each sub-area. The signal-to-noise ratio attenuation matrix is ​​calculated as follows: Where SNR(x, y) represents the signal-to-noise ratio (dB) at position (x, y), S(x, y) represents the signal strength, N(x, y) represents the noise strength, and areas below 6 decibels are marked as high distortion risk areas. Finally, the normalization process generates a standardized distortion matrix, which is calculated as Among them D norm represents the normalized distortion matrix, SNR threshold Indicates the signal-to-noise ratio threshold, SNR threshold =20dB, D norm The value range is 0 to 1, where 0 indicates no distortion and 1 indicates complete distortion. The purpose of this step is to quantitatively evaluate the impact of observation distance on the image quality of the ship name area, providing a scientific basis for subsequent image enhancement processing and ensuring targeted improvement of low-quality areas.

[0262] The specific implementation of step S06 is to calculate the deviation matrix of the third-level region of interest. This step first applies the hull attitude light reflection model to analyze the interaction between light and the hull surface, extracts the hull surface normal vector from the image, obtains the incident light angle through the illumination estimation algorithm, obtains the hull material reflection coefficient using the material recognition network, obtains the surface roughness through texture analysis, and obtains the coating aging degree based on image degradation assessment. The calculation of the hull surface normal vector is expressed as in represents the surface normal vector, represents the image gradient, represents the gradient modulus; the incident light angle estimation is expressed as in Represents the normalized incident light direction vector, S represents the pixel set in the highlight area, represents the half vector of the i-th pixel, Represents the surface normal vector of the i-th pixel. The hull surface reflection model is expressed as Among them I reflect Indicates the intensity of reflected light, k dRepresents the diffuse reflection coefficient, obtained by the material recognition network, represents the surface normal vector, calculated from image analysis, Indicates the direction of incident light, obtained from the illumination estimation algorithm, k s Represents the specular reflection coefficient, obtained by the material recognition network, Indicates the direction of reflected light, Indicates the viewing direction, n represents the roughness index, which is obtained from texture analysis and generally ranges from 1 to 200. The larger the value, the smoother the surface. The surface roughness calculation is expressed as where R roughness represents the surface roughness, σ gradient represents the local gradient standard deviation, σ gradient,max Indicates the global maximum gradient standard deviation, and the roughness value range is 0 to 1. The larger the value, the smoother the surface. The coating aging degree is calculated as Among them A coating Indicates the degree of coating aging, CV local Represents the local color variation coefficient, CV reference It represents the reference color variation coefficient. The larger the aging degree value, the more serious the coating aging. Then, a relationship model between the hull posture and perspective deformation is established to calculate the tilt angle of the hull surface relative to the camera. The relationship model between the hull posture and perspective deformation is expressed as D perspective (θ)=1+β·max(0,θ-θ0), where D perspective Indicates the text deformation coefficient, θ indicates the angle between the hull surface and the camera plane (degrees), θ0 indicates the critical angle, θ0 = 30 degrees, β indicates the deformation coefficient, β = 0.02 / degree, that is, when the angle increases by 1 degree, the deformation coefficient increases by 0.02. When θ exceeds 60 degrees, D perspective It is greater than 1+0.02·(60-30)=1.6, which meets the requirements. Then we analyze the perspective deformation degree of the text area under different posture angles and establish a quantitative index of the deformation degree. The effect of perspective deformation angle on recognition difficulty is expressed as Among them D difficulty Indicates the recognition difficulty coefficient, θ indicates the perspective deformation angle (degrees), and the recognition difficulty coefficient increases with each 10-degree increase in the perspective deformation angle. Then the effective contrast matrix of the ship name area is generated to evaluate the contrast changes between the text and the background under different lighting conditions. The calculation of the effective contrast matrix of the ship name area is expressed as:

[0263] Among them C effective represents the effective contrast matrix, I text Indicates the brightness of the text area, I bakcground Indicates the brightness of the background area, A coatingIndicates the degree of coating aging, α indicates the incident angle influence coefficient, α = 0.6, and the area with an effective contrast lower than 0.3 is marked as a high-risk area. Finally, the text edge sharpness evaluation matrix is ​​calculated to analyze the text outline clarity. The text edge sharpness evaluation matrix is ​​calculated as Among them S edge represents the edge sharpness evaluation matrix, represents the gradient amplitude, R roughness Indicates surface roughness, A coating Indicates the degree of coating aging, β indicates the roughness influence coefficient, β = 0.4, γ indicates the aging influence coefficient, γ = 0.5, and the area with edge sharpness lower than 0.4 requires special enhancement treatment. The deviation matrix calculation is expressed as B(x, y) = 1-min(C effective (x, y), S edge (x, y)), where B(x, y) represents the bias matrix, C effective Indicates the effective contrast, S edge B(x, y) represents edge sharpness, with values ​​ranging from 0 to 1, where 0 indicates no deviation and 1 indicates maximum deviation. The purpose of this step is to quantitatively assess the degree of geometric deformation of the ship name area caused by changes in the ship's posture, providing accurate parameters for subsequent perspective correction and image enhancement.

[0264] The specific implementation of step S07 is to generate a recognizable matrix based on the comprehensive calculation of the distortion matrix and the deviation matrix. This step first applies a multi-feature fusion function to integrate various influencing factors, including the distortion matrix generated in the previous step, the calculated deviation matrix, the ambient brightness matrix obtained from image analysis, the hull texture complexity obtained from texture analysis, and the ship name font feature vector obtained from the font recognition network. The multi-feature fusion function is expressed as R(x, y) = w1·D norm (x, y) + w2·B(x, y) + w3·L(x, y) + w4·T(x, y) + w5·F(x, y), where R(x, y) represents the original comprehensive score, D norm represents the normalized distortion matrix, B represents the deviation matrix, L represents the ambient brightness matrix, T represents the complexity of the hull texture, F represents the font feature vector of the ship name, w1, w2, w3, w4, and w5 represent the weights of each factor, w1 = 0.35, w2 = 0.3, w3 = 0.15, w4 = 0.1, and w5 = 0.1. The ambient brightness matrix is ​​calculated as Where L(x, y) represents the ambient brightness deviation, I(x, y) represents the brightness value at position (x, y), and I max Indicates the maximum brightness value. The closer the value of L(x, y) is to 0, the brightness is moderate, and the closer it is to 0.5, the brightness is too high or too low. The texture complexity calculation is expressed as Where T(x, y) represents the texture complexity, N represents the number of pixels in the local area, Indicates position (x i ,y i ). The comprehensive score is then calculated by weighted summation, with the weights of the distortion factor being 0.35, the deviation factor being 0.3, the brightness factor being 0.15, the texture complexity being 0.1, and the font feature being 0.1. A nonlinear mapping function is then applied to convert the original score into a standardized recognition difficulty coefficient, which is expressed as: where R norm represents the standardized recognition difficulty coefficient, R represents the original comprehensive score, k represents the Sigmoid function steepness parameter, k = 5, R0 represents the center point parameter, R0 = 0.5, R norm The value range is 0 to 1, where 0 means easy to identify and 1 means extremely difficult to identify. Then, a heat map of the recognition degree is generated based on the recognition difficulty coefficient, which intuitively shows the distribution of the recognition difficulty of each part of the ship name area. The recognition matrix is ​​generated as follows: R norm τ4 represents the standardized recognition difficulty coefficient, and the recognition difficulty threshold is τ4 = 0.7. Areas with a recognition difficulty coefficient above 0.7 are marked as high-risk areas. Finally, optimization strategies are recommended based on the recognition score, including recommended image processing methods and parameter settings. When the overall recognition score falls below 0.5, adaptive enhancement is recommended. This step aims to comprehensively assess the recognition difficulty of the ship name area under current conditions and provide a basis for decision-making on subsequent targeted enhancement processing.

[0265] The specific implementation of step S08 is to design an adaptive region enhancement algorithm. This step first identifies the region that needs to be enhanced based on the identifiable matrix, and focuses on the region with an identifiable coefficient greater than 0.6. Then, differentiated enhancement strategies are designed for different problems. For distortion-dominated problems, a deconvolution recovery algorithm is applied. The deconvolution recovery algorithm is expressed as in Denotes the restored image, h denotes the point spread function, g denotes the observed image, D denotes the difference operator, λ denotes the regularization parameter, λ = 0.02, and the convolution kernel size is 5 × 5. For the bias-dominated problem, the perspective correction algorithm is applied, which is expressed as Where H represents the perspective transformation matrix, A1 represents the four corner coordinate matrices of the source quadrilateral, and A2 represents the four corner coordinate matrices of the target rectangle. The transformed image is obtained by Calculate, where I corrected Represents the corrected image, I original represents the original image, h ijRepresents the elements of the perspective transformation matrix H. Perspective transformation is performed based on the four-point correspondence relationship, and the control points are selected at the four corners of the text. For the problem of insufficient contrast, local contrast enhancement is applied. The local contrast enhancement algorithm is expressed as I enhanced (x,y)=I(x,y)+α(x,y)·(I(x,y)-I blur (x, y)), where I enhanced represents the enhanced image, I represents the original image, and I blur represents the image after Gaussian blur, α represents the enhancement coefficient, Ensure that the maximum gain does not exceed 200% of the original contrast. The enhancement coefficient is adaptively adjusted according to the local brightness, and the maximum gain does not exceed 200% of the original contrast. Then edge enhancement processing is applied to improve the clarity of the text edge. The edge enhancement processing is represented by I edge (x,y)=I(x,y)+β(x,y)·(I(x,y)-I blur (x, y) 2 ·sign(I(x,y)-I blur (x, y)), where I edge represents the edge enhanced image, I represents the original image, and I blur represents the image after Gaussian blur, β represents the nonlinear enhancement coefficient, Where β0 represents the basic enhancement coefficient, β0=10, ∈ represents a small constant to prevent division by zero, ∈=0.01, and a nonlinear gradient enhancement method is used. The enhancement strength is inversely proportional to the edge sharpness matrix. Finally, a noise suppression filter is applied to protect the text edge while suppressing noise based on the guided filtering algorithm. The guided filtering algorithm is expressed as I fltered (x, t) = a (x, y) G (x, y) + b (x, y), where I filtered Represents the filtered image, G represents the guide image, generally the original image, a and b are the coefficients of the local linear model, by minimizing the cost function Solve, where ω (x,y) represents a window centered at (x, y), ∈ represents the regularization coefficient, ∈ = 0.1, and the filter radius is set to 3 pixels. The purpose of this step is to perform targeted enhancement on low-recognition areas based on the aforementioned analysis results, improve the overall recognizability of the ship name area, and provide high-quality input for the subsequent text recognition system.

[0266] The specific implementation of step S09 is to output the optimized ship name area image and coordinate information. This step first integrates the above processing results to generate a high-quality ship name area image; then calculates the precise coordinate information of the ship name area. The calculation of the coordinate information of the ship name area is expressed as B final ={x min ,y min , xmax ,y max ,θ,conf}, where B final Represents the final bounding box information, x min 、y min 、x max 、y max Represent the coordinates of the upper left corner and lower right corner of the bounding box, θ represents the rotation angle, and conf represents the confidence score. The confidence score calculation is expressed as Where conf represents the confidence score, |M roi3 | represents the number of pixels in the third-level interest area, P text Denotes the probability of the text region, and RM denotes the recognizable matrix. The processed image and coordinate information of the ship name region are then transmitted to the verification system. The MarineIDNet neural network is then applied to verify the ship name region. This network uses a dual-branch structure consisting of feature extraction and region verification branches, verifying the region validity through multi-layer feature extraction and comparative analysis. Finally, the boundaries of the verified region are fine-tuned. The boundary fine-tuning algorithm is represented by B adjusted =B final +ΔB, where B adjusted represents the adjusted boundary, B final represents the initial boundary, ΔB represents the boundary adjustment amount, ΔB=f adjustment (G text ), where f adjustment represents the boundary adjustment function, G text This step represents the gradient map of the text edge, with the adjustment accuracy controlled within 2 pixels of the original boundary. The purpose of this step is to provide the final optimized image of the ship name area and precise location information, providing standardized input for the subsequent text recognition system. At the same time, the verification step ensures the accuracy and reliability of the extraction results.

[0267] In order to better understand and implement the present invention, Example 2 of a specific application scenario of the present invention is provided below: In an intelligent ship monitoring project of a certain marine research institution, researchers used a ViT-based method for extracting the name and number regions of fishing vessels entering and leaving the port to automatically monitor and identify fishing vessels entering and leaving a certain fishing port. The project aims to improve the efficiency of fishing vessel management, reduce labor costs, and ensure that the ship names can be accurately identified under various complex environmental conditions. The researchers selected 50 fishing vessels of different types and collected a total of 500 high-definition images under different lighting, weather and observation distance conditions for testing to verify the effectiveness and robustness of the method. The test images are divided into three groups: a close-range group (5 to 10 meters), a medium-range group (10 to 30 meters), and a long-range group (30 to 100 meters), each group containing images under different weather and lighting conditions.

[0268] During the system initialization phase, the researchers first set parameters and pre-trained ViT. The model parameter settings are shown in Table 1:

[0269] Table 1 ViT parameter setting table

[0270]

[0271]

[0272] For a typical test image (resolution 3840×2160, medium distance group, sunny lighting, distance about 15 meters), step S01 is applied to construct a three-level pyramid feature extraction structure. After processing by three different depth encoders, the fusion feature weight coefficients are set to α1=0.2 (shallow layer), α2=0.3 (middle layer) and α3=0.5 (deep layer). The fused multi-scale feature F multi The dimension is 768, including global and local feature information.

[0273] In step S02, the self-attention mechanism is applied to determine the first-level interest area, and the average attention heat map A is calculated. map The image size is 224×224. An attention threshold of τ1 = 0.65 is set to extract high-attention regions. Connected component analysis yields three candidate regions, accounting for 23.6%, 18.7%, and 9.4% of the area, respectively. Geometric shape analysis yields contour complexity indices of 1.24, 1.76, and 2.31, respectively, and perimeter-to-area ratios of 0.11, 0.19, and 0.28, respectively. The first and second regions are retained as first-level regions of interest.

[0274] In step S03, enhanced position encoding is applied to the first-level ROI, followed by fine-grained analysis using eight attention heads (each with a dimension of 64). The attention response threshold τ2 is set to 0.78, and combined with the prior knowledge constraints in the ship database, the ship's side area on the right side of the hull is ultimately determined as the second-level ROI, with a size of 512 × 256 pixels.

[0275] In step S04, a fine-grained feature extraction network was applied to the second-level ROI, segmenting it into 4×4 pixel blocks for detailed analysis. Using a text region detection algorithm with a probability threshold of τ3 = 0.82, the ship's name region on the side of the ship was successfully detected. An adaptive threshold algorithm (with a contrast weighting factor of γ = 0.7) was applied to accurately segment the text region, ultimately determining a third-level ROI of 124×42 pixels. A boundary expansion ratio of 0.1 was set, resulting in an expanded region size of 136×46 pixels.

[0276] In step S05, the key parameters extracted when constructing the distortion matrix are shown in Table 2:

[0277] Table 2 Image quality parameters

[0278] Parameter name Parameter value Observation distance 15.3 meters Atmospheric visibility 23.5 kilometers Camera focal length 85 mm Aperture diameter 17 mm Actual size of the ship name area 0.9×0.3 meters Imaging time 14:25:36 Atmospheric refractive index structure constant <![CDATA[3.5×10 -14 ]]>

[0279] Based on the above parameters, the calculated point spread function radius is 1.2 pixels, and the lowest value in the signal-to-noise ratio attenuation matrix is ​​12.6 dB, which is higher than the 6 dB threshold, indicating that the image quality is good at this observation distance. norm The average value is 0.37.

[0280] In step S06, when calculating the deviation matrix, the key parameters extracted from the image are shown in Table 3:

[0281] Table 3 Hull light reflection parameters

[0282]

[0283]

[0284] Text deformation coefficient D caused by hull posture perspective =1.24, the influence coefficient of perspective deformation angle on recognition difficulty D difficulty =1.84, effective contrast matrix C effective The average value is 0.52, which is higher than the 0.3 threshold. edge The average value is 0.61, which is higher than the threshold of 0.4. The average value of the final calculated deviation matrix B is 0.48.

[0285] In step S07, the factors are integrated through the multi-feature fusion function. The average value of the environment brightness matrix L is 0.12, the texture complexity T is 0.28, and the similarity of the font feature vector F is 0.83. The weights w1 = 0.35 (distortion), w2 = 0.3 (bias), w3 = 0.15 (brightness), w4 = 0.1 (texture), and w5 = 0.1 (font) are used to calculate the original comprehensive score R, which is then mapped to the standardized recognition difficulty coefficient R through the Sigmoid function (parameters k = 5, R0 = 0.5). norm , the average value is 0.56, which is lower than the 0.7 threshold, but still needs to be enhanced.

[0286] In step S08, since the recognition difficulty coefficient of 0.56 is greater than 0.5, adaptive enhancement processing is required. According to the recognizable matrix analysis, the ship name area is mainly affected by perspective deformation, so the perspective correction algorithm is applied. The perspective transformation matrix H is calculated and the text area is corrected. Then, the local contrast enhancement algorithm is applied, and the enhancement coefficient is adaptively adjusted, with an average value of 1.36. In the edge enhancement process, the average value of the nonlinear enhancement coefficient is 16.39. Finally, a guided filter is applied for noise suppression, with a filter radius of 3 pixels and a regularization coefficient of 0.1. The enhancement processing effect is shown in Table 4:

[0287] Table 4 Comparison of image enhancement effects

[0288] Evaluation indicators Before treatment After processing Improvement rate Effective contrast 0.52 0.78 50.0% Edge sharpness 0.61 0.85 39.3% Signal-to-noise ratio 12.6 decibels 18.9 decibels 50.0% Text deformation coefficient 1.24 1.03 16.9% Identification difficulty coefficient 0.56 0.28 50.0%

[0289] In step S09, the final optimized image and coordinate information of the ship name area is output. The bounding box coordinates are (1562, 873, 1698, 919), the rotation angle is 2.7 degrees, and the confidence score is 0.92. MarineIDNet is used for verification, confirming that the extraction results are accurate and valid, and the final deviation after fine-tuning the boundaries is less than 1.8 pixels.

[0290] By analyzing the processing results of all 500 test images, the accuracy of ship name area extraction in different distance groups is shown in Table 5:

[0291] Table 5 Statistics of ship name area extraction accuracy under different conditions

[0292] Test conditions Number of samples Accurately extract the number Accuracy Sunny day at close range 65 64 98.5% Close-up cloudy 55 53 96.4% Close-up dusk 35 33 94.3% Medium distance sunny 70 67 95.7% Medium-range cloudy 60 56 93.3% Middle distance dusk 40 36 90.0% Long-distance sunny day 80 70 87.5% Long-distance cloudy sky 60 51 85.0% distant dusk 35 28 80.0% total 500 458 91.6%

[0293] Traditional methods for extracting fishing vessel name regions primarily rely on convolutional neural network-based target detection algorithms, such as YOLO, SSD, or Faster R-CNN. These methods suffer from several technical limitations: First, they have limited ability to detect small targets at long distances and under varying lighting conditions, with detection accuracy typically below 70% at distances beyond 30 meters. Second, traditional methods lack modeling for perspective distortion and image quality degradation, resulting in a sharp drop in performance in complex environments. Third, they fail to consider the hull material and light reflection characteristics, resulting in insufficient recognition accuracy for text on the ship's side. Finally, traditional methods typically employ fixed-parameter image enhancement strategies, making it impossible to tailor processing to different degradation types. In contrast, the present invention's ViT-based three-level pyramid feature extraction structure can better capture global and local features, significantly improving detection capabilities under long-distance conditions (by approximately 25%). The imaging process is accurately modeled through optical imaging degradation equations and hull attitude light reflection models, enabling the system to understand and respond to various complex environments. The introduction of multi-feature fusion functions and identifiable matrices enables a quantitative assessment of recognition difficulty. The adaptive region enhancement algorithm can select the optimal processing strategy based on different types of image degradation. Overall, the present invention achieves an average accuracy of 91.6% under various complex environmental conditions, an increase of approximately 20 percentage points over traditional methods, showing significant advantages in particular at long distances and under low-light conditions.

[0294] It should be noted that the variables involved in the present invention are explained in detail as shown in Tables 6, 7, 8 and 9 below.

[0295] Table 6 Variable Explanation Table (Part 1)

[0296]

[0297] Table 7 Variable Explanation Table (Part 2)

[0298]

[0299]

[0300] Table 8 Variable Explanation Table (Part 3)

[0301]

[0302]

[0303] Table 9 Variable Explanation Table (Part 4)

[0304]

[0305] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered by the scope of protection of the present invention.

Claims

1. A method for extracting the name area of ​​fishing vessels entering and leaving a port based on ViT, characterized in that: include: Design of a three-level pyramid feature extraction structure based on ViT for fishing boat image analysis; The self-attention mechanism is used to determine the first-level interest region; the second-level interest region is extracted based on the first-level interest region; the fine-grained feature extraction network is used to determine the third-level interest region; and a distortion matrix is ​​constructed for the third-level interest region; Calculate the deviation matrix of the third-level region of interest; generate the recognition matrix based on the comprehensive calculation of the distortion matrix and the deviation matrix; design an adaptive region enhancement algorithm to perform targeted image enhancement processing on low-recognition areas according to the guidance of the recognition matrix; Output the final optimized ship name area image and provide the ship name area coordinate information for subsequent text recognition and verification systems.

2. The method according to claim 1, characterized in that The first-level area of ​​interest refers to the large-scale area containing the main structure of the hull determined after preliminary analysis of the overall image of the fishing vessel by ViT. The first-level area of ​​interest covers the main outline of the fishing vessel and the surrounding area; the second-level area of ​​interest refers to the hull structure area determined after further narrowing the scope on the basis of the first-level area of ​​interest. The second-level area of ​​interest includes the key parts of the ship's side and the outer wall of the wheelhouse where the ship's name is marked; the third-level area of ​​interest refers to the precise location area of ​​the ship's name determined after fine processing. The third-level area of ​​interest only includes the ship's name text and its immediate background.

3. The method according to claim 2, characterized in that The distortion matrix is ​​a mathematical model that quantifies the effect of different observation distances on image clarity. It is constructed by analyzing the attenuation of high-frequency information in the image and is used to assess the recognition difficulty caused by distance. The deviation matrix is ​​a mathematical model that characterizes the degree of geometric deformation of the ship name area caused by changes in the ship's attitude angle. The deviation matrix primarily considers the difficulties that perspective projection transformations bring to text recognition. The identifiable matrix is ​​an evaluation index generated by integrating the distortion matrix and the deviation matrix. The identifiable matrix uses numerical values ​​to quantify the difficulty of identifying the ship name area under current observation conditions.

4. The method according to claim 3, characterized in that It also includes applying an optical imaging degradation equation to simulate the change pattern of image quality of the ship name area at different observation distances. The input of the optical imaging degradation equation includes the observation distance obtained from the image metadata, the atmospheric visibility obtained from the meteorological database, the camera optical parameters obtained from the camera parameter file, the actual size of the ship name area measured from the image analysis, and the imaging time obtained from the image metadata. The output of the optical imaging degradation equation is the distance-dependent point spread function and the signal-to-noise ratio attenuation matrix.

5. The method according to claim 4, characterized in that It also includes applying a hull posture light reflection model to analyze the impact of the light reflection characteristics of the hull surface on the recognition of the ship name. The input of the hull posture light reflection model includes the hull surface normal vector calculated from image analysis, the incident light angle obtained from the illumination estimation algorithm, the hull material reflection coefficient obtained from the material recognition network, the hull surface roughness obtained from texture analysis, and the coating aging degree obtained from the image degradation assessment. The output of the hull posture light reflection model is the effective contrast of the ship name area and the text edge sharpness evaluation matrix.

6. The method according to claim 5, characterized in that It also includes applying a multi-feature fusion function to comprehensively consider multiple factors to determine the recognizability of the ship name area. The input of the multi-feature fusion function includes the distortion matrix, the deviation matrix, the ambient brightness matrix obtained from image analysis, the hull texture complexity obtained from texture analysis, and the ship name font feature vector obtained from the font recognition network. The output of the multi-feature fusion function is the area recognizability score and optimization strategy recommendation.

7. The method according to claim 6, characterized in that The method also includes using the MarineIDNet neural network to verify the ship name area and fine-tune the boundaries. The specific structure of the MarineIDNet model is a dual-branch network architecture, including a feature extraction branch and a region verification branch. The feature extraction branch adopts an improved ViT structure, and the region verification branch adopts a four-layer convolutional neural network and two fully connected layers.

8. The method according to claim 7, characterized in that The steps for establishing a training dataset during the MarineIDNet model pre-training process specifically include collecting high-definition images of different types of fishing vessels under various lighting conditions, weather environments, and observation distances, using semi-automatic annotation tools to annotate the vessel name area with rectangular frames and transcribe the text content, and performing stratified sampling and balancing processing based on vessel size, observation distance, and weather conditions. Finally, a pre-training dataset containing 50,000 images is constructed.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program instructions, and when the program instructions are run in a computer, they are used to execute the method for extracting the name area of ​​fishing vessels entering and leaving the port based on ViT according to any one of claims 1 to 8.

10. A ViT-based system for extracting the name and number regions of fishing vessels entering and leaving a port, characterized in that: The computer-readable storage medium according to claim 9 is included, the system is any one of a computer, a server, and a single-chip microcomputer, the computer-readable storage medium is arranged in the system, and the system is provided with a microprocessor for executing program instructions stored in the computer-readable storage medium.

Citation Information

Patent Citations

  • Ship detection method based on remote sensing image

    CN119516369A

  • Complex scene-oriented ship number identification method and system

    CN119672694A

  • Intelligent inspection method and system for compliance of fishing vessels leaving port

    CN119741627A

  • Method and apparatus for generating surround view monitoring image for ship

    US20250121774A1

  • Transformer-based multi-modal and multi-region data fusion framework

    WO2025075870A1