Fishing boat name identification method and system based on image processing and deep learning
Through the method based on image processing and deep learning, the problem of low accuracy of fishing boat names recognition in the prior art and inability to adapt to environmental changes is solved, and high accuracy recognition in complex environments is achieved, and the reliability of recognition and anti-interference ability are improved.
Patent Information
- Application Number
- CN202411942920.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-09
AI Technical Summary
The existing fishing boat name recognition technology has low recognition accuracy under different lighting conditions and paint depths, lacks real-time monitoring and feedback mechanisms, cannot adjust the identification parameters in time, and fails to fully consider the impact of shooting angle and lighting conditions.
Using an image processing and deep learning method, the high accuracy of the fishing boat name is achieved through steps such as image preprocessing and enhancement, adaptive image enhancement, ship name area detection and segmentation, feature extraction and optimization, recognition result feedback and optimization, character recognition and post-processing, correction and verification, etc.
It improves the accuracy and reliability of fishing boat name recognition, can still obtain clear image data that meets the analysis needs in complex environments, has strong anti-interference ability, and adapt to changeable environmental conditions.
Smart Images

Figure CN119964167A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fishing vessel name recognition, and in particular to a fishing vessel name recognition method and system based on image processing and deep learning. Background Art
[0002] At present, the existing technology for ship name recognition usually adopts a method combining traditional optical character recognition (OCR) and image enhancement and processing. Its main implementation steps are as follows:
[0003] (1) Automatic Identification System (AIS) as an auxiliary data source: In some cases, the Automatic Identification System (AIS) is used as an auxiliary data source. AIS can provide basic information of the ship (such as ship name, port of registry, etc.);
[0004] (2) Radar-assisted image detection and acquisition: Use ship-borne or shore-based radar systems to detect and track surface targets (such as ships) to determine the position and heading of the ship. This type of radar system can be used to detect ships at a long distance and provide the outline information of the hull to a certain extent.
[0005] However, the existing solutions have the following shortcomings:
[0006] (1) Fixed image processing parameters are usually used, which have low recognition accuracy under different lighting conditions and different paint depths of fishing boats.
[0007] (2) There is a lack of real-time monitoring and feedback mechanisms, and the recognition parameters cannot be adjusted in time to adapt to environmental changes.
[0008] (3) The impact of shooting angle and lighting conditions on ship name recognition is not fully considered, which may cause the ship name in the image to be blurred due to insufficient lighting or unsatisfactory shooting angle, affecting the final recognition effect.
[0009] (4) Image processing techniques that rely on rule-based techniques, such as edge detection and region growing, do not work well under complex backgrounds or low-contrast conditions.
[0010] (5) Existing OCR technology may not perform well in complex backgrounds or low-resolution images.
[0011] (6) Lack of the ability to self-adjust and optimize based on recognition results.
[0012] (7) The lack of precise calibration and manual verification steps leads to insufficient accuracy of recognition results. Summary of the invention
[0013] The technical problem to be solved by the present invention is to provide a method and system for identifying the names of fishing vessels based on image processing and deep learning, so as to solve the problems raised in the background technology.
[0014] In order to solve the above technical problems, the technical solutions adopted by the present invention are as follows.
[0015] A method and system for identifying fishing vessel names based on image processing and deep learning, comprising the following steps:
[0016] S1. Fishing boat image collection;
[0017] S2. Image preprocessing and enhancement to achieve image noise reduction, deblurring and brightness adjustment, and image quality judgment after processing to form a closed-loop optimization;
[0018] S3. Adaptive image enhancement to achieve illumination compensation and image dynamic enhancement;
[0019] S4. Ship name area detection and segmentation;
[0020] S5. Feature extraction and optimization of ship name area;
[0021] S6. Recognition result feedback and optimization;
[0022] S7. Character recognition and post-processing;
[0023] S8. Calibration and verification;
[0024] S9. Identification log and report generation;
[0025] S10. Result display and storage.
[0026] Preferably, in step S1, images of fishing boats are collected through a multi-camera layout, and according to the position and motion trajectory of the target fishing boat, the camera with the best viewing angle is intelligently selected for image collection; and a real-time video analysis algorithm is deployed on the camera end or edge computing device to perform preliminary processing on the collected video stream, and dynamically adjust the camera parameters according to the processing information; at the same time, the collected video stream is pre-compressed while ensuring image quality.
[0027] Preferably, the brightness adjustment method in step S2 is: using a convolutional neural network to extract features from the input fishing boat image, inputting the extracted features into a fully connected layer, and using the output of the fully connected layer as the basis for adjusting the image enhancement parameters; the output of the fully connected layer is connected to two branches, one branch is used to generate a brightness adjustment coefficient, and the other branch is used to generate a contrast adjustment coefficient; specifically, a sigmoid activation function is used to map the output value to between 0 and 1 as a proportional factor for adjusting the brightness and contrast; according to the calculated brightness and contrast adjustment coefficients, the brightness and contrast of the image are adjusted pixel by pixel; for brightness adjustment, the RGB value of each pixel is multiplied by the brightness adjustment coefficient; for contrast adjustment, the average brightness of the image is first calculated, and then the difference between each pixel and the average brightness is adjusted according to the contrast adjustment coefficient to enhance the contrast of the image;
[0028] The method for judging the image quality in step S2 is as follows: using a model that learns the mapping relationship between various visual features of an image and a quality score to judge the processed image and output a quality score, and judging whether the image is suitable for character recognition based on a preset threshold; at the same time, based on a deep learning algorithm, the image quality indicators including but not limited to clarity, contrast and noise level are evaluated, and based on the evaluation results, optimization suggestions for image preprocessing and enhancement processing algorithm parameters are provided to re-preprocess images with unqualified quality indicators, thereby forming a closed-loop optimization system.
[0029] Preferably, the method of illumination compensation in step S3 is as follows: first, a multimodal image enhancement strategy is adopted for different types of fishing boat images, specifically: a deep learning model is used to classify the images to determine the modal type to which they belong, and then a corresponding image enhancement algorithm combination is applied according to the characteristics of different modalities; for low-light images at night, a denoising algorithm based on deep learning is first used to remove noise, and then an adaptive illumination compensation algorithm is applied to increase the brightness, and then a contrast enhancement algorithm is used to highlight the ship name area; for images of different hull colors, the color correction parameters are adjusted according to the detected color information to ensure that the color of the ship name has a good contrast with the background color;
[0030] The adaptive illumination compensation algorithm is a corresponding compensation strategy generated by predicting the impact of environmental data collected in real time on image quality. Specifically, the image is firstly analyzed for illumination and the illumination histogram of the image is calculated; then, according to the distribution of the illumination histogram, the area with uneven illumination in the image is judged. If the brightness value of a sub-area is significantly lower or higher than other areas, it is considered that the area has insufficient or excessive illumination problems; for the insufficiently illuminated area, the pixel value is increased to compensate, specifically, an interpolation algorithm based on neighboring pixels is adopted; for the area with excessively strong illumination, the pixel value is reduced to compensate;
[0031] The method for dynamic image enhancement in step S3 specifically comprises the following steps:
[0032] S31.Brightness and contrast adjustment;
[0033] S311. Adjust the brightness: calculate the average brightness value of the image. If the average brightness value is lower than the preset lower limit of the appropriate brightness range, increase the brightness value of each pixel by a certain proportion; if it is higher than the upper limit, reduce it proportionally;
[0034] S312. Perform contrast adjustment: calculate the brightness standard deviation of the image, and when the contrast is lower than a preset threshold, improve the contrast by enhancing the brightness difference between pixels;
[0035] S32. Angle deviation correction:
[0036] S321. Feature point detection: Use a feature point detection algorithm based on deep learning to detect representative feature points in an image. These feature points have relatively stable feature patterns when shot at different angles.
[0037] S322. Angle estimation: Calculate the shooting angle deviation of the image through geometric relationships based on the detected feature points;
[0038] S323. Image rotation: According to the calculated angle deviation, the image is rotated and corrected so that the ship name area is close to the horizontal or vertical direction, thereby optimizing the display effect of the ship name area.
[0039] Preferably, the specific method of step S4 is: first, use the target detection algorithm to locate the outside of the ship name area; then fuse the context information of the image to improve the accuracy of detection and segmentation, specifically: by analyzing the context features of the ship name area, assist in determining the boundary and position of the ship name area; after obtaining the preliminary ship name area segmentation result, use post-processing technology to refine the segmented area and remove the inaccurate parts of the edge; at the same time, establish a priori model of the shape and size of the ship name area, and constrain and correct the detection results according to the common shapes and size ranges of ship name areas of different types of fishing vessels.
[0040] Preferably, the specific method of step S5 is: introducing a multi-scale feature fusion mechanism to fuse feature maps at different levels; based on the extracted high-dimensional feature vector, applying feature selection and dimensionality reduction techniques to optimize feature representation; using feature selection methods based on, but not limited to, information gain and chi-square test, to screen out the key features that contribute most to ship name recognition, remove redundant and irrelevant features, and reduce the dimension of the feature vector; at the same time, using principal component analysis and linear discriminant analysis dimensionality reduction algorithms to project the high-dimensional feature vector into a low-dimensional space, while maintaining the main feature information, further reducing the amount of data and computational complexity, and improving the efficiency and accuracy of subsequent recognition algorithms; in the dimensionality reduction process, by including but not limited to cross-validation technology to select the best dimensionality reduction parameters, to ensure that the reduced dimensionality features have good performance in different data sets and task scenarios.
[0041] Preferably, the specific method of step S6 is: regarding the parameter adjustment of the recognition system as an action, and the recognition performance index as a reward signal, the system selects the best parameter adjustment strategy through a reinforcement learning algorithm according to the current recognition results and performance indicators; when the recognition accuracy is low, try to adjust the relevant parameters of image preprocessing, feature extraction, and character recognition, observe the impact of the adjusted recognition results on the reward signal, and continuously explore and optimize the parameter space.
[0042] Preferably, the specific method of character recognition in step S7 is: first, using the OCR engine to perform preliminary recognition on the characters in the image to obtain an initial character recognition result; then, the image is input into the convolutional recurrent neural network for re-recognition, and the convolutional recurrent neural network uses its processing ability for sequence data to learn the contextual relationship between characters, improve the recognition ability of blurred and deformed characters, and output a recognition result; finally, the recognition results of the OCR engine and the convolutional recurrent neural network are fused, and the final character recognition result is determined by a voting mechanism or a weighted fusion method based on confidence;
[0043] The specific method of post-processing in step S7 includes:
[0044] Character correction: According to the grammatical rules of fishing boat names and common character combination patterns, the recognition results are checked for syntax and semantics. For suspected erroneous characters, corrections are made based on the character shape, context, and similarity with the standard character library.
[0045] De-noising: remove noise characters in the recognition results. Specifically, a character dictionary can be established to treat low-frequency characters or obviously wrong characters in the recognition results that are not in the dictionary as noise and delete them. At the same time, the recognition results can be further optimized by combining the character position information, the relationship between adjacent characters and the context information in the image. Specifically, if the distance between a certain character and the surrounding characters is abnormal or does not conform to the common format of the ship name, it can be corrected or supplemented according to the context information;
[0046] Establish a character recognition error case library: used to analyze and summarize common error types, continuously optimize post-processing algorithms, and improve the accuracy and reliability of character recognition.
[0047] Preferably, the correction and verification method in step S8 includes automatic correction and manual verification, and when the automatic correction cannot determine the accuracy of the recognition result or encounters complex situations, it is submitted to manual verification;
[0048] The automatic correction process is specifically as follows: using image comparison technology, the image of the identified ship name area is compared with the image in the standard ship name template library; by calculating the similarity between the images, the accuracy of the recognition result is determined; if the similarity is lower than a preset threshold, the system automatically starts the correction algorithm; the correction algorithm corrects the recognition result according to the difference characteristics of the image, using technologies including but not limited to image deformation and character replacement; if a character has a large difference in shape from a standard character in the template library, the image deformation technology is used to try to adjust it to a shape closer to the standard, and then the comparison and verification are performed again until satisfactory accuracy is achieved;
[0049] The manual verification process is specifically as follows: first, the automatically identified ship name is reviewed as a whole to determine whether it conforms to common sense and the ship name format in fishery management regulations; then, the ship name area in the image is carefully compared with the recognition result, with a focus on checking the characters and areas marked as suspicious; for uncertain characters, the verifier refers to the knowledge base, similar ship name cases or consults fishery experts to make a judgment, and during the verification process, the verifier records his or her own judgment basis and operation process for subsequent tracing and analysis; if the verifier believes that the automatic recognition result is correct, it is confirmed; if an error is found, the recognition result is manually modified, and the reason for the modification is noted in the system.
[0050] A fishing vessel name recognition system based on image processing and deep learning, characterized by: comprising a fishing vessel image acquisition module, an image preprocessing and enhancement module, an image quality judgment module, an environmental condition detection and analysis module, an adaptive image enhancement module, a vessel name area detection and segmentation module, a feature extraction and optimization module, a recognition result feedback and optimization module, a character recognition and post-processing module, a correction and verification module, a recognition log and report generation module, a result display and storage module, and a data transmission and synchronization module;
[0051] The fishing boat image acquisition module includes multiple high-definition cameras set up around the dock to achieve all-round and multi-angle real-time monitoring of the fishing boat to ensure the clearest and most complete fishing boat image;
[0052] The image preprocessing and enhancement module is based on an image preprocessing model, which automatically adjusts parameters including but not limited to noise reduction, deblurring and brightness adjustment by learning a large number of fishing boat images under different environmental conditions;
[0053] The image quality judgment module is based on an image quality assessment model, the training data of which includes a large number of fishing boat images of different quality levels and corresponding manually annotated quality scores. During the training process, the image quality assessment model learns the mapping relationship between various visual features of the image and the quality score, and uses a variety of data enhancement techniques to expand the training data and a regularization method to prevent overfitting.
[0054] The environmental condition detection and analysis module integrates multiple environmental sensors and is based on a machine learning model. The machine learning model uses historical environmental data and corresponding image quality data as training samples, and the training model learns the relationship between environmental changes and image quality;
[0055] The adaptive image enhancement module is based on building a convolutional neural network model specifically for optimizing image enhancement parameters. The convolutional neural network model takes the original features of the image and environmental condition data as input and outputs the parameters required for image enhancement. The module learns how to generate the most appropriate enhancement parameters based on the input image and environmental information by training a large number of fishing boat images labeled with different environmental conditions and corresponding optimal image enhancement parameters.
[0056] The ship name region detection and segmentation module is based on the YOLOv8 model. The YOLOv8 model uses an image data set containing and annotating the fishing boat name region as a training set, and the annotation information includes the bounding box coordinates and category of the ship name region. During the training process, the YOLOv8 model calculates the error between the prediction result and the true label through the loss function according to the input image and annotation information, and updates the network parameters through the back propagation algorithm. During the training process, an optimization algorithm is used to accelerate convergence and improve model performance. During the training process, data enhancement technology is also used to increase the diversity of data and improve the model's detection and segmentation capabilities for ship name regions of different forms.
[0057] The feature extraction and optimization module is based on a deep convolutional neural network model. During the training process, the deep convolutional neural network model uses a combination of classification loss and regression loss to supervise the model's classification ability for ship name features and optimize the regression accuracy of feature vectors for information including but not limited to the shape and position of the ship name area;
[0058] The recognition result feedback and optimization module is based on an online learning model, which can learn new sample data in real time. As new fishing boat image data is continuously input, the system automatically incorporates it into the training process and updates the parameters of the online learning model. Specifically, the incremental learning technology is used to avoid retraining all data. At the same time, the performance of the online learning model is regularly evaluated. When it is found that the performance of the online learning model has declined or a new image feature pattern has appeared, the retraining or fine-tuning operation of the online learning model is triggered.
[0059] The character recognition and post-processing module is based on a character recognition model, which is based on a convolutional recurrent neural network. The character recognition model uses a large number of character sample images and annotates the corresponding character categories as a training set. The training process calculates the error between the prediction result and the true label through a loss function, updates the network parameters through back propagation, uses an optimization algorithm for optimization, and improves the generalization ability of the model through data enhancement technology.
[0060] The correction and verification module is based on image comparison and intelligent correction algorithms and a knowledge base that provides auxiliary decision-making suggestions for manual verification;
[0061] The identification log and report generation module is used to achieve detailed log recording and real-time monitoring as well as automated report generation and data analysis support;
[0062] The result display and storage module is used to realize result display and data storage;
[0063] The data transmission and synchronization module adopts data packetization and reorganization technology to realize data transmission between modules, and uses a high-precision time synchronization protocol to ensure time synchronization accuracy.
[0064] Due to the adoption of the above technical scheme, the technical progress achieved by the present invention is as follows.
[0065] The present invention uses multi-source image acquisition and data synchronization technology to ensure the time consistency of multi-source data, provide stable and continuous input data for subsequent analysis, and avoid information loss or dislocation caused by data asynchrony.
[0066] The present invention uses image preprocessing and environmental adaptive enhancement technology to enhance image quality, adapt it to changing environmental conditions, and ensure that clear image data that meets analysis requirements is obtained even in complex environments.
[0067] The present invention uses deep learning-driven key area detection and segmentation technology to achieve precise positioning and segmentation of key areas of ships, improve the accuracy of area detection, and provide high-quality input for subsequent feature extraction.
[0068] The present invention combines OCR with deep learning character recognition technology: it can still accurately recognize characters under complex backgrounds and low-definition images, has strong anti-interference ability, and ensures high accuracy of recognition results.
[0069] The present invention uses a closed-loop feedback and self-learning optimization mechanism: the recognition model and parameters can be continuously optimized, so that the system has self-learning capabilities, and the recognition effect is gradually improved with use.
[0070] The present invention adopts a collaborative mechanism of correction and manual verification: it can combine automatic correction and manual verification to achieve high-precision guarantee, is suitable for scenarios with high-precision requirements, and ensures the reliability of output results.
[0071] The present invention uses automatic report generation and historical data management technology: it can automatically generate identification reports, facilitate users to trace and analyze data, improve the practicality of the system, and provide data support for model optimization.
[0072] The present invention can be independently developed and maintained through system modules, which is convenient for function expansion and update iteration, and improves the flexibility and scalability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION
[0074] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0075] A fishing vessel name recognition method based on image processing and deep learning, combined with Figure 1 As shown, the following steps are included:
[0076] S1. Fishing boat image collection.
[0077] Multi-camera layout and intelligent switching: Carefully plan the multi-camera layout at key locations such as ports and waterways to ensure coverage without blind spots. For example, a combination of circular and linear layouts is used to set up multiple high-definition cameras around the dock, and cameras are installed at certain intervals along the waterway to achieve all-round, multi-angle real-time monitoring of fishing boats. The camera is equipped with an automatic switching function, which intelligently selects the camera with the best viewing angle for image acquisition based on the location and movement trajectory of the target fishing boat, ensuring the clearest and most complete image of the fishing boat.
[0078] HD and UHD image acquisition: HD or UHD cameras (such as 8K resolution) with high pixel density and excellent imaging performance are used to capture the tiny features of the fishing boat's name and hull details. At the same time, the camera has autofocus and autoexposure functions, which can quickly adjust the focus and exposure parameters according to the target distance and ambient light conditions to ensure clear and accurate images in different scenes. For example, it automatically reduces the exposure in strong light to prevent overexposure, and increases the exposure time and optimizes the sensitivity in low light conditions to obtain high-quality images.
[0079] Real-time processing and intelligent analysis of video streams: Deploy real-time video analysis algorithms on the camera side or edge computing devices to perform preliminary processing on the collected video streams. For example, use the target detection algorithm to quickly identify the fishing boat in the picture and extract key information such as its position and speed. Based on this information, dynamically adjust the camera parameters (such as zoom factor, shooting direction, etc.) to maintain stable tracking and clear imaging of the fishing boat. At the same time, pre-compress the video stream to reduce the data transmission volume and improve transmission efficiency while ensuring image quality.
[0080] S2. Image preprocessing and enhancement to achieve image noise reduction, deblurring and brightness adjustment, and after processing, image quality judgment is performed to form a closed-loop optimization.
[0081] For noise reduction, according to the distribution characteristics of different noise types (such as Gaussian noise, salt and pepper noise, etc.) in different environments, appropriate filtering algorithms are applied for noise reduction; for deblurring, according to the image blur degree and blur type (such as motion blur, defocus blur, etc.), appropriate deblurring algorithms (such as Wiener filtering, blind deconvolution, etc.) are selected and parameters are adjusted; in terms of brightness adjustment, according to the overall brightness distribution of the image and the brightness requirements of the target area (ship name area), the brightness adjustment coefficient is automatically calculated to ensure that the ship name area has appropriate brightness contrast.
[0082] Specifically, the brightness adjustment method is:
[0083] The convolutional neural network (CNN) is used to extract features from the input fishing boat image. The network structure of CNN can be designed as multiple convolution layers, pooling layers and fully connected layers. For example, the convolution layer uses a small 3x3 convolution kernel with a step size of 1. Different numbers of convolution kernels (such as 32, 64, etc.) are used to gradually extract different levels of features of the image. The pooling layer uses a 2x2 maximum pooling operation for downsampling. The extracted features are input into a fully connected layer, and the output of the fully connected layer is used as the basis for adjusting the image enhancement parameters. Specifically, the output of the fully connected layer can be connected to two branches, one for generating a brightness adjustment coefficient and the other for generating a contrast adjustment coefficient. For example, a sigmoid activation function is used to map the output value to between 0 and 1 as a scaling factor for adjusting brightness and contrast. According to the calculated brightness and contrast adjustment coefficients, the brightness and contrast of the image are adjusted pixel by pixel. For brightness adjustment, the RGB value of each pixel is multiplied by the brightness adjustment coefficient; for contrast adjustment, the average brightness of the image is first calculated, and then the difference between each pixel and the average brightness is adjusted according to the contrast adjustment coefficient, thereby enhancing the contrast of the image.
[0084] The method for judging image quality is as follows: using a model that learns the mapping relationship between various visual features of an image and a quality score to judge the processed image and output a quality score, and judging whether the image is suitable for character recognition based on a preset threshold; at the same time, the image quality indicators such as clarity, contrast, and noise level are evaluated based on a deep learning algorithm, and based on the evaluation results, optimization suggestions for image preprocessing and enhancement algorithm parameters are provided, so that images with unqualified quality indicators can be preprocessed again to form a closed-loop optimization system.
[0085] S3. Adaptive image enhancement to achieve illumination compensation and image dynamics enhancement.
[0086] Multi-sensor data fusion: Integrate multiple environmental sensors, such as light sensors, meteorological sensors (to measure temperature, humidity, wind speed, etc.) and color sensors (to detect the color characteristics of the hull paint). The data collected by these sensors are comprehensively analyzed through the data fusion algorithm to more comprehensively and accurately understand the environmental conditions of the fishing boat. For example, the Kalman filter algorithm is used to fuse the data of the light sensor and the meteorological sensor to predict the trend of changes in light intensity and provide a reference for the image enhancement algorithm in advance. For hull paint color detection, the color sensor collects the color information of the hull surface, and analyzes the reflection and absorption characteristics of the color in combination with the lighting conditions, providing a more accurate basis for color correction for subsequent image enhancement and ship name recognition.
[0087] The method of illumination compensation is as follows: first, for different types of fishing boat images (normal illumination images taken during the day, low-light images at night, images of different hull colors and paint conditions, etc.), a multi-modal image enhancement strategy is adopted, specifically, the deep learning model is used to classify the images, determine the modality type to which they belong, and then the corresponding image enhancement algorithm combination is applied according to the characteristics of different modalities. For low-light images at night, a denoising algorithm based on deep learning is first used to remove noise, and then an adaptive illumination compensation algorithm is applied to increase the brightness, and then the contrast enhancement algorithm is used to highlight the ship name area; for images of different hull colors, the color correction parameters are adjusted according to the color information detected by the color sensor to ensure that the color of the ship name has a good contrast with the background color.
[0088] The adaptive illumination compensation algorithm is a corresponding compensation strategy generated by predicting the impact of environmental data on image quality based on real-time collected environmental data. Specifically, the image is first analyzed for illumination and the illumination histogram of the image is calculated. Specifically, the image is divided into multiple sub-regions (such as 8x8 or 16x16 small blocks), and the brightness distribution of pixels in each sub-region is counted to obtain the illumination histogram; then, according to the distribution of the illumination histogram, the area with uneven illumination in the image is judged. If the brightness value of a sub-region is significantly lower or higher than other regions, it is considered that the area has insufficient or excessive illumination problems; for insufficiently illuminated areas, compensation is performed by increasing the pixel value. Specifically, an interpolation algorithm based on neighboring pixels can be used. For example, taking the pixel in the insufficiently illuminated area as the center, the average value or weighted average value of the surrounding neighboring pixels is taken, and the average value is appropriately increased according to the degree of insufficient illumination and then assigned to the central pixel, thereby improving the brightness of the area; for areas with excessively strong illumination, compensation is performed by reducing the pixel value. The method is similar, but the adjustment direction is opposite.
[0089] The method for dynamic image enhancement specifically comprises the following steps:
[0090] S31.Brightness and contrast adjustment.
[0091] S311. Perform brightness adjustment: Calculate the average brightness value of the image (add the brightness values of all pixels in the image and divide by the total number of pixels). If the average brightness value is lower than the preset lower limit of the appropriate brightness range, increase the brightness value of each pixel by a certain ratio (such as 10% each time); if it is higher than the upper limit, reduce it proportionally. This ratio can be determined based on actual tests and experience to achieve an appropriate brightness effect.
[0092] S312. Perform contrast adjustment: Calculate the brightness standard deviation of the image. The smaller the standard deviation, the lower the contrast. When the contrast is lower than a preset threshold, the contrast is improved by enhancing the brightness difference between pixels. One method is to multiply the difference between the brightness value of each pixel and the average brightness value by a contrast enhancement coefficient (such as 1.5 or 2), and then add the average brightness value to obtain the adjusted pixel brightness value.
[0093] S32. Angle deviation correction:
[0094] S321. Feature point detection: Use a deep learning-based feature point detection algorithm (such as a trainable convolutional neural network feature point detector) to detect representative feature points in the image, such as corner points and edge points in the ship name area. These feature points have relatively stable feature patterns when shot at different angles.
[0095] S322. Angle estimation: Calculate the shooting angle deviation of the image based on the detected feature points through geometric relationships. For example, assume that the ship name area in the image is a rectangle, determine the four sides of the rectangle based on the feature points, and estimate the shooting angle deviation by calculating the angle between the rectangle side and a preset standard direction (such as horizontal or vertical direction).
[0096] S323. Image rotation: according to the calculated angle deviation, the image is rotated and corrected. An image rotation algorithm, such as a bilinear interpolation algorithm, can be used to rotate the image around the image center by a corresponding angle so that the ship name area is as close to the horizontal or vertical direction as possible, thereby optimizing the display effect of the ship name area.
[0097] This step fully considers the influence of shooting angle and lighting conditions on ship name recognition, thereby avoiding the ship name in the image being blurred due to insufficient lighting or undesirable shooting angle, which affects the final recognition effect.
[0098] S4. Ship name region detection and segmentation.
[0099] First, the target detection algorithm is used to locate the area outside the ship name; then the context information of the image is integrated to improve the accuracy of detection and segmentation. Specifically, the boundary and position of the ship name area are determined by analyzing the context features such as the hull structure and logo pattern around the ship name area. After obtaining the preliminary segmentation results of the ship name area, post-processing techniques such as region growing and edge optimization are used to refine the segmented area and remove the inaccurate parts of the edge, making the segmentation of the ship name area more accurate and complete. At the same time, a priori model of the shape and size of the ship name area is established. According to the common shapes and size ranges of the ship name areas of different types of fishing vessels, the detection results are constrained and corrected to avoid unreasonable segmentation results.
[0100] S5. Extraction and optimization of ship name area features. The specific method is as follows:
[0101] Deep feature learning and multi-scale feature fusion: A deeper convolutional neural network (ResNet, DenseNet) is used to extract features from the ship name area to obtain a more representative and discriminative feature vector. In the network structure, a multi-scale feature fusion mechanism is introduced to fuse feature maps at different levels, so that the local detail features and overall structural features of the ship name area can be captured at the same time. For example, through upsampling and splicing operations, the high-resolution local features extracted by the shallow network are fused with the low-resolution global features extracted by the deep network to form a feature representation containing rich information. During the training process, a combination of classification loss and regression loss is used to supervise the model's classification ability for ship name features (distinguishing ship names of different fonts and styles), and to optimize the regression accuracy of the feature vector for the shape, position and other information of the ship name area.
[0102] Feature selection and dimensionality reduction optimization: Based on the extracted high-dimensional feature vectors, feature selection and dimensionality reduction techniques are applied to optimize feature representation. Feature selection methods based on information gain, chi-square test, etc. are used to screen out the key features that contribute most to ship name recognition, remove redundant and irrelevant features, and reduce the dimension of feature vectors. At the same time, principal component analysis (PCA) and linear discriminant analysis (LDA) dimensionality reduction algorithms are used to project high-dimensional feature vectors into low-dimensional space. While maintaining the main feature information, the amount of data and computational complexity are further reduced, thereby improving the efficiency and accuracy of subsequent recognition algorithms. In the dimensionality reduction process, the best dimensionality reduction parameters are selected through cross-validation and other techniques to ensure that the reduced dimensionality features have good performance in different data sets and task scenarios.
[0103] S6. Feedback and optimization of recognition results.
[0104] Reinforcement learning algorithms (Q learning or deep Q network (DQN)) are used to implement feedback and optimization of recognition results. The specific method is: the parameter adjustment of the recognition system is regarded as an action, and the performance indicators such as recognition accuracy are regarded as reward signals. The system selects the best parameter adjustment strategy through the reinforcement learning algorithm based on the current recognition results and performance indicators. When the recognition accuracy is low, try to adjust the relevant parameters in the image preprocessing, feature extraction, character recognition and other modules (increase the intensity of image enhancement, adjust the weight of the feature extraction network, etc.), observe the impact of the adjusted recognition results on the reward signal, and continuously explore and optimize the parameter space so that the system gradually learns the optimal parameter configuration and improves the recognition accuracy in the future.
[0105] S7. Character recognition and post-processing.
[0106] The specific method of deeply integrating traditional OCR technology with deep learning model (convolutional recurrent neural network (CRNN)) for character recognition and post-processing is as follows:
[0107] The specific method of character recognition is as follows: first, use the OCR engine to perform preliminary recognition on the characters in the image to obtain an initial character recognition result; then, input the image into the convolutional recurrent neural network for re-recognition. The convolutional recurrent neural network uses its ability to process sequence data to learn the contextual relationship between characters, improve the recognition ability of blurred and deformed characters, and output a recognition result; finally, the recognition results of the OCR engine and the convolutional recurrent neural network are fused, and the final character recognition result is determined through a voting mechanism or a confidence-based weighted fusion method.
[0108] Post-processing is based on error correction algorithms, including:
[0109] Character correction: According to the grammatical rules of the fishing boat names and common character combination patterns, the recognition results are checked for grammar and semantics. For example, check whether the boat name conforms to common naming specifications, whether it contains illegal characters or illogical character combinations, and correct suspected erroneous characters based on the character shape, context, and similarity with the standard character library. If a character sequence that does not conform to grammatical rules appears in the recognition result (such as two consecutive identical punctuation marks), it will be corrected according to the context and language habits.
[0110] De-noising: remove noise characters in the recognition results (such as incorrectly recognized characters due to image quality issues). Specifically, a character dictionary can be established to treat low-frequency characters or obviously incorrect characters (such as garbled characters) in the recognition results that are not in the dictionary as noise and delete them. At the same time, the recognition results can be further optimized by combining the character position information, the relationship between adjacent characters and the context information in the image. Specifically, if the distance between a certain character and the surrounding characters is abnormal or does not conform to the common format of the ship name, it can be corrected or supplemented according to the context information.
[0111] Establish a character recognition error case library: used to analyze and summarize common error types, continuously optimize post-processing algorithms, and improve the accuracy and reliability of character recognition.
[0112] S8. Calibration and verification.
[0113] The correction and verification methods include automatic correction and manual verification. When automatic correction cannot determine the accuracy of the recognition result or encounters complex situations, it is submitted to manual verification.
[0114] The specific process of automatic correction is as follows: using image comparison technology, the image of the identified ship name area is compared with the image in the standard ship name template library; by calculating the similarity between the images (based on methods such as structural similarity index (SSIM) or cosine similarity), the accuracy of the recognition result is judged; if the similarity is lower than the preset threshold, the system automatically starts the correction algorithm; the correction algorithm corrects the recognition result according to the difference characteristics of the image (such as character shape, position, color, etc.), including but not limited to image deformation and character replacement technology; if a character has a large difference in shape from the standard character in the template library, try to adjust it to be closer to the standard shape through image deformation technology, and then compare and verify again until satisfactory accuracy is achieved;
[0115] The specific process of manual verification is as follows: First, conduct an overall review of the automatically identified ship name to determine whether it conforms to common sense and the ship name format in fishery management regulations; then, carefully compare the ship name area in the image with the recognition result, and focus on checking the characters and areas marked as suspicious; for uncertain characters, the verifier refers to the knowledge base, similar ship name cases or consults fishery experts for judgment, and during the verification process, the verifier records his or her judgment basis and operation process for subsequent tracing and analysis; if the verifier believes that the automatic recognition result is correct, it will be confirmed; if an error is found, the recognition result will be manually modified, and the reason for the modification will be noted in the system. In order to facilitate manual verification, the following design is also provided:
[0116] Friendly verification interface design: Develop a special manual verification interface to present the recognition results to the verifier in an intuitive and clear way. The interface displays key data such as the fishing boat image, the automatically recognized ship name, relevant confidence information, image acquisition time and location, etc. At the same time, it provides image zooming in, zooming out, rotating and other operation functions to facilitate the verifier to carefully check the details of the ship name area. For suspected erroneous areas or characters, mark them in a striking way, such as marking the ship name area with a red frame, and highlighting suspicious characters with underscores or different colors, to guide the verifier to focus on these parts.
[0117] Detailed verification data records: During the manual verification process, the system records the verification personnel's operation behavior and decision results in detail, including verification time, verification personnel ID, judgment on each ship name (confirmed correct, modified or marked as difficult), the content of the ship name before and after modification, the reason for modification, and whether the knowledge base suggestions were referred to. These records form a rich verification data set, which provides an important basis for subsequent analysis and optimization.
[0118] Feedback and optimization loop: Regularly conduct statistical analysis on manual verification data to summarize the types, scenarios and reasons for the errors that are prone to occur in the automatic recognition system. Based on the analysis results, optimize the automatic correction algorithm, adjust the training strategy of the character recognition model, update the knowledge base or improve the image acquisition and preprocessing methods. For example, if it is found that the ship name of a certain font is often misrecognized under specific lighting conditions, the system can optimize the lighting compensation algorithm and character recognition model for this situation to improve the recognition accuracy of the system in this scenario. Through this cyclic process from automatic correction to manual verification to feedback optimization, the overall performance and reliability of the system are continuously improved to ensure the high accuracy and credibility of the fishing vessel name recognition results.
[0119] S9. Identification log and report generation.
[0120] Detailed log records and real-time monitoring: Design a comprehensive log management system to record every operation step, parameter setting, recognition result and related environmental information in the process of fishing vessel name recognition. Log records use a standardized format, including timestamp, event description, data source, processing result and other detailed information, which is convenient for subsequent query and analysis. At the same time, establish a real-time monitoring mechanism to monitor and record key indicators (such as recognition accuracy, processing time, data transmission rate, etc.) during the operation of the system in real time. Display the dynamic changes of these indicators in a visual way (such as dashboards, charts, etc.), so that system administrators can timely understand the operation status of the system, find potential problems and take timely measures to optimize them.
[0121] Automated report generation and data analysis support: Use the report generation tool (report library in Python) to automatically generate regular recognition reports based on log records and system performance indicator data. The report content includes the system's operation overview within a certain period of time (the number of fishing vessel images processed, the trend of changes in recognition accuracy, etc.), detailed recognition result statistics (recognition success rate of different types of fishing vessel names, distribution of error types, etc.) and system performance analysis (processing time share of each module, resource utilization, etc.). These reports not only provide system administrators with a comprehensive summary of the system's operation status, but also provide strong support for further data analysis and decision-making. Through in-depth analysis of the report data, the bottlenecks and problems of the system can be discovered, providing a basis for the optimization and upgrade of the system. It also helps to summarize experience and continuously improve the performance and accuracy of the fishing vessel name recognition system.
[0122] S10. Result display and storage.
[0123] Visual interface optimization and interactive design: Design an intuitive and friendly visual interface to display the results of fishing vessel name recognition. The interface adopts a concise and clear layout, presenting key information such as fishing vessel images, recognized vessel names, relevant confidence information, image acquisition time and location to users in a clear and easy-to-understand manner. At the same time, it provides interactive functions, and users can mark, annotate and review the recognition results on the interface to facilitate further confirmation and processing of the recognition results. Users can click on the vessel name area to view detailed recognition process information, such as feature extraction results, character recognition confidence distribution, etc., to help users better understand and trust the recognition results.
[0124] Efficient database management and data storage strategy: Use efficient database management systems (MySQL, PostgreSQL, etc.) to store recognition results and related data. Design a reasonable data storage structure, classify and store fishing vessel image data, recognition results, environmental condition data, operation logs and other information, and establish indexes to improve data query and retrieval efficiency. In the data storage process, use data compression technology and incremental storage strategies to reduce storage space and increase data storage speed. At the same time, regularly back up and optimize the database to ensure data security and integrity, and provide reliable data support for subsequent data analysis and system performance evaluation.
[0125] Through the above steps, the method of the present invention realizes a closed-loop feedback optimization mechanism as a whole, which is as follows:
[0126] Data collection and preprocessing
[0127] Comprehensive data collection: During the operation of the fishing vessel name recognition system, various data are collected, including the fishing vessel image data obtained by the image acquisition module, the lighting, meteorological and other environmental data recorded by the environmental condition detection and analysis module, the operating parameters of each processing module (such as the noise reduction coefficient in the image preprocessing algorithm, the brightness adjustment value of the adaptive image enhancement module, etc.) and the final recognition results (including the recognized ship name, confidence, recognition time, etc.). These data constitute the basic information source for closed-loop feedback optimization and fully reflect the operation status of the system under different conditions.
[0128] Data cleaning and labeling: Clean the collected data to remove invalid data (such as incomplete data caused by acquisition errors and transmission interruptions) and abnormal data (such as light intensity values that are obviously deviated from the normal range, incorrectly recognized ship names that do not conform to grammatical rules, etc.). At the same time, label the image data and recognition results, including the true value of the ship name in the image (for comparing the accuracy of the recognition results), the quality level of the image (based on indicators such as clarity and contrast), the category of environmental conditions (such as sunny, cloudy, daytime, night, etc.), and the initial parameter settings of the processing module when processing the image. The labeled data set will be used for subsequent model training and optimization.
[0129] Feedback generation and analysis
[0130] Calculation of recognition result evaluation indicators: Based on the annotated real ship names and the system recognition results, a series of evaluation indicators are calculated, such as accuracy (the ratio of correctly recognized ship names to the total number of recognized ship names), recall (the ratio of correctly recognized ship names to the actual number of existing ship names), F1 value (an indicator that comprehensively considers accuracy and recall), and average recognition time, etc. These indicators intuitively reflect the current recognition performance of the system and provide a quantitative basis for feedback analysis.
[0131] Error type analysis and classification: In-depth analysis of the types of recognition errors, classifying the errors into character recognition errors (such as single character misrecognition, missing or redundant characters, etc.), regional segmentation errors (such as inaccurate positioning of the ship name area, incomplete segmentation or too much background), errors caused by image quality problems (such as insufficient lighting, blurred images, etc., making it difficult to clearly identify the ship name), and other types of errors (such as system failures, algorithm anomalies, etc.). By classifying and counting the error types, we can understand the problems that may exist in the system at different stages and provide directions for targeted optimization.
[0132] Analysis of correlation between environment and parameters: Combine environmental data with the operating parameters of the processing module to analyze the correlation between environmental factors and recognition performance, as well as the impact of different parameter settings on recognition results. For example, study the relationship between light intensity and image preprocessing effects, determine the optimal noise reduction and brightness adjustment parameters under different lighting conditions; analyze the correlation between feature extraction network parameters (such as convolution kernel size, network depth, etc.) and the recognition accuracy of ship names in different fonts, and find out the most suitable parameter configuration for various ship name features. Through this correlation analysis, we can explore the potential influencing factors of system performance and provide a basis for parameter optimization.
[0133] Optimize strategy formulation and execution
[0134] Model-based parameter adjustment strategy: Based on the feedback information analysis results, an optimization model is constructed to determine the parameter adjustment strategy of the processing module. For example, a machine learning algorithm (such as support vector machine regression, neural network, etc.) is used to establish a prediction model between recognition performance and parameters, and the current evaluation indicators, environmental data and error types are used as input to output the optimal parameter adjustment value. For the image preprocessing module, if it is found that the image clarity is low under certain lighting conditions, the model may suggest increasing the number of iterations of the deblurring algorithm or adjusting the brightness adjustment parameters; for the feature extraction module, according to the recognition accuracy of different ship name fonts and styles, the network structure parameters (such as increasing or decreasing the number of convolution layers, adjusting the convolution kernel size, etc.) are adjusted to improve the feature extraction effect.
[0135] Algorithm selection and update strategy: Based on error type analysis, if it is found that a certain algorithm currently used frequently makes errors when processing specific types of ship names or environmental conditions, consider replacing or updating the algorithm. For example, if the traditional OCR algorithm has a low accuracy rate when recognizing certain artistic font ship names, the system can switch to a character recognition algorithm based on deep learning (such as convolutional recurrent neural network), or improve and optimize the existing algorithm. At the same time, pay attention to the latest research results in related fields, and introduce new and better-performing algorithms into the system in a timely manner to maintain the advancement and adaptability of the system.
[0136] Data enhancement and expansion strategy: If the analysis finds that the system has poor recognition performance in certain special situations (such as low light, long-distance shooting, etc.), it may be due to the lack of corresponding samples in the training data. At this time, formulate a data enhancement and expansion strategy to expand the training data set by simulating or actually collecting more representative sample data. For example, use image synthesis technology to generate images of fishing boats under different light intensities, angles, and hull color combinations, as well as simulate images of various blurry and noisy conditions, so that the training data can more comprehensively cover various scenarios in actual applications and improve the system's ability to handle complex situations.
[0137] Model retraining and system updates
[0138] Incremental model training: Adopting incremental learning methods, relevant models in the system (such as image preprocessing model, feature extraction model, character recognition model, etc.) are retrained using new annotated data and optimized parameters. Unlike traditional batch training methods, incremental training only updates new data and parameter adjustment parts, avoiding repeated training of the entire data set, greatly improving training efficiency, and being able to timely integrate optimized knowledge into the model, allowing the model to quickly adapt to new situations and needs.
[0139] System integration and verification: Integrate the retrained model into the fishing vessel name recognition system and verify it in the actual operating environment. During the verification process, closely monitor the various performance indicators of the system to ensure that the optimized system has improved in terms of accuracy, stability, and efficiency. If new problems are found or the performance does not meet expectations, return to the feedback information analysis stage and continue to adjust the optimization strategy until the system performance meets the requirements. Through this closed-loop feedback optimization process, the system can continuously self-learn and evolve, continuously improve the accuracy and reliability of fishing vessel name recognition, and adapt to the ever-changing actual application scenarios.
[0140] A fishing boat name recognition system based on image processing and deep learning includes a fishing boat image acquisition module, an image preprocessing and enhancement module, an image quality judgment module, an environmental condition detection and analysis module, an adaptive image enhancement module, a boat name area detection and segmentation module, a feature extraction and optimization module, a recognition result feedback and optimization module, a character recognition and post-processing module, a correction and verification module, a recognition log and report generation module, a result display and storage module, and a data transmission and synchronization module. The following is a detailed description of each module:
[0141] The fishing vessel image acquisition module includes multiple high-definition cameras set up around the dock to achieve all-round, multi-angle real-time monitoring of fishing vessels, ensuring the clearest and most complete images of fishing vessels.
[0142] The image preprocessing and enhancement module is based on the image preprocessing model, which is built based on deep learning. The image preprocessing model automatically adjusts parameters such as noise reduction, deblurring and brightness adjustment by learning a large number of fishing boat images under different environmental conditions. For example, using the conditional generation model in the generative adversarial network (GAN), environmental factors such as light intensity and weather conditions are used as conditional inputs, and the model outputs the optimal image processing parameters for the environment.
[0143] The image quality judgment module is based on the image quality assessment model, which is built on a convolutional neural network (CNN). The training data of the image quality assessment model includes a large number of fishing boat images of different quality levels and the corresponding manually annotated quality scores. During the training process, the image quality assessment model learns the mapping relationship between various visual features of the image and the quality score. In order to improve the generalization ability of the image quality assessment model, a variety of data enhancement techniques are used to expand the training data, and regularization methods (L1 and L2 regularization) are used to prevent overfitting.
[0144] In addition to a single quality score, the image quality assessment model also performs detailed analysis and assessment of multiple quality indicators of the image. The image clarity index (gradient-based clarity metric), contrast index (contrast histogram statistics), noise index (noise power spectrum estimation), etc. are calculated separately, and these indicators are presented to the user in a visual way to help the user understand the image quality status more comprehensively. At the same time, according to the evaluation results of different quality indicators, targeted optimization suggestions are provided for the image preprocessing and enhancement modules, and the parameters of the preprocessing algorithm are adjusted in time to form a closed-loop optimization system. If the image clarity is low, increase the intensity of the deblurring algorithm and process it again until the image quality reaches the preset standard; if the contrast is insufficient, it is prompted to adjust the contrast enhancement parameters, etc., and process it again until the image quality reaches the preset standard, realizing the coordinated optimization between image quality assessment and image processing.
[0145] The environmental condition detection and analysis module integrates a variety of environmental sensors and builds a machine learning model (such as support vector regression (SVR) or neural network) based on the machine learning model. The historical environmental data (including light, weather, time, etc.) and the corresponding image quality data are used as training samples to train the model to learn the relationship between environmental changes and image quality. Using the trained model, the impact of environmental data collected in real time is predicted on image quality, and the corresponding compensation strategy is generated. For example, when it is predicted that the light intensity is about to change, the light compensation parameters in the image preprocessing and enhancement module are adjusted in advance to ensure the stability of image quality.
[0146] The adaptive image enhancement module is based on building a convolutional neural network model specifically for optimizing image enhancement parameters. The convolutional neural network model takes the original features of the image (brightness histogram, gradient information, etc.) and environmental condition data (light intensity, hull paint color characteristics, etc.) as input, and outputs the parameters required for image enhancement (contrast enhancement coefficient, brightness adjustment value, etc.); and by training a large number of fishing boat images annotated with different environmental conditions and corresponding optimal image enhancement parameters, it learns how to generate the most appropriate enhancement parameters based on the input image and environmental information; in actual applications, the collected image is input into the model, the optimized image enhancement parameters are obtained in real time, and applied to image enhancement operations.
[0147] The ship name region detection and segmentation module is based on the YOLOv8 model. The YOLOv8 model increases the depth and width of the network to improve the model's ability to extract features of the ship name region under complex backgrounds. By introducing the attention mechanism, the model can pay more attention to the key features of the ship name region when processing images and reduce background interference. At the same time, using transfer learning technology, the model parameters pre-trained on other large-scale target detection datasets are migrated to the fishing boat name detection task to accelerate the convergence speed of the model and improve its generalization ability. The specific process of YOLOv8 model construction is as follows:
[0148] Network structure construction: The network structure of YOLOv8 includes a backbone network (for extracting image features), a neck network (for feature fusion and enhancement), and a head network (for predicting the location and category of the target). The backbone network can use the Darknet structure or other similar efficient convolutional neural network structures to extract deep features of the image through a series of convolutional layers, residual blocks and other components. The neck network fuses the different levels of features output by the backbone network to enhance the expressiveness of the features. The head network locates the target and predicts the category based on the fused features.
[0149] Dataset preparation: A large number of image datasets containing fishing boat name areas are collected and annotated. The annotation information includes the bounding box coordinates of the boat name area (upper left corner and lower right corner coordinates) and the category (boat name area). The dataset is divided into a training set and a validation set according to a certain ratio (e.g. 80% for training and 20% for validation).
[0150] Model training: Use the training set to train the YOLOv8 model. During the training process, the model calculates the error between the predicted result and the true label based on the input image and annotation information through the loss function (such as the cross entropy loss function for category prediction and the mean square error loss function for position prediction), and updates the network parameters through the back propagation algorithm. During the training process, some optimization algorithms (such as stochastic gradient descent (SGD) and its variants Adagrad, Adadelta, Adam, etc.) can be used to accelerate convergence and improve model performance. During the training process, data enhancement techniques (such as random cropping, flipping, scaling, etc.) can also be used to increase the diversity of data and improve the generalization ability of the model.
[0151] Model evaluation and optimization: Use the validation set to evaluate the trained model and calculate the model's accuracy, recall, mean average precision (mAP) and other indicators. Based on the evaluation results, if the model performance does not meet expectations, you can adjust the network structure parameters (such as increasing the network depth, adjusting the convolution kernel size, etc.), optimize the training parameters (such as adjusting the learning rate, number of iterations, etc.) or increase the size of the data set to further optimize the model until the model achieves better performance on the validation set.
[0152] Region segmentation: The trained and optimized YOLOv8 model can predict the input fishing boat image and output the bounding box coordinates of the ship name area. Based on these coordinates, the ship name area can be segmented from the original image to obtain an independent ship name area image, providing accurate input data for subsequent feature extraction and character recognition.
[0153] The feature extraction and optimization module is based on a deep convolutional neural network model. During the training process, the deep convolutional neural network model uses a combination of classification loss and regression loss to supervise the model's classification ability for ship name features and optimize the regression accuracy of the feature vector for information including but not limited to the shape and position of the ship name area.
[0154] The recognition result feedback and optimization module is based on the online learning model. By establishing an online learning mechanism, the system can learn new sample data in real time. As new fishing boat image data is continuously input, the system automatically incorporates it into the training process and updates the parameters of the model. Incremental learning technology is used to avoid retraining all data and improve learning efficiency. At the same time, the performance of the model is evaluated regularly. When the model performance is found to be degraded or new image feature patterns appear, the model retraining or fine-tuning operation is triggered to ensure that the model always maintains a good recognition ability for the names of different types of fishing vessels and adapts to the ever-changing actual application scenarios.
[0155] Character recognition and post-processing module: Based on the character recognition model, the character recognition model is based on a convolutional recurrent neural network. The character recognition model uses a large number of character sample images and annotates the corresponding character categories as training sets. The training process calculates the error between the predicted result and the true label through the loss function, back-propagates to update the network parameters, uses the optimization algorithm for optimization, and improves the generalization ability of the model through data enhancement technology. The specific process of character recognition model construction is as follows:
[0156] OCR engine selection and preprocessing:
[0157] Select a suitable OCR engine, such as Tesseract OCR. Before character recognition, preprocess the segmented ship name area image. Preprocessing includes grayscale processing (converting color images to grayscale images to reduce data volume and computational complexity) and binarization processing (converting grayscale images to black and white images to highlight character outlines for subsequent recognition). Binarization can use an adaptive threshold algorithm to automatically determine the threshold based on the local brightness information of the image, so that the characters are clearly separated from the background.
[0158] Deep learning model assisted recognition:
[0159] Build a character recognition model based on deep learning, such as a convolutional recurrent neural network (CRNN). The CRNN network structure includes convolutional layers (for extracting features of character images), recurrent layers (such as long short-term memory networks (LSTM) or gated recurrent units (GRU), for processing sequence information, i.e. the order of characters), and fully connected layers (for classification prediction).
[0160] Training data preparation: Collect a large number of character sample images (including ship name characters of various fonts, font sizes, and writing styles) and mark the corresponding character categories. Input the character sample images into the CRNN model for training. The training process is similar to the above-mentioned deep learning target detection model. The error between the predicted result and the true label is calculated through the loss function, and the network parameters are updated through back propagation. The optimization algorithm is used for optimization, and the generalization ability of the model is improved through data enhancement technology.
[0161] Combining OCR with deep learning models: The preprocessed ship name area image is first input into the OCR engine for preliminary recognition to obtain an initial character recognition result. Then the image is input into the trained CRNN model, which recognizes the characters again and outputs a recognition result. Finally, the recognition results of the OCR engine and the CRNN model are fused. A simple fusion method is to compare the characters in the same position in the two results. If they are consistent, the character is retained. If they are inconsistent, the character with higher confidence is selected as the final recognition result based on the confidence of the two models (which can be represented by the probability value output by the model). For example, if the confidence of the OCR engine in recognizing a character is 0.7 and the confidence of the CRNN model in recognizing the character is 0.8, the recognition result of the CRNN model is selected.
[0162] Correction and verification module: Based on image comparison and intelligent correction algorithms, and a knowledge base that provides auxiliary decision-making suggestions for manual verification, the construction and update of the knowledge base are as follows:
[0163] Establish an expert system knowledge base that includes fishery field knowledge, ship naming conventions, common error patterns and correction methods. The construction of the knowledge base is based on the analysis of a large amount of fishing vessel name data, the interpretation of fishery regulations, and the experience summary of industry experts. As the system runs and new data accumulates, the knowledge base is continuously updated to cover more ship name types and actual situations. For example, when encountering a new ship name format or special character combination, it is promptly incorporated into the knowledge base and the corresponding processing method is recorded.
[0164] During the manual verification process, when the verifier encounters difficult problems or uncertain situations, the expert system provides auxiliary decision-making and suggestions based on the information in the knowledge base. For example, if an uncommon ship name abbreviation or a character combination with special meaning is encountered, the expert system can query the knowledge base and provide possible explanations and correct recognition suggestions. At the same time, the expert system analyzes the possible types of problems based on image features, recognition results and the operation history of the verifier, and recommends corresponding verification methods and reference materials. For example, if there is a certain degree of occlusion in the ship name area in the image, the expert system will prompt the verifier to pay attention to the impact of the occluded part on character recognition, and provide empirical methods for dealing with ship name recognition under occlusion.
[0165] The identification log and report generation module is used to achieve detailed logging and real-time monitoring as well as automated report generation and data analysis support.
[0166] The result display and storage module is used to realize result display and data storage.
[0167] The data transmission and synchronization module uses data packetization and reorganization technology to achieve data transmission between modules, and uses a high-precision time synchronization protocol to ensure time synchronization accuracy. The details are as follows:
[0168] High-speed network architecture and protocol optimization: Build a high-speed and stable network architecture, such as using fiber-optic networks or high-speed wireless networks (such as 5G) to ensure the rapid transmission of image data and environmental condition data. At the same time, optimize the network transmission protocol and select a protocol suitable for real-time data transmission (such as the UDP protocol or a real-time transmission protocol optimized on the basis of the TCP protocol) to reduce transmission delays and data packet loss rates. During the data transmission process, use data segmentation and reassembly technology to divide large data packets into smaller data packets for transmission, and accurately reassemble them at the receiving end to improve the reliability and efficiency of data transmission.
[0169] Precise time synchronization mechanism: Use high-precision time synchronization protocols (such as NTP or PTP protocols) to ensure that the time synchronization accuracy between all devices (including cameras, sensors, processing units, etc.) reaches milliseconds or even higher. During data acquisition, record accurate timestamps for each image frame and environmental data. During data transmission and processing, strictly synchronize image data and environmental condition data according to timestamps to ensure data consistency and accuracy. For example, during image preprocessing and enhancement, environmental data such as light intensity and weather conditions corresponding to the image frame can be accurately obtained, providing an accurate basis for image processing.
Claims
1. A fishing vessel name recognition method based on image processing and deep learning, characterized in that: The following steps are involved: S1. Fishing boat image collection; S2. Image preprocessing and enhancement to achieve image noise reduction, deblurring and brightness adjustment, and image quality judgment after processing to form a closed-loop optimization; S3. Adaptive image enhancement to achieve illumination compensation and image dynamic enhancement; S4. Ship name area detection and segmentation; S5. Feature extraction and optimization of ship name area; S6. Recognition result feedback and optimization; S7. Character recognition and post-processing; S8. Calibration and verification; S9. Identification log and report generation; S10. Result display and storage.
2. The method for identifying fishing vessel names based on image processing and deep learning according to claim 1, characterized in that: In step S1, images of fishing boats are collected through a multi-camera layout, and the camera with the best viewing angle is intelligently selected for image collection according to the position and movement trajectory of the target fishing boat; a real-time video analysis algorithm is deployed on the camera end or edge computing device to perform preliminary processing on the collected video stream, and dynamically adjust the camera parameters according to the processing information; at the same time, the collected video stream is pre-compressed while ensuring image quality.
3. The method for identifying fishing vessel names based on image processing and deep learning according to claim 1, characterized in that: The brightness adjustment method in step S2 is as follows: using a convolutional neural network to extract features from the input fishing boat image, inputting the extracted features into a fully connected layer, and using the output of the fully connected layer as a basis for adjusting image enhancement parameters; the output of the fully connected layer is connected to two branches, one branch is used to generate a brightness adjustment coefficient, and the other branch is used to generate a contrast adjustment coefficient; specifically, a sigmoid activation function is used to map the output value to between 0 and 1 as a proportional factor for adjusting brightness and contrast; and according to the calculated brightness and contrast adjustment coefficients, the brightness and contrast of the image are adjusted pixel by pixel; For brightness adjustment, multiply the RGB value of each pixel by the brightness adjustment coefficient; For contrast adjustment, the average brightness of the image is first calculated, and then the difference between each pixel and the average brightness is adjusted according to the contrast adjustment coefficient to enhance the contrast of the image; The method for judging the image quality in step S2 is as follows: using a model that learns the mapping relationship between various visual features of an image and a quality score to judge the processed image and output a quality score, and judging whether the image is suitable for character recognition based on a preset threshold; at the same time, based on a deep learning algorithm, the image quality indicators including but not limited to clarity, contrast and noise level are evaluated, and based on the evaluation results, optimization suggestions for image preprocessing and enhancement processing algorithm parameters are provided to re-preprocess images with unqualified quality indicators, thereby forming a closed-loop optimization system.
4. The method for identifying fishing vessel names based on image processing and deep learning according to claim 1, characterized in that: The method of illumination compensation in step S3 is as follows: first, a multimodal image enhancement strategy is adopted for different types of fishing boat images, specifically: the images are classified using a deep learning model to determine the modal type to which they belong, and then corresponding image enhancement algorithm combinations are applied according to the characteristics of different modalities; for low-light images at night, a denoising algorithm based on deep learning is first used to remove noise, and then an adaptive illumination compensation algorithm is applied to increase the brightness, and then a contrast enhancement algorithm is used to highlight the ship name area; for images of different hull colors, the color correction parameters are adjusted according to the detected color information to ensure that the color of the ship name has a good contrast with the background color; The adaptive illumination compensation algorithm is a corresponding compensation strategy generated by predicting the impact of environmental data collected in real time on image quality. Specifically, the image is firstly analyzed for illumination and the illumination histogram of the image is calculated; then, according to the distribution of the illumination histogram, the area with uneven illumination in the image is judged. If the brightness value of a sub-area is significantly lower or higher than other areas, it is considered that the area has insufficient or excessive illumination problems; for the insufficiently illuminated area, the pixel value is increased to compensate, specifically, an interpolation algorithm based on neighboring pixels is adopted; for the area with excessively strong illumination, the pixel value is reduced to compensate; The method for dynamic image enhancement in step S3 specifically comprises the following steps: S31.Brightness and contrast adjustment; S311. Perform brightness adjustment: calculate the average brightness value of the image. If the average brightness value is lower than the preset lower limit of the appropriate brightness range, increase the brightness value of each pixel in a certain proportion; if it is higher than the upper limit, reduce it in proportion; S312. Perform contrast adjustment: calculate the brightness standard deviation of the image, and when the contrast is lower than a preset threshold, improve the contrast by enhancing the brightness difference between pixels; S32. Angle deviation correction: S321. Feature point detection: Use a feature point detection algorithm based on deep learning to detect representative feature points in an image. These feature points have relatively stable feature patterns when shot at different angles. S322. Angle estimation: Calculate the shooting angle deviation of the image through geometric relationships based on the detected feature points; S323. Image rotation: According to the calculated angle deviation, the image is rotated and corrected so that the ship name area is close to the horizontal or vertical direction, thereby optimizing the display effect of the ship name area.
5. The method for identifying fishing vessel names based on image processing and deep learning according to claim 1, characterized in that: The specific method of step S4 is: first, use the target detection algorithm to locate the area outside the ship name; then fuse the context information of the image to improve the accuracy of detection and segmentation, specifically: by analyzing the context features of the ship name area, assist in determining the boundary and position of the ship name area; after obtaining the preliminary ship name area segmentation result, use post-processing technology to refine the segmented area and remove the inaccurate parts of the edge; at the same time, establish a priori model of the shape and size of the ship name area, and constrain and correct the detection results according to the common shapes and size ranges of ship name areas of different types of fishing vessels.
6. The method for identifying fishing vessel names based on image processing and deep learning according to claim 1, characterized in that: The specific method of step S5 is as follows: introducing a multi-scale feature fusion mechanism to fuse feature maps at different levels; applying feature selection and dimensionality reduction techniques to optimize feature representation based on the extracted high-dimensional feature vector; using a feature selection method based on, but not limited to, information gain and chi-square test, to screen out the key features that contribute most to ship name recognition, remove redundant and irrelevant features, and reduce the dimension of the feature vector; at the same time, using principal component analysis and linear discriminant analysis dimensionality reduction algorithms to project the high-dimensional feature vector into a low-dimensional space, while maintaining the main feature information, further reducing the amount of data and computational complexity, and improving the efficiency and accuracy of subsequent recognition algorithms; in the dimensionality reduction process, selecting the best dimensionality reduction parameters by, but not limited to, cross-validation techniques, to ensure that the features after dimensionality reduction have good performance in different data sets and task scenarios.
7. The method for identifying fishing vessel names based on image processing and deep learning according to claim 1, characterized in that: The specific method of step S6 is: the parameter adjustment of the recognition system is regarded as an action, and the recognition performance index is regarded as a reward signal. The system selects the best parameter adjustment strategy through a reinforcement learning algorithm according to the current recognition results and performance indicators; when the recognition accuracy is low, try to adjust the relevant parameters of image preprocessing, feature extraction, and character recognition, observe the impact of the adjusted recognition results on the reward signal, and continuously explore and optimize the parameter space.
8. The method for identifying fishing vessel names based on image processing and deep learning according to claim 1, characterized in that: The specific method of character recognition in step S7 is as follows: first, the characters in the image are preliminarily recognized by using an OCR engine to obtain an initial character recognition result; then, the image is input into a convolutional recurrent neural network for re-recognition, and the convolutional recurrent neural network uses its ability to process sequence data to learn the contextual relationship between characters, improve the recognition ability of blurred and deformed characters, and output a recognition result; finally, the recognition results of the OCR engine and the convolutional recurrent neural network are fused, and the final character recognition result is determined by a voting mechanism or a weighted fusion method based on confidence; The specific method of post-processing in step S7 includes: Character correction: According to the grammatical rules of fishing boat names and common character combination patterns, the recognition results are checked for syntax and semantics. For suspected erroneous characters, corrections are made based on the character shape, context, and similarity with the standard character library. De-noising: remove noise characters in the recognition results. Specifically, a character dictionary can be established to treat low-frequency characters or obviously wrong characters in the recognition results that are not in the dictionary as noise and delete them. At the same time, the recognition results can be further optimized by combining the character position information, the relationship between adjacent characters and the context information in the image. Specifically, if the distance between a certain character and the surrounding characters is abnormal or does not conform to the common format of the ship name, it can be corrected or supplemented according to the context information; Establish a character recognition error case library: used to analyze and summarize common error types, continuously optimize post-processing algorithms, and improve the accuracy and reliability of character recognition.
9. The method for identifying fishing vessel names based on image processing and deep learning according to claim 1, characterized in that: The correction and verification method in step S8 includes automatic correction and manual verification. When the automatic correction cannot determine the accuracy of the recognition result or encounters complex situations, it is submitted to manual verification; The automatic correction process is specifically as follows: using image comparison technology, the image of the identified ship name area is compared with the image in the standard ship name template library; by calculating the similarity between the images, the accuracy of the recognition result is determined; if the similarity is lower than a preset threshold, the system automatically starts the correction algorithm; the correction algorithm corrects the recognition result according to the difference characteristics of the image, using technologies including but not limited to image deformation and character replacement; if a character has a large difference in shape from a standard character in the template library, the image deformation technology is used to try to adjust it to a shape closer to the standard, and then the comparison and verification are performed again until satisfactory accuracy is achieved; The manual verification process is specifically as follows: first, the automatically identified ship name is reviewed as a whole to determine whether it conforms to common sense and the ship name format in fishery management regulations; then, the ship name area in the image is carefully compared with the recognition result, with a focus on checking the characters and areas marked as suspicious; for uncertain characters, the verifier refers to the knowledge base, similar ship name cases or consults fishery experts to make a judgment, and during the verification process, the verifier records his or her own judgment basis and operation process for subsequent tracing and analysis; if the verifier believes that the automatic recognition result is correct, it is confirmed; if an error is found, the recognition result is manually modified, and the reason for the modification is noted in the system.
10. A fishing vessel name recognition system based on image processing and deep learning, characterized in that: It includes fishing boat image acquisition module, image preprocessing and enhancement module, image quality judgment module, environmental condition detection and analysis module, adaptive image enhancement module, ship name area detection and segmentation module, feature extraction and optimization module, recognition result feedback and optimization module, character recognition and post-processing module, correction and verification module, recognition log and report generation module, result display and storage module, and data transmission and synchronization module; The fishing boat image acquisition module includes multiple high-definition cameras set up around the dock to achieve all-round and multi-angle real-time monitoring of the fishing boat to ensure the clearest and most complete fishing boat image; The image preprocessing and enhancement module is based on an image preprocessing model, which automatically adjusts parameters including but not limited to noise reduction, deblurring and brightness adjustment by learning a large number of fishing boat images under different environmental conditions; The image quality judgment module is based on an image quality assessment model, the training data of which includes a large number of fishing boat images of different quality levels and corresponding manually annotated quality scores. During the training process, the image quality assessment model learns the mapping relationship between various visual features of the image and the quality score, and uses a variety of data enhancement techniques to expand the training data and a regularization method to prevent overfitting. The environmental condition detection and analysis module integrates multiple environmental sensors and is based on a machine learning model. The machine learning model uses historical environmental data and corresponding image quality data as training samples, and the training model learns the relationship between environmental changes and image quality; The adaptive image enhancement module is based on building a convolutional neural network model specifically for optimizing image enhancement parameters. The convolutional neural network model takes the original features of the image and environmental condition data as input and outputs the parameters required for image enhancement. The module learns how to generate the most appropriate enhancement parameters based on the input image and environmental information by training a large number of fishing boat images labeled with different environmental conditions and corresponding optimal image enhancement parameters. The ship name region detection and segmentation module is based on the YOLOv8 model. The YOLOv8 model uses an image dataset containing and annotating the fishing vessel name region as a training set, and the annotated information includes the bounding box coordinates and category of the ship name region. During the training process, the YOLOv8 model calculates the error between the predicted result and the true label through the loss function based on the input image and annotation information, and updates the network parameters through the back-propagation algorithm; During the training process, optimization algorithms are used to accelerate convergence and improve model performance. During the training process, data enhancement techniques are also used to increase data diversity and improve the model's ability to detect and segment ship name areas of different shapes. The feature extraction and optimization module is based on a deep convolutional neural network model. During the training process, the deep convolutional neural network model uses a combination of classification loss and regression loss to supervise the model's classification ability for ship name features and optimize the regression accuracy of feature vectors for information including but not limited to the shape and position of the ship name area; The recognition result feedback and optimization module is based on an online learning model, which can learn new sample data in real time. As new fishing boat image data is continuously input, the system automatically incorporates it into the training process and updates the parameters of the online learning model. Specifically, the incremental learning technology is used to avoid retraining all data. At the same time, the performance of the online learning model is regularly evaluated. When it is found that the performance of the online learning model has declined or a new image feature pattern has appeared, the retraining or fine-tuning operation of the online learning model is triggered. The character recognition and post-processing module is based on a character recognition model, which is based on a convolutional recurrent neural network. The character recognition model uses a large number of character sample images and annotates the corresponding character categories as a training set. The training process calculates the error between the prediction result and the true label through a loss function, updates the network parameters through back propagation, uses an optimization algorithm for optimization, and improves the generalization ability of the model through data enhancement technology. The correction and verification module is based on image comparison and intelligent correction algorithms and a knowledge base that provides auxiliary decision-making suggestions for manual verification; The identification log and report generation module is used to achieve detailed log recording and real-time monitoring as well as automated report generation and data analysis support; The result display and storage module is used to realize result display and data storage; The data transmission and synchronization module adopts data packetization and reorganization technology to realize data transmission between modules, and uses a high-precision time synchronization protocol to ensure time synchronization accuracy.
Citation Information
Cited By
Vacuum pump machining system based on groove body forming and machining device
CN120439098A
Red tide outbreak early warning method and system based on weak target detection
CN120472352A
Automobile circuit board quality monitoring method based on OCR
CN120599635A
Automatic garbage identification method and system based on data fusion
CN120655988A
Training method, medium and system for ship name area extraction model of fishing boats entering and leaving port
CN120932037A