Hydropower station gate number identification method and system

Through the improved PSENet-CRNN network model, the problem of accuracy and inefficiency in the identification of gate numbers of hydropower stations is solved, and more efficient and accurate gate numbers are achieved, which enhances the expression ability and robustness of feature maps.

CN120375347APending Publication Date: 2025-07-25聂道静
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510268490.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, the YOLO algorithm cannot accurately and efficiently identify the gate number of the hydropower station in the identification of gate numbers of hydropower stations, resulting in low identification accuracy and efficiency.

Method used

The improved PSENet-CRNN network model is used for training, and the hydropower station gate number images are collected through the camera, preprocessed and marked, and the number identification data set is built, and the improved PSENet-CRNN network model is used for training to generate a model that can accurately identify the gate number of the hydropower station.

Benefits of technology

The accuracy and efficiency of gate number identification of hydropower stations is improved. By comprehensively evaluating image clarity and introducing a standardized attention mechanism, the expression ability of feature maps is enhanced, and text areas of different sizes and shapes can be more accurately detected.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375347A_ABST
    Figure CN120375347A_ABST
Patent Text Reader

Abstract

The invention provides a hydropower station gate number identification method and system, and the method comprises the steps: collecting a number image of a hydropower station gate, and carrying out the preprocessing of the image, and obtaining standard image data; marking the standard image data to obtain a gate number identification data set; constructing a number identification model based on the network model of the improved PSENet-CRNN, inputting the gate number identification data set into the network model of the improved PSENet-CRNN for training, and obtaining a trained number identification model; and inputting the serial number image of the hydropower station gate to be identified into the trained serial number identification model, and outputting the serial number information of the hydropower station gate, the method comprises the steps of collecting the serial number image of the hydropower station gate, performing preprocessing, constructing a standard data set, performing training by using an improved PSENet-CRNN network model, and obtaining the serial number information of the hydropower station gate. The model capable of accurately identifying the hydropower station gate number can efficiently and accurately output the number information of the to-be-identified gate image, so that the accuracy and efficiency of hydropower station gate number identification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gate number recognition in hydropower stations, and particularly to a method and system for recognizing the numbers of gates in hydropower stations. Background Art

[0002] With the rapid development of China's economy, hydropower stations have become increasingly prominent in the national energy supply system. As an important part of clean energy, hydropower stations can not only provide stable power supply, but also help optimize the energy structure and reduce environmental pollution; during the construction and operation of hydropower stations, various gates, as key facilities for controlling water flow, play a crucial role. These gates include, but are not limited to, intake gates, flood discharge gates, tailwater gates, etc. Each gate has its specific function and location, jointly constituting a complex network of the hydropower station's hydraulic system; in order to ensure the efficient operation and safety management of hydropower stations, it is particularly important to accurately number and identify these gates. Gate number recognition can not only help operators quickly locate and operate specific gates, but also provide reliable data support for daily maintenance, fault troubleshooting, and emergency response.

[0003] An automatic recognition method for hydropower station gates based on YOLO with the publication number CN115862021B includes the steps of positioning the gate, positioning the gate number plate, and recognizing the gate number in sequence. The object detection model YOLO based on deep learning is used to locate the gate target in the real-time video, and the gate image is cropped; the number plate image in the intercepted gate image is located by the object detection model YOLO based on deep learning, and the target image is cropped; the gate number plate image is input into the character recognition model based on the convolutional neural network to obtain the actual content represented by the number plate and record it.

[0004] In the above solution, the YOLO algorithm is used for automatic recognition of hydropower station gate numbers. Although the YOLO algorithm can quickly locate the gate target, it cannot accurately and efficiently recognize the characters of hydropower station gate numbers, thus reducing the accuracy and efficiency of hydropower station gate number recognition. Summary of the Invention

[0005] In view of this, the present invention proposes a method and system for recognizing the numbers of gates in hydropower stations. By using an improved PSENet-CRNN network model for training, a model capable of accurately recognizing the numbers of gates in hydropower stations is obtained, which can efficiently and accurately output the number information of the gate image to be recognized, thereby improving the accuracy and efficiency of hydropower station gate number recognition.

[0006] The technical solution of the present invention is realized as follows: The present invention provides a method for recognizing the numbers of gates in hydropower stations, including the following steps:

[0007] S1. Install the camera at the end beam of the large frame of the portal machine or around the gate slot, collect the numbered images of the hydropower station gate using the camera, and preprocess the images to obtain standard image data;

[0008] S2. Label the standard image data to obtain a dataset for gate number recognition;

[0009] S3. Build a number recognition model based on the improved PSENet-CRNN network model, input the dataset for gate number recognition into the improved PSENet-CRNN network model for training, and obtain a trained number recognition model;

[0010] S4. Input the numbered image of the hydropower station gate to be recognized into the trained number recognition model, and output the number information of the hydropower station gate.

[0011] Based on the above technical solutions, preferably, step S1 includes the following sub-steps:

[0012] Install the camera at the end beam of the large frame of the portal machine or around the gate slot, collect the numbered video data of the hydropower station gate using the camera, and make the gate number located in the center of the image;

[0013] Use the video frame extraction algorithm to extract the numbered video data of the hydropower station gate at a fixed time step to obtain the numbered image of the hydropower station gate;

[0014] Perform inversion, rotation, mosaic, and cropping on the image data of the gate number to obtain standard image data.

[0015] Based on the above technical solutions, preferably, step S2 includes the following sub-steps:

[0016] S21. Extract channels from the standard image data, calculate the variance of the gray values of each channel image to obtain the first evaluation index; perform edge enhancement on each channel image using the Laplacian operator, calculate the variance of the enhanced image to obtain the second evaluation index, and use the edge detection algorithm to detect the edges of each channel image and calculate the total intensity of the edge pixels to obtain the third evaluation index;

[0017] S22. Comprehensively evaluate the clarity value of the corresponding channel image according to the first evaluation index, the second evaluation index, and the third evaluation index, obtain the channel image with the highest clarity value, and convert it into a grayscale image;

[0018] S23. Statistically analyze the pixel values in the grayscale image, calculate the sum of all pixel values, and calculate the segmentation threshold according to the sum of pixel values divided by the number of pixel points;

[0019] S24. Binarize the grayscale image according to the segmentation threshold, extract the pixel points greater than the segmentation threshold, and obtain the number recognition region;

[0020] S25. Mark the corresponding gate number in the standard image data, associate it with the number recognition region and the corresponding standard image data, generate a number recognition label, and combine the number recognition labels of all images to form a gate number recognition data set.

[0021] Based on the above technical solutions, preferably, in step S3, a number recognition model is constructed using the improved PSENet-CRNN network model. Among them, the structure of the network model includes a backbone network, a feature pyramid network, a progressive scale expansion module, a convolutional layer, a recurrent layer, and a transcription layer. Among them,

[0022] The output end of the backbone network is connected to the input end of the feature pyramid network. The backbone network uses the ResNet convolutional neural network as the basic network to extract the deep features in the image;

[0023] The output end of the feature pyramid network is connected to the input end of the progressive scale expansion module, which is used to generate multi-scale feature maps;

[0024] The output end of the progressive scale expansion module is connected to the input end of the convolutional layer, which is used to realize the progressive expansion of the text area and obtain the image feature sequence;

[0025] The output end of the convolutional layer is connected to the input end of the recurrent layer, which is used to learn each feature vector in the image feature sequence and output the predicted number recognition label distribution;

[0026] The output end of the recurrent layer is connected to the input end of the transcription layer. The transcription layer uses a loss function to convert the predicted number recognition label distribution output by the recurrent layer into the final number recognition label sequence and outputs the predicted number information result.

[0027] Based on the above technical solutions, preferably, the feature pyramid network includes an upsampling module, a downsampling module, a double-cascade receptive field module, and a convolutional fusion module. Among them,

[0028] The output end of the backbone network is connected to the input end of the downsampling module. The output end of the downsampling module is connected to the input end of the double-cascade receptive field module. The output end of the double-cascade receptive field module is connected to the input end of the upsampling module. The output end of the upsampling module is connected to the input end of the convolutional fusion module;

[0029] Both the downsampling module and the upsampling module include multiple convolutional units of different levels. The convolutional units of the same level are correspondingly arranged and connected to each other. The downsampling module performs downsampling successively from bottom to top through multiple convolutional units of different levels to obtain feature maps of multiple different scales;

[0030] The double-cascaded receptive field module includes a first convolution and a second convolution. The output end of the downsampling module is connected to the input end of the first convolution, and the output end of the first convolution is respectively connected to the input end and the output end of the second convolution. The feature maps output by the first convolution and the second convolution are fused and input into the upsampling module. The first convolutional unit is a 3×3 convolution with a dilation rate of 1, and the second convolutional unit is a 3×3 convolution with a dilation rate of 3;

[0031] The upsampling module performs upsampling successively from top to bottom on the feature maps output by the fused first convolution and second convolution through multiple convolutional units of different levels to obtain a feature map with the same size as the low-level feature map, and inputs it into the convolutional fusion module for fusion.

[0032] On the basis of the above technical solutions, preferably, a normalization attention mechanism is provided between the upsampling module and the downsampling module to correspondingly connect and fuse the feature maps of different scales obtained on the downsampling and upsampling paths. Among them,

[0033] Perform batch normalization on the feature maps of different scales after downsampling. The expression is:

[0034]

[0035] In the formula, μ B and σ B are respectively the mean and standard deviation of the mini-batch B; γ and β are trainable affine transformation parameters; ε is a very small positive constant, F in is the input feature, F o is the output feature; BN is batch normalization;

[0036] For the feature maps of different scales after batch normalization, calculate the attention weight of each channel and the sum of the attention weights, divide the absolute value of each weight by the sum of the absolute values of the weights, calculate the importance ratio of each channel, and weight the corresponding-sized feature maps after upsampling. The weighted feature map is multiplied by the corresponding-scale feature map after downsampling to obtain an enhanced feature map. The expression is:

[0037]

[0038] In the formula, F enhanced is the enhanced feature map, F in is the input feature map, C is the number of channels, |ω i| is the absolute value of the weight of the i-th channel, is the sum of the absolute values of all channel weights, F residual is the feature map of the corresponding scale after downsampling;

[0039] The convolutional fusion module fuses the enhanced feature maps of each size and inputs the fused feature map.

[0040] On the basis of the above technical solutions, preferably, the progressive scale expansion module samples the fused feature map, divides it into text kernels of multiple scales, uses the breadth-first search algorithm to identify the smallest text kernel, and gradually expands and merges from the smallest text kernel to the larger-sized text kernels to predict and output the target gate number area.

[0041] On the basis of the above technical solutions, preferably, define the loss function of the model, calculate the loss between text kernels of different scales, and the expression is:

[0042] L = λL f +(1 - λ)L c

[0043] In the formula: L f represents the text kernel of the n-th scale, Lc represents the loss of all scale text kernels, and λ is the balance coefficient between the two;

[0044] Use the Dice coefficient to calculate the loss between the predicted output target gate number area and the number recognition area in the number recognition label of the image. The expression of the coefficient D(Si, Gi) is:

[0045]

[0046] In the formula: S i,x,y and G i,x,y respectively represent the pixel values at the graphic coordinates (x, y) of the predicted output target gate number area Si and the number recognition area Gi in the number recognition label of the image;

[0047] L in the model loss function f and Lc are calculated through the Dice coefficient as follows:

[0048] L f = 1 - D(S n , G n )

[0049]

[0050] In the second aspect, the present invention also provides a number recognition system for a hydropower station gate, which is implemented by using the above-mentioned number recognition method for a hydropower station gate. The system includes:

[0051] The acquisition module is used to acquire the numbered image of the hydropower station gate and preprocess the image to obtain standard image data;

[0052] The annotation module is used to annotate the standard image data to obtain a gate number recognition data set;

[0053] The prediction module is used to construct a number recognition model based on the improved PSENet-CRNN network model, input the gate number recognition data set into the improved PSENet-CRNN network model for training, and obtain a trained number recognition model;

[0054] The output module is used to input the numbered image of the hydropower station gate to be recognized into the trained number recognition model and output the number information of the hydropower station gate.

[0055] In a third aspect, the present invention also provides a computer-readable storage medium, on which a program for a method of identifying the number of a hydropower station gate is stored, and when the program for a method of identifying the number of a hydropower station gate is executed, it implements the method of identifying the number of a hydropower station gate as described above.

[0056] The method and system for identifying the number of a hydropower station gate of the present invention have the following beneficial effects compared with the prior art:

[0057] (1) By collecting and preprocessing the numbered image of the hydropower station gate, constructing a standard data set, and training with an improved PSENet-CRNN network model, a model capable of accurately identifying the number of the hydropower station gate is obtained. This model can efficiently and accurately output the number information of the gate image to be recognized, thereby improving the accuracy and efficiency of the identification of the number of the hydropower station gate;

[0058] (2) By comprehensively evaluating the clarity of images in different channels and selecting the optimal channel for subsequent processing, the accuracy of number recognition can be effectively improved. Converting the image to a grayscale image and performing binary processing not only simplifies the image processing process but also reduces the computational complexity of subsequent machine learning;

[0059] (3) Through the double-cascaded receptive field module, the model can capture context information at different scales in the image, enhancing the expressive ability of the feature map. Moreover, a normalization attention mechanism is set between the upsampling module and the downsampling module. The model can more effectively fuse feature maps of different scales. By calculating the attention weights of each channel and weighting the upsampled feature map, the model can highlight important features and suppress unimportant features, thereby enhancing the expressive ability of the feature map. Due to the fusion of multi-scale feature maps and the enhancement of the expressive ability of the feature map, the model can more accurately detect text regions and adapt to texts of different sizes and shapes, thus improving the accuracy and robustness of the number recognition model. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0061] Figure 1 It is a flowchart of the method for identifying the number of the hydropower station gate of the present invention;

[0062] Figure 2 It is a schematic diagram of the improved PSENet-CRNN network structure for the method of identifying the number of the hydropower station gate of the present invention;

[0063] Figure 3 It is a schematic diagram of the structure of the double-cascade receptive field module for the method of identifying the number of the hydropower station gate of the present invention;

[0064] Figure 4 It is a schematic diagram of the structure of the normalized attention mechanism module for the method of identifying the number of the hydropower station gate of the present invention;

[0065] Figure 5 It is a schematic diagram of the structure of the camera installed on the end beam of the gantry crane for the method of identifying the number of the hydropower station gate of the present invention;

[0066] Figure 6 It is a schematic diagram of the structure of the camera installed at the relative positions around the gate slot for the method of identifying the number of the hydropower station gate of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0068] As Figure 1 shown, a method for identifying the number of a hydropower station gate of the present invention includes the following steps:

[0069] S1, install the camera on the end beam of the gantry crane or around the gate slot, collect the number image of the hydropower station gate by the camera, and preprocess the image to obtain standard image data.

[0070] As Figure 5 and Figure 6As shown in the figure, step S1 includes the following sub-steps:

[0071] Install the camera at the end beam of the large frame of the portal machine or around the gate slot, and use the camera to collect the video data of the gate number of the hydropower station gate, and make the gate number located in the center of the image;

[0072] Use the video frame extraction algorithm to extract the video data of the gate number of the hydropower station gate at a fixed time step to obtain the gate number image of the hydropower station;

[0073] Perform inversion, rotation, mosaic, and cropping processing on the image data of the gate number to obtain standard image data.

[0074] It should be noted that first, install a high-resolution camera at the relative position of the hydropower station gate to capture the gate number information; the position of the camera needs to be carefully adjusted to ensure that the gate number can be clearly presented in the center of the image, which can minimize the complexity and error of subsequent image processing; start the camera and perform continuous video recording on the hydropower station gate. The time length of the video recording can be set according to actual needs, but it is necessary to ensure that a clear image of the gate number can be fully captured. Use an advanced video frame extraction algorithm to extract the image frames containing the gate number from the video at a preset fixed time step. In this embodiment, one image frame is extracted every 10 seconds to reduce the data volume while retaining enough information for subsequent processing. Perform a series of preprocessing operations on the extracted gate number images, including inversion, rotation, mosaic, and cropping, to improve the image quality.

[0075] Among them, by installing the camera on the end beam of the large frame of the portal machine, the camera moves with the portal machine, and the image acquisition of the gate numbers on the movement trajectory of the portal machine can be performed. During the acquisition, the camera can be moved to the relative position of the gate number and stopped by moving the large vehicle of the portal machine, so that the gate number is located at the center position of the acquired image, and the acquisition is performed sequentially as the portal machine moves; by installing the camera around the gate slot and setting it relative to the gate number position, the acquisition is performed in a fixed one-to-one manner, and the position of the camera is adjusted so that the gate number is located at the center position of the acquired image.

[0076] S2. Annotate the standard image data to obtain the gate number recognition data set.

[0077] Among them, step S2 includes the following sub-steps:

[0078] S21. Extract channels from the standard image data, calculate the variance of the grayscale values of each channel image to obtain the first evaluation index; perform edge enhancement on each channel image using the Laplacian operator, calculate the variance of the enhanced image to obtain the second evaluation index, use an edge detection algorithm to detect the edges of each channel image, and calculate the total intensity of the edge pixels to obtain the third evaluation index;

[0079] S22. Comprehensively evaluate the sharpness values of the corresponding channel images according to the first evaluation index, the second evaluation index, and the third evaluation index, obtain the channel image with the highest sharpness value, and convert it into a grayscale image;

[0080] S23. Statistically analyze the pixel values in the grayscale image, calculate the sum of all pixel values, and calculate the segmentation threshold according to the sum of the pixel values divided by the number of pixel points;

[0081] S24. Perform binary processing on the grayscale image according to the segmentation threshold, extract the pixel points greater than the segmentation threshold, and obtain the number recognition area;

[0082] S25. Mark the corresponding gate numbers in the standard image data, associate them with the number recognition area and the corresponding standard image data, generate number recognition labels, and combine the number recognition labels of all images to form a gate number recognition data set.

[0083] It should be noted that by comprehensively evaluating the sharpness of different channel images and selecting the optimal channel for subsequent processing, the accuracy of number recognition can be effectively improved. Converting the image into a grayscale image and performing binary processing not only simplifies the image processing process but also reduces the computational complexity of subsequent machine learning.

[0084] As Figure 2 shown, S3. Build a number recognition model based on the improved PSENet-CRNN network model, input the gate number recognition data set into the improved PSENet-CRNN network model for training, and obtain the trained number recognition model.

[0085] Among them, in step S3, the number recognition model is built based on the improved PSENet-CRNN network model. The structure of the network model includes a backbone network, a feature pyramid network, a progressive scale expansion module, a convolutional layer, a recurrent layer, and a transcription layer. Among them,

[0086] The output end of the backbone network is connected to the input end of the feature pyramid network. The backbone network uses the ResNet convolutional neural network as the basic network to extract deep features in the image;

[0087] The output end of the feature pyramid network is connected to the input end of the progressive scale expansion module, which is used to generate multi-scale feature maps;

[0088] The output end of the progressive scale expansion module is connected to the input end of the convolutional layer, which is used to realize the progressive expansion of the text region and obtain the image feature sequence;

[0089] The output end of the convolutional layer is connected to the input end of the recurrent layer, which is used to learn each feature vector in the image feature sequence and output the predicted number recognition label distribution;

[0090] The output end of the recurrent layer is connected to the input end of the transcription layer. The transcription layer uses a loss function to convert the predicted number recognition label distribution output by the recurrent layer into the final number recognition label sequence and outputs the predicted number information result.

[0091] It should be noted that the backbone network is the basis of the model and is used to extract the deep features in the image. The ResNet convolutional neural network is used as the basic network. By introducing residual blocks, ResNet effectively alleviates the problems of gradient disappearance and gradient explosion in the deep neural network, improves the training efficiency and recognition performance of the model. The output end of the backbone network is connected to the input end of the feature pyramid network, providing the high-level feature representation of the image. The feature pyramid network is used to generate multi-scale feature maps to adapt to text regions of different sizes. By fusing feature maps of different levels, a feature pyramid with rich context information is constructed. The output end of the feature pyramid network is connected to the input end of the progressive scale expansion module, providing input for multi-scale text detection. The progressive scale expansion module is used to realize the progressive expansion of the text region. This module receives the output of the feature pyramid network and gradually approaches the real text region by gradually expanding the boundary of the text region. The output end of the progressive scale expansion module is connected to the input end of the convolutional layer, providing the image feature sequence after scale expansion. The convolutional layer is used to further process the image feature sequence, extract finer features, and perform convolution operations with multiple convolutional kernels to extract the key information in the image feature sequence. The recurrent layer is used to learn each feature vector in the image feature sequence and output the predicted number recognition label distribution. The long short-term memory network recurrent neural network structure is adopted to perform sequence modeling on the image feature sequence. The output end of the recurrent layer is connected to the input end of the transcription layer, providing the predicted label distribution for number recognition. The transcription layer converts the predicted number recognition label distribution output by the recurrent layer into the final number recognition label sequence, decodes the predicted label distribution into the final number recognition result using the connectionist temporal classification loss function, and the transcription layer outputs the predicted number information result as the final output of the model.

[0092] In this embodiment, by combining the progressive scale expansion characteristics of PSENet and the sequence learning ability of CRNN, the model can more accurately identify text regions of different sizes and shapes and extract the number information therein.

[0093] In addition, the feature pyramid network includes an upsampling module, a downsampling module, a double-cascaded receptive field module, and a convolutional fusion module. Among them,

[0094] the output end of the backbone network is connected to the input end of the downsampling module, the output end of the downsampling module is connected to the input end of the double-cascaded receptive field module, the output end of the double-cascaded receptive field module is connected to the input end of the upsampling module, and the output end of the upsampling module is connected to the input end of the convolutional fusion module;

[0095] both the downsampling module and the upsampling module include multiple convolutional units at different levels. The convolutional units at the same level are correspondingly arranged and connected to each other. The downsampling module performs downsampling successively from bottom to top through multiple convolutional units at different levels to obtain feature maps of multiple different scales;

[0096] It should be noted that the downsampling module includes multiple convolutional units at different levels. Each level of convolutional unit is correspondingly arranged and connected to each other. These convolutional units perform downsampling operations successively from bottom to top, thereby generating feature maps of multiple different scales. Through the downsampling module, the model can capture the feature information of the image at different scales, providing rich context information for subsequent multi-scale text detection.

[0097] As Figure 3 shown, the double-cascaded receptive field module includes a first convolution and a second convolution. The output end of the downsampling module is connected to the input end of the first convolution. The output end of the first convolution is respectively connected to the input end and the output end of the second convolution. The feature maps output by the first convolution and the second convolution are fused and input into the upsampling module. The first convolutional unit is a 3×3 convolution with a dilation rate of 1, and the second convolutional unit is a 3×3 convolution with a dilation rate of 3.

[0098] It should be noted that the double-cascaded receptive field module includes a first convolution and a second convolution. The first convolution uses a 3×3 convolution with a dilation rate of 1 to capture local features; the second convolution uses a 3×3 convolution with a dilation rate of 3 to capture more extensive context information. The output of the first convolution serves as both an input and a part of the output of the second convolution for feature map fusion. Through the double-cascaded receptive field module, the model can capture context information at different scales in the image, enhancing the expressive ability of the feature map and contributing to improving the accuracy and robustness of text detection.

[0099] The upsampling module performs upsampling successively from top to bottom on the feature maps output by the fused first convolution and second convolution through multiple convolutional units at different levels to obtain feature maps of the same size as the low-level feature maps, and inputs them into the convolutional fusion module for fusion.

[0100] It should be noted that the upsampling module also includes multiple convolutional units at different levels. These convolutional units perform upsampling operations sequentially in a top-down manner. Through the upsampling module, the model can fuse high-level feature maps with low-level feature maps to generate multi-scale feature maps with rich context information. These feature maps not only contain information about the image at different scales but also integrate feature information at different levels, which helps improve the accuracy and adaptability of text detection.

[0101] In addition, the convolutional fusion module uses 1×1 convolutions for the fusion operation of feature maps. 1×1 convolutions can not only reduce the number of channels of the feature maps but also achieve linear combinations between different feature maps. Through the convolutional fusion module, the model can effectively fuse multi-scale feature maps to generate feature pyramids with rich context information and strong expressive power. These feature pyramids provide strong support for subsequent multi-scale text detection.

[0102] In this embodiment, the downsampling module, the double-cascaded receptive field module, the upsampling module, and the convolutional fusion module are used to generate multi-scale feature maps with rich context information and strong expressive power. These feature maps not only improve the accuracy and robustness of text detection but also enhance the model's adaptability to text regions of different scales.

[0103] As Figure 4 shown, a normalization attention mechanism is set between the upsampling module and the downsampling module in this embodiment to connect and fuse the feature maps of different scales obtained on the downsampling and upsampling paths. Among them,

[0104] batch normalization is performed on the feature maps of different scales after downsampling. The expression is:

[0105]

[0106] In the formula, μ B and σ B are the mean and standard deviation of the mini-batch B respectively; γ and β are trainable affine transformation parameters; ε is a very small positive constant; F in is the input feature, and F o is the output feature; BN is batch normalization;

[0107] It should be noted that batch normalization is performed on the feature maps of different scales after downsampling. The purpose of batch normalization is to reduce internal covariate shift, thereby accelerating the convergence speed of the model and improving the stability of the model.

[0108] For the feature maps of different scales after batch normalization, calculate the attention weights of each channel and the sum of the attention weights. Divide the absolute value of each weight by the sum of the absolute values of the weights to calculate the importance ratio of each channel, and weight the feature map of the corresponding size after upsampling. Multiply the weighted feature map by the feature map of the corresponding scale after downsampling to obtain the enhanced feature map. The expression is as follows:

[0109]

[0110] In the formula, F enhanced is the enhanced feature map, F in is the input feature map, C is the number of channels, |ω i | is the absolute value of the weight of the i-th channel, is the sum of the absolute values of all channel weights, F residual is the feature map of the corresponding scale after downsampling;

[0111] The convolution fusion module fuses the enhanced feature maps of each size and inputs the fused feature map.

[0112] It should be noted that by introducing the normalized attention mechanism, the model can more effectively fuse feature maps of different scales. By calculating the attention weights of each channel and weighting the feature map after upsampling, the model can highlight important features and suppress unimportant features, thereby enhancing the expression ability of the feature map. Since the multi-scale feature maps are fused and the expression ability of the feature map is enhanced, the model can more accurately detect the text region and adapt to texts of different sizes and shapes, thus improving the performance of the number recognition model.

[0113] The progressive scale expansion module in this embodiment samples the fused feature map, divides it into text kernels of multiple scales, uses the breadth-first search algorithm to identify the smallest text kernel, and gradually expands and merges from the smallest text kernel to larger-sized text kernels to predict and output the target gate number region.

[0114] This embodiment also includes defining the loss function of the model and calculating the loss between text kernels of different scales. The expression is as follows:

[0115] L = λL f +(1 - λ)L c

[0116] In the formula: L f represents the text kernel of the n-th scale, Lc represents the loss of all-scale text kernels, and λ is the balance coefficient between the two;

[0117] Calculate the loss between the predicted output target gate number area and the number recognition area in the number recognition label of the image using the Dice coefficient. The expression of the coefficient D(Si, Gi) is as follows:

[0118]

[0119] In the formula: S i,x,y and G i,x,y respectively represent the pixel values of the predicted output target gate number area Si and the number recognition area Gi in the number recognition label of the image at the graphic coordinates (x, y);

[0120] L f and Lc in the model loss function are calculated through the Dice coefficient as follows:

[0121] L f = 1 - D(S n , G n )

[0122]

[0123] It should be noted that in the number recognition task, in order to evaluate the performance of the model and optimize its parameters, a loss function needs to be defined. The loss function in this embodiment considers two main aspects: the loss between text kernels of different scales, and the loss between the predicted output target gate number area and the true number recognition area. In order to capture text information of different scales in the image, the model usually generates text kernels of multiple scales, and these text kernels contain text areas of different sizes and shapes. Therefore, it is necessary to calculate the loss between them to ensure that the model can accurately identify text of all scales. In order to evaluate the accuracy of the model's predicted output target gate number area, it is necessary to calculate the loss between the predicted area and the true number recognition area. By optimizing the loss function, the accuracy and robustness of the model in the number recognition task can be improved.

[0124] S4. Input the number image of the hydropower station gate to be recognized into the trained number recognition model, and output the number information of the hydropower station gate.

[0125] Secondly, the present invention also provides a number recognition system for hydropower station gates, which is implemented by using the above-mentioned number recognition method for hydropower station gates. The system includes:

[0126] An acquisition module, configured to acquire the number image of the hydropower station gate and preprocess the image to obtain standard image data;

[0127] A labeling module, configured to label the standard image data to obtain a gate number recognition data set;

[0128] A prediction module, which is used to construct a number recognition model based on an improved PSENet-CRNN network model, input a sluice gate number recognition data set into the improved PSENet-CRNN network model for training, and obtain a trained number recognition model;

[0129] An output module, which is used to input the number image of the hydropower station sluice gate to be recognized into the trained number recognition model and output the number information of the hydropower station sluice gate.

[0130] It should be noted that this system corresponds to the above-mentioned method for recognizing the numbers of hydropower station sluice gates. All implementation manners in the above method embodiments are applicable to the embodiments of this system and can achieve the same technical effects.

[0131] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0132] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described system and modules can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.

[0133] In the embodiments provided by the present invention, it should be understood that the disclosed system and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0134] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0135] In addition, in each embodiment of the present invention, each functional unit may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit.

[0136] If the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0137] In addition, it should be noted that in the system and method of the present invention, obviously, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present invention. And, the steps of performing the above series of processes can naturally be executed in chronological order according to the described order, but it is not necessary to execute them in chronological order. Certain steps can be executed in parallel or independently of each other. For those of ordinary skill in the art, it is understandable that all or any steps or components of the method and device of the present invention can be implemented in any computing device (including a processor, a storage medium, etc.) or a network of computing devices in the form of hardware, firmware, software, or a combination thereof, which can be achieved by those of ordinary skill in the art using their basic programming skills after reading the description of the present invention.

[0138] Therefore, the object of the present invention can also be achieved by running a program or a set of programs on any computing system. The computing system may be a well-known general system. Therefore, the object of the present invention can also be achieved only by providing a program product containing program codes for implementing the method or device. That is to say, such a program product also constitutes the present invention, and the storage medium storing such a program product also constitutes the present invention. Obviously, the storage medium can be any well-known storage medium or any storage medium developed in the future. It should also be noted that in the device and method of the present invention, obviously, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present invention. And, the steps of performing the above series of processes can naturally be executed in chronological order according to the described order, but it is not necessary to execute them in chronological order. Certain steps can be executed in parallel or independently of each other.

[0139] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for identifying the number of a gate of a hydropower station, characterized in that, It includes the following steps: S1. Install the camera at the end beam of the gantry crane frame or around the gate slot, use the camera to collect the numbered images of the hydropower station gate, and preprocess the images to obtain standard image data; S2. Annotate the standard image data to obtain the gate number recognition data set; S3. Build a number recognition model based on the improved PSENet-CRNN network model, input the gate number recognition data set into the improved PSENet-CRNN network model for training, and obtain the trained number recognition model; S4. Input the numbered image of the hydropower station gate to be recognized into the trained number recognition model, and output the number information of the hydropower station gate.

2. The number identification method of the hydropower station gate according to claim 1, characterized in that: The step S1 includes the following sub-steps: Install the camera at the end beam of the gantry crane frame or around the gate slot, use the camera to collect the numbered video data of the hydropower station gate, and make the gate number located in the center of the image; Use the video frame extraction algorithm to extract the numbered video data of the hydropower station gate at a fixed time step to obtain the numbered image of the hydropower station gate; Perform inversion, rotation, mosaic and cropping processing on the image data of the gate number to obtain the standard image data.

3. The number identification method of the hydropower station gate according to claim 1, characterized in that: The step S2 includes the following sub-steps: S21. Extract the channels of the standard image data, calculate the variance of the gray values of each channel image to obtain the first evaluation index; perform edge enhancement on each channel image using the Laplace operator, calculate the variance of the enhanced image to obtain the second evaluation index, use the edge detection algorithm to detect the edges of each channel image, and calculate the total intensity of the edge pixels to obtain the third evaluation index; S22. Comprehensively evaluate the clarity value of the corresponding channel image according to the first evaluation index, the second evaluation index and the third evaluation index, obtain the channel image with the highest clarity value, and convert it into a gray image; S23. Statistically calculate the pixel values in the gray image, calculate the sum of all pixel values, and calculate the segmentation threshold according to the sum of pixel values divided by the number of pixel points; S24. Perform binary processing on the gray image according to the segmentation threshold, extract the pixel points greater than the segmentation threshold, and obtain the number recognition area; S25. Annotate the corresponding gate number in the standard image data, associate it with the number recognition area and the corresponding standard image data, generate a number recognition label, and combine the number recognition labels of all images to form a gate number recognition data set.

4. The number identification method of the hydropower station gate according to claim 3, characterized in that: In step S3, the number recognition model is built based on the improved PSENet-CRNN network model, where the structure of the network model includes a backbone network, a feature pyramid network, a progressive scale expansion module, a convolutional layer, a recurrent layer and a transcription layer, where The output end of the backbone network is connected to the input end of the feature pyramid network. The backbone network uses the ResNet convolutional neural network as the basic network to extract the deep features in the image; The output end of the feature pyramid network is connected to the input end of the progressive scale expansion module to generate multi-scale feature maps; The output end of the progressive scale expansion module is connected to the input end of the convolutional layer to realize the progressive expansion of the text area and obtain the image feature sequence; The output end of the convolutional layer is connected to the input end of the recurrent layer, which is used to learn each feature vector in the image feature sequence and output the predicted number recognition label distribution; The output end of the recurrent layer is connected to the input end of the transcription layer. The transcription layer uses a loss function to convert the predicted number recognition label distribution output by the recurrent layer into a final number recognition label sequence and outputs the predicted number information result.

5. The method for number identification of the hydropower station gate according to claim 4, wherein: The feature pyramid network includes an upsampling module, a downsampling module, a double-cascade receptive field module, and a convolutional fusion module. Among them, The output end of the backbone network is connected to the input end of the downsampling module, the output end of the downsampling module is connected to the input end of the double-cascade receptive field module, the output end of the double-cascade receptive field module is connected to the input end of the upsampling module, and the output end of the upsampling module is connected to the input end of the convolutional fusion module; Both the downsampling module and the upsampling module include multiple convolutional units at different levels. The convolutional units at the same level are correspondingly set and connected to each other. The downsampling module successively downsamples from bottom to top through multiple convolutional units at different levels to obtain feature maps of multiple different scales; The double-cascade receptive field module includes a first convolution and a second convolution. The output end of the downsampling module is connected to the input end of the first convolution, and the output end of the first convolution is respectively connected to the input end and the output end of the second convolution to fuse the feature maps output by the first convolution and the second convolution and input them into the upsampling module. The first convolutional unit is a 3×3 convolution with a dilation rate of 1, and the second convolutional unit is a 3×3 convolution with a dilation rate of 3; The upsampling module successively upsamples the feature maps output by the fused first convolution and second convolution from top to bottom through multiple convolutional units at different levels to obtain a feature map with the same size as the low-level feature map and inputs it into the convolutional fusion module for fusion.

6. The number identification method of the hydropower station gate according to claim 5, characterized in that: A normalization attention mechanism is set between the upsampling module and the downsampling module to correspondingly connect and fuse the feature maps of different scales obtained on the downsampling and upsampling paths. Among them, Perform batch normalization processing on the feature maps of different scales after downsampling. The expression is: where, μ B and σ B are the mean and standard deviation of the minimum batch B respectively; γ and β are trainable affine transformation parameters; ε is a very small positive constant, F in is the input feature, and F o is the output feature; BN is batch normalization processing; For the feature maps of different scales after batch normalization processing, calculate the attention weight and the sum of attention weights for each channel, divide the absolute value of each weight by the sum of the absolute values of the weights to calculate the importance ratio of each channel, and weight the corresponding-size feature maps after upsampling. Multiply the weighted feature maps by the corresponding-scale feature maps after downsampling to obtain enhanced feature maps. The expression is: Wherein, F enhanced is the enhanced feature map, F in is the input feature map, C is the number of channels, |ω i | is the absolute value of the weight of the i-th channel, is the sum of the absolute values of all channel weights, F residual is the feature map of the corresponding scale after downsampling; The convolutional fusion module fuses the enhanced feature maps of each size and inputs the fused feature map.

7. The number identification method of the hydropower station gate according to claim 6, characterized in that: The progressive scale expansion module samples the fused feature map, divides it into text kernels of multiple scales, uses the breadth-first search algorithm to identify the smallest text kernel, and gradually expands and merges from the smallest text kernel to larger-size text kernels to predict and output the target gate number area.

8. The number identification method for the gate of a hydropower station according to claim 1, characterized in that: Define the loss function of the model and calculate the loss between text kernels of different scales. The expression is: L = λL f +(1 - λ)L c where: L f represents the text kernel of the nth scale, Lc represents the loss of the text kernels of all scales, and λ is the balance coefficient between the two; Calculate the loss between the predicted output target gate number area and the number recognition area in the number recognition label of the image using the Dice coefficient. The expression of the coefficient D(Si, Gi) is as follows: where: S i,x,y and G i,x,y respectively represent the pixel values of the predicted output target gate number area Si and the number recognition area Gi in the number recognition label of the image at the graphic coordinates (x, y); L in the model loss function f 、The calculation expression of Lc through the Dice coefficient is as follows: L f = 1 - D(S n , G n ) 9. A number identification system for a hydropower station gate, implemented by using the number identification method for a hydropower station gate according to any one of claims 1-8, characterized in that, The system includes: An acquisition module, which is used to acquire the number image of the hydropower station gate, preprocess the image, and obtain standard image data; A labeling module, which is used to label the standard image data to obtain a gate number recognition data set; A prediction module, which is used to construct a number recognition model based on the improved PSENet-CRNN network model, input the gate number recognition data set into the improved PSENet-CRNN network model for training, and obtain a trained number recognition model; An output module, which is used to input the number image of the hydropower station gate to be recognized into the trained number recognition model and output the number information of the hydropower station gate.

10. A computer-readable storage medium, characterized in that, A program for a method of recognizing the number of a hydropower station gate is stored on the storage medium. When the program for a method of recognizing the number of a hydropower station gate is executed, it implements a method of recognizing the number of a hydropower station gate according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Automatic identification method of hydropower station gate based on YOLO

    CN115862021B