Power equipment model identification method and device, storage medium and computer equipment

Through feature extraction, fusion and enhancement processing of the pre-trained text detection model, the problem of low accuracy in model recognition of distribution network equipment is solved, and efficient model recognition in complex environments is achieved.

CN120635875APending Publication Date: 2025-09-12QINZHOU POWER SUPPLY BUREAU OF GUANGXI POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510512342.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In the prior art, the accuracy of model recognition of distribution network equipment is low, especially when the appearance of the equipment changes in different usage environments, which makes recognition difficult.

Method used

Using a pre-trained text detection model, this method processes images of power equipment and identifies equipment models through feature extraction, fusion, enhancement, and output modules. The feature extraction module includes multiple convolutional layers and dilated convolutional layers. The fusion module performs feature fusion. The enhancement module uses a long short-term memory neural network to process feature sequences. The output module performs text box information prediction and character recognition.

Benefits of technology

The accuracy of power equipment model recognition is improved, and the equipment model can be accurately identified in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635875A_ABST
    Figure CN120635875A_ABST
Patent Text Reader

Abstract

The invention discloses a power equipment model identification method and device, a storage medium and computer equipment, and relates to the field of power systems and image identification. The method comprises the following steps: acquiring a target power equipment image, and performing feature extraction on the target power equipment image based on a feature extraction module in a pre-trained text detection model to obtain multi-layer image features; based on a feature fusion module in the text detection model, performing feature fusion on the multiple layers of image features to obtain fused features; based on a feature enhancement module in the text detection model, performing enhancement processing on the fused features to obtain an enhanced feature sequence; based on an output module in the text detection model, predicting the enhanced feature sequence to obtain textbox information corresponding to a text region in the target power equipment image; and performing character recognition on the textbox information to obtain a target power equipment model. The accuracy of power equipment model identification can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power systems and image recognition technology, and in particular to a method, device, storage medium and computer equipment for identifying the model of power equipment. Background Art

[0002] As a critical link in the power system directly serving users, the stable operation of the distribution network is crucial for ensuring the reliability and continuity of power supply. The distribution network encompasses numerous types of equipment, such as transformers, switchgear, distribution boxes, and circuit breakers. The proper operation of these devices directly impacts the safety and reliability of the network. As the scale of distribution networks continues to expand and the number of devices increases, efficient and accurate management of these devices has become a significant challenge. Distribution network equipment is widely distributed and installed in a variety of environments, including indoors, outdoors, at high altitudes, and underground. Different installation environments can affect the appearance of the equipment. For example, outdoor equipment is subject to erosion from natural factors like sunlight and rain, resulting in surface stains and discoloration, making it difficult to identify the device model.

[0003] Some identification methods based on single features (such as the device's appearance, shape, color, etc.) can assist in identifying device models to a certain extent. However, since distribution networks may have similarities in appearance and the same model of device may have appearance changes in different usage environments (such as aging, maintenance, etc.), the accuracy of device model identification is low. Summary of the Invention

[0004] In view of this, the present application provides a method, device, storage medium and computer equipment for identifying the model of an electric power device, the main purpose of which is to solve the technical problem of low accuracy in identifying the model of an electric power device.

[0005] According to a first aspect of the present invention, a method for identifying a model of an electric power device is provided, the method comprising:

[0006] Acquire a target power equipment image, and perform feature extraction on the target power equipment image based on a feature extraction module in a pre-trained text detection model to obtain multi-layer image features;

[0007] Based on the feature fusion module in the text detection model, the multi-layer image features are fused to obtain fused features;

[0008] Based on the feature enhancement module in the text detection model, the fused features are enhanced to obtain an enhanced feature sequence;

[0009] Based on the output module in the text detection model, the enhanced feature sequence is predicted to obtain text box information corresponding to the text area in the target power equipment image;

[0010] Character recognition is performed on the text box information to obtain the target power equipment model.

[0011] Optionally, the feature extraction module in the text detection model includes multiple first convolutional layers and one second convolutional layer; the feature extraction module in the pre-trained text detection model performs feature extraction on the target power equipment image to obtain multi-layer image features, including: the target power equipment image is sequentially subjected to feature extraction through multiple first convolutional layers and one second convolutional layer in the feature extraction module to obtain image features of different sizes corresponding to each convolutional layer; wherein, the second convolutional layer includes an original convolution branch and multiple hole convolution branches, each of the hole convolution branches is provided with a different hole rate, and after the image features extracted by the original convolution branch and the hole convolution branch are fused, the image features corresponding to the second convolution layer are obtained.

[0012] Optionally, the feature fusion module in the text detection model performs feature fusion on multiple layers of image features to obtain fused features, including: merging the image features extracted from each convolutional layer layer by layer from bottom to top through feature connection and convolution until the final fused features are generated.

[0013] Optionally, the feature enhancement module based on the text detection model enhances the fused features to obtain an enhanced feature sequence, including: serializing the fused features to obtain a one-dimensional feature vector sequence, and grouping the feature vector sequence according to a preset batch size to obtain multiple groups of feature sequences; inputting the multiple groups of feature sequences into a long short-term memory neural network in sequence for bidirectional processing and splicing them in sequence to form a new one-dimensional feature vector sequence to obtain the enhanced feature sequence.

[0014] Optionally, the output module in the text detection model includes a text area score branch and a text box parameter branch; based on the output module in the text detection model, the enhanced feature sequence is predicted to obtain text box information corresponding to the text area in the target power equipment image, including: processing the enhanced feature sequence through the convolution layer of the text area score branch to obtain an image score map, wherein the pixel value in the image score map represents the score of the pixel belonging to the text area; predicting the distance between each pixel in the image score map and the text area boundary and predicting the rotation angle of the text area through the convolution layer of the text box output branch.

[0015] Optionally, the training process of the text detection model includes: obtaining multiple images containing power equipment model information, rotating, scaling, and adding noise to the images to obtain multiple sample images; and training each module in the text detection model based on the sample images.

[0016] Optionally, the training process of the text detection model further includes: optimizing the training process using a Dice soft loss function and an enhanced cosine loss function.

[0017] According to a second aspect of the present invention, there is provided a device for identifying a model of an electric power device, the device comprising:

[0018] A feature fusion module is used to fuse the multiple layers of image features to obtain fused features;

[0019] A feature enhancement module, configured to enhance the fused features to obtain an enhanced feature sequence;

[0020] A text output module, configured to predict the enhanced feature sequence to obtain text box information corresponding to the text area in the target power equipment image;

[0021] The model recognition module is used to perform character recognition on the text box information to obtain the target power equipment model.

[0022] According to a third aspect of the present invention, a storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the above-mentioned method for identifying the model of an electric power device is implemented.

[0023] According to a fourth aspect of the present invention, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method for identifying the model of electric power equipment when executing the program.

[0024] The present invention provides a method, device, storage medium and computer equipment for identifying the model of electric power equipment. The method obtains a target electric power equipment image, performs feature extraction on the target electric power equipment image based on a feature extraction module in a pre-trained text detection model, and obtains multi-layer image features; performs feature fusion on the multi-layer image features based on a feature fusion module in the text detection model to obtain fused features; performs enhancement processing on the fused features based on a feature enhancement module in the text detection model to obtain an enhanced feature sequence; predicts the enhanced feature sequence based on an output module in the text detection model to obtain text box information corresponding to the text area in the target electric power equipment image; performs character recognition on the text box information to obtain the target electric power equipment model. By means of the above technical solution, the text area in the image is accurately extracted through the text detection model, and then character recognition is performed on the text information in the text area to obtain the model of the electric power equipment, thereby improving the accuracy of electric power equipment model recognition.

[0025] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0027] Figure 1 A schematic diagram showing a flow chart of a method for identifying a model of an electric device provided by an embodiment of the present invention;

[0028] Figure 2 A schematic diagram of the structure of a text detection model provided by an embodiment of the present invention is shown;

[0029] Figure 3 A schematic structural diagram of a device for identifying a model of an electric power device provided by an embodiment of the present invention is shown;

[0030] Figure 4 A schematic structural diagram of a computer device provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0031] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of the present application can be combined with each other.

[0032] The present application embodiment provides a method for identifying the model of an electric device. In one embodiment, Figure 1 As shown, the method includes the following steps:

[0033] 101. Obtain a target power equipment image, and perform feature extraction on the target power equipment image based on a feature extraction module in a pre-trained text detection model to obtain multi-layer image features.

[0034] 102. Based on the feature fusion module in the text detection model, perform feature fusion on the multiple layers of image features to obtain fused features.

[0035] 103. Based on the feature enhancement module in the text detection model, enhance the fused features to obtain an enhanced feature sequence.

[0036] 104. Based on the output module in the text detection model, predict the enhanced feature sequence to obtain text box information corresponding to the text area in the target power equipment image.

[0037] 105. Perform character recognition on the text box information to obtain the target power equipment model.

[0038] As a critical link in the power system directly serving users, the stable operation of the distribution network is crucial for ensuring the reliability and continuity of power supply. The distribution network encompasses numerous types of equipment, such as transformers, switchgear, distribution boxes, and circuit breakers. The proper operation of these devices directly impacts the safety and reliability of the network. As the scale of distribution networks continues to expand and the number of devices increases, efficient and accurate management of these devices has become a significant challenge. Distribution network equipment is widely distributed and installed in a variety of environments, including indoors, outdoors, at high altitudes, and underground. Different installation environments can affect the appearance of the equipment. For example, outdoor equipment is subject to erosion from natural factors like sunlight and rain, resulting in surface stains and discoloration, making it difficult to identify the device model.

[0039] Some identification methods based on single features (such as the device's appearance, shape, color, etc.) can assist in identifying device models to a certain extent. However, since distribution networks may have similarities in appearance and the same model of device may have appearance changes in different usage environments (such as aging, maintenance, etc.), the accuracy of device model identification is low.

[0040] To address the above issues, this embodiment uses image and text recognition to identify the model of a target electrical device. Specifically, an image of the target electrical device is captured. A text detection model is then used to identify the text region within the image. This region corresponds to the model label or nameplate. Finally, character recognition is used to identify the characters within the text region to obtain the model of the target electrical device.

[0041] During text region recognition, image features are first extracted using a feature extraction module. This module comprises multiple convolutional layers, each with a different convolution kernel size. This allows for the extraction of image features at multiple scales. The feature fusion module then fuses these image features at different scales to produce fused features. The fused features are then enhanced using a feature enhancement module to produce an enhanced feature sequence. Finally, the output module predicts the enhanced feature sequence to obtain the text box information corresponding to the text region. This information includes the pixel position of the text region and the rotation angle of the text box. This allows the text region to be extracted from the target power equipment image, and character recognition is then performed based on the text region. The image features of varying sizes extracted by the feature extraction module yield rich feature information, providing a data foundation for subsequent processing. Feature enhancement of the fused features strengthens the correlation between the features, resulting in more accurate text region localization results based on the enhanced features, which in turn improves the accuracy of character recognition in text regions.

[0042] The electric power equipment model recognition method provided in this embodiment obtains a target electric power equipment image, performs feature extraction on the target electric power equipment image based on the feature extraction module in the pre-trained text detection model, and obtains multi-layer image features; performs feature fusion on the multi-layer image features based on the feature fusion module in the text detection model to obtain fusion features; performs enhancement processing on the fusion features based on the feature enhancement module in the text detection model to obtain an enhanced feature sequence; predicts the enhanced feature sequence based on the output module in the text detection model to obtain text box information corresponding to the text area in the target electric power equipment image; performs character recognition on the text box information to obtain the target electric power equipment model. By means of the above technical solution, the text area in the image is accurately extracted through the text detection model, and then the text information in the text area is subjected to character recognition to obtain the model of the electric power equipment, thereby improving the accuracy of electric power equipment model recognition.

[0043] Furthermore, in order to fully illustrate the implementation process of this embodiment, the specific implementation methods of the above embodiment are refined and expanded below. Specifically, in one embodiment, the feature extraction module in the text detection model includes multiple first convolutional layers and one second convolutional layer; the feature extraction module in the pre-trained text detection model in step 101 performs feature extraction on the target power equipment image to obtain multi-layer image features, which can be specifically implemented in the following way: the target power equipment image is sequentially subjected to feature extraction through multiple first convolutional layers and one second convolutional layer in the feature extraction module to obtain image features of different sizes corresponding to each convolutional layer; wherein, the second convolutional layer includes an original convolution branch and multiple hole convolution branches, each of the hole convolution branches is provided with a different hole rate, and the image features extracted by the original convolution branch and the hole convolution branch are fused to obtain the image features corresponding to the second convolution layer.

[0044] In the above embodiment, the text detection model can be implemented based on the EAST (Efficient and Accurate SceneText) model. Specifically, in one example, the structure of the text detection model can be as follows: Figure 2 As shown. The feature extraction module includes multiple first convolutional layers, and the first convolutional layer can be, for example, the second to fourth convolutional modules in the VGG16 (Visual Geometry Group) model ( Figure 2 The second convolutional layer in the feature extraction model is implemented using Conv stages 2 to 4 in the VGG16 network. Considering the computational complexity of the model, the features extracted by the first convolutional module (Conv stage 1) in the VGG16 network are not included in the calculation to improve the model's lightweightness. The second convolutional layer in the feature extraction model can be implemented using an ASPP (Atrous Spatial Pyramid Pooling) network. The second convolutional layer consists of an original convolution branch and multiple dilated convolution branches. Each dilated convolution branch has a different dilation rate.

[0045] The target power equipment image is input into the feature extraction module, and features are extracted through multiple first convolutional layers and second convolutional layers in sequence, and finally multi-layer image features are extracted. For example, Figure 2 In the text detection model shown, the feature extraction module includes three first convolutional layers. The resulting image features include the three layers of image features extracted by the first convolutional layers plus the image features extracted by the second convolutional layer. The features extracted by the second convolutional layer are a fusion of the features extracted by the original convolution branch and the features extracted by multiple dilated convolution branches. Figure 2 The second convolutional layer shown in the example includes one original convolution branch and three dilated convolution branches. Figure 2Taking the structure of the second convolutional layer in as an example, the convolution kernel size of the original convolution branch is 1×1 and the stride is 1. The convolution kernel of the dilated convolution branch is 3×3. According to the dilation rate, it can be divided into three dilation rate branches: small, medium, and large, with dilation rates of 3, 6, and 9 respectively.

[0046] The feature output by the original convolution branch is Y1:

[0047]

[0048] The features output by the dilated convolution are Y2, Y3, and Y4, and the calculation formula is:

[0049]

[0050] In the above formula, i and j represent the index of the row and column of the output feature map X, respectively, k represents the channel index of the output feature map, and W k,l (m,n) represents the convolution kernel weight, m, n, l are the convolution indexes; d_rate is the dilation rate, d = 2, 3, 4.

[0051] The features extracted by the original convolution branch and the hole convolution branch of the second convolution layer are spliced ​​in the channel dimension and fused to obtain the image features extracted by the second convolution layer. Based on the multi-layer image features extracted by multiple first convolution layers and the image features extracted by the second convolution layer, the final multi-layer image features of different scales are obtained. In the above embodiment, the second convolution layer is provided with the original convolution branch and the hole convolution branches with different hole rates, and the receptive field of the convolution kernel is expanded by step-by-step expansion. Combined with the multiple image features extracted by the first convolution layer, multiple layers of image features of different scales are obtained, thereby obtaining rich feature information of the target power equipment image, laying the foundation for accurate identification of text areas.

[0052] In one embodiment, the feature fusion module in the text detection model described in step 102 performs feature fusion on multiple layers of image features to obtain fused features, which can be specifically achieved in the following way: the image features extracted from each convolutional layer are merged layer by layer from bottom to top through feature connection and convolution until the final fused features are generated.

[0053] In the above embodiment, the multiple image features extracted by the feature extraction module are merged layer by layer in a bottom-up manner. Figure 2 For example, the image features extracted by the second convolutional layer are upsampled by 2 times and then combined with the second first convolutional layer ( Figure 2 The image features extracted by Conv stage 3) are connected, and then the channels are merged through the 1×1 convolution kernel and the 3×3 convolution kernel. The merged image features are then upsampled by 2 times and then combined with the first convolution layer ( Figure 2The image features extracted in Convstage 2) are connected, and then the channels are merged through a 1×1 convolution kernel and a 3×3 convolution kernel, and then processed by a 3×3 convolution kernel to obtain the final fusion feature.

[0054] In one embodiment, the feature enhancement module in the text detection model described in step 103 enhances the fused features to obtain an enhanced feature sequence, which can be specifically achieved in the following manner: serializing the fused features to obtain a one-dimensional feature vector sequence, and grouping the feature vector sequence according to a preset batch size to obtain multiple groups of feature sequences; inputting the multiple groups of feature sequences into a long short-term memory neural network in sequence for bidirectional processing and splicing them in sequence to form a new one-dimensional feature vector sequence to obtain the enhanced feature sequence.

[0055] In the above embodiment, before enhancing the fused features, the fused features are first serialized. Each row in the fused feature's two-dimensional feature map is expanded sequentially in row order, and each element is concatenated sequentially to form a one-dimensional feature vector sequence. The feature vector sequence is then grouped according to a preset batch size to form batch data, which is sequentially input into the feature enhancement module for processing. The feature enhancement module can be implemented using an LC-BLSTM, a recurrent long short-term memory neural network composed of a bidirectional BLSTM neural network. Specifically, during initialization, a bidirectional processing method and calculation logic are determined to capture past and future information of the sequence respectively. The two hidden feature sequences obtained by the bidirectional calculation (the hidden state sequence in the past direction and the hidden state sequence in the future direction) are sequentially spliced ​​according to corresponding positions to form a new one-dimensional feature sequence, resulting in an enhanced feature sequence. The enhanced feature sequence integrates past and future information of the sequence, making it more reasonable, uniform, and connected. The LC-BLSTM network directly acts on the sequence elements through its bidirectional structure and gating mechanism. By capturing long-term dependencies and bidirectional information in the sequence, it enhances the inherent logic of the feature sequence, providing more connected feature information for device model identification.

[0056] In one embodiment, the output module in the text detection model includes a text area score branch and a text box parameter branch; in step 104, the output module in the text detection model is used to predict the enhanced feature sequence to obtain text box information corresponding to the text area in the target power equipment image, which can be specifically achieved in the following manner: the enhanced feature sequence is processed through the convolution layer of the text area score branch to obtain an image score map, and the pixel value in the image score map represents the score of the pixel belonging to the text area; the distance between each pixel in the image score map and the text area boundary and the rotation angle of the predicted text area are predicted through the convolution layer of the text box output branch.

[0057] After feature extraction, feature fusion and feature enhancement, the output layer receives a feature sequence with rich information and uses the fully connected layer to locate the text area. Figure 2 As shown, the output layer can include a text region score branch and a text box parameter branch. The text region score branch can process the enhanced feature sequence through a 1×1 convolution kernel to obtain a score map of image features. Each pixel value in the score map is used to represent the score / probability of the pixel belonging to the text region, based on which the text region can be located. The text box parameter branch can calculate the distance between each pixel and the text region boundary in each direction (up, down, left, and right) through a 1×1 convolution kernel, and calculate the rotation angle of the text region through a 1×1 convolution kernel. Based on the above information, the text region is extracted from the target power equipment image, and then the text region is subjected to character recognition, such as through OCR, to identify the model of the target power equipment.

[0058] In one embodiment, the training process of the text detection model includes: obtaining multiple images containing power equipment model information, rotating, scaling, and adding noise to the images to obtain multiple sample images; and training each module in the text detection model based on the sample images. This increases the diversity of training data and improves the generalization ability of the model.

[0059] Among them, the Dice soft loss function and the enhanced cosine loss function are used for optimization during the training process. Specifically, when performing text area recognition on the target power equipment image, the text area accounts for a small proportion of the image, and the background accounts for a large proportion. In order to better solve this uneven distribution problem, the class-balanced cross entropy loss is usually used as the loss function for the text area score. However, the cross entropy loss mainly focuses on the overlapping part of the predicted results and the true results. When the background accounts for a large proportion, the model only reduces the overall loss by simply reducing the prediction of small targets, which will result in missed detection of small targets. Therefore, this embodiment uses the Dice soft loss function and the enhanced cosine loss function to comprehensively measure the overlapping proportion of positive and negative samples based on the predicted probability distribution, wherein the formula of the Dice softloss function is:

[0060]

[0061] Among them, p i 、g i are the predicted value and true value in probability form respectively.

[0062] The enhanced cosine loss function is introduced to maximize inter-class differences and minimize intra-class differences. By eliminating radial variations, it focuses on the essential structure and directional characteristics of the text. In the feature space output by the model, the sample feature vector and the category center vector are analyzed using cosine similarity. On this basis, a fixed parameter, cosine margin value τ, is added to further maximize the decision boundary of the learned features in angular space, thereby more accurately detecting text information. The formula of the enhanced cosine loss function is:

[0063]

[0064] Where N is the number of samples, θ j is the angle between the weight vector and the eigenvector, y i is the true category label of the i-th sample, j is the category index, j≠y i represents the categories other than the true category, s is the sensitivity of the adjustment loss to the similarity difference, and cos(·) is the cosine similarity.

[0065] The electric power equipment model recognition method provided in the embodiment of the present application adopts the new end-to-end EAST text detection algorithm in the recognition task of the electric power equipment model, reduces the proportion of redundant parts in the detection process, and directly predicts the text content. The hole convolution in the ASPP network is used to enhance the receptive field of the extracted features. At the same time, the balanced rationality and connectivity of the feature sample sequence are strengthened based on the recurrent long short-term memory neural network, thereby improving the robustness of the model for paragraph text processing. Finally, the Dice softloss function and the enhanced cosine loss function are designed to strengthen the model's attention to sample distribution and classification, further enhance the effectiveness of the model in detecting text data, and improve the accuracy of equipment model recognition.

[0066] Further, as Figure 1 As well as the specific implementation of the method shown in the above embodiment, this embodiment provides a device for identifying the model of an electric device, such as Figure 3 As shown, the device includes: a feature extraction module 31, a feature fusion module 32, a feature enhancement module 33, a text output module 34, and a model identification module 35.

[0067] The feature extraction module 31 can be used to obtain a target power equipment image, perform feature extraction on the target power equipment image, and obtain multi-layer image features;

[0068] A feature fusion module 32 is configured to fuse multiple layers of image features to obtain fused features;

[0069] A feature enhancement module 33 is configured to enhance the fused features to obtain an enhanced feature sequence;

[0070] A text output module 34 is configured to predict the enhanced feature sequence and obtain text box information corresponding to the text area in the target power equipment image;

[0071] The model recognition module 35 can be used to perform character recognition on the text box information to obtain the model of the target power equipment.

[0072] In a specific application scenario, the feature extraction module in the text detection model includes multiple first convolutional layers and one second convolutional layer; the feature extraction module in the pre-trained text detection model performs feature extraction on the target power equipment image to obtain multi-layer image features. The feature extraction module 31 can be specifically used to sequentially extract the target power equipment image through multiple first convolutional layers and one second convolutional layer in the feature extraction module to obtain image features of different sizes corresponding to each convolutional layer; wherein, the second convolutional layer includes an original convolution branch and multiple hole convolution branches, each of the hole convolution branches is provided with a different hole rate, and after the image features extracted by the original convolution branch and the hole convolution branch are fused, the image features corresponding to the second convolution layer are obtained.

[0073] In a specific application scenario, the feature fusion module based on the text detection model performs feature fusion on multiple layers of image features to obtain fused features. The feature fusion module 32 can be used to merge the image features extracted from each convolution layer layer by layer from bottom to top through feature connection and convolution until the final fused features are generated.

[0074] In a specific application scenario, the feature enhancement module based on the text detection model enhances the fused features to obtain an enhanced feature sequence. The feature enhancement module 33 can be specifically used to serialize the fused features to obtain a one-dimensional feature vector sequence, and group the feature vector sequence according to a preset batch size to obtain multiple groups of feature sequences; the multiple groups of feature sequences are sequentially input into the long short-term memory neural network for bidirectional processing and sequentially spliced ​​to form a new one-dimensional feature vector sequence to obtain the enhanced feature sequence.

[0075] In a specific application scenario, the output module in the text detection model includes a text area score branch and a text box parameter branch; the output module based on the text detection model predicts the enhanced feature sequence to obtain the text box information corresponding to the text area in the target power equipment image, and the text output module 34 can be specifically used to process the enhanced feature sequence through the convolution layer of the text area score branch to obtain an image score map, and the pixel value in the image score map represents the score of the pixel belonging to the text area; through the convolution layer of the text box output branch, the distance between each pixel in the image score map and the text area boundary and the predicted rotation angle of the text area are predicted.

[0076] In a specific application scenario, the device also includes a model training module. During the training process of the text detection model, the model training module can be specifically used to: obtain multiple images containing power equipment model information, rotate, scale, and add noise to the images to obtain multiple sample images; train each module in the text detection model based on the sample images; and use the Dice softloss function and the enhanced cosine loss function to optimize the training process.

[0077] It should be noted that for other corresponding descriptions of the functional units involved in the power equipment model identification device provided in this embodiment, please refer to Figure 1 As well as the corresponding descriptions in the above embodiments, they will not be repeated here.

[0078] The present application also provides a computer device, such as Figure 4 As shown, the computer device can specifically be a personal computer, a server, a network device, etc. The computer device includes a system bus, a processor, a memory, and a communication interface, and may also include an input / output interface and a display device. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store location information. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the steps in each method embodiment are implemented.

[0079] Those skilled in the art will understand that the structure of the above-mentioned computer device is only a partial structure related to the solution of the present application and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components, or combine certain components, or have a different component arrangement.

[0080] In one embodiment, a computer-readable storage medium is provided. The computer-readable storage medium may be non-volatile or volatile, and stores a computer program thereon. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0081] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0082] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, and the like.

[0083] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0084] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for identifying the model of an electric power device, characterized in that: The method comprises: Acquire a target power equipment image, and perform feature extraction on the target power equipment image based on a feature extraction module in a pre-trained text detection model to obtain multi-layer image features; Based on the feature fusion module in the text detection model, the multi-layer image features are fused to obtain fused features; Based on the feature enhancement module in the text detection model, the fused features are enhanced to obtain an enhanced feature sequence; Based on the output module in the text detection model, the enhanced feature sequence is predicted to obtain text box information corresponding to the text area in the target power equipment image; Character recognition is performed on the text box information to obtain the target power equipment model.

2. The method according to claim 1, characterized in that The feature extraction module in the text detection model includes multiple first convolutional layers and one second convolutional layer; the feature extraction module in the pre-trained text detection model performs feature extraction on the target power equipment image to obtain multi-layer image features, including: The target power equipment image is sequentially subjected to feature extraction by a plurality of first convolutional layers and a second convolutional layer in the feature extraction module, to obtain image features of different sizes corresponding to each convolutional layer; Among them, the second convolutional layer includes an original convolution branch and multiple hollow convolution branches, each of the hollow convolution branches is set with a different hollow rate, and the image features extracted by the original convolution branch and the hollow convolution branch are fused to obtain the image features corresponding to the second convolutional layer.

3. The method according to claim 2, characterized in that The feature fusion module in the text detection model performs feature fusion on multiple layers of image features to obtain fused features, including: The image features extracted by each convolutional layer are merged layer by layer from bottom to top through feature connection and convolution until the final fusion features are generated.

4. The method according to claim 3, characterized in that The feature enhancement module in the text detection model is based on which the fusion feature is enhanced to obtain an enhanced feature sequence, including: Serializing the fused features to obtain a one-dimensional feature vector sequence, and grouping the feature vector sequence according to a preset batch size to obtain multiple groups of feature sequences; Multiple groups of feature sequences are sequentially input into a long short-term memory neural network for bidirectional processing and sequentially spliced ​​to form a new one-dimensional feature vector sequence, thereby obtaining the enhanced feature sequence.

5. The method according to claim 1, characterized in that The output module in the text detection model includes a text area score branch and a text box parameter branch; the output module in the text detection model is used to predict the enhanced feature sequence to obtain text box information corresponding to the text area in the target power equipment image, including: Processing the enhanced feature sequence through the convolutional layer of the text region score branch to obtain an image score map, wherein the pixel value in the image score map represents the score of the pixel belonging to the text region; The convolutional layer of the text box output branch is used to predict the distance between each pixel in the image score map and the text area boundary and the rotation angle of the text area.

6. The method according to claim 1, wherein The training process of the text detection model includes: Acquire multiple images containing model information of electric power equipment, and perform processing such as rotating, scaling, and adding noise on the images to obtain multiple sample images; Each module in the text detection model is trained based on the sample image.

7. The method according to claim 6, characterized in that The training process of the text detection model also includes: The Dice softloss function and enhanced cosine loss function are used to optimize the training process.

8. A device for identifying the model of an electric power device, characterized in that: The device comprises: A feature extraction module is used to obtain a target power equipment image, perform feature extraction on the target power equipment image, and obtain multi-layer image features; A feature fusion module is used to fuse the multiple layers of image features to obtain fused features; A feature enhancement module, configured to enhance the fused features to obtain an enhanced feature sequence; A text output module, configured to predict the enhanced feature sequence to obtain text box information corresponding to the text area in the target power equipment image; The model recognition module is used to perform character recognition on the text box information to obtain the target power equipment model.

9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.