An intelligent inspection LED digital meter reading method based on improved SAR

Through improved SAR algorithms and lightweight models, combined with inspection robots and deep learning technology, efficient and accurate readings of LED digital instruments are achieved, identification problems and edge deployment problems in complex scenarios are solved, and instrument monitoring efficiency and accuracy are improved.

CN117173384BActive Publication Date: 2025-07-22BEIJING GREEN VALLEY TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311151406.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-07
Publication Date
2025-07-22
Estimated Expiration
2043-09-07

AI Technical Summary

Technical Problem

In the prior art, the reading method of LED digital instruments has low recognition accuracy in complex scenarios, large amount of parameters in deep learning models, difficult to deploy at the edge, low manual monitoring efficiency and prone to errors.

Method used

The improved SAR algorithm is used to combine ultra-fast and lightweight object detection model and LED digital instrument recognition SAR model, and images are collected through inspection robots, and feature extraction and character sequence recognition are used using resnet34, two-dimensional attention module and LSTM encoding/decoder to reduce the amount of model parameters and improve robustness.

Benefits of technology

It realizes high-precision recognition of LED digital instruments in complex scenarios, supports edge deployment, improves monitoring efficiency and recognition accuracy, solves the inefficiency problem of manual monitoring, and has real-time recognition capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117173384B_ABST
    Figure CN117173384B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent inspection LED digital instrument reading method based on improved SAR. By using an ultra-fast lightweight object detection model to identify and analyze the target instrument image collected by the inspection robot, the digital dial image is cut out from the target instrument image. Furthermore, the trained LED digital instrument recognition SAR model established by using the improved SAR algorithm is used to identify the digital dial image to obtain a character sequence, and finally the instrument reading of the target instrument is obtained by using the character sequence; the model is lightweight and can be deployed to the edge side, with a real-time recognition speed, solving the problems of rapid positioning and intelligent reading of digital instruments, and solving the problem of low efficiency of manual instrument monitoring, avoiding adverse factors caused by manual operation, improving the instrument monitoring efficiency, and improving the instrument reading recognition accuracy under conditions such as uneven illumination and instrument tilt. It not only has a high degree of automation but also is easy to implement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of instrument status monitoring, and in particular to an intelligent inspection LED digital instrument reading method based on improved SAR. Background Art

[0002] With the rapid development of the Chinese power grid, the scale and structure of the power grid have also undergone earth-shaking changes. The operation of monitoring instrument parameters has also experienced explosive growth. Among them, a large number of LED digital instruments are widely used in power scenarios such as substations. However, while the instrument operation brings convenience to power status monitoring, the large-scale instrument status monitoring task faces huge challenges.

[0003] Traditional methods collect and record the monitored data of LED digital instruments through manual operations. Facing a large number of instruments that need to collect and record data in power plants, manual operation has low efficiency and cannot always maintain work efficiency, and is prone to misrecording and misreading, resulting in errors in subsequent work.

[0004] With the development of computer vision-related technologies, especially the rapid progress of deep learning algorithms, the recording work of power plant instrument data has begun to develop in the direction of automation. This not only improves the data collection efficiency but also reduces the safety risks of data collectors during operation.

[0005] Currently, the main methods for reading LED digital instruments include template matching, which uses pre-defined digital templates to match the numbers in the input image. This method is sensitive to image deformation and noise, has a large amount of calculation, and is not suitable for complex scenarios; methods based on feature extraction and classification, which extract the features of numbers from the input image through edge detection, morphological operations, image segmentation, etc., and then use classifiers such as support vector machines (SVM) and neural networks to identify the numbers. This method is difficult to handle diverse digital styles in complex scenarios; methods based on deep learning, which use deep neural networks to extract features and identify numbers in the image. This method has good accuracy and robustness, but the number of parameters of the deep learning model is often large, and the performance of edge-end hardware devices needs to meet certain requirements, otherwise it will affect the speed of digital recognition, and even the model is difficult to be deployed at the edge end. Summary of the Invention

[0006] The purpose of the present invention is to provide an intelligent inspection LED digital instrument reading method based on improved SAR, which solves the above-mentioned technical problems pointed out in the prior art.

[0007] The present invention provides an intelligent inspection LED digital instrument reading method based on improved SAR, including the following operation steps:

[0008] The inspection robot collects and obtains the target instrument image of the target instrument; the target instrument is a digital instrument; the target instrument image is a digital instrument image;

[0009] Based on the target instrument image, the digital dial image is obtained through detection by a trained ultra-fast lightweight object detection model;

[0010] Based on the digital dial image, a character sequence is obtained through recognition by a trained LED digital instrument recognition SAR model;

[0011] According to the character sequence, the instrument reading corresponding to the target instrument is obtained.

[0012] Preferably, the structure of the trained LED digital instrument recognition SAR model includes resnet34, a two-dimensional attention module, and an LSTM encoder / decoder.

[0013] Preferably, the obtaining of the digital dial image by detecting the target instrument image through the trained ultra-fast lightweight object detection model includes the following operation steps:

[0014] Train the initially established ultra-fast lightweight object detection model to obtain a trained ultra-fast lightweight object detection model;

[0015] Input the target instrument image into the trained ultra-fast lightweight object detection model, and output and obtain the instrument detection result;

[0016] Intercept the digital dial image from the target instrument image according to the instrument detection result.

[0017] Preferably, the training of the initially established ultra-fast lightweight object detection model to obtain a trained ultra-fast lightweight object detection model includes the following operation steps:

[0018] Obtain a plurality of original instrument images; the original instrument images are collected and obtained by the inspection robot;

[0019] Annotate the original instrument images to obtain annotated instrument images;

[0020] Divide the annotated instrument images into a training set and a test set according to a preset ratio;

[0021] Use the training set and the test set to train the ultra-fast lightweight object detection model to obtain a trained ultra-fast lightweight object detection model.

[0022] Preferably, the output image size of the trained ultra-fast lightweight object detection model is 8500×33.

[0023] Preferably, the character sequence is obtained by identifying the LED digital instrument SAR model based on the digital dial image, and the operation steps are as follows:

[0024] Extract the feature sequence data from the digital dial image through the convolutional layer and the pooling layer;

[0025] Convert the feature sequence data into vector information through the encoder;

[0026] Analyze the feature sequence data to obtain the initial weight information corresponding to each area in the digital dial image; and obtain the extreme left area in the digital dial image and the extreme left initial weight information corresponding to the extreme left area;

[0027] The extreme left area is the leftmost area in the digital dial image;

[0028] After decoding and outputting the first character according to the extreme left initial weight information and the vector information with the same time step, use the extreme left initial weight information corresponding to the first character as the second weight information of the first character, and calculate and obtain the second weight information corresponding to each area by using the initial weight information of each area, the previous character, and the second weight information corresponding to the previous character; and decode and output the characters of each area according to the second weight information corresponding to each area and the vector information with the same time step;

[0029] Search the characters through the sequence generation strategy and output the character sequence through the regularization operation.

[0030] Preferably, the feature sequence data includes low-level feature data, high-level feature data, image association information, and resolution levels.

[0031] Preferably, the above regularization operation includes duplicate removal operation and empty segment removal operation.

[0032] Preferably, the above sequence generation strategy is the prefix beam search strategy.

[0033] Compared with the prior art, the embodiments of the present invention have at least the following technical advantages:

[0034] Analyzing the above-mentioned intelligent inspection LED digital meter reading method based on improved SAR provided by the present invention, it can be seen that in specific applications, first, the target meter image is collected by the inspection robot, and then the ultra-fast and lightweight object detection model is used to identify and analyze the target meter image, so as to cut out the digital dial image from the target meter image; the ultra-fast and lightweight object detection model has fewer parameters and faster inference speed, thus accelerating the detection and interception speed of the digital dial image; furthermore, the trained LED digital meter recognition SAR model established by using the improved SAR algorithm is used to identify the digital dial image to obtain the character sequence, which not only has better recognition accuracy in regular text recognition, but also can recognize irregular text (such as tilted text, etc.), and has stronger robustness. It can not only support the recognition of diverse styles of numbers in complex scenarios (such as tilted and irregular numbers caused by the shooting angle, etc.), solving the problem of low recognition accuracy of the model in irregular numbers in the existing deep learning-based technical methods; and the improved model is lightweight and can be deployed to the edge side, with the speed of real-time recognition, solving the problems of rapid positioning and intelligent reading of digital meters, and solving the problem of low efficiency of manual monitoring of meters, avoiding the adverse factors caused by manual operation, improving the meter monitoring efficiency, and improving the meter reading recognition accuracy under conditions such as uneven illumination and tilted meters. It not only has a high degree of automation but also is easy to implement; finally, the meter reading of the target meter is obtained by using the character sequence. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0036] Figure 1 It is a schematic diagram of the overall operation steps of an intelligent inspection LED digital meter reading method based on improved SAR provided in Embodiment 1 of the present invention;

[0037] Figure 2 It is a schematic simulation diagram of the target meter image collected by an intelligent inspection LED digital meter reading method based on improved SAR provided in Embodiment 1 of the present invention;

[0038] Figure 3 It is a schematic diagram of the operation steps for detecting and obtaining the digital dial image in an intelligent inspection LED digital meter reading method based on improved SAR provided in Embodiment 1 of the present invention;

[0039] Figure 4It is a schematic diagram of the simulation of the instrument detection result in an intelligent inspection LED digital instrument reading method based on improved SAR provided in the first embodiment of the present invention;

[0040] Figure 5 It is a schematic diagram of the simulation of the digital dial image obtained by intercepting in an intelligent inspection LED digital instrument reading method based on improved SAR provided in the first embodiment of the present invention;

[0041] Figure 6 It is a schematic diagram of the operation steps of obtaining a trained ultra-fast lightweight object detection model in an intelligent inspection LED digital instrument reading method based on improved SAR provided in the first embodiment of the present invention;

[0042] Figure 7 It is a schematic diagram of the NanoDet-Plus-m-1.5x model in an intelligent inspection LED digital instrument reading method based on improved SAR provided in the first embodiment of the present invention;

[0043] Figure 8 It is a schematic diagram of the operation steps of obtaining a character sequence in an intelligent inspection LED digital instrument reading method based on improved SAR provided in the first embodiment of the present invention. Detailed implementation manners

[0044] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0045] Next, the present invention will be further described in detail through specific embodiments in conjunction with the accompanying drawings.

[0046] Embodiment 1

[0047] As Figure 1 shown, the present invention proposes an intelligent inspection LED digital instrument reading method based on improved SAR, including the following operation steps:

[0048] Step S10: The inspection robot collects and obtains the target instrument image of the target instrument;

[0049] The target instrument is a digital instrument; the target instrument image is a digital instrument image;

[0050] It should be noted that in the above embodiments of the present application, after the inspection robot reaches the instrument target point based on the positioning and navigation algorithm (laser SLAM positioning and navigation algorithm), the point parameters of the instrument target point are extracted, and the camera pose of the inspection robot is adjusted according to the parameters of the instrument target point, and then the image acquisition instruction of the camera is triggered, so as to acquire the target instrument image (such as Figure 2 as shown);

[0051] The above positioning and navigation algorithm, extracting the point parameters of the instrument target point, and adjusting the camera pose are all prior arts, and will not be elaborated in the embodiments of the present application.

[0052] Step S20: Detect and obtain the digital dial image based on the target instrument image through the trained ultra-fast lightweight object detection model;

[0053] Step S30: Recognize and obtain the character sequence based on the digital dial image through the trained LED digital instrument recognition SAR model;

[0054] The structure of the trained LED digital instrument recognition SAR model includes resnet34, two-dimensional attention module, and LSTM encoder / decoder (the resnet34 structure can deepen the depth of the backbone network, improve the ability of the backbone network to extract feature information, and provide richer semantic information extracted by the backbone network for the subsequent LSTM encoder / decoder and 2D attention module, thus realizing the improvement of the model's recognition accuracy for LED digits);

[0055] It should be noted that the SAR (Show, Attend and Read: A Simple and Strong Baseline for Irregular Text Recognition) algorithm is an end-to-end deep neural network based on the attention mechanism. Compared with the current digital recognition methods based on deep learning (such as CRNN+CTC), etc., SAR not only has better recognition accuracy in regular text recognition, but also can recognize irregular texts (such as tilted texts), and has stronger robustness;

[0056] The above trained LED digital instrument recognition SAR model needs to first establish the LED digital instrument recognition SAR model, and then obtain a large amount of original instrument image data through the inspection robot; manually annotate the instrument rectangular frames for the original instrument image data; after the annotation process, divide the manually annotated instrument rectangular frames into a training set and a test set; and train the LED digital instrument recognition SAR model through the training set and the test set to obtain the trained LED digital instrument recognition SAR model; the establishment and training process of the above trained LED digital instrument recognition SAR model is based on the improvement of the above SAR algorithm;

[0057] The above-mentioned LED digital instrument recognition SAR model finally outputs a vector through softmax. Among the predicted vectors of each character, the one with the highest probability is taken as the representative to obtain a character sequence. The CTC loss is directly used to solve the alignment problem, allowing the model to freely insert blank tags and repeated tags when generating text, so as to better adapt to the length difference between the input and the output. In contrast, in the original SAR algorithm, the cross-entropy loss function is used to compare the difference between the final character sequence and the true label. This method has the disadvantage of difficult alignment. Specifically,

[0058] Difficult alignment: Cross-entropy loss usually requires character-level alignment, that is, precisely aligning each output time step with the corresponding part of the input image. However, in character recognition tasks, there may be spaces, overlaps, or irregular alignments between characters, which makes character-level alignment difficult and thus affects the recognition accuracy of the model. Therefore, using cross-entropy loss may require additional alignment steps or dealing with complex alignment problems, increasing the complexity and development cost of the system;

[0059] In the above-mentioned recognition and analysis of the digital dial image by the trained LED digital instrument recognition SAR model of this application, the SAR algorithm is improved, and the improved SAR algorithm is used for digital recognition. The specific operations are shown in the following steps S31 - S35;

[0060] The SAR algorithm (i.e., the above-mentioned trained LED digital instrument recognition SAR model) in the embodiment of this application can reduce the length of the character set that the SAR model needs to recognize according to the characteristic that the target scene character set of the digital instrument is small, from the original 9 digits, 26 lowercase letters, 26 uppercase letters, and 33 punctuation marks to 10 digits (0 - 9) and 2 punctuation marks (negative sign "-", decimal point "."). By narrowing the range of the character set that the model focuses on, the reading accuracy of the model on LED digital characters is improved. Moreover, this improvement also reduces the number of parameters of the SAR model and enhances the recognition speed of the model;

[0061] Step S40: Obtain the instrument reading corresponding to the target instrument according to the character sequence.

[0062] It should be noted that the method for identifying LED digital meters using deep learning technology in the above embodiments of the present application solves the problems of rapid positioning and intelligent reading of digital meters, and solves the problem of low efficiency of manual meter monitoring. It avoids adverse factors caused by manual operations, improves the meter monitoring efficiency, and improves the accuracy of meter reading recognition under conditions such as uneven illumination and inclined meters. It not only has a high degree of automation but is also easy to implement. At the same time, the method for identifying LED digital meters using deep learning technology solves the problems of large model parameter quantities, insufficient accuracy, and slow recognition speed in digital meter reading based on deep learning. The method proposed by the present invention significantly improves the recognition accuracy and has very good robustness in scenarios where the numbers are angled and tilted.

[0063] Specifically, as Figure 3 shown, in step S20, based on the target meter image, a trained ultra-fast lightweight object detection model is used for detection to obtain a digital dial image, including the following operation steps:

[0064] Step S21: Establish an ultra-fast lightweight object detection model; train the ultra-fast lightweight object detection model to obtain a trained ultra-fast lightweight object detection model;

[0065] Step S22: Input the target meter image into the trained ultra-fast lightweight object detection model, and output the meter detection result;

[0066] Step S23: Intercept the digital dial image from the target meter image according to the meter detection result.

[0067] It should be noted that the meter detection result in the above embodiments of the present application is represented by (x1, y1, x2, y2) as Figure 4 shown, where (x1, y1) represents the coordinates of the upper left corner point of the digital dial image, and (x2, y2) represents the coordinates of the lower right corner point of the digital dial image. Thus, the digital dial image can be intercepted from the target meter image according to the meter detection result (as Figure 5 shown);

[0068] The embodiments of the present application use the NanoDet-Plus object detection algorithm for meter detection, which has fewer parameters and faster inference speed, thus accelerating the detection and interception speed of the digital dial image.

[0069] Specifically, as Figure 6 shown, in step S21, training the ultra-fast lightweight object detection model to obtain a trained ultra-fast lightweight object detection model includes the following operation steps:

[0070] Step S211: Obtain multiple original instrument images; the original instrument images are collected by an inspection robot;

[0071] Step S212: Annotate the original instrument images to obtain annotated instrument images;

[0072] It should be noted that in the above embodiments of the present application, the digital dial images in the original instrument images are annotated by manual annotation (annotating the upper left point coordinates and lower right point coordinates of the digital dial images) to obtain the instrument rectangular frames.

[0073] Step S213: Divide the annotated instrument images into a training set and a test set according to a preset ratio;

[0074] Step S214: Use the training set and the test set to train the ultra-fast lightweight object detection model to obtain a trained ultra-fast lightweight object detection model;

[0075] The output image size of the trained ultra-fast lightweight object detection model is 8500×33.

[0076] It should be noted that the NanoDet-Plus detection model (i.e., the above ultra-fast lightweight object detection model) in the embodiments of the present application samples the NanoDet-Plus-m-1.5x model of open-source NanoDet-Plus (with fewer parameters and faster inference speed) as a prototype. As follows Figure 7 shown, the input is an image with a size of 416*416*3, and its output is modified to 8500*33 (originally 8500*112). Since there is only 1 detection category in the embodiments of the present application, the original structure supports 80 categories.

[0077] Specifically, as Figure 8 shown, in step S30, based on the digital dial image, the trained LED digital instrument recognition SAR model is used for recognition to obtain a character sequence, including the following operation steps:

[0078] Step S31: Extract feature sequence data (i.e., the image features of the digital dial image) from the digital dial image through a convolutional layer and a pooling layer;

[0079] The feature sequence data includes low-level feature data, high-level feature data, image association information, and resolution levels;

[0080] It should be noted that the above low-level features refer to the edge, texture, and color information in the digital dial image, and these features are very important for the representation of the basic structure and texture of the image;

[0081] The above-mentioned high-level features refer to the fact that as the depth of the network increases, higher-level features will be gradually extracted, such as the shape, part combination, and overall structure of the objects in the digital dial image;

[0082] The above-mentioned image correlation information (image correlation information or image context information) refers to obtaining the image correlation information of the image regions in the digital dial image through multi-layer processing. These image correlation information can help the model understand the relationships between different objects, as well as the position and size relationships of the objects in the scene;

[0083] The above-mentioned resolution hierarchy refers to reducing the spatial resolution of the feature map through successive downsampling operations, which results in obtaining broader image correlation information at lower feature levels, while higher feature levels pay more attention to local details.

[0084] Step S32: Convert the feature sequence data into vector information through an encoder;

[0085] It should be noted that in the present application, the above-mentioned LSTM is used as an encoder to serialize the extracted image features into a vector representation of a fixed length. LSTM captures the image correlation information in the image through step-by-step processing of the feature sequence data;

[0086] For example: When extracting the feature sequence data in the above-mentioned step S31, assume that the height of the feature map is 100, the width is 100, and the number of channels is 5; thus, in step S32, the feature sequence data is converted into vector information. By downsampling the height through pooling to 1 while keeping the width and the number of channels unchanged, 100 vector information (each vector is 5-dimensional) is obtained; LSTM performs step-by-step processing on the previously obtained vector information to capture the image correlation information in the image;

[0087] The above-mentioned feature map is the feature map output by the resnet34 network. The resnet34 network consists of many convolutional layers and pooling layers) and outputs this feature map. Before LSTM parses / analyzes this feature, it needs to serialize it into a vector (achieved through max pooling).

[0088] Step S33: Analyze the feature sequence data to obtain the initial weight information corresponding to each region in the digital dial image; and obtain the extreme left region in the digital dial image and the extreme left initial weight information corresponding to the extreme left region;

[0089] The extreme left region is the leftmost region in the digital dial image;

[0090] Step S34: Decode according to the extreme left initial weight information and the vector information at the same time step to output the first character;

[0091] It should be noted that the above-described embodiment of the present application introduces a two-dimensional attention module at the decoder end, which is used to focus on different regions of the input image when generating each character; this two-dimensional attention module helps the trained LED digital meter recognition SAR model focus on the text region and suppress background interference; the two-dimensional attention module calculates the attention weights based on the output of the encoder and the previously generated characters to select the image region corresponding to the current character.

[0092] When generating each character, the above two-dimensional attention module adjusts the attention region of the trained LED digital meter recognition SAR model for the input image (i.e., the above digital dial image) through weights, so that the trained LED digital meter recognition SAR model focuses as much as possible on the region of the target character to be generated currently in the image.

[0093] The above adjusts the text region that the trained LED digital meter recognition SAR model focuses on and suppresses background interference through weight parameters. The initial weights are obtained by convolving the image features (feature map) through some convolutional layers. These convolutional layers are trained to implement this attention mechanism, and the generated second weights can help the model focus on the text region and suppress background interference.

[0094] When the decoder decodes the output vector sequence of the encoder, it will parse this vector information multiple times (the specific number of times is a hyperparameter set manually). The parsing is from left to right (which can be understood as parsing the picture from left to right). Each time the decoder parses, it uses the content parsed in the previous time according to the memory mechanism of the LSTM model. Through the content parsed in the previous time and the image features (feature map), the two-dimensional attention module generates weight information, and uses this weight information to help the model focus more on the image region corresponding to the character to be parsed in this parsing.

[0095] Among them, in the first parsing, the attention module only uses the image features to generate weight information.

[0096] Step S35: Obtain the initial weight information corresponding to the second region; calculate and obtain the second weight information corresponding to the second region according to the initial weight information corresponding to the second region, the first character, and the extremely left initial weight information corresponding to the first character; and decode and output the second character according to the second weight information and the vector information at the same time step; return the second character and the second weight information corresponding to the second character, and repeat the above operations to obtain the characters corresponding to each region in the digital dial image; until all the characters in all regions are output.

[0097] The second region is the first region to the right of the extremely left region.

[0098] It should be noted that returning the second character and the second weight information corresponding to the second character as described above, and repeating the above operations to obtain the characters corresponding to each region in the digital dial image means obtaining the initial weight information of the third region (i.e., the first region to the right of the second region), and calculating the second weight information corresponding to the third region based on the second character, the second weight information corresponding to the second character, and the initial weight information of the third region; then decoding and outputting the third character according to the second weight information of the third region and the vector information with the same time step; obtaining the initial weight information of the fourth region (i.e., the first region to the right of the third region), calculating the second weight information corresponding to the fourth region based on the third character, the second weight information corresponding to the third character, and the initial weight information of the fourth region; and decoding and outputting the fourth character according to the second weight information of the fourth region and the vector information with the same time step...

[0099] In the above embodiment of the present application, LSTM is used as the decoder, which receives the output vector information of the encoder and the output weight information of the two-dimensional attention module as inputs. The decoder gradually generates a character sequence, generates a character at each time step, and adjusts according to the previously generated characters and attention weights;

[0100] Among them, different weight information will be generated each time decoding is performed, which will play an adjustment role on the weight information. The specific generated characters are the outputs of the decoder decoding (the outputs jointly generated by the LSTM module and the attention module), which are composed of multiple vectors.

[0101] Step S36: Search for the characters through a sequence generation strategy and then output a character sequence through a regularization operation;

[0102] The above regularization operations include duplicate removal operations, empty segment removal operations, etc.;

[0103] The above sequence generation strategy is a prefix beam search strategy;

[0104] It should be noted that the above uses prefix beam search to decode the predicted probability matrix, improving the credibility of the final recognition; the original SAR model used greedy search or beam search algorithm to decode the predicted probability matrix; among them, the greedy search algorithm is to find the character with the highest probability in each time step as the decoding output, and then output the final recognized character after regularization. This method does not consider the relationship between time steps before and after, and there may be errors in probability selection;

[0105] The beam search algorithm is to find the N (here N is a hyperparameter, and when N = 1, it is equivalent to the greedy search algorithm) characters with the highest probability in each time step as the decoding output, which can utilize the information of the previous and subsequent time steps to a certain extent, but may also not be able to cover the global optimal solution and fall into a local optimal solution.

[0106] This application uses the prefix beam search algorithm to decode the predicted probability matrix by generating prefix sequences, which can better utilize the image correlation information, improve the overall recognition accuracy and coherence. And generating multiple optimal prefixes provides more choices and can better handle the complex situations in the industrial environment;

[0107] Compared with the original SAR algorithm for character text recognition, especially for irregular character text recognition, it shows good performance. However, the number of parameters in its LSTM encoding module and decoding model is too large, resulting in a bloated model that is difficult to deploy to edge devices, and the recognition speed is slow, making it difficult to meet real-time recognition;

[0108] Based on the steps of the SAR for digital text recognition, this application improves and optimizes the SAR model, realizes the lightweight of the SAR model, can be deployed to edge devices, improves the detection speed, can meet real-time recognition, and improves the recognition accuracy of digital characters;

[0109] According to the characteristics of the short length of the target scene character text area of the LED digital meter in the above embodiments of this application, the LSTM encoder and decoder are improved. By reducing the number of input channels of the LSTM structure and the dimension of the hidden state of the LSTM unit block, while ensuring the unchanged recognition accuracy of the LED digital in the model, the number of model parameters is reduced, and the recognition speed of the model is improved;

[0110] The technical solution adopted in the above embodiments of this application can achieve a recognition accuracy of 99% on the test set using the improved SAR model, including irregular LED digital images caused by shooting angles.

[0111] In summary, an intelligent inspection method for LED digital meter readings based on improved SAR proposed in the embodiments of the present invention collects a target meter image (the target meter image in this application specifically refers to a digital meter image) through an inspection robot, and identifies the target meter image through an ultra-fast lightweight object detection model, so as to cut out the digital dial image. With the characteristics of fewer parameters and faster inference speed of the ultra-fast lightweight object detection model, the detection and interception speed of the digital dial image is accelerated;

[0112] Further, the digital dial image is subjected to feature sequence acquisition, encoding processing, decoding processing output, sequence generation strategy, and regularization operation through a trained LED digital instrument recognition SAR model trained by an improved SAR algorithm, and then a character sequence (i.e., the target character sequence) is output; among them, aiming at the characteristic of the small character set of the digital instrument image, the model's attention character set range is narrowed, the reading accuracy of the model on LED digital characters is improved, and at the same time, the number of SAR model parameters is reduced, and the model recognition speed is increased; and the resnet34 structure is used to deepen the depth of the backbone network, improve the ability of the backbone network to extract feature information, and provide richer semantic information extracted by the backbone network for the subsequent LSTM encoder and 2D attention module, thus realizing the improvement of the model's recognition accuracy of LED digits; according to the characteristic that the length of the target scene character text area of the LED digital instrument is short, by reducing the number of input channels of the LSTM structure and the dimension of the hidden state of the LSTM unit block, while ensuring that the model's recognition accuracy of LED digits remains unchanged, the number of model parameters is reduced, and the model's recognition speed is increased; the CTC loss is used to solve the alignment problem, allowing the model to freely insert blank markers and repeated markers when generating text, so as to better adapt to the length difference between the input and the output; the prefix beam search is used to decode the predicted probability matrix, improving the credibility of the final recognition;

[0113] Finally, the target character sequence obtained through recognition and analysis is returned to the target instrument image to obtain the instrument reading.

[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; those of ordinary skill in the art can modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An intelligent inspection LED digital instrument reading method based on improved SAR, characterized in that, It includes the following operation steps: The inspection robot collects and obtains the target instrument image of the target instrument; the target instrument is a digital instrument; the target instrument image is a digital instrument image; Based on the target instrument image, the digital dial image is obtained by detecting through a trained ultra-fast lightweight object detection model; Based on the digital dial image, the character sequence is obtained by identifying through a trained LED digital instrument recognition SAR model; According to the character sequence, the instrument reading corresponding to the target instrument is obtained; The structure of the trained LED digital instrument recognition SAR model includes resnet34, a two-dimensional attention module, and an LSTM encoder / decoder; The step of obtaining the digital dial image by detecting through a trained ultra-fast lightweight object detection model based on the target instrument image includes the following operation steps: Train the initially established ultra-fast lightweight object detection model to obtain a trained ultra-fast lightweight object detection model; Input the target instrument image into the trained ultra-fast lightweight object detection model, and output the instrument detection result; Intercept the digital dial image from the target instrument image according to the instrument detection result; The step of obtaining the character sequence by identifying through a trained LED digital instrument recognition SAR model based on the digital dial image includes the following operation steps: Based on the digital dial image, feature sequence data is extracted through a convolutional layer and a pooling layer; The feature sequence data is converted into vector information through an encoder; According to the feature sequence data analysis, the initial weight information corresponding to each region in the digital dial image is obtained; and the extreme left region in the digital dial image and the extreme left initial weight information corresponding to the extreme left region are obtained; After decoding and outputting the first character according to the extreme left initial weight information and the vector information with the same time step, use the extreme left initial weight information corresponding to the first character as the second weight information of the first character, and calculate and obtain the second weight information corresponding to each region by using the initial weight information of each region and the previous character and the second weight information corresponding to the previous character; and decode and output the characters of each region according to the second weight information corresponding to each region and the vector information with the same time step; The characters are searched through a sequence generation strategy and then output as a character sequence through a regularization operation.

2. The intelligent inspection LED digital instrument reading method based on improved SAR according to claim 1, characterized in that, The step of training the initially established ultra-fast lightweight object detection model to obtain a trained ultra-fast lightweight object detection model includes the following operation steps: Obtain multiple original instrument images; the original instrument images are collected and obtained by the inspection robot; Annotate the original instrument images to obtain annotated instrument images; Divide the annotated instrument images into a training set and a test set according to a preset ratio; Use the training set and the test set to train the ultra-fast lightweight object detection model to obtain a trained ultra-fast lightweight object detection model.

3. The intelligent inspection LED digital instrument reading method based on improved SAR according to claim 2, wherein The feature sequence data includes low-level feature data, high-level feature data, image association information, and resolution levels.

4. The intelligent inspection LED digital meter reading method based on improved SAR according to claim 3, characterized in that, The extreme left region is the leftmost region in the digital dial image.

5. A method for reading LED digital instrument readings in intelligent patrol inspection based on improved SAR according to claim 4, characterized in that, The above regularization operations include duplicate removal operations and empty segment removal operations.

6. The intelligent inspection LED digital instrument reading method based on improved SAR according to claim 5, characterized in that, The above sequence generation strategy is a prefix beam search strategy.

Citation Information

Patent Citations

  • Automatic reading method of pointer instrument, electronic equipment and storage medium

    CN114782940A