A nameplate recognition method and device

By employing a recognition method that combines nameplate image feature extraction with a multi-strategy collaborative mechanism, the problem of low accuracy in nameplate recognition under complex backgrounds is solved. This enables high-precision automated nameplate recognition and information entry, improving the objectivity and intelligence of the detection process.

CN121904770BActive Publication Date: 2026-06-19NANCHANG KECHEN ELECTRIC POWER TEST & RES CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-03-25
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

In complex backgrounds or low-contrast situations, the recognition accuracy of appliance nameplates is low, and missegmentation, missed detection, or misalignment are prone to occur, making it impossible to distinguish the positions of different semantic fields.

Method used

By acquiring nameplate images, feature extraction and matching of target nameplate templates are performed. Combined with horizontal and vertical transition detection, connected component analysis, and table line detection, the content region is determined, and target detection and content recognition are performed to improve recognition accuracy.

Benefits of technology

Under complex imaging conditions, it improves the accuracy of nameplate recognition, realizes the automated and accurate entry of sample information, overcomes the low efficiency of manual inspection, and supports the data objectivity and traceability of blind sample inspection process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121904770B_ABST
    Figure CN121904770B_ABST
Patent Text Reader

Abstract

This application discloses a nameplate recognition method and apparatus, belonging to the field of image processing technology. The method includes: acquiring a nameplate image of a nameplate to be recognized; extracting features from the nameplate image and matching a target nameplate template based on the extracted features; detecting and extracting the content region of the nameplate image according to the content layout of the target nameplate template to obtain a content image; performing horizontal and vertical transition detection and connected component analysis on the content image to determine the content region, and performing horizontal and vertical table line detection on the content image to obtain a deprojected region; determining the content localization result of the nameplate to be recognized based on the content region and the deprojected region; and performing target detection and content recognition on the content localization result based on the content image to obtain the recognition result of the nameplate to be recognized. This method improves the recognition accuracy of nameplates under complex imaging conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of image processing technology, and particularly relates to a nameplate recognition method and apparatus. This invention can be widely applied to quality inspection and asset management in fields such as power equipment and industrial manufacturing, and is especially suitable for blind sample inspection systems, serving as a key technical means to achieve automated and accurate entry of sample information. Background Technology

[0002] Electrical nameplates are an important component of the electronics industry, playing an increasingly important role in modern electronic equipment. With the advent of the information age and the mass production of electrical nameplates, they play a significant role in the packaging, inventory, and verification of electrical equipment.

[0003] Unlike ordinary document characters with a white background, electrical appliance nameplate characters typically appear against a colored background. These characters include text, numbers, and letters, and recognizing them is crucial for equipment management.

[0004] However, in complex backgrounds or low-contrast situations, missegmentation, missed detection, or misaligned recognition can easily occur, making it impossible to distinguish the positions of different semantic fields, resulting in low recognition accuracy of appliance nameplates. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention proposes a nameplate recognition method and device, which improves the recognition accuracy of nameplates under complex imaging conditions.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] In a first aspect, the present invention provides a nameplate identification method, the method comprising:

[0008] Obtain the image of the nameplate to be identified;

[0009] Feature extraction is performed on the nameplate image, and the target nameplate template of the nameplate to be identified is matched based on the extracted nameplate features;

[0010] Based on the content layout of the target nameplate template, the content area of ​​the nameplate image is detected and extracted to obtain the content image;

[0011] The content image is subjected to horizontal and vertical jump detection and connected component analysis to determine the content region, and the content image is subjected to horizontal and vertical table line detection to obtain the deprojected region.

[0012] Based on the content area and the deprojection area, the content positioning result of the nameplate to be identified is determined;

[0013] Based on the content image, target detection and content recognition are performed on the content location result to obtain the recognition result of the nameplate to be identified.

[0014] Furthermore, the nameplate recognition method can be integrated into a blind sample inspection system for equipment, serving as a key technology for automatically acquiring sample information. In this application scenario, acquiring the nameplate image of the nameplate to be identified specifically includes: during the sample receiving stage of the blind sample inspection system, photographing the nameplate of the actual equipment to be inspected using an image acquisition device. After obtaining the recognition result, the method further includes: using the recognition result as structured nameplate information to automatically fill in or generate the digital sample information of the sample to be inspected in the blind sample inspection system.

[0015] Secondly, this application provides a nameplate identification device, the device comprising:

[0016] The acquisition module is used to acquire the image of the nameplate to be identified;

[0017] The first processing module is used to extract features from the nameplate image and match the target nameplate template of the nameplate to be identified based on the extracted nameplate features.

[0018] The second processing module is used to detect and extract the content area of ​​the nameplate image according to the content layout of the target nameplate template to obtain the content image;

[0019] The third processing module is used to perform horizontal and vertical jump detection and connected component analysis on the content image to determine the content region, and to perform horizontal and vertical table line detection on the content image to obtain the deprojected region.

[0020] The fourth processing module is used to determine the content positioning result of the nameplate to be identified based on the content area and the deprojection area.

[0021] The fifth processing module is used to perform target detection and content recognition on the content location result based on the content image to obtain the recognition result of the nameplate to be recognized.

[0022] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the nameplate recognition method as described in the first aspect above.

[0023] The advantages of this invention compared to existing technologies are as follows:

[0024] (1) By leveraging the prior characteristics of fixed nameplate content layout, structure-guided recognition can be achieved, which can suppress complex interferences common in industrial sites such as table lines, shadows, and reflections. Through a multi-strategy collaborative mechanism of content image fusion, horizontal and vertical jump detection, connected component analysis, and table line detection, the content area and deprojection area are accurately generated, and fine-grained text area localization is performed. Combined with layout constraints to eliminate invalid interference areas, it can focus on high-confidence effective text, improve the recognition accuracy of nameplates under complex imaging conditions, overcome the shortcomings of low efficiency in manual inspection, and make online inspection work more objective, standardized, and intelligent.

[0025] (2) By integrating color statistical characteristics and texture structure information to construct a comprehensive feature vector with high discriminative power, the limitations of single features in complex industrial scenarios are overcome: color features can distinguish nameplates of different materials or spraying processes, texture features are highly sensitive to structural differences such as character layout and border style, and through similarity discrimination mechanism, the correct nameplate template can still be accurately matched under conditions of light change, slight dirt or shooting angle shift.

[0026] (3) The nameplate recognition method provided by the present invention improves the recognition accuracy under complex conditions by integrating horizontal and vertical jump detection, connected component analysis and table line detection. When the method is applied to the blind sample detection system of equipment, it can realize the automated and accurate entry of sample information, avoid manual entry errors from the source, and effectively support the data objectivity and traceability of the blind sample detection process. Attached Figure Description

[0027] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0028] Figure 1 This is one of the flowcharts illustrating a nameplate identification method provided in this application embodiment;

[0029] Figure 2 This is a second schematic flowchart of a nameplate identification method provided in an embodiment of this application;

[0030] Figure 3 This is a schematic diagram illustrating the effect of removing horizontal and vertical table lines according to an embodiment of this application;

[0031] Figure 4 This is a flowchart illustrating the method for constructing a nameplate template library provided in an embodiment of this application;

[0032] Figure 5 This is a schematic diagram illustrating the positional relationship between image sub-blocks and overlapping windows provided in an embodiment of this application;

[0033] Figure 6This is a schematic diagram of the process for obtaining a nameplate image provided in an embodiment of this application;

[0034] Figure 7 This is a schematic diagram of the structure of a nameplate recognition device provided in an embodiment of this application;

[0035] Figure 8 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0036] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0037] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0038] The nameplate recognition method, nameplate recognition device, electronic device, and readable storage medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0039] The nameplate recognition method can be applied to the terminal, specifically by the nameplate recognition software in the terminal.

[0040] like Figure 1 As shown, the nameplate identification method includes:

[0041] Step 110: Obtain the image of the nameplate to be identified.

[0042] Understandably, the nameplate to be identified is a metal or plastic label attached to the surface of industrial equipment such as motors, transformers, and distribution boxes, used to identify structured information such as model, parameters, and manufacturer.

[0043] Nameplate images are obtained by preprocessing captured images, such as enhancement, correction, and segmentation.

[0044] In step 110, as Figure 2As shown, under good lighting conditions, the user accesses and controls the camera to take a picture of the appliance nameplate through the nameplate recognition software. The nameplate area must be complete, containing at least all character information, and there should be no characters stuck together, rusted characters, or folds. The nameplate to be recognized needs to have a corresponding nameplate template library. Through image enhancement, correction, and segmentation, a clear and upright nameplate image is obtained. The nameplate recognition software then recognizes the nameplate image, obtains the nameplate information, and outputs it.

[0045] Step 120: Extract features from the nameplate image and match the target nameplate template of the nameplate to be identified based on the extracted nameplate features.

[0046] Understandably, feature extraction involves calculating discriminative numerical vectors from an image, and the resulting nameplate features are used to characterize the visual attributes of the image.

[0047] Nameplate features include color histograms, color distribution such as mean, texture patterns such as HOG directional gradient histograms, and other global or local statistics.

[0048] The target nameplate template is matched from a pre-built nameplate template library. In the nameplate template library, each standard template contains a fixed content layout, field positions, field semantics, and corresponding feature vectors.

[0049] Since different manufacturers and equipment use different nameplate types, all nameplates in the samples need to be categorized and managed according to manufacturer and equipment type. Specifically, a nameplate template library is established based on manufacturer and equipment name type to facilitate subsequent algorithm processing and increase the algorithm's relevance and accuracy. When photographing nameplates, background interference and shooting angle can affect the image, requiring preprocessing through nameplate positioning and correction. If there are subsequent nameplate type updates, samples with good image quality can be preprocessed and used as standard templates in the nameplate template library, or the samples can be deformed and tilted by marking vertices and calculating perspective / affine transformation parameters.

[0050] like Figure 4 As shown, when constructing the nameplate template library, in addition to labeling all nameplates as different types based on the manufacturer's name and equipment name, for each type of standard template nameplate, it is also necessary to complete the layout detection of the corresponding nameplate image and define the area to be identified and its physical meaning. Through image preprocessing operations such as grayscale conversion, smoothing and noise filtering, edge detection, local adaptive binarization, horizontal and vertical table line removal, horizontal and vertical jump detection, morphological noise filtering, and connected component analysis, the layout on the nameplate is automatically detected, including tables, manufacturer's name, equipment type, physical quantity name, reading, unit, number, etc.

[0051] The system automatically detects the layout of the nameplate template, including tables, manufacturer names, equipment types, physical quantity names, readings, units, and serial numbers.

[0052] After binarization, images lose a lot of original information. The color depth of different areas in the nameplate is different. Fixed threshold binarization methods cannot preserve the texture and edge information in the image. A local adaptive threshold correction method is needed. For each pixel, threshold segmentation is performed within the surrounding rectangle with the current pixel as the center.

[0053] After image binarization, text localization needs to be performed on the image. Based on the pixel transition information, a distance threshold is taken, and points in the same row whose distance between two adjacent transition points is less than the threshold are connected to form a white line segment. The start and end columns of each line segment in each row are recorded. The same steps are performed for row transitions. Based on the reliability of the connected component analysis region, the text regions that meet the requirements are determined, and the coarse localization results are obtained.

[0054] like Figure 3 As shown, the coarse positioning result is projected horizontally and vertically using the projection method. Since each character region has been segmented and tilt correction has been performed, the projection regions with large gradient changes can be removed using horizontal and vertical projection, thus achieving the effect of removing horizontal and vertical table lines. Based on the horizontal and vertical jump detection and table line detection and removal results, the fine character positioning result of the nameplate template is obtained.

[0055] In step 120, the nameplate recognition software calculates color features such as the color histogram and three-channel mean in the hue, saturation, and brightness (HSV) space, as well as texture features such as the gradient histogram in the 8-direction HOG direction for the nameplate image. The two are concatenated into a comprehensive feature vector. The comprehensive feature vector is then compared with the feature vector of each pre-stored standard template in the nameplate template library using weighted cosine similarity to calculate a similarity score. The standard template with the highest score that exceeds the similarity threshold is selected as the target nameplate template. If no template meets the condition, an exception handling process is triggered, prompting manual review or template addition.

[0056] Step 130: Based on the content layout of the target nameplate template, detect and extract the content area of ​​the nameplate image to obtain the content image.

[0057] Understandably, content layout refers to the position, size, and arrangement of each text field predefined in the target nameplate template.

[0058] The content area is the sub-region in the nameplate image that actually carries the text information, located within the field box specified in the target nameplate template.

[0059] The content image contains only sub-images with valid text content, without borders, logos, or blank backgrounds.

[0060] In step 130, the coordinates of each field area recorded in the target nameplate template are read and mapped onto the original nameplate image. If there is a shooting angle deviation in the nameplate image, the template and the image are aligned through affine transformation or perspective correction. All field areas are merged to generate a compact mask area, and pixel data is extracted from it to form a content image.

[0061] Step 140: Perform horizontal and vertical jump detection and connected component analysis on the content image to determine the content region, and perform horizontal and vertical table line detection on the content image to obtain the deprojected region.

[0062] Understandably, horizontal and vertical transition detection detects abrupt changes in grayscale or gradient along the horizontal and vertical directions to estimate character row and column boundaries.

[0063] The content region is a set of high-confidence text regions after joint verification by transitions and connected components.

[0064] The deprojection region is a set of interference intervals formed by the expansion and projection of table lines, used to correct the projection.

[0065] In step 140, the gray-level difference sequences in the horizontal and vertical directions of the content image are calculated, and the gradient direction consistency is calculated in combination with multi-directional Sobel gradients. Jump points with gradient direction consistency higher than the threshold are retained. At the same time, connected component analysis is performed on the binarized image. Connected components composed of adjacent foreground pixels in the binary image are marked and attributed. Connected components that meet the character shape constraints are selected. The connected components are projected onto the row and column directions to generate density curves to eliminate false jump lines without character support, thereby determining reliable content regions. The horizontal and vertical table lines are identified by the line segment detection algorithm. Their normals are expanded and the parts covering the characters are removed to eliminate interference. The horizontal and vertical interference intervals are then projected to generate the final deprojected region.

[0066] Step 150: Determine the content positioning result of the nameplate to be identified based on the content area and the deprojection area.

[0067] Understandably, the content location result is a structured collection of regions that precisely mark the location of each text field, represented in the form of rectangles or masks.

[0068] In step 150, spatial logic operations are performed on the content area and the deprojection area to exclude the areas in the content area that overlap with the deprojection area and are interfered with by table lines; then, the remaining area is divided into grids according to the jump lines, and only the cells containing valid characters and whose grid size conforms to the a priori row height / column width of the nameplate are retained; each retained cell is fine-tuned based on the character centroid shrinkage to generate a set of non-overlapping, semantically complete text positioning boxes as the content positioning result.

[0069] Step 160: Based on the content image, perform target detection and content recognition on the content location result to obtain the recognition result of the nameplate to be recognized.

[0070] Understandably, object detection involves locating the position of a single character or block of text within an image.

[0071] Content recognition is the process of converting text in an image into a machine-readable string.

[0072] The recognition result is to output structured nameplate information from the character recognition results of each area, such as model: ABC-123, voltage: 380V.

[0073] In step 160, each bounding box in the content localization result is taken as a region of interest (ROI), and the corresponding sub-image is cropped from the content image; for each sub-image, a lightweight optical character recognition (OCR) engine based on PP-OCR is used to perform character recognition to obtain the original string; the recognition result is mapped to fields based on the field semantics of the target nameplate template, for example, the second row and first column should be the model number, and post-processing error correction is performed using a domain dictionary and regular rules to generate and output a structured nameplate recognition result.

[0074] According to the nameplate recognition method provided in the embodiments of this application, the structure-guided recognition is achieved by leveraging the prior characteristics of the fixed layout of the nameplate content. This method can suppress complex interferences commonly found in industrial environments, such as table lines, shadows, and reflections. Through a multi-strategy collaborative mechanism that integrates horizontal and vertical jump detection, connected component analysis, and table line detection of the content image, the method accurately generates the content region and the deprojection region, performs fine-grained text region localization, and eliminates invalid interference regions by combining layout constraints. This allows the method to focus on high-confidence valid text, improves the recognition accuracy of nameplates under complex imaging conditions, overcomes the inefficiency of manual inspection, and makes online inspection work more objective, standardized, and intelligent.

[0075] In some embodiments, the nameplate features include color features and texture features, and the step of extracting features from the nameplate image and matching the target nameplate template of the nameplate to be identified based on the extracted nameplate features includes:

[0076] The color features are obtained by extracting the color histogram and color mean from the nameplate image.

[0077] The texture features are obtained by extracting the directional gradient histogram from the nameplate image;

[0078] Based on the color features and the texture features, a comprehensive feature vector is constructed;

[0079] The similarity between the comprehensive feature vector and the template feature vector of each nameplate template is calculated using a discriminant function.

[0080] Based on the similarity, the target nameplate template is determined among the multiple nameplate templates.

[0081] Understandably, color features are vectors used to describe the statistical characteristics of image color in a nameplate image, such as color histograms and mean values. A color histogram is a vector representing the frequency distribution of pixel values ​​in each of the red, green, and blue (RGB) color channels in the system nameplate image; the color mean is the average value of all pixels in the nameplate image across all channels in the RGB color space, representing the overall hue.

[0082] Texture features are vectors used to describe local grayscale variation patterns in an image, reflecting surface roughness, directionality, etc. Histogram of Oriented Gradients (HOG) is a texture descriptor that characterizes edge and shape structure by statistically analyzing the distribution of gradient directions in local regions.

[0083] The comprehensive feature vector is a high-dimensional vector formed by concatenating or weighting color features and texture features, and is used to uniquely represent the corresponding nameplate to be identified.

[0084] Discriminant functions are mathematical functions used to measure the similarity between two feature vectors, such as cosine similarity and Euclidean distance.

[0085] The feature vector of the nameplate template is a comprehensive feature vector that is pre-calculated offline and stored for each standard nameplate template.

[0086] Similarity is a numerical value output by the discriminant function, reflecting the degree of matching between the current nameplate image and the nameplate template.

[0087] The target nameplate template is the nameplate template with the highest similarity to the current image among all candidate templates.

[0088] In actual execution, the input nameplate image is converted from RGB space to Hue Saturation Value (HSV) space, which is more in line with human visual perception; the normalized histograms of the hue (H) and saturation (S) channels are calculated to form a color histogram; at the same time, the pixel mean values ​​of the entire image in the red (R), green (G), and blue (B) channels are calculated to obtain a 3D color mean vector; the color histogram and the mean vector are concatenated to form the color feature.

[0089] The Histogram of Oriented Gradients (HOG) algorithm is applied to the grayscale nameplate image. For example, the image is divided into 8×8 pixel image cells, and the gradient histograms of 9 directions from 0° to 180° are calculated for each image cell. Every 2×2 image cells form an image block, and the histograms within the image block are normalized using the L2-Hys norm. The gradient orientation histogram vectors of all image blocks are concatenated to form high-dimensional texture features, which can capture geometric texture information such as the character arrangement and border structure on the nameplate surface.

[0090] Color and texture features are directly concatenated. To balance the differences in dimensionality, the color and texture components can be standardized using Z-scores to ensure that they contribute equally to the similarity calculation, thus forming a comprehensive feature vector.

[0091] The nameplate recognition software pre-calculates and stores the comprehensive feature vector of each nameplate template in the nameplate template library offline as the template feature vector. During recognition, the weighted cosine similarity is used as the discriminant function to calculate the similarity between the comprehensive feature vector of the nameplate image and the template feature vector of each nameplate template.

[0092] Iterate through the similarity scores of all nameplate templates and select the nameplate template with the highest score that exceeds the similarity threshold as the target nameplate template; if no template meets the standard, trigger the unknown type processing flow and prompt manual review or template addition; if a target nameplate template meets the standard, use the field layout parameters of the template to locate the content area of ​​the nameplate image.

[0093] In this embodiment, a highly discriminative comprehensive feature vector is constructed by fusing color statistical characteristics and texture structure information, overcoming the limitations of single features in complex industrial scenarios: color features can distinguish nameplates of different materials or spraying processes, texture features are highly sensitive to structural differences such as character layout and border style, and through a similarity discrimination mechanism, the correct nameplate template can still be accurately matched under conditions of lighting changes, slight dirt or shooting angle shift.

[0094] In some embodiments, the step of performing horizontal and vertical transition detection and connected component analysis on the content image to determine the content region includes:

[0095] The content image is subjected to jump response calculations along the horizontal and vertical directions, and local maxima detection is performed on the jump response to obtain a set of horizontal jump lines and a set of vertical jump lines.

[0096] The content image is binarized and connected component analysis is performed to obtain multiple connected regions;

[0097] Based on the three-dimensional distribution model of character shapes, the probability density value of each connected region in the feature space is calculated, and the connected regions with probability density values ​​greater than the density threshold are regarded as valid character regions. The feature space includes radius, aspect ratio and duty cycle.

[0098] The effective character region is projected to the horizontal and vertical directions respectively to generate a dual-resolution adaptive projection curve. The projection includes a coarse resolution projection and a fine resolution projection. The coarse resolution projection is used to determine the typical line height, and the window width of the fine resolution projection is determined according to the typical line height.

[0099] Based on the spatial distribution of the horizontal jump line set, the vertical jump line set, and the effective character region, a jump line diagram structure is constructed.

[0100] Based on the aforementioned jump line graph structure, with jump lines as nodes and the weighted sum of the average offset of the character centroids between jump lines and the line spacing as edge weights, the minimum spanning tree algorithm is used to filter out the subset of jump lines with the highest structural consistency.

[0101] Divide the grid cells according to the subset of jump lines;

[0102] Based on the valid character area within the grid cell and the distance between the boundary of the grid cell and adjacent jump lines, the grid cell is filtered to obtain the content area.

[0103] It is understandable that jump response calculation is to calculate the intensity of grayscale change between adjacent pixels or rows / columns in a certain direction of an image, which is used to reflect the position of edges or character boundaries.

[0104] The horizontal and vertical directions represent the X-axis row direction and Y-axis column direction of the image, respectively.

[0105] Local maxima detection identifies local peaks in a transition response sequence, with the local peaks corresponding to the boundaries of character rows or columns.

[0106] The set of horizontal jump lines is a set of candidate text line separators in the horizontal direction, and the set of vertical jump lines is a set of candidate text column separators in the vertical direction.

[0107] Binarization converts a grayscale image into an image containing only 0 (black) and 1 (white), facilitating the separation of foreground characters from the background.

[0108] Connected component analysis is a technique in binary images that aggregates adjacent foreground pixels into independent regions. The resulting connected regions are closed blocks in the image composed of continuous foreground pixels, corresponding to a single character or character fragments.

[0109] The character shape 3D distribution model is an offline statistical model that describes the distribution density of real characters in the three-dimensional feature space of radius-aspect ratio-duty ratio.

[0110] The radius is the radius of the circumcircle of the connected region or its equivalent radius; the aspect ratio is the ratio of the height to the width of the connected region; and the duty cycle is the ratio of the area of ​​the connected region to the area of ​​its circumcircle rectangle, used to reflect the tightness of the filling.

[0111] The probability density value is calculated using kernel density estimation (KDE) to determine the likelihood of a given feature vector under the given distribution model.

[0112] The valid character region is the connected region that is determined to be a real character after statistical verification.

[0113] Projection is the process of accumulating the number of pixels in a two-dimensional region along a certain direction to form a one-dimensional distribution curve. For example, horizontal projection reflects the character density of each line.

[0114] Coarse resolution projection projects the downsampled image to quickly estimate global structural parameters, while fine resolution projection projects the image at the original resolution, but the window parameters are dynamically adjusted.

[0115] Typical line height is the average height of the character lines in the nameplate, used to guide subsequent segmentation accuracy.

[0116] The window width is the size of the sliding window used during projection smoothing or detection.

[0117] The jump line graph structure is represented as a graph theory model, with jump lines as nodes, and the connection relationships between nodes are determined according to the character distribution.

[0118] Spatial distribution refers to the coordinate position of the effective character region in the image and its centroid.

[0119] Edge weight is the numerical value assigned to the edge connecting two nodes in a graph, and it is used to measure the rationality of the structure.

[0120] The Minimum Spanning Tree (MST) algorithm is used to select the subset of edges with the minimum total weight that connects all nodes from a jump graph structure.

[0121] The highest structural consistency is characterized by the jump line layout best matching the actual character distribution, avoiding fragmentation or misalignment caused by interference lines.

[0122] A grid cell is a rectangular area enclosed by the intersection of adjacent horizontal and vertical jump lines, corresponding to a potential field. For example, the potential field is model: ABC123.

[0123] In actual execution, the absolute value of the gray-level difference between adjacent pixel columns is calculated row by row for the content image, forming the horizontal transition response for each row; similarly, the absolute value of the gray-level difference between adjacent pixel rows is calculated column by column, forming the vertical transition response. The horizontal transition response is smoothed by a sliding window along the vertical direction, and the locations of local maxima are detected and recorded as horizontal transition lines; the vertical transition response is detected by detecting local maxima along the horizontal direction and recorded as vertical transition lines, thus forming the sets of horizontal transition lines and the sets of vertical transition lines, respectively.

[0124] The Sauvola adaptive thresholding algorithm is used to locally binarize the content image to address uneven illumination or low contrast issues. Then, the binary image is traversed using the four-neighbor or eight-neighbor connectivity criterion to mark all unconnected foreground pixel clumps, resulting in multiple connected regions. The geometric attributes of each connected region, such as its bounding rectangle, area, and centroid, are recorded to form a list of connected regions.

[0125] The system extracts connected regions of all real characters from multiple (e.g., 10,000) labeled nameplate images, calculates their radius, aspect ratio, and duty cycle, and constructs a three-dimensional kernel density estimation model using a data-driven approach with Gaussian kernels and automatic bandwidth selection, serving as the three-dimensional distribution model of character shapes. During the recognition phase, the same three features are calculated for each connected region and substituted into the three-dimensional distribution model of character shapes to obtain the probability density value. If the probability density value is greater than the density threshold, it is retained as a valid character region; otherwise, it is discarded as noise, blemishes, or table line fragments.

[0126] The content image is scaled up, and the effective character area is horizontally projected onto this low-resolution image. The typical line height is estimated by statistically analyzing the peak spacing. Then, the image is returned to the original resolution image, and the effective character area is weighted horizontally and vertically by the window width to generate a smooth dual-resolution adaptive projection curve for subsequent jump line verification.

[0127] Treat each line in the set of horizontal jump lines as a node in the graph, and process the set of vertical jump lines in the same way. For any two adjacent horizontal jump lines, calculate the average vertical distance from the centroid of all valid character regions between them to these two lines. Similarly, calculate the average horizontal offset between vertical jump lines. This distance information will be used to define edge weights, thereby constructing a complete undirected graph structure.

[0128] Based on the average offset of character centroids, line spacing, and preset weights, edge weights are defined for each pair of adjacent transition lines. The Kruskal algorithm is run on the horizontal and vertical transition line graphs respectively to generate the minimum spanning tree. The transition lines contained in the tree are retained, and the transition lines corresponding to isolated or high-weight edges are removed to obtain the structurally optimal subset of transition lines.

[0129] Sort the filtered horizontal jump lines by y-coordinate and the vertical jump lines by x-coordinate; take two adjacent horizontal lines and two adjacent vertical lines in turn to form a rectangle, and each rectangle is a grid cell; all combinations constitute the complete grid division result.

[0130] Traverse each grid cell and check if it contains at least one valid character region. At the same time, determine whether its height falls within the range determined based on the typical line height and whether its width is reasonable. When both the character and size compliance conditions are met, retain the grid cell as the final content region, discard the remaining cells, and output the content location result with high confidence.

[0131] In this embodiment, by fusing multi-level transition responses, a weighted collaborative semantic edge probability map of grayscale abrupt changes and character semantic edge priors is performed at the pixel level as a domain prior for offline training, which accurately enhances the response strength of real character edges without relying on end-to-end localization; the set of horizontal and vertical transition lines output by local maximum detection has both geometric accuracy and semantic credibility, improving the recall rate and anti-interference ability of nameplate text localization.

[0132] In some embodiments, the step of calculating the jump response of the content image along the horizontal and vertical directions respectively, and performing local maxima detection on the jump response to obtain a set of horizontal jump lines and a set of vertical jump lines includes:

[0133] The gray-level abrupt response is obtained by calculating the absolute value sequence of gray-level differences between adjacent pixels along the horizontal and vertical directions of the content image.

[0134] Multi-directional gradient calculation is performed on the content image to obtain a gradient direction consistency index, and pixel positions with gradient direction consistency greater than the consistency threshold are retained as gradient level transition responses.

[0135] Based on the semantic edge probability map, the gray-level jump response and the gradient-level jump response are weighted and fused at the pixel level to obtain an enhanced jump response map. The semantic edge probability map is generated by a binary classification segmentation network and is used to characterize the semantic prior of character edges.

[0136] Local maxima detection is performed on the enhanced jump response map to obtain the set of transverse jump lines and the set of longitudinal jump lines.

[0137] It is understandable that the gray-level difference absolute value sequence represents the absolute value of the gray-level difference between adjacent pixels in the horizontal or vertical direction, which is used to identify edges.

[0138] Gradient direction consistency metrics are used to measure the consistency of gradient directions around a point in an image, to highlight significant directional changes, such as character boundaries.

[0139] Semantic edge probability maps provide probability estimates of which regions in an image might be character edges.

[0140] Binary classification segmentation networks are deep learning models that can classify each pixel in a content image into one of two categories, such as edge and non-edge, to generate semantic edge probability maps.

[0141] In enhanced abrupt response maps, different types of response maps are assigned different weights to highlight locations that are more likely to be real character edges.

[0142] Local maxima detection is the process of finding local peaks in the response graph, which correspond to the boundaries of potential character rows or columns.

[0143] In actual execution, the content image is traversed in both the horizontal and vertical directions for each row, and the gray-level difference between each pair of adjacent pixels is calculated. The absolute value of the difference is taken as the absolute value of the gray-level difference. The resulting sequence is the gray-level transition response, which is used to reflect the edge information that may exist in the image.

[0144] Using the Sobel operator, the gradient magnitude and direction of the image are calculated in multiple directions such as 0°, 45°, 90°, and 135°. Based on the gradient information in these directions, the consistency of the gradient direction at each pixel is evaluated, that is, whether the gradient direction near the point is concentrated in a specific direction. If the gradient direction consistency of a pixel exceeds a preset threshold, the point is considered to belong to a meaningful edge part and is marked as part of the gradient level transition response.

[0145] An enhanced transition response map is generated by combining gray-level transition response, gradient-level transition response, and semantic edge probability map through weighted fusion.

[0146] A local maximum detection algorithm is applied to the enhanced transition response map to find local maxima points in the response map. For the horizontal direction, local maxima are found within each column to determine the horizontal transition lines that serve as character column boundaries. For the vertical direction, local maxima are found within each row to determine the vertical transition lines that serve as character column boundaries. The detected local maxima form the sets of horizontal and vertical transition lines.

[0147] In this embodiment, grayscale difference is used to capture basic brightness abrupt changes, multi-directional gradients suppress non-directional texture interference, and gradient direction consistency effectively suppresses false responses caused by non-structural interference such as metal reflections and background textures. The semantic edge probability map enhances the response intensity of real text edges from the character semantic level. The enhanced jump response map formed after pixel-level weighted fusion improves the signal-to-noise ratio, so that the horizontal and vertical jump line sets output by subsequent local maximum detection are more accurately aligned with the real character row / column boundaries, thereby providing reliable geometric guidance for high-precision content area positioning and solving the problem of false edge detection caused by reflections, etching noise, low contrast or complex backgrounds in industrial nameplate images.

[0148] In some embodiments, the step of detecting horizontal and vertical table lines in the content image to obtain the deprojected region includes:

[0149] The content image is subjected to line segment detection along the horizontal and vertical directions to obtain a set of horizontal line segments and a set of vertical line segments.

[0150] Based on the length, continuity, and vertical distance from the character area of ​​the line segments in the set of horizontal and vertical line segments, the line segments are filtered to obtain candidate horizontal and vertical table lines.

[0151] The candidate horizontal table lines and the candidate vertical table lines are expanded bidirectionally along their respective normal directions to generate horizontal interference bands and vertical interference bands.

[0152] Morphological closing operations are performed on the horizontal and vertical interference bands to determine and fill the internal holes. A logical AND operation is then performed with the binarization result of the content image to remove the parts containing valid characters, thus obtaining the table line area.

[0153] The table line area is projected in the horizontal and vertical directions respectively to generate horizontal interference projection curves and vertical interference projection curves.

[0154] The intervals in the lateral interference projection curve where the amplitude is greater than the lateral amplitude threshold are marked as lateral deprojection intervals, and the intervals in the longitudinal interference projection curve where the amplitude is greater than the longitudinal amplitude threshold are marked as longitudinal deprojection intervals.

[0155] The image region jointly defined by the horizontal deprojection interval and the vertical deprojection interval is determined as the deprojection region.

[0156] It is understandable that the horizontal line segment set and the vertical line segment set are lists of line segments detected in the image that are approximately horizontal or vertical.

[0157] The line segment length is the Euclidean distance between the two endpoints of the line segment, used to measure whether it is a valid table line.

[0158] Continuity is evaluated through pixel connectivity or fitting residuals and is used to characterize whether a line segment is complete and without obvious breaks.

[0159] The vertical distance from the character region is the shortest vertical distance from the line segment to the boundary of the nearest valid character region. It is used to determine whether a line is a distracting line that is close to the text, such as a field underline.

[0160] The normal direction is the vector direction perpendicular to the direction of the line segment. The normal to a horizontal line is in the vertical direction; the normal to a vertical line is in the horizontal direction.

[0161] Bidirectional expansion is the expansion of line segments along the normal direction to both sides, forming a band-shaped area of ​​a certain width, which serves as a lateral or longitudinal interference band to cover the surrounding area that the table line may affect.

[0162] Morphological closing operations are operations that first expand and then erode, used to fill small holes and connect adjacent areas.

[0163] The horizontal interference projection curve is used to characterize the degree of interference in each row, and the vertical interference projection curve is used to characterize the degree of interference in each column.

[0164] The deprojection interval is a continuous interval on the projection curve that is judged to be disturbed by the table lines. It is used to suppress misjudgments in traditional projection segmentation.

[0165] The deprojection region is a two-dimensional area defined by the intersection of the horizontal and vertical deprojection intervals, used to represent the interference area that should be excluded from text positioning.

[0166] In actual implementation, the Hough Transform is used to identify and extract line segments with approximately straight-line shapes on the content image. The extracted line segments are classified according to their angles, and horizontal and vertical line segments are constructed respectively. For example, angles less than 5° with the horizontal axis are considered horizontal, and angles greater than 85° are considered vertical.

[0167] Traverse all line segments and remove those with a length less than the length threshold or those with obvious breaks due to low continuity scores; at the same time, calculate the vertical distance of each line segment to the nearest valid character area. If the distance is too small, it is judged as a field underline or decorative line and removed; the remaining line segments are used as candidate table lines.

[0168] For each candidate horizontal table line, expand it upwards and downwards by N pixels along the vertical normal direction to form a rectangular strip area; for the vertical table line, expand it left and right along the horizontal normal direction to generate a vertical interference band.

[0169] Morphological closing operations are performed on the dilated interference bands. For example, the structuring element is 3×3, and the breaks caused by noise are filled. The processed interference bands are then ANDed with the binarized result of the content image to retain the overlapping part of the two binary images. In the binarized result, the foreground is the character. If a certain interference band area overlaps with the character, it means that the part is actually a character rather than a pure table line, and it is removed. The area that is finally retained is the pure table line area.

[0170] By accumulating pixel values ​​along the vertical direction of the table line area, a one-dimensional horizontal interference projection curve is obtained; by accumulating along the horizontal direction, a one-dimensional vertical interference projection curve is obtained.

[0171] For example, if both the horizontal and vertical amplitude thresholds are set to an equivalent value of 10 pixels in width, continuous intervals on the horizontal or vertical interference projection curves that exceed the corresponding amplitude thresholds are marked as deprojection intervals. This indicates that these row / column regions are severely interfered with by the table lines and are not suitable for direct use in text projection segmentation.

[0172] All horizontal and vertical deprojection intervals are combined by Cartesian product to generate multiple rectangular sub-regions, which constitute the deprojection region and are used to shield interference during subsequent content localization.

[0173] In this embodiment, through a multi-stage table line detection and purification process, interference caused by metal etching, printed borders or structural dividing lines in the nameplate image can be accurately identified and removed, effectively avoiding misjudging character underlines or decorative lines close to the text as interference sources; the table line area is transformed into a deprojection interval in a one-dimensional projection space, which can actively skip or correct the dividing points in the interference area, thus solving the problem of line-text confusion in content recognition.

[0174] In some embodiments, determining the content location result of the nameplate to be identified based on the content area and the deprojection area includes:

[0175] Spatial relationship analysis is performed on the content area and the deprojection area to determine the overlapping and non-overlapping parts;

[0176] Based on the overlapping portion, the content region is corrected to obtain the corrected content region;

[0177] Based on the relative positional relationship between the non-overlapping portion and the modified content area, the modified content area is adjusted to obtain the content positioning result.

[0178] Understandably, the content region is a preliminary set of text candidate regions containing valid characters, determined after jump detection and connected component analysis, and represented in the form of a rectangle or mask.

[0179] The deprojection region is a set of two-dimensional regions generated by table line detection, representing the disturbed areas, and is used to indicate the locations where traditional projection methods are prone to errors.

[0180] Spatial relationship analysis involves geometric calculations of the positional relationship between two sets of regions in an image coordinate system, including operations such as intersection and difference.

[0181] The overlapping area is the part where the content area and the deprojection area intersect in space, used to represent areas that appear to be text but are actually affected by table lines.

[0182] The non-overlapping portion is the part of the content area that does not intersect with the deprojection area; it is a high-confidence valid text area.

[0183] Relative positional relationships refer to the vertical, horizontal, or adjacency or alignment relationships between the non-overlapping parts and the corrected area, which are used to guide the restoration of structural integrity.

[0184] The content localization result is a text field localization area with a complete structure and resistance to interference, which is used for subsequent optical character recognition (OCR).

[0185] In actual execution, the content region, represented as a set of rectangular boxes, and the deprojection region, represented as a set of rectangles or binary masks, are loaded onto the same coordinate system; the overlapping part is obtained by calculating the intersection of the two, which is used as a sub-region that belongs to both the content region and the deprojection region; the non-overlapping part is obtained by subtracting the overlapping part from the content region, which is used as a reliable text region that is not disturbed by the table lines.

[0186] Remove all overlapping parts from the original content area. If the content area is a mask, perform a logical AND-NOT operation. If it is a rectangle, calculate the difference between each rectangle and the deprojected area to generate multiple sub-regions that may split. Only retain the sub-regions with an area greater than the minimum character size to form the corrected content area.

[0187] Analyze the spatial adjacency between the non-overlapping portion and the corrected content area. If a non-overlapping area is directly above the corrected area and horizontally aligned, it is inferred that the original content area may have been incorrectly segmented due to table lines crossing it. In this case, the two are merged into a larger positioning box. In addition, based on the field layout prior of the target nameplate template, the corrected area that is too short or too narrow is expanded vertically / horizontally to ensure that it conforms to the typical field size. This yields a content positioning result that is structurally complete and semantically coherent.

[0188] In this embodiment, the deprojection area is used as an interference suppression guide signal to accurately remove overlapping areas contaminated by table lines, avoiding the misidentification of scratches or borders as text; the undisturbed non-overlapping parts are used as contextual clues, and the correction area is adjusted in combination with the layout rules of the nameplate fields to restore the field breakage problem caused by excessive removal; the integrity and accuracy of content positioning in strong interference scenarios are improved.

[0189] In some embodiments, the step of performing target detection and content recognition on the content location result based on the content image to obtain the recognition result of the nameplate to be identified includes:

[0190] The image corresponding to the content location result is divided into image blocks, pixels are redistributed, and pixel grayscale values ​​are reconstructed to obtain an enhanced image;

[0191] The enhanced image is subjected to tilt correction, projection segmentation, and segmentation correction to determine the character segmentation result of the nameplate to be identified;

[0192] The character segmentation results are horizontally projected to determine the character boundaries of the nameplate to be identified;

[0193] Based on the character boundaries, wavelet transform is used to extract the spectral energy of the characters in each direction from the content image;

[0194] Based on the spectral energy and directional gradient histogram, the texture information of the character is obtained;

[0195] Based on the texture information, a similarity metric matrix is ​​constructed between the texture information and the character template to obtain the first recognition result;

[0196] The character segmentation result is identified by a character recognition model, which is built based on a deep learning network, to obtain a second recognition result.

[0197] The first identification result and the second identification result are fused to obtain a fused result;

[0198] Based on character arrangement rules, a character sequence is obtained by performing permutation and combination analysis on the characters in the fusion result using a Markov chain.

[0199] The character sequence is corrected to obtain the recognition result.

[0200] Understandably, the content localization result is the text region coordinates output by the previous steps, which can be represented as a rectangle to indicate the location of the character to be recognized.

[0201] Image segmentation divides the localized area into several sub-blocks, which facilitates local enhancement processing.

[0202] Pixel redistribution adjusts the pixel spatial arrangement through interpolation, resampling, and other methods to improve resolution or align with the grid.

[0203] Pixel grayscale reconstruction is based on neighborhood information to rebuild pixel intensity, which is used for noise reduction, contrast enhancement, etc.

[0204] Skew correction aligns text lines horizontally by estimating and correcting their rotation angle.

[0205] Projection segmentation uses the valley positions of horizontal / vertical projection curves to segment characters or fields.

[0206] Segmentation correction is a post-processing step on the projection segmentation results, merging connected characters or splitting broken characters.

[0207] Horizontal projection accumulates pixel values ​​along the vertical direction to generate the character density distribution for each line, which is used for precise delimitation.

[0208] Wavelet transform, through time-frequency analysis, can extract the spectral energy of an image at multiple scales and in multiple directions.

[0209] Spectral energy is the sum of squares of wavelet coefficients, reflecting the textural activity of a character at a specific direction / scale.

[0210] The directional gradient histogram is used to statistically analyze the gradient direction distribution in a local region, representing the shape and contour features of a character.

[0211] The similarity metric matrix is ​​obtained by calculating the similarity between the character to be identified and the feature vectors of each character in the template library.

[0212] The character recognition model is an end-to-end optical character recognition (OCR) model trained using deep learning based on Transformer.

[0213] The character arrangement rules are the syntax rules for the nameplate fields. For example, voltage = number + V, and the model number starts with a letter.

[0214] In practice, the Limiting Contrast Adaptive Histogram Equalization (CLAHE) algorithm can enhance the visibility of local image details by increasing the contrast of local regions, while also suppressing some noise generation.

[0215] like Figure 5As shown, the image corresponding to the content localization result is first divided into M×N non-overlapping sub-blocks. The non-overlapping areas can have their grid effect removed through inter-region interpolation. Introducing a certain amount of overlapping areas makes the gray-level mapping function of each sub-block more reasonable. For each sub-block, its histogram is first calculated. For the histogram portion of each sub-block exceeding the set contrast, it is cropped and evenly distributed to each gray level. Then, the new histogram is equalized. Due to the independence of each sub-block, the pixel connection between each block is not smooth after the above operation. However, the pixel value at the center of each block completely conforms to the original definition. Therefore, for each pixel in the image, bilinear interpolation can be performed through the mapping function of its neighboring blocks. Then, Otsu's threshold segmentation algorithm is used to perform binarization segmentation on each region to obtain the enhanced image. After fine localization of the area to be identified on the nameplate and image enhancement, the content and location information of each area to be identified can be obtained. The areas to be identified mainly include text, spaces between text, letters, numbers, and characters.

[0216] Each character is segmented using projection images and thresholds for character width and spacing. To facilitate subsequent character recognition, each character image is normalized to ensure consistent size. Although the nameplate area was previously tilted, characters are not always vertical; special cases, such as italics, require further tilt correction using Hough transform on the binarized character area. After character correction, gaps remain between adjacent characters. Therefore, vertical projection is used, selecting consecutive columns projected to a value greater than 0 as the initial segmentation result. Width statistics are performed on the initial segmentation results, identifying the most widely distributed width ranges as references. These reference widths, character gaps, and the relative positions of adjacent characters are then used to refine the width of the initial segmentation. Finally, the characters divided into two parts are joined, and any adhering characters are further segmented to obtain the final character segmentation result.

[0217] The corrected segmentation result is then horizontally projected. The upper and lower blank edges are precisely truncated according to the start and end positions of the projection, the upper and lower character boundaries are corrected, and pixels outside the boundaries are removed to obtain the character boundaries of the nameplate to be identified.

[0218] A two-dimensional discrete wavelet transform is performed on each character image to extract high-frequency sub-bands in three directions: horizontal, vertical, and diagonal. The squared L2 norm of each sub-band coefficient is calculated as the spectral energy in that direction, forming a multi-directional energy vector.

[0219] By concatenating the wavelet spectral energy vector with the directional gradient histogram (HOG) vector, a fused texture feature vector is formed, which comprehensively represents the structure and edge characteristics of the character.

[0220] In a pre-built standard character template library, each template corresponds to the same texture feature vector. The cosine similarity between the texture vector of the current character and all template vectors is calculated, and the template character corresponding to the highest score is taken as the first recognition result.

[0221] The character segmentation results are input into GoogleNet and ResNet models that have been trained on a large amount of industrial nameplate data. The output character category probability distribution is then used as the second recognition result.

[0222] If the first identification result and the second identification result are consistent, they are adopted directly; if they are inconsistent, their confidence levels are compared and the one with the higher confidence level is selected as the fusion result. Alternatively, the first identification result can be corrected and supplemented by the second identification result.

[0223] Since the arrangement of a certain character in the nameplate follows a certain pattern, the joint probability of different character sequences can be calculated using a Markov chain through the character transition probability matrix of the built-in nameplate field, and the most likely legal sequence can be selected as the character sequence.

[0224] Using Markov chains, we can calculate the probability that a permutation of k candidate characters forms a term S:

[0225] ;

[0226] in, It's probability. It is a character sequence; Let be the i-th character in the character sequence, k be the length of the character sequence, and n be the order of the Markov chain, representing the current character. It depends on the history of the preceding n-1 characters; N is the counting function; π is the chain rule of probability.

[0227] Based on the probability of different permutations forming different terms, these probabilities are arranged from high to low, and the term with the highest probability is selected. The identified character sequence is then corrected using grammar, dictionary, or rules. For example, the number 0 is corrected to the letter O. The result after the misjudgment correction is used as the final nameplate recognition result.

[0228] In this embodiment, the handcrafted texture features based on wavelets and histogram of directional gradients (HOG) are complemented by a deep learning model to effectively address issues such as low contrast, etched blur, or partial occlusion. Character arrangement rules and Markov chains are introduced to perform semantic rationality verification on the recognition sequence, significantly reducing field errors caused by misidentification of isolated characters. Through fusion and correction strategies, high-accuracy and high-compliance nameplate structured recognition is achieved without relying on large-scale labeled data, meeting the stringent requirements for reliable automatic reading in industrial sites such as power and manufacturing.

[0229] In some embodiments, correcting the character sequence to obtain the recognition result includes:

[0230] The character sequence is matched with the semantic rules in the nameplate semantic knowledge base, which is constructed based on the semantic rules of common nameplate fields, the logical relationships between fields, and the value range of fields.

[0231] If any field in the character sequence does not conform to the corresponding semantic rules, the field is marked as a suspicious field according to the logical relationship and value range in the nameplate semantic knowledge base;

[0232] Based on the contextual information of any of the fields and the overall semantic logic of the nameplate to be identified, the suspicious fields are corrected to determine the final identification result.

[0233] It is understandable that a character sequence is the raw string or structured field sequence output after optical character recognition (OCR).

[0234] The nameplate semantic knowledge base is a structured rule base stored in JSON or database format, containing prior knowledge such as naming conventions, value formats, logical constraints, and numerical ranges of common nameplate fields.

[0235] Semantic rules are used to describe the language or numeric format that a field should meet. For example, a voltage field must begin with a number and end with V.

[0236] The logical relationships between fields are dependencies or consistency constraints between different fields. For example, the product of rated power (kW) and power factor is less than or equal to input power (kVA).

[0237] The range of values ​​is the upper and lower limits or an enumeration set of the legal values ​​of the field. For example, the frequency can only be 50Hz or 60Hz.

[0238] Suspicious fields are those that violate semantic rules, logical relationships, or exceed the range of values, and require further verification or correction.

[0239] Contextual information consists of the contents of other fields adjacent to or related to the suspicious field, used to help infer the correct value.

[0240] The overall semantic logic is a semantic consistency system composed of all fields of the nameplate, used for global rationality verification.

[0241] In actual execution, the character sequence output by OCR is parsed into a structured dictionary according to the preset field names; the nameplate semantic knowledge base is loaded, in which each rule contains a field name, regular expression template, list of legal values ​​and logical formulas with other fields; rule matching is performed field by field, and the compliance status of each field is recorded.

[0242] If a field does not meet the regular expression format, for example, 22OV does not match the voltage regular expression, the value is out of range, or it violates the logical relationship, it is marked as a suspicious field; at the same time, the type of rule it violates and the related fields are recorded to facilitate subsequent targeted correction. The rule type can be format, range, logic, etc.

[0243] The following multi-strategy correction mechanism is activated for suspicious fields.

[0244] Perform character-level error correction for easily confused characters, such as the letter O and the number 0, the letter I and the number 1, and find the closest item in the set of legal values ​​based on edit distance;

[0245] In contextual reasoning, if the input voltage is 380V and the frequency is identified as "SOHz", then, in accordance with the Chinese power grid standard (50Hz), the letter SO is corrected to the number 50.

[0246] In logical consistency repair, if the rated power is 10kW and the power factor is 0.8, but the apparent power is 8kVA, while the apparent power should be 12.5kVA, then the high confidence field is trusted first, and the low confidence field is corrected.

[0247] After correction, global semantic consistency is verified again, and the final compliant and structured recognition result is output.

[0248] In this embodiment, by introducing a rule-based post-processing mechanism driven by a nameplate semantic knowledge base, not only are obvious anomalies filtered out using the format and value constraints of the fields themselves, but intelligent error correction is also achieved through logical relationships between fields and contextual reasoning. This avoids the shortcomings of traditional pure character recognition models that lack domain knowledge, and solves the semantic error problem caused by poor image quality, shallow character etching, or optical distortion in optical character recognition OCR in industrial scenarios. It is suitable for industrial automation scenarios such as power and machinery where the accuracy of nameplate parameters is strictly required.

[0249] In some embodiments, obtaining the nameplate image of the nameplate to be identified includes:

[0250] Obtain the image to be processed of the nameplate to be identified;

[0251] The image to be processed is cropped, converted to grayscale, and denoised to obtain a preprocessed image;

[0252] Edge detection and fitting are performed on the preprocessed image to obtain contour information;

[0253] The nameplate area corresponding to the contour information is subjected to image binarization processing to obtain a binarized image;

[0254] The Hough transform is used to detect straight line information in the binarized image, and the binarized image is tilted based on the straight line information to obtain the nameplate image.

[0255] It is understandable that the image to be processed is the original nameplate scene image directly captured by imaging devices such as industrial cameras and mobile phones, which may contain interference such as background, noise, and tilt.

[0256] In actual implementation, such as Figure 6 As shown, since the resolution of the images captured by the camera may be too large, it will cause significant problems to the processing speed and memory usage of the algorithm. Therefore, bicubic interpolation is used to scale the original image and extract a sub-region containing the main body of the nameplate from the image to be processed, thereby reducing the processing range.

[0257] Due to significant color differences between different nameplates, this application does not consider color spaces such as RGB or HSV. Generally, captured images are RGB images; therefore, the images are converted to grayscale.

[0258] ;

[0259] in, This is the image after grayscale conversion; R, G, and B represent the red, green, and blue pixel values ​​in the RGB channels of the image, respectively.

[0260] Image noise is the deviation between the real signal and the ideal signal. The noise in the image to be processed is mostly random impulse noise, which is a high-frequency component. In order to reduce the impact of noise on the vertical / horizontal edges and ensure efficiency, threshold neighborhood averaging filtering and median filtering are used, which can be determined based on the actual experimental results.

[0261] The basic principle of median filtering is to replace the value of a point in a digital image or digital sequence with the median value of all points in its neighborhood, so that the surrounding pixel values ​​are close to the true value, thereby eliminating isolated noise points.

[0262] Threshold neighborhood averaging filtering is an improved method based on mean filtering:

[0263] ;

[0264] in, It is the filtered output value; The cropped image is in position Pixel value at; It is a window The total number of pixels in; It is the local mean of the pixels within the window; It is a preset noise judgment threshold.

[0265] The Laplacian edge detection method is used to detect the polygonal contours of the nameplate image. Then, based on constraints such as the geometric relationship, aspect ratio, area size, and texture features of the quadrilateral contours of the nameplate area, interfering areas are removed. Finally, Hough line detection is used to fit the quadrilateral contours of the nameplate area, thus obtaining a coarse localization of the nameplate area.

[0266] The Hough transform is a coordinate transformation of an image, causing it to exhibit a peak at a specific location in polar coordinate space. First, the nameplate area is locally binarized, the binarized image is then thinned, and finally, Hough line detection is performed on the thinned image.

[0267] Points on the same straight line will have a common intersection point when they are transformed into all curves in Hough space. This will cause points on straight lines with similar directions and positions to form a dense region of points in Hough space after transformation. Then, by setting a certain threshold, the approximate outline of the quadrilateral of the nameplate area can be obtained through fitting.

[0268] Various factors, such as shooting angle and installation method, can cause tilt distortion in the captured nameplate image, posing challenges for subsequent character segmentation and recognition. Therefore, before character segmentation, it is necessary to correct the tilt of the nameplate area image. This involves rotating, affine, and perspective transforming the image to restore its horizontal position and original quadrilateral shape. Specifically, based on the previously fitted quadrilateral lines in the nameplate area, the coordinates of the four vertices of the quadrilateral are automatically selected, and these coordinates are used as parameters to perform a global perspective transformation on the image, thus completing the tilt correction of the nameplate area image.

[0269] In this embodiment, edge fitting is used to accurately separate the nameplate body from the background, avoiding interference from irrelevant areas; adaptive binarization is combined to effectively address common problems of industrial nameplates such as reflection and shallow etching; and tilt correction driven by Hough transform ensures that the text lines are strictly horizontal, providing ideal input for subsequent geometrically sensitive operations such as jump detection and projection segmentation, thereby improving the first frame availability and positioning accuracy in real industrial environments.

[0270] The nameplate recognition method provided in this application can be executed by a nameplate recognition device. This application uses a nameplate recognition device executing the nameplate recognition method as an example to illustrate the nameplate recognition device provided in this application.

[0271] This application also provides a nameplate recognition device.

[0272] like Figure 7 As shown, the nameplate identification device includes:

[0273] The acquisition module 710 is used to acquire the nameplate image of the nameplate to be identified;

[0274] The first processing module 720 is used to extract features from the nameplate image and match the target nameplate template of the nameplate to be identified based on the extracted nameplate features.

[0275] The second processing module 730 is used to detect and extract the content area of ​​the nameplate image according to the content layout of the target nameplate template to obtain a content image;

[0276] The third processing module 740 is used to perform horizontal and vertical jump detection and connected component analysis on the content image to determine the content region, and to perform horizontal and vertical table line detection on the content image to obtain the deprojected region.

[0277] The fourth processing module 750 is used to determine the content positioning result of the nameplate to be identified based on the content area and the deprojection area.

[0278] The fifth processing module 760 is used to perform target detection and content recognition on the content location result based on the content image to obtain the recognition result of the nameplate to be recognized.

[0279] According to the nameplate recognition device provided in the embodiments of this application, the structure-guided recognition is achieved by leveraging the prior characteristics of the fixed layout of the nameplate content. This can suppress complex interferences commonly found in industrial environments, such as table lines, shadows, and reflections. Through a multi-strategy collaborative mechanism of content image fusion, horizontal and vertical jump detection, connected component analysis, and table line detection, the device accurately generates content regions and deprojection regions, performs fine-grained text region localization, and eliminates invalid interference regions by combining layout constraints. This allows the device to focus on high-confidence valid text, improves the recognition accuracy of nameplates under complex imaging conditions, overcomes the inefficiency of manual inspection, and makes online inspection work more objective, standardized, and intelligent.

[0280] The nameplate recognition device provided in this application embodiment can realize the various processes implemented in the nameplate recognition method embodiment as described above. To avoid repetition, it will not be described again here.

[0281] In some embodiments, such as Figure 8 As shown, this application embodiment also provides an electronic device 800, including a processor 801, a memory 802, and a computer program stored in the memory 802 and executable on the processor 801. When the program is executed by the processor 801, it implements the various processes of the above-described nameplate identification method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0282] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0283] The nameplate recognition method described in this invention can be efficiently integrated into the blind sample detection system of equipment. As a core technology for realizing the automated and accurate entry of sample information, it fundamentally solves the problems of low efficiency, easy error and poor traceability of traditional manual entry.

[0284] In a specific scenario of blind sample testing of equipment, the application process of this invention is as follows:

[0285] During the sample receiving process at the testing institution, operators use the terminal equipment of the blind sample testing system to scan the primary blind sample code provided by the submitter and affixed to the packaging of the sample to be tested (such as a transformer or switchgear). This primary blind sample code is an anonymous identifier assigned by the submitter and used for the transfer of the sample between the submitter and the testing institution; it does not contain any identifiable information.

[0286] After scanning the blind sample code once, the system automatically activates the nameplate recognition function integrated with this invention. The operator uses a terminal camera to photograph the nameplate on the sample body under inspection, following the prompts, to obtain an image of the nameplate to be identified. Subsequently, the system executes the nameplate recognition method described in this invention. This process includes: extracting features from the nameplate image; matching the extracted nameplate features to a target nameplate template for the nameplate to be identified; detecting and extracting the content region of the nameplate image based on the content layout of the target nameplate template to obtain a content image; performing horizontal and vertical transition detection and connected component analysis on the content image to determine the content region, and performing horizontal and vertical table line detection on the content image to obtain a deprojected region; determining the content location result of the nameplate to be identified based on the content region and the deprojected region; and performing target detection and content recognition based on the content image to obtain the recognition result of the nameplate to be identified.

[0287] After identification, the system outputs the identification result of the nameplate to be identified. This result is structured nameplate information presented in the form of "field name-field value" (e.g., {"Equipment Model": "S11-M-200 / 10", "Rated Voltage": "10kV", "Rated Capacity": "200kVA"}). The system automatically performs the following operations: associating and binding the structured nameplate information with the scanned primary blind sample code, and storing it in the system database. After binding, the system immediately logically disables the access permissions for the bound structured nameplate information and the primary blind sample code, ensuring that they cannot be accessed arbitrarily in subsequent testing processes. Based on this binding relationship, the system generates a brand new secondary blind sample code for internal circulation only. The operator prints this secondary blind sample code label, replacing or covering the original primary blind sample code on the sample. From this point on, the sample to be tested, carrying the secondary blind sample code, enters the blind sample state and is transferred to the subsequent testing task assignment, laboratory testing, and data upload stages. When all testing items are completed and the report needs to be signed, authorized personnel (such as report auditors) scan the secondary blind sample code on the sample to be tested in the system, triggering the "unblinding" process. After the system verifies the permissions, it automatically retrieves the accurate nameplate information previously isolated and stored and identified by the method of this invention, fills it into the standard test report template, and generates a final test report with real sample information, completing the entire data loop and reliable traceability from physical nameplate to authoritative report.

[0288] Through the above applications, this invention deeply integrates high-precision nameplate recognition technology with a rigorous blind sample testing management process. This invention not only significantly improves the efficiency and accuracy of sample information entry, eliminating the subjective errors or tampering risks that may be introduced by manual entry, but also strengthens the principles of blind sample testing through technical means, ensuring the objectivity and impartiality of test results. It is a key technological innovation driving the digital and intelligent transformation of the testing industry.

[0289] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0290] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0291] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.

Claims

1. A nameplate identification method, characterized in that, include: Obtain the image of the nameplate to be identified; Feature extraction is performed on the nameplate image, and the target nameplate template of the nameplate to be identified is matched based on the extracted nameplate features; Based on the content layout of the target nameplate template, the content area of ​​the nameplate image is detected and extracted to obtain the content image; The content image is subjected to horizontal and vertical jump detection and connected component analysis to determine the content region, and the content image is subjected to horizontal and vertical table line detection to obtain the deprojected region. The specific steps include: The content image is subjected to line segment detection along the horizontal and vertical directions to obtain a set of horizontal line segments and a set of vertical line segments. Based on the length, continuity, and vertical distance from the character area of ​​the line segments in the set of horizontal and vertical line segments, the line segments are filtered to obtain candidate horizontal and vertical table lines. The candidate horizontal table lines and the candidate vertical table lines are expanded bidirectionally along their respective normal directions to generate horizontal interference bands and vertical interference bands. Morphological closing operations are performed on the horizontal and vertical interference bands to determine and fill the internal holes. A logical AND operation is then performed with the binarization result of the content image to remove the parts containing valid characters, thus obtaining the table line area. The table line area is projected in the horizontal and vertical directions respectively to generate horizontal interference projection curves and vertical interference projection curves. The intervals in the lateral interference projection curve where the amplitude is greater than the lateral amplitude threshold are marked as lateral deprojection intervals, and the intervals in the longitudinal interference projection curve where the amplitude is greater than the longitudinal amplitude threshold are marked as longitudinal deprojection intervals. The image region jointly defined by the horizontal deprojection interval and the vertical deprojection interval is determined as the deprojection region; Based on the content area and the deprojection area, the content positioning result of the nameplate to be identified is determined; Based on the content image, target detection and content recognition are performed on the content location result to obtain the recognition result of the nameplate to be identified.

2. The nameplate identification method according to claim 1, characterized in that, The nameplate features include color features and texture features. The step of extracting features from the nameplate image and matching the extracted nameplate features to the target nameplate template for the nameplate to be identified includes: The color features are obtained by extracting the color histogram and color mean from the nameplate image. The texture features are obtained by extracting the directional gradient histogram from the nameplate image; Based on the color features and the texture features, a comprehensive feature vector is constructed; The similarity between the comprehensive feature vector and the template feature vector of each nameplate template is calculated using a discriminant function. Based on the similarity, the target nameplate template is determined among the multiple nameplate templates.

3. The nameplate identification method according to claim 1, characterized in that, The step of performing horizontal and vertical transition detection and connected component analysis on the content image to determine the content region includes: The content image is subjected to jump response calculations along the horizontal and vertical directions, and local maxima detection is performed on the jump response to obtain a set of horizontal jump lines and a set of vertical jump lines. The content image is binarized and connected component analysis is performed to obtain multiple connected regions; Based on the three-dimensional distribution model of character shapes, the probability density value of each connected region in the feature space is calculated, and the connected regions with probability density values ​​greater than the density threshold are regarded as valid character regions. The feature space includes radius, aspect ratio and duty cycle. The effective character region is projected to the horizontal and vertical directions respectively to generate a dual-resolution adaptive projection curve. The projection includes a coarse resolution projection and a fine resolution projection. The coarse resolution projection is used to determine the typical line height, and the window width of the fine resolution projection is determined according to the typical line height. Based on the spatial distribution of the horizontal jump line set, the vertical jump line set, and the effective character region, a jump line diagram structure is constructed. Based on the aforementioned jump line graph structure, with jump lines as nodes and the weighted sum of the average offset of the character centroids between jump lines and the line spacing as edge weights, the minimum spanning tree algorithm is used to filter out the subset of jump lines with the highest structural consistency. Divide the grid cells according to the subset of jump lines; Based on the valid character area within the grid cell and the distance between the boundary of the grid cell and adjacent jump lines, the grid cell is filtered to obtain the content area.

4. The nameplate identification method according to claim 3, characterized in that, The step involves calculating the abrupt response of the content image along both the horizontal and vertical directions, and performing local maxima detection on the abrupt response to obtain a set of horizontal abrupt lines and a set of vertical abrupt lines, including: The gray-level abrupt response is obtained by calculating the absolute value sequence of gray-level differences between adjacent pixels along the horizontal and vertical directions of the content image. Multi-directional gradient calculation is performed on the content image to obtain a gradient direction consistency index, and pixel positions with gradient direction consistency greater than the consistency threshold are retained as gradient level transition responses. Based on the semantic edge probability map, the gray-level jump response and the gradient-level jump response are weighted and fused at the pixel level to obtain an enhanced jump response map. The semantic edge probability map is generated by a binary classification segmentation network and is used to characterize the semantic prior of character edges. Local maxima detection is performed on the enhanced jump response map to obtain the set of transverse jump lines and the set of longitudinal jump lines.

5. The nameplate identification method according to claim 1, characterized in that, The step of determining the content location result of the nameplate to be identified based on the content area and the deprojection area includes: Spatial relationship analysis is performed on the content area and the deprojection area to determine the overlapping and non-overlapping parts; Based on the overlapping portion, the content region is corrected to obtain the corrected content region; Based on the relative positional relationship between the non-overlapping portion and the modified content area, the modified content area is adjusted to obtain the content positioning result.

6. The nameplate identification method according to claim 1, characterized in that, The step of performing target detection and content recognition on the content location result based on the content image to obtain the recognition result of the nameplate to be recognized includes: The image corresponding to the content location result is divided into image blocks, pixels are redistributed, and pixel grayscale values ​​are reconstructed to obtain an enhanced image; The enhanced image is subjected to tilt correction, projection segmentation, and segmentation correction to determine the character segmentation result of the nameplate to be identified; The character segmentation results are horizontally projected to determine the character boundaries of the nameplate to be identified; Based on the character boundaries, wavelet transform is used to extract the spectral energy of the characters in each direction from the content image; Based on the spectral energy and directional gradient histogram, the texture information of the character is obtained; Based on the texture information, a similarity metric matrix is ​​constructed between the texture information and the character template to obtain the first recognition result; The character segmentation result is identified by a character recognition model, which is built based on a deep learning network, to obtain a second recognition result. The first identification result and the second identification result are fused to obtain a fused result; Based on character arrangement rules, a character sequence is obtained by performing permutation and combination analysis on the characters in the fusion result using a Markov chain. The character sequence is corrected to obtain the recognition result.

7. The nameplate identification method according to claim 6, characterized in that, The step of correcting the character sequence to obtain the recognition result includes: The character sequence is matched with the semantic rules in the nameplate semantic knowledge base, which is constructed based on the semantic rules of common nameplate fields, the logical relationships between fields, and the value range of fields. If any field in the character sequence does not conform to the corresponding semantic rules, the field is marked as a suspicious field according to the logical relationship and value range in the nameplate semantic knowledge base; Based on the contextual information of any of the fields and the overall semantic logic of the nameplate to be identified, the suspicious fields are corrected to determine the final identification result.

8. The nameplate identification method according to claim 1, characterized in that, The nameplate recognition method is applied to the equipment blind sample detection system, and the recognition result of the nameplate recognition method is used to automatically generate or fill in the sample information of the sample to be tested.

9. A nameplate identification device, characterized in that, include: The acquisition module is used to acquire the image of the nameplate to be identified; The first processing module is used to extract features from the nameplate image and match the target nameplate template of the nameplate to be identified based on the extracted nameplate features. The second processing module is used to detect and extract the content area of ​​the nameplate image according to the content layout of the target nameplate template to obtain the content image; The third processing module is used to perform horizontal and vertical jump detection and connected component analysis on the content image to determine the content region, and to perform horizontal and vertical table line detection on the content image to obtain the deprojected region; wherein, the content image is subjected to line segment detection along the horizontal and vertical directions respectively to obtain the set of horizontal line segments and the set of vertical line segments. Based on the length, continuity, and vertical distance from the character area of ​​the line segments in the set of horizontal and vertical line segments, the line segments are filtered to obtain candidate horizontal and vertical table lines. The candidate horizontal table lines and the candidate vertical table lines are expanded bidirectionally along their respective normal directions to generate horizontal interference bands and vertical interference bands. Morphological closing operations are performed on the horizontal and vertical interference bands to determine and fill the internal holes. A logical AND operation is then performed with the binarization result of the content image to remove the parts containing valid characters, thus obtaining the table line area. The table line area is projected in the horizontal and vertical directions respectively to generate horizontal interference projection curves and vertical interference projection curves. The intervals in the lateral interference projection curve where the amplitude is greater than the lateral amplitude threshold are marked as lateral deprojection intervals, and the intervals in the longitudinal interference projection curve where the amplitude is greater than the longitudinal amplitude threshold are marked as longitudinal deprojection intervals. The image region jointly defined by the horizontal deprojection interval and the vertical deprojection interval is determined as the deprojection region; The fourth processing module is used to determine the content positioning result of the nameplate to be identified based on the content area and the deprojection area. The fifth processing module is used to perform target detection and content recognition on the content location result based on the content image to obtain the recognition result of the nameplate to be recognized.

Citation Information

Patent Citations

  • Multi-character characteristic fused license plate positioning method

    CN102375982A

  • Feature extraction method and system for power equipment nameplate

    CN110569848A