Machine Vision-Based Cable Number Recognition System
Through the combination of deep learning features and significance scores, dot-shaped text features on the cable are automatically extracted, solving the problem of low recognition accuracy of small-diameter cables in traditional methods, and achieving efficient cable number recognition.
Patent Information
- Application Number
- CN202510098466.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-01-22
AI Technical Summary
传统方法在识别小直径线缆上的点状文字时,识别精度显著下降,导致电气控制箱接线错误,存在安全隐患。
Using a cable number recognition system based on machine vision, the dot-shaped text features on the cable are automatically extracted through deep learning features and significance scores, and character segmentation and recognition are performed.
It improves the adaptability to complex character forms and scenes, enhances the geometric consistency and recognition accuracy of characters, reduces background interference, and significantly improves the recognition effect.
Smart Images

Figure CN119580279B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a number recognition system, specifically a cable number recognition system based on machine vision. Background Art
[0002] The electrical control box is an important part of the marine control system, and its wiring quality is directly related to the stability, safety and operating efficiency of equipment operation. During the production process of the electrical control box, the production of the circuit module requires multiple processes including sorting cables, sleeving, wire pressing, inserting terminals, and pressing into the circuit. Among them, the sorting and number matching of cables are key links. Traditional automated production usually relies on manual cable sorting and sleeving. However, due to the possibility of human recognition errors, it often leads to wiring errors in the circuit module, which in turn causes unstable operation of the finally produced electrical control box, and may even cause damage to the equipment, bringing serious safety hazards and economic losses. The patent document with the patent publication number CN118366174A discloses a method for fast recognition of low-computing-power characters based on machine vision, which performs character recognition based on the closed-loop contour of the obtained connected component, realizing accurate and fast recognition of characters. However, in the above patent technology and the prior art, their character recognition requires accurate image recognition through high-quality images, while the numbers on the cables are usually presented in the form of dot spraying.
[0003] Traditional methods for recognizing dot-shaped characters mainly include image processing steps such as grayscale processing, binarization processing, and median filtering for noise reduction, converting the dot-shaped characters into continuous lines, then retaining the text area through image cutting, and combining matrix rotation algorithms for skew correction, and finally performing character segmentation and template matching. This method has high requirements for the quality of the original image and is more suitable for planar dot-shaped characters or dot-shaped characters on cables with a larger diameter. When the cable diameter is small, the dot-shaped characters are prone to deformation due to insufficient shooting angle or resolution, resulting in a significant decrease in recognition accuracy. Summary of the Invention
[0004] In view of the deficiencies of the prior art, the present invention provides a cable number recognition system based on machine vision, which automatically extracts the dot-shaped text features on the cable through the features of deep learning and the saliency scores in the features, and can also efficiently complete character segmentation and recognition, thus solving the technical problems proposed in the background art.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions:
[0006] A cable number recognition system based on machine vision, comprising:
[0007] A line character acquisition module for acquiring the line character image of the cable number;
[0008] A deep feature map extraction module for extracting a deep feature map of the cable number from the line character image;
[0009] A feature vector module for flattening the deep feature map of the line character image into a high-dimensional feature vector of the line character image;
[0010] A linear mapping module for inputting the high-dimensional feature vector of the line character image into a fully connected layer, and the fully connected layer performs a linear mapping on it to obtain a one-dimensional vector representation of the line character image;
[0011] A classification module for passing the one-dimensional vector representation of the line character image through a softmax function to generate a classification label for each line character image;
[0012] A cable number generation module for obtaining the classification labels of all line character images and combining them in the order of the characters in the image to generate the cable number.
[0013] In some embodiments, obtaining the line character image of the cable number includes:
[0014] S1-1. Obtaining a cable image containing the cable number; wherein, the cable number is composed of dot characters;
[0015] S1-2. Performing object detection on the cable image and extracting the bounding box coordinates of each dot character in the cable image;
[0016] S1-3. Preprocessing the cable image after object detection so that the dot characters in the cable image are converted into regular line characters;
[0017] S1-4. According to the bounding box coordinates extracted in the object detection, cropping out the regions containing each regular line character from the preprocessed cable image to obtain the line character image.
[0018] In some embodiments, extracting the deep feature map of the cable number from the line character image includes:
[0019] S2-1. Receiving the line character image as the input of a convolutional model, and extracting its local features by sliding a convolutional kernel to obtain N local feature maps in the line character image; wherein, each local feature map represents a type of visual feature of the line character image;
[0020] S2-2. Performing non-linear activation on the N local feature maps to generate N activated feature maps;
[0021] S2-3. Perform pooling on N activated feature maps, and select M significant feature maps from several activated feature maps; among them, the significant feature maps retain the high-level abstract features with the maximum local feature significance scores through the max pooling operation;
[0022] S2-4. Define the set of M significant feature maps as the depth feature map of the line character image.
[0023] In some embodiments, for object detection of the cable image, extracting the bounding box coordinates of each dot character in the cable image includes:
[0024] S1-2-1. Input the cable image into a pre-trained object detection model;
[0025] Among them, the object detection model is iteratively generated based on a convolutional neural network through supervised learning pre-training, and its iterative goal is: extracting the image features of the cable image and generating candidate box feature vectors;
[0026] S1-2-2. Divide the cable image input into the object detection model into B×B grid regions;
[0027] S1-2-3. Extract features for each grid region to generate a candidate box feature vector corresponding to each grid region; among them, the candidate box feature vector includes: candidate box boundary coordinates, the category probability that the candidate box belongs to a dot character, and the category confidence;
[0028] S1-2-4. Select the candidate box feature vectors with category confidence higher than a preset threshold from the B×B candidate box feature vectors;
[0029] S1-2-5. Perform non-maximum suppression on the selected candidate box feature vectors, remove overlapping candidate boxes, retain the candidate box boundary coordinates where the dot character is located, and generate the bounding box coordinates of the dot character.
[0030] In some embodiments, preprocess the cable image after object detection to convert the dot characters in the cable image into regular line characters, including:
[0031] S1-3-1. Perform grayscale processing on the cable image after object detection to convert the cable image into a grayscale image;
[0032] S1-3-2. Perform binarization of the grayscale image, and divide the pixels in the grayscale image into foreground characters and background based on the binarization threshold;
[0033] S1-3-3. Perform mean filtering denoising on the binarized grayscale image to remove the regional noise of the foreground characters and obtain a denoised image;
[0034] S1-3-4. Perform edge detection on the denoised image, extract the character contours, and obtain an image with clear character contours.
[0035] S1-3-5. Perform dilation processing on the image with clear character contours to fill the holes inside the contours caused by noise, and convert the dot-like characters into regular line characters.
[0036] In some embodiments, according to the bounding box coordinates extracted in object detection, crop the region containing each regular line character from the preprocessed cable image to obtain a line character image, including:
[0037] S1-4-1. Read the bounding box coordinates of each dot-like character one by one.
[0038] S1-4-2. Determine the cropping range of the dot-like character according to the bounding box coordinates; wherein, if the bounding box coordinates exceed the boundary of the cable image, adjust the cropping range to the image boundary.
[0039] S1-4-3. Intercept the regular line character region from the preprocessed cable image according to the cropping range to obtain the line character image.
[0040] In some embodiments, the significant feature map retains the high-level abstract features of the maximum local feature significance score through max pooling operation, including:
[0041] S2-3-1. Extract the pixel activation values of the activated features from each activated feature map, calculate and define them as significance scores.
[0042] S2-3-2. Based on the pixel resolution of the activated feature map, predefine the size and stride of the covering window.
[0043] S2-3-3. Slide the covering window to the first covering area according to the stride of the covering window.
[0044] S2-3-4. Calculate the significance scores of each activated feature map within the first covering area of the covering window.
[0045] S2-3-5. Compare the significance scores within the first covering area, and select the activated feature corresponding to the maximum significance score.
[0046] S2-3-6. Define the activated feature corresponding to the maximum significance score as the significant feature, and combine them to generate a significant feature map.
[0047] S2-3-7. Repeat steps S2-3-1 to S2-3-6 to sequentially generate multiple significant feature maps of the line character image.
[0048] In some of these embodiments, the pixel activation values of the activated features are extracted from each activated feature map, calculated and defined as the saliency score, including:
[0049] S2-3-1-1. Extract all the pixel activation values in the activated feature map after convolution kernel and non-linear activation processing, and combine them to generate a set of pixel activation values;
[0050] The expression of the pixel activation value is:
[0051] ;
[0052] Where, represents the pixel activation value, characterizing the position The feature intensity after processing by the convolution kernel and the non-linear activation function; represents the cumulative sum of the convolution kernel during sliding, used to traverse each position of the convolution kernel; represents the weight value of the convolution kernel, represents the two-dimensional index of the convolution kernel; represents the pixel value of the line character image; b represents the bias value, used to adjust the overall offset of the pixel activation value; f represents the non-linear activation function;
[0053] S2-3-1-2. Calculate the saliency score of the activated feature map according to the set of pixel activation values;
[0054] The expression for calculating the saliency score of the activated feature map is:
[0055] ;
[0056] Where, S represents the saliency score, represents the pixel average value of the set of pixel activation values, represents the standard deviation of the set of pixel activation values, used to measure the distribution dispersion degree of the pixel activation values, represents the weight coefficient of the pixel average value, is the weight coefficient of the standard deviation; represents the i-th pixel activation value, and n represents the total number of pixel activation values.
[0057] In some of these embodiments, the one-dimensional vector representation of the line character image is passed through the softmax function to generate the classification label of each line character image, including:
[0058] S5-1. Input the one-dimensional vector of the line character image into the softmax function to convert it into a probability distribution; where, the probability distribution represents the probability that the one-dimensional vector of the line character image belongs to category j.
[0059] S5-2. Select the category j with the highest probability from the probability distribution as the classification label of the line character image.
[0060] In some of these embodiments, obtaining the classification labels of all line character images and combining them in the order of the characters in the image to generate the cable number includes:
[0061] S6-1. Read the classification label of each line character image one by one and store it as a label reading sequence L1 in the reading order;
[0062] The expression of the label reading sequence is:
[0063] ;
[0064] where L1 represents the label reading sequence and k represents the total number of line character images;
[0065] S6-2. According to the spatial order of the line character images in the original cable image, convert the reading order into the spatial order to reorder the label reading sequence and generate a label spatial sequence;
[0066] The expression of the label spatial sequence is:
[0067] ;
[0068] where represents the label spatial sequence, represents the k-th sorted classification label, s represents the index mapping function of the sorted line character images, and represents the mapping from the label reading sequence to the label spatial sequence;
[0069] S6-3. Concatenate the classification labels in the label spatial sequence in sequence to form the cable number N.
[0070] The expression of the cable number N is:
[0071] ;
[0072] where Concat represents the concatenation operation.
[0073] The present invention provides a cable number recognition system based on machine vision, which has the following beneficial effects:
[0074] By preprocessing the dot-shaped characters and converting them into line character images with clear outlines and coherent shapes, it provides optimized data for subsequent deep feature extraction. At the same time, the present invention uses an object detection model to locate the bounding boxes of the cable images, effectively segment the character regions, reduce background interference, and enhance the character morphological characteristics through image preprocessing steps such as grayscale conversion, binarization, noise reduction, edge detection, and dilation processing, improving the character coherence and geometric consistency. Compared with the traditional rule-based matching method, the present invention significantly improves the ability to adapt to complex character shapes and scenarios.
[0075] Furthermore, the significance score of the present invention combines weighted average and standard deviation to quantitatively evaluate the pixel activation value set of the feature map after activation, achieving precise screening of feature importance. Compared with the traditional global pooling method, this method can retain the high-level abstract features with the largest significance score in the local features, thereby generating a significant feature map with higher expressiveness. Finally, combined with convolution kernel extraction and non-linear activation operations, the cable number is generated by combining the classification labels, thus completing the recognition of the cable number. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 is a structural block diagram of the cable number recognition system based on machine vision of the present invention;
[0077] Figure 2 is a schematic diagram of the recognition steps of the cable number recognition system based on machine vision of the present invention;
[0078] Figure 3 is a schematic diagram of the calculation steps of the significance score described in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0079] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0080] Embodiment 1: Please refer to Figures 1 to 3 , the present invention provides a cable number recognition system based on machine vision, which is characterized by including:
[0081] A line character acquisition module for acquiring line character images of cable numbers; wherein, the line character images represent line character images formed after preprocessing the dot characters in the cable numbers, retaining the geometric features and spatial layout of the dot characters and enhancing the coherence of the characters. Compared with the dot characters, the line character images have a more coherent shape and clearer edges, which are suitable for deep feature extraction and classification processing.
[0082] A deep feature map extraction module for extracting a deep feature map of the cable number from the line character image; wherein, the deep feature map represents high-level abstract features extracted from the line character image through a convolutional neural network.
[0083] A feature vector module for flattening the deep feature map of the line character image into a high-dimensional feature vector of the line character image; wherein, the flattening process maps the two-dimensional features in the deep feature map to a one-dimensional space one by one, realizes the linear arrangement of all feature values, and constructs a flat high-dimensional feature vector that can represent the overall features of the line character image.
[0084] A linear mapping module for inputting the high-dimensional feature vector of the line character image into a fully connected layer, and the fully connected layer performs a linear mapping on it to obtain a one-dimensional vector representation of the line character image; wherein, the linear mapping of the fully connected layer multiplies the high-dimensional feature vector by the weight of the fully connected layer and then sums it with the bias of the fully connected layer, thereby mapping it to a one-dimensional vector representation.
[0085] A classification module for generating a classification label for each line character image by passing the one-dimensional vector representation of the line character image through a softmax function; the softmax function is a normalization function mainly used in multi-classification tasks. It calculates the proportion of the exponential value of each component of the input vector to the total exponential value, converts any real value into a probability value ranging from 0 to 1, and ensures that the sum of all probability values is 1. Then, the classification label is determined by selecting the probability.
[0086] A cable number generation module for obtaining the classification labels of all line character images and combining them in the order of the characters in the image to generate the cable number.
[0087] In summary, based on the application scenario of cable number recognition, this embodiment provides a cable number recognition system based on machine vision. The system generates line character images by preprocessing dot characters, making the character shape clearer and more coherent, then combines with deep learning to extract its high-level deep features, and then converts the deep features into corresponding classification labels until a cable number based on machine recognition is generated.
[0088] Among them, the acquisition steps of the line character acquisition module include:
[0089] S1-1. Obtain a cable image containing a cable number; wherein, the cable number is composed of dot-shaped characters;
[0090] S1-2. Perform object detection on the cable image, and extract the bounding box coordinates of each dot-shaped character in the cable image;
[0091] S1-3. Preprocess the cable image after object detection so that the dot-shaped characters in the cable image are converted into regular line characters;
[0092] S1-4. According to the bounding box coordinates extracted in object detection, crop the area containing each regular line character from the preprocessed cable image to obtain a line character image.
[0093] In summary, in this embodiment, by obtaining a cable image containing a cable number, performing object detection on it, and extracting the bounding box coordinates of each dot-shaped character; then, preprocessing the cable image after object detection to convert the dot-shaped characters into regular line characters; finally, according to the extracted bounding box coordinates, crop the area containing each regular line character from the preprocessed cable image to obtain a line character image. This embodiment effectively converts dot-shaped characters into line character images that are more suitable for recognition, laying a data foundation for deep feature extraction.
[0094] Further, the extraction steps of the deep feature map extraction module include:
[0095] S2-1. Receive the line character image as the input of the convolutional model, and extract its local features by sliding the convolutional kernel to obtain N local feature maps in the line character image; wherein, each local feature map represents a type of visual feature of the line character image;
[0096] S2-2. Perform non-linear activation on the N local feature maps to generate N activated feature maps;
[0097] S2-3. Perform pooling on the N activated feature maps to obtain M significant feature maps from several activated feature maps; wherein, the significant feature maps retain the high-level abstract features of the maximum local feature significance scores through the max pooling operation;
[0098] S2-4. Define the set of M significant feature maps as the deep feature map of the line character image.
[0099] In summary, in this embodiment, the convolutional model is used to extract deep features from the line character image. First, the convolutional kernel is slid to extract local features, generating multiple local feature maps. Then, the local feature maps are non-linearly activated to generate more expressive post-activation feature maps. Subsequently, through the max-pooling operation, the features with the largest significance scores are selected from the post-activation feature maps to extract the significant feature maps with high-level abstract meanings. Finally, all the significant feature maps are combined to form the deep feature map. Through the deep feature extraction, this embodiment effectively retains the key visual features in the line character image and can more accurately describe the high-level semantic information of the cable number characters.
[0100] Further, step S1-2 of extracting the bounding box coordinates specifically further includes:
[0101] S1-2-1: Input the cable image into a pre-trained object detection model;
[0102] Among them, the object detection model is iteratively generated based on a convolutional neural network through supervised learning pre-training, and its iterative goal is to extract the image features of the cable image and generate candidate box feature vectors;
[0103] S1-2-2: Divide the cable image input into the object detection model into B×B grid regions;
[0104] S1-2-3: Extract features from each grid region to generate candidate box feature vectors corresponding to each grid region; among them, the candidate box feature vectors include: candidate box boundary coordinates, the category probability that the candidate box belongs to a dot character, and the category confidence;
[0105] S1-2-4: Select the candidate box feature vectors with category confidence higher than a preset threshold from the B×B candidate box feature vectors;
[0106] S1-2-5: Perform non-maximum suppression on the selected candidate box feature vectors, remove overlapping candidate boxes, and retain the candidate box boundary coordinates where the dot characters are located to generate the bounding box coordinates of the dot characters.
[0107] Specifically, in the output of the object detection model, there may be multiple overlapping candidate bounding boxes for the same object (such as dot-shaped characters), and these candidate bounding boxes may correspond to different confidence levels. Non-maximum suppression is a commonly used post-processing algorithm for retaining the optimal candidate bounding boxes from the overlapping candidate bounding boxes and eliminating redundant bounding boxes. It sorts all candidate bounding boxes by confidence level, preferentially retaining the bounding boxes with high confidence levels. It calculates the overlap degree between each candidate bounding box and the currently highest-confidence bounding box. If the overlap degree exceeds a preset threshold, then this candidate bounding box is removed; the above steps are repeated for the remaining candidate bounding boxes until all candidate bounding boxes are processed. Through non-maximum suppression, the candidate bounding boxes with high overlap degrees are removed, and only the candidate bounding box with the highest confidence level and a relatively small overlap degree with other bounding boxes is retained. After removing the redundant bounding boxes, each target area (dot-shaped character) will only correspond to one final candidate bounding box, that is, the boundary coordinates of the candidate bounding box where the dot-shaped character is located are retained.
[0108] Furthermore, the conversion step S1-3 of the regular line characters specifically further includes:
[0109] S1-3-1. Perform grayscale processing on the cable image after object detection, and convert the cable image into a grayscale image; thereby reducing the computational complexity and weakening the influence of illumination.
[0110] Specifically, the original cable image is an RGB color image, where each pixel has three channels (red, green, and blue), and the pixel values of each channel need to be processed separately. The grayscale image is converted into a single-channel image (that is, each pixel only retains one grayscale value) by performing a weighted average on the values of the three RGB channels. After reducing from three channels to one channel, the amount of data for processing each pixel is reduced to one-third of the original, thereby being able to reduce the computational complexity.
[0111] Also, because the RGB color image is sensitive to illumination conditions, for example, the same cable may present different colors and brightnesses under different lighting conditions, resulting in unstable image processing results. Grayscaling compresses the color information into brightness information and removes the influence of the color channels. Only the brightness is retained as the basis for processing, making the processing results less sensitive to different lighting conditions.
[0112] S1-3-2. Perform binarization on the grayscale image, and divide the pixels in the grayscale image into foreground characters and background based on the binarization threshold; thereby enhancing the contrast between the dot-shaped characters and the background through binarization.
[0113] S1-3-3. Perform mean filtering denoising on the grayscale image after binarization to remove the regional noise in the foreground character area and obtain a denoised image; thereby enhancing the overall coherence of the character area.
[0114] S1-3-4. Perform edge detection on the denoised image, extract the character contours, and obtain an image with clear character contours;
[0115] S1-3-5. Perform dilation processing on the image with clear character contours to fill the holes inside the contours caused by noise, and convert the dot-shaped characters into regular conventional line characters.
[0116] In summary, this embodiment discloses the specific method of preprocessing. It sequentially performs grayscale processing, binarization processing, mean filtering for noise reduction, edge detection, and dilation processing on the image, effectively converting the dot-shaped characters into regular conventional line characters. Grayscale processing reduces the computational complexity and weakens the influence of illumination; binarization enhances the contrast between the characters and the background; mean filtering for noise reduction improves the overall coherence of the character region; edge detection makes the character contours clearer; the dilation operation fills the holes in the character contours and generates a coherent and regular line character form. This embodiment significantly enhances the geometric shape and coherence of the characters, providing optimized data for subsequent deep feature extraction and recognition.
[0117] Further, step S1-4 specifically further includes:
[0118] S1-4-1. Read the bounding box coordinates of each dot-shaped character one by one;
[0119] Specifically, the bounding box is a rectangular box used to represent the target area detected in the image. Each bounding box is described by a set of coordinates for its position and size. In the object detection task, the model will generate a corresponding bounding box for each object (such as dot-shaped characters) to describe the position of the character in the image. In this embodiment, the object detection model outputs the rectangular box coordinates of the bounding box, that is, the upper left corner coordinates and the lower right corner coordinates; it is the result generated by the pre-trained object detection model; the object detection model (such as YOLO algorithm, FasterR-CNN, etc.) can automatically identify the objects in the image through training and generate corresponding bounding boxes for each object. Of course, the YOLO algorithm is generally represented by the center point coordinate method, including the center point coordinate , width w and height h; and then through the center point coordinate conversion formula, convert it into the upper left corner coordinate
[0120] and the lower right corner coordinate ;
[0121] The center point coordinate conversion formula is:
[0122] ;
[0123] S1-4-2. Determine the cropping range of the dot-shaped characters according to the bounding box coordinates; wherein, if the bounding box coordinates exceed the boundary of the cable image, adjust the cropping range to the image boundary;
[0124] S1-4-3. Intercept the regular line character area from the preprocessed cable image according to the cropping range to obtain the line character image.
[0125] In summary, in this embodiment, the boundary box coordinates are extracted, and the area containing each regular line character is cropped from the preprocessed cable image, effectively realizing the segmentation and extraction of the character area.
[0126] In this embodiment, the steps for obtaining the significant feature map specifically include:
[0127] S2-3-1. Extract the pixel activation values of the activated features from each activated feature map, calculate and define them as the significance scores.
[0128] S2-3-2. Based on the pixel resolution of the activated feature map, pre-define the size and stride of the coverage window.
[0129] S2-3-3. Slide the coverage window to the first coverage area according to the stride of the coverage window.
[0130] S2-3-4. Calculate the significance scores of each activated feature map within the first coverage area of the coverage window.
[0131] S2-3-5. Compare the significance scores within the first coverage area and select the activated feature corresponding to the maximum significance score.
[0132] S2-3-6. Define the activated feature corresponding to the maximum significance score as the significant feature, and combine them to generate the significant feature map.
[0133] S2-3-7. Repeat steps S2-3-1 to S2-3-6 to sequentially generate multiple significant feature maps of the line character image.
[0134] In summary, in this embodiment, the significance scores of each activated feature map are calculated through pixel activation values, and the significance scores within the coverage area are calculated by gradually sliding the coverage window; the feature corresponding to the maximum significance score within the coverage area is selected as the significant feature, and finally the significant feature map is generated by combination. Through the extraction of the significant feature map, the abstract expression ability of the deep features is improved.
[0135] Furthermore, the calculation steps of the significance scores in this embodiment include:
[0136] S2-3-1-1. Extract all the pixel activation values after convolution kernel and non-linear activation processing within the activated feature map, and combine them to generate a pixel activation value set.
[0137] The expression of the pixel activation value is:
[0138] ;
[0139] Among them, represents the pixel activation value, characterizing the position The feature intensity after being processed by the convolutional kernel and the non-linear activation function; represents the cumulative sum of the convolutional kernel during sliding, used to traverse each position of the convolutional kernel; represents the weight value of the convolutional kernel, represents the two-dimensional index of the convolutional kernel; represents the pixel value of the line character image; b represents the bias value, used to adjust the overall offset of the pixel activation value; f represents the non-linear activation function;
[0140] S2-3-1-2. Calculate the significance score of the activated feature map according to the set of pixel activation values;
[0141] The expression for calculating the significance score of the activated feature map is:
[0142] ;
[0143] Among them, S represents the significance score, represents the pixel average value of the set of pixel activation values, represents the standard deviation of the set of pixel activation values, used to measure the distribution dispersion degree of the pixel activation values, represents the weight coefficient of the pixel average value, is the weight coefficient of the standard deviation; represents the i-th pixel activation value, and n represents the total number of pixel activation values.
[0144] In summary, in this embodiment, the pixel activation values extracted by the convolutional kernel are combined with the non-linear activation function to form a set of pixel activation values, and the significance score formula considering the average value and the dispersion degree is used to efficiently quantify the significance score of the activated feature map.
[0145] Embodiment 2: The technical solution of Embodiment 2 is different from that of Embodiment 1 in that the specific steps for generating a cable number based on the one-dimensional vector of the line character image in Embodiment 1 are disclosed, including:
[0146] S5-1. Input the one-dimensional vector of the line character image into the softmax function to convert it into a probability distribution; among them, the probability distribution represents the probability that the one-dimensional vector of the line character image belongs to category j.
[0147] S5-2. Select the category j with the largest probability from the probability distribution as the classification label of the line character image.
[0148] In summary, in this embodiment, by performing probability distribution transformation on the one-dimensional vector representation of the line character image, using the softmax function for normalization to generate a classification probability distribution, and by selecting the category corresponding to the maximum probability value, the classification label is accurately generated.
[0149] Among them, the steps of obtaining the cable number according to the classification label include:
[0150] S6-1. Read the classification label of each line character image one by one , and store them in the label reading sequence L1 in the reading order;
[0151] The expression of the label reading sequence is:
[0152] ;
[0153] Among them, L1 represents the label reading sequence, and k represents the total number of line character images;
[0154] S6-2. According to the spatial order of the line character image in the original cable image, convert the reading order into the spatial order to reorder the label reading sequence and generate the label space sequence;
[0155] The expression of the label space sequence is:
[0156] ;
[0157] Among them, represents the label space sequence, represents the k-th sorted classification label, s represents the index mapping function of the sorted line character image, representing the mapping from the label reading sequence to the label space sequence,
[0158] S6-3. Concatenate the classification labels in the label space sequence in sequence to obtain the cable number N.
[0159] The expression of the cable number N is:
[0160] ;
[0161] Among them, Concat represents the concatenation operation.
[0162] In summary, in this embodiment, by combining the order of the classification labels of the line character images, the cable number is finally generated, realizing the automatic recognition of the cable number based on machine vision. This system is based on the dot characters in the cable image. Through target detection and image preprocessing, the dot characters are converted into line characters with coherent shapes and clear edges. Subsequently, based on the deep learning model, local features are extracted by sliding the convolution kernel, the feature expression ability is enhanced through non-linear activation, and the activated features are screened using the saliency score formula to extract the saliency feature map to construct the deep feature map. In the saliency score formula, the weighted average is combined with the standard deviation, and by quantitatively evaluating the distribution characteristics of the pixel activation value set, the effective extraction of high-level features is achieved, enhancing the reliability and accuracy of the recognition process.
[0163] After that, the classification labels are generated through linear mapping and the softmax function; finally, the cable number is generated according to the order combination of the classification labels. By introducing the saliency score formula and the feature screening of the deep learning model, this embodiment improves the automation accuracy of cable number recognition and is particularly suitable for intelligent management in complex cable scenarios.
[0164] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired (such as infrared, wireless, microwave, etc.) manner.
[0165] The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0166] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and systems can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a division of a channel underwater terrain change analysis system and its system logic functions. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0167] As described above, the foregoing is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.
Claims
1. A cable number recognition system based on machine vision, characterized in that Including: A line character acquisition module, configured to acquire a line character image of a cable number; A depth feature map extraction module, configured to extract a depth feature map of the cable number from the line character image; A feature vector module, configured to flatten the depth feature map of the line character image into a high-dimensional feature vector of the line character image; A linear mapping module, configured to input the high-dimensional feature vector of the line character image into a fully connected layer, and the fully connected layer performs linear mapping on it to obtain a one-dimensional vector representation of the line character image; A classification module, configured to pass the one-dimensional vector representation of the line character image through a softmax function to generate a classification label for each line character image; A cable number generation module, configured to acquire the classification labels of all line character images and combine them in the order of characters in the image to generate the cable number; Extracting the depth feature map of the cable number from the line character image includes: Receiving the line character image as the input of a convolutional model, and extracting its local features by sliding a convolutional kernel to obtain N local feature maps in the line character image; wherein, each local feature map represents a type of visual feature of the line character image; Performing non-linear activation on the N local feature maps to generate N activated feature maps; Performing pooling on the N activated feature maps to select M significant feature maps from several activated feature maps; wherein, the significant feature maps retain the high-level abstract features of the maximum local feature significance score through a max pooling operation; Defining the set of M significant feature maps as the depth feature map of the line character image; The selecting M significant feature maps from several activated feature maps includes: Extracting the pixel activation values of the activated features from each activated feature map, and calculating and defining them as significance scores; The calculating and defining them as significance scores includes: Extracting all pixel activation values after convolution kernel and non-linear activation processing in the activated feature map, and combining them to generate a pixel activation value set; The expression of the pixel activation value is: ; Among them, represents the pixel activation value, characterizing the position The feature intensity after being processed by the convolution kernel and the non-linear activation function; represents the cumulative sum during the sliding of the convolution kernel, used to traverse each position of the convolution kernel; represents the weight value of the convolution kernel, represents the two-dimensional index of the convolution kernel; represents the pixel value of the line character image; b represents the bias value, used to adjust the overall offset of the pixel activation value; f represents the non-linear activation function; Calculating the significance score of the activated feature map according to the pixel activation value set; The expression for calculating the significance score of the activated feature map is: ; Among them, S represents the significance score, represents the pixel average value of the set of pixel activation values, represents the standard deviation of the set of pixel activation values, which is used to measure the distribution dispersion degree of the pixel activation values, represents the weight coefficient of the pixel average value, is the weight coefficient of the standard deviation; represents the i-th pixel activation value, and n represents the total number of pixel activation values.
2. The cable number recognition system based on machine vision according to claim 1, characterized in that, Acquiring the line character image of the cable number includes: S1-1. Acquiring a cable image containing the cable number; wherein, the cable number is composed of dot characters; S1-2. Performing object detection on the cable image, and extracting the bounding box coordinates of each dot character in the cable image; S1-3. Preprocessing the cable image after object detection to convert the dot characters in the cable image into regular line characters; S1-4. According to the bounding box coordinates extracted in object detection, cropping out the region containing each regular line character from the preprocessed cable image to obtain a line character image.
3. The cable number recognition system based on machine vision according to claim 2, characterized in that, Performing object detection on the cable image and extracting the bounding box coordinates of each dot character in the cable image includes: S1-2-1. Inputting the cable image into a pre-trained object detection model; Among them, the target detection model is iteratively generated based on a convolutional neural network through supervised learning pre-training, and its iterative goal is to extract the image features of the cable image and generate candidate box feature vectors. S1-2-2: Divide the cable image input into the target detection model into B×B grid regions. S1-2-3: Extract features from each grid region to generate a candidate box feature vector corresponding to each grid region. Among them, the candidate box feature vector includes: candidate box boundary coordinates, the category probability that the candidate box belongs to a dot character, and the category confidence. S1-2-4: Select candidate box feature vectors with a category confidence higher than a preset threshold from the B×B candidate box feature vectors. S1-2-5: Perform non-maximum suppression on the selected candidate box feature vectors, eliminate overlapping candidate boxes, retain the candidate box boundary coordinates where the dot characters are located, and generate the boundary box coordinates of the dot characters.
4. The cable number recognition system based on machine vision according to claim 2, characterized in that Preprocess the cable image after target detection to convert the dot characters in the cable image into regular line characters, including: S1-3-1: Perform grayscale processing on the cable image after target detection to convert the cable image into a grayscale image. S1-3-2: Perform binarization of the grayscale image, and divide the pixels in the grayscale image into foreground characters and background based on the binarization threshold. S1-3-3: Perform mean filtering denoising on the binarized grayscale image to remove the regional noise of the foreground characters and obtain a denoised image. S1-3-4: Perform edge detection on the denoised image to extract the character contour and obtain an image with clear character contours. S1-3-5: Perform dilation processing on the image with clear character contours to fill the holes inside the contour caused by noise, and convert the dot characters into regular line characters.
5. The cable number recognition system based on machine vision according to claim 2, wherein According to the boundary box coordinates extracted in target detection, crop the regions containing each regular line character from the preprocessed cable image to obtain line character images, including: S1-4-1: Read the boundary box coordinates of each dot character one by one. S1-4-2: Determine the cropping range of the dot character according to the boundary box coordinates. Among them, if the boundary box coordinates exceed the cable image boundary, adjust the cropping range to the image boundary. S1-4-3: Intercept the regular line character region from the preprocessed cable image according to the cropping range to obtain the line character image.
6. The cable number recognition system based on machine vision according to claim 3, characterized in that The selection of M significant feature maps from several activated feature maps further includes: Based on the pixel resolution of the activated feature map, predefine the size and step of the coverage window. According to the step of the coverage window, slide the coverage window to the first coverage area. Calculate the significance score of each activated feature map in the first coverage area of the coverage window. Compare the significance scores in the first coverage area and select the activated feature corresponding to the maximum significance score. Define the activated feature corresponding to the maximum significance score as a significant feature, and combine them to generate a significant feature map. Repeat to generate M significant feature maps of the line character image.
7. The cable number recognition system based on machine vision according to claim 6, characterized in that Generate the classification label for each line character image by passing the one-dimensional vector representation of the line character image through the softmax function, including: S5-1. Input the one-dimensional vector of the line character image into the softmax function to convert it into a probability distribution; wherein, the probability distribution represents the probability that the one-dimensional vector of the line character image belongs to category j; S5-2. Select the category j with the highest probability from the probability distribution as the classification label of the line character image.
8. The cable number recognition system based on machine vision according to claim 7, wherein Obtain the classification labels of all line character images, and combine them in the order of the characters in the image to generate the cable number, including: S6-1. Read the classification label of each line character image one by one , and store them in the label reading sequence L1 in the reading order; The expression of the label reading sequence is: ; wherein, L1 represents the label reading sequence, and k represents the total number of line character images; S6-2. According to the spatial order of the line character images in the original cable image, convert the reading order into the spatial order to reorder the label reading sequence and generate the label space sequence; The expression of the label space sequence is: ; Among them, represents the tag space sequence, represents the k-th sorted classification tag, s represents the index mapping function of the sorted line character image, and represents the mapping from the tag reading sequence to the tag space sequence; S6-3. Concatenate the classification labels in the label space sequence in sequence to obtain the cable number N; The expression of the cable number N is: ; wherein, Concat represents the concatenation operation.
Citation Information
Patent Citations
Low-computing-power character rapid recognition method based on machine vision
CN118366174A
Power transmission line data acquisition method based on Yolov3 and CRNN algorithm
CN115331240A
Container number and container threshold weight identification method and device based on deep learning
CN119251824A