Blind Braille Image Recognition Method Based on Object Detection

Through the target detection method, the identification of Braille images is solved by using the feature pyramid structure, and the problems of cumbersome and inefficient traditional methods are solved, and more efficient Braille recognition is achieved.

CN116612490BActive Publication Date: 2025-05-27LANZHOU UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310649217.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-02
Publication Date
2025-05-27
Estimated Expiration
2043-06-02

AI Technical Summary

Technical Problem

Traditional Braille image recognition methods are cumbersome, inefficient and require a lot of manual effort.

Method used

Using a Braille image recognition method based on object detection, the Braille data set and object detection model are constructed, and the feature pyramid structure is used for identification.

Benefits of technology

It reduces the unnecessary steps in the identification process, improves the efficiency of Braille recognition, and realizes a more concise identification process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116612490B_ABST
    Figure CN116612490B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for braille image recognition based on target detection, comprising: constructing a braille data set; constructing a target detection model based on a pyramid feature fusion structure; training the target detection model based on the braille data set; and recognizing braille based on the trained target detection model. The braille image recognition method proposed by the present invention reduces redundant steps in the recognition process and combines most of the cumbersome steps into a whole. The present invention not only reduces the recognition steps, but also improves the accuracy of braille recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image recognition, and particularly relates to a method for recognizing braille images based on object detection. Background Art

[0002] Braille, also known as dot writing, has two font description forms: raised characters and recessed characters. It is a kind of writing specially designed for the tactile perception of the blind. Braille signs can be seen everywhere in real life, such as at elevator buttons, toilet signs, and on the packaging boxes of medical drugs. Usually, braille is composed of different combinations of dot characters engraved on paper manually or by a dot board, etc.

[0003] Traditional braille image recognition methods, as Figure 1 shown, can be roughly divided into the following processes: First, preprocess the braille image, including operations such as denoising and filtering; second, correct the braille image in the form of artificial auxiliary lines; then use the threshold method to segment the braille dots and the background to detect the braille dots; after that, detect and locate the braille square, extract the features of the braille square, and finally send them into the classifier for training and recognition to obtain the final recognition result.

[0004] The implementation of traditional braille image recognition methods is relatively cumbersome, requiring multiple complex recognition steps to be divided. Therefore, the braille recognition efficiency is relatively low, and at the same time, it requires a lot of manual effort and machine assistance. Summary of the Invention

[0005] To solve the above technical problems, the present invention proposes a method for recognizing braille images based on object detection, which reduces the redundant steps in the recognition process and combines most of the cumbersome steps into an overall one.

[0006] To achieve the above object, the present invention proposes a method for recognizing braille images based on object detection, including:

[0007] A method for recognizing braille images based on object detection, characterized by including:

[0008] Construct a braille data set;

[0009] Based on the gold pyramid feature fusion structure, construct an object detection model;

[0010] Based on the braille data set, train the object detection model;

[0011] Based on the trained object detection model, recognize the braille.

[0012] Optionally, the braille data set includes: a character braille data set, a single-sided braille data set, and a double-sided braille data set.

[0013] Optionally, after constructing the braille dataset, it includes: performing image preprocessing on the braille dataset;

[0014] The image preprocessing includes: image binarization, brightness adjustment, image deskewing, and size adjustment.

[0015] Optionally, the object detection model is constructed using a feature pyramid structure;

[0016] The object detection model includes: a ResNet-50 module, an Upsample module, and a Class+BoxSubnet module connected in sequence.

[0017] Optionally, training the object detection model includes:

[0018] Based on the ResNet-50 module, extracting features from the braille images in the braille dataset to obtain features at different levels;

[0019] Based on the Upsample module, magnifying the features at different levels and performing feature fusion;

[0020] Based on the Class+Box Subnet module, identifying the image after feature fusion.

[0021] Optionally, the ResNet-50 module includes 4 stage sub-modules, namely: stage0, stage1, stage2, and stage3; among them, for the 4 stage modules, the first stage0 does not have a residual structure, while stage1 uses 3 residual structures, stage2 uses 4 residual structures, combines the Conv4 and Conv5 structures, and stage3 uses 9 residual structures.

[0022] Optionally, the Upsample module includes: several 1×1 convolutional kernels and an upsampling part;

[0023] In the Upsample module, after the features pass through the 1×1 convolutional kernels, the size of the feature map is changed, and the feature map is sent to the upsampling part to be added to the feature map of the previous layer. After addition, the final feature map of the corresponding layer is obtained.

[0024] Optionally, the Class+Box Subnet module includes: a bounding box regression subnet and a classification subnet;

[0025] Both the bounding box regression subnet and the classification subnet include: several layers of 3×3 convolutional layers and the activation function ReLU; among them, the number of channel dimensions of the last convolutional layer in the bounding box regression subnet is a multiple of 64, and the number of channel dimensions of the last convolutional layer in the classification subnet is a multiple of 4.

[0026] Optionally, the recognition of the image after feature fusion includes:

[0027] Based on the box regression sub-network, obtain the box position of the image after feature fusion;

[0028] Based on the classification sub-network, obtain the category information of the image after feature fusion.

[0029] Compared with the prior art, the present invention has the following advantages and technical effects:

[0030] The braille image recognition method proposed by the present invention reduces the redundant steps in the recognition process and combines most of the cumbersome steps into an integral whole. In this way, the braille image recognition method is more like a black box, which inputs from one end, passes through the designed braille recognition black box, and outputs the recognition result of the braille from the other end. Description of the Drawings

[0031] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0032] Figure 1 is a schematic flow chart of a traditional braille image recognition method;

[0033] Figure 2 is a schematic diagram of different scale fusion methods in the pyramid of the embodiment of the present invention; wherein, Figure 2 (a) is a specified picture pyramid, Figure 2 (b) is a single-layer feature map, Figure 2 (c) is the hierarchical pyramid feature, Figure 2 (d) is a feature pyramid structure;

[0034] Figure 3 is a schematic diagram of the overall framework of the object detection model in the embodiment of the present invention;

[0035] Figure 4 is a two-line interpolation statistical chart in the embodiment of the present invention;

[0036] Figure 5 is a schematic diagram of the Class+Box Subnet module in the embodiment of the present invention;

[0037] Figure 6 is a schematic diagram of three braille data sets in the embodiment of the present invention; wherein, Figure 6 (a) is a character braille data set, Figure 6 (b) is a single-sided braille data set, Figure 6 (c) is a double-sided braille data set;

[0038] Figure 7 Schematic diagram of partial character braille images according to an embodiment of the present invention;

[0039] Figure 8 Schematic diagram of partial single-sided braille data according to an embodiment of the present invention;

[0040] Figure 9 Schematic diagram of the DBSI dataset according to an embodiment of the present invention;

[0041] Figure 10 Schematic diagram of the category and data display of the AngelinaDataset dataset according to an embodiment of the present invention;

[0042] Figure 11 Schematic diagram of image preprocessing - binarization operation according to an embodiment of the present invention;

[0043] Figure 12 Schematic diagram of image preprocessing - image debiasing operation according to an embodiment of the present invention;

[0044] Figure 13 Schematic diagram of the grid of single-sided braille image recognition results according to an embodiment of the present invention;

[0045] Figure 14 Schematic diagram of individual pictures stained, creased, or worn in the braille dataset according to an embodiment of the present invention;

[0046] Figure 15 Schematic diagram of the grid of double-sided braille image recognition results according to an embodiment of the present invention;

[0047] Figure 16 Schematic diagram of the braille recognition method based on object detection according to an embodiment of the present invention. Detailed implementation manners

[0048] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments may be combined with each other. The following will describe this application in detail with reference to the drawings and in combination with the embodiments.

[0049] It should be noted that the steps shown in the flowchart of the drawings may be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.

[0050] The braille image recognition method based on object detection proposed by the present invention, as Figure 16 shown, performs recognition in an overall manner. First, image preprocessing and image correction operations are carried out, and then the specific information of each braille value in the label information and the braille box information are input into the Feature Pyramid PyramidNet, and finally the object detection result is obtained. It includes:

[0051] Construct a Braille dataset;

[0052] Based on the pyramid feature fusion structure, construct an object detection model;

[0053] Based on the Braille dataset, train the object detection model;

[0054] Based on the trained object detection model, recognize Braille.

[0055] Furthermore, the Braille dataset includes: a character Braille dataset, a single-sided Braille dataset, and a double-sided Braille dataset.

[0056] Furthermore, perform image preprocessing on the Braille dataset;

[0057] The image preprocessing includes: image binarization, brightness adjustment, image debiasing, and size adjustment.

[0058] Furthermore, the object detection model is constructed using a feature pyramid structure;

[0059] The object detection model includes: an improved ResNet-50 module, an Upsample module, and a Class+Box Subnet module connected in sequence.

[0060] Furthermore, training the object detection model includes:

[0061] Based on the ResNet-50 module, extract features from the Braille images in the Braille dataset to obtain features at different levels;

[0062] Based on the Upsample module, magnify the features at different levels and perform feature fusion;

[0063] Based on the Class+Box Subnet module, recognize the image after feature fusion.

[0064] Furthermore, the improved ResNet-50 module contains 4 stage modules. The first stage0 does not have a residual structure, while stage1 uses 3 residual structures, stage2 uses 4 residual structures, combines the Conv4 and Conv5 structures, and stage3 uses 9 residual structures.

[0065] Furthermore, the Upsample module includes: several 1×1 convolutional kernels and an upsampling part;

[0066] In the Upsample module, after the features pass through the 1×1 convolutional kernel, the size of the feature map is changed, and it is sent to the upsampling part to be added to the feature map of the previous layer. After addition, the final feature map of the corresponding layer is obtained.

[0067] Furthermore, the Class+Box Subnet module includes: a box regression subnet and a classification subnet;

[0068] Both the box regression subnet and the classification subnet include: several 3×3 convolutional layers and the activation function ReLU; among them, the number of channel dimensions of the last convolutional layer in the box regression subnet is a multiple of 64, and the number of channel dimensions of the last convolutional layer in the classification subnet is a multiple of 4.

[0069] Furthermore, identifying the image after feature fusion includes:

[0070] Based on the box regression subnet, obtaining the box position of the image after feature fusion;

[0071] Based on the classification subnet, obtaining the category information of the image after feature fusion.

[0072] In this embodiment, the constructed object detection model is the PyramidNet model with a pyramid structure based on a convolutional neural network. This model is an innovative object detection model based on the RetinaNet model structure and the FPN model structure, and is used to detect small object structures. (FPN is a typical pyramid model, which makes full use of image information and is suitable for small target detection. The RetinaNet model is improved based on the FPN model. The model in this embodiment is based on the RetinaNet model and the FPN model, and the improvement part includes increasing the number of features in the model, making full use of the features of the blind side, and improving the accuracy of the model). When designing this model, the present invention studied Figure 2 different scale fusion methods in the pyramid. Figure 2 The (a) method of is the specified picture pyramid (Featurized image pyramid), where the picture is cropped into sub-pictures of different sizes, and the feature maps of the corresponding sub-pictures are calculated and the results are predicted respectively. Although this method can accurately predict each piece of information, it requires a large amount of computing resources and the process is cumbersome. Figure 2 The (b) method of is the single feature map (Sigle feature map), where the deep feature map containing rich semantic information is calculated and the results are predicted. This method will lose some semantic information. Figure 2 The (c) method of is the pyramidal feature hierarchy, where the information of different feature scales is calculated and the results are predicted respectively. This method will lose a small amount of semantic information in the feature maps at different levels. Therefore, by synthesizing the above three different scale fusion methods of the pyramid, this embodiment uses Figure 2 the (d) method of , that is, the pyramid feature fusion structure.

[0073] PyramidNet belongs to one-stage object detection algorithms, which directly extract features through a convolutional neural network and predict the classification and localization of objects. Different from two-stage object detection algorithms (which first generate candidate regions and then classify and localize the candidate regions through a convolutional neural network), this model directly outputs class and location information through the backbone network, simplifying the model parameters and improving the training efficiency. As Figure 3 shown, PyramidNet adopts the Figure 2 (d) method of

[0074] , that is, it uses the feature maps extracted by each stage architecture for interpolation and fusion operations and sends them to the Class+Box Subnet part. This model is mainly divided into three parts: the ResNet Backone module, the Upsample module, and the Class+Box Subnet module. 2 -p 7 The overall process of the braille recognition method based on object detection is as follows: First, the preprocessed braille data passes through the improved RestNet-50 to obtain starting feature maps of different sizes. After the feature maps of different levels are fused and added through the Upsample part, convolution is performed to obtain different levels of p 2 -p 7 total feature map part. Finally, the final braille recognition result is obtained through the sub-network Class+BoxSubnet module. One of the innovation points of the PyramidNet model is that the model uses p 2 -p 7 feature map part. The feature maps of the p 7 part are obtained by performing convolution and activation function operations on the p 6 feature maps. The feature maps of the p 6 part are obtained by performing convolution and activation function operations on the p 5 feature maps. The feature maps of the p 2 part are the p 3 feature maps obtained by convolution and activation function. The feature maps of p 3 , p 4 , p 5 are all obtained by adding the feature maps extracted by the backbone network and the feature maps after bilinear interpolation. This increases the rich semantic information of braille images at different levels and improves the accuracy of braille image recognition.

[0075] The main steps for constructing the PyramidNet model are as follows:

[0076] (1) ResNet Backbone module

[0077] The PyramidNet backbone network part completely uses the ResNet-50 model. The ResNet-50 model using batch normalization has three advantages: First, it stabilizes the braille image data input to each layer of the network; second, it reduces the parameter sensitivity in the network, making the model more stable; third, it has a regularization effect in feature extraction.

[0078] In this embodiment, to save more computing resources and extract the main features of the braille image, an improved ResNet-50 structure is adopted, which includes 4 stage modules. The first stage0 does not have a residual structure, while stage1 uses 3 residual structures, stage2 uses 4 residual structures, the Conv4 and Conv5 structures are merged, and stage3 uses 9 residual structures. The specific network parameters are shown in Table 1:

[0079] Table 1 ResNet structure

[0080]

[0081]

[0082] (2) Upsample module

[0083] The main function of the Upsample (upsampling) part is to enlarge the features of different levels extracted from the ResNet-50 backbone network to the same size as the previous feature layer for feature fusion operations. The purpose of using multiple 1×1 convolutional kernels in the Upsample part is not to change the size of the feature map, but to change the number of feature maps, which has the effect of reducing the dimension in the model and fusing and adding with the interpolated feature map. There are three common upsampling algorithms: nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation. Since single linear interpolation calculates using the corresponding values of two points, the resulting error is relatively large. To avoid the above problems, this embodiment uses bilinear interpolation as the image frame interpolation algorithm to change the size of the feature map without changing the features themselves. Bilinear interpolation is a linear interpolation extension of the interpolation function with two variables. The core idea of this algorithm is to perform linear interpolation once in both the horizontal and vertical directions. Figure 4 For the specific process of interpolation.

[0084] According to the bilinear interpolation statistical chart, to obtain the pixel value at the position of point p(x, y), it is necessary to know the Q points around p 11 (x 1 , y 1 ), Q 12 (x 1 , y 2 ), Q 21 (x 2 , y 1)、Q 22 (x 2 ,y 2 ) pixel values of four points. According to the calculation formula of the intersection point of the line segments between these four points, the pixel value of point p is finally calculated to enlarge the size of the feature map on the braille image. First, linear interpolation is performed in the x-axis direction of the coordinate axis, so R 1 、R 2 values in the x direction are obtained:

[0085]

[0086] Then, linear interpolation is performed in the y-axis direction of the coordinate axis, and the value p can be obtained:

[0087]

[0088] Combining the above three formulas and substituting them into each other, the final result of bilinear interpolation is p(x, y):

[0089]

[0090] (3) Class+Box Subnet module

[0091] This module is a bounding box regression network (Box SubNet) and a classification network (Class SubNet). Its main function is to send the feature map after linear interpolation and feature fusion addition into the box regression and classification subnets with the same weights to obtain all box positions and class information. For the braille image, it is to obtain braille information of 64 categories. As shown in Figure 5 , both the box regression subnet and the classification subnet are composed of four 3×3 convolutional layers and the activation function ReLU, and the length of their feature maps remains unchanged. The difference between the two is that when passing through the last convolutional layer, the number of channel dimensions of the classification subnet is a multiple of 64, while the number of channel dimensions of the box regression subnet is a multiple of 4 when passing through the last convolutional layer.

[0092] Further, the construction of the braille dataset is also disclosed in this embodiment;

[0093] The datasets used in this embodiment are all self-built or publicly available datasets. The pictures in the datasets are sourced from publicly available braille pictures, obtained by mobile phone shooting, and related picture cropping. The types of data are mainly divided into two categories: printed or handwritten. The character braille dataset and the single-sided braille dataset used in this embodiment are all concave or convex points. In the double-sided braille dataset, these points include both concave points and convex points, and there is a certain inclination angle between them, and the distance between the braille points is relatively close. These three datasets have a common feature, that is, there is an approximately equal distance between each individual braille square and the surrounding braille squares. AsFigure 6 Diagrams showing three Braille datasets; among them, Figure 6 (a) of shows the character Braille dataset, Figure 6 (b) of shows the single-sided Braille dataset, Figure 6 (c) of shows the double-sided Braille dataset, among which, recognizing double-sided Braille data is the most difficult. In this embodiment, the task of recognizing the character Braille dataset is regarded as a classification task, and the tasks of recognizing the single-sided Braille dataset and the double-sided Braille dataset are regarded as classification tasks or detection tasks.

[0094] Character Braille dataset:

[0095] The character Braille dataset (CharsDataSet) is a self-built dataset. The labels of this dataset are all composed of a - z and 0 - 9. Each picture is a Braille square dataset composed of 6 dots. After the pictures in the dataset are cropped, they are composed of color images with a size of 28×28 pixels. This dataset contains a total of 1583 color or gray - white pictures, from 36 different categories. Each category is identified by the name of the picture, and each picture corresponds to a label. To increase the number of pictures, each picture is processed by rotating clockwise or counterclockwise with the same quality, changing the brightness of the picture, randomly cropping, adding noise, etc. After final sorting, some character Braille images are as shown in Figure 7 shown.

[0096] Single - sided Braille dataset:

[0097] The single - sided Braille dataset (SigleDataSet) is a self - built dataset. Due to a large amount of Braille square recognition workload, this dataset is a line dataset. The pictures in the dataset come from two different sources: one is to crop the whole Braille image, and the other is to crawl data from the Braille database. This dataset consists of two parts: a training set and a test set. The training set consists of 53×67 color character Braille pictures of 36 classes, a total of 26878 pictures; the test set consists of 312×42 color single - sided Braille pictures of 36 classes, a total of 100 pictures. The overall data is shown in Figure 8 shown.

[0098] Double - sided Braille dataset:

[0099] This double - sided Braille dataset (DBSI) is made by a scanner and provided on Github. This dataset consists of 114 color double - sided Braille pictures, which come from 6 different Braille books and 1 ordinary printed Braille document respectively. The entire DBSI provides both the single - sided Braille dataset and the double - sided Braille dataset.

[0100] The same number of double-sided Braille images, front Braille images, back Braille images, and corresponding txt annotation files in each type of file. For each txt file, the first line is the inclination angle of the Braille image. When the number is positive, the image is rotated clockwise; if the angle is negative, the Braille image is rotated counterclockwise. The second line is the position of the vertical lines of each Braille cell. According to the characteristics of the Braille cell, the number of vertical lines must be uniform. The third line is the distance of the horizontal lines. According to the characteristics of the Braille dots, the number of horizontal lines is a multiple of 3. The remaining multiple lines are the positions of the Braille squares and their corresponding labels. Each line has a total of 8 numbers. The first 2 digits correspond to the position of the Braille square, and the last 6 digits are the labels of the Braille square (composed of 0 and 1). The entire DBSI provides both a single-sided Braille dataset and a double-sided Braille dataset. The DBSI dataset is as Figure 9 shown.

[0101] Another double-sided Braille dataset (AngelinaDataset) is obtained through web crawling to supplement the quantity of the DBSI dataset. This dataset consists of 250 color pictures, among which 212 pictures are from Braille books, and the remaining 38 pages are from handwritten Braille pictures by students. This dataset is divided into 191 training pictures (80%) and 49 test images (20%). In addition, to increase the diversity of the data, 44 pictures of non-Braille texts from the Internet are added to the training set as negative examples.

[0102] Each picture corresponds to a corresponding json file. The content in the json file records the information of each Braille square. The first line shows the label of the Braille square, which can be converted into Russian Braille in a certain encoding way. The second and third lines both represent the colors of the lines and dots of the Braille square. The third line represents the position coordinate information of the Braille square in the picture, which is convenient for subsequent code to detect the position of the entire Braille square box. The categories and data display of this dataset are as Figure 10 shown.

[0103] To enrich the double-sided dataset, in this embodiment, the DBSI dataset and the AngelinaDataset dataset are merged into a large dataset, with a total of 360 double-sided Braille color pictures.

[0104] Furthermore, in this embodiment, image preprocessing is performed on the images of the three Braille datasets to make it easier for the backbone network to extract the main features of the Braille images. The image preprocessing operations include: image binarization, brightness adjustment, image deskewing, size adjustment, etc. This embodiment focuses on a detailed introduction to image binarization and image deskewing.

[0105] Image binarization operation:

[0106] To more easily capture Braille dot information, the Braille image needs to be binarized on the basis of grayscale conversion. First, the color image is grayscale processed, that is, it is converted into an image that only contains brightness information and does not contain color information. The characteristic of this grayscale image is that the brightness changes continuously from dark to bright. Each pixel of the grayscale image only needs one byte to store the grayscale value, and the grayscale range is 0 - 255. The grayscale conversion formula is as follows:

[0107] Y = 0.229R + 0.599G + 0.112B (5)

[0108] Where Y represents the grayscale value, and R, G, B represent the red, green, and blue components of the color Braille image.

[0109] Image binarization is the process of setting the pixel values of the Braille grayscale image to 0 or 255. By selecting an appropriate threshold, the target is separated from the background, so that the whole image presents an obvious black and white effect. In the binarized image, the Braille dots appear white and the background area appears black. For the threshold segmentation method of image binarization, there are three common methods: One is global binarization. This method uses a specified threshold T to divide all pixel points in the Braille grayscale image area into two categories. If the pixel value is greater than the threshold T, the pixel value is set to 255, otherwise it is set to 0. The mathematical principle of this method is simple, but the grayscale characteristics of the Braille image edge and the center position are very different. Using the specified threshold to segment the background and the target will result in unreasonable segmentation of a large number of areas. This method is suitable for simple images. The second is local binarization. This method is an improvement over global binarization. It divides a Braille image into multiple sub - Braille images, and each sub - Braille image selects a threshold according to its corresponding characteristics, and each area is separately threshold - segmented to perform global binarization of multiple small images. Compared with the global binarization method, this method has a more uniform segmentation and better effect.

[0110] The third is local adaptive binarization, which improves the rationality of threshold selection. This method combines numerical values such as the average value and variance of pixels in each part, and obtains the optimal threshold through corresponding operations. This calculation method improves the binarization effect of the details of the image. Therefore, in this paper, the local adaptive binarization method is used to pre - process the Braille image, Figure 11 is the result of local adaptive binarization. The Braille image is on the left and the binarized image is on the right.

[0111] Image deskewing operation

[0112] The image debiasing operation is the most important operation in the preprocessing of Braille images. Although the arrangement of Braille dots and Braille cells is standard in the horizontal or vertical direction, through image binarization or visual inspection, it is found that most Braille images have a slight tilt angle, which will increase the difficulty of Braille grid division and character recognition. Now, a large number of researchers have conducted extensive research on the image debiasing operation, including classical methods such as linear regression and Hough transform. Among them, Manners J et al. proposed calculating the sum deviation of all rows to calculate the image rotation angle, and the image is continuously tilted by one pixel difference in the vertical direction until there is no row sum deviation; the Chinese Academy of Sciences team proposed a two-stage angle debiasing method from coarse to fine using the variance maximization of the horizontal and vertical projections of the binary image.

[0113] This embodiment attempts to use linear regression to find the most fitting straight line for image debiasing operation for two reasons: one is that all Braille dot coordinate information is presented in the dataset, which is convenient for directly calculating the angle; the other is that the operation is simple and the processing time of Braille images is relatively fast.

[0114] First, set the reference line:

[0115] Y = Bx + a (6)

[0116] Secondly, the value of B needs to be obtained from the coordinate values of the Braille dots, that is, the image rotation angle is known:

[0117]

[0118] where n is the number of pixels, x i and y i are the coordinates of the Braille dots in the Braille image. After obtaining the image rotation angle according to the formula, the Rotate function is directly used to rotate the image. Figure 12 The result of the image debiasing operation is shown. On the left is the Braille image, and on the right is the corrected Braille image. In terms of the processing efficiency of Braille images, an average of 2s can correct one Braille image.

[0119] This embodiment mainly focuses on three datasets, namely the Braille character dataset, the single-sided Braille dataset, and the double-sided Braille dataset. In order to verify the accuracy and rationality of the model, this embodiment divides these three datasets into two types of datasets, large and small, according to the number of pictures, size (memory size occupied by the image), category, and language. Table 2 shows the statistical information of the small dataset, and Table 3 shows the statistical information of the large dataset. In order to further improve the recognition accuracy of Braille images, ablation experiments are used in some experiments.

[0120] Table 2 Statistical Information of the Small Dataset

[0121]

[0122] Statistics of Information Related to the Big Data Set

[0123]

[0124]

[0125] It can be seen from Table 2 and Table 3 that the character braille data set and the single-sided braille data set have sufficient data volumes. However, for the double-sided data set, the number of pictures is significantly insufficient. Therefore, in this embodiment, the DBSI data set and the AngelinaDataset data set are combined with each other to make up for the number of training set images. Therefore, the data in the double-sided braille experiment is mixed data.

[0126] Result Display and Experimental Analysis of the Character Braille Data Set:

[0127] The CharsDataSet data set is in the form of "one Figure 1 label", that is, one braille picture corresponds to one label value. In this experiment, the models that have been popular in recent years are compared with the model proposed in this paper, including the ResNet hierarchical series, the MobileNetV2 model, and the GhostNet model. A comparison is also made with the traditional braille image recognition model LBP+SVM, laying a foundation for the subsequent hybrid innovation model. During the process of verifying the optimal model, in this embodiment, the Adam optimizer is used for training, and the accuracy rate is used as the evaluation criterion, where Acc represents the accuracy rate and Loss represents the loss value, indicating the degree of model convergence. Then, the results of the training set and the test set are respectively displayed. Among them, Table 4 shows the model running results on the small data set, and Table 5 shows the model running results on the large data set.

[0128] Table 4 Model Running Results on the Small Data Set

[0129]

[0130] Table 5 Model Running Results on the Large Data Set

[0131]

[0132]

[0133] Experimental result parameters are shown based on the CharDataSet size data set in Tables 4 and 5. First, the size of the data set affects the change in model accuracy. The accuracies of LBP+SVM, ResNet-50, ResNet-101, MobileNetV2, GhostNet, Vit, and PyramidNet on the large data set are increased by 1.02%, 0.69%, 0.71%, 9.9%, 2.53%, 6.48%, and 0.35% respectively compared to the small data set. Compared with the small data set, the large data set contains richer Braille dot information, is suitable for processing more complex models, and has strong model generalization ability.

[0134] Second, different models also affect the accuracy of character Braille recognition. Compared with the popular models in recent years, LBP+SVM, RestNet-50, RestNet-101, MobileNetV2, GhostNet, and Vit, PyramidNet is increased by 8.97%, 0.75%, 5.83%, 0.64%, 15.38%, and 7.33% respectively on the small data set, and by 8.30%, 0.41%, 0.28%, 5.83%, 7.19%, and 1.20% respectively on the large data set. From the difference in the result data of the models, it is concluded that the experimental accuracy of the PyramidNet model is the highest, while the experimental accuracy of the traditional Braille recognition model LBP+SVM is the lowest, followed by the GhostNet model. The experiment on this character Braille data set proves the effectiveness and rationality of the PyramidNet model for Braille recognition. In addition, for the character Braille data set, the choice of optimizer will affect the results of different models, and the accuracy error is within 15%.

[0135] Result Display and Experimental Analysis of Single-Sided Braille Data Set

[0136] Compared with the character Braille data set (CharDataSet), the single-sided Braille data set (SigleDataSet) has a larger number of positive Braille squares and occupies more memory. To further explore the change in the accuracy of the PyramidNet model and R-vit in the SigleDataSet data set, the GhostNet model with poor character recognition results is abandoned in this embodiment. Since the result accuracies of the RestNet-50 and RestNet-101 models are close, the more accurate and complex RestNet-101 model is selected in this embodiment. Based on the experimental results of the character Braille data set and compared with other popular models in recent years, MobileNetV2, Vit, and the traditional Braille recognition model, the experimental result parameters of the large and small data sets in Tables 6 and 7 are obtained:

[0137] Table 6 Model Running Results on Small Data Sets

[0138]

[0139] Table 7 Model running results on large datasets

[0140]

[0141] The experimental result parameters on the SigleDataSet-sized datasets in Table 6 and Table 7 are shown. First, the size of the dataset affects the change in model accuracy. The accuracies of LBP+SVM, ResNet-101, MobileNetV2, Vit, and PyramidNet on large datasets are increased by 1.15%, 2.00%, 1.92%, 0.02%, and 0.03% respectively compared to small datasets. Compared with small datasets, large datasets contain richer Braille dot information and Braille square information, are suitable for processing more complex models, and the model generalization ability is strong.

[0142] Second, different models also affect the accuracy of character Braille recognition. Compared with the popular models in recent years, LBP+SVM, RestNet-101, MobileNetV2, and Vit, PyramidNet is increased by 11.2%, 2.27%, 6.57%, and 3.79% respectively on small datasets, and by 10.08%, 0.30%, 4.68%, and 3.80% respectively on large datasets. From the difference in the result data of the models, it is concluded that the experimental accuracy of the R-vit model is the highest, followed by the PyramidNet model, while the experimental accuracy of the traditional Braille recognition model LBP+SVM is the lowest, followed by the MobileNetV2 model. The experiment on this single-sided Braille dataset proves the effectiveness and rationality of the Braille recognition model PyramidNet model.

[0143] To further check the authenticity of the experimental results, this embodiment uses Matplotlib to draw the recognition results after segmentation as Figure 13 shown.

[0144] Results display and experimental analysis of double-sided Braille dataset

[0145] Compared with the single-sided Braille dataset (SigleDataSet), the double-sided Braille dataset (DBSI) contains more Braille squares on both sides, occupies more memory, and has stronger publicity. In this paper, the accuracy of the model BraUNet studied at home and abroad in the DBSI dataset is compared with the accuracy of the model proposed in this paper, and only the Braille dots on the front of the Braille are recognized. The experimental result parameters in Tables 8 and 9 in the table are obtained:

[0146] Table 8 Model running results on small datasets

[0147]

[0148] Table 9 Model Running Results on Large Datasets

[0149]

[0150]

[0151] According to the experimental result parameters on the DBSI large and small datasets in Table 8 and Table 9. First, the size of the dataset affects the change in model accuracy. The accuracies of LBP+SVM, ResNet-101, Vit, BraUNet, and PyramidNet on the large dataset are 1.96%, 0.83%, 0.91%, 0.03%, and 0.53% higher than those on the small dataset respectively. Compared with the small dataset, the large dataset contains richer Braille dot information and Braille square information, is suitable for processing more complex models, and has strong model generalization ability.

[0152] Second, different models also affect the accuracy of character Braille recognition. Compared with the popular models LBP+SVM, RestNet-101, Vit, and BraUNet in recent years, PyramidNet has increased by 18.17%, 2.00%, 3.07%, and 0.31% on the small dataset respectively, and by 16.74%, 1.70%, 2.69%, and 0.81% on the large dataset respectively. From the difference in the result data of the models, it is concluded that the experimental accuracy of the PyramidNet model is the highest, while the experimental accuracy of the traditional Braille recognition model LBP+SVM is the lowest, followed by the Vit model. The experiment on this single-sided Braille dataset proves the effectiveness and rationality of the PyramidNet model for Braille recognition, while the traditional Braille recognition model LBP+SVM is difficult to recognize complex double-sided Braille pictures. The reason why the PyramidNet model cannot achieve 100% recognition accuracy in the test set is that the model cannot recognize relatively blurred or contaminated pictures. As Figure 14 shown.

[0153] To further check the authenticity of the experimental results, in this embodiment, Matplotlib is used to draw the recognition results after segmentation as Figure 15 shown.

[0154] Experimental Comparison of Braille Recognition Methods:

[0155] This embodiment proposes a blind braille image recognition based on object detection, and makes a comparison with the braille image recognition method based on image segmentation and the traditional braille recognition method. This experiment makes a comparison of these three methods in terms of model accuracy, the number of model parameters, the types of recognition tasks, and Dpi (the number of grid points printed on the braille image), and proves the effectiveness of the two-stage braille image classification recognition and the object detection-based braille recognition method. The comparative experiments are all carried out on the DBSI double-sided braille mixed dataset.

[0156] Table 10 Comparison of Braille Recognition Process Results

[0157]

[0158] Among them, the method based on traditional braille recognition LBP+SVM belongs to image classification, and its accuracy is the lowest. The BraUNet model represents the recent braille recognition method based on image segmentation, which belongs to the image segmentation task. The PyramidNet model represents the object detection-based image recognition, which belongs to the object detection task, and its recognition accuracy is the highest. From the results in Table 10, in terms of implementation complexity, the BraUNet model needs to label each braille image with expert data to obtain the label image, with a relatively high implementation complexity and cost. For the PyramidNet model, only txt or json is required for content annotation, with a relatively low implementation complexity, the highest recognition accuracy, and it basically meets the current braille image task recognition.

[0159] To make up for the shortage of the braille dataset and verify the rationality of the algorithm in this paper, this embodiment combines the characteristics of braille character information and creates two new datasets: the character braille dataset and the single-sided braille dataset by means of photographing, cropping, and website collection. At the same time, to increase the number of the double-sided braille dataset, this embodiment increases the number of the double-sided braille dataset by rotating, changing brightness, and adding non-braille photos. In particular, the AngelinaDataset part of the dataset is added to the DBSI dataset in the training set, laying a foundation for verifying the accuracy of the model algorithm.

[0160] In this embodiment, based on the characteristics of braille characters, a braille image recognition method based on the object detection algorithm is proposed. On this basis, the feature pyramid algorithm model PyramidNet is proposed, and the accuracy rates on the CharDataSet, SigleDataSet, and DBSI mixed data sets reach 98.42%, 98.37%, and 98.84%, respectively. This paper studies the relationship between the data volume and the model performance. The experimental results show that the increase in data improves the performance of the model proposed in this paper. This embodiment studies the comparison between the traditional braille image recognition method and the proposed recognition method. Under the conditions of model accuracy, model parameters, and implementation complexity, the experimental results show that the two-stage braille image classification and recognition method is the most efficient. This embodiment studies the ablation experiment of the Focalloss loss function. In the experiments where the PyramidNet model uses and does not use Facal loss respectively, the results show that PyramidNet combined with Focalloss achieves the best effect.

[0161] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for blind braille image recognition based on object detection, characterized in that, it includes: Constructing a braille dataset; Based on the gold pyramid feature fusion structure, constructing an object detection model; Based on the braille dataset, training the object detection model; Based on the trained object detection model, recognizing braille; The object detection model is constructed using a feature pyramid structure; The object detection model includes: an improved ResNet-50 module, an Upsample module, and a Class+Box Subnet module connected in sequence; Training the object detection model includes: Based on the ResNet-50 module, extracting features from the braille images in the braille dataset to obtain features at different levels; Based on the Upsample module, magnifying the features at different levels and performing feature fusion; Based on the Class+Box Subnet module, recognizing the image after feature fusion; The ResNet-50 module includes 4 stage sub-modules, namely: stage0, stage1, stage2, and stage3; among them, for the 4 stage modules, the first stage0 does not have a residual structure, while stage1 uses 3 residual structures, stage2 uses 4 residual structures, merging the Conv4 and Conv5 structures, and stage3 uses 9 residual structures; The Upsample module includes: several 1×1 convolutional kernels and an upsampling part; In the Upsample module, after the features pass through a 1×1 convolutional kernel, the size of the feature map is changed, and the feature map is sent to the upsampling part and added to the feature map of the previous layer. After the addition, multiple types of fused p corresponding to the layer are obtained. 2 -p 7 Final feature map; The Class+Box Subnet module includes: a box regression sub-network and a classification sub-network; Both the box regression sub-network and the classification sub-network include: several layers of 3×3 convolutional layers and the activation function ReLU; among them, the number of channel dimensions of the last convolutional layer in the box regression sub-network is a multiple of 64, and the number of channel dimensions of the last convolutional layer in the classification sub-network is a multiple of 4.

2. The method for blind braille image recognition based on object detection according to claim 1, characterized in that, The braille dataset includes: a character braille dataset, a single-sided braille dataset, and a double-sided braille dataset.

3. The method for blind braille image recognition based on object detection according to claim 1, characterized in that, After constructing the braille dataset, it includes: performing image preprocessing on the braille dataset; The image preprocessing includes: image binarization, brightness adjustment, image debiasing, and size adjustment.

4. The method for blind braille image recognition based on object detection according to claim 1, characterized in that, Recognizing the image after feature fusion includes: Based on the box regression sub-network, obtaining the box position of the image after feature fusion; Based on the classification sub-network, obtaining the category information of the image after feature fusion.