Optical character recognition method based on multi-view assistance

Through the multi-view angle-assisted optical character recognition method, principal component analysis and DBSCAN clustering screen character point clouds are used, and multi-view angle fusion and dynamic parameter optimization are combined to generate efficient three-dimensional character line point clouds, solving the optical character recognition problem in complex scenarios and achieving high-precision and robust recognition effects.

CN120340039APending Publication Date: 2025-07-18XIAN HULUMIBAO DATA TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510390687.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing optical character recognition technology is difficult to effectively utilize three-dimensional point cloud data in complex scenarios, which has problems such as difficulty in dimensionality reduction, noise interference, parameter selection inappropriateness and insufficient fusion of multi-view angles, resulting in low recognition accuracy.

Method used

A multi-view assisted method is used to filter character point clouds through principal component analysis, Hough transformation and DBSCAN clustering, combining multi-view fusion and dynamic parameter optimization to generate efficient three-dimensional character line point clouds, driving the deep learning OCR engine for recognition.

Benefits of technology

It significantly improves the accuracy of optical character recognition in complex environments, enhances noise suppression ability and parameter adaptability, and improves the robustness and accuracy of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340039A_ABST
    Figure CN120340039A_ABST
Patent Text Reader

Abstract

The invention discloses an optical character recognition method based on multi-view assistance, and the method comprises the steps: obtaining three-dimensional point cloud data through a laser radar, carrying out the dimension reduction through principal component analysis, and removing background points through elevation filtering. The point cloud is projected to the aerial view, noise points are screened out through Hough transform, and then the number of vertical character lines is calculated through DBSCAN clustering. And reconstructing character lines coarsely and finely in sequence to generate a complete three-dimensional point cloud for character recognition. Hough transform and DBSCAN parameters can be dynamically adjusted according to the point cloud density in the process. And a multi-view fusion verification step is also provided, and if the reconstruction result does not meet the requirement, the coarse reconstruction parameters are returned and adjusted. Finally, the generated three-dimensional character line point cloud drives an optical character recognition engine based on deep learning, and character segmentation and recognition are assisted by using three-dimensional structure information. According to the method, through multi-view fusion, three-dimensional reconstruction and intelligent parameter optimization, the difficulty of character recognition in a complex environment is systematically solved, and a high-precision and high-robustness solution is provided for the fields of document scanning, automatic driving and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to an optical character recognition method based on multi-view assistance. Background Art

[0002] In modern optical character recognition (OCR) technology, character recognition in complex scenarios remains a challenging problem. Traditional two-dimensional image-based OCR methods often suffer from a decrease in segmentation and recognition accuracy due to the lack of depth information when facing illumination changes, low contrast, character deformation, or occlusion. With the development of three-dimensional perception technology, using multi-view assisted three-dimensional point cloud data has become an important direction to improve the robustness of OCR. However, existing methods still have limitations in processing three-dimensional point clouds: on the one hand, the three-dimensional point cloud data volume is huge and the structure is complex, and traditional dimensionality reduction and feature extraction algorithms are difficult to efficiently retain the key information of characters; on the other hand, the character line reconstruction process is easily interfered by noise, resulting in the subsequent recognition model being unable to accurately analyze the character morphology.

[0003] Current mainstream three-dimensional OCR methods mainly rely on single-view point cloud analysis and are difficult to comprehensively capture the spatial structure characteristics of characters. For example, the processing method based on bird's-eye view projection simplifies the computational complexity but ignores the elevation differences in the vertical direction, resulting in the loss of character line continuity or misjudgment. In addition, in the existing technology, parameter selection in the noise point screening and character line reconstruction links depends on empirical settings and lacks an adaptive adjustment mechanism, making it difficult to adapt to point cloud data with different densities and distributions. For example, if the distance threshold of the Hough transform and the neighborhood radius of DBSCAN clustering are not dynamically optimized according to the scene, it will directly affect the screening accuracy of character point clouds, thereby reducing the reliability of the recognition results.

[0004] The lack of a multi-view fusion verification mechanism is also a significant shortcoming of the existing technology. Traditional methods only verify the reconstruction results through a single view and cannot effectively detect the breakage or redundancy problems of character lines in different directions. At the same time, the approximate processing of the curve shape (such as using broken lines to replace curves) and the mean filling strategy of elevation information during the character line reconstruction process may introduce structural distortion, further affecting the performance of subsequent deep learning-based recognition models. In addition, the existing technology lacks the ability to dynamically respond to changes in point cloud density, resulting in computational redundancy in high-density point cloud scenarios and possible loss of character features due to insufficient information in sparse point cloud scenarios. These problems restrict the generalization ability of the OCR system in complex environments. Summary of the Invention

[0005] In order to overcome the disadvantages and deficiencies of the existing technology, the present invention provides an optical character recognition method based on multi-view assistance.

[0006] An optical character recognition method based on multi-view assistance includes the following steps:

[0007] Step S1, Background point removal: Use lidar to obtain three-dimensional point cloud data containing character regions. Perform dimensionality reduction on the original point cloud through the principal component analysis algorithm. Extract the principal component directions based on the singular value decomposition of the covariance matrix. Apply the elevation filtering technique, set the minimum and maximum values of the elevation, and filter out the background point clouds that are not within this elevation range;

[0008] Step S2, Noise point screening: Project the processed point cloud onto a bird's-eye view. Use the Hough transform to detect line features, calculate the distance from each point to the detected lines, set a distance threshold, and screen out the points whose distance to the line is less than this threshold. These points are the point clouds belonging to the characters;

[0009] Step S3, Calculation of the number of character lines in the vertical direction: In the bird's-eye view, apply the DBSCAN clustering algorithm to the screened point cloud. Determine the number of character lines in the vertical direction based on the aggregation characteristics of the elevation distribution of the point cloud;

[0010] Step S4, Coarse reconstruction of character lines: Apply the DBSCAN clustering again to the point cloud of a single character line, divide the blank areas therein, generate filling points by evenly dividing the coordinate differences of the two end points of the reconstruction area, and optimize the coordinates of the reconstruction points in combination with the local centroid calculation method;

[0011] Step S5, Fine reconstruction of character lines: In the bird's-eye view, divide the character lines after coarse reconstruction into multiple evenly distributed points. For each point, take the average value of the elevations of several points closest to it in the character line where it is located as the elevation coordinate of this point, and generate the three-dimensional point cloud of the complete character line for subsequent optical character recognition.

[0012] Furthermore, in the step S1, calculate the covariance matrix of the original point cloud. This matrix is obtained by operating on the difference between the coordinate vector of each point and the mean vector of the point cloud. Perform singular value decomposition on the covariance matrix to obtain the eigenvalues and the corresponding eigenvectors. Since the eigenvalues have different magnitudes, select the eigenvectors corresponding to the two larger eigenvalues. The directions represented by these two eigenvectors are the principal component directions. Perform dimensionality reduction on the original point cloud based on these two principal component directions to retain the main structural information of the point cloud and reduce the data volume for subsequent processing.

[0013] Further, in the step S2, the point cloud data in the bird's-eye view is converted into a binary image, the contour in the image is extracted by using an edge detection algorithm, then the Hough transform is performed on the contour, the parameters of the straight line are mapped into the accumulator space, the significant straight lines with higher vote counts are identified by counting the votes in the accumulator space, for each detected straight line, the distances from all points in the point cloud to this straight line are calculated, a distance threshold is set, and the points with distances less than or equal to the threshold are retained, and these points form the character point cloud.

[0014] Further, in the step S3, the value range of the neighborhood radius is set between 0.1 meter and 0.5 meter, and this value range will be dynamically adjusted according to the density of the point cloud. When the point cloud density is large, the neighborhood radius is appropriately increased; when the point cloud density is small, the neighborhood radius is decreased. The minimum number of points is set between 3 and 5, and this parameter is used to ensure the effectiveness of clustering. In the clustering result, each clustering cluster corresponds to a character line in the vertical direction, and different character lines are distinguished by analyzing the elevation range difference of the point cloud.

[0015] Further, in the step S4, the two end points of the reconstruction area are determined, the coordinate differences of these two end points in the horizontal direction and the vertical direction are calculated, the coordinate difference in the horizontal direction is divided by the preset number of points to obtain the spacing between adjacent filling points in the horizontal direction; the spacing between adjacent filling points in the vertical direction is calculated, and according to these two spacings, a plurality of filling points are sequentially incremented in the horizontal and vertical directions to form a broken line with the filling points, approximately replacing the character line that may originally be a curve, and the preliminary reconstruction of the character line is completed.

[0016] Further, in the step S4, for each point, its weight is determined according to its distance to the reconstruction straight line. The closer the distance is, the greater the weight is. The coordinates of all points are weighted and summed, and the weighted sum is divided by the sum of the weights of all points respectively, and the result obtained is the coordinate of the reconstruction point.

[0017] Further, in the step S5, in the bird's-eye view, the roughly reconstructed character line is divided into multiple equally spaced points according to the preset rules. For each divided point, several points closest to it are searched in the three-dimensional point cloud, the average elevation of these closest points is calculated, and this average value is used as the elevation value of this divided point.

[0018] Further, the method further includes step S6, a multi-view fusion verification step: projecting the reconstructed three-dimensional character line point cloud onto multiple different views such as the front view and the side view, binarizing the images in each view so that the images have only two colors, black and white, performing morphological processing to remove small noises and holes in the images, and judging whether the reconstruction result meets the preset requirements by observing the continuity and integrity of the character lines in each view. If it is found that there are breaks or redundant parts in the character lines, return to step S4, readjust the parameters of the rough reconstruction, and perform the rough reconstruction and subsequent processing of the character lines again until the multi-view verification result meets the preset requirements.

[0019] Further, the method further includes step S7, a dynamic parameter optimization mechanism: automatically adjusting the distance threshold in the Hough transform and the neighborhood radius and minimum number of points in the DBSCAN clustering algorithm according to the density of the point cloud. When the point cloud density is large, appropriately increase the distance threshold of the Hough transform and the neighborhood radius of DBSCAN; when the point cloud density is small, decrease the distance threshold of the Hough transform and the neighborhood radius of DBSCAN. Use a machine learning model to predict the most suitable parameter combination in the current scenario based on historical data and the point cloud characteristics in different scenarios.

[0020] Further, the generated three-dimensional character line point cloud is used to drive an optical character recognition engine. The engine uses a recognition model based on deep learning. When performing character recognition, it uses the three-dimensional structure information provided by the three-dimensional character line point cloud to assist in character segmentation and recognition.

[0021] Beneficial effects:

[0022] The present invention proposes an optical character recognition method based on multi-view assistance. This method significantly improves the OCR recognition performance in complex scenarios through multi-view fusion and three-dimensional point cloud processing technologies.

[0023] Specific advantages are as follows:

[0024] 1. Robust noise suppression ability: Using principal component analysis (PCA) and elevation filtering to remove background point clouds, and combining the Hough transform and DBSCAN clustering to screen character points, effectively filtering environmental interferences (such as occlusion, blur), ensuring the integrity and accuracy of the character line point cloud.

[0025] 2. Efficient character line reconstruction: Through two-step optimization of rough reconstruction (coordinate difference equal division method) and fine reconstruction (neighborhood elevation mean filling), the continuity and high density of the character lines are achieved, especially suitable for scenarios of curved text, low contrast, or partially missing characters.

[0026] 3. Multi - perspective verification mechanism: Project the 3D point cloud onto multiple perspectives such as the front view and side view for morphological verification, ensuring the consistency of character lines in different directions, avoiding the risk of misjudgment in a single perspective, and improving the reliability of the reconstruction results.

[0027] 4. Dynamic parameter adaptability: Based on the point cloud density, adaptively adjust the Hough transform threshold and DBSCAN parameters, and combine machine learning to predict the optimal parameter combination, enabling the method to flexibly adapt to different scenarios (such as dense characters, sparse point clouds) and enhancing robustness.

[0028] 5. 3D information - enhanced recognition: Input the reconstructed 3D point cloud into a deep - learning OCR engine (such as CRNN), and use depth information to optimize character segmentation and feature extraction, breaking through the limitations of traditional 2D OCR in scenarios such as overlapping characters and tilted text, and significantly improving the recognition accuracy.

[0029] In summary, through multi - perspective fusion, 3D reconstruction, and intelligent parameter optimization, this method systematically solves the difficulties of character recognition in complex environments, providing a high - precision and high - robustness solution for fields such as document scanning and autonomous driving. Brief Description of the Drawings

[0030] Figure 1 It is a flowchart of the method steps of the present invention. Detailed Embodiments

[0031] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following further describes this application in detail with reference to the drawings and specific embodiments.

[0032] As Figure 1 shown, an optical character recognition method based on multi - perspective assistance, the method includes:

[0033] Step S1, background point removal: Use lidar to obtain 3D point cloud data containing the character area, perform dimensionality reduction processing on the original point cloud through the principal component analysis algorithm, extract the principal component direction based on the singular value decomposition of the covariance matrix, and apply elevation filtering technology to set the minimum and maximum values of elevation, filtering out the background point cloud that is not within this elevation range.

[0034] Specifically, a lidar is used to obtain three-dimensional point cloud data containing character regions. The original point cloud data has a high dimension, which is not conducive to subsequent processing. It is dimensionally reduced through the principal component analysis algorithm (PCA). Specifically, the covariance matrix of the original point cloud is calculated, which is obtained by operating on the difference between the coordinate vector of each point and the mean vector of the point cloud. The covariance matrix is subjected to singular value decomposition to obtain eigenvalues and corresponding eigenvectors. Since the eigenvalues have different magnitudes, the two eigenvectors corresponding to the larger two eigenvalues are selected. The directions represented by these two eigenvectors are the principal component directions. Based on these two principal component directions, the original point cloud is dimensionally reduced, and the main structural information of the point cloud can be retained. Then, elevation filtering technology is used to set the minimum and maximum elevation values, and the background point cloud outside this elevation range is filtered out.

[0035] In the process of calculating the covariance matrix and performing singular value decomposition, it mainly involves the coordinate information of the point cloud. The minimum and maximum elevation values during elevation filtering need to be set according to the approximate elevation range where the characters are located in the actual scene. For example, in some road sign recognition scenarios, if the height of the sign from the ground is approximately between 2 - 5 meters, the minimum elevation value can be initially set to 1.5 meters and the maximum elevation value to 5.5 meters, and then fine-tuned according to the actual effect.

[0036] Dimensionality reduction processing can greatly reduce the amount of data for subsequent processing and improve processing efficiency. By removing the background point cloud through elevation filtering, a large amount of noise data irrelevant to the characters can be effectively excluded, enabling subsequent processing to focus on the point cloud related to the characters, and improving the accuracy and reliability of character recognition. For example, in a complex urban environment, without removing the background points, a large amount of background point clouds such as buildings and trees will interfere with character recognition, while after this step of processing, these interference factors can be greatly reduced.

[0037] Step S2, noise point screening: Project the processed point cloud onto a bird's-eye view perspective, use the Hough transform to detect line features, calculate the distance from each point to the detected line, set a distance threshold, and screen out the points whose distance to the line is less than this threshold. These points are the point cloud belonging to the characters.

[0038] Specifically, project the point cloud after background point removal onto a bird's-eye view perspective, convert the point cloud data in the bird's-eye view perspective into a binary image, use an edge detection algorithm to extract the contours in the image, and then perform the Hough transform on the contours. The Hough transform maps the parameters of the line to the accumulator space, and by counting the votes in the accumulator space, the significant lines with higher votes are identified. For each detected line, calculate the distance from all points in the point cloud to this line, set a distance threshold, and retain the points whose distance is less than or equal to this threshold. These points constitute the character point cloud.

[0039] In the Hough transform, the setting of the distance threshold is crucial. For example, for some clear characters with regular strokes, the distance threshold can be set to 0.05 meters; if the characters may have certain deformations or the quality of the point cloud data is slightly poor, the threshold can be appropriately increased to 0.1 meters. Classic algorithms such as Canny can be selected for the edge detection algorithm, and its parameters such as the high-to-low threshold ratio will also affect the edge detection effect. Generally, the high threshold can be set to about 3 times the low threshold.

[0040] Detecting line features through the Hough transform and screening the point cloud according to the distance can effectively separate the character point cloud from the complex point cloud data. Compared with directly performing character recognition on the original point cloud, after removing noise points, the character point cloud is purer, reducing the interference of irrelevant points on character recognition, thereby improving the recognition accuracy. For example, when recognizing license plate characters, the point cloud of parts such as the surrounding vehicle body may interfere with the recognition, and this step can effectively remove these interfering points.

[0041] Step S3: Calculating the number of character lines in the vertical direction: From the bird's-eye view perspective, apply the DBSCAN clustering algorithm to the screened point cloud, and determine the number of character lines in the vertical direction based on the aggregation characteristics of the elevation distribution of the point cloud.

[0042] Specifically, from the bird's-eye view perspective, apply the DBSCAN clustering algorithm to the screened point cloud. The DBSCAN algorithm clusters based on the density connectivity of the point cloud. The neighborhood radius and the minimum number of points are important parameters of the DBSCAN algorithm. The value range of the neighborhood radius is set between 0.1 meters and 0.5 meters, and this value range will be dynamically adjusted according to the density of the point cloud. When the point cloud density is large, the neighborhood radius is appropriately increased; when the point cloud density is small, the neighborhood radius is decreased. The minimum number of points is set between 3 and 5, and this parameter is used to ensure the effectiveness of clustering. In the clustering result, each cluster corresponds to a character line in the vertical direction, and different character lines are distinguished by analyzing the elevation range difference of the point cloud.

[0043] The neighborhood radius and the minimum number of points need to be dynamically adjusted according to the point cloud density. For example, when the point cloud density is large, such as in the scenario of high-precision character point cloud obtained at close range, the neighborhood radius can be set to 0.3 meters and the minimum number of points can be set to 5; in the scenario of low-density point cloud obtained at a long distance, the neighborhood radius is set to 0.15 meters and the minimum number of points is set to 3.

[0044] Accurately calculating the number of character lines in the vertical direction is crucial for subsequent character recognition. Through the DBSCAN clustering algorithm, the character lines can be automatically recognized according to the distribution characteristics of the point cloud. Compared with manual annotation or other simple segmentation methods, it has higher automation and accuracy. This provides a basis for subsequent processing of each single character line, ensuring that subsequent processing can be carried out for each character line separately, improving the pertinence and effectiveness of the processing.

[0045] Step S4, Character line rough reconstruction: Apply DBSCAN clustering to the point cloud of a single character line again to divide the blank areas therein. Generate filling points by evenly dividing the coordinate differences between the two endpoints of the reconstruction area, and optimize the coordinates of the reconstruction points in combination with the local centroid calculation method.

[0046] Apply DBSCAN clustering to the point cloud of a single character line again to divide the blank areas therein. Determine the two endpoints of the reconstruction area, calculate the coordinate differences of these two endpoints in the horizontal and vertical directions, divide the coordinate difference in the horizontal direction by the preset number of points to obtain the spacing between adjacent filling points in the horizontal direction; calculate the spacing between adjacent filling points in the vertical direction, and generate multiple filling points in sequence in the horizontal and vertical directions according to these two spacings. Use the filling points to form a broken line to approximately replace the character line that may originally be a curve, and complete the preliminary reconstruction of the character line. For each point, determine its weight according to the distance from the point to the reconstruction line. The closer the distance, the greater the weight. Perform a weighted sum of the coordinates of all points, and divide the weighted sum by the sum of the weights of all points respectively. The result obtained is the coordinate of the reconstruction point, and the coordinates of the reconstruction points are optimized in this way.

[0047] The preset number of points will affect the refinement degree of the reconstruction. Generally, it can be set according to the length of the character line and the expected reconstruction accuracy. For example, for a character line with a length of 0.5 meters, if a high-precision reconstruction is expected, the number of points can be set to 50. When calculating the weight, the weight calculation function can adopt the form of inverse square of the distance, that is, weight w = 1 / d 2 where d is the distance from the point to the reconstruction line.

[0048] Character line rough reconstruction can perform preliminary repair and supplementation on the character line, making the character line more complete. By dividing the blank area and generating filling points, part of the character line information missing due to noise or other reasons can be restored. Optimizing the coordinates of the reconstruction points can make the reconstructed character line more conform to the shape of the actual character, providing a better basis for subsequent fine reconstruction and character recognition. For example, in some character recognition scenarios with wear or partial occlusion, rough reconstruction can effectively restore the general shape of the character.

[0049] Step S5, Character line fine reconstruction: From the bird's-eye view perspective, divide the character line after rough reconstruction into multiple evenly distributed points. For each point, take the average value of the elevations of several points closest to it in the character line where it is located as the elevation coordinate of this point, and generate a complete three-dimensional point cloud of the character line for subsequent optical character recognition.

[0050] Specifically, from the bird's-eye view perspective, the character lines after rough reconstruction are divided into multiple equally spaced points according to a preset rule. For each divided point, several points closest to it are searched in the 3D point cloud, and the average elevation value of these closest points is calculated. This average value is used as the elevation value of the divided point, generating a complete 3D point cloud of the character lines for subsequent optical character recognition.

[0051] The equal-spacing interval of the divided points needs to be set according to the length of the character line and the desired fine reconstruction accuracy. For example, for a character line with a length of 1 meter, if a high accuracy is desired, the interval can be set to 0.01 meter. The number of closest points searched also affects the fine reconstruction effect, and generally can be set to 5 - 10, and adjusted according to the density and distribution of the point cloud.

[0052] The fine reconstruction further refines the 3D information of the character lines. By obtaining the accurate elevation value of each point, the generated 3D point cloud of the character lines more accurately reflects the actual shape of the characters. This is of great significance for improving the accuracy of optical character recognition, because more accurate character shape information can help the recognition model better perform character segmentation and recognition, especially for some scenarios with high requirements for character shape details, such as high-precision document recognition, etc.

[0053] Step S6: Multi-view fusion verification. The reconstructed 3D character line point cloud is projected onto multiple different views such as the front view and the side view respectively. The images in each view are binarized so that the images have only two colors, black and white. Morphological processing is performed to remove small noises and holes in the images. By observing the continuity and integrity of the character lines in each view, it is judged whether the reconstruction result meets the preset requirements. If it is found that there are breaks or redundant parts in the character lines, return to step S4 to re-adjust the parameters of the rough reconstruction, and perform the rough reconstruction and subsequent processing of the character lines again until the multi-view verification result meets the preset requirements.

[0054] Specifically, in the binarization process, the setting of the threshold affects the binarization effect, and generally automatic threshold algorithms such as Otsu can be used. The size of the structuring element for the erosion and dilation operations in morphological processing needs to be set according to factors such as the width of the character line. For example, for a character line with a width of 0.05 meter, the size of the structuring element can be set to 0.01 meter. The preset requirements can be determined according to the actual application scenario, such as the break length of the character line cannot exceed 5% of the total length, and the area of the redundant part cannot exceed 10% of the area of the character line, etc.

[0055] Multi - perspective fusion verification can comprehensively check the reconstruction results of character lines from multiple perspectives. Compared with single - perspective verification, it greatly improves the accuracy and reliability of verification. By morphological processing, noise and holes are removed, making the verification results more reliable. When it is found that the reconstruction results do not meet the requirements, adjusting the parameters in a timely manner can continuously optimize the reconstruction results and improve the success rate of final character recognition. For example, in some complex industrial identification scenarios, multi - perspective verification can effectively avoid recognition errors caused by reconstruction defects that cannot be detected from a single perspective.

[0056] Step S7: Dynamic parameter optimization mechanism. According to the density of the point cloud, automatically adjust the distance threshold in the Hough transform and the neighborhood radius and minimum number of points in the DBSCAN clustering algorithm. When the point cloud density is large, appropriately increase the distance threshold of the Hough transform and the neighborhood radius of DBSCAN; when the point cloud density is small, decrease the distance threshold of the Hough transform and the neighborhood radius of DBSCAN. Use a machine - learning model to predict the most suitable parameter combination in the current scenario based on historical data and the point - cloud characteristics in different scenarios.

[0057] Specifically, the machine - learning model can be a decision tree, neural network, etc. Its training data needs to include the optimal parameter combinations under different - density point clouds and the corresponding point - cloud feature information. For example, for the point - cloud density feature, the number of point clouds per unit volume can be used to represent it. The adjustment range of the Hough - transform distance threshold and DBSCAN parameters needs to be determined according to the actual effect, generally increasing or decreasing by 20% - 50% based on the original parameters.

[0058] The dynamic parameter optimization mechanism enables the algorithm to automatically adjust parameters according to the characteristics of different point - cloud data, improving the adaptability and robustness of the algorithm. Compared with fixed - parameter settings, it can maintain good performance in various complex actual scenarios. By predicting the parameter combination through a machine - learning model, it can continuously learn the rules in historical data, further optimizing the accuracy of parameter adjustment, thus improving the overall effect of character recognition, especially having significant advantages in scenarios with large differences in different environments and data quality.

[0059] The generated three - dimensional character - line point cloud is used to drive an optical character recognition engine. This engine uses a deep - learning - based recognition model. When performing character recognition, it uses the three - dimensional structure information provided by the three - dimensional character - line point cloud to assist in character segmentation and recognition. Compared with traditional two - dimensional - image - based character - recognition methods, this method uses three - dimensional point - cloud information, which can more comprehensively reflect the shape and position information of characters, improving the accuracy of character recognition, especially having obvious advantages in the recognition of characters with complex spatial structures or in three - dimensional scenarios.

[0060] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various equivalent changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalent scope.

Claims

1. An optical character recognition method based on multi-view assistance, characterized in that, Including the following steps: Step S1, background point removal: Use a lidar to obtain three-dimensional point cloud data containing the character area. Perform dimensionality reduction on the original point cloud through the principal component analysis algorithm. Extract the principal component directions based on the singular value decomposition of the covariance matrix. Apply the elevation filtering technique, set the minimum and maximum values of the elevation, and filter out the background point clouds not within this elevation range; Step S2, noise point screening: Project the processed point cloud onto the bird's-eye view perspective. Use the Hough transform to detect line features. Calculate the distance from each point to the detected line. Set a distance threshold and screen out the points whose distance to the line is less than this threshold. These points are the point clouds belonging to the characters; Step S3, calculation of the number of character lines in the vertical direction: Under the bird's-eye view perspective, apply the DBSCAN clustering algorithm to the screened point cloud. Determine the number of character lines in the vertical direction based on the aggregation characteristics of the elevation distribution of the point cloud; Step S4, rough reconstruction of the character line: Use the DBSCAN clustering again for the point cloud of a single character line, divide the blank areas therein, generate filling points by evenly dividing the coordinate differences of the two endpoints of the reconstruction area, and optimize the coordinates of the reconstruction points in combination with the local centroid calculation method; Step S5, fine reconstruction of the character line: Under the bird's-eye view perspective, divide the roughly reconstructed character line into multiple uniform points. For each point, take the average value of the elevations of several points closest to it in the character line where it is located as the elevation coordinate of this point, and generate the three-dimensional point cloud of the complete character line for subsequent optical character recognition.

2. The optical character recognition method based on multi-view assistance according to claim 1, characterized in that, In the said step S1, calculate the covariance matrix of the original point cloud. This matrix is obtained by operating on the difference between the coordinate vector of each point and the mean vector of the point cloud. Perform singular value decomposition on the covariance matrix to obtain eigenvalues and corresponding eigenvectors. Since the eigenvalues have different magnitudes, select the eigenvectors corresponding to the two larger eigenvalues. The directions represented by these two eigenvectors are the principal component directions. Perform dimensionality reduction on the original point cloud based on these two principal component directions to retain the main structural information of the point cloud and reduce the data volume for subsequent processing.

3. The optical character recognition method based on multi-view assistance according to claim 1, wherein In the said step S2, convert the point cloud data in the bird's-eye view perspective into a binary image. Use an edge detection algorithm to extract the contours in the image. Then perform the Hough transform on the contours, map the parameters of the lines to the accumulator space, identify the significant lines with higher vote counts by counting the votes in the accumulator space. For each detected line, calculate the distance from all points in the point cloud to this line. Set a distance threshold and retain the points whose distance is less than or equal to this threshold. These points constitute the character point cloud.

4. The optical character recognition method based on multi-view assistance according to claim 1, wherein In the said step S3, the value range of the neighborhood radius is set between 0.1 meter and 0.5 meter. This value range will be dynamically adjusted according to the density of the point cloud. When the point cloud density is large, appropriately increase the neighborhood radius; when the point cloud density is small, decrease the neighborhood radius. The minimum number of points is set between 3 and 5. This parameter is used to ensure the effectiveness of clustering. In the clustering result, each cluster corresponds to a character line in the vertical direction. Distinguish different character lines by analyzing the elevation range differences of the point cloud.

5. The optical character recognition method based on multi-perspective assistance according to claim 1, characterized in that In step S4, determine the two endpoints of the reconstruction region, calculate the coordinate differences of these two endpoints in the horizontal and vertical directions, divide the coordinate difference in the horizontal direction by the preset number of points to obtain the spacing between adjacent filling points in the horizontal direction; Calculate the spacing between adjacent filling points in the vertical direction. According to these two spacings, generate multiple filling points in sequence in the horizontal and vertical directions, and use the filling points to form a polyline to approximately replace the character line that may originally be a curve, thus completing the preliminary reconstruction of the character line.

6. The optical character recognition method based on multi - perspective assistance according to claim 1, wherein, In step S4, for each point, determine its weight according to its distance to the reconstruction line. The closer the distance, the greater the weight. Perform a weighted sum of the coordinates of all points, and divide the weighted sum by the sum of the weights of all points. The result obtained is the coordinate of the reconstruction point.

7. An optical character recognition method based on multi-perspective assistance according to claim 1, characterized in that In step S5, from the bird's-eye view perspective, divide the roughly reconstructed character line into multiple equally spaced points according to the preset rules. For each divided point, search for several points closest to it in the 3D point cloud, calculate the average elevation of these closest points, and use this average value as the elevation value of this divided point.

8. The optical character recognition method based on multi-view assistance according to claim 1, wherein The method further includes step S6, the multi-view fusion verification step: project the reconstructed 3D character line point cloud onto multiple different views such as the front view and the side view respectively, perform binarization processing on the image in each view so that the image has only two colors, black and white, perform morphological processing to remove small noises and holes in the image, and judge whether the reconstruction result meets the preset requirements by observing the continuity and integrity of the character line in each view. If it is found that the character line has breaks or redundant parts, return to step S4, readjust the parameters of the rough reconstruction, and perform the rough reconstruction and subsequent processing of the character line again until the multi-view verification result meets the preset requirements.

9. A method for optical character recognition based on multi - perspective assistance according to claim 1, characterized in that, The method further includes step S7, the dynamic parameter optimization mechanism: automatically adjust the distance threshold in the Hough transform and the neighborhood radius and minimum number of points in the DBSCAN clustering algorithm according to the density of the point cloud. When the point cloud density is large, appropriately increase the distance threshold of the Hough transform and the neighborhood radius of DBSCAN; when the point cloud density is small, decrease the distance threshold of the Hough transform and the neighborhood radius of DBSCAN. Use a machine learning model to predict the most suitable parameter combination in the current scenario based on historical data and the point cloud characteristics in different scenarios.

10. An optical character recognition method based on multi-view assistance according to claim 1, characterized in that, The generated 3D character line point cloud is used to drive an optical character recognition engine. The engine uses a recognition model based on deep learning. When performing character recognition, it uses the 3D structural information provided by the 3D character line point cloud to assist in character segmentation and recognition.