A Method for Identifying Compass Scales of an Optoelectronic Azimuth Instrument Based on Computer Vision
Through the computer vision-based photoelectric azimuth compass scale recognition method, combined with deep learning algorithms to automatically identify the scale readings of ship compass, the problem of traditional mechanical compass requiring manual readings is solved, and fast and accurate azimuth indicator monitoring is achieved.
Patent Information
- Application Number
- CN202310695855.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-13
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-06-13
AI Technical Summary
The traditional mechanical compass equipped with existing ships requires manual reading, which is inconvenient to operate and cannot quickly obtain orientation results.
The compass scale recognition method based on computer vision is adopted to acquire dial images through a short-focus optical lens and a camera, and image preprocessing, polar coordinate transformation, text segmentation and recognition are combined with deep learning algorithms to automatically obtain digital readings of the compass scale.
It realizes fast and accurate identification of the compass scale of the photoelectric azimuth compass, monitors the ship's azimuth indicators in real time, and improves operational efficiency and recognition accuracy.
Smart Images

Figure CN116682103B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence technology and optical character recognition technology, and particularly relates to a method for recognizing the compass scale of a photoelectric azimuth instrument based on computer vision. Background Art
[0002] A photoelectric azimuth instrument is an angle measuring device based on optical and electronic technologies. It can measure the azimuth information of a target and output an angle value. This device is mainly applied in the military field, such as unmanned aerial vehicles, ships, etc.
[0003] However, the compasses equipped on existing ships often still have a traditional mechanical dial with 360-degree scales surrounding it. It rotates to the corresponding angle according to the bow position of the ship in navigation, and the specific reading of the compass is mainly manually completed by the shipboard operators, which is inconvenient to operate and unable to quickly obtain the azimuth result.
[0004] Compass image scale recognition is a commonly used method, which can calculate the deflection angle by the position of the pointer in the compass image, so as to automatically obtain the azimuth information of the target. In addition, this device can also realize night low-light target detection and tracking. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for recognizing the compass scale of a photoelectric azimuth instrument based on computer vision to solve the problems in the background art.
[0006] To solve the above technical problems, the present invention provides a method for recognizing the compass scale of a photoelectric azimuth instrument based on computer vision, including:
[0007] Using a short-focus optical lens and a camera to collect and capture a local image near the dial pointer, with the center of the field of view aligned with the pointer, and the dial rotating 360 degrees while the pointer remains stationary;
[0008] For images under different ambient light conditions, evaluating the average brightness and variance information of the images to determine whether the current image illumination is appropriate, and completing the preprocessing of the images;
[0009] For the numbers and scales surrounding the circular dial area, performing a polar coordinate transformation on the entire dial area to re-transform the arc-shaped text area into a horizontal arrangement;
[0010] According to the illumination result of the image evaluation in the above preprocessing, performing binary segmentation on the image using different thresholds to roughly segment the dial text area;
[0011] Using the morphological dilation and erosion algorithms to denoise and connect the binary image of the text after rough segmentation, removing the interference of dust around the text and the upper and lower arc-shaped line segments, and at the same time connecting multiple numbers together;
[0012] Find the connected regions for the denoised and connected text regions, sort the multiple connected regions in descending order of pixel area, find the minimum bounding rectangle of the region with the largest area, and segment the text image from the original image according to the coordinate information;
[0013] Scale the segmented text image proportionally to a height of 32 pixels and pad with a fixed pixel value on the right until the image width is 160 pixels; input it into the deep learning-based text recognition network CRNN, extract features through the convolutional neural network and predict the probability of each digit through the recurrent neural network to obtain the digital readings of the dial scale near the pointer;
[0014] Perform a preliminary rule check on the recognized reading results, including whether it is all digits, whether the number of digits is less than 3, and whether the number range is compliant; if all the check results are satisfied, consider the current recognition result correct; if any one of the checks is not satisfied, select the region with the second largest area and re-recognize;
[0015] Take the correct text recognition result and calculate the angle between the center of the minimum bounding rectangle of the corresponding text region of the correct result and the center of the pointer, and convert it to the accurate scale reading corresponding to the current pointer;
[0016] Use the comprehensive determination of multiple frame recognition results and report them to the control end at a frequency of once per second; the rotation of the photoelectric azimuth instrument panel angle is a slow and continuous process. If the difference between the recognition results of two adjacent frames within one second is greater than 5, it is considered that the current recognition is incorrect and the result of the previous second is maintained; otherwise, the result is updated normally.
[0017] In one implementation, the preprocessing of the image includes:
[0018] Convert the input color picture to a grayscale image, traverse each pixel of the grayscale image, use the grayscale value of each pixel point as the abscissa and the frequency of the grayscale value appearing on the whole picture as the ordinate to calculate the grayscale histogram among 256 pixel points;
[0019] Calculate the average value and standard deviation of the grayscale histogram, and determine whether it is too dark or too bright by the ratio of the average value and the standard deviation; if the picture is too bright, set the threshold according to the size of the average value and return 1; if the picture is too dark, also set the threshold according to the size of the average value and return -1; if the picture is neither too bright nor too dark, return 0, indicating normal brightness.
[0020] In one implementation, the following steps are used to transform the arc-shaped text region back to a horizontal layout:
[0021] Select the center of the dial as the transformation center point. For each pixel, calculate its polar radius and polar angle relative to the transformation center point. Among them, the polar radius is the distance from the pixel to the transformation center point, and the polar angle is the angle between the line connecting the pixel and the transformation center point and the reference axis.
[0022] Convert the polar coordinates to Cartesian coordinates. By calculating the polar radius and polar angle of each pixel in the polar coordinate system, convert them into coordinate values in the Cartesian coordinate system.
[0023] Perform interpolation processing. Since the pixels of the original image do not necessarily fall on integer coordinates, interpolation processing needs to be performed on the converted Cartesian coordinates to obtain the final polar coordinate image.
[0024] In one implementation, using the morphological dilation and erosion algorithms to denoise and connect the binary image of the text after rough segmentation includes:
[0025] Perform an erosion operation on the thresholded binary image to remove the influence of noise near the text and the thin arcs above and below.
[0026] Perform a dilation operation on the image after the previous step to connect multiple digits together to form a whole.
[0027] Perform an erosion operation on the image after the previous step to shrink the overly dilated area in the left-right and up-down directions so that it can just surround the text.
[0028] In one implementation, the text recognition network CRNN is a deep learning algorithm for text recognition, composed of a convolutional neural network CNN and a recurrent neural network RNN. It mainly includes three steps: convolutional feature extraction, sequence modeling, and transcription output, automatically converting the text in the image into readable text.
[0029] The convolutional neural network CNN extracts the spatial features in the image by sliding the convolutional kernel on the image and performing convolutional operations at each position, enabling the convolutional neural network CNN to capture the local features of the image.
[0030] The convolutional neural network CNN will input its convolutional features into the LSTM to use the LSTM to model the sequence. Among them, the LSTM is a special RNN with memory units and gating mechanisms, capable of processing variable-length sequence data. In the text recognition network CRNN, the LSTM is responsible for performing temporal modeling on the features extracted by the convolutional neural network CNN and updating its state according to the output of the previous time slice.
[0031] The text recognition network CRNN uses a fully connected layer to map the features in the sequence to the text output.
[0032] In one embodiment, the method for obtaining the scale reading corresponding to the current pointer is as follows: When the recognition result is correct, obtain the center coordinates of the minimum bounding rectangle of the corresponding text region, and calculate the angle between the center coordinates (x1, y1) of the current bounding rectangle and the center coordinates (x2, y1) of the horizontal pointer. r is the radius of the dial;
[0033] The obtained angle exactly corresponds to the actual scale. According to whether the coordinates of the text region are on the left or right side of the pointer, use the result of text box recognition plus or minus θ to obtain the final pointer reading result.
[0034] A method for identifying the compass scale of an optoelectronic azimuth instrument based on computer vision provided by the present invention has the following beneficial effects:
[0035] (1) Compared with the traditional manual reading scheme, the present invention can monitor various azimuth indicators of the optoelectronic azimuth instrument in real time;
[0036] (2) Through the traditional image processing algorithm plus the deep learning algorithm, the rapid positioning and accurate recognition of the text on the optoelectronic azimuth instrument panel are realized. Compared with the traditional optical character scheme that segments and recognizes single characters in sequence, the present invention can recognize all characters of variable length at one time through the CRNN algorithm, and can effectively improve the recognition accuracy and suppress the influence of noise through a large number of training data of numbers;
[0037] (3) In practical applications, the significance of automatic recognition of the compass image scale is mainly manifested in the aspect of enemy situation monitoring, which can quickly and accurately measure the azimuth of enemy situation targets and provide important data support for subsequent combat decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a schematic flow chart of a method for identifying the compass scale of an optoelectronic azimuth instrument based on computer vision provided by the present invention.
[0039] Figure 2 is a schematic structural diagram of the text recognition network CRNN. DETAILED DESCRIPTION OF THE INVENTION
[0040] The following further elaborates in detail on a method for identifying the compass scale of an optoelectronic azimuth instrument based on computer vision proposed by the present invention in combination with the accompanying drawings and specific embodiments. According to the following description, the advantages and features of the present invention will be clearer. It should be noted that the accompanying drawings are all in a very simplified form and use non-precise scales, only for the purpose of facilitating and clearly assisting in explaining the purpose of the embodiments of the present invention.
[0041] The present invention provides a method for identifying the compass scale of an optoelectronic azimuth instrument based on computer vision, which uses traditional image processing and combines deep learning technology to automatically identify and report the readings of the mechanical dial of the optoelectronic azimuth instrument. The specific step process is as Figure 1 shown, and includes the following steps:
[0042] Step S1: Use a short-focus optical lens and a camera to collect and capture a local image near the dial pointer. Align the center of the field of view with the pointer, and rotate the dial 360 degrees while the pointer remains stationary.
[0043] Step S2: For images under different ambient light conditions, evaluate the average brightness and variance information of the images to determine whether the current image illumination is appropriate, which is divided into three categories: too weak light intensity, appropriate illumination, and too strong illumination, and complete the preprocessing of the images.
[0044] Specifically included in this step S2 are: (1) Convert the input color picture into a grayscale image. (2) Traverse each pixel of the grayscale image, and use the grayscale value of each pixel point as the abscissa and the frequency of the grayscale value appearing on the entire picture as the ordinate to calculate the grayscale histogram among 256 pixel points. (3) Calculate the average value and standard deviation of the grayscale histogram, and determine whether it is a too dark or too bright situation through the ratio of the average value and the standard deviation. (4) If the picture is too bright, set a threshold according to the size of the average value and return 1; if the picture is too dark, also set a threshold according to the size of the average value and return -1; if the picture is neither too bright nor too dark, return 0, indicating normal brightness.
[0045] The numbers and scales on the dial are arranged around the circular dial area, and multiple numbers are arranged in an arc and not horizontally. Direct identification will affect the identification accuracy. Therefore, it is necessary to perform a polar coordinate transformation on the entire dial area to re-transform the arc-shaped text area into a horizontal arrangement.
[0046] Perform a polar coordinate transformation on the entire dial area. The polar coordinate transformation of an image is to convert an image from a rectangular coordinate system to a polar coordinate system, which is achieved through the following steps:
[0047] (1) Determine the transformation center point, usually choosing the center of the dial as the transformation center point. (2) Calculate the polar radius and polar angle. For each pixel point, calculate its polar radius and polar angle relative to the transformation center. Among them, the polar radius is the distance from the pixel point to the transformation center, and the polar angle is the angle between the line connecting the pixel point and the transformation center and the reference axis. (3) Convert the polar coordinates to Cartesian coordinates. By calculating the polar radius and polar angle of each pixel point in the polar coordinate system, they can be converted into coordinate values in the Cartesian coordinate system. (4) Perform interpolation processing. Since the pixel points of the original image do not necessarily fall on integer coordinates, interpolation processing needs to be performed on the converted Cartesian coordinates to obtain the final polar coordinate image. Through the above algorithms and steps, the dial image can be converted from the Cartesian coordinate system to the polar coordinate system, better presenting the characteristics and forms of the dial.
[0048] Step S4: According to the illumination result of the image evaluation preprocessed in Step S2, perform binary segmentation on the image using different thresholds to roughly segment the dial text area.
[0049] Step S5: Use the morphological dilation and erosion algorithms to denoise and connect the binary text image after rough segmentation, removing the interference of dust around the text and the upper and lower arc segments, and at the same time enabling multiple digits to be connected together.
[0050] First, use a rectangular structuring element filter with a width of 7 and a height of 3 to perform an erosion operation on the thresholded binary image to remove the noise near the text and the influence of the thin arc segments above and below; then use a rectangular structuring element filter with a width of 23 and a height of 3 to perform a dilation operation on the image after the previous step, enabling multiple digits to be connected together to form a whole; finally, use a rectangular structuring element with a width of 15 and a height of 9 to perform an erosion operation on the image after the previous step, shrinking the overly dilated area in the left-right and up-down directions so that it can just surround the text around.
[0051] Step S6: Obtain the connected regions of the text area after denoising and connection, sort the multiple connected regions in the order of pixel area size, find the minimum bounding rectangle of the region with the largest area, and segment the text image from the original image according to the coordinate information.
[0052] Step S7: Scale the segmented text image proportionally to a height of 32 pixels, and fill a fixed pixel value on the right to make the image width 160 pixels; input it into the deep learning-based text recognition network CRNN, extract features through the convolutional neural network and predict the probability of each digit through the recurrent neural network to obtain the digital readings of the dial scale near the pointer.
[0053] In recent years, deep learning has shined brightly in various fields, constantly refreshing people's imagination and beautiful vision of artificial intelligence. Such asFigure 2 The following is a schematic diagram of the structure of the text recognition network CRNN. CRNN (Convolutional Recurrent Neural Network) is a deep learning algorithm for text recognition. CRNN consists of a CNN (Convolutional Neural Network) and an RNN (Recurrent Neural Network), and mainly includes three steps: convolutional feature extraction, sequence modeling, and transcription output, which automatically converts the text in the image into readable text.
[0054] First, the CNN is used to extract the spatial features in the image. The CNN achieves this by sliding the convolutional kernel over the image and performing convolutional operations at each position. This enables the CNN to capture the local features of the image. Next, the CNN feeds its convolutional features into the LSTM to model the sequence using the LSTM. The LSTM is a special type of RNN with memory cells and gating mechanisms that can handle variable-length sequence data. These gating mechanisms enable the LSTM to better capture long-term dependencies. In CRNN, the LSTM is responsible for performing temporal modeling on the features extracted by the CNN and updating its state based on the output of the previous time slice. Finally, CRNN uses a fully connected layer to map the features in the sequence to the text output.
[0055] CRNN performs well in text recognition because it can utilize the local feature extraction of the CNN and the long-term dependency capture of the LSTM to solve the recognition problem. CRNN is an end-to-end text recognition algorithm that can directly learn the recognition task from the raw data without the need for manually designed features; it can be easily applied to various OCR (Optical Character Recognition) scenarios and has been widely used in practical applications.
[0056] Step S8: Conduct a preliminary rule check on the recognized reading result, including whether it is all digits, whether the number of digits is less than 3, and whether the number range is compliant; if all the check results are satisfied, the current recognition result is considered correct; if any one of the checks is not satisfied, repeat steps S6 and S7 to select the second-largest area and re-recognize.
[0057] Step S9: Take the correct text recognition result from the previous step and calculate the angle between the center of the minimum bounding rectangle of the corresponding text area of this correct result and the center of the pointer, and convert it to the accurate scale reading corresponding to the current pointer.
[0058] Obtain the scale reading corresponding to the current pointer. The specific process is as follows: Obtain the center coordinates of the minimum bounding rectangle of the text area corresponding to the correct result, and calculate the angle between the center coordinates (x1, y1) of the current bounding rectangle and the center coordinates (x2, y1) of the horizontal pointer. The radius of the dial is r. The obtained angle exactly corresponds to the actual scale. According to whether the coordinates of the text area are on the left or right side of the pointer, add or subtract θ to the result recognized by the text box to obtain the final pointer reading result.
[0059] Step S10: Use the multi-frame recognition results for comprehensive determination and report them to the control end at a frequency of once per second. The rotation of the angle of the optoelectronic azimuth instrument panel is a slow and continuous process. If the difference between the recognition results of two adjacent frames within one second is greater than 5, it is considered that the current recognition is incorrect, and the result of the previous second is maintained for output; otherwise, the result is updated normally.
[0060] The above description is only a description of the preferred embodiments of the present invention and does not limit the scope of the present invention in any way. Any changes and modifications made by those of ordinary skill in the art of the present invention according to the above disclosure fall within the protection scope of the claims.
Claims
1. A method for identifying the compass scale of an optoelectronic azimuth instrument based on computer vision, characterized in that, Including: Use a short - focus optical lens and a camera to collect and capture a local image near the dial pointer. Align the center of the field of view with the pointer, and rotate the dial 360 degrees while the pointer remains stationary; For images under different ambient light conditions, evaluate the average brightness and variance information of the image to determine whether the current image illumination is appropriate, and complete the pre - processing of the image; For the numbers and scales surrounding the circular dial area, perform a polar coordinate transformation on the entire dial area to re - transform the arc - shaped text area into a horizontal arrangement; According to the illumination result of the image evaluation in the above pre - processing, use different thresholds to perform binary segmentation on the image to roughly segment the dial text area; Use the morphological dilation and erosion algorithms to denoise and connect the binary text image after rough segmentation, remove the interference of dust around the text and the upper and lower arc - shaped line segments, and at the same time connect multiple numbers together; Find the connected regions of the denoised and connected text area, sort the multiple connected regions in the order of pixel area size, find the minimum bounding rectangle of the region with the largest area, and segment the text image from the original image according to the coordinate information; Scale the segmented text image proportionally to a height of 32 pixels, and fill a fixed pixel value on the right until the image width is 160 pixels; input it into the deep - learning - based text recognition network CRNN. Extract features through the convolutional neural network and predict the probability of each digit through the recurrent neural network to obtain the digital reading of the dial scale near the pointer; Perform a preliminary rule check on the recognized reading results, including whether it is all digits, whether the number of digits is less than 3, and whether the number range is compliant; if all the check results are met, the current recognition result is considered correct; if any one of the checks is not met, select the region with the second - largest area to re - perform the recognition; Take the correct text recognition result, calculate the angle between the center of the minimum bounding rectangle of the corresponding text area of the correct result and the center of the pointer, and convert it to obtain the accurate scale reading corresponding to the current pointer; Use the comprehensive determination of multiple - frame recognition results and report them to the control end at a frequency of once per second; the rotation of the photoelectric azimuth instrument panel angle is a slow and continuous process. If the difference between the recognition results of two adjacent frames within one second is greater than 5, it is considered that the current recognition is incorrect, and the result of the previous second is maintained; otherwise, the result is updated normally; The method for obtaining the scale reading corresponding to the current pointer is as follows: when the recognition result is correct, obtain the center coordinates of the minimum bounding rectangle of the corresponding text region, and calculate the center coordinates of the current bounding rectangle and the center coordinates of the horizontal pointer to obtain the included angle , where is the radius of the dial The obtained angle exactly corresponds to the actual scale. Depending on whether the coordinates of the text area are to the left or right of the pointer, add or subtract the result of text box recognition by to obtain the final pointer reading result.
2. The method for identifying the compass scale of an optoelectronic azimuth instrument based on computer vision according to claim 1, wherein The above - mentioned completion of the image pre - processing includes: Convert the input color picture into a grayscale image, traverse each pixel of the grayscale image, use the grayscale value of each pixel point as the abscissa and the frequency of the grayscale value appearing on the whole picture as the ordinate, and calculate the grayscale histogram among 256 pixel points; Calculate the average value and standard deviation of the grayscale histogram, and determine whether it is over - bright or over - dark by the ratio of the average value and the standard deviation; if the picture is over - bright, set the threshold according to the size of the average value and return 1; if the picture is over - dark, also set the threshold according to the size of the average value and return - 1; if the picture is neither over - bright nor over - dark, return 0, indicating normal brightness.
3. The method for identifying the compass scale of the optoelectronic azimuth instrument based on computer vision according to claim 1, wherein, The re - transformation of the arc - shaped text area into a horizontal arrangement is achieved through the following steps: Select the center of the dial as the transformation center point. For each pixel, calculate its polar radius and polar angle relative to the transformation center point. Among them, the polar radius is the distance from the pixel to the transformation center point, and the polar angle is the angle between the line connecting the pixel and the transformation center point and the reference axis. Convert the polar coordinates to Cartesian coordinates. By calculating the polar radius and polar angle of each pixel in the polar coordinate system, convert them into coordinate values in the Cartesian coordinate system. Perform interpolation processing. Since the pixel points of the original image do not necessarily fall on integer coordinates, interpolation processing needs to be performed on the converted Cartesian coordinates to obtain the final polar coordinate image.
4. The method for identifying the compass scale of an optoelectronic azimuth instrument based on computer vision according to claim 1, characterized in that, Use the morphological dilation and erosion algorithms to denoise and connect the binary image of the text after rough segmentation, including: Perform an erosion operation on the thresholded binary image to remove the noise near the text and the influence of the thin arcs above and below. Perform a dilation operation on the image after the previous step to connect multiple digits together to form a whole. Perform an erosion operation on the image after the previous step to shrink the overly dilated area in the left-right and up-down directions so that it can just surround the text.
5. The method for identifying the compass scale of the optoelectronic azimuth instrument based on computer vision according to claim 1, wherein The text recognition network CRNN is a deep learning algorithm for text recognition, composed of a convolutional neural network CNN and a recurrent neural network RNN. It mainly includes three steps: convolutional feature extraction, sequence modeling, and transcription output, which automatically converts the text in the image into readable text. The convolutional neural network CNN extracts the spatial features in the image by sliding the convolutional kernel on the image and performing convolutional operations at each position, enabling the convolutional neural network CNN to capture the local features of the image. The convolutional neural network CNN inputs its convolutional features into the LSTM to use the LSTM to model the sequence. Among them, the LSTM is a special RNN with memory units and gating mechanisms, capable of processing variable-length sequence data. In the text recognition network CRNN, the LSTM is responsible for performing temporal modeling on the features extracted by the convolutional neural network CNN and updating its state according to the output of the previous time slice. The text recognition network CRNN uses a fully connected layer to map the features in the sequence to the text output.
Citation Information
Patent Citations
Pointer type instrument reading identification method based on computer vision and deep learning
CN111046881A
Instrument angle correction and reading identification method based on deep learning
CN116188756A