A method, device, equipment and storage medium for identifying instrument dial readings

The integration of YOLOv5 and DeepLabV3+ models with feature pyramid networks addresses the issue of low accuracy and reliability in instrument dial reading by improving the robustness and precision of dial reading recognition across diverse environments.

CN119206699BActive Publication Date: 2025-07-15厦门市政环境科技股份有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411352941.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-26
Publication Date
2025-07-15
Estimated Expiration
2044-09-26

AI Technical Summary

Technical Problem

In the prior art, the recognition accuracy and reliability of the instrument panel reading are insufficient, which mainly due to poor image quality, resulting in low accuracy and poor reliability of image processing.

Method used

The dial area is positioned using the YOLOv5 model, combined with the DeepLabV3+ model for feature extraction and fusion, and the shallow and deep features are fused through the feature pyramid network to achieve semantic segmentation and reading recognition.

Benefits of technology

Improves the recognition accuracy and reliability of instrument panel readings, and can accurately identify readings under complex backgrounds and different lighting conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119206699B_ABST
    Figure CN119206699B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device, equipment and storage medium for identifying the readings of an instrument dial. First, an image to be recognized collected by an image acquisition device is obtained, and a pre-trained YOLOv5 model is called to recognize the image to be recognized to locate the dial area image. Then, a pre-trained DeepLabV3+ model is called to extract features from the dial area image to generate deep features and shallow features. Finally, the deep features and the shallow features are fused to generate fused features, and semantic segmentation is performed on the fused features to identify and output the reading result of the dial. The problem of insufficient recognition accuracy and reliability caused by poor image quality of the captured image is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of instrument recognition, and particularly to a method, device, equipment and storage medium for recognizing instrument dial readings. Background Art

[0002] Traditional instrument dial readings are realized by manual reading and transmitting the data to the cloud for recording. Traditional manual reading has low accuracy and high labor costs. With the development of technology, instrument panel readings gradually adopt intelligent recognition technology, and intelligent recognition of instrument dial readings is achieved through computer vision technology. However, most current reading recognitions are based on traditional image processing technology. There are too many styles of instrument dials, and the pictures taken of many instrument dials in the environment are not clear enough. In many cases, high quality requirements for image processing are required, resulting in low accuracy and poor reliability of image processing of instrument dials.

[0003] In view of this, this application is proposed. Summary of the Invention

[0004] The present invention discloses a method, device, equipment and storage medium for recognizing instrument dial readings, aiming to solve the problem of insufficient recognition accuracy and reliability due to poor quality of the captured images.

[0005] The first embodiment of the present invention provides a method for recognizing instrument dial readings, including:

[0006] Obtain the image to be recognized collected by the image acquisition device, and call the pre-trained YOLOv5 model to recognize the image to be recognized to locate the dial area image;

[0007] Call the pre-trained DeepLabV3+ model to extract features from the dial area image to generate deep features and shallow features, where the shallow features include: the specific shape of the pointer, the exact position of the scale line, and the contour of the number, and the deep features are the shape and type of the instrument panel;

[0008] Fuse the deep features and the shallow features to generate fused features, perform semantic segmentation on the fused features and recognize and output the reading result of the dial. The fusion process is: using a feature pyramid network to fuse the shallow features and the deep features in a top-down and lateral connection manner.

[0009] Preferably, the step of calling the pre-trained YOLOv5 model to recognize the image to be recognized to locate the dial area image is specifically:

[0010] Divide the image to be recognized into n*n grids, and generate a confidence evaluation result based on whether there is an instrument in each grid. Specifically, if there is, the confidence Pc = 1; if not, the confidence Pc = 0.

[0011] Call the image annotation tool, draw a prediction box on the image to be recognized according to the confidence evaluation result, and locate the dial area image based on the position information of the prediction box.

[0012] Preferably, call the pre-trained DeepLabV3+ model to extract features from the dial area image to generate deep features and shallow features. Specifically:

[0013] Perform shallow feature extraction on the dial area image through the resnet deep residual network algorithm to generate the basic shallow features and feature pictures of the dial area image.

[0014] Extract features from the feature pictures through three different atrous convolutions and global image features to generate the feature layer of the picture, and then obtain the deep features through a 1*1 convolutional layer.

[0015] Preferably, perform semantic segmentation on the fused features and recognize and output the reading result of the dial. Specifically:

[0016] Based on the DeepLabV3+ algorithm, recognize the fused features to recognize the starting scale line and digital display, and the ending scale line and digital display.

[0017] According to the coordinates of the three points of the scale line, the pointer reading position, and the pointer axis center, generate the mathematical equation of the scale line circle, the included angle between scales, and the included angle between the scale and the pointer, and then generate the pointer reading.

[0018] Preferably, before calling the pre-trained YOLOv5 model to recognize the image to be recognized to locate the dial area image, it also includes training the YOLOv5 model. The training process is specifically as follows:

[0019] Obtain the instrument panel dataset, label the instrument panel area in the instrument panel dataset, and label the pictures after area labeling and non-instrument panel pictures as the first sample set. The instrument panel dataset includes instrument panels of different styles and environments and pictures similar to the instrument panels.

[0020] Use the first sample set as model data, and extract training set features by inputting different picture data.

[0021] Based on the training set features, generate an instrument panel recognition model based on the YOLOv5 algorithm.

[0022] Label the images that have not been collected before as the second sample set, further collect the second sample set, input the second sample set into the trained instrument dial recognition model, and further optimize the model. If the output result is incorrect, optimize the weight parameters of the model.

[0023] Preferably, taking the first sample set as model data, by inputting different image data, the training set features are extracted as follows:

[0024] Based on the acquisition algorithm matrix, each grid algorithm matrix is set to 1*1*6, and the algorithm matrix of all grids in the image is 19*19*6. Among them, the expression of the acquisition algorithm matrix is:

[0025] y=[P c b x b y b h b w -1

[0026] Where: P c is the confidence level, b x is the x coordinate of the center point of the prediction box in the overall image, b y is the y coordinate of the center point of the prediction box in the overall image, b h is the height of the prediction box, b w is the length of the prediction box;

[0027] Map all grid algorithms through CNN. Use the Resize and Normalize functions to adjust the spatial dimensions of the input image to a fixed size and perform normalization processing on the homogeneous variance statistically obtained from the dataset.

[0028] Preferably, before calling the pre-trained DeepLabV3+ model to extract features from the dial area image, it also includes training the DeepLabV3+ model. The training process is as follows:

[0029] Obtain the regional image after extracting the instrument dial area, identify the starting scale line and digital display, and the ending scale line and digital display of the instrument, and use the image after regional identification as the third sample set;

[0030] Extract the features of the third sample set and generate a dashboard reading model based on the DeepLabV3+ model,

[0031] Label the instrument dial area images different from those collected before as the fourth sample set, input the fourth sample set into the dashboard reading model. If the input result is incorrect, optimize the weight parameters of the instrument dial intelligent recognition model, etc. ​

[0032] The second embodiment of the present invention provides a device for identifying the readings of an instrument dial, including:

[0033] A dial area image positioning unit, configured to obtain an image to be recognized collected by an image acquisition device, and call a pre-trained YOLOv5 model to recognize the image to be recognized, so as to locate the dial area image;

[0034] A feature extraction unit, configured to call a pre-trained DeepLabV3+ model to extract features from the dial area image to generate deep features and shallow features, wherein the shallow features include: the specific shape of the pointer, the precise position of the scale lines, and the contour of the numbers, and the deep features include the shape and type of the instrument panel;

[0035] A fusion unit, configured to fuse the deep features and the shallow features to generate fusion features, perform semantic segmentation on the fusion features, and recognize and output the reading result of the dial. The fusion process is: using a feature pyramid network to fuse the shallow features and the deep features in a top-down and lateral connection manner.

[0036] The third embodiment of the present invention provides an apparatus for identifying the readings of an instrument dial, including a memory and a processor. A computer program is stored in the memory, and the computer program can be executed by the processor to implement a method for identifying the readings of an instrument dial as described in any one of the above.

[0037] The fourth embodiment of the present invention provides a computer-readable storage medium storing a computer program, and the computer program can be executed by a processor of a device where the computer-readable storage medium is located to implement a method for identifying the readings of an instrument dial as described in any one of the above.

[0038] Based on the method, device, apparatus, and storage medium for identifying the readings of an instrument dial provided by the present invention, by first obtaining an image to be recognized collected by an image acquisition device, calling a pre-trained YOLOv5 model to recognize the image to be recognized to locate the dial area image; then, calling a pre-trained DeepLabV3+ model to extract features from the dial area image to generate deep features and shallow features; finally, fusing the deep features and the shallow features to generate fusion features, performing semantic segmentation on the fusion features, and recognizing and outputting the reading result of the dial. The problem of insufficient recognition accuracy and reliability caused by poor image quality of the captured image is solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a schematic flowchart of a method for identifying the readings of an instrument dial provided by the first embodiment of the present invention;

[0040] Figure 2 It is a schematic diagram of the process for identifying the readings on the instrument dial provided by the present invention;

[0041] Figure 3 It is a schematic diagram of the process for extracting features of the training set provided by the present invention;

[0042] Figure 4 It is a schematic diagram of the modules of an apparatus for identifying the readings on the instrument dial provided by the second embodiment of the present invention. Detailed implementation manners

[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0044] For a better understanding of the technical solutions of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0045] It should be clear that the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0046] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0047] It should be understood that the term "and / or" used herein is only a kind of association relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0048] Depending on the context, the word "if" used herein may be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (stated condition or event)" may be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".

[0049] The "first / second" mentioned in the embodiments is only used to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged in a specific order or sequence when permitted. It should be understood that the objects distinguished by the "first / second" can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.

[0050] The following will describe in detail the specific embodiments of the present invention with reference to the accompanying drawings.

[0051] The present invention discloses a method, device, equipment and storage medium for identifying instrument dial readings, aiming to solve the problem of insufficient recognition accuracy and reliability caused by poor quality of the captured images.

[0052] Please refer to Figure 1 , the first embodiment of the present invention provides a method for identifying instrument dial readings, which can be executed by an instrument dial reading identification device (hereinafter referred to as the identification device). Specifically, it is executed by one or more processors in the identification device to at least implement the following steps:

[0053] S101, obtain the to-be-identified image collected by the image acquisition device, and call the pre-trained YOLOv5 model to identify the to-be-identified image to locate the dial area image;

[0054] In this embodiment, the identification device can be a terminal with data processing capabilities such as a desktop computer, a laptop computer, a server, a workstation, etc. The identification device can be installed with corresponding operating systems and application software, and the functions required in this embodiment can be realized through the combination of the operating system and the application software; the identification device can establish communication with the image acquisition device, obtain the image collected by the image acquisition device, and identify the image to determine whether it is a dashboard image;

[0055] Specifically, in this embodiment, the to-be-identified image is divided into n*n grids, and a confidence evaluation result is generated based on whether there is an instrument in each grid. Among them, if there is, the confidence Pc = 1, and if not, the confidence Pc = 0;

[0056] Call the image annotation tool, draw a prediction box on the to-be-identified image according to the confidence evaluation result, and locate the dial area image based on the position information of the prediction box.

[0057] It should be noted that in this embodiment, the image to be recognized can be divided into 19*19 grids, and evaluating each grid can reduce the cases of misrecognition and missed recognition. Using binary confidence evaluation (1 or 0) simplifies the calculation process, which can improve the processing speed while ensuring accuracy. At the same time, this grid evaluation method can handle images under various complex backgrounds and different lighting conditions. Even if the position and size of the instrument in the image change, it can be effectively located.

[0058] S102, call the pre-trained DeepLabV3+ model to extract features from the dial area image to generate deep features and shallow features;

[0059] Specifically, in this embodiment: the shallow features of the dial area image can be extracted by the resnet deep residual network algorithm to generate the basic shallow features and feature pictures of the dial area image;

[0060] Among them, the shallow features are the features extracted from the network levels closer to the input layer. They have a higher resolution and are closer to the original image. In dashboard recognition, the shallow features can capture the specific shape of the pointer, the precise position of the scale lines, and the outline of the numbers, etc. The resnet deep residual network algorithm can effectively solve the problems of gradient disappearance and gradient explosion in the training of deep neural networks, so as to extract rich feature information in deeper networks. At the same time, resnet can provide reliable shallow features in image recognition tasks;

[0061] Among them, the deep features are the features extracted from the network levels farther from the input layer. These features are more abstract and global, contain higher-level semantic information, have less noise, but lower resolution. In dashboard recognition, the deep features can help identify the layout and overall structure of the entire dashboard, such as distinguishing different types of dashboards (round, square) or identifying the overall patterns on the dashboard (such as speedometer, fuel gauge). The feature pictures are extracted by three different dilated convolutions and global image features to generate the feature layers of the pictures, and then the deep features are obtained through a 1*1 convolutional layer.

[0062] It should be noted that in this embodiment, multiple feature maps of the same size can be extracted, specifically through 1*1Conv, 3*3Convrate6, 3*3Convrate12, 3*3Convrate18, and image pooling to extract, so as to expand the receptive field of the convolutional kernel without increasing the number of parameters and computational complexity, thereby capturing a larger range of context information; then, after splicing the feature maps of multiple feature maps of the same size, a 1*1 convolutional layer is used to obtain deep features, which can effectively compress and integrate the information of the feature layer, map multi-channel features to a low-dimensional space, reduce the amount of calculation and improve the representation ability of features. Further, global feature extraction can capture the global information of the image and make up for the details that may be missed by local convolution.

[0063] S103, fuse the deep features and the shallow features to generate fused features, and perform semantic segmentation on the fused features and identify and output the reading result of the dial.

[0064] Specifically, in this embodiment, based on the DeepLabV3+ algorithm, the fused features are identified to identify the starting scale line and digital display, and the ending scale line and digital display.

[0065] According to the coordinates of the three points of the scale line, the pointer reading position, and the pointer axis center, the mathematical equation of the scale line circle, the included angle between the scales, and the included angle between the scale and the pointer are generated, and then the pointer reading is generated.

[0066] In this embodiment, the specific recognition principle is:

[0067] Let the center point of the instrument dial be (0, 0), the coordinate of the starting scale line of the instrument be (x1, y1), the digital display be M, the coordinate of the ending scale line of the instrument be (x2, y2), the digital display be N, the current coordinate of the position where the instrument pointer is located be (α, β), and the internal scale of the instrument dial be elliptical or circular. Through the three points, the rectangular coordinate equation of the ellipse can be known as:

[0068] Ax 2 +Bxy+Cy 2 +Dx+Ey+F=0

[0069] The rectangular coordinate equation of the ellipse can be converted into a polar coordinate equation:

[0070] x=ρcosθ; y=ρsinθ

[0071] The included angle between the position where the instrument pointer is located and the starting scale line is:

[0072]

[0073] The included angle between the starting scale line and the ending scale line of the instrument is:

[0074]

[0075] The starting scale line number is displayed as M, and the ending scale line number is displayed as N. The reading X of the pointer position can be obtained through the angle between the coordinates of the three:

[0076]

[0077] Please combine Figure 3 , in a possible implementation manner of the present invention, before calling the pre-trained YOLOv5 model to identify the image to be recognized to locate the dial area image, it further includes training the YOLOv5 model, and the training process is as follows:

[0078] Collect instrument dial images of various different styles and in different environments from multiple sources to ensure the diversity and wide coverage of the data set. Among them, the data set includes not only clear instrument dial images, but also pictures similar to instrument dials, which can increase the discrimination ability of the model.

[0079] Manually annotate the dial area in the collected instrument dial images to generate pictures after area identification. Generate the first sample set, including the pictures of the instrument dial area after identification and non-instrument dial pictures.

[0080] Use the pictures in the first sample set as model training data and input them into the YOLOv5 model. Different picture data can be preprocessed first, including size adjustment and normalization processing, and scale the image pixel values from the range [0, 255] to the range [0, 1] or other ranges to help the model converge better. Among them, the normalization formula is as follows:

[0081]

[0082] In the formula, x is the original pixel value, and mean and std are the mean and standard deviation of the data set respectively.

[0083] After preprocessing the pictures, the YOLOv5 model automatically extracts features from the preprocessed pictures through a convolutional neural network, and finally generates training set features.

[0084] Among them, each convolutional layer scans the image through a filter (or called a kernel) to extract local features. The convolution operation can be expressed by the following formula:

[0085]

[0086] In the formula, W is the weight matrix, X is the input image, and * represents the convolution operation.

[0087] Among them, pooling operations usually use max-pooling or average-pooling to reduce the data dimension while retaining the most important information. The formula for max-pooling can be expressed as:

[0088]

[0089] In the formula, stride is the stride of pooling.

[0090] Based on the extracted training set features, the YOLOv5 algorithm is used for model training to generate a preliminary instrument dial recognition model. During the training process, data augmentation techniques such as rotation, scaling, flipping, etc. can be adopted to improve the generalization ability of the model.

[0091] After the preliminary training of the model is completed, new pictures are further collected from the unused picture sources to generate a second sample set. It includes more diverse dial images to further improve the recognition accuracy and robustness of the model. The second sample set is input into the trained instrument dial recognition model for further training and tuning. If the model output result is not ideal, the weight parameters of the model are optimized and adjusted. Through continuous iterative training and optimization, the recognition accuracy of the model is gradually improved. The model performance is evaluated after each iteration until the expected recognition effect is achieved.

[0092] It should be noted that optimizing and adjusting the weight parameters of the model first requires using a loss function to measure the difference between the model prediction value and the actual value. Here, the cross-entropy loss function is used:

[0093]

[0094] In the formula: y i is the true label, is the probability predicted by the model.

[0095] Then, use the gradient descent optimization algorithm to minimize the loss function by iteratively updating the weights:

[0096] w new = w old - η·abla w L

[0097] In the formula: w new and w old are the updated and pre-updated weights respectively, η is the learning rate, and bla w L is the gradient of the loss function with respect to the weights.

[0098] Next, use the method of backpropagation to calculate the gradient, and pass the error from the output layer to the input layer through the chain rule to calculate the gradient of the weights of each layer. For the weights of the l-th layer, its gradient can be expressed as:

[0099]

[0100] wherein, is the weighted input of the j-th neuron in the l-th layer.

[0101] Finally, after obtaining the gradients of all weights, update the weights using the gradient descent rule:

[0102]

[0103] Repeat the above process, update the weights with new data in each iteration, and evaluate the model performance on the validation set. If the performance does not meet the expectation, hyperparameters can be adjusted or more rounds of training can be continued.

[0104] In a possible implementation manner of the present invention, taking the first sample set as model data, by inputting different picture data, extract the training set features as follows:

[0105] Divide the image to be recognized into small grids of n*n, and each grid is represented by an algorithm matrix.

[0106] The size of the algorithm matrix for each grid is 1*1*6, and the total generated matrix is 19*19*6;

[0107] The expression of the algorithm matrix is as follows:

[0108] y = [P c b x b y b h b w -1

[0109] wherein: P c is the confidence, b x is the x coordinate of the center point of the prediction box in the overall picture, b y is the y coordinate of the center point of the prediction box in the overall picture, b h is the height of the prediction box, b w is the length of the prediction box;

[0110] Furthermore, the algorithm matrices of all grids can be mapped through a convolutional neural network (CNN). Through the stacking operation of the CNN, important features in the picture are extracted. The Resize function is used to adjust the spatial size of the input image to a fixed size to meet the input requirements of the model. The Normalize function is used to normalize the image, and the pixel values are adjusted to a unified range. Among them, the normalization process uses the homogeneous variance statistically obtained from the dataset, making the numerical distribution of the image more uniform and improving the training effect of the model.

[0111] Input the preprocessed picture data into the model, and extract features through CNN to generate training set features. Among them, the extracted features contain the spatial information and context information of the picture, providing comprehensive data support for the training of the model.

[0112] In a possible implementation manner of the present invention, before calling the pre-trained DeepLabV3+ model to extract features from the dial area image, it further includes training the DeepLabV3+ model, and the training process is as follows:

[0113] Extract pictures of the dial area from instrument dial images of various different types and environments.

[0114] Identify the starting scale lines and digital displays, and the ending scale lines and digital displays on the extracted pictures of the dial area to ensure that the model can recognize and distinguish these key areas. Use the pictures after the above identification as the third sample set for the initial training of the model.

[0115] Use deep learning algorithms to extract features from the third sample set. By extracting these features, the model can learn the key areas and identification information in the pictures. In the DeepLabV3+ model, the extraction of features is mainly achieved through the atrous convolution and spatial pyramid pooling modules:

[0116] Among them, the shallow features first use atrous convolution to expand the receptive field. Atrous convolution expands the effective receptive field of the convolution kernel by inserting blank spaces (holes) between standard convolution kernels, while maintaining a high-resolution feature map. For a given dilation rate d, the convolution operation can be expressed as:

[0117]

[0118] In the formula, x[i] is the value at the i-th position of the input feature map, w[k] is the k-th weight of the convolution kernel, and K is the size of the convolution kernel.

[0119] Secondly, use the spatial pyramid pooling (ASPP) module to use parallel atrous convolutions with different dilation rates to capture multi-scale context information. The output of ASPP is the combined result of these parallel convolution operations:

[0120] y[i] = Concat(y1[i], y2[i],..., y n [i])

[0121] In the formula, y1, y2,..., y n are the outputs of atrous convolutions with different dilation rates.

[0122] After ASPP, there is usually a 1×1 1×1 convolutional layer to reduce the number of channels and fuse features:

[0123]

[0124] Where y[i,j] is the value of the j-th channel at the i-th position of the ASPP output, C is the number of channels, and w[j] is the weight of the 1×1 1×1 convolutional kernel.

[0125] Through these steps, the DeepLabV3+ model can extract rich shallow features while maintaining high resolution, which is crucial for capturing fine details in the image (such as the pointers and scale lines on the dashboard). These features will then be used in deeper network layers for further feature abstraction and semantic understanding.

[0126] Among them, the deep features also first use dilated convolution to expand the receptive field to capture larger context information while maintaining a high-resolution feature map. Then, the ASPP module uses parallel dilated convolutional layers with different dilation rates to capture multi-scale features, and obtains the overall image features through the global average pooling layer. The calculation formula of global average pooling in the ASPP module:

[0127]

[0128] Where H and W are the height and width of the feature map respectively.

[0129] Then, the feature map after dilated convolution and ASPP processing will pass through a 1x1 convolutional layer for dimensionality reduction and feature fusion to generate the final deep features. Finally, by combining dilated convolution, ASPP, and the 1x1 convolutional layer, the DeepLabV3+ model can effectively extract highly abstract and global deep features from network layers far from the input layer. These deep features are crucial for identifying the overall layout and structure of the dashboard, and help to distinguish different types of dashboards or identify the overall patterns on the dashboard.

[0130] For the shallow features and deep features, the Feature Pyramid Network (FPN) is used to fuse the shallow features and deep features formula through a top-down and lateral connection method:

[0131] 1. Top-down path: Upsample the high-level feature map to align its spatial size with the low-level feature map. Assume the current layer is the i-th layer, and its feature map is C i , then the upsampled feature map is denoted as P i . The upsampling operation can be achieved through interpolation methods;

[0132] 2. Lateral connection: At each level, the feature maps obtained by upsampling from top to bottom are laterally connected with the feature maps of the corresponding level obtained by bottom-up processing. This connection method can effectively combine the strong semantic information of the high levels and the detailed information of the low levels to generate fused feature maps. Assume the feature map at the current level is C i-1 , then the fused feature map P i can be expressed as:

[0133] P i = C i-1 + Upsample(P i+1 )

[0134] In the formula, Upsample represents the upsampling operation.

[0135] Based on the DeepLabV3+ model, by training the extracted features, a model capable of accurately reading and recognizing the dashboard readings is generated. New pictures of the dial area are further collected from instrument dial images of different types and environments. These pictures are different from the previous sample sets and provide more diverse data for model optimization.

[0136] The fourth sample set is input into the already generated dashboard reading model for verification and testing. If the result output by the model is incorrect, the weight parameters of the model are optimized. By using an iterative approach, the model is continuously adjusted and optimized to ensure that it can accurately recognize and read the readings in instrument dial images of different environments and different styles.

[0137] Please refer to Figure 4 , the second embodiment of the present invention provides an identification device for instrument dial readings, including:

[0138] A dial area image positioning unit 201, configured to obtain the image to be recognized collected by the image acquisition device, and call a pre-trained YOLOv5 model to recognize the image to be recognized to locate the dial area image;

[0139] A feature extraction unit 202, configured to call a pre-trained DeepLabV3+ model to extract features from the dial area image to generate deep features and shallow features, where the shallow features include: the specific shape of the pointer, the exact position of the scale lines, and the contour of the numbers, and the deep features include the shape and type of the dashboard;

[0140] A fusion unit 203, configured to fuse the deep features and the shallow features to generate fused features, perform semantic segmentation on the fused features and recognize and output the reading result of the dial, where the fusion process is: using a feature pyramid network to fuse the shallow features and the deep features in a top-down and lateral connection manner.

[0141] The third embodiment of the present invention provides an identification device for instrument dial readings, including a memory and a processor. A computer program is stored in the memory and can be executed by the processor to implement an identification method for instrument dial readings as described in any one of the above.

[0142] The fourth embodiment of the present invention provides a computer-readable storage medium storing a computer program that can be executed by a processor of a device where the computer-readable storage medium is located to implement an identification method for instrument dial readings as described in any one of the above.

[0143] Based on an identification method, device, equipment, and storage medium for instrument dial readings provided by the present invention, by first obtaining a to-be-identified image collected by an image acquisition device, calling a pre-trained YOLOv5 model to identify the to-be-identified image to locate the dial area image; then, calling a pre-trained DeepLabV3+ model to perform feature extraction on the dial area image to generate deep features and shallow features; finally, fusing the deep features and the shallow features to generate fused features, performing semantic segmentation on the fused features, and identifying and outputting the reading result of the dial. The problem of insufficient recognition accuracy and reliability caused by poor image quality of the captured image is solved.

[0144] Exemplarily, the computer program described in the third and fourth embodiments of the present invention can be divided into one or more modules. The one or more modules are stored in the memory and executed by the processor to complete the present invention. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the identification device for instrument dial readings. For example, the device described in the second embodiment of the present invention.

[0145] The so-called processor may be a Central Processing Unit (CPU), or it may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the method for identifying the readings of an instrument dial, and uses various interfaces and circuits to connect the whole to implement various parts of the method for identifying the readings of an instrument dial.

[0146] The memory can be used to store the computer programs and / or modules. The processor realizes various functions of the method for identifying the readings of an instrument dial by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, a text conversion function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, text message data, etc.). In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0147] Among them, if the implemented module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0148] It should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative work.

[0149] As described above, only the preferred specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for identifying the reading of an instrument dial, characterized in that, Including: Obtain the image to be recognized collected by the image acquisition device, and call the pre-trained YOLOv5 model to recognize the image to be recognized to locate the dial area image. Specifically: divide the image to be recognized into n*n grids, and generate a confidence evaluation result based on whether there is an instrument in each grid. Among them, if there is, the confidence Pc = 1, and if not, the confidence Pc = 0; call the image annotation tool, and draw a prediction box on the image to be recognized according to the confidence evaluation result, and locate the dial area image based on the position information of the prediction box. The training process of the YOLOv5 model is: obtain the instrument panel dataset, identify the instrument dial area in the instrument panel dataset, and label the pictures after area identification and the pictures of non-instrument dials as the first sample set. Among them, the instrument panel dataset includes instrument dials of different styles and different environments and pictures similar to instrument dials. Using the first sample set as model data, by inputting different picture data, training set features are extracted. Specifically: based on the acquisition algorithm matrix, each grid algorithm matrix is set to 1*1*6, and the algorithm matrix of all grids of the picture is 19*19*6. Among them, the expression of the acquisition algorithm matrix is: , where: is the confidence level, is the x coordinate of the center point of the prediction box in the overall picture, is the y coordinate of the center point of the prediction box in the overall picture, is the height of the prediction box, is the length of the prediction box; map all grid algorithms through CNN, use the Resize and Normalize functions to adjust the spatial size of the input image to a fixed size, and perform normalization processing on the homogeneous variance statistically obtained from the dataset; Based on the training set features, generate an instrument dial recognition model using the YOLOv5 algorithm; label the pictures that have not been collected before as the second sample set, further collect the second sample set, input the second sample set into the trained instrument dial recognition model, and further optimize the model. If the output result is incorrect, optimize the weight parameters of the model. Call the pre-trained DeepLabV3+ model to perform feature extraction on the dial area image to generate deep features and shallow features. Among them, the shallow features include: the specific shape of the pointer, the exact position of the scale line, and the outline of the number. The deep features include the shape and type of the instrument panel. Fuse the deep features and the shallow features to generate fused features, perform semantic segmentation on the fused features and identify and output the reading result of the dial. Specifically: based on the DeepLabV3+ algorithm, identify the starting scale line and digital display, the ending scale line and digital display in the fused features. According to the coordinates of the three points of the scale line, the pointer reading position, and the pointer axis center, generate the mathematical equation of the scale line circle, the angle between scales, and the angle between the scale and the pointer, and then generate the pointer reading. Among them, the fusion process is: use the feature pyramid network to fuse the shallow features and the deep features in a top-down and lateral connection manner.

2. The identification method of an instrument dial reading according to claim 1, characterized in that, The specific process of calling the pre-trained DeepLabV3+ model to perform feature extraction on the dial area image to generate deep features and shallow features is as follows: Perform shallow feature extraction on the dial area image through the resnet deep residual network algorithm to generate the basic shallow features and feature pictures of the dial area image. Perform feature extraction on the feature pictures through three different atrous convolutions and the global features of the picture to generate the feature layer of the picture, and then obtain the deep features through a 1*1 convolutional layer.

3. The identification method of the instrument dial reading according to claim 1, characterized in that, Before calling the pre-trained DeepLabV3+ model to perform feature extraction on the dial area image, it also includes training the DeepLabV3+ model. The training process is specifically as follows: Obtain the regional image after extracting the instrument dial area, identify the starting scale line and digital display, and the ending scale line and digital display of the instrument, and use the image after regional identification as the third sample set; Extract the features of the third sample set, and generate a dashboard reading model based on the DeepLabV3+ model; Label the instrument dial area images different from those collected before as the fourth sample set, input the fourth sample set into the dashboard reading model. If the input result is incorrect, optimize the weight parameters of the instrument dial intelligent recognition model.

4. An identification device for instrument dial readings, characterized in that It includes: A dial area image positioning unit, which is used to obtain the image to be recognized collected by the image acquisition device, call the pre-trained YOLOv5 model to recognize the image to be recognized, so as to locate the dial area image. Specifically, it is used to divide the image to be recognized into n*n grids, and generate a confidence evaluation result based on whether there is an instrument in each grid. Among them, if there is, the confidence Pc = 1, if not, the confidence Pc = 0; call the image annotation tool, and draw a prediction box on the image to be recognized according to the confidence evaluation result, and locate the dial area image based on the position information of the prediction box. The training process of the YOLOv5 model is: obtain the dashboard data set, identify the instrument dial area in the dashboard data set, and label the images after regional identification and non-instrument dial images as the first sample set. Among them, the dashboard data set includes instrument dials of different styles and environments and images similar to instrument dials. Taking the first sample set as model data, by inputting different picture data, training set features are extracted. Specifically: Based on the acquisition algorithm matrix, each grid algorithm matrix is set to 1*1*6, and the algorithm matrices of all grids of the picture are 19*19*6. Among them, the expression of the acquisition algorithm matrix is: , where: is the confidence level, is the x coordinate of the center point of the prediction box in the overall picture, is the y coordinate of the center point of the prediction box in the overall picture, is the height of the prediction box, is the length of the prediction box; Mapping all grid algorithms is realized through CNN. The spatial dimensions of the input image are adjusted to a fixed size by using the Resize and Normalize functions, and normalization processing is performed on the homogeneous variance statistically calculated from the dataset; Based on the training set features, generate an instrument dial recognition model using the YOLOv5 algorithm; label the images that have not been collected before as the second sample set, further collect the second sample set, input the second sample set into the trained instrument dial recognition model, and further optimize the model. If the output result is incorrect, optimize the weight parameters of the model; A feature extraction unit, which is used to call the pre-trained DeepLabV3+ model to extract features from the dial area image to generate deep features and shallow features. Among them, the shallow features include: the specific shape of the pointer, the precise position of the scale line, and the contour of the number, and the deep features include the shape and type of the dashboard; A fusion unit, which is used to fuse the deep features and the shallow features to generate fusion features, perform semantic segmentation on the fusion features and identify and output the reading result of the dial. Specifically, it is used to identify the starting scale line and digital display, and the ending scale line and digital display based on the DeepLabV3+ algorithm. According to the coordinates of the scale line, the pointer reading point, and the pointer axis center, generate the mathematical equation of the scale line circle, the angle between scales, and the angle between the scale and the pointer, and then generate the pointer reading; among them, the fusion process is: use the feature pyramid network to fuse the shallow features and the deep features in a top-down and horizontal connection manner.

5. An identification device for instrument dial readings, characterized in that, It includes a memory and a processor. A computer program is stored in the memory and can be executed by the processor to implement a method for identifying instrument dial readings as described in any one of claims 1 to 3.

6. A computer-readable storage medium, characterized in that, A computer program is stored and can be executed by the processor of the device where the computer-readable storage medium is located to implement a method for identifying instrument dial readings as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Oral health management system for adjusting electric toothbrush based on artificial intelligence image recognition

    CN113244009A

  • Method for identifying lightweight pointer type instrument locally deployed by inspection robot

    CN115457262A

  • Physical experiment scale instrument reading method based on target detection and semantic segmentation

    CN117456154A