Intelligent industrial instrument identification system based on deep learning

By adopting an adaptive deep learning model and improved angle reading method in industrial instrument recognition systems, the problems of low instrument recognition accuracy and difficulty in automation in the prior art are solved, and an intelligent identification system with high accuracy and scalability are realized.

CN120032355APending Publication Date: 2025-05-23SHANGHAI JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510117727.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When existing industrial instrument identification methods face electromagnetic interference and mechanical vibration in complex industrial environments, the reading accuracy is low and it is difficult to achieve automated and intelligent identification.

Method used

The intelligent industrial instrument recognition system based on deep learning is adopted to adaptively select the object detection model suitable for edge devices, combined with the improved YOLOv10 model of hybrid convolution and defogging module, dial positioning and instrument type recognition are realized, and reading accuracy is improved through improved angle reading method.

Benefits of technology

提高了仪表识别的准确率和系统的可扩展性,降低了读数误差,实现了在工业环境中的实时、自动化识别,满足了工业实践的可行性要求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032355A_ABST
    Figure CN120032355A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent industrial instrument identification system based on deep learning. The system comprises a positioning part which is used for positioning a dial plate part from an original instrument picture; the reading part is used for reading according to the instrument type by applying a corresponding algorithm, the reading precision is improved through an improved angle reading method, and an extensible design is reserved for a new instrument type; and the interaction part is used for providing a simple instrument reading webpage interface for a user. The problem that manual inspection or manual photo reading is needed in industrial operation and maintenance of the transformer substation is solved, the operation and maintenance efficiency is improved through an automatic instrument reading scheme, and errors generated by manual reading are reduced. According to the invention, on the basis of ensuring compatibility of an instrument identification algorithm and enterprise edge equipment, relatively high reading accuracy is realized, so that urgent demands of the industry for automatic and intelligent instrument identification are fully met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of industrial instruments, and in particular to an intelligent industrial instrument identification system based on deep learning. Background Art

[0002] At a time when industrial automation and intelligence are developing rapidly, industrial instruments are key tools for monitoring and controlling the operating status of various types of equipment. There are various types of industrial instruments, including ammeters, voltmeters, lightning arresters, etc., which are suitable for various industrial scenarios. The dial design is mainly divided into two types: pointer type and digital type. Traditionally, manual inspections are generally used to read and monitor the data of industrial instruments. However, with the increasing complexity of the industrial environment, especially in strong electromagnetic scenarios represented by substations, many unfavorable factors have begun to emerge. Among them, factors such as electromagnetic interference and mechanical vibration will have a significant impact on the reading accuracy of digital instruments, resulting in frequent data distortion problems. In contrast, pointer instruments are still the preferred instrument type for the industry to monitor the operating status of equipment such as power and temperature control due to their high anti-interference ability and stability.

[0003] Traditional manual inspection methods have a series of problems such as low efficiency and high safety risks. On the one hand, manual inspection is highly dependent on human resources, which is not only time-consuming and labor-intensive, but also difficult to meet the needs of large-scale industrial facilities for real-time monitoring. On the other hand, there are safety hazards such as electric shock and burns during the inspection process, especially in high-voltage environments, the personal safety of workers will be greatly threatened. In addition, with the continuous expansion of the scale of industrial equipment and the increasing complexity, manual inspections seem to be unable to ensure the consistency and comprehensiveness of data collection, and thus cannot achieve accurate monitoring of the operating status of equipment.

[0004] In recent years, robot inspection technology has emerged. It uses mobile cameras to collect instrument images and transmits the collected instrument images to operation and maintenance personnel through the network by presetting inspection points or remote control. This method reduces manual participation to a certain extent and improves inspection efficiency and safety. However, most existing solutions still rely on manpower to read the photos taken. How to achieve fully automated and intelligent instrument recognition is still a difficult problem that the industry needs to overcome.

[0005] At present, the existing instrument recognition methods are mainly divided into two categories: image recognition algorithms based on traditional image processing and deep learning. Traditional methods, such as Hough transform and other image processing technologies, have been widely used in instrument recognition. Although such methods are simple to apply and have relatively low overhead, they are less robust to image distortion and noise, and are easily affected by complex backgrounds, lighting conditions, and shooting angles in industrial environments, resulting in low overall recognition accuracy. With the rapid development of deep learning technology, it has been widely used in the field of computer vision, especially in the direction of target detection. Algorithms based on deep learning have greatly improved the accuracy and generalization of image recognition through models such as convolutional neural networks. Convolutional Neural Networks (CNN), as a deep learning model specifically used to process data with similar grid structures, are centered on convolution operations and weight sharing. The convolution operation locally connects the input data by sliding the convolution kernel to extract local features; weight sharing enables the same convolution kernel to be reused on the entire input data, thereby greatly reducing model parameters and improving training efficiency. The YOLO (You Only Look Once) series of models transforms the target detection problem into a single regression problem. After multiple versions of iterative optimization, its detection accuracy and speed have been greatly improved, making it the most commonly used method in current target detection tasks.

[0006] For instrument recognition tasks, deep learning models can effectively cope with the diverse changes in instrument images, including different instrument types, dial designs, pointer positions, lighting conditions, and background interference. By selecting a suitable target detection model and designing a reasonable instrument reading process, accurate positioning, classification, segmentation, and reading of various instruments can be achieved. Existing work often uses models such as YOLO to complete the positioning of instruments, pointers, and scales, and then uses mathematical methods to obtain reading results based on the pointer angle and range information. However, deep learning models face many limitations and challenges when they are actually applied in industrial environments. On the one hand, industrial instrument recognition needs to run in real time on edge devices, which places higher requirements on the lightweight and computational efficiency of the model. On the other hand, the diversity and complexity of data collection in industrial environments require the model to have a high degree of generalization and robustness. However, existing solutions have problems to varying degrees in terms of algorithm overhead or recognition accuracy, and most of them are difficult to meet the feasibility requirements of industrial practice.

[0007] Through searching patent documents, it was found that the invention patent with publication number CN112949564B discloses a method for automatic reading of pointer-type instruments based on deep learning, which includes the following steps: S1. Determine the instrument type and instrument location; S2. Extract the dial from the pointer-type instrument image; S3. Perform pointer detection and calculate the pointer angle; S4. Detect the angle between the 0 scale and the maximum scale and the maximum range digital area; S5. Identify the maximum range digital; S6. Perform the reading calculation of the pointer-type instrument to obtain the reading result. This patent only uses the SSD model, focuses on the automatic reading method of the pointer-type instrument, lacks the interactive part, does not retain the scalable design for new instrument types, and does not mention the scalable reading design such as the strategy mode, and the function is relatively simple.

[0008] To sum up, in the field of industrial instrument identification, whether it is the traditional manual inspection method or the existing instrument identification method, there are many problems that need to be solved urgently. Researching an intelligent industrial instrument identification system based on deep learning has become a key task that needs to be solved urgently. Summary of the invention

[0009] In view of the defects in the prior art, the object of the present invention is to provide an intelligent industrial instrument identification system based on deep learning.

[0010] According to the present invention, a deep learning-based intelligent industrial instrument recognition system is provided, comprising:

[0011] Positioning part: Adaptively select the appropriate target detection model according to the hardware conditions of the edge device, process the original instrument image through the target detection model, and obtain the dial image and instrument type;

[0012] Reading part: Based on the instrument type and dial image, the corresponding algorithm is used to read the instrument and obtain the instrument reading result;

[0013] Interactive part: Provides a web user interface to users to display instrument reading results.

[0014] Preferably, the target detection model includes a rescalable backbone architecture, a hybrid convolution module and a defogging module.

[0015] Preferably, according to the hardware conditions of the edge device, the positioning part adopts an adaptive model scale selection method to automatically select an adapted target detection model.

[0016] Preferably, the adaptive model scale selection method includes three stages: an environmental assessment stage, a reasoning fine-tuning stage, and a real-time monitoring stage.

[0017] Environmental assessment phase: Evaluate the hardware and software information of edge devices, measure the model calculation amount with floating-point operations based on the hardware performance and software compatibility of the device, and use heuristic algorithms to select the appropriate model type and size;

[0018] Inference fine-tuning phase: Use the selected model to perform small-scale inference attempts, collect runtime data including inference latency, device load, and instrument recognition effect, and fine-tune the model size based on performance and overhead;

[0019] Real-time monitoring stage: Real-time performance monitoring is performed during the operation of the instrument identification system. The software and hardware conditions are checked at regular intervals. If the inference time is higher than the threshold or the device hardware load is too high, the inference is suspended and the model scale selection logic is re-executed.

[0020] Preferably, the positioning part uses an improved YOLOv10 model to locate the dial area and determine the instrument type from the original instrument image, and according to the dial area, a partial image of the dial is captured from the original instrument image, and then the partial image of the dial is preprocessed to obtain a dial image.

[0021] Preferably, based on the instrument type and the dial image, the reading part is read by an improved angle reading method, which includes: first, using an improved lightweight and directional target detection model to locate the pointer and scale, and then, according to the polar coordinate relationship between the pointer and the scale and the center of the dial, selecting the two scales closest to the pointer; finally, the angle between the pointer and the two scales and the scale markings is used to calculate the actual reading, thereby reducing the angle calculation error.

[0022] Preferably, the reading part implements extensible instrument reading in a strategy mode through an abstract reading strategy class AbstractReadingStrategy, a concrete reading strategy class ConcreteReadingStrategy and a context class Context.

[0023] Preferably, the interactive part includes an input part and an output part. The input part includes two uploading methods: picture and camera, as well as model architecture, image size, and confidence adjustment options; the output part includes instrument reading results and model annotation schematics.

[0024] Preferably, the interactive part uses the Gradiu framework component to build an end-to-end graphical user interface, and calls the meter reading algorithm deployed in the backend through the network protocol.

[0025] Preferably, the front-end components of the graphical user interface include:

[0026] Image, upload images as local files or URLs;

[0027] Textbox, multi-line text input and output;

[0028] Slider, input value or range for parameter adjustment;

[0029] Dropdown, drop-down box option, used to select the model;

[0030] Interface, quickly create interactive applications, customize functions and their input and output;

[0031] Block / Row / Column, structural layout elements of web pages.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] 1. The present invention proposes an improved YOLOv10 model that integrates hybrid convolution and defogging modules, which improves the accuracy of existing instrument recognition schemes and reduces the reading error of pointer instruments through an improved angle reading method.

[0034] 2. The present invention proposes an adaptive model scale adjustment algorithm, which enhances the compatibility of the instrument identification solution with edge devices, enabling the system to achieve a good balance between the balance model and resource overhead; in addition, by formulating special instrument classification standards and adopting unique code design patterns, the scalability of the solution is improved.

[0035] 3. The present invention proposes a simple user interface that can be quickly deployed. While hiding the implementation details, the interface provides users with flexible input options and intuitive output forms, thereby improving the usability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0037] Figure 1 is a framework diagram of an intelligent industrial instrument identification system based on deep learning in an embodiment of the present invention;

[0038] Figure 2 Schematic diagram of a flow chart of an adaptive model scale selection method in an embodiment of the present invention;

[0039] Figure 3 A schematic diagram of the image preprocessing process in an embodiment of the present invention;

[0040] Figure 4 This is an implementation class diagram of the extensible reading design in an embodiment of the present invention. DETAILED DESCRIPTION

[0041] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0042] The present invention provides an intelligent industrial instrument recognition system based on deep learning, wherein the system includes: a positioning part: used to locate the dial part from the original instrument image, firstly, according to the hardware conditions of the edge device, a deep target detection model of appropriate scale is adaptively selected, and then the dial area is identified through model reasoning, the instrument type is determined, and the image is preprocessed using perspective transformation, Gaussian blur and other methods, and the process has the characteristics of high precision and strong real-time performance; the reading part is used to apply the corresponding algorithm to read according to the instrument type, including circular pointer instruments, square pointer instruments, lightning arresters and other types, and the reading accuracy is improved by the improved angle reading method, and the scalable design is retained for new instrument types; the interactive part is used to provide users with a concise instrument reading web page interface, the input includes two modes of picture uploading and camera shooting, and also has model architecture, image size, and confidence adjustment options, and the output includes instrument reading results and model annotation schematics, and the operation is intuitive and the implementation details are hidden. The present invention solves the problem of manual inspection or manual reading of photos in the industrial operation and maintenance of substations. The automatic instrument reading solution not only improves the operation and maintenance efficiency, but also reduces the errors caused by manual reading. The present invention is committed to achieving a higher reading accuracy rate on the basis of ensuring the compatibility of the instrument identification algorithm with the enterprise edge devices, so as to fully meet the urgent needs of the industry for automated and intelligent instrument identification.

[0043] Embodiment 1:

[0044] Figure 1 This is a framework diagram of a deep learning-based intelligent industrial instrument recognition system in an embodiment of the present invention.

[0045] like Figure 1 As shown, this embodiment provides an intelligent industrial instrument recognition system based on deep learning, which includes a positioning part, a reading part and an interaction part. The positioning part includes the basic architecture and related optimization modules of the target detection model, the reading part includes the related objects and their relationships of the extensible reading design pattern; the interaction part includes the input and output content of the user interface. Among them, Figure 1 The solid arrows in the middle indicate the process sequence of the system, and the dotted arrows indicate the relationship between classes.

[0046] Specifically, the intelligent industrial instrument recognition system based on deep learning includes:

[0047] Positioning part: Adaptively select the appropriate target detection model according to the hardware conditions of the edge device, process the original instrument image through the target detection model, and obtain the dial image and instrument type.

[0048] Specifically, the object detection model includes a resizable backbone architecture, a hybrid convolution module, and a defogging module. The object detection model has an optimized architecture design that includes a lightweight hybrid convolution and defogging and denoising modules.

[0049] Among them, the variable-scale backbone architecture changes the number of parameters and computational complexity of the model by adjusting the scale factors of the depth, width and maximum number of channels of the model, and choosing whether to integrate the optimization module to adapt to edge devices of different specifications and reading delay requirements.

[0050] The hybrid convolution module proposes a hybrid feature convolution method to improve the problem of insufficient feature extraction capability of the depthwise separable convolution (DSC) of the existing model. Specifically, standard convolution and depthwise separable convolution are used for general input features respectively. Assuming that the number of input and output channels is C in , C out , the input channel first undergoes a standard convolution to generate C out / 2 channel output; this output undergoes a depth-separable convolution and also produces C out / 2 channel output; these two outputs are spliced ​​and randomly shuffled to finally generate C out The computational complexity of the mixed convolution is O((C out ×K×K×H×W) / 2×(C in +1)), which is about half of the standard convolution, but its learning performance can reach a level close to that of the standard convolution. Then, a more complex feature fusion module is built based on this hybrid convolution logic. The computational complexity of the hybrid convolution module is about half of that of the standard convolution, but its performance is close to that of the standard convolution.

[0051] The dehazing module proposes two basic attention modules, channel attention CA and pixel attention PA, which process the two types of feature information in different ways. CA focuses on the difference in weights of different channel features. First, the global spatial information at the channel level is converted into channel descriptors through global average pooling, then processed through two convolutional layers and Sigmoid and ReLU activation functions, and finally the input is multiplied element by element with the channel weight. PA focuses on the uneven distribution of blur levels on different pixels, and directly inputs the feature map into two convolutional layers, changing its shape from C×W×H to 1×W×H. Next, a basic attention module (BAM) is constructed based on feature attention and local residual learning modules, where the role of residual learning is to skip less important information such as clearer or hazy areas, so that the network can focus more on effective information. Connecting multiple consecutive BAMs to form a combination can further increase the network's expressive power; these combinations can also be added to the network according to performance requirements, and then feature fusion is performed by multiplying with adaptive weights. Finally, two convolutional layers are used as the restoration part to obtain a pre-processed haze-free image.

[0052] Specifically, the dehazing module addresses the blur and distortion problems of photos taken by mobile inspection robots and effectively handles the uneven blur in images through a specially designed attention mechanism. The dehazing module includes a channel attention mechanism that pays more attention to channel weight differences and a pixel attention mechanism that pays more attention to the uneven distribution of blur levels at the pixel granularity. The combination of channel attention, pixel attention and other convolution and activation logic can significantly enhance the clarity and contrast of the image.

[0053] Figure 2 The figure is a flowchart of the adaptive model scale selection method in an embodiment of the present invention.

[0054] like Figure 2 As shown in the figure, according to the hardware conditions of the edge device, the positioning part uses an adaptive model scale selection method to automatically select an adaptive model architecture. The adaptive model scale selection method includes three stages: environmental assessment, inference fine-tuning, and real-time monitoring to achieve a balance between accuracy and algorithm overhead.

[0055] Among them, the environmental assessment stage: evaluate the hardware and software information of the edge device. The hardware information includes the model, main frequency, number of cores, whether the CPU supports hyperthreading technology, etc.; the model, frequency, number of floating-point units, video memory frequency, size, etc. of the graphics processor GPU; the size and frequency of the memory RAM. The software information mainly includes the operating system type, version, and CUDA environment version. Check whether there is an independent GPU. If not, select a lightweight CPU model; if there is an independent GPU, check the compatibility of the system and CUDA environment, and then select a model of the corresponding scale based on the GPU performance, video memory size, and memory size. The model scale is measured by the amount of calculation (GFLOPs, floating-point operations) and the number of parameters. The correspondence between the model scale and the hardware performance is calculated using a heuristic algorithm.

[0056] Inference fine-tuning phase: After the model is determined, a small-scale inference is performed on a given image dataset to collect runtime data including inference latency, device load, and instrument recognition effect, and the model size is fine-tuned according to performance and overhead. If the model inference latency is too high (such as more than 1s) or the device load is too high (such as high CPU and GPU occupancy and high temperature), try to reduce the model size; if the device load is low and the model recognition accuracy is low, try to increase the model size or deepen the model level. This phase will repeat the above process until the inference performance and overhead reach a reasonable level. If it still cannot be achieved within a certain timeout threshold, the system will stop running and provide feedback information.

[0057] Real-time monitoring stage: When the instrument identification system is running, real-time performance monitoring is performed, and the system operation status is checked at regular intervals (such as one minute). If the model inference time exceeds the threshold or the device hardware load is too high, it means that the current device instance is occupied by other applications or the concurrent access volume is too high, and the current model can no longer meet the demand. Therefore, the system operation is suspended and the model selection logic is re-executed.

[0058] In this embodiment, the positioning part uses an improved YOLOv10 model to locate the dial area and determine the instrument type from the original instrument image. According to the dial area, a partial image of the dial is captured from the original instrument image, and then the partial image of the dial is preprocessed to obtain a dial image.

[0059] Specifically, since the target of the pointer and scale recognition task is relatively fine and the positioning accuracy is required to be high, in order to reduce the image noise interference, a series of preprocessing operations need to be performed on the dial part image according to actual needs. The preprocessing operation includes four stages: image correction, noise reduction, grayscale / binary image conversion, and morphological operation.

[0060] Figure 3 Schematic diagram of the image preprocessing process in an embodiment of the present invention.

[0061] like Figure 3 As shown in the figure, image correction: Most original images cannot be taken directly facing the dial. The tilted angle will cause inaccurate center positioning and pointer-scale angle calculation, which will eventually lead to large reading errors. For three-dimensional space, perspective transformation is used to deal with the problem of instrument tilt and rotation. The perspective transformation matrix is ​​calculated by the least squares method through the four corner coordinates of the original image and the four corner coordinates of the target image, and then the tilted image is corrected to a square image.

[0062] Noise reduction: Choose traditional methods or deep learning methods based on the performance of edge devices. The traditional method uses Gaussian blur, which performs weighted averaging on each pixel in the image through the Gaussian function to blur the image, thereby achieving the effect of image smoothing and noise reduction; and the closer the pixel, the greater the impact on the current pixel. The deep learning method is a feature fusion module based on channel attention and pixel attention. By effectively retaining shallow feature information, it has a good defogging effect on images with uneven blur distribution, and significantly enhances the clarity and contrast of the image.

[0063] Convert to grayscale image / binary image: Most instrument images only need to determine the scale and pointer position according to the boundary to read, and do not need to use color information. The three channels of the color image are weighted averaged to obtain the grayscale image. Since the human eye is more sensitive to green, the formula used is Y=0.2989R+0.5870G+0.1140B. Furthermore, the grayscale image is converted to a binary image (i.e., black and white image) according to the set threshold, that is, the points with pixel values ​​greater than the threshold are set to white, and the points with pixel values ​​less than the threshold are set to black. In order to effectively deal with the problem of uneven brightness of the captured image, an adaptive threshold method based on the Gaussian function is adopted.

[0064] Morphological operation: Select morphological operations such as dilation and erosion as needed. Dilation increases the white area in the image, and erosion removes small noise or smoothes the edge of the object. This operation is implemented using a 3×3 structural element. If the structural element intersects with any pixel area in the image, the value of that position in the output image is white, otherwise the pixel value is not changed.

[0065] Reading part: Based on the instrument type and dial image, the corresponding algorithm is used to read the instrument to obtain the instrument reading result.

[0066] Specifically, the instrument types include round pointer instruments, square pointer instruments, lightning arresters, etc. The algorithm corresponding to the instrument type is determined through the configuration file.

[0067] In this embodiment, the instrument type is a pointer instrument. Based on the instrument type and the dial image, the reading part adopts an improved angle reading method for reading. The improved angle reading method includes: first, using an improved lightweight and directional (OBB) target detection model to locate the pointer and scale, and then, according to the polar coordinate relationship between the pointer and the scale and the center of the dial, the two scales closest to the pointer are selected; finally, the angle between the pointer and the two scales and the scale markings are used to calculate the actual reading to reduce the angle calculation error.

[0068] Figure 4 This is a class diagram for implementing the extensible reading method in the reading part design proposed in the present invention proposal.

[0069] like Figure 4 As shown in the figure, the reading part uses the abstract reading strategy class AbstractReadingStrategy, the concrete reading strategy class ConcreteReadingStrategy and the context class Context to implement extensible instrument reading in a strategy mode. Among them, the abstract reading strategy class is a strategy interface, which contains an abstract method read(), which declares the implementation paradigm of the method for reading instrument data.

[0070] Specifically, the AbstractReadingStrategy class includes:

[0071] read(self), abstract reading method, classes that implement this interface must implement this method;

[0072] The ConcreteReadingStrategy class includes:

[0073] read(self), specific reading method, implements corresponding reading logic according to the instrument type determined by the positioning part;

[0074] The Context class includes:

[0075] strategy, the reading strategy currently used;

[0076] reading_result, current reading result, including value or failure information;

[0077] set_strategy(self,strategy:ReadingStrategy), a method to dynamically change strategies;

[0078] get_reading_result(self), gets the reading according to the strategy;

[0079] The three classes have inheritance and composition relationships. The specific reading strategy class implements the reading method of the abstract strategy interface. A specific reading strategy should be implemented for each instrument type. The context class contains the current reading strategy. The reading module completes the instrument reading by maintaining the context class.

[0080] For example, for a round pointer meter, define a RoundPointerMeterStategy class to implement the AbstractReadingStrategy interface and the abstract method read(). First, run the target detection model inference for the pointer and scale. Then, according to the center position of the pointer and scale and the center position of the circle, select the two nearest scales on the left and right sides of the pointer according to the polar coordinate relationship. Finally, calculate the meter indication value according to the scale number and the angle ratio between the scale and the pointer. When the solution is running, use the reading context class Context to maintain the current reading strategy and the current reading result. After determining the reading strategy, call the current reading method through get_reading_result to read the meter and save the reading result in the reading_result variable. This variable is a structure with a Boolean variable indicating whether the reading is successful. If successful, read the reading result of the numeric type. If failed, read the error message of the string type. When the meter recognition category changes or the strategy needs to be actively changed, call the set_strategy method to dynamically change the reading strategy to achieve a flexible and scalable strategy mode. The input of the reading part is the dial part image cropped by the positioning part and the determined instrument type. The corresponding reading method is called according to the type and the reading result or error message is returned.

[0081] Interactive part: Provides a web user interface to users to display instrument reading results.

[0082] Specifically, the interactive part includes an input part and an output part. The input part includes two upload methods: pictures and cameras, as well as model architecture, image size, and confidence adjustment options; the output part includes instrument reading results and model annotation diagrams. The interactive part supports quick deployment and simple operation, making it convenient for users to check the reading process while hiding the implementation details.

[0083] Furthermore, the interaction part uses the Gradiu framework to build an end-to-end graphical user interface, and calls the instrument reading algorithm deployed in the backend through the network protocol. The front-end components of the graphical user interface include:

[0084] Image, upload images as local files or URLs;

[0085] Textbox, multi-line text input and output;

[0086] Slider, input value or range for parameter adjustment;

[0087] Dropdown, drop-down box option, used to select the model;

[0088] Interface, quickly create interactive applications, customize functions and their input and output;

[0089] Block / Row / Column, structural layout elements of web pages.

[0090] Those skilled in the art know that, in addition to realizing the system and its various devices, modules, and units provided by the present invention in a purely computer-readable program code, it is entirely possible to realize the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a hardware component, and the devices, modules, and units included therein for realizing various functions can also be regarded as structures within the hardware component; the devices, modules, and units for realizing various functions can also be regarded as both software modules for realizing the method and structures within the hardware component.

[0091] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. An intelligent industrial instrument recognition system based on deep learning, characterized in that: include: Positioning part: Adaptively select an appropriate target detection model according to the hardware conditions of the edge device, process the original instrument image through the target detection model, and obtain the dial image and instrument type; Reading part: based on the meter type and the dial image, a corresponding algorithm is used to read the meter to obtain a meter reading result; Interactive part: Provides a web user interface to the user to display the meter reading results.

2. According to claim 1, a deep learning-based intelligent industrial instrument identification system is characterized in that: The target detection model includes a rescalable backbone architecture, a hybrid convolution module and a defogging module.

3. According to claim 2, a deep learning-based intelligent industrial instrument identification system is characterized in that: According to the hardware conditions of the edge device, the positioning part adopts an adaptive model scale selection method to automatically select an adaptive target detection model.

4. According to claim 3, a deep learning-based intelligent industrial instrument identification system is characterized in that: The adaptive model scale selection method includes three stages: environmental assessment stage, reasoning fine-tuning stage and real-time monitoring stage. The environmental assessment phase: evaluate the hardware and software information of the edge device, measure the model calculation amount with floating-point operations according to the hardware performance and software compatibility of the device, and use a heuristic algorithm to select the appropriate model type and size; Inference fine-tuning phase: Use the selected model to perform small-scale inference attempts, collect runtime data including inference latency, device load, and instrument recognition effect, and fine-tune the model size based on performance and overhead; Real-time monitoring stage: Real-time performance monitoring is performed during the operation of the instrument identification system. The software and hardware conditions are checked at regular intervals. If the inference time is higher than the threshold or the device hardware load is too high, the inference is suspended and the model scale selection logic is re-executed.

5. The deep learning-based intelligent industrial instrument recognition system according to claim 4 is characterized in that: The positioning part uses an improved YOLOv10 model to locate the dial area and determine the instrument type from the original instrument picture, and according to the dial area, a partial image of the dial is captured from the original instrument picture, and then the partial image of the dial is preprocessed to obtain a dial image.

6. According to the deep learning-based intelligent industrial instrument recognition system of claim 1, it is characterized in that: Based on the instrument type and the dial image, the reading part is read by an improved angle reading method, which includes: first, using an improved lightweight and directional target detection model to locate the pointer and the scale, and then, according to the polar coordinate relationship between the pointer and the scale and the center of the dial, selecting the two scales closest to the pointer; finally, the angle between the pointer and the two scales and the scale markings is used to calculate the actual reading to reduce the angle calculation error.

7. The deep learning-based intelligent industrial instrument identification system according to claim 1 is characterized in that: The reading part implements extensible instrument reading in a strategy mode through an abstract reading strategy class AbstractReadingStrategy, a specific reading strategy class ConcreteReadingStrategy and a context class Context.

8. The deep learning-based intelligent industrial instrument recognition system according to claim 1 is characterized in that: The interactive part includes an input part and an output part. The input part includes two upload methods: picture and camera, as well as model architecture, image size, and confidence adjustment options; the output part includes instrument reading results and model annotation schematics.

9. The deep learning-based intelligent industrial instrument identification system according to claim 4 is characterized in that: The interactive part uses the Gradiu framework to construct an end-to-end graphical user interface and calls the instrument reading algorithm deployed in the back end through the network protocol.

10. The deep learning-based intelligent industrial instrument recognition system according to claim 4 is characterized in that: The front-end components of the graphical user interface include: Image, upload images as local files or URLs; Textbox, multi-line text input and output; Slider, input value or range for parameter adjustment; Dropdown, drop-down box option, used to select the model; Interface, quickly create interactive applications, customize functions and their input and output; Block / Row / Column, structural layout elements of web pages.

Citation Information

Patent Citations

  • A Deep Learning-Based Method for Automatic Reading of Pointer Instruments

    CN112949564B