Data processing method and device for instrument detection, equipment and storage medium
By automatically generating training data from power distribution room images and introducing an attention mechanism and a high-resolution detection head module, the problems of data scarcity and difficulty in identifying small targets in power distribution room scenarios are solved, and high-precision digital instrument detection is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 广州市扬新技术研究有限责任公司
- Filing Date
- 2025-12-16
- Publication Date
- 2026-05-01
AI Technical Summary
In the context of power distribution rooms, traditional manual inspections are inefficient and prone to errors. Deep learning target detection models are difficult to adapt to the scarcity of data and the difficulty in identifying small targets, resulting in insufficient detection accuracy and failing to meet the accuracy requirements of industrial-grade inspections.
By automatically generating digital characters and overlaying them onto images of real power distribution room backgrounds, a large-scale training dataset is constructed. An attention mechanism and a high-resolution detection head module are introduced on the YOLOv8 framework to improve the model's feature extraction capabilities and small object detection accuracy.
It achieves high-precision and robust detection of small-sized digital instruments in complex power distribution room environments, improves detection accuracy and model training efficiency, and provides a low-cost industrial vision inspection solution.
Smart Images

Figure CN121963231A_ABST
Abstract
Description
Data processing methods, devices, equipment, and storage media for instrument testing Technical Field
[0001] This application relates to the field of digital instrument testing technology, and in particular to a data processing method, apparatus, equipment and storage medium for instrument testing. Background Technology
[0002] In locations such as subways and substations, power distribution rooms contain numerous digital instruments, such as voltmeters and ammeters. The readings of these instruments are checked through inspections (e.g., manual inspections, video inspections) to determine the equipment's operational status. However, traditional manual inspection methods are inefficient and prone to errors. While automated reading recognition using deep learning object detection models is feasible, it faces challenges in terms of data and model adaptation to the power distribution room scenario. Specifically, in this high-risk area, personnel face difficulties in accessing the site to collect images, resulting in limited access to large-scale real-world data. The sample data obtained by related technologies is also limited and difficult to acquire, making it difficult to provide sufficient training samples for the models. Furthermore, when images are acquired through long-distance photography or video surveillance, the small display size of digital instruments makes it easy for deep learning object detection models to misdetect small targets, leading to insufficient detection accuracy. Therefore, deep learning-based automated reading methods in power distribution room scenarios face technical bottlenecks due to data scarcity and the difficulty in identifying small targets, resulting in insufficient detection accuracy and failing to meet the accuracy requirements of industrial-grade inspections. Summary of the Invention
[0003] This application provides a data processing method, apparatus, device, and storage medium for instrument detection, which solves the problem of insufficient detection accuracy of digital instruments in related technologies. This solution can generate a large amount of sample data based on the original image and train the target detection model with an attention mechanism to deploy a more accurate model on the detection device, thereby improving the detection accuracy of characters in digital instruments.
[0004] In a first aspect, this application provides a data processing method for instrument detection, comprising: upon acquiring a processing image of a corresponding instrument target in the detection location, adding a string with randomly determined font parameters within a preset parameter range to each processing image, and simultaneously performing annotation processing during string generation to generate sample data suitable for a pre-built target detection model; training the target detection model based on the sample data, so that the target detection model enhances the features of characters in the instrument target associated with the sample data through a convolutional attention module and locates and identifies characters in the instrument target of the sample data through multiple detection head modules corresponding to different resolution scales; and deploying the trained target detection model in a detection device so that the detection device can call the target detection model to perform character detection on the input image.
[0005] Secondly, this application also provides a data processing device for instrument detection, comprising: a sample generation module configured to, upon acquiring a processing image of the corresponding instrument target in the detection location, add a string with randomly determined font parameters within a preset parameter range to each processing image, and simultaneously perform annotation processing during string generation to generate sample data suitable for a pre-built target detection model; a model training module configured to train the target detection model based on the sample data, so that the target detection model enhances the features of characters in the instrument target associated with the sample data through a convolutional attention module and locates and identifies characters in the instrument target of the sample data through multiple detection head modules corresponding to different resolution scales; and a model deployment module configured to deploy the trained target detection model in the detection device, so that the detection device can call the target detection model to perform character detection on the input image.
[0006] Thirdly, this application also provides an electronic device comprising: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the data processing method for instrument detection of this application.
[0007] Fourthly, this application also provides a storage medium for storing computer-executable instructions, which, when executed by a processor, are used to perform the data processing method for instrument detection of this application.
[0008] This application's solution efficiently generates high-quality synthetic images and corresponding label text as sample data by overlaying randomly generated characters onto images. It also ensures the authenticity and diversity of the sample data, which is beneficial to improving the training efficiency and recognition accuracy of the target detection model. Furthermore, by introducing a higher resolution detection head module and a convolutional attention mechanism, this solution can achieve high-precision and robust detection of small-sized, low-contrast digital instruments in the complex environment of a power distribution room, effectively improving the detection accuracy of characters in digital instruments. Attached Figure Description
[0009] Figure 1 is a schematic diagram of the steps of a data processing method for instrument detection provided in an embodiment of this application.
[0010] Figure 2 is a schematic diagram of the steps for generating an image to be labeled according to an embodiment of this application.
[0011] Figure 3 is a schematic diagram of the generated image to be labeled provided in an embodiment of this application.
[0012] Figure 4 is a schematic diagram of the target detection model provided in an embodiment of this application.
[0013] Figure 5 is a schematic diagram of the recognition effect of the target detection model shown in Figure 4.
[0014] Figure 6 is a schematic diagram of the attention mechanism for a model provided in an embodiment of this application.
[0015] Figure 7 is a schematic diagram of the steps for deploying a target detection model according to an embodiment of this application.
[0016] Figure 8 is a schematic diagram of the structure of a data processing device for instrument detection provided in an embodiment of this application.
[0017] Figure 9 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0018] The embodiments of this application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely illustrative of the embodiments of this application and are not intended to limit the scope of this application. Furthermore, it should be noted that, for ease of description, the accompanying drawings only show the parts related to the embodiments of this application, not all structures. Those skilled in the art, after reading this specification, should be able to conceive that any combination of technical features can constitute an optional implementation method, provided that the technical features do not contradict each other.
[0019] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects, not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, not limited in number; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship. In the description of this application, "multiple" means two or more, and "several" means one or more.
[0020] Numerous digital instruments, such as voltmeters and ammeters, are installed in power distribution rooms in locations like subways and substations. Data from these instruments is read through routine inspections to determine the equipment's operational status. For these locations, manual inspections, where personnel check each instrument individually, are labor-intensive, inefficient, and prone to errors. In computer vision, particularly OCR (Optical Character Recognition) and object detection tasks, model training heavily relies on large amounts of high-quality, accurately labeled image data. However, in industrial automation scenarios, acquiring digital instrument image data is difficult. Traditional dataset construction processes involve on-site photography, data cleaning, image enhancement, and manual annotation—a cumbersome and inefficient process. While some technologies utilize deep learning and reinforcement learning for automatic annotation, these methods still require complex deep learning and reinforcement learning algorithms and only address the annotation stage. Furthermore, the solution of using deep learning detection models for automated reading recognition to achieve inspection is difficult to meet the needs of power distribution room scenarios. It's conceivable that in such high-risk areas, it's difficult for personnel to enter the site to collect images, making it hard to obtain large-scale, real-world data. Moreover, if images are acquired through long-distance photography or surveillance video, the small display size of digital instruments makes traditional target detection algorithms prone to false detections, resulting in insufficient detection accuracy. Therefore, traditional deep learning algorithms lack the accuracy for detecting digital instruments and are prone to false detections. In power distribution room scenarios, deep learning-based automated reading methods face technical bottlenecks such as data scarcity and difficulty in identifying small targets, leading to insufficient detection accuracy and failing to meet the accuracy requirements of industrial-grade inspections.
[0021] To address the aforementioned challenges of obtaining large-scale training datasets and the low detection rate of digital instruments in complex power distribution room environments using traditional object detection algorithms, this application provides an algorithm for generating large-scale training datasets based on real-world scenarios, along with a high-precision, low-latency digital instrument target detection scheme for instrument detection. The data processing method for instrument detection provided in this application is based on a procedural image generation mechanism. It automatically generates digital characters and overlays them onto images of the real power distribution room background, automatically generating digital instrument reading images and their corresponding target locations, bounding boxes, and category annotation files as sample data. Furthermore, this scheme introduces an attention mechanism on the YOLOv8 framework to enhance feature extraction capabilities in low-light environments, replaces the backbone feature network to reduce model size and improve inference speed, and adds a detection head to improve the detection accuracy of small targets at long distances. The method provided in this application not only achieves efficient automatic dataset generation but also achieves high-precision recognition of digital instrument readings in complex power distribution room scenarios, providing a scalable and low-cost solution for industrial vision inspection and intelligent patrol.
[0022] This method can be applied to electronic devices such as servers and host computers. This solution can generate high-quality sample data based on the acquired images and improve the detection accuracy of the target detection model to adapt to the detection of characters in digital instruments in complex industrial scenarios. This allows for the deployment of a model trained on the sample data on the detection equipment to perform character detection on the digital instruments. Figure 1 is a schematic diagram of the steps of a data processing method for instrument detection provided in an embodiment of this application. As shown in the figure, after acquiring the image to be processed corresponding to the instrument target, a large number of training samples are constructed by overlaying strings on the image to be processed. This achieves both the diversification and realism of the sample data, and also enables high-precision and robust detection of small-sized, low-contrast digital instruments in complex environments such as power distribution rooms in the detection site, thereby recognizing the characters in the digital instruments. The specific steps are as follows, including steps S110-S130.
[0023] Step S110: After obtaining the image to be processed of the corresponding instrument target in the area to be detected, add a string with randomly determined font parameters within a preset parameter range to each image to be processed, and perform annotation processing simultaneously when generating the string to generate sample data suitable for the pre-built target detection model.
[0024] The images to be processed for the corresponding instrument targets within the testing location can be actual images from a power distribution room in a location such as a subway or substation. These images display digital instruments that serve as the instrument targets, such as those on a distribution cabinet or a testing instrument. It is conceivable that the number of images acquired for processing is less than the number required as sample data. Therefore, by overlaying different strings onto the images to obtain different images, the data scale can be expanded to meet the quantity requirements for use as sample data.
[0025] Furthermore, a preset parameter range is configured for the font parameters, allowing for the random determination of specific font parameters to add corresponding strings to the image to be processed. Optionally, if the characters displayed on the digital instrument use a digital tube font, the generated string will also use a digital tube font, and the randomly determined font parameters correspond to parameters such as font size, color, and position, thus generating the string to be added. Optionally, the determined parameters can also be flip angle, distortion parameters, etc., where distortion parameters are used to describe the distortion caused by camera stillness, including radial distortion parameters and tangential distortion parameters, to simulate the angle and image content deformation caused by the lens when shooting images with real lenses through flip and distortion processing. It is conceivable that, in order to adapt to more types of digital instruments, strings of different font types can be randomly generated and added to the image to be processed, in order to generate richer sample data and help improve the detection accuracy of the model. Thus, by using a small number of images to be processed, after the above processing, the data scale can be expanded, thereby obtaining a large set of image data to enrich the sample data.
[0026] In computer vision, data annotation refers to the labeling and annotation of image or video data to provide information needed for training machine learning models. These labels can include identifying objects in an image, locating their positions, and describing their attributes. Data annotation plays a crucial role in computer vision, serving as the foundation for model learning and inference, and is essential for improving model accuracy and performance. Through data annotation, object detection algorithms can be provided with rich training samples, enabling them to learn features, patterns, and relationships within images. The resulting annotation information helps the model better understand image content, thus achieving more accurate object detection. Annotation processing is performed simultaneously when generating strings, such as labeling instrument targets and characters in images to highlight areas of greater interest. The generated sample data is adapted to pre-built object detection models. For example, the sample data may include bounding box annotation information. Bounding box annotation is a commonly used method in image data annotation, used to mark the location of target objects in an image. Furthermore, by drawing rectangles on the image, the boundary range of the target object is determined, and a category label is assigned to it to generate the corresponding bounding box annotation information.
[0027] Optionally, after generating sample images, simulation processing can be performed on the images, such as reducing brightness to simulate a low-light environment or adding highlights to simulate reflections in the environment. This adds interference from the real environment to the images, simulating image degradation in the real environment, which helps to improve the robustness of the target detection model.
[0028] Step S120: Based on the sample data, train the target detection model so that the target detection model enhances the features of the characters in the instrument targets associated with the sample data through the convolutional attention module and locates and identifies the characters in the instrument targets in the sample data through multiple detection head modules corresponding to different resolution scales.
[0029] As a neural network model used to detect characters within instrument targets in images, the target detection model integrates a convolutional attention module and multiple detection head modules. These detection heads correspond to different resolution scales and can perform multi-level segmentation of the feature maps generated by convolutional operations on sample images. This allows for more effective localization and recognition of character regions of different sizes, improving the detection coverage of small-sized instrument targets. In this regard, our target detection model introduces a higher-resolution detection head module and a convolutional attention mechanism to segment small-sized instrument targets and enhance the features associated with characters within the instrument targets, thereby better recognizing the characters. In the model structure, the detection head module utilizes higher-resolution shallow feature maps to better locate and recognize small-sized instrument targets, thus improving the detection capability for small-sized digital instruments (e.g., those occupying only tens of pixels). Furthermore, a convolutional attention module is inserted after the feature extraction and fusion module output and before the input detection head module. Furthermore, in applications involving digital instrument detection in power distribution rooms, the model focuses on the LCD / digital tube display area of the instrument, automatically suppressing interference from complex backgrounds such as power distribution cabinets, cables, and metal casings. This enhances the model's ability to extract key features (i.e., characters within the instrument target) under complex lighting conditions, such as low light. To this end, after acquiring sample data automatically generated in the above manner, the model is trained on a pre-built target detection model to adjust its parameters, enabling it to more accurately detect and recognize characters on instrument targets in images.
[0030] Step S130: Deploy the trained target detection model in the detection device so that the detection device can call the target detection model to perform character detection on the input image.
[0031] The detection equipment and electronic equipment are connected via a network or line. After the electronic equipment completes model training, the model is deployed, and data is then sent to the detection equipment via the network or line to load the target detection model. The detection equipment then uses the target detection model to perform character detection on the input image (such as image data captured by a live camera). Optionally, the model is also adaptively adjusted to be compatible with the detection equipment, thereby avoiding situations where the detection equipment struggles to stably run the target detection model for character detection.
[0032] Therefore, this solution efficiently generates high-quality synthetic images and corresponding label text as sample data by overlaying randomly generated characters onto images. It also ensures the authenticity and diversity of the sample data, which is beneficial to improving the training efficiency and recognition accuracy of the target detection model. Furthermore, by introducing a higher resolution detection head module and a convolutional attention mechanism, this solution can achieve high-precision and robust detection of small-sized, low-contrast digital instruments in the complex environment of a power distribution room, effectively improving the detection accuracy of characters in digital instruments.
[0033] Figure 2 is a schematic diagram of the steps for generating an image to be labeled according to an embodiment of this application. In one embodiment, during the process of generating the image to be labeled, a string is added to the image to be processed, such as overlaying the generated string as an image onto the image to be processed. Thus, different images can be obtained by adding different strings to the same image to be processed and used as images to be labeled. The specific steps are as follows: Step S210: Randomly generate the target text to be added based on a preset generation algorithm.
[0034] Step S220: Based on the size and color determined from the font parameters, adjust the characters in the target text to serve as the first string to be added.
[0035] Step S230: Based on the position determined from the font parameters, add the first string to the image to be processed to obtain the image to be labeled.
[0036] Understandably, target text is randomly generated, consisting of several characters such as numbers and decimal points to form a numerical value. Font parameters include size, color, and position, each with a corresponding parameter range. The size is related to the font size; a value is randomly selected from the configured parameter range as the size of the character in the target text. Similarly, the color is selected based on the corresponding parameter range, thus adjusting the size and color of the characters in the target text to obtain the first string. Correspondingly, the position configured in the font parameters has a parameter range related to the size of the image to be processed. The first string is then added to the image according to the position determined by the font parameters, for example, using the coordinates corresponding to the position determined by the font parameters as the vertex coordinates of the first string. It is conceivable that during the process of adding the first string to the image according to the position determined by the font parameters, the vertex coordinates corresponding to the area occupied by the first string are used as a criterion. The vertex coordinates corresponding to the area occupied by the first string are compared with the size range of the image to determine whether the first string exceeds the image size range, and readjustment is performed if the first string exceeds the image size range. Optionally, multiple strings can be added to the image to be processed. To this end, after adjusting the characters in the generated target texts, when adding multiple target texts to the image, positional constraints are applied to the different target texts. For example, by determining whether different first strings overlap, the position of at least one of the first strings is readjusted if overlap exists. It is conceivable that for the same image to be processed, different strings can be added to obtain different labeled images—that is, these images have the same background but display different text—to enrich the displayed characters in the sample data, making the object detection model focus more on the characters in the instrument targets. By controlling the randomness of color, font size, and position, this scheme can generate large-scale, diverse, and uniformly formatted high-quality training datasets, effectively reducing the manual cost and time consumption of data preparation and helping to improve the generalization performance of model training.
[0037] As shown in Figure 3, which is a schematic diagram of the generated image to be labeled according to an embodiment of this application, the figure shows the image to be processed and the generated image to be labeled. For the same image to be processed, several strings (in the figure, numerical characters and decimal points are used to form the corresponding strings) can be added to the image to generate the image to be labeled. Moreover, as shown in the figure, any number of strings can be added to the image, and the added strings are random in terms of font size, color, and position. Furthermore, different strings can be added to the same image to be processed to generate different images to be labeled, that is, images with the same background but different added characters. In the figure, ellipses indicate that other images to be labeled can be generated based on the same image to be processed, and the added strings are shown in dashed boxes. It is worth noting that the dashed boxes are used to highlight and more clearly demonstrate the effect of this embodiment, and are not added in actual applications.
[0038] Therefore, this solution can efficiently generate high-quality synthetic instrument reading images and corresponding label text by superimposing randomly generated different characters on real images. This not only ensures the authenticity and diversity of the data, but also significantly improves the training efficiency and recognition accuracy of the target detection model.
[0039] Optionally, characters in the target text can be randomly selected and generated in a loop. For example, when generating text corresponding to a numerical value, a preset number of loops can be set, and a decimal point can be added to the text in one of the loops, so that a complete numerical value is formed after the preset number of loops ends. Specifically, when in the process of a preset number of loops, the number of loops is accumulated and it is determined whether the number of loops is a preset number threshold. It can be understood that within one loop cycle of the preset number of loops, there is a parameter value used to accumulate the number of loops, and by judging whether the currently accumulated number of loops has reached the preset number threshold, it can be conceivable that the loop cycle corresponding to the preset number threshold is used to configure the decimal point. To this end, a numerical value can be randomly selected within the range of the preset number as the preset number threshold, and then a decimal point can be configured in the string when entering the loop cycle corresponding to the preset number threshold, while adding numeric characters in other loop cycles.
[0040] Furthermore, by determining whether the number of iterations is within a preset threshold, the action to be performed in the current iteration is determined. Specifically, if the number of iterations is not within the preset threshold, a digit character is randomly selected from a randomly generated list of digit characters and added to the second string. If the number of iterations is within the preset threshold, the target character with the corresponding decimal point is added to the second string. It's understandable that the digit characters in the list can be randomly generated, and their order can be randomized. Alternatively, the digits corresponding to the digit characters can be arranged according to their numerical value. In other iterations, a randomly selected digit character is added to the second string. It's conceivable that if a corresponding digit character was added in the previous iteration, the digit character to be added in the current iteration is set after the digit character added in the previous iteration. The target character is added in the same way as digit characters. It's worth noting that the second string serves as an intermediate variable set throughout the loop, with corresponding characters added in each iteration; that is, as the number of iterations accumulates, the content of the second string also changes.
[0041] Furthermore, upon exiting the loop, it checks whether the second string conforms to a numerical format. If the second string conforms, it is used as the target text. Understandably, after generating the corresponding second string through a preset number of loops, it's necessary to further determine whether the numerical value formed by the second string conforms to a numerical format. This involves checking whether the decimal point position meets the format requirements, such as having digits before and after the decimal point. By passing this format check, the second string that conforms to the numerical format is used as the target text. To address this, this solution uses random generation of the target text, which helps improve the authenticity and diversity of the data, allowing for the use of diverse sample data to improve the training efficiency and recognition accuracy of the object detection model.
[0042] In this scheme, the generated sample data includes images and label text. In one embodiment, the corresponding label text is generated simultaneously when generating numeric strings. The label text represents the string added to the image and its location. Therefore, the label text corresponding to the image can be generated simultaneously when the string is generated, and the label text corresponds one-to-one with the generated image. Specifically, during the generation of the current string, the bounding box coordinates of each character in the string are extracted. The bounding box coordinates correspond to the coordinates of each vertex within the bounding box, used to determine the character's position in the image. Optionally, a minimum bounding rectangle matching the character size is preset as the bounding box. Then, the category index and bounding box coordinates of each character in the string are concatenated to generate the target format label text. The category index is associated with the character's category. Optionally, for numeric characters, the category index can be represented numerically. Referring to Figure 3, if the string "551.93" exists in the image to be labeled, the generated label text records the category index "5" of the first character "5" and its bounding box coordinates. Then, the information for the second character "5" is recorded in another line, thus recording all the characters of the string and their corresponding positions in the label text. Optionally, the bounding box coordinates recorded in the label text can be normalized coordinates. This solution generates the label text synchronously during the string generation process; that is, it provides the corresponding label text for each image in the sample data simultaneously, forming a complete sample data set. Therefore, this method can save a significant amount of annotation work, achieving more efficient sample generation while ensuring the accuracy of the sample data.
[0043] Figure 4 is a schematic diagram of the target detection model provided in an embodiment of this application. As shown in the figure, the target detection model includes a backbone network, a neck network, and a multi-scale detection head. Specifically, the backbone network includes at least four downsampling convolutional layers. Each downsampling convolutional layer consists of a convolutional module (Conv), a C2f feature extraction and fusion module (C2f), and a convolutional attention module (CBAM). The convolutional module is used to initially extract the spatial features of the input image; the C2f feature extraction and fusion module is based on cross-stage partial connections and effectively captures multi-scale semantic information through multi-branch depthwise separable convolution; the convolutional attention module applies attention mechanisms in the channel dimension and spatial dimension respectively, adaptively enhancing the feature response sensitive to the instrument character region.
[0044] In the neck network portion, the model employs a feature pyramid structure. An upsampling module restores the spatial resolution of high-level semantic feature maps, and a concatenation module fuses these features with low-level features from the corresponding layers of the backbone network along the channel dimension, thereby constructing a multi-scale feature representation with rich semantics and high spatial accuracy. An optional ground sampling module can be implemented using bilinear interpolation.
[0045] Furthermore, the model incorporates a pooling module (SPPF, Spatial Pyramid Pooling-Fast) deep within the backbone network as a core component for multi-scale contextual modeling. This module, through sequential max-pooling operations, aggregates contextual information from different receptive fields without significantly increasing computational overhead, effectively improving the model's robustness to instrument targets of varying sizes while balancing detection accuracy and inference efficiency.
[0046] Ultimately, multiple detector heads are applied to fused feature maps at different scales, and lightweight convolutional layers are used to perform bounding box regression and character category prediction. Each detector head corresponds to a different spatial resolution, enabling collaborative coverage of instrument character targets ranging from large to small, with a particularly enhanced ability to detect small characters.
[0047] Referring to Figure 4, the target detection model includes four detection head modules, each corresponding to a feature map at a different scale to detect targets of varying sizes. Furthermore, a convolutional attention module is inserted after each feature extraction and fusion module's output and before it enters the detection head module. The convolutional attention module helps the model better focus on key areas of the instrument target (such as numbers and pointers in the image) and suppress background noise before making the final prediction. Moreover, in scenarios like power distribution rooms where there are small instruments, reflections, and low light conditions, the convolutional attention module can amplify these subtle but important features, significantly improving detection accuracy. Figure 5 is a schematic diagram of the recognition effect of the target detection model shown in Figure 4. As shown in Figure 5, the figure shows the original image 101, the intermediate feature map of the model 102, and the model detection result 103. During the character detection process of the target detection model on the original image 101, the feature map obtained in the intermediate processing is shown in the intermediate feature map 102 of the model in Figure 4. Referring to Figure 4, it can be seen from left to right that the shallow layer of the model can only capture limited low-level features (such as edges and textures). However, as the depth of the model network increases, the layer-by-layer abstraction makes the high-level semantic features gradually richer and the representation ability significantly enhanced, as shown in the model detection result 103 of Figure 4. Finally, it can accurately locate and recognize the characters on the instrument.
[0048] Optionally, Figure 6 is a schematic diagram of the attention mechanism added to the model according to an embodiment of this application. In one embodiment, the convolutional attention module includes a channel attention module 201 and a spatial attention module 202. The channel attention module 201 is used to configure channel weights to select feature channels that are robust to the current illumination. Illumination robustness is a key performance indicator in the fields of object detection and image recognition, which is used to indicate that the model can still maintain stable performance (such as detection accuracy and recognition accuracy) when illumination conditions change drastically. The spatial attention module 202 is used to configure spatial weights to suppress non-digit regions in the image. As shown in the figure, the input feature map first enters the channel attention module 201 for processing. The processed result is fused with the original input feature map. The fused data then enters the spatial attention module 202 for processing. After another fusion operation, the final output feature map is obtained.
[0049] In response, some channels in the feature map respond to the contour edges of the detected target and are independent of illumination; these channels can serve as illumination-robust feature channels. Understandably, in digital instrument detection, certain spatial regions are more important, such as the LCD (Liquid Crystal Display) area and LED (Light-Emitting Diode) area, which can be used as regions of interest. In the convolutional attention module, channel attention module 201 learns channel weights and spatial attention module 202 learns spatial weights, achieving adaptive feature recalibration. It is conceivable that metallic reflections and border textures interfere with digital instrument detection; spatial attention module 202 can suppress non-digital areas and focus more on the LCD / LED screen position. Furthermore, overexposure, underexposure, and shadows also interfere with digital instrument detection; channel attention module 201 can select feature channels robust to the current illumination, such as those corresponding to edge positions and high-brightness areas. Therefore, this scheme, by introducing an attention mechanism into the model, better focuses on the regions of interest in the feature map, resulting in richer extracted features.
[0050] Figure 7 is a schematic diagram of the steps for deploying a target detection model according to an embodiment of this application. In one embodiment, in order to enable the target detection model to run in real time on the detection device, for example, the detection device is a low-power industrial control computer in a power distribution cabinet, in order to avoid the situation where the industrial control computer cannot run the target detection model for a long time due to its small storage capacity and limited computing power, the model can be adjusted through model pruning during the deployment process to reduce unimportant channels, thereby ensuring the stability of the model operation. The specific steps are as follows: Step S310: Determine the importance score of the convolution kernel corresponding to each convolutional layer in the target detection model and the score threshold corresponding to the preset pruning rate. The preset pruning rate is associated with the detection device.
[0051] Step S320: Select channels with importance scores less than the score threshold as target channels.
[0052] Step S330: Set the weights of the corresponding target channels in the target detection model to zero to update the target detection model, and deploy the updated target detection model on the detection device.
[0053] Understandably, different types of detection devices have pre-set pruning rates. When pruning the model, an importance score is determined, representing the importance of the input channels. For example, the importance of input channels can be scored based on the L1 norm of the input channel weights in each convolutional layer. The score threshold corresponding to the pre-set pruning rate is then compared with the calculated importance score. Channels with importance scores less than the threshold are selected as target channels. It's conceivable that the target channels, being of lower importance, contribute little to feature extraction and may even be redundant. To better adapt to the detection device, these target channels are pruned, meaning their weights in the target detection model are reset to zero. By resetting the channel weights to zero, the target channels are removed, and the model is updated. This updated model removes channels with low contribution and retains channels that effectively extract features, reducing model size while avoiding a significant decrease in detection accuracy.
[0054] Optionally, in determining the importance score of each convolutional kernel in the object detection model, L1 regularization can be used to score channel importance. For example, the L1 norm of the weights corresponding to all input channels and all convolutional positions of each convolutional kernel in the same convolutional layer can be determined and used as the importance score. The corresponding calculation formula is as follows:
[0055] in, Let be the weights of the j-th convolutional kernel in layer l at input channel i and convolution position (k1, k2). Let K represent the total number of input channels and K represent the kernel size. This can be understood as performing an L1 norm summation by calculating all weights of a single convolutional kernel, specifically by iterating through input channels i, i.e., from 1 to... This process covers all input channels corresponding to the convolutional kernel and traverses the spatial positions (k1, k2) of the kernel, i.e., from k1=1 to K and k2=1 to K, covering all spatial pixels of the kernel, and sums the absolute values. From this, we can determine the sum of the absolute values of the weights of the j-th convolutional kernel across all input channels and all spatial positions, which is the importance score of the convolutional kernel. .
[0056] Then, a preset pruning rate is used as the quantile ratio, and the score value of the corresponding quantile is determined from the importance scores of all convolutional kernels in the same convolutional layer as the scoring threshold. Specifically, the Quantile function can be used to determine this scoring threshold. The Quantile function is a quantile function that, given a dataset and a quantile ratio, returns the value of the split point in the dataset corresponding to that ratio, where the proportion of data with a value less than that split point is equal to the quantile ratio. The corresponding formula is as follows:
[0057] Where p is the preset pruning rate, which is a value between 0 and 1. Let τ represent the dataset corresponding to all importance scores, and τ be the scoring threshold. To achieve this, quantiles are calculated, and a score is selected from all importance scores as the scoring threshold. The target channel is then determined by comparing the importance scores with this threshold. Furthermore, this solution employs model pruning to enable the detection equipment to run the model efficiently in real-time, thus meeting the inspection requirements.
[0058] Figure 8 is a schematic diagram of the structure of a data processing device for instrument testing provided in an embodiment of this application. The device is used to execute the data processing method for instrument testing provided in the above embodiment, and the device also has a functional module for executing the method and beneficial effects. As shown in the figure, the data processing device for instrument testing includes a sample generation module 401, a model training module 402 and a model deployment module 403.
[0059] The sample generation module 401 is configured to, upon acquiring the image to be processed of the corresponding instrument target in the area to be detected, add a string with randomly determined font parameters within a preset parameter range to each image to be processed, and simultaneously perform annotation processing during string generation to generate sample data suitable for the pre-built target detection model; the model training module 402 is configured to train the target detection model based on the sample data, so that the target detection model enhances the features of characters in the instrument target associated with the sample data through the convolutional attention module, and locates and recognizes the characters in the instrument target of the sample data through multiple detection head modules corresponding to different resolution scales; the model deployment module 403 is configured to deploy the trained target detection model in the detection device, so that the detection device can call the target detection model to perform character detection on the input image.
[0060] Based on the above embodiments, the sample generation module 401 is specifically configured as follows: randomly generating target text to be added based on a preset generation algorithm; adjusting the characters in the target text to serve as the first string to be added based on the size and color determined from the font parameters; and adding the first string to the image to be processed based on the position determined from the font parameters to obtain the image to be labeled.
[0061] Based on the above embodiments, the image generation module 401 is further configured to: when in a loop process with a preset number of iterations, accumulate the number of iterations and determine whether the number of iterations is a preset number threshold; if the number of iterations is not a preset number threshold, randomly select a numeric character from a randomly generated list of numeric characters to add to the second string; if the number of iterations is a preset number threshold, add a target character with a corresponding decimal point to the second string; when exiting the loop, determine whether the second string conforms to a numerical format, so that if the second string conforms to a numerical format, the second string is used as the target text.
[0062] Based on the above embodiments, the sample generation module 401 is further configured to: extract the bounding box coordinates of each character in the string during the process of generating the current string; and concatenate the category index and bounding box coordinates of each character in the string to generate target format label text, wherein the category index is associated with the category of the character.
[0063] Based on the above embodiments, the target detection model includes a backbone network, a neck network, and multiple detection heads at different scales. The backbone network includes at least four downsampling convolutional layers. Each downsampling convolutional layer includes a convolution module, a C2f feature extraction and fusion module, and a convolutional attention module. The convolution module is used to initially extract spatial features of the input image. The C2f feature extraction and fusion module is used to obtain multi-scale semantic information through multi-branch depthwise separable convolution. The convolutional attention module is used to apply an attention mechanism. The neck network includes an upsampling module and a stitching module. The upsampling module is used to restore the spatial resolution of the high-level semantic feature map. The stitching module is used to fuse the high-level semantic features with the low-level features of the backbone network in the channel dimension to obtain a fused feature map. The multiple detection heads are applied to the fused feature map at different scales.
[0064] Based on the above embodiments, the convolutional attention module includes a spatial attention module and a channel attention module. The spatial attention module is used to configure spatial weights to suppress non-digital regions in the image, and the channel attention module is used to configure channel weights to select feature channels that are robust to the current illumination.
[0065] Based on the above embodiments, the model deployment module 403 is specifically configured as follows: determining the importance score of the convolution kernel corresponding to each convolutional layer in the target detection model and the score threshold corresponding to the preset pruning rate, wherein the preset pruning rate is associated with the detection device; selecting channels with importance scores less than the score threshold as target channels; resetting the weights of the corresponding target channels in the target detection model to zero to update the target detection model, and deploying the updated target detection model on the detection device.
[0066] Based on the above embodiments, the model deployment module 403 is further configured to: determine the L1 norm of the weights corresponding to all input channels and all convolution positions of each convolution kernel in the same convolutional layer, and use it as an importance score; use the preset pruning rate as the quantile ratio, and determine the score value of the corresponding quantile in the importance scores of all convolution kernels in the same convolutional layer as a score threshold.
[0067] It is worth noting that in the embodiments of the above-mentioned device, the modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each module are only for easy differentiation and are not used to limit the protection scope of the embodiments of this application.
[0068] Figure 9 is a schematic diagram of an electronic device provided in an embodiment of this application. This device is used to execute the data processing method for instrument detection provided in the above embodiments, and has corresponding functional modules and beneficial effects for executing the method. As shown in the figure, the device includes a processor 501, a memory 502, an input device 503, and an output device 504. The number of processors 501 can be one or more; one processor 501 is shown as an example in the figure. The processor 501, memory 502, input device 503, and output device 504 can be connected via a bus or other means; a bus connection is shown as an example in the figure. The memory 502, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the data processing method for instrument detection in the embodiments of this application. The processor 501 executes various corresponding functional applications and data processing by running the software programs, instructions, and modules stored in the memory 502, thereby realizing the above-mentioned data processing method for instrument detection.
[0069] The memory 502 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data recorded or created during use. Furthermore, the memory 502 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 502 may further include memory remotely configured relative to the processor 501, which can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof. The input device 503 can be used to input corresponding numerical or character information to the processor 501 and to generate key signal inputs related to user settings and function control of the device; the output device 504 can be used to send or display key signal outputs related to user settings and function control of the device.
[0070] This application also provides a storage medium storing computer-executable instructions, which, when executed by a processor, are used to perform relevant operations in the data processing method for instrument detection provided in any embodiment of this application.
[0071] Computer-readable storage media include both permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0072] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0073] Note that the above description is merely a preferred embodiment and the technical principles employed in this application. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments. Many other equivalent embodiments may be included without departing from the concept of this application, and the scope of this application is determined by the scope of the appended claims.
Claims
1. A data processing method for instrument testing, characterized in that, include: When the image to be processed of the corresponding instrument target in the area to be detected is obtained, a string with font parameters randomly determined within a preset parameter range is added to each image to be processed, and annotation processing is performed simultaneously when the string is generated to generate sample data suitable for the pre-built target detection model. Based on the sample data, the target detection model is trained so that the target detection model enhances the features of the characters in the instrument targets associated with the sample data through the convolutional attention module and locates and identifies the characters in the instrument targets of the sample data through multiple detection head modules corresponding to different resolution scales. The trained target detection model is deployed in the detection device so that the detection device can call the target detection model to perform character detection on the input image.
2. The data processing method for instrument testing according to claim 1, characterized in that, The step of adding a string with randomly determined font parameters within a preset parameter range to each of the images to be processed includes: randomly generating target text to be added based on a preset generation algorithm; adjusting the characters in the target text to serve as a first string to be added based on the size and color determined from the font parameters; and adding the first string to the image to be processed based on the position determined from the font parameters to obtain an image to be labeled.
3. The data processing method for instrument testing according to claim 2, characterized in that, The method of randomly generating target text to be added based on a preset generation algorithm includes: when in a loop process of a preset number of iterations, accumulating the number of iterations and determining whether the number of iterations is a preset threshold; if the number of iterations is not the preset threshold, randomly selecting a numeric character from a randomly generated list of numeric characters to add to the second string; if the number of iterations is the preset threshold, adding a target character with a corresponding decimal point to the second string; and when exiting the loop, determining whether the second string conforms to a numerical format, so that if the second string conforms to a numerical format, the second string is used as the target text.
4. The data processing method for instrument testing according to claim 1 or 2, characterized in that, The step of simultaneously performing annotation processing during the generation of the string to generate sample data suitable for the pre-built target detection model includes: extracting the bounding box coordinates of each character in the string during the generation of the current string; concatenating the category index of each character in the string with the bounding box coordinates to generate target format label text, wherein the category index is associated with the category of the character.
5. The data processing method for instrument testing according to claim 1 or 2, characterized in that, The target detection model includes a backbone network, a neck network, and multiple detection heads at different scales. The backbone network includes at least four downsampling convolutional layers. Each downsampling convolutional layer includes a convolution module, a C2f feature extraction and fusion module, and a convolutional attention module. The convolution module is used to initially extract spatial features from the input image. The C2f feature extraction and fusion module is used to obtain multi-scale semantic information through multi-branch depthwise separable convolution. The convolutional attention module is used to apply an attention mechanism. The neck network includes an upsampling module and a stitching module. The upsampling module is used to restore the spatial resolution of the high-level semantic feature map. The stitching module is used to fuse the high-level semantic features with the low-level features of the backbone network in the channel dimension to obtain a fused feature map. The multiple detection heads are applied to the fused feature map at different scales.
6. The data processing method for instrument testing according to claim 5, characterized in that, The convolutional attention module includes a spatial attention module and a channel attention module. The spatial attention module is used to configure spatial weights to suppress non-digital regions in the image, and the channel attention module is used to configure channel weights to select feature channels that are robust to the current illumination.
7. The data processing method for instrument testing according to claim 1 or 2, characterized in that, The step of deploying the trained target detection model in the detection device, so that the detection device can call the target detection model to perform character detection on the input image, includes: determining the importance score of the convolution kernel corresponding to each convolutional layer in the target detection model and the score threshold corresponding to the preset pruning rate, wherein the preset pruning rate is associated with the detection device; selecting the channel with an importance score less than the score threshold as the target channel; resetting the weight of the corresponding target channel in the target detection model to zero to update the target detection model, and deploying the updated target detection model in the detection device.
8. The data processing method for instrument testing according to claim 7, characterized in that, The step of determining the importance score of each convolutional kernel in the target detection model and the score threshold corresponding to the preset pruning rate includes: determining the L1 norm of the weights of each convolutional kernel in the same convolutional layer across all input channels and all convolutional positions, and using it as the importance score; using the preset pruning rate as the quantile ratio, and determining the score value corresponding to the quantile among the importance scores of all convolutional kernels in the same convolutional layer as the score threshold.
9. An electronic device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the data processing method for instrument detection as described in any one of claims 1-8.
10. A storage medium for storing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a processor, are used to perform the data processing method for instrument detection as described in any one of claims 1-8.