Instrument identification method, computer device, storage medium and computer program product

CN122676489APending Publication Date: 2026-09-01YANTAI RAYTRON TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610998264.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

这种非正面姿态会带来两个问题:一是刻度几何畸变,原本均匀的刻度在图像中变得疏密不均,导致指针指向与真实数值之间的映射关系被破坏

Benefits of technology

[0008] Thirdly, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the instrument identification method described in the above embodiments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122676489A_ABST
    Figure CN122676489A_ABST
Patent Text Reader

Abstract

The application provides an instrument identification method, computer equipment, a storage medium and a computer program product. The method first crops the instrument region in the inspection image, removes background interference, and obtains an image highlighting the instrument. Then, the cropped image is input into an end-to-end correction model to directly generate a front view image through a single forward propagation, avoiding error accumulation of traditional multi-stage methods and improving the robustness of the correction from the architecture. In the model training stage, the reading result of the front view image output by the correction model is taken as a supervision signal by a multi-modal large language model, which on the one hand distills the geometric and semantic understanding ability of the large model to the lightweight model, and on the other hand makes the correction model directly optimize the reading accuracy instead of only focusing on image similarity. This training strategy significantly improves the performance of the model under complex working conditions such as inclination, occlusion and dramatic changes in illumination. Finally, based on the corrected front view image, the reading recognition is performed, which greatly improves the accuracy and stability of the instrument reading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an instrument recognition method, computer equipment, storage medium, and computer program product. Background Technology

[0002] With the development of machine vision and deep learning technologies, instrument recognition technology has evolved from the traditional manual feature extraction mode to a data-driven intelligent recognition mode. Its core process usually includes key steps such as instrument area positioning, instrument posture correction, and reading recognition.

[0003] Instrument posture correction is a core prerequisite for determining reading accuracy. This is because instruments often exhibit varying degrees of tilt, rotation, or perspective distortion due to factors such as limited installation location, non-direct observation angle, bracket deformation, or vibration. This non-direct orientation leads to two problems: first, geometric distortion of the scale, where originally uniform graduations become unevenly spaced in the image, disrupting the mapping between the pointer's direction and the actual value; second, local occlusion, where the dial edge or digital wheel may be obstructed by its own casing, reflections, or dust, directly causing reading failure or significant deviation. Without posture correction, subsequent reading algorithms will be based on a distorted image, resulting in erroneous instrument readings.

[0004] Existing instrument calibration techniques mainly fall into two categories: one uses keypoint extraction algorithms (such as YOLO, HRNet, MobileNet, etc.) to locate instrument keypoints and then calculates the homography matrix for perspective transformation. The other directly uses a deep learning model to regress the homography matrix. However, these methods are all multi-stage algorithms with inherent drawbacks: on the one hand, keypoint detection errors and matrix estimation errors are amplified at each stage, ultimately resulting in severe distortion of the corrected orthographic image. On the other hand, extreme conditions in some production environments (direct sunlight, uneven nighttime lighting, and corrosive dust adhesion) further interfere with keypoint extraction, and the homography matrix is ​​essentially a planar perspective transformation, which cannot handle nonlinear distortions caused by curved dial surfaces, glass refraction, or local shadows. Therefore, existing methods experience a sharp decline in calibration effectiveness under conditions such as tilted instruments, drastic changes in lighting, and scale obstruction, leading to low accuracy of instrument readings. Summary of the Invention

[0005] To address the existing technical problems, this application provides an instrument identification method, computer equipment, storage medium, and computer program product that can improve the accuracy of instrument readings.

[0006] Firstly, a method for instrument identification is provided, the method comprising: Acquire inspection images of instruments at inspection points collected by the inspection equipment; Detect the instruments in the inspection image, and crop the image based on the position of the instruments in the inspection image to obtain an instrument image; The instrument image is input into a pre-trained correction model, and the correction model generates a front view image of the instrument. The correction model is an end-to-end model, and during training, the reading results of the front view image output by the correction model from a multimodal large language model are used as a supervision signal. The instrument reading is obtained by identifying the front view image of the instrument.

[0007] In a second aspect, a computer device is provided, including a processor and a memory connected to the processor, wherein the memory stores a computer program executable by the processor, and the computer program, when executed by the processor, implements the steps of the instrument identification method described in the above embodiments.

[0008] Thirdly, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the instrument identification method described in the above embodiments.

[0009] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, it implements the steps of the instrument identification method described in the above embodiments.

[0010] The instrument recognition method provided in the above embodiments first crops the instrument area in the inspection image to remove background interference and obtain an image highlighting the instrument. Then, the cropped image is input into an end-to-end correction model, which directly generates a frontal view image through a single forward propagation, avoiding error accumulation in traditional multi-stage methods and improving the robustness of the correction architecture. During the model training phase, the reading results of the frontal view image output by the correction model from a multimodal large language model are used as a supervision signal. This distills the geometric and semantic understanding capabilities of the large model into a lightweight model, and allows the correction model to directly optimize reading accuracy rather than just focusing on image similarity. This training strategy significantly improves the model's performance under complex conditions such as tilt, occlusion, and drastic changes in lighting. Finally, reading recognition is performed based on the corrected frontal view image, eliminating distortion and interference in the original image and presenting the dial in a frontal view shape, thereby greatly improving the accuracy and stability of instrument readings.

[0011] The computer equipment, storage medium, and computer program products provided in the above embodiments belong to the same concept as the corresponding instrument identification method embodiments, and thus have the same technical effects as the corresponding instrument identification method embodiments, which will not be repeated here. Attached Figure Description

[0012] Figure 1 This is a flowchart of an instrument identification method in one embodiment.

[0013] Figure 2 This is a schematic diagram of the functional modules of an inspection device in one embodiment.

[0014] Figure 3 This is a flowchart illustrating a method for training a correction model in one embodiment.

[0015] Figure 4 This is a schematic diagram of the structure of the model to be trained in one embodiment.

[0016] Figure 5 This is a flowchart of the steps for obtaining instrument readings based on the front view image of an instrument in one embodiment.

[0017] Figure 6 This is a schematic diagram of the front view image of an instrument in one embodiment.

[0018] Figure 7 To Figure 6 The key points for identifying the front view image of the instrument, and a schematic diagram of the first and second included angles.

[0019] Figure 8 This is a schematic diagram illustrating the training method of the correction model in one embodiment.

[0020] Figure 9 This is a flowchart of an instrument identification method in one embodiment.

[0021] Figure 10 This is a flowchart of the process for correcting instrument images in one implementation series.

[0022] Figure 11 This is a schematic diagram comparing an instrument image and an instrument front view image in one embodiment. Detailed Implementation

[0023] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0024] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0025] In the following description, the phrase "some embodiments" refers to a subset of all possible embodiments. It should be noted that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0026] In the following description, the terms "first, second, and third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, and third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0027] In production settings such as the petrochemical industry, various industrial instruments, including pressure gauges, thermometers, and level gauges, are core sensing carriers for monitoring equipment operation and ensuring production safety. The real-time performance and accuracy of instrument readings are directly related to process parameter control, equipment fault early warning, and safe production decisions. Taking the petrochemical industry as an example, due to the extreme environmental characteristics of petrochemical plants, such as high temperature, high pressure, flammable and explosive materials, corrosive gases, and dust, manual inspection is not only labor-intensive and risky, but also prone to misinterpretation or missed detection due to visual fatigue and subjective judgment bias. Therefore, inspection robots have gradually become the mainstream execution carrier for instrument monitoring in petrochemical scenarios.

[0028] With the development of machine vision and deep learning technologies, instrument recognition technology has evolved from the traditional manual feature extraction mode to a data-driven intelligent recognition mode. Its core process usually includes key steps such as instrument area positioning, instrument posture correction, and reading recognition.

[0029] Instrument posture correction is a core prerequisite for determining reading accuracy. This is because instruments often exhibit varying degrees of tilt, rotation, or perspective distortion due to factors such as limited installation location, non-direct observation angle, bracket deformation, or vibration. This non-direct orientation leads to two problems: first, geometric distortion of the scale, where originally uniform graduations become unevenly spaced in the image, disrupting the mapping between the pointer's direction and the actual value; second, local occlusion, where the dial edge or digital wheel may be obstructed by its own casing, reflections, or dust, directly causing reading failure or significant deviation. Without posture correction, subsequent reading algorithms will be based on a distorted image, resulting in erroneous instrument readings.

[0030] Existing instrument calibration techniques mainly fall into two categories: one uses keypoint extraction algorithms (such as YOLO, HRNet, MobileNet, etc.) to locate instrument keypoints and then calculates the homography matrix for perspective transformation. The other directly uses a deep learning model to regress the homography matrix. However, these methods are all multi-stage algorithms with inherent drawbacks: on the one hand, keypoint detection errors and matrix estimation errors are amplified at each stage, resulting in severe distortion of the corrected image. On the other hand, extreme conditions in some production environments (direct sunlight, uneven nighttime lighting, and corrosive dust adhesion) further interfere with keypoint extraction, and the homography matrix is ​​essentially a planar perspective transformation, which cannot handle nonlinear distortions caused by curved dial surfaces, glass refraction, or local shadows. Therefore, existing methods experience a sharp decline in calibration performance under conditions such as tilted instruments, drastic changes in lighting, and scale obstruction, leading to low accuracy of instrument readings.

[0031] To address this problem, this application provides an instrument identification method, such as... Figure 1 As shown, the method includes: Step 102: Obtain the inspection images of the instruments at the inspection points collected by the inspection equipment.

[0032] In one embodiment, the inspection equipment may be an autonomous mobile inspection robot. For example... Figure 2 As shown, the inspection equipment mainly includes a motion module 201, a data acquisition module 202, a network module 203, and an optional edge computing unit 204.

[0033] Among them, the motion module 201 is specifically a tracked or wheeled chassis, which enables the inspection equipment to navigate autonomously under different application conditions (such as anti-slip and corrosion-resistant ground) and move according to the preset inspection route.

[0034] The acquisition module 202 is specifically a 1-25x optical zoom rotating explosion-proof pan-tilt camera installed on the inspection equipment, meeting the explosion-proof rating Ex d IIC T6. The camera has a maximum field of view of 60°. Through the pan-tilt rotation and zoom control of the pan-tilt unit, the instrument under inspection can be adjusted to the center of the image, ensuring that the instrument area occupies the main proportion of the screen.

[0035] The network module 203 supports 5G / wired Ethernet and is used to send the collected inspection images to the AI ​​server in real time. It can also receive control commands through the network module 203.

[0036] The edge computing unit 204, which can be optionally configured on the inspection equipment, is specifically a GPU module that can run lightweight instrument recognition programs directly on the device, reducing network latency.

[0037] During inspection, the inspection robot is controlled by the motion module 201 to reach the inspection point along the preset path. Then, the acquisition module 202 automatically aligns with the target instrument and adjusts the magnification factor to make the instrument clear and centered, thus completing the acquisition of the inspection image.

[0038] In one embodiment, the inspection image can be sent to the AI ​​server via the network module 203. After receiving the image, the server stores the image and associates it with metadata such as the inspection point ID, timestamp, and robot number. Then, the server-side instrument recognition program identifies the instrument readings in the image.

[0039] In one embodiment, if the inspection equipment is equipped with an edge computing unit 204, the inspection image can be sent to the edge computing unit 204 and the instrument recognition program in the edge computing unit can be used for local recognition to reduce network dependence and transmission latency.

[0040] That is, the instrument identification method of this application can be implemented by the edge computing unit 204 of the inspection equipment, or by a server connected to the inspection equipment.

[0041] Step 104: Detect the instruments in the inspection image and crop them based on their positions in the inspection image to obtain the instrument image.

[0042] Specifically, firstly, an instrument target detection algorithm (such as YOLO, SSD, or a Transformer-based detection network) is used to locate the instruments in the inspection image, outputting the bounding box coordinates of each instrument (including the coordinates of the top left and bottom right corners). Then, based on the bounding boxes, the corresponding image regions are cropped from the inspection image to obtain the instrument image. This cropping operation effectively removes interference information from the background, allowing the instrument target to dominate the image, thereby reducing the difficulty of subsequent image correction and instrument reading recognition, and improving the overall recognition accuracy.

[0043] If a single inspection image contains multiple instruments, each instrument is cropped and then sent to the subsequent processing flow in sequence.

[0044] Among them, the instrument target detection algorithm is obtained through supervised training based on an image training set with pre-annotated instrument bounding boxes. Its main purpose is to accurately identify the location of the instrument in the inspection image, so as to provide a reliable area input for subsequent cropping, correction and reading.

[0045] Step 106: Input the instrument image into the pre-trained correction model and generate the instrument front view image through the correction model; wherein, the correction model is an end-to-end model, and during training, the reading results of the front view image output by the multimodal large language model are used as the supervision signal.

[0046] An end-to-end model is an artificial intelligence technique that can directly generate the final output from raw input data, omitting intermediate processing steps in traditional workflows. The correction model in this embodiment is an end-to-end model. When using this model for correction processing, the input instrument image directly outputs the corresponding front view image of the instrument, without the need to extract key points or calculate transformation matrices.

[0047] In contrast, most existing image correction methods employ a multi-stage process. For example, keypoint extraction algorithms (YOLO, HRNet, MobileNet, etc.) are first used to locate instrument keypoints, then the homography matrix is ​​calculated, and finally, the corrected image is obtained through matrix transformation. Alternatively, a deep learning model can be used to regress the homography matrix, and then the corrected image is obtained through matrix transformation. The common problem with these multi-stage methods is that errors are propagated and amplified at each stage, and the accuracy bottleneck of intermediate steps limits the final correction effect.

[0048] In contrast, the end-to-end correction model of this application directly generates the corrected image through a single forward propagation, eliminating the accumulation of errors in intermediate stages and improving the robustness and accuracy of correction from an architectural perspective.

[0049] Furthermore, during the model training phase, this application also uses the readings of the front view image output by the multimodal large language model as a supervision signal.

[0050] A multimodal large language model refers to a large-scale language model capable of simultaneously processing multiple modal inputs such as images and text, and generating text output. For example, a multimodal large language model can output instrument numerical text based on an input instrument image and prompts for reading the instrument reading. In this embodiment, the specific type of multimodal large language model is not limited; any existing mainstream model with image understanding and text generation capabilities can be used, such as Qwen3-VL.

[0051] This training method differs from traditional correction model training. Traditional correction model training typically uses image reconstruction quality as the optimization objective, evaluating whether the output corrected image looks like a front view. However, visually realistic corrected images may still contain minor geometric distortions, leading to subsequent reading errors. Therefore, this embodiment additionally inputs the corrected image output by the training correction model into a multimodal large language model during training. The large model is required to output instrument readings, and the cross-entropy loss between this reading and the actual reading is used as one of the supervision signals for training.

[0052] In this way, the powerful capabilities of the multimodal large language model in understanding the geometric structure and semantics of instruments can be transferred to the lightweight correction model, achieving knowledge distillation. At the same time, the correction model not only pursues visual similarity of images, but also directly optimizes the accuracy of the final reading, thereby significantly improving the performance of the model under complex conditions such as tilt, occlusion, and drastic changes in lighting.

[0053] By organically combining the aforementioned end-to-end architecture with large-model supervised training, the correction model of this application can output high-precision instrument front view images while improving real-time inference efficiency.

[0054] Step 108: Based on the front view image of the instrument, identify and obtain the instrument reading.

[0055] Since the instrument front view image output by the correction model has effectively eliminated interference such as tilt, perspective distortion, local occlusion and uneven lighting in the original image, making the dial present a front view shape, reading recognition on the instrument front view image can significantly improve the accuracy and stability of recognition.

[0056] In one embodiment, if the instrument is a digital instrument, the digital display area in the front view image of the instrument is detected, and optical character recognition (OCR) technology is applied to this area to extract the digital string as the instrument reading. Frontal visualization processing ensures that the numbers are arranged neatly and without perspective distortion, which can significantly improve the success rate of OCR recognition.

[0057] In one embodiment, if the instrument is a pointer instrument, the coordinates of key pointer points (such as the pointer tip and the axis of rotation) and key range points (such as the zero mark and the full-scale mark) in the front view image of the instrument are detected. The pointer deflection angle is calculated based on the coordinates of each point, and the instrument reading is determined in combination with the range of the instrument. The front view image ensures that the scale is evenly distributed and the angle mapping is accurate and reliable.

[0058] The aforementioned instrument recognition method first crops the instrument area in the inspection image to remove background interference, obtaining an image that highlights the instrument. Then, the cropped image is input into an end-to-end correction model, directly generating a frontal view image through a single forward propagation. This avoids error accumulation in traditional multi-stage methods, improving the robustness of the correction architecture. During model training, the reading results of the frontal view image output by the correction model from a multimodal large language model are used as a supervision signal. This distills the geometric and semantic understanding capabilities of the large model into a lightweight model, and allows the correction model to directly optimize reading accuracy rather than solely focusing on image similarity. This training strategy significantly improves the model's performance under complex conditions such as tilt, occlusion, and drastic changes in lighting. Finally, reading recognition is performed based on the corrected frontal view image, eliminating distortion and interference in the original image and presenting the dial in a frontal view form, thereby greatly improving the accuracy and stability of instrument readings.

[0059] In one embodiment, such as Figure 4 As shown, the methods for training the correction model include: Step 302: Based on the image training set, train the end-to-end model to be trained to obtain the preliminary training base model.

[0060] The image training set includes instrument images and labeled data for the instrument images. The labeled data includes at least the front view image corresponding to the instrument image and the actual reading of the instrument image.

[0061] The base model in this embodiment is an end-to-end model initially trained using an image training set. The base model has the ability to generate corrected orthographic images from instrument images.

[0062] Step 304: Input the instrument images from the image training set into the base model and output the predicted corrected front view image.

[0063] Using the basic model obtained in step 302, forward inference is performed on the instrument images in the training image set to obtain the corresponding predicted and corrected images.

[0064] Step 306: Calculate the training loss based on the prediction results of the base model and the corresponding labeled data.

[0065] The training loss is used to measure the difference between the corrected image output by the base model and the true front view image. In this embodiment, the training loss ensures that the corrected image is structurally close to the true front view.

[0066] The training loss varies depending on the structure of the base model. For example, when the base model is simply a generator, the training loss typically only includes pixel reconstruction loss (such as L1 loss). When the base model is a generative adversarial network (GAN), the training loss includes both pixel reconstruction loss and adversarial loss. If an auxiliary head (such as keypoint heatmap prediction) is added on top of this, the training loss must also include the geometric loss of the auxiliary head.

[0067] Step 308: Input the predicted corrected front view image and the reading prompt words into the multimodal large language model to obtain the instrument reading predicted by the multimodal large language model.

[0068] The reading prompt is something like, "Please read the meter reading in this image." After receiving the corrected front view image and the reading prompt, the multimodal large language model outputs the reading text as the prediction result.

[0069] Step 310: Calculate the reading loss based on the predicted instrument readings and the actual readings.

[0070] The image training set contains labeled actual instrument readings. The predicted reading output in step 308 is compared with the actual reading, and the reading loss is calculated. In one embodiment, the reading loss can be the cross-entropy loss between the predicted and actual readings, used to measure the difference between the predicted and actual readings.

[0071] Step 312: Iteratively execute the step of jointly fine-tuning the base model based on the training loss and reading loss until the training termination condition is met, and obtain the trained correction model.

[0072] In this step, the training loss obtained in step 306 and the reading loss obtained in step 310 are weighted and combined, and the parameters of the base model are updated through backpropagation. Through multiple iterations, the model not only learns to generate visually similar front views, but also directly optimizes the accuracy of the final reading.

[0073] The training termination condition can be that the total loss is below a threshold or that a preset number of iterations has been reached. For example, the training termination condition is met when the total loss is below 0.5 or when the training reaches 200 epochs.

[0074] Traditional correction models prioritize image reconstruction quality, often resulting in visually realistic corrections with minor geometric distortions, leading to subsequent reading errors. This application introduces reading feedback from a multimodal large language model as additional supervision during the training phase, offering two advantages. First, it transfers the powerful capabilities of the large model in understanding instrument geometry and semantics to a lightweight correction model. Second, it enables the correction model to directly optimize reading accuracy, rather than solely pursuing image similarity.

[0075] By organically combining end-to-end architecture with large-model supervised training, the correction model of this application significantly improves the accuracy and stability of instrument readings under complex working conditions such as tilt, occlusion, and drastic changes in lighting while maintaining real-time inference efficiency.

[0076] In one embodiment, the model to be trained includes a feature extraction network, and the model to be trained also includes a training auxiliary head connected to an intermediate layer of the feature extraction network; the training auxiliary head uses a convolutional neural network; the labeled data also includes a key point heatmap of the instrument image.

[0077] A keypoint heatmap is a two-dimensional visualization tool used in deep learning and computer vision to represent the probability distribution of keypoint locations in an image. It encodes the location of each keypoint (such as the center of an instrument panel, corner points of a border, etc.) as a continuous probability distribution map, rather than discrete coordinates, thus facilitating model learning of spatial structure and improving detection accuracy. In one embodiment, a multimodal large language model can be used to automatically generate a keypoint heatmap of an instrument panel image as a supervisory signal.

[0078] In this embodiment, the base model is trained based on the model to be trained. For example... Figure 4 As shown, the model to be trained includes a feature extraction network and a training auxiliary head connected to its intermediate layers. Specifically, the training auxiliary head consists of multiple convolutional layers, batch normalization (BN) layers, ReLU activation functions, and a final sigmoid output layer. The processing is as follows: the high-dimensional semantic feature tensor output from the intermediate layers of the feature extraction network is input into the auxiliary head, sequentially passing through convolutional layers for dimensionality reduction, BN layers to accelerate convergence, ReLU to introduce non-linearity, and then being mapped to the number of keypoint categories through convolutional layers. Finally, the sigmoid function outputs the probability that each pixel location belongs to a keypoint (e.g., 5 categories: a center point and four border points), thereby generating a keypoint heatmap.

[0079] like Figure 3 As shown, step 302, namely, training the end-to-end model to be trained based on the image training set to obtain the preliminary trained base model, includes: Step 3021: Input the instrument images from the image training set into the model to be trained, and output the predicted corrected front view image.

[0080] The model to be trained processes the input image using a forward propagation method to generate the corresponding corrected image.

[0081] Step 3023: Calculate the first training loss based on the corrected front view image output by the model to be trained and the labeled front view image.

[0082] The first training loss measures the pixel-level difference between the corrected image output by the model and the real front view image. In this embodiment, the first training loss uses L1 loss or L2 loss to ensure that the corrected image is close to the real front view in terms of overall structure and brightness.

[0083] Step 3025: Input the feature tensor output from the intermediate layer into the training auxiliary head, and output the predicted key point heatmap through the convolution operation of the training auxiliary head.

[0084] This operation is performed only during the training phase. For example... Figure 4 As shown, the training auxiliary head receives the feature extraction network (such as the feature tensor of the bottleneck layer of the encoder of the generator (e.g., the size is 16×16×512)), and after several convolutional layers, the dimensionality is reduced, and the number of output channels is equal to the number of key point categories (e.g., 5), generating a heatmap that is proportionally reduced to the size of the input image (e.g., 64×64×5).

[0085] Step 3027: Calculate the second training loss based on the predicted keypoint heatmap and the labeled keypoint heatmap.

[0086] The labeled keypoint heatmap is generated by labeling keypoints on the front view image using a multimodal large language model or manually, resulting in a downsampled Gaussian heatmap (σ = 2-3 pixels). The second training loss can be the mean squared error loss (MSE Loss), which calculates the squared difference between each pixel in the predicted heatmap and the labeled heatmap.

[0087] in, To train the auxiliary head to predict key point heatmaps, Heatmap of key points marked.

[0088] The intermediate layers of a feature extraction network, such as the bottleneck layer of the encoder, are typically where the feature map size is smallest and the semantics are most abstract. Applying keypoint heatmap supervision at this point forces the encoder's hidden space to explicitly encode the spatial structure of the dashboard (center, borders, pointer axis, etc.), avoiding geometric blurring caused by pixel-level loss alone. It's important to emphasize that the training auxiliary head is only activated during model training and does not participate in computation during inference, thus not adding any extra overhead after deployment.

[0089] Step 3029: Iteratively execute the steps of adjusting the model to be trained according to the training loss until the training termination condition is met, and obtain the initial training base model; the training loss includes the first training loss and the second training loss.

[0090] Training loss is defined as: in This is the first training loss. For the second training loss, These are the weight coefficients. The parameters of the feature extraction network and the training auxiliary head are updated simultaneously through backpropagation. This process is iteratively executed until both the first and second training losses converge (e.g., the loss decreases less than a threshold for 10 consecutive epochs) or the preset maximum number of epochs is reached.

[0091] By introducing the aforementioned training auxiliary head, the base model gains geometric supervision during the initial training phase, enabling it to more accurately recover the frontal view of the dashboard and providing a good initial state for subsequent semantic fine-tuning of the multimodal large language model.

[0092] In this embodiment, the training process of the base model is divided into two stages, and the training loss of both stages includes a first training loss and a second training loss. The specific logic is as follows: Phase 1: Initial training of the basic model.

[0093] This stage uses the image training set to perform end-to-end supervised learning on the model to be trained. The total loss includes the first training loss. Second training loss Weighted composition.

[0094] The first training loss measures the pixel-level difference between the predicted corrected image and the labeled front view image, while the second training loss measures the geometric consistency between the predicted keypoint heatmap and the labeled heatmap. By jointly optimizing these two losses, the base model not only learns the pixel mapping from the tilted image to the front view image, but also forces the encoder to model the spatial structure of the dashboard (center, border, etc.) through the auxiliary head, thus achieving high geometric correction accuracy in the initial training stage.

[0095] Phase 2: Joint fine-tuning based on reading loss of multimodal large language models.

[0096] After the initial training of the basic model is completed, reading feedback from a multimodal large language model is introduced as an additional supervision signal. This stage retains the first training loss from the first stage. Second training loss At the same time, the reading loss is increased. (That is, the cross-entropy loss between the predicted and actual readings of the large model). The total loss is expanded to: in These are the weighting coefficients. Through joint fine-tuning, the model further optimizes the accuracy of the final reading while maintaining its original geometric correction capability, distilling the understanding of instrument semantics from the large model into a lightweight correction model.

[0097] Whether in the initial training of the basic model or in the subsequent fine-tuning stage using the reading loss of a large model, the first and second training losses always exist as core constraints. The former ensures the visual quality of the image, the latter maintains the correctness of the geometric structure, while the reading loss serves as a task-driven enhancement. The three work together to enable the final correction model to possess both high geometric accuracy and high reading accuracy under complex conditions.

[0098] In one embodiment, the model to be trained is a generative adversarial network (GAN), which includes a generator and a discriminator.

[0099] The generator uses an encoder-decoder structure (such as U-Net) to generate an orthographic image from a tilted instrument image; the discriminator uses a patchGAN structure to distinguish the generated orthographic image from the real orthographic image.

[0100] The process of training a basic model based on a generative adversarial network includes the following steps: Step 1: Input the instrument images from the image training set into the pre-trained generator in the generative adversarial network, and output the predicted corrected front view image.

[0101] In this embodiment, the generator is first pre-trained based on an image training set to enable it to generate orthographic images from tilted images. After pre-training, the pre-trained generator is used as the initial generator for subsequent adversarial training to output the predicted corrected orthographic images.

[0102] Step 2: Calculate the pixel reconstruction loss of the generator based on the difference between the corrected orthographic image output by the pre-trained generator and the labeled orthographic image.

[0103] Step 3: Input the corrected orthographic image and the labeled orthographic image into the discriminator to obtain the discrimination result, and calculate the adversarial loss based on the discrimination result and the true category.

[0104] In this embodiment, the first training loss includes pixel reconstruction loss and adversarial loss. The pixel reconstruction loss is calculated based on the pixel-level difference between the corrected front view image output by the pre-trained generator and the labeled front view image. The corrected front view image and the labeled front view image output by the generator are respectively input into the discriminator to obtain the discrimination result. The adversarial loss is calculated based on the discrimination result and the true category (true for the real image, false for the generated image).

[0105] The first training loss is used to alternately update the generator and discriminator, prompting the generator to output a more realistic front view image.

[0106] Step 4: Input the feature tensor output from the intermediate layer into the training auxiliary head, and output the predicted key point heatmap through the convolution operation of the training auxiliary head.

[0107] Step 5: Calculate the second training loss based on the predicted keypoint heatmap and the labeled keypoint heatmap.

[0108] Step 6: Iteratively execute the steps to adjust the model to be trained based on the training loss until the training termination condition is met, thus obtaining the initial trained base model. The training loss includes the first training loss and the second training loss.

[0109] In this embodiment, a generative adversarial network (GAN) is used to train the base model. The adversarial loss of the discriminator forces the generator to output texture details that are closer to the real front view, avoiding blurring or artifacts and making the corrected instrument image look more natural. The base model trained with GANs already has good visual quality and geometric accuracy. On this basis, a large model reading loss is introduced for fine-tuning, which can further optimize the reading accuracy.

[0110] In one embodiment, the pre-training method of the generator includes: inputting instrument images from the image training set into the generator to be trained and outputting a corrected front view image; calculating the pixel reconstruction loss of the generator based on the difference between the corrected front view image output by the generator to be trained and the labeled front view image; iteratively executing the step of updating the generator parameters based on the pixel reconstruction loss of the generator until the training termination condition is met, and obtaining the pre-trained generator.

[0111] In this embodiment, the training process of the base model is divided into three stages.

[0112] Phase 1: Pre-trained generator.

[0113] Training loss includes pixel reconstruction loss. At this point, the discriminator, training auxiliary head, and multimodal large language model are all frozen. This stage enables the generator to have basic correction capabilities.

[0114] Phase 2: Training the base model. At this stage, the discriminator is activated and the auxiliary head is trained, while the multimodal large language model remains frozen.

[0115] The training loss includes a first training loss and a second training loss. By training the generator and discriminator alternately, and by connecting a training auxiliary head to the intermediate layer of the generator to apply geometric supervision, the generator can improve the realism of the image while maintaining geometric correctness. After this stage of training is completed, a preliminary trained base model is obtained.

[0116] The first training loss includes the pixel reconstruction loss of the generator and the adversarial loss of the discriminator.

[0117] The second training loss is the MSE loss between the keypoint heatmap output by the auxiliary head and the labeled heatmap.

[0118] In the second stage of training the base model, under the constraint of training loss, the generator simultaneously receives pixel reconstruction loss (to ensure structural similarity), adversarial loss (to improve texture realism) and geometric loss (to constrain key point positions), forcing the generator to maintain the perfect circle shape of the dashboard, uniform scale, and accurate pointer axis while outputting high-fidelity images.

[0119] The third stage: using reading feedback from the multimodal large language model as a supervision signal, the basic model obtained in the second stage is fine-tuned.

[0120] In this phase, the total loss includes the first training loss, the second training loss, and the reading loss, which is the cross-entropy loss between the predicted readings and the actual readings of the multimodal large language model.

[0121] In this stage, constrained by the total loss, the generator not only pursues visual similarity but also directly optimizes reading accuracy. After training, only the generator is retained as the final correction model. This generator is an end-to-end lightweight model that takes an instrument image as input and directly outputs a corrected front view image, without requiring a discriminator, auxiliary head, or large model for inference, thus meeting real-time requirements.

[0122] In one embodiment, the instrument recognition method further includes: acquiring a basic image training set, which includes instrument images and labeled data of the instrument images; inputting the instrument images from the basic image training set into a multimodal large language model to generate an expanded image with interference features; and obtaining an image training set based on the expanded image, the basic image, and the corresponding labeled data.

[0123] In this embodiment, a basic image training set is obtained, which includes instrument images and their corresponding labeled data, such as front view images, actual readings, and key point heat maps.

[0124] The instrument images from the basic image training set are input into a multimodal large language model (e.g., Qwen3-VL), and corresponding interference is generated to provide prompts. This allows the large model to add interference features simulating real-world operating conditions to the images while maintaining the main structure of the instrument. These features include, but are not limited to, motion blur, rain / snow coverage, direct sunlight, uneven nighttime lighting, and dust occlusion. The multimodal large language model then outputs the corresponding augmented images.

[0125] The final image training set is composed of the augmented images, the base images, and the corresponding labeled data. The augmented images inherit the labeling information of the original base images (such as frontal views, actual readings, etc.), requiring no additional manual labeling.

[0126] In this embodiment, the image generation and editing capabilities of a multimodal large language model are utilized to enrich the diversity of training data at low cost and on a large scale, enabling the model to be exposed to complex working conditions such as blur, rain, snow, and strong light during the training phase, thereby significantly improving the robustness and generalization ability of the correction model in real petrochemical scenarios.

[0127] It should be noted that the multimodal large language model used to generate the augmented image in this embodiment can be the same model or a different model from the large language model that provides reading feedback during subsequent training; this application does not impose any restrictions on this. Both can be selected independently based on actual computing power, deployment costs, and generation quality requirements.

[0128] In one embodiment, such as Figure 5 As shown, the instrument reading is obtained by identifying the instrument's front view image, including: Step 502: Identify the type of instrument.

[0129] In one embodiment, the instrument detection model detects the instrument's location and also outputs the instrument's category, such as pointer-type or digital-type instrument. This detection and classification function can be implemented through a multi-task network and pre-trained on an image set labeled with instrument locations and types. In other embodiments, the instrument type for each inspection point can be pre-configured, and when acquiring inspection images, the instrument type can be directly obtained by looking up a table based on the inspection point information, without relying on image classification.

[0130] If the instrument is a pointer instrument, then proceed to step 504 to identify the coordinates of the pointer key point and the range key point in the instrument's front view image.

[0131] In one embodiment, a keypoint detection network is used to infer the front view image of the instrument and output the pixel coordinates of key points. The keypoint detection network can be HRNet, SimpleBaseline, or the lightweight MobileNetV3+ keypoint header structure.

[0132] When training the keypoint detection network, annotation tools can be used to annotate the keypoints (including pointer keypoints and range keypoints) and bounding boxes of each instrument, and to set the instrument label category. The trained keypoint detection network is then trained on the image data with the annotated keypoints using any of the aforementioned network structures.

[0133] In one embodiment, the key point categories include the range start point, range end point, pointer rotation axis, and pointer tip. In this embodiment, the coordinates of the identified range key points may include the coordinates of the range start point (e.g., the 0 mark) on the dial. , ), coordinates of the end point of the measurement range (such as the full-scale position) , The coordinates of the pointer key points can include the coordinates of the pointer rotation axis. , ) and pointer tip coordinates ( , ).

[0134] After step 504, step 506 is executed, whereby the ratio of the pointer indication position to the full scale of the instrument is calculated based on the coordinates of the pointer key point and the range key point, and the instrument reading is determined.

[0135] Specifically, identifying such as Figure 6 The measuring range shown in the instrument's front view image is calculated based on the coordinates of the starting point of the measuring range, the ending point of the measuring range, the coordinates of the pointer's rotation axis, and the coordinates of the pointer tip. Figure 7The pointer is shown to be at the first angle θ between the starting point of the range and the second angle α between the starting point and the ending point of the range. Based on the ratio of the first angle to the second angle, the ratio of the pointer position to the full scale of the instrument is determined. Based on the range and the ratio of the pointer position to the full scale of the instrument, the instrument reading is calculated.

[0136] Specifically, first, identify the measurement range in the instrument's front view image. The measurement range is specifically the maximum measurement value. and minimum range The difference. Among them, the minimum range value. and maximum value The range can be obtained by searching a preset database based on the instrument model, or by automatically resolving the range by performing OCR recognition on the numbers on the dial.

[0137] Then, rotate the axis with the pointer. , As the pole, calculate the first angle between the pointer and the starting point of the range. The second included angle between the start and end points of the measurement range .

[0138] Considering the image coordinate system In the axial direction downwards, to ensure that the angle increases as the pointer rotates clockwise (consistent with the actual instrument reading direction), the calculation needs to be adjusted accordingly. The components are negative, thus mapping the image coordinates to a mathematical coordinate system.

[0139] The formula for calculating the included angle is as follows: Finally, the actual instrument readings are calculated using a linear scaling mapping: in, Indicates the measurement range. This indicates the first angle between the pointer and the starting point of the range. This represents the second included angle between the start and end points of the measurement range. This indicates the minimum value of the measurement range.

[0140] Because the front view image of the instrument has eliminated perspective distortion and the scale is distributed in a standard circular arc, high-precision readings can be obtained based on the range and the ratio of the pointer indication position to the full scale of the instrument.

[0141] like Figure 5 As shown, if the instrument is a digital instrument, then step 508 is executed to detect the digital display area in the front view image of the instrument, perform optical character recognition on the digital display area, and obtain the instrument reading.

[0142] This method can automatically identify different types of instruments. Based on the corrected image, it uses a sub-strategy for reading analysis. Specifically, it employs angle mapping for pointer instruments and optical character recognition for digital instruments. This significantly improves the reading accuracy and stability of various instruments under complex operating conditions.

[0143] In one embodiment, the instrument recognition method of this application is described using a correction model trained with a generative adversarial network as an example. The method comprises four stages: instrument target detection model training, correction model training, instrument key point detection network training, and instrument recognition application.

[0144] First, the instrument target detection model is trained.

[0145] The image training set consists of N sets of collected field images and images acquired from the network. The field images are visible light images collected by inspection equipment from field instruments, supplemented by publicly available instrument images acquired from the network. According to the model training requirements, the original images are scaled proportionally to 640×640 pixels, with insufficient short sides filled with gray.

[0146] For the image training set, the annotation stage uses annotation tools (such as X-AnyLabeling) to annotate each meter with a bounding box, uniformly categorized as "meter". To improve annotation efficiency, a large language model (such as the LLMDet model) is introduced to assist in annotation. After specifying the target category to be annotated, the large language model automatically identifies and generates the corresponding bounding box, achieving rapid pre-annotation. This is then manually reviewed and corrected before final annotation, forming the final annotated data.

[0147] In one embodiment, the instrument target detection model adopts the YOLOv11 architecture, and its network structure includes a Backbone, a Neck (PAN-FPN), and a Head (decoupled detection head). It is trained using a labeled image training set and can output bounding box coordinates and class confidence.

[0148] Secondly, correct the model training.

[0149] Obtain the basic image training set. The basic images mainly consist of tilted instrument images (instrument areas after target detection and cropping) collected by the inspection equipment. Their labeled data includes paired front view images, actual readings, and key point heatmaps. Front view images can be obtained by manually adjusting the instrument angle during shooting, or by using a multimodal large language model to correct the tilted images and generate pseudo-front view images. Key point heatmaps can be generated by labeling key points (center and four border points) on the front view images using a multimodal large language model or manually, producing a Gaussian heatmap (64×64 pixels, matching the bottleneck layer output).

[0150] The instrument images from the base image training set are input into a multimodal large language model (such as Qwen3-VL) to generate augmented images with interference features, such as motion blur, rain and snow coverage, direct sunlight, uneven nighttime lighting, and dust occlusion. Based on the augmented images, the base images, and the corresponding labeled data, the final image training set is obtained by merging them.

[0151] In the image training set preparation stage, the image generation and editing capabilities of the multimodal large language model are utilized to enrich the diversity of training data at low cost and on a large scale, enabling the model to be exposed to complex conditions such as blur, rain, snow, and strong light during the training stage, thereby significantly improving the robustness and generalization ability of the correction model under different complex conditions.

[0152] In one embodiment, the overall architecture of the correction model is a pix2pix generative adversarial network. The generator uses U-Net, the discriminator uses patchGAN, and the multimodal large language model Qwen3-VL is used as a high-level discriminator during training to improve the semantic and geometric accuracy of the training. Figure 8 As shown, the training process is divided into three stages: Phase 1: Training the generator.

[0153] In this phase, the discriminator patchGAN module, the Qwen3-VL large model, and the training auxiliary head are frozen, and the generator is trained solely on the image training set. Specifically, the instrument images from the image training set are input into the generator to be trained, and the generator outputs a corrected front view image. Based on the difference between the corrected front view image output by the generator and the labeled front view image, the pixel reconstruction loss of the generator is calculated. The steps of updating the generator's parameters based on the pixel reconstruction loss are iteratively executed until the training termination condition is met, resulting in a pre-trained generator.

[0154] This stage uses only the pixel reconstruction loss L1 loss.

[0155] in, For the input instrument image, For the corresponding front view image, The corrected image output by the generator.

[0156] The training objective of this stage is to enable the generator to learn the basic geometric transformation contours from a tilted image to a frontal image.

[0157] Phase 2: Training the base model.

[0158] In this stage, the discriminator and training auxiliary head are activated, while the multimodal large language model remains frozen. Specifically, a parallel convolutional sub-network is introduced as a training auxiliary head from the intermediate layer of the generator encoder (e.g., the bottleneck layer of the encoder). The high-dimensional semantic feature tensor of the intermediate layer is input into the training auxiliary head, and the channel dimension is mapped to the number of keypoint categories through convolution operations. A keypoint heatmap is generated, with 5 keypoints: one center point and four bounding box points.

[0159] Specifically, the instrument images from the image training set are input into a pre-trained generator, which outputs a predicted corrected front view image. The pixel reconstruction loss of the generator is calculated based on the difference between the corrected front view image and the labeled front view image. The corrected and labeled front view images are input into a discriminator to obtain a discrimination result, and an adversarial loss is calculated based on the discrimination result and the true class. The high-dimensional semantic feature tensor output from the intermediate layer is input into a training auxiliary head, and a predicted keypoint heatmap is output through convolution operations. A second training loss is calculated based on the predicted and labeled keypoint heatmaps. The process of adjusting the model to be trained based on the training losses is iteratively executed until the training termination condition is met, resulting in a pre-trained base model. The training losses in this stage include pixel reconstruction loss, adversarial loss, and the second training loss.

[0160] This can be represented as: in, , , These are the weighting coefficients, To combat the loss, a discriminator provides a mechanism to ensure that the generated image textures are realistic. ; For pixel reconstruction loss, The second training loss is generated by the training auxiliary head and can be the mean square error between the predicted heatmap and the labeled heatmap.

[0161] This stage enables the generator to significantly improve the texture realism of the output image while maintaining the basic geometric contours, and forces the encoder to learn the spatial geometry of the instrument (center position, dial boundary) by training the auxiliary head, so that it can still output a geometrically correct front view under conditions such as tilt and occlusion.

[0162] The third stage: semantic fine-tuning stage.

[0163] Input the instrument images from the image training set into the base model and output the predicted corrected front view image; calculate the training loss based on the prediction results of the base model and the corresponding labeled data; input the predicted corrected front view image and the reading prompt words into the multimodal large language model to obtain the instrument reading predicted by the multimodal large language model; calculate the reading loss based on the predicted instrument reading and the actual reading; iteratively execute the step of jointly fine-tuning the base model based on the training loss and the reading loss until the training termination condition is met, and obtain the trained corrected model.

[0164] In this phase, the total loss includes training loss and reading loss. Reading loss is the cross-entropy loss between the predicted instrument reading and the actual reading.

[0165] in, For reading loss, This is the weighting coefficient for reading loss.

[0166] In this phase, the gradient of the total loss is fed back to the generator through the Input Embedding layer of the multimodal large language model to update the generator parameters. Training is complete when the total loss is below 0.5 or the training reaches 200 epochs.

[0167] After the model is trained, only the generator is retained as the final corrected model.

[0168] Next, the instrument key point detection network is trained.

[0169] The image training set includes corrected front-view images of the instruments output by the model (or manually photographed front-view images), as well as raw visible light images (uncorrected) to enhance robustness. According to model training requirements, the images are scaled to 640×640 pixels (maintaining aspect ratio, with short sides filled in gray). Using annotation tools, four key point categories (range start, range end, pointer rotation axis, and pointer tip) and bounding boxes are labeled for each instrument, and instrument label categories are assigned.

[0170] In one embodiment, the keypoint detection network adopts a YOLOv11-pose structure, which includes a backbone, a neck (PAN-FPN), and a head (decoupled detection head). Trained using a labeled image training set, it can output the coordinates of four keypoints of the instrument.

[0171] Finally, instrument recognition applications.

[0172] like Figure 9As shown, the inspection equipment captures images of the instruments at the inspection points. The equipment then transmits these images to an AI server via a network module. The AI ​​server is equipped with the trained instrument target detection model, correction model, and instrument keypoint detection network. The AI ​​server processes the instrument images as follows: The AI ​​server performs instrument detection processing on the inspection images. Specifically, the AI ​​server uses a deployed instrument target detection model to detect instruments in the inspection images and crops them based on their positions in the images to obtain instrument images.

[0173] The AI ​​server performs instrument correction processing on the instrument images. Specifically, such as... Figure 10 As shown, the instrument image is input into the pre-trained correction model, which then generates a front view image of the instrument. Figure 11 As shown, the correction model can be used to correct tilted instrument images (such as...). Figure 11 (Left side) converted to instrument front view image (e.g.) Figure 11 (Right side) effectively eliminates geometric distortion caused by perspective distortion and tilt angle, making the dial present a perfect circle shape and the scale evenly distributed, providing high-quality input for subsequent reading recognition.

[0174] The AI ​​server processes the readings from the instrument's front view image. Specifically, it identifies the instrument's reading based on the front view image. Specifically, it identifies the instrument type; if the instrument is an analog instrument, the AI ​​server uses a deployed instrument key point detection network to identify the coordinates of the pointer key points and the range key points in the instrument's front view image; based on the coordinates of the pointer and range key points, it calculates the ratio of the pointer's indicated position to the instrument's full scale to determine the instrument reading. If the instrument is a digital instrument, it detects the digital display area in the instrument's front view image and performs optical character recognition on the digital display area to obtain the instrument reading.

[0175] This method can automatically identify different types of instruments. Based on the corrected image, it uses a differentiated reading analysis strategy, such as angle mapping for pointer instruments and OCR for digital instruments. This significantly improves the reading accuracy and stability of various instruments under complex operating conditions.

[0176] The instrument identification method of this application has the following advantages: (1) By using inspection equipment and rotating gimbal, it is possible to collect instrument images under different working conditions at the application site, and cover instruments in different locations in the factory around the clock and in all scenarios.

[0177] (2) Distill the end-to-end image correction capability of the multimodal large language model into a lightweight end-to-end model to improve the speed of single instrument recognition and improve inspection efficiency.

[0178] (3) Use an end-to-end correction model to restore the instrument front view image information (including scale, pointer, numbers, etc.) to the greatest extent possible. Even if some scales are obscured, they can be effectively restored, thus improving the accuracy of readings.

[0179] (4) Enhance the instrument correction training data using a multimodal large language model (such as adding interference such as blur, rain, snow, and strong light) to enhance the generalization ability and robustness of the model, making it suitable for extreme working conditions such as rain and snow, blurry images, and strong light interference.

[0180] (5) Based on the joint supervision of generative adversarial network (generator + discriminator) and multimodal large language model, pixel reconstruction loss, adversarial loss, geometric heatmap loss and large model reading loss are integrated to improve the training accuracy of lightweight correction model.

[0181] In another aspect, this application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described instrument identification method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0182] In another aspect, this application also provides a computer device, which includes a processor and a memory connected to the processor. The memory stores a computer program that can be executed by the processor. When the computer program is executed by the processor, it implements the various processes of the instrument identification method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0183] In another aspect, this application also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the various processes of the instrument recognition method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0184] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0185] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0186] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for instrument identification, characterized in that, The method includes: Acquire inspection images of instruments at inspection points collected by the inspection equipment; Detect the instruments in the inspection image, and crop the image based on the position of the instruments in the inspection image to obtain an instrument image; The instrument image is input into a pre-trained correction model, and the correction model generates a front view image of the instrument. The correction model is an end-to-end model, and during training, the reading results of the front view image output by the correction model from a multimodal large language model are used as a supervision signal. The instrument reading is obtained by identifying the front view image of the instrument.

2. The instrument identification method according to claim 1, characterized in that, The methods for training the correction model include: Based on the image training set, the end-to-end training model is trained to obtain a preliminary training base model; the image training set includes instrument images and labeled data of the instrument images, the labeled data including at least the front view image corresponding to the instrument image and the actual reading of the instrument image. The instrument images in the image training set are input into the base model, and the predicted corrected front view image is output. Based on the prediction results of the base model and the corresponding labeled data, calculate the training loss; The predicted corrected front view image and the reading prompt words are input into the multimodal large language model to obtain the instrument reading predicted by the multimodal large language model; Calculate the reading loss based on the predicted instrument reading and the actual reading; The training loss and the reading loss are used to iteratively fine-tune the base model until the training termination condition is met, at which point the trained corrected model is obtained.

3. The instrument identification method according to claim 2, characterized in that, The model to be trained includes a feature extraction network, and the model to be trained also includes a training auxiliary head connected to the intermediate layer of the feature extraction network; the training auxiliary head uses a convolutional neural network; the labeled data also includes a heatmap of key points of the instrument image; The process of training the end-to-end model based on the image training set to obtain a preliminary training base model includes: The instrument images in the image training set are input into the model to be trained, and the predicted corrected front view image is output. The first training loss is calculated based on the corrected front view image output by the model to be trained and the labeled front view image. The feature tensor output from the intermediate layer is input into the training auxiliary head, and the predicted key point heatmap is output through the convolution operation of the training auxiliary head. The second training loss is calculated based on the predicted keypoint heatmap and the labeled keypoint heatmap; The training loss is adjusted according to the training loss in an iterative manner until the training termination condition is met, and a preliminary training base model is obtained; the training loss includes the first training loss and the second training loss.

4. The instrument identification method according to claim 3, characterized in that, The model to be trained is a generative adversarial network, which includes a generator and a discriminator. The step of inputting the instrument images in the image training set into the model to be trained and outputting the predicted corrected front view image includes: inputting the instrument images in the image training set into the pre-trained generator in the generative adversarial network and outputting the predicted corrected front view image. The step of calculating the first training loss based on the corrected orthographic image predicted by the model to be trained and the labeled orthographic image includes: The pixel reconstruction loss of the generator is calculated based on the difference between the corrected orthographic image output by the generator and the labeled orthographic image. The corrected orthographic image and the labeled orthographic image are input into the discriminator to obtain a discrimination result. The adversarial loss is calculated based on the discrimination result and the true category. The first training loss includes the pixel reconstruction loss and the adversarial loss.

5. The instrument identification method according to claim 4, characterized in that, The pre-training methods for the generator include: The instrument images in the image training set are input into the generator to be trained, and the corrected front view image is output. The pixel reconstruction loss of the generator is calculated based on the difference between the corrected orthographic image output by the generator to be trained and the labeled orthographic image. The process of iteratively updating the generator's parameters based on the generator's pixel reconstruction loss continues until the training termination condition is met, resulting in a pre-trained generator.

6. The instrument identification method according to any one of claims 2 to 5, characterized in that, The method further includes: Obtain a basic image training set, which includes instrument images and labeled data of the instrument images; The instrument images from the basic image training set are input into the multimodal large language model to generate augmented images with interference features; The image training set is obtained based on the augmented image, the base image, and the corresponding annotation data.

7. The instrument identification method according to claim 1, characterized in that, The step of identifying the instrument reading based on the front view image includes: Identify the type of the instrument; If the instrument is a pointer instrument, then identify the coordinates of the pointer key point and the range key point in the front view image of the instrument; Based on the coordinates of the pointer key point and the range key point, calculate the ratio of the pointer indication position to the full scale of the instrument, and determine the instrument reading.

8. The instrument identification method according to claim 7, characterized in that, The key points of the measurement range include the start and end points of the measurement range, and the key points of the pointer include the pointer rotation axis and the pointer tip; the step of calculating the ratio of the pointer indication position to the full scale of the instrument based on the coordinates of the key points of the pointer and the key points of the measurement range, and determining the instrument reading, includes: Identify the measuring range of the instrument; Based on the coordinates of the starting point of the range, the coordinates of the ending point of the range, the coordinates of the pointer rotation axis, and the coordinates of the pointer tip, calculate the first angle between the pointer and the starting point of the range, and the second angle between the starting point of the range and the ending point of the range; Based on the ratio of the first included angle to the second included angle, determine the ratio of the pointer indication position to the full scale of the instrument; The instrument reading is calculated based on the range and the ratio of the pointer position to the full scale of the instrument.

9. The instrument identification method according to claim 7, characterized in that, The method further includes: If the instrument is a digital instrument, then the digital display area in the front view image of the instrument is detected, and optical character recognition is performed on the digital display area to obtain the instrument reading.

10. A computer device, characterized in that, The device includes a processor and a memory connected to the processor, the memory storing a computer program executable by the processor, the computer program being executed by the processor to implement the steps of the instrument identification method as described in any one of claims 1 to 9.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the instrument identification method as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the instrument identification method as described in any one of claims 1 to 9.