Pointer instrument identification method and system based on multi-modal large model

By fusing image and text features through multimodal large model, the problem of insufficient generalization capability and high complexity in pointer instrument recognition is solved, and the instrument reading recognition with higher accuracy and robustness is achieved.

CN120260050APending Publication Date: 2025-07-04SICHUAN ZHONGDIAN AOSTAR INFORMATION TECHNOLOGIES CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510247215.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When identifying pointer instruments in the prior art, there are problems such as limited generalization capabilities of the model, complex implementation process, and low recognition accuracy and insufficient robustness due to relying on single-modal data.

Method used

The multimodal big model is adopted, through the fusion of image and text features, the pre-trained multimodal big model GLM-4V is fine-tuned, and the identification information is filtered and verified in combination with text description templates and regular expressions to achieve accurate identification of instrument readings.

Benefits of technology

It improves the robustness and accuracy of instrument recognition, reduces the workload of preprocessing and feature engineering, and can adapt to pointer instruments of different models and layouts, quickly locate pointer positions and extract readings, and significantly improves recognition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260050A_ABST
    Figure CN120260050A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-modal large model-based pointer instrument identification method and system. The method comprises the following steps of classifying instruments according to measurement objects, and collecting and preprocessing image data; the method comprises the following steps: marking images, recording key information such as a measuring range, a pointer position and a reading, and constructing an image-text training set; the multi-mode large model is finely adjusted by using the set, so that the multi-mode large model is good at instrument reading identification. After the model outputs key information and readings, the reading accuracy is verified based on the key information, the readings are corrected and updated if necessary, and higher precision is ensured. According to the invention, the instrument reading is calibrated through the instrument key information and the instrument reading is corrected according to the calibration result, so that more accurate pointer type instrument reading identification is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for identifying pointer-type instruments based on a multimodal large model, belonging to the technical field of pattern recognition and image processing. Background Art

[0002] Instruments play a key role in industry. In order to read and record the readings of instruments in various industrial scenarios, in addition to the method of manually reading instrument readings, in more cases, the instrument readings are obtained by combining instrument detection and instrument segmentation technologies. However, these methods have obvious disadvantages, which are specifically as follows: 1. The model generalization ability is limited. Pointer-type instruments of different types in industrial scenarios vary greatly in appearance, size, dial layout, etc. The pointer-type instrument recognition method implemented by a combination of multiple technologies needs to separately adjust the parameters and algorithms required for each technology for each instrument type to ensure that the accuracy of the obtained instrument readings meets industrial requirements. 2. The implementation process is complex. A complete instrument reading recognition method based on deep learning usually includes multiple technologies such as object detection and pointer scale segmentation. Each technology involved needs to be separately parameter-adjusted or trained, and all technology modules need to be coupled and debugged according to logical requirements. In addition, errors will accumulate and time will be consumed due to the connection between different technology modules. 3. Instrument features rely on single-modal data. Usually, the instrument features obtained when training the model in existing instrument reading recognition methods mainly come from instrument images. Compared with multimodal large models that can extract features from multiple modal data such as text and images at the same time, the data features obtained are single, and the results generated by the model are easily affected by the image quality.

[0003] The patent document with the patent number "CN116416620A" discloses a method and system for instrument reading based on semantic segmentation. The problems of this method are as follows: It mainly relies on a semantic segmentation model to extract information such as dials, pointers, and scale lines, and determines the readings through complex image processing steps (such as line detection, angle calculation, text recognition, etc.). This method has extremely high requirements for the accuracy and robustness of the model, and is prone to errors in complex scenarios. It is necessary to perform preprocessing operations such as background removal, cropping, and binarization on the image. These steps increase the complexity and computational cost of the system, and may also cause information loss or misjudgment due to improper preprocessing. Summary of the Invention

[0004] In order to solve the problems existing in the above-mentioned prior art, the present invention proposes a method and system for identifying pointer-type instruments based on a multimodal large model.

[0005] The technical solution of the present invention is as follows:

[0006] On the one hand, the present invention provides a method for identifying pointer-type instruments based on a multi-modal large model, including the following steps:

[0007] Divide the instrument to be measured into several types to be recognized according to the measurement object of the pointer-type industrial instrument, collect images of the instrument to be measured of each type to be recognized, and preprocess the images of the instrument to be measured. The preprocessing includes Contrast Limited Adaptive Histogram Equalization (CLAHE) and sharpening enhancement using the Laplacian of Gaussian operator;

[0008] Fill in the information of the image of the instrument to be measured into the text description template according to the preset text description template to obtain the text description of the image of the instrument to be measured. The information includes the range of the instrument, the scale position where the instrument pointer is located, and the reading of the instrument at this time; Store the name, storage path, and text description of the image of the instrument to be measured in a json file as the training set;

[0009] Use the training set to fine-tune the pre-trained multi-modal large model GLM-4V. After the fine-tuning is completed, it is used to identify pointer-type instruments;

[0010] Obtain the recognition information output by the multi-modal large model GLM-4V. The recognition information includes the instrument reading;

[0011] Verify the recognition information and correct the instrument reading according to the verification result.

[0012] As a preferred implementation manner, the number of images collected for each type of instrument to be measured is not less than 400, and the images of the instrument to be measured of any type to be recognized need to cover different lighting conditions, angles, distances, background environments, and pointer positions.

[0013] As a preferred implementation manner, the fine-tuning steps of the multi-modal large model GLM-4V are as follows:

[0014] Use the encoder ViT to extract the image features of the instrument images to be measured in the training set, and map the image features to the same feature space as the text features through the feature adapter;

[0015] Extract the text features of the training set through the language encoder, fuse the image features and text features based on the attention mechanism fusion method, and input the fused features into the multi-modal large model GLM-4V to complete the model fine-tuning.

[0016] As a preferred implementation manner, design a regular expression according to the text description template to filter the recognition information output by the multi-modal large model GLM-4V;

[0017] Obtain the key information in the filtered recognition information, use the name of the image of the instrument to be measured as the key, and convert the key information into a preset storage format as the value for storage; the key information includes the type of the image of the instrument to be measured, the starting value and the ending value of the instrument range, which small scale between which two large scales the pointer is located on the instrument, how many small scales are there between two large scales, and the reading of the instrument recognized by the multi-modal large model GLM-4V.

[0018] As a preferred implementation, it also includes a verification method for the recognition result:

[0019] Obtain the text description saved in the image of the instrument to be measured, and compare it with the key information output by the multi-modal large model GLM-4V. If they are inconsistent, no verification is performed. If they are consistent, calculate the theoretical value of the instrument:

[0020] Theoretical value of the instrument = large scale of the instrument * 1 + small scale of the instrument * m;

[0021] Where, m represents the difference between small scales;

[0022] Verify the recognition result through the theoretical value of the instrument. The method is:

[0023]

[0024] Update the instrument reading according to the verification result. The update method of the instrument reading is:

[0025]

[0026] On the other hand, the present invention also provides a pointer-type instrument recognition system based on a multi-modal large model, including:

[0027] Data acquisition module: Divide the instruments to be measured into several types to be recognized according to the measurement objects of the pointer-type industrial instruments, collect the images of the instruments to be measured of each type to be recognized, and preprocess the images of the instruments to be measured. The preprocessing includes Contrast Limited Adaptive Histogram Equalization (CLAHE) and sharpening enhancement with Gaussian-Laplacian operator;

[0028] Data processing module: Fill the information of the image of the instrument to be measured into the text description template according to the preset text description template to obtain the text description of the image of the instrument to be measured. The information includes the range of the instrument, the scale position where the instrument pointer is located, and the reading of the instrument at this time; store the name, storage path and text description of the image of the instrument to be measured in a json file as a training set;

[0029] Model recognition module: Fine-tune the pre-trained multi-modal large model GLM-4V using the training set, and after the fine-tuning is completed, use it to recognize the pointer-type instrument;

[0030] Recognition and verification module: Obtain the recognition information output by the multimodal large model GLM-4V, where the recognition information includes instrument readings; verify the recognition information and correct the instrument readings according to the verification result.

[0031] As a preferred implementation, the number of images collected for each type of instrument to be recognized is not less than 400, and the images of the instruments to be recognized for any type need to cover different lighting conditions, angles, distances, background environments, and pointer positions.

[0032] As a preferred implementation, the fine-tuning steps of the multimodal large model GLM-4V are as follows:

[0033] Use the encoder ViT to extract the image features of the instrument images to be measured in the training set, and map the image features to the same feature space as the text features through a feature adapter;

[0034] Extract the text features of the training set through a language encoder, fuse the image features and text features based on the attention mechanism fusion method, and input the fused features into the multimodal large model GLM-4V to complete model fine-tuning.

[0035] As a preferred implementation, design a regular expression according to the text description template to filter the recognition information output by the multimodal large model GLM-4V;

[0036] Obtain the key information in the filtered recognition information, use the name of the instrument image to be measured as the key, and convert the key information into a preset storage format as the value for storage; the key information includes the type of the instrument image to be measured, the starting value and ending value of the instrument range, which small scale between which two large scales the pointer is located, how many small scales are there between two large scales, and the reading of the instrument recognized by the multimodal large model GLM-4V.

[0037] As a preferred implementation, it also includes a verification method for the recognition result:

[0038] Obtain the text description saved in the instrument image to be measured, and compare it with the key information output by the multimodal large model GLM-4V. If they are inconsistent, no verification is performed. If they are consistent, calculate the theoretical value of the instrument:

[0039] Theoretical value of the instrument = large scale of the instrument * 1 + small scale of the instrument * m;

[0040] Where m represents the difference between small scales;

[0041] Verify the recognition result through the theoretical value of the instrument. The method is:

[0042]

[0043] Update the instrument reading according to the verification result. The method for updating the instrument reading is as follows:

[0044]

[0045] The present invention has the following beneficial effects:

[0046] Through the multi-modal large model, the present invention can fuse image and text information, and the recognition of pointer-type instruments no longer depends on a single image feature, thus showing higher robustness in complex environments such as uneven illumination, complex backgrounds, or blurred pointers. Through large-scale data pre-training, the model can adapt to pointer-type instruments of different models and layouts without the need for separate optimization for each instrument. The multi-modal large model can automatically learn the deep features of images and texts without the need to manually design complex feature extraction algorithms, reducing the workload of preprocessing and feature engineering. Through multi-modal fusion, the model can quickly locate the pointer position and extract the reading, and the recognition efficiency is significantly improved compared with traditional methods. Verify the instrument reading through the key information of the instrument and correct the instrument reading according to the verification result to achieve more accurate recognition of pointer-type instrument readings. Description of the Drawings

[0047] Figure 1 It is a flowchart of the method implementation of the present invention. Detailed Embodiments

[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0049] It should be understood that the step numbers used in the text are only for convenience of description and do not limit the execution order of the steps.

[0050] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0051] The terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.

[0052] The term "and / or" refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0053] Embodiment 1:

[0054] Refer to Figure 1 , the present invention provides a pointer-type instrument recognition method based on a multimodal large model, including the following steps:

[0055] Divide the instrument to be measured into several types to be recognized according to the measurement object of the pointer-type industrial instrument, and the types to be recognized include gas pressure gauges and wattmeters; collect images of the instruments to be measured of each type to be recognized, and preprocess the images of the instruments to be measured, and the preprocessing includes Contrast Limited Adaptive Histogram Equalization (CLAHE) and sharpening enhancement with the Laplacian of Gaussian operator;

[0056] Fill in the information of the image of the instrument to be measured into the text description template according to the preset text description template to obtain the text description of the image of the instrument to be measured, and the information includes the range of the instrument, the scale position where the instrument pointer is located, and the reading of the instrument at this time; store the name, storage path and text description of the image of the instrument to be measured in a json file as a training set;

[0057] Use the training set to fine-tune the pre-trained multimodal large model GLM-4V, and after the fine-tuning is completed, use it to recognize the pointer-type instrument;

[0058] Obtain the recognition information output by the multimodal large model GLM-4V, and the recognition information includes the instrument reading;

[0059] Verify the recognition information, and correct the instrument reading according to the verification result.

[0060] In this embodiment, the preset text description template is:

[0061]

[0062] As a preferred implementation method, the number of images collected for each type of instrument to be measured is not less than 400, and the images of the instruments to be measured of any type to be recognized need to cover different lighting conditions, angles, distances, background environments, and pointer positions.

[0063] As a preferred implementation method, the fine-tuning steps of the multimodal large model GLM-4V are:

[0064] Use the encoder ViT to extract the image features of the images of the instruments to be measured in the training set, and map the image features to the same feature space as the text features through a feature adapter;

[0065] Extract the text features of the training set through a language encoder, fuse the image features and text features based on the attention mechanism fusion method, and input the fused features into the multimodal large model GLM-4V to complete model fine-tuning.

[0066] As a preferred implementation, design a regular expression according to the text description template to filter the recognition information output by the multimodal large model GLM-4V;

[0067] Obtain the key information in the filtered recognition information, use the name of the image of the instrument to be measured as the key, and convert the key information into a preset storage format as the value for storage; the key information includes the type of the image of the instrument to be measured, the starting value and ending value of the instrument range, which small scale between which two large scales the pointer is located, how many small scales are there between two large scales, and the reading of the instrument recognized by the multimodal large model GLM-4V.

[0068] The preset storage format in this embodiment is:

[0069] {

[0070] "Instrument category": "What kind of instrument",

[0071] "Starting value of instrument range": "Value",

[0072] "Ending value of instrument range": "Value",

[0073] "Large scale 1": "The smaller value of the two large scales",

[0074] "Large scale 2": "The larger value of the two large scales",

[0075] "Number of small scales": "How many small scales are there between two large scales",

[0076] "Small scale": "Which small scale between two large scales",

[0077] "Instrument reading": "The instrument reading obtained by the multimodal large model"

[0078] As a preferred implementation, it also includes a verification method for the recognition result:

[0079] Obtain the text description saved in the image of the instrument to be measured, and compare it with the key information output by the multimodal large model GLM-4V. If they are inconsistent, no verification is performed. If they are consistent, calculate the theoretical value of the instrument:

[0080] Theoretical value of the instrument = instrument large scale * 1 + instrument small scale * m;

[0081] Where m represents the difference between small scales;

[0082] Verify the recognition result through the theoretical value of the instrument. The method is as follows:

[0083]

[0084] Update the instrument reading according to the verification result. The method for updating the instrument reading is as follows:

[0085]

[0086] Embodiment 2:

[0087] The present invention also provides a pointer-type instrument recognition system based on a multimodal large model, including:

[0088] Data acquisition module: Divide the instrument to be measured into several types to be recognized according to the measurement object of the pointer-type industrial instrument, collect the images of the instrument to be measured of each type to be recognized, and preprocess the images of the instrument to be measured. The preprocessing includes Contrast Limited Adaptive Histogram Equalization (CLAHE) and sharpening enhancement using the Laplacian of Gaussian operator;

[0089] Data processing module: Fill in the information of the image of the instrument to be measured into the text description template according to the preset text description template to obtain the text description of the image of the instrument to be measured. The information includes the range of the instrument, the scale position where the instrument pointer is located, and the reading of the instrument at this time; Store the name, storage path, and text description of the image of the instrument to be measured in a json file as the training set;

[0090] Model recognition module: Fine-tune the pre-trained multimodal large model GLM-4V using the training set, and after the fine-tuning is completed, use it to recognize the pointer-type instrument;

[0091] Recognition verification module: Obtain the recognition information output by the multimodal large model GLM-4V. The recognition information includes the instrument reading; Verify the recognition information and correct the instrument reading according to the verification result.

[0092] As a preferred implementation method, the number of images collected for each type of instrument to be measured is not less than 400, and the images of the instrument to be measured of any type to be recognized need to cover different lighting conditions, angles, distances, background environments, and pointer positions.

[0093] As a preferred implementation method, the fine-tuning steps of the multimodal large model GLM-4V are as follows:

[0094] Use the encoder ViT to extract the image features of the images of the instrument to be measured in the training set, and map the image features to the same feature space as the text features through the feature adapter;

[0095] Extract the text features of the training set through a language encoder, fuse the image features and text features based on the attention mechanism fusion method, and input the fused features into the multimodal large model GLM-4V to complete model fine-tuning.

[0096] As a preferred implementation, design a regular expression according to the text description template to filter the recognition information output by the multimodal large model GLM-4V;

[0097] Obtain the key information in the filtered recognition information, use the name of the image of the instrument to be measured as the key, and convert the key information into a preset storage format as the value for storage; the key information includes the type of the image of the instrument to be measured, the starting value and ending value of the instrument range, which small scale between which two large scales the pointer is located, how many small scales are there between two large scales, and the reading of the instrument recognized by the multimodal large model GLM-4V.

[0098] As a preferred implementation, it also includes a verification method for the recognition result:

[0099] Obtain the text description saved for the image of the instrument to be measured, and compare it with the key information output by the multimodal large model GLM-4V. If they are inconsistent, no verification is performed. If they are consistent, calculate the theoretical value of the instrument:

[0100] Theoretical value of the instrument = large scale of the instrument * 1 + small scale of the instrument * m;

[0101] Where m represents the difference between small scales;

[0102] Verify the recognition result through the theoretical value of the instrument. The method is:

[0103]

[0104] Update the instrument reading according to the verification result. The update method of the instrument reading is:

[0105]

[0106] In the embodiments of the present application, "at least one" means one or more, and "multiple" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent the situation of A existing alone, A and B existing simultaneously, and B existing alone. Where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.

[0107] Those of ordinary skill in the art can realize that the various units and algorithm steps described in the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0108] Those skilled in the art can clearly understand that for the sake of convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0109] In several embodiments provided in this application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM for short), random access memories (RAM for short), magnetic disks, or optical discs that can store program codes.

[0110] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A method for identifying pointer-type instruments based on a multi-modal large model, characterized in that It includes the following steps: Divide the instrument to be measured into several types to be recognized according to the measurement object of the pointer industrial instrument, collect the images of the instrument to be measured of each type to be recognized, and preprocess the images of the instrument to be measured. The preprocessing includes Contrast Limited Adaptive Histogram Equalization (CLAHE) and sharpening enhancement using the Laplacian of Gaussian operator; Fill in the information of the image of the instrument to be measured into the text description template according to the preset text description template to obtain the text description of the image of the instrument to be measured. The information includes the range of the instrument, the scale position where the instrument pointer is located, and the reading of the instrument at this time; Store the name, storage path, and text description of the image of the instrument to be measured in a json file as the training set; Use the training set to fine-tune the pre-trained multimodal large model GLM-4V. After the fine-tuning is completed, it is used to recognize the pointer instrument; Obtain the recognition information output by the multimodal large model GLM-4V. The recognition information includes the instrument reading; Verify the recognition information and correct the instrument reading according to the verification result.

2. The method for identifying pointer-type instruments based on a multi-modal large model according to claim 1, wherein The number of images collected for each type of instrument to be measured is not less than 400, and the images of the instrument to be measured of any type to be recognized need to cover different lighting conditions, angles, distances, background environments, and pointer positions.

3. The pointer-type instrument recognition method based on a multi-modal large model according to claim 1, characterized in that The fine-tuning steps of the multimodal large model GLM-4V are: Use the encoder ViT to extract the image features of the images of the instrument to be measured in the training set, and map the image features to the same feature space as the text features through the feature adapter; Extract the text features of the training set through the language encoder, fuse the image features and text features based on the attention mechanism fusion method, and input the fused features into the multimodal large model GLM-4V to complete the model fine-tuning.

4. The method for identifying pointer-type instruments based on a multi-modal large model according to claim 1, wherein, Design a regular expression according to the text description template to filter the recognition information output by the multimodal large model GLM-4V; Obtain the key information in the filtered recognition information, use the name of the image of the instrument to be measured as the key, and convert the key information into a preset storage format as the value for storage; The key information includes the type of the image of the instrument to be measured, the starting value and ending value of the instrument range, which small scale between which two large scales the pointer is located, how many small scales are there between two large scales, and the reading of the instrument recognized by the multimodal large model GLM-4V.

5. The pointer-type instrument recognition method based on a multi-modal large model according to claim 4, wherein The verification method of the recognition result: Obtain the text description saved in the image of the instrument to be measured, and compare it with the key information output by the multimodal large model GLM-4V. If they are inconsistent, no verification is performed. If they are consistent, calculate the theoretical value of the instrument: Theoretical value of the instrument = large scale of the instrument * 1 + small scale of the instrument * m; Where, m represents the difference between small scales; Verify the recognition result through the theoretical value of the instrument. The method is: Update the instrument reading according to the verification result. The update method of the instrument reading is:

6. A pointer-type instrument recognition system based on a multimodal large model, characterized in that, It includes: Data acquisition module: Divide the instrument to be measured into several types to be recognized according to the measurement object of the pointer industrial instrument, collect the images of the instrument to be measured of each type to be recognized, and preprocess the images of the instrument to be measured. The preprocessing includes Contrast Limited Adaptive Histogram Equalization (CLAHE) and sharpening enhancement using the Laplacian of Gaussian operator; Data processing module: Fill the information of the instrument image to be measured into the text description template according to the preset text description template to obtain the text description of the instrument image to be measured, where the information includes the range of the instrument, the scale position where the instrument pointer is located, and the reading of the instrument at this time; Store the name, storage path, and text description of the instrument image to be measured in a json file as a training set; Model recognition module: Use the training set to fine-tune the pre-trained multi-modal large model GLM-4V, and after the fine-tuning is completed, use it to recognize pointer-type instruments; Recognition verification module: Obtain the recognition information output by the multi-modal large model GLM-4V, where the recognition information includes the instrument reading; Verify the recognition information and correct the instrument reading according to the verification result.

7. The pointer-type instrument recognition system based on a multi-modal large model according to claim 6, characterized in that, The number of instrument images to be measured for each type to be recognized is not less than 400, and the instrument images to be measured of any type to be recognized need to cover different lighting conditions, angles, distances, background environments, and pointer positions.

8. The pointer-type instrument recognition system based on a multi-modal large model according to claim 6, wherein The fine-tuning steps of the multi-modal large model GLM-4V are as follows: Use the encoder ViT to extract the image features of the instrument images to be measured in the training set, and map the image features to the same feature space as the text features through a feature adapter; Extract the text features of the training set through a language encoder, fuse the image features and text features based on the attention mechanism fusion method, and input the fused features into the multi-modal large model GLM-4V to complete the model fine-tuning.

9. The pointer-type instrument recognition system based on a multi-modal large model according to claim 6, characterized in that Design a regular expression according to the text description template to filter the recognition information output by the multi-modal large model GLM-4V; Obtain the key information in the filtered recognition information, use the name of the instrument image to be measured as the key, and convert the key information into a preset storage format as the value for storage; the key information includes the type of the instrument image to be measured, the starting value and ending value of the instrument range, which small scale between which two large scales the pointer is located on the instrument, how many small scales are there between two large scales, and the reading of the instrument recognized by the multi-modal large model GLM-4V.

10. The pointer-type instrument recognition system based on a multi-modal large model according to claim 9, wherein It also includes a method for verifying the recognition result: Obtain the text description saved in the instrument image to be measured and compare it with the key information output by the multi-modal large model GLM-4V. If they are inconsistent, no verification is performed. If they are consistent, calculate the theoretical value of the instrument: Theoretical value of the instrument = large scale of the instrument * 1 + small scale of the instrument * m; where m represents the difference between small scales; Verify the recognition result through the theoretical value of the instrument. The method is as follows: Update the instrument reading according to the verification result. The update method of the instrument reading is:

Citation Information

Cited By

  • General instrument image processing method and device and electronic equipment

    CN120976908A