Intelligent inspection alarm explainable method and system based on multi-modal large model multi-agent

CN122616751APending Publication Date: 2026-08-21HANGZHOU HARMONYCLOUD TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610785752.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

[0004]在工业设备智能运维与巡检领域,传统方案存在显著技术瓶颈:一方面,人工巡检依赖经验判断,效率低且易受环境干扰,而单一基于深度学习的视觉识别模型仅能完成仪表读数提取,无法关联业务场景进行告警分析;另一方面,现有智能巡检系统多采用 “识别- 告警” 的线性流程,缺乏多模态信息融合能力,难以处理仪表图像、设备文本台账等跨模态数据,也无法通过业务知识解释告警原因,导致运维人员仅知 “告警” 而不知 “为何告警、影响范围”

Benefits of technology

1、以多模态数据(仪表图像、文本)为基础,通过大模型微调实现仪表类型与数值的精准识别并输出结构化结果;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122616751A_ABST
    Figure CN122616751A_ABST
Patent Text Reader

Abstract

The application discloses a kind of intelligent inspection alarm explainable method and system based on multimodal large model multi-agent, belong to computer technology field;The method comprises: collecting multimodal data;The multimodal data is preprocessed, and the multimodal data after preprocessing is obtained;The multimodal data after preprocessing is enhanced and expanded, and the multimodal data after enhancement and expansion is obtained;The multimodal data after enhancement and expansion is identified, and the instrument reading is obtained;The multimodal data after enhancement and expansion and instrument reading are input into multimodal large model and are fine-tuned training, and the fine-tuned model is obtained;Based on fine-tuned model, instrument type identification agent and reading identification agent are constructed;The multimodal data to be identified is input into instrument type identification agent and reading identification agent, and multimodal recognition result is obtained;The application is based on multimodal data, and the accurate identification of instrument type and value is realized by large model fine-tuning, and structured result is output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, specifically to an interpretable intelligent inspection and alarm method based on a multimodal large model and multiple agents. Background Technology

[0002] With the development of large-scale intelligent technology, the scale of industrial equipment and the complexity of inspection scenarios are constantly expanding. At the same time, the demand for inspection efficiency, alarm accuracy and interpretability of results is also constantly increasing. It is particularly important to efficiently complete the intelligent inspection of large-scale industrial equipment, accurately determine the alarm status and clearly explain the cause and scope of the alarm.

[0003] In recent years, the development and increasing maturity of multimodal large-scale models and multi-agent technologies have provided new solutions to the core needs of complex industrial inspection scenarios. Based on multimodal large-scale models, key inspection objects such as industrial instruments are uniformly identified, achieving accurate extraction of instrument types and values. Furthermore, by combining multi-agent collaborative technology with modules such as rule engine alarm judgment, RAG knowledge base retrieval enhancement generation, and online incremental learning based on user feedback, multimodal information fusion processing, automated scheduling of inspection processes, accurate matching of alarm rules, and continuous optimization of system performance are achieved. This enables efficient inspection, accurate alarming, and interpretable analysis of large-scale industrial equipment. Multimodal large-scale models and multi-agent frameworks possess mature cross-modal understanding and collaborative decision-making capabilities. Due to their powerful semantic reasoning, flexible module linkage capabilities, and comprehensive technological ecosystem, they demonstrate significant advantages in the field of industrial intelligent inspection. More and more industrial enterprises are gradually upgrading their intelligent inspection systems based on multimodal large-scale models and multi-agent technologies. Building intelligent inspection alarm interpretable systems based on multimodal large-scale model multi-agent technology is an inevitable trend.

[0004] In the field of intelligent operation and maintenance and inspection of industrial equipment, traditional solutions have significant technical bottlenecks: On the one hand, manual inspection relies on experience-based judgment, which is inefficient and easily affected by environmental interference. On the other hand, a single deep learning-based visual recognition model can only extract instrument readings and cannot be associated with business scenarios for alarm analysis. On the other hand, existing intelligent inspection systems mostly adopt a linear process of "recognition-alarm", lacking the ability to fuse multimodal information. They are unable to process cross-modal data such as instrument images and equipment text ledgers, and cannot explain the cause of alarms through business knowledge. As a result, maintenance personnel only know "alarm" but do not know "why the alarm occurred and the scope of impact".

[0005] Meanwhile, the existing systems mostly use statically deployed models, which cannot dynamically optimize recognition accuracy based on user feedback. They are not robust enough in the face of extreme conditions such as complex lighting and occlusion. Alarm judgment relies only on simple threshold rules and does not combine business knowledge base for scenario-based reasoning. Furthermore, each functional module (recognition, rule matching, and analysis) is independent of each other and lacks a collaborative decision-making mechanism. The final output results lack interpretability and are difficult to support accurate operation and maintenance decisions in industrial scenarios.

[0006] In addition, the demand for "interpretability" of inspection results in industrial settings is becoming increasingly urgent. However, existing solutions cannot logically correlate information such as instrument readings, alarm rules, and business impacts, making it difficult for maintenance personnel to quickly locate the root cause of problems and hindering the application of intelligent inspection systems in complex industrial environments. Summary of the Invention

[0007] The purpose of this invention is to provide an intelligent inspection alarm interpretable system based on a multimodal large model and multiple agents. Based on a multimodal large model fine-tuned using actual industrial data, it can automatically identify instrument types and values, retrieve MySQL rule bases to logically complete alarm determination, and finally combine a pre-built knowledge base to complete the explanation of business scenarios and impact scope analysis. Simultaneously, a user feedback loop is introduced, high-quality labeled samples are built, and the model is incrementally optimized online, achieving system self-evolution and long-term stable availability. To address the aforementioned technical problems, this invention provides an interpretable intelligent inspection and alarm method based on a multimodal large model and multiple agents, comprising the following steps: Collect multimodal data; The multimodal data is preprocessed to obtain preprocessed multimodal data; The preprocessed multimodal data is enhanced and augmented to obtain enhanced and augmented multimodal data; The enhanced and expanded multimodal data are identified to obtain instrument readings; The enhanced and expanded multimodal data and instrument readings are input into the large multimodal model for fine-tuning training to obtain the fine-tuned model; Based on the fine-tuning model, construct intelligent agents for instrument type recognition and reading recognition; The multimodal data to be identified is input into the instrument type identification agent and the reading identification agent to obtain the multimodal identification result; the multimodal identification result includes the type label and the final reading; The multimodal recognition results are input into the alarm determination agent to determine the alarm information; the alarm information includes whether an alarm has been triggered and the alarm level. The alarm information is interpreted by a knowledge base retrieval and an interpretable generative agent, and then sent to the relevant users.

[0008] Preferably, the multimodal data is preprocessed to obtain preprocessed multimodal data, specifically including the following steps: The multimodal data is standardized, cut into uniform size specifications, and then formatted to obtain preprocessed multimodal data.

[0009] Preferably, the preprocessed multimodal data is enhanced and augmented to obtain enhanced and augmented multimodal data, specifically including the following steps: The preprocessed multimodal data is parametrically modeled to simulate the occlusion effects of overexposure, shadows, and highlight glare on the scale and pointer by considering ambient light intensity, color temperature, main light direction, and specular reflection from the instrument glass cover. Based on the mapping relationship between the instrument range scale and pointer angle in the preprocessed multimodal data, a pointer posture consistent with the target reading is generated, and the reading value, pointer angle, key points, and visibility mask are output simultaneously to obtain the instrument reading. Through the mechanism of pointer self-occlusion, external occlusion, and local scale loss caused by perspective projection, a difficult sample set covering extreme working conditions is generated to obtain enhanced and expanded multimodal data.

[0010] Preferably, the fine-tuning training of the multimodal large model specifically includes the following steps: Instrument type classification: The model takes enhanced and expanded multimodal data as the main input and integrates text information to output predefined instrument type categories. It is supervised by cross-entropy loss to enable the model to stably distinguish instruments with different structures and scale shapes, and output structured JSON output. Numerical Recognition Regression: Based on the enhanced and expanded multimodal data, the ROI of the instrument panel is selected as the input, and the continuous reading value corresponding to the pointer is output. A regression head is added to the top layer of the model to predict the continuous value of the fused visual features, and SmoothL1 / Huber regression loss is used to improve the robustness to noisy labels and difficult samples. The regression target is designed as the normalized reading and combined with the range information to map back to the actual value. For dual-needle and dual-ring instruments, it is expanded to a multi-output regression head to output the final reading of the instrument.

[0011] Preferably, the multimodal data to be identified is input into the instrument type identification agent and the reading identification agent to obtain the multimodal identification result; specifically, it includes the following steps: The multimodal data to be identified and the corresponding user text input type recognition agent are used to determine the type label; The multimodal data to be identified and the type label are input into the reading recognition agent for identification, and the final reading is obtained.

[0012] Preferably, the multimodal recognition results are input into the alarm determination agent to determine the alarm information, specifically including the following steps: Input the type label, final reading and corresponding user text into the alarm determination agent, retrieve applicable rules from the MySQL database and perform multi-level threshold calculations to determine whether an alarm is triggered and the alarm level. The MySQL database includes multi-level rules. For a matched rule, the alarm decision agent compares the readings with the threshold range defined by the rule and the pre-defined prompt rule. Interval thresholds: Triggering an alarm when the threshold exceeds [min, max], with different levels corresponding to segmented intervals; otherwise, no alarm is triggered. Upper limit thresholds: reading_value >= the maximum value of the normal threshold, triggering a low-level warning; reading_value >= 5% of the maximum value of the normal threshold, triggering a general warning; reading_value >= 10% of the maximum value of the normal threshold, triggering a severe warning. Lower threshold: reading_value <= minimum normal threshold, triggers low-level warning; reading_value <= 5% of minimum normal threshold, triggers general warning; reading_value <= 10% of minimum normal threshold, triggers severe warning.

[0013] Preferably, the alarm information is interpreted through knowledge base retrieval and an interpretable generative agent, and then sent to the corresponding user. This specifically includes the following steps: The knowledge base retrieval and interpretable generative agent orchestration layer packages the user-input business scenario text, multimodal recognition results, and alarm information into unified input information for the knowledge base retrieval and interpretable generative agent; Structure and vectorize the knowledge base; Input information is fed into the retrieval module for retrieval: The business scenario text input by the user, the multimodal recognition results, and the alarm information are combined into a semantic query vector to recall TopK candidate fragments; the device, model, and fault keywords of the business scenario text input by the user are used for inverted index retrieval; the results of multi-path recall are obtained, and the results of multi-path recall are merged and deduplicated to form candidate evidence; The candidate evidence is filtered out by the reordering and relevance judgment module to obtain the filtered evidence by filtering low-relevance segments from the set threshold. Based on the interpretable generation phase generation module, alarm explanations are generated from filtered evidence. Send alarm information and alarm explanations to the user.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Based on multimodal data (instrument images, text), the instrument type and values ​​are accurately identified and structured results are output through fine-tuning of a large model; 2. An innovative multi-agent collaborative architecture is designed, which breaks down the recognition task into two specialized intelligent agents, classification and reading, which are executed sequentially to enhance recognition accuracy; 3. Integrating rule base judgment and knowledge base retrieval capabilities, alarm level determination is completed through multi-level threshold rules, and interpretable conclusions supported by evidence are generated by combining knowledge from operation and maintenance documents, historical cases, etc. 4. Construct a user feedback-driven incremental learning mechanism to continuously optimize the model's recognition and inference capabilities through data cleaning, quality assessment, and incremental training. Attached Figure Description

[0015] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0016] Figure 1 This is a flowchart illustrating an intelligent inspection and interpretable alarm method based on a multimodal large model and multiple agents. Figure 2 This is a flowchart illustrating the multi-agent collaborative identification module; Figure 3 This is a flowchart illustrating the process of the rule base and the alarm determination agent; Figure 4 This is a flowchart illustrating the user feedback and incremental learning process. Detailed Implementation

[0017] Numerous specific details are set forth in the following description to provide a full understanding of the invention. However, the invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0018] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0019] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0020] The present invention will now be described in further detail with reference to the accompanying drawings: To better illustrate the technical effects of the present invention, the present invention provides the following specific embodiments to illustrate the above technical process: Example 1, such as Figure 1 As shown, the intelligent inspection and interpretable alarm system based on a multimodal large model and multiple agents mainly includes the following steps: Step 1: Fine-tuning of the multimodal large model. Collect equipment data and construct multimodal data such as instrument images and text. Through fine-tuning training, enable the basic model to have the ability to classify instrument types and accurately identify values. Finally, output the instrument type and current value in JSON structured results.

[0021] Step 2: Construct a multi-agent collaborative recognition module, decompose the recognition task into classification agents and reading agents, and execute multiple agents serially to improve recognition accuracy.

[0022] Step 3: Rule base and alarm judgment agent. Based on the rule table of equipment, instruments, operating conditions and readings, multi-level threshold rule calculation is completed, and the results are used together with the output of the recognition agent to determine whether an alarm is triggered and the alarm level.

[0023] Step 4: Knowledge base retrieval and interpretable intelligent agent generation. Based on the business scenario and equipment information input by the user, and combined with some equipment descriptions, operation manuals, historical fault cases, etc. from the historical operation and maintenance process, vectorized retrieval and reordering are performed to output citationable evidence fragments. By integrating the recognition results, rule judgments and knowledge evidence, interpretable conclusions of equipment instruments are output.

[0024] Step 5: User feedback and incremental learning. Record user confirmations or corrections, perform data cleaning, quality assessment, and sample construction, and continuously optimize the recognition and reasoning capabilities of the large model using an incremental training strategy.

[0025] I. Fine-tuning of a multimodal large model Multimodal large model fine-tuning is mainly divided into two parts: one is the construction and enhancement of the dataset, and the other is the fine-tuning of the model.

[0026] 1. Data preprocessing: Multimodal data from all devices is collected via camera capture. All image data is input into the image preprocessing module, which standardizes the instruments in the images, crops them to a uniform size, and formats them into images that conform to the input specifications of the multimodal large model.

[0027] After the basic data is constructed, data enhancement and expansion are performed. The collected basic data only meets the lighting environment under the current shooting conditions. It is also necessary to expand the original dataset by employing controllable data enhancement methods oriented towards physical imaging and instrument geometry to simulate different states of instruments in the simulated field, making the dataset cover as many situations as possible in the working environment. This method simulates the occlusion effects of overexposure, shadows, and specular reflection on the scale and pointer by parametrically modeling optical factors such as ambient light intensity, color temperature, main light direction, and specular reflection from the instrument's glass cover. Simultaneously, based on the mapping relationship between the instrument's range scale and pointer angle, a pointer posture consistent with the target reading is generated, and annotation information such as the reading value, pointer angle, key points, and visibility mask is output synchronously. Furthermore, through mechanisms such as pointer self-occlusion, external occlusion, and local scale loss caused by perspective projection, a difficult sample set covering extreme working conditions is generated. Through these enhancement methods, high-coverage training samples can be constructed without relying on a large amount of manual annotation, effectively improving the model's robustness and generalization ability in reading recognition under complex lighting, occlusion, and changing viewing angle conditions.

[0028] 2. The final generated image data, the corresponding instrument type labels, instrument readings, and text content are combined into a training dataset for a multimodal large model. A "freeze the base and fine-tune" strategy is adopted to freeze the underlying parameters of the model's visual encoder and language model, and only fine-tune the top-level task head related to the task, so that the model can quickly adapt to industry tasks of "instrument type + reading".

[0029] The training task was broken down into two sub-tasks and completed using a phased training approach.

[0030] The first task is instrument type classification: taking instrument images as the main input and fusing text information, the output is a predefined instrument type category, specifically including single-circle single-pointer instruments, double-circle dual-pointer instruments, and double-circle single-pointer instruments. Ultimately, a classification head is added on top of the multimodal fusion features, and cross-entropy loss is used for supervision, enabling the model to stably distinguish instruments with different structures and scale shapes, and outputting structured JSON output to provide structural priors for subsequent readings.

[0031] The second task is numerical recognition regression: using the instrument image as a baseline, the ROI of the instrument dial is selected as input, and the continuous reading value corresponding to the pointer is output. A regression head is added to the top layer of the model to predict continuous values ​​of the fused visual features, and SmoothL1 / Huber regression loss is used to improve robustness to noisy annotations and difficult samples. To enhance the generalization ability across measurement ranges, the regression objective is designed as a normalized reading combined with range information to map back to the actual value; for dual-needle and dual-ring instruments, it is expanded to a multi-output regression head to improve accuracy and convergence efficiency, and finally outputs the final reading of the instrument.

[0032] II. Constructing a Multi-Agent Collaborative Identification Module The multi-agent collaborative identification module can be divided into two core intelligent agents: the instrument type identification agent (TypeAgent) and the reading identification agent (Reading Agent). The specific process is as follows: Figure 2 As shown, the overall process is as follows: Multimodal input (mainly images, supplemented by text) first enters the Type Agent to complete end-to-end type determination; "type label + original image" is passed to the Reading Agent as structured input; The Reading Agent works when it obtains the accurate instrument type and outputs the final reading according to the rules set by the system prompt based on clarity priority. The type determination and the reading rules are independent of each other, and the reading rules are enforced in the Reading Agent.

[0033] (1) Prompt Construction and Role Isolation. The Type Agent's Prompt should "only classify, not interpret," and its output should be strictly limited to the pre-defined label set to ensure downstream routing stability. The Reading Agent's Prompt needs to fix the business rules as hard constraints: only process instruments of the preparation type, perform clarity verification first, read only the inner circle for double-circle instruments, read only the red pointer for double-pointer instruments, output only numerical values ​​or fixed failure text, and explicitly "ignore non-type directly." Through such strong constraint prompts, the model illusion, misreading, and multiple outputs caused by the model acting on its own on complex images are reduced.

[0034] (2) Type Agent Design: The input is image and text, and the output is a single type_label. It is recommended to limit the type label to: double circle single pointer / single circle single pointer / double circle double pointer / type cannot be determined. The Type Agent only makes the judgment of "what kind of table is this", does not read the data, and does not output any explanatory text, thereby ensuring that its output can be parsed programmatically and routed reliably.

[0035] (3) Reading Agent: Receives type_label and image. When type_label does not belong to the set instrument label type, it is ignored and an empty string is output without any recognition. If the type is the specified instrument type, "clarity verification priority" is strictly enforced. If any condition is found, such as indistinguishable inner scale, unclear pointer, or severe distortion due to reflection, the value is directly output as unreadable, and estimation is prohibited. The value of the pointer on the inner scale is only read when the clarity is satisfied, and the outer scale information is completely ignored; the final output must only contain the inner scale value itself, without units or explanations.

[0036] III. Rule Base and Alarm Judgment Agent In addition to the existing "Type Agent → Reading Agent" link, a rule base and an alarm determination agent (Alarm Agent) are added in parallel: First, the meter type and reading are identified. Then, the "reading + user text (device type and meter logic type)" are input into the Alarm Agent. Applicable rules are retrieved from the MySQL database, multi-level threshold calculations are performed, and finally, "whether an alarm is triggered + alarm level" is output. The overall process is as follows: Figure 3 As shown.

[0037] (1) Multi-level alarm rule base creation. The rules are broken down into fields such as "Device description + Normal threshold definition + Red / yellow / blue alarm threshold below normal range + Low / normal / critical alarm threshold above normal range". The threshold definition supports multiple levels.

[0038] (2) Rule retrieval. The device_id and device_type used as matching conditions are parsed from the text input by the user. Based on the parsed values, the matching alarm rules are filtered from the rule base using a database retrieval tool.

[0039] (3) Multi-level threshold logic judgment. The reading of the instrument value is received from the Reading Agent. The agent will only output two forms: a numerical string or a fixed failure message. The Alarm Agent needs to convert it into a usable structure: reading_value. If the reading is unreadable, the Alarm Agent will not perform threshold comparison.

[0040] For a matched rule, the Alarm Agent compares the readings based on the threshold range defined by the rule and the pre-defined prompt rule: Interval threshold: Triggered when exceeding [min, max], with different levels corresponding to segmented intervals; otherwise, no alarm.

[0041] Upper limit thresholds: reading_value >= the maximum value of the normal threshold, triggering a low-level warning; reading_value >= 5% of the maximum value of the normal threshold, triggering a general warning; reading_value >= 10% of the maximum value of the normal threshold, triggering a severe warning. Lower threshold: reading_value <= minimum normal threshold, triggers low-level warning; reading_value <= 5% of minimum normal threshold, triggers general warning; reading_value <= 10% of minimum normal threshold, triggers severe warning.

[0042] IV. Knowledge Base Retrieval and Explainable Agent Generation Add a "Knowledge Base Retrieval and Explainable Generation Agent (ExplainAgent)" after AlarmAgent. Its goal is not to re-evaluate alarms, but rather, given that AlarmAgent has already provided the alarm level and hit rules, it retrieves and aligns "user business scenario + device and instrument information + identification readings + alarm rule results" with internal enterprise knowledge, outputs citationable evidence fragments, and generates explanatory conclusions for on-site and maintenance personnel. These conclusions include why the alarm occurred, possible causes, recommended handling methods, and what additional information is needed. The overall flowchart is shown in Figure 4.

[0043] The overall process is divided into four steps: input aggregation, retrieval recall, re-ranking and relevance judgment, and interpretable generation.

[0044] 1. The orchestration layer packages the user-input business scenario text, multimodal recognition results (meter type, reading), and AlarmAgent output (whether an alarm is triggered, alarm level, hit rules, threshold information) into a unified input for the Explain Agent. This input is standardized according to a predefined prompt and forms a "retrieval query description," which is used to locate the most relevant knowledge evidence in the knowledge base later.

[0045] 2. The knowledge base needs to be structured and vectorized in advance. Data sources include: equipment manuals and operation manuals, standard operating procedures, historical maintenance work orders and fault cases, training materials and experience summaries, etc. The knowledge base also establishes two types of indexes: one is a vector index for semantic recall, solving the matching of "synonymous expressions and colloquial descriptions", and the other is a keyword / structured index for precise filtering, solving the precise positioning of "model, tag number, chapter, terminology", thereby ensuring that the retrieval is both comprehensive and accurate.

[0046] The retrieval module employs a multi-path recall process. The first path is semantic vector retrieval: it constructs a semantic query by combining "user scenario + device information + instrument type + reading + alarm level / rule hit" to recall Top K candidate fragments. The second path is keyword-based recall: it performs an inverted search on device, model, and fault keywords to fill in business-critical fragments that might be overlooked by the vector. The results of the multi-path recall are merged and deduplicated to form a candidate evidence pool, providing higher-quality input for subsequent fine-tuning.

[0047] 3. The reordering and relevance assessment module filters low-relevance fragments from candidate evidence based on set thresholds. A consistency check is then performed: it checks whether the evidence matches the current context. After passing the consistency check, evidence deduplication and coverage control are performed: this avoids multiple fragments repeatedly expressing the same statement; it also ensures that the final evidence covers at least two to three categories of information from the explanation of the phenomenon / threshold meaning, possible causes, and handling suggestions or precautions, ensuring that the generated conclusions are both interpretable and actionable.

[0048] 4. The core constraint of the interpretable generation module is: it should not rewrite the alarm conclusions of AlarmAgent (whether the alarm, level, and hit rule ID remain consistent), but should only explain "why it was triggered", "what are the common causes", "how to handle it", and "what information still needs to be confirmed". During generation, the conclusions are broken down into multiple paragraphs with a fixed structure: the first paragraph is an alarm judgment summary; the second paragraph is the cause explanation; the third paragraph is the handling suggestion; and the fourth paragraph provides a list of evidence citations, allowing operations personnel to trace the source.

[0049] V. User Feedback and Incremental Learning Based on basic interpretability, this alarm system builds a feedback loop capability by adding a user feedback module to collect information such as identification errors and alarm result verification errors. The user feedback results are recorded, and user preferences and identification values ​​are aggregated. The high-quality feedback after statistics is precipitated as reliable labeled data to continuously generate new training samples. Then, the large model is iteratively optimized through online incremental learning to form a closed-loop mechanism of "feedback-training-update", enabling the system to continuously evolve and improve itself in real business scenarios.

[0050] This invention enhances the practicality and reliability of industrial intelligent inspection. The multimodal large model and multi-agent collaborative mechanism significantly improve the accuracy and stability of instrument identification. The combination of multi-level threshold rules and knowledge base retrieval ensures the accuracy of alarm judgment and clearly presents the alarm cause, impact range, and evidence through interpretable conclusions, reducing the understanding and decision-making costs for maintenance personnel. The incremental learning mechanism enables the system's self-evolution, continuously optimizing performance based on user feedback to adapt to complex and changing industrial conditions. It significantly improves inspection efficiency, while interpretable output addresses the pain point of traditional intelligent inspection systems that "only issue alarms without explanations," providing strong support for precise operation and maintenance of industrial equipment.

[0051] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules, units, or units is merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units, modules, or components may be combined or integrated into another device, or some features may be ignored or not executed.

[0052] The units may or may not be physically separate. The components shown as units can be one or more physical units, meaning they can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0053] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0054] In particular, according to embodiments disclosed in this invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), it performs the functions defined in the methods of this invention. It should be noted that the computer-readable medium described above in this invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof.

[0055] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0056] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions within the technical scope disclosed in the present invention should be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An interpretable intelligent inspection and alarm method based on a multimodal large model and multiple agents, characterized in that, Includes the following steps: Collect multimodal data; The multimodal data is preprocessed to obtain preprocessed multimodal data; The preprocessed multimodal data is enhanced and augmented to obtain enhanced and augmented multimodal data; The enhanced and expanded multimodal data are identified to obtain instrument readings; The enhanced and expanded multimodal data and instrument readings are input into the large multimodal model for fine-tuning training to obtain the fine-tuned model; Based on the fine-tuning model, construct intelligent agents for instrument type recognition and reading recognition; The multimodal data to be identified is input into the instrument type identification agent and the reading identification agent to obtain the multimodal identification result; the multimodal identification result includes the type label and the final reading; The multimodal recognition results are input into the alarm determination agent to determine the alarm information; the alarm information includes whether an alarm has been triggered and the alarm level. The alarm information is interpreted by a knowledge base retrieval and an interpretable generative agent, and then sent to the relevant users.

2. The interpretable intelligent inspection and alarm method based on a multimodal large model and multiple agents according to claim 1, characterized in that, Preprocessing the multimodal data to obtain preprocessed multimodal data includes the following steps: The multimodal data is standardized, cut into uniform size specifications, and then formatted to obtain preprocessed multimodal data.

3. The interpretable intelligent inspection and alarm method based on a multimodal large model and multiple agents according to claim 2, characterized in that, The preprocessed multimodal data is augmented and expanded to obtain augmented and expanded multimodal data. The specific steps include: The preprocessed multimodal data is parametrically modeled to simulate the occlusion effects of overexposure, shadows, and highlight glare on the scale and pointer by considering ambient light intensity, color temperature, main light direction, and specular reflection from the instrument glass cover. Based on the mapping relationship between the instrument range scale and pointer angle in the preprocessed multimodal data, a pointer posture consistent with the target reading is generated, and the reading value, pointer angle, key points, and visibility mask are output simultaneously to obtain the instrument reading. Through the mechanism of pointer self-occlusion, external occlusion, and local scale loss caused by perspective projection, a difficult sample set covering extreme working conditions is generated to obtain enhanced and expanded multimodal data.

4. The interpretable intelligent inspection and alarm method based on a multimodal large model and multiple agents according to claim 3, characterized in that, The fine-tuning training of the multimodal large model specifically includes the following steps: Instrument type classification: The model takes enhanced and expanded multimodal data as the main input and integrates text information to output predefined instrument type categories. It is supervised by cross-entropy loss to enable the model to stably distinguish instruments with different structures and scale shapes, and output structured JSON output. Numerical Recognition Regression: Based on the enhanced and expanded multimodal data, the ROI of the instrument panel is selected as the input, and the continuous reading value corresponding to the pointer is output. A regression head is added to the top layer of the model to predict the continuous value of the fused visual features, and SmoothL1 / Huber regression loss is used to improve the robustness to noisy labels and difficult samples. The regression target is designed as the normalized reading and combined with the range information to map back to the actual value. For dual-needle and dual-ring instruments, it is expanded to a multi-output regression head to output the final reading of the instrument.

5. The interpretable intelligent inspection and alarm method based on a multimodal large model and multiple agents according to claim 4, characterized in that, The multimodal data to be identified is input into the instrument type identification agent and the reading identification agent to obtain the multimodal identification results; Specifically, the following steps are included: The multimodal data to be identified and the corresponding user text input type recognition agent are used to determine the type label; The multimodal data to be identified and the type label are input into the reading recognition agent for identification, and the final reading is obtained.

6. The interpretable intelligent inspection and alarm method based on a multimodal large model and multiple agents according to claim 5, characterized in that, The multimodal recognition results are input into the alarm determination agent to determine the alarm information, specifically including the following steps: Input the type label, final reading and corresponding user text into the alarm determination agent, retrieve applicable rules from the MySQL database and perform multi-level threshold calculations to determine whether an alarm is triggered and the alarm level. The MySQL database includes multi-level rules. For a matched rule, the alarm decision agent compares the readings with the threshold range defined by the rule and the pre-defined prompt rule. Interval thresholds: Triggering an alarm when the threshold exceeds [min, max], with different levels corresponding to segmented intervals; otherwise, no alarm is triggered. Upper limit thresholds: reading_value >= the maximum value of the normal threshold, triggering a low-level warning; reading_value >= 5% of the maximum value of the normal threshold, triggering a general warning; reading_value >= 10% of the maximum value of the normal threshold, triggering a severe warning. Lower threshold: reading_value <= minimum normal threshold, triggers low-level warning; reading_value <= 5% of minimum normal threshold, triggers general warning; reading_value <= 10% of minimum normal threshold, triggers severe warning.

7. The interpretable intelligent inspection and alarm method based on a multimodal large model and multiple agents according to claim 6, characterized in that, The alarm information is interpreted through knowledge base retrieval and interpretable generative agents, and then sent to the relevant users. This process includes the following steps: The knowledge base retrieval and interpretable generative agent orchestration layer packages the user-input business scenario text, multimodal recognition results, and alarm information into unified input information for the knowledge base retrieval and interpretable generative agent; Structure and vectorize the knowledge base; Input information is fed into the retrieval module for retrieval: The business scenario text input by the user, the multimodal recognition results, and the alarm information are combined into a semantic query vector to recall TopK candidate fragments; the device, model, and fault keywords of the business scenario text input by the user are used for inverted index retrieval; the results of multi-path recall are obtained, and the results of multi-path recall are merged and deduplicated to form candidate evidence; The candidate evidence is filtered out by the reordering and relevance judgment module to obtain the filtered evidence by filtering low-relevance segments from the set threshold. Based on the interpretable generation phase generation module, alarm explanations are generated from filtered evidence. Send alarm information and alarm explanations to the user.

8. An interpretable intelligent inspection alarm system based on a multimodal large model and multiple agents, used to implement the interpretable intelligent inspection alarm method based on a multimodal large model and multiple agents as described in any one of claims 1-7, characterized in that, include: The acquisition module is used to acquire multimodal data; The preprocessing module is used to preprocess the multimodal data to obtain preprocessed multimodal data; The enhancement and expansion module is used to enhance and expand the preprocessed multimodal data to obtain enhanced and expanded multimodal data. The identification module is used to identify the enhanced and expanded multimodal data to obtain instrument readings; The fine-tuning module is used to input the enhanced and expanded multimodal data and instrument readings into the large multimodal model for fine-tuning training, thereby obtaining the fine-tuned model; The building module is used to construct instrument type recognition agents and reading recognition agents based on the fine-tuning model; The multimodal recognition module is used to input the multimodal data to be recognized into the instrument type recognition agent and the reading recognition agent to obtain the multimodal recognition result; the multimodal recognition result includes the type label and the final reading; The alarm module is used to input the multimodal recognition results into the alarm judgment agent to determine the alarm information; the alarm information includes whether an alarm has been triggered and the alarm level. The execution module is used to interpret alarm information by retrieving it from the knowledge base and generating an interpretable intelligent agent, and then send the alarm information to the corresponding user.