Supervision process-oriented multi-modal engineering document intelligent structured extraction and association system

The intelligent structured extraction and association system for multimodal engineering documents solves the problems of unprofessional parsing and insufficient association of supervision documents, realizes efficient and accurate management and risk prediction of the engineering supervision process, and improves the level of intelligence in supervision document processing.

CN121935358APending Publication Date: 2026-04-28BEIJING HAICE ENG CONSULTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING HAICE ENG CONSULTING CO LTD
Filing Date
2025-12-22
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In traditional supervision work, the lack of a unified parsing and association mechanism for multimodal documents limits the value mining of supervision data. The closed-loop management of engineering problems from discovery to review relies on manual tracking, which is prone to management loopholes due to information delays and unclear relationships.

Method used

Design a multimodal engineering document intelligent structured extraction and association system, including a multimodal perception and understanding module, a dynamic event graph construction module, a closed-loop state tracking module, and an interactive application module. Through technologies such as domain large model, chain reasoning, assisted measurement, dynamic graph update, and intelligent early warning, it realizes accurate parsing, dynamic association, and process tracking of multimodal data.

Benefits of technology

It significantly improves the accuracy of problem identification, the completeness of correlation tracing, and the timeliness of risk prediction, promoting the transformation of supervision document processing from fragmented to intelligent full-process control, and providing more efficient and reliable technical support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121935358A_ABST
    Figure CN121935358A_ABST
Patent Text Reader

Abstract

The invention discloses a supervision-process-oriented multi-modal engineering document intelligent structured extraction and association system, and relates to the technical field of engineering supervision informatization. A multi-modal perception understanding module accurately analyzes texts, images and voice documents through a supervision field fine-tuning large model, multi-step chain reasoning and geometric parameter measurement; the dynamic affair graph construction module is used for defining relationships between nodes such as problems and rectification and triggering and response, establishing association according to an analysis result and dynamically updating attributes and weights; the closed-loop state tracking module determines and rectifies a flow state rule, and realizes overtime and associated risk early warning and flow evolution prediction; the interaction application module provides a visual billboard, voice interaction and automatic report generation functions; the model optimization module acquires new case data, calculates process deviation and adjusts model parameters, and the system improves supervision document processing efficiency and rectification management and control precision and supports intelligent decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information technology in engineering supervision, and more specifically, to a multimodal intelligent structured extraction and association system for engineering documents oriented towards the supervision process. Background Technology

[0002] With the deepening of digital transformation in the engineering construction field, supervision, as the core control link of engineering quality and safety, faces the dual challenges of a surge in multimodal data and the need for refined process control in its document management. In traditional supervision work, multimodal documents such as text-based inspection batch reports, image-based hidden works photos, and audio-based on-site briefing records are fragmented and lack a unified parsing and association mechanism, which limits the value mining of supervision data. At the same time, the closed-loop control of engineering problems from discovery, rectification to review relies on manual tracking, which is prone to control loopholes due to information delays and unclear relationships. This problem is particularly prominent in large and complex projects.

[0003] Breakthroughs in multimodal intelligent processing and knowledge graph technology provide technical support for upgrading the management and control of supervision documents. Although multimodal large models can achieve cross-modal analysis in general fields, the professional terminology system and engineering problem representation rules in the field of supervision are highly domain-specific, and direct application is prone to problem identification bias. The application of process visualization technology has shown initial results, but the rectification process of supervision problems is dynamic and uncertain. Static graphs are difficult to adapt to the relationship changes and status update needs in the process evolution. In addition, the detailed needs such as geometric parameter calculation in engineering documents and accurate early warning of rectification process require deep integration of professional modules and general intelligent technologies to meet them.

[0004] The core demands of the supervision industry for document processing have shifted from simple storage management to a fully intelligent process encompassing parsing, association, tracking, and optimization. Existing technologies struggle to simultaneously address the professionalism of multimodal document parsing, the dynamism of process association, and the accuracy of control decisions: text parsing easily overlooks engineering semantics, image measurement accuracy is affected by shooting angle, and speech transcription lacks domain terminology correction; the association of data across modalities is limited to label matching, failing to construct causal links between problems and rectification; and rectification process tracking relies on manual status entry, lacking data-driven automatic deduction and early warning mechanisms. Therefore, building a multimodal intelligent parsing and association system adapted to the supervision field has become crucial for overcoming the bottlenecks in supervision document management and improving engineering management efficiency. Summary of the Invention

[0005] To overcome the problems of unprofessional multimodal parsing, label matching as the only association method, and lack of automatic inference and early warning in process tracking, this invention discloses an intelligent structured extraction and association system for multimodal engineering documents oriented towards the supervision process, which can effectively solve the above-mentioned technical problems.

[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: A multimodal engineering document intelligent structured extraction and association system for the supervision process includes: a multimodal perception and understanding module, a dynamic event graph construction module, a closed-loop state tracking module, an interactive application module, and a model optimization module; The multimodal perception and understanding module is used to intelligently parse multimodal engineering documents, which include text documents, image documents, and voice documents. The multimodal perception and understanding module includes a domain large model unit, a chain reasoning unit, and an auxiliary measurement unit. The domain large model unit is generated based on the multimodal large model and fine-tuned with supervision domain data. The chain reasoning unit is configured with multi-step reasoning logic to realize problem identification and instruction generation. The auxiliary measurement unit is used to accurately measure engineering geometric parameters. The dynamic event graph construction module includes a graph pattern definition unit, a node relationship creation unit, and a graph update unit. The graph pattern definition unit presets several node types and relationship types. The node types include problem event nodes, rectification instruction nodes, rectification feedback nodes, and review nodes. The relationship types include trigger relationships, response relationships, and verification relationships. The node relationship creation unit is used to create nodes and relationships between nodes based on the parsing results of the multimodal perception and understanding module. The graph update unit is used to dynamically update node attributes and relationship weights based on the process progress status. The closed-loop state tracking module includes a state machine configuration unit, an intelligent early warning unit, and a process deduction unit. The state machine configuration unit defines state transition rules for each problem rectification process. The state transition rules include the transition conditions for states to be assigned, in rectification, pending review, closed, and timed out. The intelligent early warning unit is used to generate timeout warnings, associated risk warnings, and review failure warnings based on preset thresholds. The process deduction unit predicts the process evolution path based on a dynamic event graph. The interactive application module includes a visual dashboard unit, a voice interaction unit, and a report generation unit. The visual dashboard unit displays the entire process of problem rectification in the form of a timeline and flowchart. The voice interaction unit supports natural language query and command input. The report generation unit automatically generates a supervision report based on dynamic event graph data. The model optimization module includes a sample acquisition unit, a deviation correction unit, and a parameter update unit. The sample acquisition unit is used to collect new supervision case data. The deviation correction unit is used to calculate the deviation value between the actual process and the predicted process. The parameter update unit adjusts the correlation parameters of the domain large model and the dynamic reasoning graph according to the deviation value.

[0007] Preferably, the domain large model unit is configured with a modal fusion algorithm, which is used to perform weighted fusion of text features, image features and speech features to generate a unified multimodal feature vector. The multimodal feature vector is used for attribute extraction of problem events, and the attributes include problem type, engineering location, severity and responsible party.

[0008] Preferably, the chain-based reasoning unit includes a description subunit, an identification subunit, a quantization subunit, and an instruction subunit. The description subunit is used to extract key elements from multimodal engineering documents. The identification subunit is used to determine whether the key element constitutes a quality or safety issue. The quantization subunit is used to numerically represent the characteristic parameters of the problem. The instruction subunit is used to generate rectification instruction text that conforms to the specifications.

[0009] Preferably, the auxiliary measurement unit includes an image preprocessing subunit, a feature point detection subunit, and a size calculation subunit. The image preprocessing subunit is used to perform noise reduction and distortion correction on the engineering images. The feature point detection subunit is used to identify reference points and feature points to be detected in the image. The size calculation subunit calculates the actual physical size based on the principle of perspective transformation.

[0010] Preferably, the node relationship creation unit is configured with a confidence evaluation algorithm, which is used to calculate the reliability value of the association relationship between nodes. The reliability value is positively correlated with the consistency, information integrity and time correlation of multimodal data. When the reliability value is greater than a preset confidence threshold, the creation of the corresponding association relationship is confirmed.

[0011] Preferably, the map update unit includes an attribute update subunit and a relation evolution subunit. The attribute update subunit is used to update the node's status label, timestamp, and processing result in real time. The relationship evolution subunit is used to adjust the strength value of the association relationship according to the progress of the process. The intensity value is related to the causal relationship between nodes and the time interval.

[0012] Preferably, the intelligent early warning unit includes a timeout calculation subunit, a correlation analysis subunit, and a pattern recognition subunit. The timeout calculation subunit is used to calculate the difference between the duration of the current state and the preset time limit. The correlation analysis subunit is used to mine frequently occurring node combinations in the dynamic event graph. The pattern recognition subunit is used to identify problem feature patterns that fail multiple times during re-examination.

[0013] Preferably, the process deduction unit is configured with a path prediction model. The path prediction model takes the current node status and historical process data as input and outputs possible future state transition paths and probability values ​​of each path. The probability values ​​are calculated and generated based on the relationship weights of the dynamic process graph and the historical transition frequency.

[0014] Preferably, the report generation unit includes a template configuration subunit, a data filling subunit, and a format optimization subunit. The template configuration subunit has multiple preset supervision report templates. The data filling subunit extracts the corresponding field data from the dynamic reasoning graph. The format optimization subunit is used to adjust the report's layout and chart display format.

[0015] Preferably, the deviation correction unit includes a process comparison subunit and a correction coefficient generation subunit. The process comparison subunit is used to compare the differences between the actual process node sequence and the predicted process node sequence. The correction coefficient generation subunit generates corresponding model correction coefficients based on the degree of difference. These correction coefficients are used to adjust the inference weights of the chain inference unit and the relationship parameters of the dynamic reasoning graph.

[0016] Compared with existing technologies, the beneficial effects of this invention are as follows: This system precisely addresses the pain points of existing technologies through modular collaborative design: For the problem of insufficient professionalism in multimodal analysis, the domain-specific large model unit is finely tuned based on data from the supervision domain, and combined with a modal fusion algorithm, weighted and fused text, image, and voice features into a unified vector. This, combined with the four-step logic of the chain-based reasoning unit—extracting key elements, identifying quality and safety issues, quantifying feature parameters, and generating standard instructions—can accurately analyze professional terms such as insufficient rebar cover thickness. The auxiliary measurement unit achieves high-precision calculation of engineering geometric parameters through image noise reduction and correction, feature point detection, and perspective transformation calculation, solving the problem of insufficient professional adaptability of general models and making multimodal information analysis more aligned with the needs of supervision scenarios. Addressing the limitation of label-only association, the dynamic event graph construction module presets node types such as problem events and rectification instructions, as well as relationship types such as triggers and responses. The node relationship creation unit is based on multimodal data consistency, information integrity, and temporal correlation. The system calculates reliable values, retaining only high-confidence associations. The graph update unit dynamically adjusts node attributes and relationship strength as the process progresses, upgrading associations from static labels to dynamic causal chains. This avoids information fragmentation and allows data from each stage to form an organic whole. Addressing the lack of automatic deduction and early warning in process tracking, the closed-loop status tracking module defines standardized state transition rules through a state machine configuration unit. The intelligent early warning unit calculates timeout differences in real time, identifies high-frequency risk node combinations, and recognizes failure patterns during review. The process deduction unit predicts path trends based on dynamic event graphs and historical data, and continuously improves accuracy through deviation correction by the model optimization module. This transforms the system from passive recording to proactive early warning, enabling supervisors to anticipate risks and intervene promptly. Ultimately, through multi-module collaboration, the system significantly improves the accuracy of problem identification, the completeness of association tracing, and the timeliness of risk prediction, propelling the processing of supervision documents from fragmented to intelligent full-process control, and providing more efficient and reliable technical support for engineering supervision. Attached Figure Description

[0017] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other embodiments can be derived from the provided drawings without creative effort.

[0018] Figure 1 This is a system structure diagram of the present invention. Detailed Implementation

[0019] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this patent. To better illustrate this embodiment, some parts in the accompanying drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions; It will be understood by those skilled in the art that certain well-known structures and their descriptions may be omitted in the accompanying drawings.

[0020] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0021] Example 1

[0022] This system aims to achieve efficient management and precise control of the engineering supervision process through functional modules such as intelligent analysis of multimodal data, construction of dynamic event graphs, closed-loop state tracking, interactive applications, and model optimization. The system takes multimodal engineering documents as input, covering various forms such as text, images, and voice. Through the collaborative operation of various modules, it completes intelligent processing of the entire process from data analysis to problem rectification, and finally outputs a supervision report and continuously optimizes its own performance to adapt to the ever-changing engineering supervision scenarios.

[0023] Please see Figure 1 A multimodal engineering document intelligent structured extraction and association system for the supervision process includes: a multimodal perception and understanding module, a dynamic event graph construction module, a closed-loop state tracking module, an interactive application module, and a model optimization module; The multimodal perception and understanding module is used to intelligently parse multimodal engineering documents, which include text documents, image documents, and voice documents. The multimodal perception and understanding module includes a domain large model unit, a chain reasoning unit, and an auxiliary measurement unit. The domain large model unit is generated based on the multimodal large model and fine-tuned with supervision domain data. The chain reasoning unit is configured with multi-step reasoning logic to realize problem identification and instruction generation. The auxiliary measurement unit is used to accurately measure engineering geometric parameters. The dynamic event graph construction module includes a graph pattern definition unit, a node relationship creation unit, and a graph update unit. The graph pattern definition unit presets several node types and relationship types. The node types include problem event nodes, rectification instruction nodes, rectification feedback nodes, and review nodes. The relationship types include trigger relationships, response relationships, and verification relationships. The node relationship creation unit is used to create nodes and relationships between nodes based on the parsing results of the multimodal perception and understanding module. The graph update unit is used to dynamically update node attributes and relationship weights based on the process progress status. The closed-loop state tracking module includes a state machine configuration unit, an intelligent early warning unit, and a process deduction unit. The state machine configuration unit defines state transition rules for each problem rectification process. The state transition rules include the transition conditions for states to be assigned, in rectification, pending review, closed, and timed out. The intelligent early warning unit is used to generate timeout warnings, associated risk warnings, and review failure warnings based on preset thresholds. The process deduction unit predicts the process evolution path based on a dynamic event graph. The interactive application module includes a visual dashboard unit, a voice interaction unit, and a report generation unit. The visual dashboard unit displays the entire process of problem rectification in the form of a timeline and flowchart. The voice interaction unit supports natural language query and command input. The report generation unit automatically generates a supervision report based on dynamic event graph data. The model optimization module includes a sample acquisition unit, a deviation correction unit, and a parameter update unit. The sample acquisition unit is used to collect new supervision case data. The deviation correction unit is used to calculate the deviation value between the actual process and the predicted process. The parameter update unit adjusts the correlation parameters of the domain large model and the dynamic reasoning graph according to the deviation value.

[0024] The domain large model unit is configured with a modal fusion algorithm, which is used to perform weighted fusion of text features, image features and speech features to generate a unified multimodal feature vector. The multimodal feature vector is used for attribute extraction of problem events, and the attributes include problem type, engineering location, severity and responsible party.

[0025] The chain-based reasoning unit includes a description subunit, an identification subunit, a quantization subunit, and an instruction subunit. The description subunit is used to extract key elements from multimodal engineering documents. The identification subunit is used to determine whether the key element constitutes a quality or safety issue. The quantization subunit is used to numerically represent the characteristic parameters of the problem. The instruction subunit is used to generate rectification instruction text that conforms to the specifications.

[0026] The auxiliary measurement unit includes an image preprocessing subunit, a feature point detection subunit, and a size calculation subunit. The image preprocessing subunit is used to perform noise reduction and distortion correction on the engineering images. The feature point detection subunit is used to identify reference points and feature points to be detected in the image. The size calculation subunit calculates the actual physical size based on the principle of perspective transformation.

[0027] The node relationship creation unit is configured with a confidence evaluation algorithm, which is used to calculate the reliability value of the association relationship between nodes. The reliability value is positively correlated with the consistency, information integrity and time correlation of multimodal data. When the reliability value is greater than the preset confidence threshold, the creation of the corresponding association relationship is confirmed.

[0028] The graph update unit includes an attribute update subunit and a relation evolution subunit. The attribute update subunit is used to update the node's status label, timestamp, and processing result in real time. The relationship evolution subunit is used to adjust the strength value of the association relationship according to the progress of the process. The intensity value is related to the causal relationship between nodes and the time interval.

[0029] The intelligent early warning unit includes a timeout calculation subunit, a correlation analysis subunit, and a pattern recognition subunit. The timeout calculation subunit is used to calculate the difference between the duration of the current state and the preset time limit. The correlation analysis subunit is used to mine frequently occurring node combinations in the dynamic event graph. The pattern recognition subunit is used to identify problem feature patterns that fail multiple times during re-examination.

[0030] The process deduction unit is equipped with a path prediction model. The path prediction model takes the current node status and historical process data as input and outputs possible future state transition paths and probability values ​​of each path. The probability values ​​are calculated and generated based on the relationship weights of the dynamic process graph and the historical transition frequency.

[0031] The report generation unit includes a template configuration subunit, a data filling subunit, and a format optimization subunit. The template configuration subunit has multiple preset supervision report templates. The data filling subunit extracts the corresponding field data from the dynamic reasoning graph. The format optimization subunit is used to adjust the report's layout and chart display format.

[0032] The deviation correction unit includes a process comparison subunit and a correction coefficient generation subunit. The process comparison subunit is used to compare the differences between the actual process node sequence and the predicted process node sequence. The correction coefficient generation subunit generates corresponding model correction coefficients based on the degree of difference. These correction coefficients are used to adjust the inference weights of the chain inference unit and the relationship parameters of the dynamic reasoning graph.

[0033] The domain-specific large model unit collects a large amount of text, image, and voice data from the supervision field, including but not limited to project acceptance reports, construction site photos, and supervision meeting recordings. It performs preprocessing operations such as word segmentation and part-of-speech tagging on text data to remove stop words; it performs operations such as cropping and scaling on image data to make it meet the input size requirements; and it performs noise reduction and channel separation on voice data to improve voice quality.

[0034] The preprocessed data is divided into training set, validation set and test set, with the training set accounting for 70%, the validation set accounting for 15% and the test set accounting for 15%.

[0035] Choose a pre-trained multimodal large model, such as the CLIP (Contrastive Language-Image Pre-training) model, which has the ability to jointly understand text and images and can be used as the basic architecture.

[0036] Fine-tuning of a multimodal large model on data from the supervision field involves the following steps: inputting training data into the model, adjusting the model's weight parameters to better adapt the model to the specific tasks of the supervision field, using the cross-entropy loss function to measure the difference between the model's predictions and the true labels during fine-tuning, updating the model parameters through the backpropagation algorithm, evaluating the fine-tuned model on the validation set, and adjusting the fine-tuning strategy based on the evaluation results, such as adjusting the learning rate and adding regularization terms, to prevent the model from overfitting.

[0037] After multiple iterations and fine-tuning, a domain-specific large model unit was obtained. This unit is able to perform a preliminary understanding of the input multimodal engineering documents, extract key information from the text, images, and speech, and map them to a unified feature space, providing data for chained reasoning and assisted measurement.

[0038] A modal fusion algorithm is configured, which uses a weighted fusion method to fuse text features, image features, and speech features. First, the text, image, and speech features are normalized so that their values ​​are within the range of [0, 1]. Then, according to the actual needs of the supervision field and the importance of each modality, different weights are assigned to the text features, image features, and speech features. For example, in the scenario of detecting engineering quality problems, the weight of image features may be higher because images can intuitively show engineering defects; while in the scenario of conveying supervision instructions, the weight of speech features may be higher. By using a weighted summation method, the features of the three modalities are fused into a unified multimodal feature vector.

[0039] The attributes of problem events are extracted using the fused multimodal feature vectors. The attributes of problem events are defined as follows: problem type (e.g., quality problem, safety problem, schedule problem), engineering location (e.g., foundation engineering, main structure engineering, decoration and renovation engineering), severity (e.g., minor, moderate, serious), and responsible party (e.g., construction unit, material supplier, supervision unit). By designing corresponding feature extraction network structures, such as convolutional neural networks (CNN) for image feature extraction and recurrent neural networks (RNN) for text and speech sequence feature extraction, these attribute information are extracted from the fused feature vectors to provide a problem description for engineering problem processing.

[0040] The chained reasoning unit describes how the subunit extracts key elements from the input multimodal engineering documents. For text documents, it uses the Named Entity Recognition (NER) algorithm from natural language processing to identify key entity information such as project name, construction unit, supervisor, and construction date. For image documents, it uses object detection algorithms, such as Faster R-CNN, to detect key objects such as engineering components, construction equipment, and safety facilities in the images and extract their location and category information. For speech documents, it uses speech recognition technology to convert speech into text and then uses text analysis methods to extract key elements.

[0041] The extracted key elements are integrated to form a structured data set containing key elements of text, images, and speech, providing input data for recognition and quantification operations.

[0042] The identification sub-unit, based on a predefined knowledge base and rule base for the supervision domain, judges the extracted key elements. The knowledge base stores standard descriptions, characteristic manifestations, and corresponding judgment criteria for various engineering quality and safety issues. For example, for concrete structure engineering, if honeycomb or pitting is detected on the concrete surface in the image, it can be judged as a quality issue based on the rules in the knowledge base; if the voice recording mentions that construction workers are not wearing safety helmets or other violations of safety regulations, it can be judged as a safety issue.

[0043] For complex combinations of key elements, classification algorithms in deep learning, such as support vector machines (SVM) or deep neural networks (DNN), are used to classify and judge them. The feature vectors of the key elements are input into the classifier, and the classifier outputs the judgment result of whether the key elements constitute a quality or safety problem based on the mapping relationship between features and categories learned during the training process.

[0044] For key elements identified as quality or safety issues, the quantification sub-unit further quantifies their characteristic parameters. For example, for the problem of insufficient concrete strength, the concrete strength value is estimated by using texture analysis algorithms in image analysis technology and combining empirical formulas between concrete strength and image texture features. For the problem of construction schedule delays, the construction schedule plan and actual progress data are extracted from text documents to calculate the specific number of days of schedule delay.

[0045] For some problem characteristics that cannot be directly represented by numerical values, fuzzy quantification methods are adopted. For example, for appearance defects in construction quality, such as flatness and smoothness, they are divided into several levels such as excellent, good, medium and poor based on the experience of supervisors and industry standards, and assigned corresponding numerical ranges, such as excellent as [90, 100], good as [70, 90), etc., thereby realizing the quantification of problem characteristic parameters.

[0046] The instruction subunit generates a rectification instruction text that conforms to the specifications based on the quantified characteristic parameters and the standard requirements of the supervision field. First, it selects the appropriate template from the predefined instruction template library according to the problem type and severity. For example, for quality problems, the template may be a problem of [problem location] [problem type], asking the construction unit to rectify according to the [rectification standard] within the [rectification period] and submit a rectification report.

[0047] Then, the quantified feature parameters and specific problem information are filled into the template to generate specific rectification instruction text. For example, for the problem of insufficient concrete strength, the generated instruction text is: "For the problem of insufficient concrete strength in foundation engineering, please have the construction unit re-pour the concrete according to the design requirements within 7 days and submit a concrete strength test report." At the same time, the instruction sub-unit can also refine and classify the instruction according to the complexity of the problem and the responsible party to ensure the pertinence and operability of the rectification instruction.

[0048] The auxiliary measurement unit and the image preprocessing subunit perform noise reduction processing on the input engineering image. They adopt the median filtering algorithm, which takes a neighborhood window centered on each pixel in the image, sorts the pixel values ​​in the neighborhood window by size, and replaces the value of the center pixel with the median value, thereby removing random noise in the image. For example, in construction site photos, noise may be generated due to changes in lighting or equipment failure. Median filtering can effectively reduce the impact of noise on image quality.

[0049] Image distortion correction is performed. If the image has barrel distortion or pincushion distortion, a distortion correction algorithm is used to perform geometric transformation on the image based on the camera's distortion parameter model to restore the image's true shape. The distortion parameters can be obtained in advance through camera calibration experiments, or some deep learning-based distortion correction networks can be used to automatically learn the distortion features of the image and perform correction.

[0050] The feature point detection subunit uses feature point detection algorithms, such as SIFT (Scale-Invariant Feature Transform) or ORB (Oriented Fast and Rotated BRIEF), to identify reference points and feature points to be measured in the image. Reference points can be coordinate points in engineering drawings, marker points on the construction site, etc., used to determine the reference coordinate system of the image. Feature points to be measured are feature points of engineering components that need to be dimensionally measured, such as the endpoints and inflection points of the components.

[0051] For complex engineering images, such as those containing a large number of repetitive textures and similar shapes, a multi-scale feature point detection method is adopted. First, the image is decomposed into pyramids at different scales. Then, feature point detection is performed at each scale layer. Finally, the feature points at different scale layers are fused to obtain a more comprehensive and accurate set of feature points.

[0052] The size calculation subunit calculates the actual physical size based on the principle of perspective transformation. First, based on the reference point and the feature point to be measured detected by the feature point detection subunit, the perspective transformation relationship in the image is determined. By calculating the ratio between the pixel size of the reference object of known size in the image and the actual physical size, the scale factor of the image is obtained.

[0053] Then, using the scale factor and the pixel coordinates of the feature points to be measured in the image, the actual physical distance between the feature points to be measured is calculated. For example, when measuring the length of a building wall, the actual length of the wall can be obtained by finding the two endpoint feature points of the wall in the image, calculating the pixel distance between them, and then multiplying it by the scale factor. At the same time, considering the influence of factors such as the image shooting angle and the shape of the object on the size measurement, an error compensation algorithm is used to correct the measurement results and improve the accuracy of the size measurement.

[0054] The graph pattern definition unit pre-defines several node types, including problem event nodes, rectification instruction nodes, rectification feedback nodes, and review nodes. Problem event nodes are used to represent various problems found during the engineering supervision process, such as quality problems and safety problems; rectification instruction nodes are used to represent rectification instructions issued in response to problem events; rectification feedback nodes are used to represent feedback from the construction unit on the implementation of rectification instructions; and review nodes are used to represent the review status of rectification results by the supervisors.

[0055] Several relationship types are preset, including trigger relationships, response relationships, and verification relationships. Trigger relationships indicate that a problem event node triggers the generation of a rectification instruction node; response relationships indicate that a rectification instruction node triggers the generation of a rectification feedback node; and verification relationships indicate that a rectification feedback node triggers the execution of a review node. These node types and relationship types together constitute the basic structural framework of the dynamic event graph, which is used to describe the problem rectification process in the engineering supervision process.

[0056] Define corresponding attributes for each node type. The attributes of the problem event node include problem type, project location, severity, responsible party, and discovery time; the attributes of the rectification instruction node include instruction content, rectification period, responsible person, and issuance time; the attributes of the rectification feedback node include feedback content, rectification completion status, and feedback time; and the attributes of the review node include review result and review time.

[0057] For each type of relationship, corresponding attributes are defined. The attributes of triggering relationships include trigger time and trigger conditions; the attributes of response relationships include response time and response method; and the attributes of verification relationships include verification time and verification criteria. These attribute information can provide richer semantic information for graph updates and process deduction.

[0058] The node relationship creation unit receives the parsing results from the multimodal perception and understanding module, including the attribute information of the problem event, the rectification instruction text, and the rectification feedback content. It then classifies these parsing results according to the predefined node types and extracts the corresponding node attribute values. For example, it extracts attribute values ​​such as the problem type being insufficient concrete strength, the engineering location being the foundation engineering, the severity being severe, the responsible party being the construction unit, and the discovery time being 2023-04-15 from the attribute information of the problem event, and generates a problem event node.

[0059] Based on the logical relationships between problem events and rectification instructions, rectification instructions and rectification feedback, and rectification feedback and review in the analysis results, corresponding node relationships are created. For example, when there is a problem event and a rectification instruction for the problem event in the analysis results, a trigger relationship is created to connect the problem event node and the rectification instruction node, and the trigger time is set to the time when the rectification instruction is issued.

[0060] Configure a confidence assessment algorithm to calculate the reliability value of the relationship between nodes. The calculation of the reliability value considers three factors: consistency of multimodal data, information integrity, and temporal correlation. Consistency refers to whether the descriptions of the same event or relationship in different modal data are consistent, such as whether the problem event described in text is consistent with the problem phenomenon shown in the image. Information integrity refers to whether the node attributes and relationship attributes are filled in completely and whether they contain key information. Temporal correlation refers to whether the time sequence of events and relationships is reasonable and conforms to the logic of the engineering supervision process.

[0061] The specific calculation formula is: Reliability value = α × consistency score + β × information integrity score + γ × time correlation score, where α, β, and γ are weighting coefficients that are adjusted according to actual needs to satisfy α + β + γ = 1. When the reliability value is greater than the preset confidence threshold (e.g., 0.8), the corresponding correlation relationship is confirmed to be created. For example, for a problem event node and a rectification instruction node, if their text descriptions are consistent, the information is complete, and the time sequence is reasonable, the calculated reliability value is 0.85, which is greater than the confidence threshold, and the triggering relationship is confirmed to be created.

[0062] The graph update unit and the attribute update subunit update the node's status label, timestamp, and processing result in real time. For example, when the status of the rectification instruction node changes from pending allocation to rectification in progress, its status label is updated to rectification in progress, and the current time is recorded as the timestamp. When rectification feedback is received from the construction unit, the processing result attribute of the rectification feedback node is updated, and the specific situation of rectification completion is recorded, such as rectification completed, problem solved, or rectification not completed and further processing required.

[0063] For newly generated nodes, their attribute values ​​are initialized based on their initial state and related information at the time of creation. For example, for a newly generated review node, its initial state label is "pending review", the timestamp is the start time of the review plan, and the processing result attribute is empty. It will be updated after the review is completed.

[0064] The relationship evolution subunit adjusts the strength value of the relationship according to the progress of the process. The strength value is related to the causal relationship between nodes and the time interval. For example, if there is a direct causal relationship between a problem event and a rectification instruction, and the rectification instruction is issued in a timely manner with a short time interval, the strength value is high; if the time interval is long, it may affect the timeliness and effectiveness of rectification, and the strength value will decrease accordingly.

[0065] The specific adjustment method is as follows: when the node status changes or the time interval changes, the strength value of the correlation is recalculated. The formula for calculating the strength value is: Strength value = f(causal correlation coefficient, time interval coefficient). The causal correlation coefficient is preset according to the strength of the causal relationship between nodes, and the time interval coefficient is dynamically adjusted according to the length of the time interval. For example, if the causal correlation coefficient is 0.9 and the time interval coefficient is 1 - (time interval / preset maximum time interval), when the time interval is half of the preset maximum time interval, the time interval coefficient is 0.5, and the strength value is 0.9 × 0.5 = 0.45. By dynamically adjusting the strength value of the correlation in this way, the graph can more accurately reflect the actual progress of the engineering supervision process.

[0066] The state machine configuration unit defines state transition rules for each problem rectification process, including the transition conditions for states such as pending assignment, in progress, pending review, closed, and timed out. The pending assignment state indicates that the problem event has been discovered, but a person responsible for rectification has not yet been assigned; the in progress state indicates that the person responsible for rectification has been identified and is carrying out the rectification task; the pending review state indicates that the rectification task has been completed and is waiting for the supervisor to review it; the closed state indicates that the review result meets the requirements and the problem rectification process ends; and the timed out state indicates that the rectification task was not completed within the specified time.

[0067] The specific conversion conditions are as follows: The condition for changing from the pending assignment status to the rectification status is that the person responsible for rectification is assigned to the problem event, and the rectification period has been set. For example, if the supervisor assigns the project manager of the construction unit as the person responsible for rectification for the problem event in the system and sets a 7-day rectification period, the status of the problem rectification process will change from pending assignment to rectification.

[0068] The conditions for changing the status from "Under Rectification" to "Pending Review" are: the rectification period has expired, or the construction unit has submitted rectification feedback. If the rectification period has expired, the system will automatically change the status to "Pending Review". If the construction unit completes the rectification ahead of schedule and submits feedback, the supervisor can also manually change the status to "Pending Review" after reviewing the feedback.

[0069] The conditions for changing the status from pending review to closed are: after the supervisor reviews and confirms that the problem has been resolved, if the supervisor finds that the problem has indeed been rectified as required, the supervisor clicks the closed button and the status changes to closed.

[0070] The conditions for changing the status from pending review to timed out are: the review period has expired, but the problem is still not resolved. If the supervisor fails to complete the review within the specified time, or if the review results show that the problem is still not resolved and the rectification period has expired, the system will automatically change the status to timed out.

[0071] The conditions for changing the status from "overdue" to "under rectification" are: the supervisor reassigns the rectification task or extends the rectification period. If the supervisor believes that the problem can still be rectified, the status will be changed back to "under rectification" after the supervisor reassigns the person responsible for rectification or extends the rectification period.

[0072] Based on the defined state transition rules, a state machine model is constructed. Taking the problem rectification process as an example, the state machine model monitors the state changes of each node in the process in real time. When a certain state transition condition is met, the state machine automatically triggers the corresponding state transition action and records the time, reason and other relevant information of the state transition.

[0073] The system interface displays the state machine's operating status in a visual way, for example, using different colored icons to represent different states: green indicates pending assignment, blue indicates rectification in progress, yellow indicates pending review, red indicates timeout, and gray indicates closure. Supervisors can intuitively understand the current status and historical status change trajectory of each problem rectification process through the interface, which facilitates overall control of the entire supervision process.

[0074] The intelligent early warning unit and the timeout calculation subunit calculate the difference between the current state duration and the preset time limit. For each problem rectification process, the time difference is calculated in real time based on its current state and the corresponding preset time limit. For example, for a problem rectification process in the rectification stage, the preset rectification period is 7 days. Starting from the rectification start time, the time difference between the current time and the rectification start time is calculated. If the time difference is greater than 7 days, it means that the rectification has exceeded the time limit.

[0075] When the time difference approaches the preset time limit, if the remaining time is less than 24 hours, the timeout calculation subunit sends a timeout warning signal to the intelligent warning unit. The intelligent warning unit generates a timeout warning message based on the warning signal, reminding the supervisors to pay attention to the progress of the rectification process and take timely measures.

[0076] The correlation analysis subunit mines frequently occurring node combinations in the dynamic event graph. By statistically analyzing the correlation relationships of nodes in the graph, it identifies frequently occurring node combination patterns. For example, it finds that a certain problem event node and a certain rectification instruction node frequently have a trigger relationship, and each trigger is accompanied by the generation of a rectification feedback node and a review node. This indicates that this combination pattern of problem events and rectification instructions is relatively common in the process of engineering supervision.

[0077] Based on the high-frequency node combination pattern, potential associated risks are analyzed. If a rectification feedback node frequently shows that rectification is incomplete, or a review node frequently fails review, the association analysis subunit will send an associated risk warning signal to the intelligent early warning unit. The intelligent early warning unit will generate associated risk warning information based on the warning signal, reminding the supervisor to pay attention to the deeper risks that may exist in this problem event, such as insufficient rectification capabilities of the construction unit or the complexity of the problem exceeding expectations.

[0078] The pattern recognition subunit identifies the characteristic patterns of problems that fail multiple reviews. When a problem fails to pass after multiple rectifications and reviews, the pattern recognition subunit analyzes the characteristics of the problem and extracts key features from the attributes of the problem, such as problem type, engineering location, and severity. Combined with the rectification instructions and feedback, the subunit identifies the characteristic patterns that lead to the failure to pass the review.

[0079] For example, if a foundation project fails to pass a re-inspection after multiple rectifications due to insufficient concrete strength, the pattern recognition subunit may identify the following characteristic patterns: problem type: insufficient concrete strength; project location: foundation project; severity: severe; rectification instruction: re-pouring concrete; but rectification feedback shows that the construction unit did not strictly follow the design requirements for pouring. Based on the identified characteristic patterns, the pattern recognition subunit sends a re-inspection failure warning signal to the intelligent early warning unit. The intelligent early warning unit generates warning information, reminding the supervisor to focus on tracking and handling the problem, and to take stricter rectification measures if necessary.

[0080] The process simulation unit is configured with a path prediction model. Taking the current node status and historical process data as input, the path prediction model can adopt a Markov chain model or a recurrent neural network (RNN) model based on deep learning. The Markov chain model assumes that the next state of the problem rectification process is only related to the current state. By calculating the state transition probability matrix, it predicts the possible future state transition paths. The RNN model can take into account the time series characteristics of historical process data and model and predict complex process evolution paths.

[0081] The path prediction model is trained using historical engineering supervision data, including the node state sequence, state transition time, and correlation information of completed problem rectification processes. Through training, the model learns the transition rules and path probability distribution between different states. For example, the model can learn that after a rectification is in progress, there is a 70% probability of it being converted to a pending review state, a 20% probability of it being converted to a timed-out state, and a 10% probability of it being directly converted to a closed state (possibly because the construction unit completed the rectification in advance and obtained the approval of the supervisor).

[0082] Based on the current node state and the trained path prediction model, the model outputs possible future state transition paths and the probability values ​​of each path. For example, if the current problem rectification process is in the rectification stage, the path prediction model predicts the following possible paths and their probability values: Under rectification → Pending re-inspection, probability is 70%.

[0083] Under rectification → Timeout has occurred, with a probability of 20%.

[0084] Under rectification → Closed, probability is 10%.

[0085] The probability value is calculated based on the relation weights and historical transition frequencies of the dynamic reasoning graph. The relation weights in the dynamic reasoning graph reflect the strength of the association between nodes and the reliability of the causal relationship. The historical transition frequency represents the ratio of the number of times a certain state transition occurs to the total number of times in historical data. The path prediction model takes into account these factors and calculates the probability value of each path, providing supervisors with predictive information on the future development trend of the problem rectification process and helping them to take countermeasures in advance.

[0086] The visual dashboard unit displays the entire problem rectification process in the form of a timeline. On the timeline, key nodes in the problem rectification process are arranged in chronological order, including the time of problem discovery, the time of rectification instruction issuance, the time of rectification feedback, and the time of review. Each node is represented by an icon, and the color and shape of the icon are distinguished according to the status of the node. For example, green indicates pending assignment, blue indicates rectification in progress, yellow indicates pending review, red indicates timeout, and gray indicates closure. The timeline also marks the duration between each node, so that supervisors can intuitively understand the time progress of the problem rectification process.

[0087] The entire process of problem rectification is presented in the form of a flowchart. The flowchart starts with the problem event node, connects to the rectification instruction node through trigger relationships, then connects to the rectification feedback node through response relationships, and finally connects to the review node through verification relationships. In the flowchart, arrows are used to represent the relationships between nodes. The thickness of the arrows can be adjusted according to the strength value of the relationship; the greater the strength value, the thicker the arrow. At the same time, detailed information about each node is displayed next to the flowchart, such as the attribute information of the problem event, the content of the rectification instruction, and the specific details of the rectification feedback, so that supervisors can gain a deeper understanding of the details of problem rectification.

[0088] It provides interactive functions, allowing supervisors to view detailed information about the rectification process of different issues through clicks, drags, and other operations. For example, supervisors can click on a node on the timeline to bring up a detailed information window that displays the node's attribute information, relationship information, and related documents, such as rectification instruction text and rectification feedback images. Supervisors can also drag nodes on the timeline to view the status changes of the rectification process in different time periods, or zoom in and out of the flowchart to view the global structure or local details of the entire rectification process.

[0089] It supports filtering and sorting functions, allowing supervisors to filter problem rectification processes based on problem type, project location, responsible party, and other conditions to quickly find the problems of concern. They can also sort the problem rectification processes based on indicators such as the severity of the problem and the progress of rectification, prioritizing the handling of serious problems and processes that are lagging behind.

[0090] The voice interaction unit is equipped with voice recognition technology, supporting natural language queries and command input. When supervisors issue query commands to the system via voice, such as querying the rectification status of quality issues in basic engineering, the voice recognition module converts the voice signal into text information, and then transmits the text information to the system's query processing module. The query processing module parses the query intent and relevant parameters based on the text content, such as if the query question is a quality issue and the query project location is basic engineering. Then, it extracts the relevant information from the dynamic context graph and returns the results to the voice synthesis module in text form.

[0091] Configure speech synthesis technology to convert the text information returned by the system into speech signals, and provide feedback on the query results to the supervisors in the form of voice. For example, the system will announce the rectification status of the quality problem of the foundation project as follows: the problem is insufficient concrete strength, the rectification instruction has been issued, the rectification period is 7 days, it is currently in the rectification stage, and the expected completion time is 2023-04-22, etc., so that the supervisors can obtain information when it is inconvenient to look at the screen.

[0092] Employing natural language processing technology, the system performs semantic understanding and intent recognition on the natural language input via voice. It can understand various expressions used by supervisors, such as "checking the rectification progress of the concrete strength issue" and "how is the concrete strength insufficient now?", which represent the same query intent. By building a dialogue management system, the system can engage in multi-round dialogues with supervisors to clarify ambiguous query intents and obtain more accurate query parameters. For example, if a supervisor only says "check the rectification progress," the system can ask, "Which part of the project's rectification progress do you want to check?" Only after receiving a clear answer from the supervisor can the query operation be performed.

[0093] It supports voice command input, allowing supervisors to operate the system directly via voice commands, such as marking the status of a problem event [problem number] as closed. After recognizing the voice command, the system updates the status of the corresponding node in the dynamic event graph and provides voice feedback on the operation result, improving the work efficiency of supervisors.

[0094] The report generation unit and template configuration subunit have preset multiple supervision report templates, including engineering quality supervision report templates, engineering safety supervision report templates, and engineering progress supervision report templates. Each template contains a fixed structure and format, such as cover, table of contents, main text (including problem overview, rectification status, review results, etc.), conclusion, and attachments. The content areas in the template are represented by placeholders, and the data filling subunit extracts the corresponding field data from the dynamic event graph to fill them.

[0095] The templates are customized according to different report types and purposes. For example, the engineering quality supervision report template will describe in more detail the detection methods, rectification standards, and re-inspection results of quality problems; while the engineering progress supervision report template will focus on showing the comparison between the engineering progress plan and the actual progress, the analysis of the causes of progress deviations, and adjustment measures.

[0096] The data population subunit extracts corresponding field data from the dynamic event graph and, according to the requirements of the report template, populates the relevant locations in the report template with the attribute information of the problem event, the content of the rectification instruction, the specific details of the rectification feedback, and the review results. For example, in an engineering quality supervision report, the attribute values ​​of the problem type, engineering location, and severity of the problem event node are populated in the problem overview section; the attribute values ​​of the instruction content and rectification period of the rectification instruction node are populated in the rectification status section; and the attribute values ​​of the review results and review time of the review node are populated in the review results section.

[0097] During the data filling process, the data is formatted to meet the report layout requirements. For example, date data is formatted as YYYY-MM-DD, and numerical data is kept to two decimal places. At the same time, the extracted data is logically checked to ensure its integrity and consistency. If any missing or abnormal data is found, a prompt message is sent to the supervisor in a timely manner, requesting the supplementation or correction of the data.

[0098] The format optimization sub-unit adjusts the report's layout and chart display format. Based on the amount and importance of the report content, the position and size of elements such as text, tables, and charts are arranged reasonably. For example, important rectification data is displayed using prominent colors or bold fonts; complex progress data is displayed intuitively using Gantt charts or line graphs. In terms of layout, the principles of simplicity, clarity, and aesthetics are followed to avoid pages that are too crowded or too sparse.

[0099] The charts in the report are optimized and designed, and appropriate chart types are selected to display different types of data. For example, bar charts can be used to show the quantity distribution of different problem types, pie charts can be used to show the proportion of problem severity, and line charts can be used to show the trend of problem rectification progress over time. At the same time, elements such as titles, legends, and axis labels are added to the charts to ensure that the charts can clearly convey information. After the report is generated, the supervisors can preview and modify the report. The system provides simple editing functions, such as adjusting text content and modifying chart data, to meet the personalized needs of the supervisors.

[0100] The sample collection unit collects new supervision case data, including text documents such as supervision logs and acceptance reports from completed engineering supervision projects, image documents such as construction site photos and engineering drawings, and audio documents such as recordings of supervision meetings and on-site instructions. During the collection process, the diversity and representativeness of the data are ensured, covering engineering supervision projects of different types, scales, and stages.

[0101] The collected supervision case data is labeled and organized. The labeling content includes the attribute information of the problem event, such as the problem type, project location, severity, etc.; the status information of the rectification process, such as the status transition time, rectification period, etc.; the correlation information, such as the trigger relationship, response relationship, etc.; and the final supervision result, such as whether the problem is resolved and whether the rectification is closed. The labeling work can be completed by professional supervision personnel or data labeling teams to ensure the accuracy and consistency of the labeling.

[0102] The labeled supervision case data is stored in the system's database, and a data index is established to facilitate subsequent data query and use. The database adopts a distributed storage architecture, which can efficiently store and manage large-scale multimodal data, while also backing up and protecting the data to prevent data loss or tampering.

[0103] Regularly update and maintain the data in the database, delete outdated or invalid data, add new supervision case data, and manage data versions through the data management system, recording the data update history and changes for easy traceability and auditing.

[0104] The deviation correction unit and the process comparison subunit compare the differences between the actual process node sequence and the predicted process node sequence. They compare the actual problem rectification process node sequence recorded in the dynamic event graph with the process node sequence predicted by the path prediction model one by one, and analyze the differences between the two in terms of node status, state transition time, and correlation. For example, in the actual process, a certain problem event directly enters the closed state after rectification, while in the predicted process, the problem event needs to go through a review node before it can be closed. This difference is the result of the comparison.

[0105] The degree of difference between the actual process and the predicted process is quantified by calculating the difference degree. The difference degree can be calculated using various methods, such as the edit distance algorithm and the cosine similarity algorithm. The edit distance algorithm calculates the minimum number of operations required to convert the predicted process node sequence into the actual process node sequence, such as inserting, deleting, and replacing nodes. The fewer the number of operations, the smaller the difference degree. The cosine similarity algorithm calculates the cosine value of the angle between the actual process node sequence and the predicted process node sequence. The closer the cosine value is to 1, the smaller the difference degree.

[0106] The correction coefficient generation subunit generates corresponding model correction coefficients based on the degree of difference. The correction coefficients are used to adjust the inference weights of the chain inference unit and the relationship parameters of the dynamic reasoning graph. When the degree of difference is large, it means that the model's prediction results deviate significantly from the actual situation, and the model needs to be adjusted more significantly. When the degree of difference is small, it means that the model's prediction results are relatively accurate, and only fine-tuning is needed.

[0107] The correction coefficient can be generated using empirical formulas or machine learning-based methods. Empirical formulas generate correction coefficients according to a certain proportional relationship based on the magnitude of the difference. For example, when the difference is greater than 0.5, the correction coefficient is 0.8; when the difference is between 0.2 and 0.5, the correction coefficient is 0.9; and when the difference is less than 0.2, the correction coefficient is 1.0. Machine learning-based methods involve training a correction coefficient generation model, taking the difference as input, and outputting the correction coefficient. The correction coefficient generation model can use linear regression models, neural network models, etc., and learns the mapping relationship between the difference and the correction coefficient through training with a large amount of sample data.

[0108] The generated correction coefficients are applied to the chain-based reasoning unit and the dynamic reasoning graph. For the chain-based reasoning unit, the reasoning weights are adjusted to make the reasoning process more consistent with the actual situation. For example, if the predicted rectification result of a certain problem event deviates significantly from the actual situation, the weight of that problem event in the reasoning process is reduced, and the weights of other related factors are increased. For the dynamic reasoning graph, the relationship parameters, such as relationship weights and strength values, are adjusted so that the graph can more accurately reflect the actual logical relationship of the problem rectification process.

[0109] The parameter update unit adjusts the correlation parameters of the domain large model unit, chain reasoning unit, and dynamic reasoning graph based on the correction coefficients generated by the deviation correction unit. For the domain large model unit, the weight parameters of the model are adjusted to better adapt to new supervision case data. For example, the weights of the text feature extraction part, image feature extraction part, and speech feature extraction part of the multimodal large model are adjusted according to the correction coefficients to improve the model's ability and accuracy in processing different modal data.

[0110] For chain-based reasoning units, the reasoning weights of the description subunit, identification subunit, quantization subunit, and instruction subunit are adjusted according to the correction coefficient. For example, if the key elements extracted by the description subunit of a problem event deviate significantly from the actual situation, the weight of the description subunit is reduced, while the weights of the identification and quantization subunits are increased, making the reasoning process more dependent on the judgment and quantification results of subsequent stages.

[0111] For dynamic event graphs, the parameters of node type and relationship type in the graph pattern definition unit, as well as the parameters of relationship weight and strength value in the node relationship creation unit and graph update unit, are adjusted according to the correction coefficient. For example, if the triggering relationship between a problem event and a rectification instruction is stronger in the actual process than predicted, the weight and strength value of the triggering relationship are increased so that the graph can more accurately reflect the causal relationship and dynamic changes in the problem rectification process.

[0112] After adjusting the model parameters, the model is retrained using new supervision case data. During the retraining process, the new data is mixed with the original training data, and the training set, validation set, and test set are divided according to a certain ratio. The same training algorithm and optimization strategy as the initial training are used to iteratively train the model until the model's performance indicators, such as accuracy, recall, and F1 score, reach a satisfactory level.

[0113] The retrained model is validated on the validation set to evaluate the performance improvement. If the model's performance metrics are good on the validation set, it indicates that the model optimization is effective and the model can be applied to the actual engineering supervision process. If the model's performance metrics are still not ideal, it is necessary to further analyze the reasons, adjust the model structure or optimization strategy, and retrain and validate until the model reaches the expected performance requirements.

[0114] System Deployment and Operating Environment For hardware environment configuration, a high-performance server is selected as the host for the system, with sufficient processor cores, memory capacity and storage space to meet the needs of complex computing tasks such as multimodal data processing, dynamic event graph construction and model optimization. For example, the server can use a multi-core processor with a main frequency of 3.0GHz or higher, memory capacity of 128GB or higher, and solid-state drive storage space of 1TB or more for fast data reading and writing.

[0115] Configure a graphics processing unit (GPU) accelerator card to accelerate the training and inference process of deep learning models. For example, use NVIDIA's Tesla series GPU accelerator cards, which have powerful parallel computing capabilities and high video memory capacity, and can improve the training speed and inference efficiency of the model.

[0116] To ensure a stable network connection for the system, smooth uploading and downloading of multimodal engineering documents and data transmission between the system and external devices, such as monitoring cameras and sensors at the construction site, the network bandwidth should meet the data traffic requirements during system operation to avoid system response delays or data loss due to network congestion.

[0117] Software environment configuration Install an operating system on the server, such as Linux or Windows Server. Choose the appropriate operating system version based on the specific needs of the system. The operating system should have good stability and compatibility and be able to support the operation of various software components of the system.

[0118] Install deep learning frameworks, such as TensorFlow and PyTorch, to build and train domain-specific large model units, chained inference units, and other deep learning models. The deep learning framework should be compatible with the system's hardware environment and be able to fully utilize the computing power of the GPU accelerator card.

[0119] Install a database management system, such as MySQL, PostgreSQL, or MongoDB, to store the system's data, including multimodal engineering documents, dynamic reasoning graphs, and supervision case data. The database management system should have efficient data storage and query capabilities and support the management of large-scale data.

[0120] Install web server software, such as Apache or Nginx, to build the system's interactive application modules. This will provide web page access services with functions such as visual dashboards, voice interaction, and report generation. The web server software should have good concurrency processing capabilities and security, and be able to support multiple users accessing the system simultaneously.

[0121] Regularly clean and organize the data in the system, deleting expired or invalid data, such as temporary data and duplicate data from completed engineering supervision projects, to free up storage space and improve system operating efficiency.

[0122] Newly collected supervision case data are labeled and organized, and updated to the system database in a timely manner to ensure that the system can obtain the latest data for model optimization and analysis. At the same time, the quality of the data labeling is checked and reviewed, and any labeling errors or inconsistencies are corrected in a timely manner.

[0123] Monitor database performance, perform regular backup and recovery tests to ensure data security and reliability. If database performance degradation is detected, such as slower query speeds or insufficient storage space, optimize and upgrade in a timely manner, such as adjusting the database index structure or adding storage devices.

[0124] Regularly evaluate and update the models in the system. Based on the new data and feedback information collected during system operation, analyze the model's performance indicators, such as accuracy, recall, and F1 score, to determine whether the model needs to be updated. If the model's performance indicators decline, or if the model is found to perform poorly in certain specific scenarios, initiate the model update process.

[0125] During model updates, incremental learning is used when retraining the model. This combines new data with existing training data, avoiding a complete retraining of the model, saving time and computing resources. At the same time, the model's structure and parameters are optimized and adjusted. Based on the actual needs and performance of the system, a suitable model architecture and optimization strategy are selected.

[0126] After the model is updated, rigorous testing and verification are conducted to ensure that the updated model can run normally in various scenarios and that the performance indicators meet the expected requirements. The testing and verification process includes multiple stages such as unit testing, integration testing and system testing to comprehensively check the model's functionality and performance.

[0127] Regularly inspect and maintain each functional module of the system to ensure stable operation. For the multimodal perception and understanding module, check whether the functions such as image preprocessing, feature point detection, and speech recognition are normal. If algorithm failure or performance degradation is found, optimize and repair it in time. For the dynamic event graph construction module, check whether the graph creation, update, and query functions are normal to ensure that the graph can accurately reflect the problem rectification process in the engineering supervision process. For the closed-loop state tracking module, check whether the state machine operation status and early warning function are normal, and promptly handle problems such as abnormal state transitions or missing early warning information.

[0128] Based on user feedback and issues discovered during system operation, the system's functionality was optimized and improved. For example, if users reported that the visualization dashboard was not intuitive enough, the interface design and display method of the visualization dashboard were optimized, and the types of charts and interactive functions were added. If users reported that the recognition accuracy of voice interaction was low, the voice recognition algorithm was optimized to improve recognition accuracy and robustness.

[0129] Regularly release system updates to incorporate optimized functional modules and fixed vulnerabilities. Before releasing an update, conduct comprehensive testing and verification to ensure its stability and compatibility. Provide users with detailed update instructions and operation guides to help them update the system and use the new features.

[0130] The same or similar labels correspond to the same or similar parts; The terms used to describe positional relationships in the accompanying drawings are for illustrative purposes only and should not be construed as limiting this patent. Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. For those skilled in the art, other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all implementation methods here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the claims of the present invention.

Claims

1. A multimodal engineering document intelligent structured extraction and association system for the supervision process, characterized in that, include: The module includes a multimodal perception and understanding module, a dynamic event graph construction module, a closed-loop state tracking module, an interactive application module, and a model optimization module. The multimodal perception and understanding module is used to intelligently parse multimodal engineering documents, which include text documents, image documents, and voice documents. The multimodal perception and understanding module includes a domain large model unit, a chain reasoning unit, and an auxiliary measurement unit. The domain large model unit is generated based on the multimodal large model and fine-tuned with supervision domain data. The chain reasoning unit is configured with multi-step reasoning logic to realize problem identification and instruction generation. The auxiliary measurement unit is used to accurately measure engineering geometric parameters. The dynamic event graph construction module includes a graph pattern definition unit, a node relationship creation unit, and a graph update unit. The graph pattern definition unit presets several node types and relationship types. The node types include problem event nodes, rectification instruction nodes, rectification feedback nodes, and review nodes. The relationship types include trigger relationships, response relationships, and verification relationships. The node relationship creation unit is used to create nodes and relationships between nodes based on the parsing results of the multimodal perception and understanding module. The graph update unit is used to dynamically update node attributes and relationship weights based on the process progress status. The closed-loop state tracking module includes a state machine configuration unit, an intelligent early warning unit, and a process deduction unit. The state machine configuration unit defines state transition rules for each problem rectification process. The state transition rules include the transition conditions for states to be assigned, in rectification, pending review, closed, and timed out. The intelligent early warning unit is used to generate timeout warnings, associated risk warnings, and review failure warnings based on preset thresholds. The process deduction unit predicts the process evolution path based on a dynamic event graph. The interactive application module includes a visual dashboard unit, a voice interaction unit, and a report generation unit. The visual dashboard unit displays the entire process of problem rectification in the form of a timeline and flowchart. The voice interaction unit supports natural language query and command input. The report generation unit automatically generates a supervision report based on dynamic event graph data. The model optimization module includes a sample acquisition unit, a deviation correction unit, and a parameter update unit. The sample acquisition unit is used to collect new supervision case data. The deviation correction unit is used to calculate the deviation value between the actual process and the predicted process. The parameter update unit adjusts the correlation parameters of the domain large model and the dynamic reasoning graph according to the deviation value.

2. The intelligent structured extraction and association system for multimodal engineering documents oriented towards the supervision process as described in claim 1, characterized in that, The domain large model unit is configured with a modal fusion algorithm, which is used to perform weighted fusion of text features, image features and speech features to generate a unified multimodal feature vector. The multimodal feature vector is used for attribute extraction of problem events, and the attributes include problem type, engineering location, severity and responsible party.

3. The intelligent structured extraction and association system for multimodal engineering documents oriented towards the supervision process as described in claim 1, characterized in that, The chain-based reasoning unit includes a description subunit, an identification subunit, a quantization subunit, and an instruction subunit; The description subunit is used to extract key elements from multimodal engineering documents; The identification subunit is used to determine whether the key element constitutes a quality or safety issue; The quantization subunit is used to numerically represent the characteristic parameters of the problem. The instruction subunit is used to generate rectification instruction text that conforms to the specifications.

4. The intelligent structured extraction and association system for multimodal engineering documents oriented towards the supervision process as described in claim 1, characterized in that, The auxiliary measurement unit includes an image preprocessing subunit, a feature point detection subunit, and a size calculation subunit; The image preprocessing subunit is used to perform noise reduction and distortion correction on the engineering images; The feature point detection subunit is used to identify the reference points and the feature points to be detected in the image; The size calculation subunit calculates the actual physical size based on the principle of perspective transformation.

5. The intelligent structured extraction and association system for multimodal engineering documents oriented towards the supervision process as described in claim 1, characterized in that: The node relationship creation unit is configured with a confidence evaluation algorithm, which is used to calculate the reliability value of the association between nodes. The reliability value is positively correlated with the consistency, information integrity and time correlation of multimodal data. When the reliability value is greater than the preset confidence threshold, the creation of the corresponding association is confirmed.

6. The intelligent structured extraction and association system for multimodal engineering documents oriented towards the supervision process as described in claim 1, characterized in that: The graph update unit includes an attribute update subunit and a relation evolution subunit; The attribute update subunit is used to update the node's status label, timestamp, and processing result in real time. The relationship evolution subunit is used to adjust the strength value of the association relationship according to the progress of the process; The intensity value is related to the causal relationship between nodes and the time interval.

7. The intelligent structured extraction and association system for multimodal engineering documents oriented towards the supervision process as described in claim 1, characterized in that: The intelligent early warning unit includes a timeout calculation subunit, a correlation analysis subunit, and a pattern recognition subunit; The timeout calculation subunit is used to calculate the difference between the duration of the current state and the preset time limit; The correlation analysis subunit is used to mine high-frequency node combinations in the dynamic event graph. The pattern recognition subunit is used to identify problem feature patterns that fail multiple times during re-examination.

8. The intelligent structured extraction and association system for multimodal engineering documents oriented towards the supervision process as described in claim 1, characterized in that: The process deduction unit is equipped with a path prediction model. The path prediction model takes the current node status and historical process data as input and outputs possible future state transition paths and probability values ​​for each path. The probability values ​​are calculated and generated based on the relationship weights of the dynamic process graph and the historical transition frequency.

9. The intelligent structured extraction and association system for multimodal engineering documents oriented towards the supervision process as described in claim 1, characterized in that: The report generation unit includes a template configuration subunit, a data filling subunit, and a format optimization subunit; The template configuration subunit has multiple preset supervision report templates; The data filling subunit extracts corresponding field data from the dynamic event graph; The format optimization subunit is used to adjust the report's layout and chart display format.

10. The intelligent structured extraction and association system for multimodal engineering documents oriented towards the supervision process as described in claim 1, characterized in that: The deviation correction unit includes a process comparison subunit and a correction coefficient generation subunit; The process comparison subunit is used to compare the differences between the actual process node sequence and the predicted process node sequence. The correction coefficient generation subunit generates corresponding model correction coefficients based on the degree of difference. These correction coefficients are used to adjust the inference weights of the chain inference unit and the relationship parameters of the dynamic reasoning graph.