Equipment inspection method and system based on space intelligence
By establishing spatial grouping and association chains for equipment, and combining visual acquisition devices and carriers, high-precision mapping between physical equipment and digital models and real-time adjustment of inspection queues were achieved. This solved the problem of difficulty in determining mapping relationships during equipment inspection and improved inspection efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING FEIDU TECH CO LTD
- Filing Date
- 2026-03-19
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies struggle to accurately determine the mapping relationship between physical equipment and digital model systems, and cannot adjust the equipment inspection queue through the connections between equipment systems. This results in low inspection efficiency, reliance on personal experience, and a lack of closed-loop feedback and drill mechanisms.
By establishing spatial grouping and association chains of equipment models, an equipment inspection queue is generated. Using semantic recognition and multimodal data processing, a high-precision mapping between equipment entities and digital models is achieved. The inspection queue is adjusted in real time through the connection between equipment systems, and automated inspection is carried out in combination with vision acquisition devices and vehicles.
It improves the accuracy and efficiency of equipment inspection, reduces reliance on personal experience, achieves high-precision mapping and closed-loop feedback between physical equipment and digital models, and enhances the adaptability of inspection tasks.
Smart Images

Figure CN121884476A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of model information processing technology, and in particular to a method and system for equipment inspection based on spatial intelligence. Background Technology
[0002] In the current field of building operation and maintenance and equipment management, with the advancement of smart city and Industry 4.0 concepts, digital inspection has become a core means to ensure the safe operation of large-scale facilities. However, in-depth analysis of existing inspection technology solutions reveals four dimensions of technical bottlenecks and information gaps in practical application scenarios.
[0003] First, there is a serious "information gap." In traditional inspection processes, inspectors observe the operational status of entities in physical space, but the necessary reference information (such as real-time IoT sensor data of the equipment, geometric data of the BIM model during the installation phase, technical specifications, etc.) is often stored in isolated backend systems. When inspectors encounter complex electromechanical equipment, it is difficult to achieve high-precision, real-time mapping between physical entities and their corresponding digital twins on mobile devices. Although some systems have introduced QR code scanning technology, it lacks spatial context awareness and cannot achieve an intuitive "see-and-align" interactive mode.
[0004] Secondly, the production efficiency of inspection records is extremely low. Currently, the mainstream inspection record-keeping method remains at the stage of "manual filling + simple photo uploading." The data generated by this method is mostly unstructured and scattered information, lacking in-depth automated analysis. After completing their on-site work, inspectors typically need to spend a significant amount of time (usually 30% to 50% of total working hours) organizing photos, verifying parameters, and filling out reports that conform to industry standards. This lagging and inefficient recording method leads to a serious delay in data feedback, making it difficult for managers to capture subtle changes in equipment operation trends in a timely manner.
[0005] Furthermore, the operation and maintenance system relies excessively on individual experience. The power distribution, central air conditioning, and water supply systems within large buildings are complex, and the fault logic varies between different brands and models of equipment. Younger or less experienced inspectors often cannot obtain timely expert-level technical support when faced with complex faults (such as abnormal noises, atypical vibrations, or multi-parameter coupling anomalies). While existing expert systems or rule-based alarm systems can handle simple threshold alarms, they struggle with complex scenarios requiring cross-modal logical reasoning (such as combining visual wear characteristics with numerical spectral characteristics for diagnosis).
[0006] Finally, there is a lack of closed-loop feedback and drill mechanisms. Most existing digital twin systems focus on "visualization," that is, unidirectionally reflecting physical data onto a model. After a fault is detected, the system can usually only provide textual suggestions, unable to proactively simulate the impact of maintenance steps on the overall system in the digital space. Inspectors lack virtual verification methods before actual operation, which poses significant safety risks, especially in maintenance involving high-voltage electricity or precision chemical pipelines.
[0007] Chinese Patent Publication No. CN120853023B discloses a method, device, equipment, and medium for safety inspection at construction sites. It utilizes a mobile, wearable, portable camera to capture real-time images of the target construction site, providing a comprehensive view of the site. The real-time images are aligned with a Building Information Model (BIM) based on camera extrinsic parameters to ensure subsequent inspections are performed at the same viewpoint and pixel count, improving accuracy. An improved Transformer network architecture is used to detect differences between real-time image frames and BIM frames, solving the problems of low accuracy and efficiency in manual inspections. Based on camera extrinsic parameters and the depth map of the real-time images, a difference mask is mapped onto the 3D engineering coordinate system of the BIM to obtain 3D coordinate points. These 3D coordinate points are then used to identify different components and generate a 3D heatmap, achieving comprehensive, accurate, and efficient safety inspection of the target construction site.
[0008] Chinese Patent Publication No. CN120411795B discloses a method, system, equipment, and medium for engineering progress analysis based on BIM models and AI video, relating to the field of construction engineering management technology. This includes acquiring regional image data and performing semantic segmentation and environmental modeling. A multi-scale spatial feature network and dynamic skeleton modeling engine are constructed. Three-dimensional environment maps, structurally identifiable area maps, and obstacle probability distribution maps are input into the multi-scale spatial feature network and dynamic skeleton modeling engine, and the output performs structural recognition and defect / anomaly annotation. An engineering response mechanism is triggered based on a risk event scoring function, and the construction path is remotely and dynamically adjusted. The method described in this invention achieves high-precision structural recognition of construction site images by combining semantic segmentation and three-dimensional environment modeling, effectively solving the problems of unclear structural recognition and blurred component boundaries in traditional video surveillance.
[0009] However, the above methods still cannot solve the problem of accurately determining the mapping relationship between the physical equipment and the digital model system during the inspection process. At the same time, they cannot adjust the inspection queue of the equipment through the connection between the equipment systems. Summary of the Invention
[0010] To address this, the present invention provides a spatial intelligence-based equipment inspection method and system to overcome the problems in the prior art of accurately determining the mapping relationship between the equipment entity and the digital model system during the inspection process, and the inability to adjust the equipment inspection queue through the connection between equipment systems.
[0011] To achieve the above objectives, in one aspect, the present invention provides a spatial intelligence-based equipment inspection method, comprising: Establish several equipment models and building models; The devices are grouped to form several device groups; In response to any device being retrieved, a corresponding device inspection queue is generated based on the device grouping. Its characteristic is that, when grouping devices, it includes, The equipment models are grouped according to their spatial location to form spatial groups; The equipment models are grouped according to their equipment relationships to form a chain of relationships; Each spatial group or each association chain has a corresponding association semantic; When any device event occurs within any acquisition period, the event semantics of the device event are decomposed, and several decomposed semantics are generated. The response deconstruction semantics are spatial semantics, invoking the spatial grouping of the device. The response decomposition semantics are numerical semantics, and the associated chain of the device is invoked; The device inspection queue for this device in the current acquisition cycle is generated based on the semantic weight of the disassembly semantics.
[0012] Specifically, the steps for generating a single association chain include: Periodically collect some operating data or some fault data from each device and record them according to the collection cycle; Based on the connection relationships of the devices, a basic association chain is formed for each device; Based on the device events, generate several data sample chains and fault sample chains corresponding to the device. Generate a corresponding association chain based on several data sample chains and / or fault sample chains corresponding to a single device; In the new data collection cycle, the corresponding association chain is invoked based on the device event; The device events include at least operational data and / or fault data; In response to changes in individual operational data in a consistent manner, the underlying correlation chain is adjusted to form a data sample chain; In response to changes in individual fault data, the underlying correlation chain is adjusted to form a fault sample chain.
[0013] Specifically, the steps for generating a single spatial group include: Set the equipment model in the corresponding position on the building model; The architectural model is divided into several continuous spaces; The device model is divided into corresponding spatial groups based on the continuous space; The continuous space is defined by dividing the building model according to its descriptive attributes.
[0014] Specifically, the steps for generating a single spatial group include: The device model is used to form a physical link diagram based on the connection relationships; Based on the physical link diagram, each device model is divided into several continuous links according to the connection type; The device model is divided into corresponding spatial groups based on the continuous links; The connection types are categorized according to the supply relationships between devices.
[0015] Specifically, for a single device group, it corresponds to at least one descriptive attribute; When forming the device group, the following is included: Set device groups that include the same descriptive attributes as associated groups; Set each device in the associated group as an attribute-associated device; Each device is linked together according to its spatial relationship with the grouping of devices. For a single device, the association chains formed by its corresponding attribute-related devices are sorted according to their distance from the device.
[0016] Specifically, the description attribute and the connection type correspond to at least one associated semantic. A single related semantic corresponds to at least one decomposition semantic; For any device event, if it corresponds to only a single device group, the device is determined to be an independent device. If it corresponds to several device groups, the device is determined to be a combined device.
[0017] Specifically, for the independent device, the device group of the device is called as the device inspection queue; For the combined device, determine the semantic weight of each device group corresponding to the device, and generate the corresponding device inspection queue; The steps for determining the semantic weights include: The decomposed semantics are merged to generate the corresponding fused semantics; The decomposed semantics are sorted from high to low according to the recurrence frequency of the fused semantics to generate a semantic queue; The devices corresponding to each semantic are grouped and sorted according to the semantic queue to form corresponding semantic weights.
[0018] In particular, the present invention also provides another equipment inspection method based on spatial intelligence, including: Inspect the initial equipment and determine its operating data and / or fault data, record and upload them as equipment data; The building model and equipment model are invoked, and corresponding event semantics are generated based on the equipment data, and a basic equipment inspection queue is generated. Remove the initial device from the basic equipment inspection queue and add the initial device to the inspection stack; New equipment is inspected according to the equipment inspection queue, and the basic equipment inspection queue is adjusted according to the new equipment data to form an equipment inspection queue. Retrieve the inspection stack and remove the corresponding devices from the equipment inspection queue; Record and upload data from each device and the corresponding inspection data.
[0019] On the other hand, the present invention also provides a space-based intelligent equipment inspection system, comprising: Several data acquisition devices are used to collect images, operating values, and sounds from the device and generate corresponding device parameters; The server is connected to each collector to record and process the device parameters; The feature is that the server further forms a device inspection queue corresponding to the device based on the device parameters; Several vehicles, each equipped with a corresponding data collector, are used to respond to the equipment inspection queue to inspect each device; The data collectors include a vision data collector, a positioning data collector, a vibration data collector, and a device data collector. The device collector can connect to the corresponding device to read the device's operating parameters.
[0020] Specifically, the server includes: A visual processing module for processing information about the appearance of the device and the scene where the vehicle is located; A numerical processing module used to determine equipment parameters and inspection cycles; Tag processing module used to determine operating status and / or fault status; A positioning processing module used to locate buildings and architectural models based on visual information; as well as, A semantic recognition module used to perform semantic recognition on the device information output by each processing module; An inspection queue module used to generate the equipment inspection queue based on the semantic recognition results; The semantic recognition module is trained and encapsulated using the previous inspection cycle before performing semantic recognition. The semantic recognition also records and categorizes information about devices with the same semantic meaning.
[0021] Based on this, the present invention also provides a model and deployment method that support semantic recognition, which constructs and organizes samples through training data, and obtains a semantic recognition model through offline training before deployment, including: The training samples are divided into stages, including cross-modal alignment pre-training, multi-task supervision on the equipment operation and maintenance dataset, and alignment based on preferences or constraints. Multimodal spatial semantic alignment includes running a lightweight feature extraction model on the terminal side to output key points, local descriptors or geometric features such as edges / contours in real time for subsequent tracking and 3D reconstruction, continuously tracking and mapping the inspection scene, and mapping points in the local 3D coordinate system to the global model. After spatial alignment is completed, the system determines the equipment instance based on the unique identifier of the equipment component in the BIM and / or twin base, and establishes semantic links.
[0022] Compared with the prior art, the beneficial effects of the present invention are that by grouping building equipment, the equipment can be associated with each other. When any equipment malfunctions during operation, the associated equipment can be inspected synchronously through the grouping of equipment. This avoids the inability of traditional inspection methods to determine equipment failures caused by system problems. At the same time, by using merging semantics to associate equipment, the problem of overlooking equipment operation or equipment failure due to subjective judgment is avoided. This effectively and accurately determines the mapping relationship between the equipment entity and the digital model system. Furthermore, by using the above method, the inspection queue of equipment can be adjusted through the connection between equipment systems, thereby effectively improving the accuracy of building equipment inspection.
[0023] Furthermore, by implementing semantic-driven approaches based on the spatial location of equipment and the relationships between equipment, the problem of non-targeted equipment retrieval caused by relying on static scheduling and fixed templates for inspection is avoided. At the same time, by decomposing equipment events into spatial semantics and numerical semantics and assigning different semantic weights, the inspection queue is adjusted in real time, thereby effectively improving the adaptability of inspection tasks to sudden failures and complex related scenarios.
[0024] Furthermore, by training semantic recognition, the inspection results are unified, and by monitoring operational data, the basic association chain is automatically adjusted and a data sample chain or fault sample chain is generated, enabling the system to continuously iterate and thus further improve the accuracy of building equipment inspection.
[0025] Furthermore, by utilizing the perspective of the inspection and monitoring device itself, proactive inspections are performed on the inspection equipment. By invoking semantic retrieval, the proactive inspection method is adjusted, which effectively improves the targeting of proactive inspections and the accuracy of building equipment inspections.
[0026] On the other hand, by setting up several data collectors, vehicles, and servers, equipment identification and positioning, real-time data association, inspection queue recommendation, fault risk prediction, and report generation can be automatically completed at the inspection site. This improves the consistency and efficiency of inspection operations, effectively reduces reliance on personal experience, and at the same time, effectively enhances the accuracy of building equipment inspection.
[0027] This invention can also achieve the following functions: (1) Align the images / videos and pose coordinates of the inspection site to the building coordinate system, automatically match them to the device instances in the BIM, and associate the device codes with IoT points to achieve rapid binding of “location-device-data”.
[0028] (2) Automatically select features related to the current operating conditions from real-time / historical operating parameters, evaluate the priority of each device, recommend the next device to be inspected, and generate an inspection queue.
[0029] (3) Integrate visual features, sensor values and equipment technical data to output fault types, probability / risk levels and handling suggestions for different equipment, and output the inspection conclusions and evidence in a formatted manner according to preset fields to form an archiveable structured professional report; when necessary, combine digital twin simulation to verify the handling plan. Attached Figure Description
[0030] Figure 1 This is flowchart a of the equipment inspection method based on spatial intelligence according to the present invention; Figure 2 This is a flowchart illustrating the steps involved in generating a single associated chain according to the present invention. Figure 3 The flowchart r is the equipment inspection method based on spatial intelligence of the present invention; Figure 4 This is a connection diagram of the equipment inspection system based on spatial intelligence according to the present invention; Figure 5 This is a schematic diagram of the overall logical architecture of the system according to an embodiment of the present invention; Figure 6 This is a logic diagram of the visual-coordinate-semantic alignment algorithm according to an embodiment of the present invention; Figure 7 This is a diagram of the multimodal large model inference engine architecture according to an embodiment of the present invention; Figure 8This is a flowchart of the digital twin closed-loop feedback mechanism according to an embodiment of the present invention; Figure 9 This is a logical template structure diagram for the automatic inspection report generation in an embodiment of the present invention. Detailed Implementation
[0031] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0032] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0033] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.
[0034] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0035] For ease of understanding, the symbols and explanations involved in this solution are described in Table 1, which is a reference table of symbols and expressions of this invention: Table 1. Symbol and Expression Comparison Table
[0036] Please see Figure 1 As shown, it is a flowchart a of the equipment inspection method based on spatial intelligence of the present invention, including: Step Sa1: Create several equipment models and building models; Step Sa2: Group the devices to form several device groups; Step Sa3: In response to any device being retrieved, generate the corresponding device inspection queue based on the device group; Specifically, when grouping devices, this includes: Step Sa21: Group the device models according to their spatial location to form spatial groups; Step Sa22: Group the device models according to their device relationships to form a relationship chain; Each spatial group or each association chain has a corresponding association semantic; When any device event occurs within any acquisition period, the event semantics of the device event are decomposed, and several decomposed semantics are generated. The response deconstruction semantics are spatial semantics, invoking the spatial grouping of the device. The response decomposition semantics are numerical semantics, and the associated chain of the device is invoked; The device inspection queue for this device in the current acquisition cycle is generated based on the semantic weight of the disassembly semantics.
[0037] Compared with the prior art, the beneficial effects of the present invention are that by grouping building equipment, the equipment can be associated with each other. When any equipment malfunctions during operation, the associated equipment can be inspected synchronously through the grouping of equipment. This avoids the inability of traditional inspection methods to determine equipment failures caused by system problems. At the same time, by using merging semantics to associate equipment, the problem of overlooking equipment operation or equipment failure due to subjective judgment is avoided. This effectively and accurately determines the mapping relationship between the equipment entity and the digital model system. Furthermore, by using the above method, the inspection queue of equipment can be adjusted through the connection between equipment systems, thereby effectively improving the accuracy of building equipment inspection.
[0038] Please see Figure 2 As shown, this is the generation step of a single association chain in this invention, including: Step Sa221: Periodically collect some operating data or some fault data from each device and record them according to the collection cycle; Step Sa222: Based on the connection relationships of each device, form the basic association chain corresponding to each device; Step Sa223: Based on the device events, generate several data sample chains and fault sample chains corresponding to the device. Step Sa224: Generate the corresponding association chain based on several data sample chains and / or fault sample chains corresponding to a single device; Step Sa225: In the new acquisition cycle, the corresponding association chain is invoked based on the device event; Among them, equipment events include at least operational data and / or fault data; In response to changes in individual operational data, the underlying correlation chain is adjusted to form a data sample chain; In response to changes in individual fault data, the underlying correlation chain is adjusted to form a fault sample chain.
[0039] Specifically, the steps for generating a single spatial group include: Set the equipment model in the corresponding position on the building model; The architectural model is divided into several continuous spaces; The device model is divided into corresponding spatial groups based on the continuous space; The continuous space is divided according to the descriptive attributes of the building model during the design process.
[0040] Specifically, the steps for generating a single spatial group include: The device models are organized into a physical link diagram based on their connection relationships; Based on the physical link diagram, each device model is divided into several continuous links according to the connection type; The device model is divided into corresponding spatial groups based on the continuous links; The connection type is classified according to the supply relationship between devices.
[0041] Specifically, for a single device group, there is at least one descriptive attribute; When forming device groups, the following are included: Set device groups that include the same descriptive attributes as associated groups; Set each device in the associated group as an attribute-associated device; Each device is linked together according to its spatial relationship with the grouping of devices. For a single device, the association chains formed by its corresponding attribute-related devices are sorted according to their distance from the device.
[0042] By enabling semantic-driven inspections based on the spatial location of devices and the relationships between devices, the problem of non-targeted device retrieval caused by relying on static scheduling and fixed templates is avoided. At the same time, by decomposing device events into spatial semantics and numerical semantics and assigning different semantic weights, the inspection queue is adjusted in real time, thereby effectively improving the adaptability of inspection tasks to sudden failures and complex related scenarios.
[0043] Specifically, the description attributes and connection types correspond to at least one association semantic; A single related semantic corresponds to at least one decomposition semantic; For any device event, if it corresponds to only a single device group, the device is determined to be an independent device. If it corresponds to several device groups, the device is determined to be a combined device.
[0044] Specifically, for an independent device, the device group is used as a device inspection queue; For combined devices, determine the semantic weights of each device group corresponding to the device, and generate the corresponding device inspection queue; The steps for determining semantic weights include: The decomposed semantics are merged to generate the corresponding fused semantics; The decomposed semantics are sorted from high to low according to the recurrence frequency of the fused semantics to generate a semantic queue; The devices corresponding to each semantic are grouped and sorted according to the semantic queue to form corresponding semantic weights.
[0045] By training semantic recognition, the inspection results are unified, and by monitoring operational data, the basic association chain is automatically adjusted and a data sample chain or fault sample chain is generated, enabling the system to continuously iterate and thus further improve the accuracy of building equipment inspection.
[0046] Please see Figure 3 As shown, it is a flowchart r of the equipment inspection method based on spatial intelligence of the present invention, including: Step Sr1: Check the initial equipment and determine the equipment's operating data and / or fault data, and record and upload them as equipment data; Step Sr2: Invoke the building model and equipment model, generate corresponding event semantics based on equipment data, and generate a basic equipment inspection queue; Step Sr3: Delete the initial device from the basic equipment inspection queue and add the initial device to the inspection stack; Step Sr4: Check the new equipment according to the equipment inspection queue, and adjust the basic equipment inspection queue according to the new equipment data to form an equipment inspection queue. Step Sr5: Retrieve the inspection stack and remove the corresponding devices from the device inspection queue; Step Sr6: Record and upload the data of each device and the corresponding inspection data.
[0047] By utilizing the perspective of the inspection and monitoring device itself, proactive inspections are performed on the inspection equipment. Furthermore, by invoking semantic retrieval, the proactive inspection method is adjusted, which effectively improves the targeting of proactive inspections and the accuracy of building equipment inspections.
[0048] Please see Figure 4 As shown, it is a connection diagram of the equipment inspection system based on spatial intelligence of the present invention, including: Several data acquisition devices are used to collect images, operating values, and sounds from the device and generate corresponding device parameters; The server is connected to each collector to record and process device parameters; Its characteristic is that the server also forms a device inspection queue corresponding to the device based on the device parameters; Several vehicles, each equipped with a corresponding data collector, are used to respond to the equipment inspection queue to inspect each device; The data collectors include vision data collectors, positioning data collectors, vibration data collectors, and equipment data collectors. The device data acquisition unit can connect to the corresponding device to read the device's operating parameters.
[0049] Specifically, the server includes: A visual processing module used to process information about the appearance of the equipment and the scene in which the vehicle is located; A numerical processing module used to determine equipment parameters and inspection cycles; Tag processing module used to determine operating status and / or fault status; A positioning processing module used to locate buildings and architectural models based on visual information; as well as, A semantic recognition module used to perform semantic recognition on the device information output by each processing module; Inspection queue module used to generate equipment inspection queues based on semantic recognition results; Among them, the semantic recognition module is trained and encapsulated using the previous inspection cycle before performing semantic recognition; Semantic recognition also records and categorizes information from devices with the same semantic meaning.
[0050] By setting up several data collectors, carriers, and servers, the system can automatically complete equipment identification and positioning, real-time data association, inspection queue recommendation, fault risk prediction, and report generation at the inspection site. This improves the consistency and efficiency of inspection operations, effectively reduces reliance on personal experience, and significantly enhances the accuracy of building equipment inspections.
[0051] The technical solution proposed in this invention includes the following overall process: large-scale offline training and deployment, multimodal spatial semantic alignment, multimodal fusion inference, automatic parameter selection and inspection queue generation, and simulation verification and compliance report generation based on digital twins. These modules organize data around device instances (GUIDs) to ensure consistency across the entire "location-device-status-disposal-archiving" chain.
[0052] S0 Multimodal Large Model Training and Deployment The multimodal large model described in this invention is obtained through offline training before deployment: The training objective is to enable the model to jointly utilize visual, numerical, and textual modalities to complete device identification, automatic parameter selection, fault prediction, inspection queue planning, and report formatting generation, provided that the semantics of "location-device-state" are consistent.
[0053] (1) Training data construction and sample organization The system collects multi-source data from the digital twin platform and the operation and maintenance system, and aligns it using the device unique identifier (GUID) and timestamp as the primary key to form training sample triples / quadruples: Visual Sample I (corresponding symbols are the same as below, i.e., I represents visual sample, and the symbols below are the same): key frames / video clips of device appearance, cropping of local abnormal areas, thermal images (optional). Numerical sample X: Multi-sensor time series (temperature, pressure, current, vibration, flow rate, etc.) associated with the device and its spectral characteristics; Text sample T: Equipment specifications, manual clauses, historical maintenance records, alarm logs, etc.; Tag / Monitoring Signal Y (optional): Fault category, risk level, handling steps, simulation verification conclusion, standardized report fields, etc.
[0054] Meanwhile, the system generates location keywords and equipment keywords based on the equipment's location in the BIM global coordinate system (floor / room / grid segment) and equipment attributes, and writes them into samples as searchable text features (e.g., "B2F_Chiller Room_Grid 3-4 / AB" "Centrifugal Chiller_CH-01") for training retrieval enhancement and formatted generation.
[0055] (2) Phased training method Phase A: Cross-modal alignment pre-training. Contrastive learning and mask modeling are performed on (I,X,T) to align different modalities within a unified semantic space. An example objective function can be written as: , where L_contrast is the InfoNCE-style contrast loss, and L_mask is the mask reconstruction loss (reconstructing image patches, numerical patches, or text tokens).
[0056] Phase B: Supervised Fine-Tuning (SFT). Multi-task supervision is performed on the equipment operation and maintenance dataset, including fault classification / regression (outputting fault probability and risk score), remedial suggestion generation (text-to-text / multimodal-to-text), and report field extraction and filling (multimodal-to-structured text). An example objective function is: .
[0057] Phase C: Preference- or Constraint-Based Alignment (Optional). Model the data using a DPO / reward system to align the output with safety constraints and standard format requirements, prioritizing "passing simulation verification and penalizing violations."
[0058] (3) Example of training parameters (one of the specific implementation methods that can be used to review evidence) In one specific implementation, a text backbone model with a parameter scale of 7B is selected, combined with a visual encoder (ViT class) and a numerical encoder (Time-series Transformer), and the fusion layer adopts a cross-attention structure. Training uses the AdamW optimizer with a base learning rate of 2×10^-4, weight decay of 0.1, warmup ratio of 3%, training epochs of 3-5, maximum sequence length of 2048, gradient clipping threshold of 1.0, and mixed precision of BF16 / FP16. If efficient parameter fine-tuning is used, LoRA can be used with rank=16, alpha=32, dropout=0.05, and the q / k / v / o projection layer is used as the injection module.
[0059] (4) Assessment and deployment After training, the model is evaluated on the offline validation set using metrics such as fault classification F1, risk score MAE, and report field completeness / consistency rate. On the deployment side, quantization (e.g., INT8) and distillation can be used to meet the inference latency requirements of the terminal or edge side.
[0060] S1 Multimodal Spatial Semantic Alignment (Sensing & Alignment) One of the core innovations of this solution lies in achieving a unified alignment of "visual-coordinate-semantic" data, establishing a stable one-to-one correspondence between the physical equipment at the inspection site and the BIM equipment instances in the digital twin base. The technical implementation path is as follows: (1) Visual acquisition and feature extraction The inspection terminal (smartphone, tablet, or AR glasses) captures images of the equipment via a camera. The system runs a lightweight feature extraction model on the terminal side, outputting key points, local descriptors, or geometric features such as edges / contours in real time for subsequent SLAM tracking and 3D reconstruction.
[0061] (2) SLAM reconstruction yields local 3D point sets and camera pose. The system uses SLAM (optionally combined with SfM) to continuously track and map the inspection scene, and outputs: The pose sequence of the camera in the SLAM local coordinate system, T_cam(t); Local 3D feature point set / sparse point cloud obtained by fusing multiple frames of observations , where p_i_SLAM is a three-dimensional point (the unit can be a meter or a millimeter).
[0062] Through this process, the original 2D image observation is elevated to a 3D geometric observation that can be used for alignment, avoiding the unsolvable problem caused by "lack of depth in image coordinates".
[0063] (3) 3D-3D similarity transformation (seven-parameter Helmert) to achieve coordinate unification The system maps a point p_SLAM in the SLAM local 3D coordinate system to a point p_BIM in the BIM global building coordinate system using a seven-parameter similarity transformation, with the following linear form:
[0064] in: It is a translation vector; To transpose into a column vector; s is the scaling factor; R is the rotation matrix.
[0065] In some implementations, R can be generated by Euler angles, and the rotation order can be explicitly specified, for example:
[0066] And agree that the coordinate system is right-handed (or specify the coordinate system and axis definition you use in the text).
[0067] In some implementations, the system updates the transformation parameters using a sliding window approach (e.g., every 1 second or every 30 frames). Alignment is considered valid when the updated parameter changes meet a stability threshold: for example, the magnitude of the translation vector change. Scale factor change Rotation angle change When any threshold is exceeded, control point rematching or re-initialization of alignment is triggered to suppress SLAM drift.
[0068] Example of parameter variation values: In a computer room scenario, the initial solution yields... , , , , The maximum change in the window update during the subsequent 5-minute inspection process is: , , , , The alignment residual RMS is 0.021 m.
[0069] The above parameters can be solved by matching the corresponding 3D control points / geometric feature point set P_BIM in the BIM model with P_SLAM. The solution method can be least squares or robust estimation (e.g., using RANSAC to remove outliers when there are mismatched points) to ensure alignment stability. After alignment, the system can obtain the transformation T_SLAM_to_BIM from SLAM coordinates to BIM coordinates, thereby determining the position of the inspection terminal in the building's global coordinates and the spatial range of the equipment instances within sight.
[0070] (4) Instance confirmation and GUID semantic mounting After spatial alignment is completed, the system determines the equipment instance based on the unique identifier (GUID) of the equipment component in the BIM / twin base and establishes a semantic link. Then, it retrieves the dynamic metadata and real-time numerical data corresponding to the GUID from the unified data base (CDE) (e.g., voltage, outlet pressure, speed, etc. obtained via MQTT or OPC UA) to form a multimodal context of "geometry-attribute-state", providing a consistent data entry point for subsequent inference.
[0071] In one specific implementation, the system further transforms the "actual location - model instance" association result into a searchable set of keywords: Based on T_SLAM_to_BIM, the system obtains the location of the inspection terminal in the building's global coordinates and generates location keywords by combining BIM spatial zoning information (floors, rooms, grid segments, system zoning); simultaneously, it reads the type code, equipment number, system affiliation, and GUID of the equipment instance to generate equipment keywords. Location keywords and equipment keywords can be used for: ① accurately retrieving documents / historical records associated with the equipment from the unified data foundation; ② serving as prompts or search enhancement (RAG) queries for multimodal large models; ③ serving as a source for formatted population of the "location description / equipment identifier" field in reports.
[0072] S2 Multimodal Large Model Inference Engine (Reasoning Engine) The inference engine proposed in this invention is based on a multi-level Transformer architecture and is used to collaboratively analyze heterogeneous signals from the field and the digital twin base, and output fault cause analysis, risk rating and standardized handling suggestions.
[0073] (1) Input mode definition and organization The input to the inference engine includes at least three main modalities: Visual modalities: RGB images / video frames, point clouds, heat maps, etc., are used to characterize visual evidence such as corrosion, leakage, temperature anomalies, and appearance damage; Numerical modalities: IoT sensor time series, frequency spectrum, discrete state variables, etc., used to characterize operating performance, efficiency shift, fatigue characteristics, and abnormal sound / vibration spectrum, etc. (it is recommended to classify "abnormal sound spectrum" as a numerical / vibration feature). Text modality: BIM structured attributes, unstructured manuals / logs / specifications, etc., used to represent design parameters, maintenance rules, historical failure cases and constraints.
[0074] (2) Numerical Modal Patching and Tokenization For numerical modes, the system introduces a time-series patching operator to slice the one-dimensional sensor sequence into a two-dimensional token sequence according to a fixed time window, so as to adapt to the Transformer structure and align with visual tokens and text tokens.
[0075] (3) Cross-modal alignment and fusion reasoning (linear formula expression) The system utilizes a cross-modal alignment network to map multimodal features to a unified semantic space and achieves cross-modal information convergence through cross-attention. A fusion attention approach using visual tokens as queries can be written as:
[0076] in: The query matrix is obtained by projecting the visual token. , These are the key matrices obtained by projecting numerical tokens and text tokens, respectively. The "projection" is implemented through a trainable linear layer (or an MLP with non-linearity), transferring the latent representations of tokens from different modalities... Dimension mapping to attention required 3D key space, thus obtaining and The dimensional relationship can be represented as: , .
[0077] A value matrix (which can be obtained by concatenating or weighting multimodal value vectors); The attention dimension (which can also be denoted as) Attention modules typically employ a multi-head structure to satisfy... ;in The number of attention heads. The aforementioned and The value is a preset hyperparameter of the network structure, which can be determined through offline experiments based on the deployed computing power, latency constraints, input token length, and target accuracy.
[0078] Through the above mechanism, the model can perform associative reasoning on "visual evidence - numerical trends - textual rules / experience", and output: Failure cause analysis (optional summary of key evidence); Risk rating (e.g., grade or score); Standardized handling recommendations (step-by-step, executable, and verifiable by the twin verification module).
[0079] (4) Automatic parameter selection and inspection queue generation To avoid omissions and biases caused by "human subjective parameter selection", the system is equipped with an automatic parameter selection module: after encoding the multi-sensor sequence of each candidate device and its context text, the importance of each parameter is calculated based on indicators such as attention weight, gradient contribution or mutual information, and the Top-k key parameter set K is selected as the reference input for this decision; where k can be 8 to 20, or filtered by an importance threshold τ (e.g., τ=0.02).
[0080] After obtaining the risk score / failure probability of each device, the system generates an inspection queue by combining spatial distance and task constraints. Let the current inspection device be i, and the candidate devices be j. The model outputs the risk score r_j (0-1), failure probability p_j (multi-class vector or scalar), and key parameter set K_j for device j. Simultaneously, the path distance d(i,j) from i to j is calculated based on the BIM spatial map or indoor navigation map. The system selects the next device based on the following comprehensive score:
[0081] Where w_j is the equipment importance weight (obtained from equipment level, downtime cost, planned maintenance cycle, etc.), Norm(·) is the distance normalization function, and λ_r, λ_p, λ_w, and λ_d are configurable weights. The system can use greedy algorithm, bundle search, or dynamic programming to iteratively optimize the score and output the inspection queue. Each queue element is then used to generate a readable entry in the form of "location keyword + device keyword" for mobile / AR devices to guide execution item by item.
[0082] S3 Closed-Loop Feedback and Automatic Reporting This invention emphasizes a closed-loop feedback mechanism from "virtual to real" to ensure the feasibility and security of the recommendations and to form compliant and traceable operation and maintenance records.
[0083] (1) Twin simulation prior verification and secondary reasoning After the inference engine generates maintenance suggestions, the system first performs virtual instruction verification on the digital twin model, demonstrates the operation with 3D animation or simulation process, and uses the built-in physical engine of the twin base (such as hydraulic calculation model, thermal simulation model) to calculate the changes in key states after the operation.
[0084] In one specific implementation, twin simulation employs a three-layer parameterization approach: "BIM structure - physical model - real-time boundary conditions". ① Geometric and topological parameters: Extract the topology, pipe diameter, length, valve and pump / fan connection relationship from BIM to form a directed graph G(V,E) and edge attributes (length L, diameter D, roughness ε, etc.). ② Equipment performance parameters: Pump curves (head-flow rate), valve flow coefficient K_v, heat exchanger UA value, chiller COP curve, frequency-speed mapping of inverter, etc. can be read from the equipment manual or nameplate; ③ Real-time boundary conditions: The supply and return water temperatures, inlet and outlet pressures, flow rates, and currents are input from IoT data as initial values and boundary conditions for the simulation, and are maintained or updated by interpolation within the simulation window.
[0085] The simulation execution process can be as follows: S3-1 Generate operation script (e.g., valve opening degree) Pump frequency Refrigeration load ratio Water supply set temperature S3-2 Set the simulation step size dt (e.g., 1 s) and the prediction time domain H (e.g., 300 s); S3-3 Run the physical solution (which can be a steady-state solution or an unsteady-state iteration), and output the sequence of key state variables (pressure P(t), flow rate Q(t), temperature T(t), electrical power W(t), etc.); S3-4 Calculate the degree of constraint violation and generate an interpretable summary (maximum exceedance, duration, trigger node / device).
[0086] Example of constraint parameters: Setting an upper pressure limit for a chilled water system Traffic limit Upper limit of water outlet temperature Motor current limit If the simulation prediction satisfies the following conditions within the H window: If the condition is met, the result is "passed"; otherwise, the result is "failed" and the violation item and numerical evidence are returned for secondary reasoning.
[0087] The system compares the simulation results with a set of preset constraints (such as upper pressure limit, upper temperature limit, lower flow rate limit, and safe operating range of equipment). If the comparison fails (e.g., predicting downstream overload or exceeding limits), the system intercepts the suggestion and triggers secondary reasoning to generate an alternative solution that meets the constraints; if the comparison succeeds, the system outputs executable steps and initiates on-site execution / guidance.
[0088] (2) Automatic report generation and standard mapping (compliance structured database entry) After the inspection task is completed, the system initiates an automatic report generation process: automatically summarizing anomalies, evidence images, key parameter comparisons, and handling records, and matching them with the classification codes of national standards (such as GB / T 51269-2017). In some implementations, this matching is achieved jointly through a "standard code mapping table / rule base + model classification results," and the report fields are checked for completeness; missing fields can be filled in by BIM attributes or historical records.
[0089] In one specific implementation, the report generation adopts a formatted text generation method of "structured schema constraints + field validation": the report JSON schema is predefined (fields include device identifier, location keywords, key parameter summary, anomaly evidence, root cause candidates, handling steps, simulation verification conclusion, standard encoding, etc.), and the large model is generated using restricted decoding or template filling to ensure that the output is parsable formatted text.
[0090] Example of formatted text (snippet): {device_id: GUID, location: location keywords, key_params: [{name, value, unit, delta}], diagnosis: [{fault_type, prob,evidence}], action: [{step, tool, safety}], sim_check: {pass, violated_constraints}, standard_code: …}. This formatted text can be directly stored in the database and used for subsequent statistical analysis and accountability.
[0091] The final report is synchronized to the asset history of the digital twin platform in structured JSON format, and a PDF version is generated for management and approval, thereby realizing a closed loop of "inspection - diagnosis - verification - disposal - archiving".
[0092] Furthermore, compared with existing building equipment inspection technologies, this invention, based on the overall process of "multimodal spatial semantic alignment—multimodal reasoning diagnosis—digital twin closed-loop verification and automatic reporting," produces the following beneficial effects in terms of information consistency, diagnostic reliability, execution security, and record traceability: 1) Achieve stable spatial semantic binding between physical devices and twin instances to improve the consistency and availability of information retrieval. This invention utilizes a spatial intelligent alignment method—"SLAM reconstruction followed by 3D-3D alignment + seven-parameter similarity transformation"—to unify the 3D geometric observations obtained on-site by the inspection terminal into the BIM global building coordinate system, and further completes instance-level semantic mounting using the device GUID. As a result, when inspection personnel observe physical equipment, the system can synchronously generate a multimodal context corresponding to that device in the background, encompassing "geometry—attributes—status—history." This reduces data inconsistencies caused by multi-system switching, information gaps, and incorrect instance selection, enabling the aggregation and retrieval of on-site observations and digital twin data on the same object.
[0093] 2) Achieve cross-modal evidence fusion for diagnostic and decision-making outputs, enhancing interpretability and verifiability in complex fault scenarios. This invention constructs a multimodal large-model inference engine that aligns and integrates visual evidence (appearance damage / leakage / thermal anomalies, etc.), numerical evidence (current / pressure / vibration and their spectral characteristics, etc.), and textual evidence (BIM attributes, manual clauses, historical maintenance records, etc.) within a unified semantic space, outputting candidate fault causes, risk levels, and standardized handling suggestions. Compared to single-modal threshold alarms or rule systems, this invention can form a more stable inference chain when multi-source evidence corroborates each other, and improves the interpretability of the results through methods such as "key evidence summaries / corresponding clause citations," facilitating review by maintenance personnel and auditing by managers.
[0094] 3) By using twin simulation for prior verification, a "virtual-to-real" closed-loop control is formed, reducing the risk of misoperation and improving the executability of the response plan. This invention introduces a digital twin simulation verification step after the inference output, transforming the proposed solutions into executable operation scripts for the twin. The twin's underlying physics engine predicts key operational states and compares them against preset safety constraints (such as pressure, flow rate, temperature, and equipment safety operating range). For solutions that do not meet the constraints, the system can intercept and trigger secondary inference to generate alternative solutions; for solutions that meet the constraints, it outputs step-by-step execution guidance. This closed-loop mechanism reduces the systemic risks caused by relying solely on experience for direct operation and improves the executability and safety of proposed solutions in real-world scenarios.
[0095] 4) Implement structured archiving and standardized mapping of inspection results to improve the standardization, exchangeability, and asset traceability of operation and maintenance records. This invention automatically summarizes information such as spatial location, key parameter comparison, abnormal evidence, and handling records after an inspection. It then maps equipment and abnormality categories to the classification codes of national standards (such as GB / T 51269-2017) through a pre-set mapping table / rule base (optionally combined with model classification results), completing field integrity verification and missing information completion. Finally, it generates structured JSON and reviewable PDF reports. This transforms an inspection from "scattered, unstructured records" into "standardized, searchable, and traceable" asset history events, facilitating subsequent statistical analysis, maintenance plan optimization, and accountability.
[0096] 5) Establish a unified end-to-end process framework to support extended deployments with different device types and data source combinations. The alignment, reasoning, and closed-loop modules of this invention organize data and processes around "device instance (GUID)" and allow access to different sensor combinations, different terminal forms, and different twin simulation models under different implementation methods, thus possessing good engineering scalability and being applicable to equipment inspection and maintenance management scenarios in large public buildings, smart parks, data center computer rooms, and complex industrial environments.
[0097] Please see Figure 5 As shown, it is a schematic diagram of the overall logical architecture of the system in an embodiment of the present invention.
[0098] The diagram illustrates the entire data flow from the physical perception layer (cameras, sensors) to the logic processing layer (alignment algorithms, large model engines) and then to the application feedback layer (AR displays, automatic reports).
[0099] Please see Figure 6 As shown, it is a logic diagram of the visual-coordinate-semantic alignment algorithm of an embodiment of the present invention.
[0100] The figure details the mathematical logic of feature point matching, pose matrix calculation, and BIM GUID semantic mounting.
[0101] Please see Figure 7 The diagram shown is an architecture diagram of the multimodal large model inference engine according to an embodiment of the present invention. It demonstrates how the Transformer-based cross-attention mechanism handles the fusion of image tokens, numerical tokens, and text tokens.
[0102] Please see Figure 8 The diagram shown is a flowchart of the digital twin closed-loop feedback mechanism according to an embodiment of the present invention. It illustrates how AI instructions are simulated, verified, and fed back within the virtual twin.
[0103] Please see Figure 9The diagram shown is a logical template structure diagram for the automatic inspection report generation in an embodiment of the present invention. It demonstrates how the system converts unstructured AI output into structured reports according to the GB / T 51269-2017 standard.
[0104] Taking the central air conditioning room of a large public building as an example: The data center contains centrifugal chillers, circulating water pumps, cooling towers, and related valve and piping networks. On-site sensors for temperature, pressure, current, and vibration are deployed, and historical maintenance records and equipment manuals are available. The digital side pre-contains a BIM model of the building and its lightweight digital twin base, with each equipment component having a unique identifier (GUID). Inspection terminals utilize mobile devices with cameras or AR glasses, possessing local inference or edge computing access capabilities, and can access a unified data base (CDE) via the network.
[0105] To ensure spatial alignment stability, this embodiment pre-selects several repeatable fixed geometric features (such as wall corners, beam and column edges, and equipment base corners) in the computer room as three-dimensional control features, and records their three-dimensional coordinates in the BIM coordinate system for subsequent 3D-3D alignment solution and verification.
[0106] S1: Multimodal spatial semantic alignment (3D-3D alignment after SLAM reconstruction) Visual acquisition and feature extraction After entering the equipment room, the inspector uses an inspection terminal to perform a surround scan of the target chiller unit, continuously acquiring a sequence of equipment images. A lightweight feature extraction network runs on the terminal, outputting key points and local descriptor subsets for subsequent SLAM tracking and mapping.
[0107] SLAM mapping and local 3D point set generation The system initiates SLAM, generating a sparse point cloud / 3D feature point set P_SLAM in the local coordinate system of the computer room based on continuous tracking and feature triangulation of multiple frames, and outputs the camera pose sequence T_cam(t) in the SLAM coordinate system. The points obtained here are 3D points p_SLAM (with real or relative scale; if monocular SLAM is used, the scale can be unified through scale estimation in subsequent steps).
[0108] 3D-3D Seven-Parameter Similarity Transformation Solution and Global Coordinate Unification The system establishes matching pairs between the corresponding 3D control point set P_BIM in P_SLAM and the BIM model (which can be manually initialized or automatically generated, and can use robust strategies to remove outliers), and solves the similarity transformation parameters from the SLAM coordinate system to the BIM coordinate system in real time, so that the 3D points satisfy the linear transformation relationship: Where t is the translation vector, s is the scale factor, and R is the rotation matrix. The solved T_SLAM_to_BIM is used to uniformly express the current camera pose and line of sight in the BIM global building coordinate system, thereby determining the inspector's location and the spatial range of the line of sight coverage.
[0109] To facilitate reproduction, this embodiment provides a set of example values for alignment parameters: after matching 8 control points and removing 2 outliers using RANSAC, the following values are obtained: , , , , The alignment residual RMS is 0.021 m. A 1-second sliding window is used for updates during the inspection process, and the maximum parameter drift is... , .
[0110] Device instance verification and GUID semantic mounting After completing coordinate unification, the system performs spatial intersection determination between the view frustum coverage area and the 3D bounding box / geometry of the equipment components in the BIM, identifies the set of equipment instances "currently visible," and selects the one with the highest confidence as the target instance for this inspection. Subsequently, the system reads the GUID of this instance and establishes a semantic link in the twin base: using the GUID as an index, it retrieves the real-time IoT data and historical operation and maintenance records, manual clauses, etc. of the device from the CDE, forming a multimodal context object (geometry—attributes—status—history) for the device.
[0111] Stage output (S1 output): T_SLAM_to_BIM, target device GUID, alignment confidence / matching inlier rate (optional), and the multimodal context corresponding to the GUID (real-time sensor data + text reference + instance attributes).
[0112] S2: Multimodal Large Model Inference Engine (Cross-modal Association Diagnosis) Multimodal input organization The system constructs the input for this inference, which includes at least: Visual modalities: keyframe images / cropped local regions of the target device, thermal images (if any), point clouds or depth segments (if any); Numerical modes: time series and spectral characteristics of sensors such as current, vibration, pressure, and temperature (including abnormal sound / vibration spectra, classified as numerical / vibration features); Text modality: The specifications of the equipment in BIM, historical maintenance records, manual chapters or fault case entries related to this model.
[0113] Numerical sequence patching and tokenization For time series data in numerical modalities, the system employs a "time window slicing" patching strategy to generate numerical tokens: for example, the sequence is segmented with a fixed window length W and a step size H, and each segment is normalized / feature extracted to form a token sequence, adapting to the input format of the Transformer. This token sequence, along with visual tokens and text tokens, is input into the inference engine.
[0114] Cross-modal fusion and inference output The inference engine maps multimodal features to a unified semantic space and achieves cross-modal information convergence through a fusion attention mechanism. The fusion can be expressed in the following linear form:
[0115] In this embodiment, the system performs joint reasoning on "visual evidence (local equipment anomalies / oil stains / part misalignment, etc.)," "numerical trends (current increase, vibration peak, slight pressure drop, etc.)," and "textual constraints (manual fault correspondence, allowable operating range, maintenance steps)," and outputs: 1. Top-k candidate causes of failure and their key evidence summaries; 2. Risk level or risk score (used to determine whether to take immediate action / re-inspect / shut down the machine); 3. Standardized handling recommendations (step-by-step, easy to be verified by subsequent twin verification), and may include necessary tools / security tips.
[0116] Stage outputs (S2 outputs): root cause candidate list (including evidence points), risk level / score, and actionable treatment recommendations (structured fields for easy access to S3).
[0117] S3: Closed-loop feedback and automatic reporting (simulation verification - execution - archiving) Prior verification and constraint determination in twin simulation The system converts the handling suggestions output by S2 into executable operation scripts for the digital twin (e.g., valve opening adjustment, bypass strategy, shutdown / load reduction strategy, etc.) and conducts virtual drills within the digital twin model. The twin's built-in physics engine predicts and calculates key states (e.g., trends in pressure, flow, and temperature) and judges them against a pre-set set of safety constraints, such as: In this embodiment, the simulation model focuses on the chilled water pipe network: the topological relationships and geometric parameters (pipe diameter, length, elevation) of pumps, valves, heat exchangers, and pipe segments are extracted from BIM; pump curves and valve K_v are read from the equipment manual; and real-time boundary conditions are obtained from sensors (inlet and outlet pressure, supply and return water temperature, flow rate, and current). Simulation step size. Predicting the time domain .
[0118] When the proposed action includes "adjust valve opening from 40% to 55%" and "reduce pump frequency from 45 Hz to 40 Hz", the simulation output is... , and The change curve is compared with the threshold constraint. If the prediction exists... or or or (Equivalent land, can be written as) If the threshold is exceeded, the test result will be rejected, and the excess (e.g., P peak 1.72 MPa, lasting 18 s) will be returned as evidence to trigger secondary inference.
[0119] Pressure constraints: .
[0120] Flow constraints: .
[0121] Temperature constraints: .
[0122] Current constraint: Furthermore, the equipment operating conditions must be within the range permitted by the manual.
[0123] If the simulation prediction violates any constraint, the system marks the suggestion as "not passed" and uses "violated constraint + simulation result summary" as new evidence to trigger secondary reasoning to generate an alternative solution that satisfies the constraint; if passed, it enters the execution / guidance phase.
[0124] On-site implementation guidance and result feedback: For schemes that pass simulation verification, the system outputs operation steps to the inspector in the form of AR guidance or mobile step cards, and collects execution result data after execution (such as whether key sensor indicators have fallen back, whether the anomaly has disappeared, etc.). The comparison record of "before execution - after execution" and the final disposal conclusion are written back to the asset history corresponding to the equipment GUID.
[0125] Automatic report generation and standard mapping archiving: After the inspection is completed, the system will automatically summarize: 1. Equipment identification and spatial positioning (results of T_SLAM_to_BIM and GUID binding from S1); 2. Key operating parameter comparison charts (from IoT data sequences and patched summary statistics); 3. Abnormal evidence images (visual keyframes and labeled areas, edge / contrast enhancement optional); 4. Root cause analysis and treatment records (from S2 output and S3 execution write-back).
[0126] The system maps equipment types and anomaly categories to standard codes such as GB / T 51269-2017 based on a "standard coding mapping table / rule base (optionally combined with model classification results)". It also performs completeness checks on report fields, allowing missing fields to be filled in using BIM attributes or historical records. Final output: 1. Structured JSON (used for system database storage and retrieval tracking); 2. PDF report (for management approvals and archiving).
[0127] 3. Stage Outputs (S3 Outputs): Simulation verification conclusions (pass / fail + constraints), execution result write-back records, structured JSON report, and PDF report.
[0128] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
[0129] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A spatial intelligence-based equipment inspection method, comprising an association chain for generating an inspection queue, characterized in that, The steps for generating a single association chain include: Periodically collect some operating data or some fault data from each device and record them according to the collection cycle; Based on the connection relationships of the devices, a basic association chain is formed for each device; Based on the device events, generate several data sample chains and fault sample chains corresponding to the device. Generate a corresponding association chain based on several data sample chains and / or fault sample chains corresponding to a single device; In the new data collection cycle, the corresponding association chain is invoked based on the device event; The device events include at least operational data and / or fault data; In response to changes in individual operational data in a consistent manner, the underlying correlation chain is adjusted to form a data sample chain; In response to changes in individual fault data, the underlying correlation chain is adjusted to form a fault sample chain.
2. The equipment inspection method based on spatial intelligence according to claim 1, comprising: Establish several equipment models and building models; The devices are grouped to form several device groups; When any device is retrieved, a corresponding device inspection queue is generated based on the device grouping. Its characteristic is that, when grouping devices, it includes, The equipment models are grouped according to their spatial location to form spatial groups; The equipment models are grouped according to their equipment relationships to form a chain of relationships; Each spatial group or each association chain has a corresponding association semantic; When any device event occurs within any acquisition period, the event semantics of the device event are decomposed, and several decomposed semantics are generated. The response deconstruction semantics are spatial semantics, invoking the spatial grouping of the device. The response decomposition semantics are numerical semantics, and the associated chain of the device is invoked; The device inspection queue for this device in the current acquisition cycle is generated based on the semantic weight of the disassembly semantics.
3. The equipment inspection method based on spatial intelligence according to claim 2, characterized in that, The steps for generating a single spatial group include: Set the equipment model in the corresponding position on the building model; The architectural model is divided into several continuous spaces; The device model is divided into corresponding spatial groups based on the continuous space; The continuous space is defined by dividing the building model according to its descriptive attributes.
4. The equipment inspection method based on spatial intelligence according to claim 2, characterized in that, The steps for generating a single spatial group include: The device model is used to form a physical link diagram based on the connection relationships; Based on the physical link diagram, each device model is divided into several continuous links according to the connection type; The device model is divided into corresponding spatial groups based on the continuous links; The connection types are categorized according to the supply relationships between devices.
5. The equipment inspection method based on spatial intelligence according to claim 3 or 4, characterized in that, For a single device group, there is at least one descriptive attribute; When forming the device group, the following is included: Set device groups that include the same descriptive attributes as associated groups; Set each device in the associated group as an attribute-associated device; Each device is linked together according to its spatial relationship with the grouping of devices. For a single device, the association chains formed by its corresponding attribute-related devices are sorted according to their distance from the device.
6. The equipment inspection method based on spatial intelligence according to claim 5, characterized in that, The description attribute and connection type correspond to at least one related semantic; A single related semantic corresponds to at least one decomposition semantic; For any device event, if it corresponds to only a single device group, the device is determined to be an independent device. If it corresponds to several device groups, the device is determined to be a combined device.
7. The equipment inspection method based on spatial intelligence according to claim 6, characterized in that, For the independent device, the device group of the device is called as the device inspection queue; For the combined device, determine the semantic weight of each device group corresponding to the device, and generate the corresponding device inspection queue; The steps for determining the semantic weights include: The decomposed semantics are merged to generate the corresponding fused semantics; The decomposed semantics are sorted from high to low according to the recurrence frequency of the fused semantics to generate a semantic queue; The devices corresponding to each semantic are grouped and sorted according to the semantic queue to form corresponding semantic weights.
8. A spatial intelligence-based equipment inspection method, characterized in that, include: Inspect the initial equipment and determine its operating data and / or fault data, record and upload them as equipment data; The building model and equipment model are invoked, and corresponding event semantics are generated based on the equipment data, and a basic equipment inspection queue is generated. Remove the initial device from the basic equipment inspection queue and add the initial device to the inspection stack; New equipment is inspected according to the equipment inspection queue, and the basic equipment inspection queue is adjusted according to the new equipment data to form an equipment inspection queue. Retrieve the inspection stack and remove the corresponding devices from the equipment inspection queue; Record and upload data from each device and the corresponding inspection data.
9. A spatial intelligence-based equipment inspection system, which applies the method described in any one of claims 1-8, comprising: Several data acquisition devices are used to collect images, operating values, and sounds from the device and generate corresponding device parameters; The server is connected to each collector to record and process the device parameters; The feature is that the server further forms a device inspection queue corresponding to the device based on the device parameters; Several vehicles, each equipped with a corresponding data collector, are used to respond to the equipment inspection queue to inspect each device; The data collectors include a vision data collector, a positioning data collector, a vibration data collector, and a device data collector. The device collector can connect to the corresponding device to read the device's operating parameters.
10. The equipment inspection system based on spatial intelligence according to claim 9, characterized in that, The server includes: A visual processing module used to process information about the appearance of the equipment and the scene in which the vehicle is located; A numerical processing module used to determine equipment parameters and inspection cycles; Tag processing module used to determine operating status and / or fault status; A positioning processing module used to locate buildings and architectural models based on visual information; as well as, A semantic recognition module used to perform semantic recognition on the device information output by each processing module; An inspection queue module used to generate the equipment inspection queue based on the semantic recognition results; The semantic recognition module is trained and encapsulated using the previous inspection cycle before performing semantic recognition. The semantic recognition also records and categorizes information about devices with the same semantic meaning.
Citation Information
Patent Citations
Methods, systems, equipment, and media for analyzing project progress based on BIM models and AI video
CN120411795B
Construction site safety detection method, device, equipment and medium
CN120853023B
Method and system for patrol inspection of operation and maintenance of construction equipment based on BIM model
CN109117531A
A method, system, storage medium and electronic device for inspecting electric equipment
CN119787620A
Power plant equipment three-dimensional virtual inspection method and system based on BIM
CN119918304A