Construction management and BIM component intelligent association method based on semantic space analysis

By using a semantic space parsing-based approach combined with acoustic enhancement and 3D indexing technology, automated and precise binding of BIM components is achieved, solving the problems of low information association efficiency and high matching error rate in existing technologies. This method is applicable to construction quality inspection of various building projects.

CN121979918APending Publication Date: 2026-05-05CHINA CONSTR FOURTH ENG DIV CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA CONSTR FOURTH ENG DIV CORP LTD
Filing Date
2025-12-31
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

In construction quality inspection, existing BIM models suffer from low information association efficiency, high manual matching error rate, poor stability due to reliance on external hardware, and no direct mapping between on-site fuzzy language and model coding, resulting in low inspection efficiency and poor accuracy.

Method used

By employing a semantic space parsing-based approach, which integrates acoustic enhancement, domain-adaptive semantic parsing, and 3D spatial indexing technologies, we can achieve precise matching between fuzzy voice commands and BIM components at the construction site. This includes voice signal acquisition, noise reduction, speech recognition, entity extraction, spatial index construction, and multi-dimensional matching algorithms, automatically binding construction inspection behaviors with BIM components.

Benefits of technology

It achieves a fully automated process for construction inspection, significantly improves information association efficiency, reduces matching error rate, adapts to high-noise environments at construction sites, is compatible with existing BIM platforms, requires no modification to basic model attributes, and is suitable for residential, public, and municipal engineering projects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121979918A_ABST
    Figure CN121979918A_ABST
Patent Text Reader

Abstract

The invention discloses a construction management and BIM component intelligent association method based on semantic space analysis. The method comprises the steps that original voice signals are collected and processed; converting the voice signal into text information; performing entity recognition on the text information, and extracting entities; extracting space coordinates and attribute information of the components, and constructing a three-dimensional space index; searching all components in the spatial position range by utilizing a three-dimensional space index, and filtering to obtain candidate components; sorting the candidate components based on the space coordinates and the component types of the candidate components, and generating a mapping relationship between relative sequence numbers and component identifiers; mapping the relative sequence number of the entity into a component identifier by utilizing the mapping relationship; calculating the matching degree of each candidate component in multiple dimensions and the total matching degree to obtain a target BIM component; the detection result and the target BIM component are stored in an associated mode, and the attribute or state of the corresponding component in the BIM model is updated. The information association efficiency and the matching accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of building engineering information technology, and in particular to a method for intelligent association of construction management and BIM components based on semantic space parsing. Background Technology

[0002] In construction quality inspection, there are significant differences between the operating habits of on-site personnel and the attribute specifications of the BIM model. Specific issues and technical details are as follows: 1. Low Information Association Efficiency: Existing BIM models use a unique identifier for components: "Floor + Area + Type + Code" (e.g., "3F-A-COL-005"). This requires inspectors to sequentially complete multiple steps within the BIM software environment, including floor selection, area location, component type filtering, and code lookup, to associate the results. This multi-level retrieval method is time-consuming in complex projects and cannot meet the immediate needs of high-frequency on-site inspections. The reason for this is that the BIM component identification system is optimized for the design and modeling phases, lacking a rapid retrieval mechanism and natural language mapping capabilities specific to this work scenario.

[0003] 2. High error rate in manual matching: Multiple components with highly similar appearances and specifications exist on the same floor or in the same area (e.g., 10 columns or beams of the same model on the same floor). When manually distinguishing specific components from visual or textual lists, absolute codes (such as "3F-A-COL-005" and "3F-A-COL-006") are easily confused, leading to incorrect binding of detection data. This is because existing coding and identification schemes favor structured chemically unique identifiers in information expression, failing to directly support on-site semantic descriptions (such as "the second column on the left") and lacking visualization or semantic verification mechanisms.

[0004] 3. Poor stability due to reliance on external hardware: Tagging and positioning solutions such as QR codes and RFID require pre-labeling components during the construction phase. However, dust covering the labels in the construction environment, labels falling off during component hoisting, and the need for multiple labels on large components (such as shear walls) all affect the stability and integrity of information reading. This is because the physical carriers of these hardware labeling methods are easily affected by the construction environment and have high maintenance costs. When label information is missing or damaged, it is impossible to reliably associate it with BIM model attributes.

[0005] 4. Lack of direct mapping between on-site fuzzy language and model coding: BIM models only contain absolute coding attributes (such as component ID, axis coordinates), lacking relative sequence attributes based on spatial arrangement (such as "the Xth root," "the nth from the left"), causing on-site personnel's natural language descriptions to be unable to directly match with model identifiers. This necessitates manual secondary conversion of location and number, increasing operation time and the probability of errors. The root cause is that the current BIM data structure lacks a relative spatial positioning field for construction inspection scenarios and lacks a mechanism combining semantic markup and spatial resolution. Summary of the Invention

[0006] In view of this, the purpose of this invention is to propose a construction management and BIM component intelligent association method based on semantic space parsing. This method does not require modification of the underlying data structure of the BIM model. By integrating acoustic enhancement, domain-adaptive semantic parsing, and three-dimensional spatial indexing technology, it achieves intelligent association of fuzzy voice commands on the construction site with precise matching of BIM components.

[0007] To achieve the above-mentioned technical objectives, the technical solution adopted by this invention is as follows: This invention provides a method for intelligent association of construction management and BIM components based on semantic space parsing, comprising the following steps: Step 1: Collect the original voice signals of on-site personnel in the construction environment, and perform noise reduction processing on the original voice signals to obtain the noise-reduced voice signals; Step 2: Convert the denoised speech signal into text information; perform entity recognition on the text information to extract entities containing spatial location, component type, relative sequence number, detection parameters, and detection results; Step 3: Parse the target building's project IFC file to extract the spatial coordinates and attribute information of the components; construct a 3D spatial index based on the spatial coordinates and attribute information of the components; Step 4: Based on the spatial location of the entity, use the three-dimensional spatial index to retrieve all components within the spatial location range, and filter them according to the component type of the entity to obtain candidate components; Step 5: Sort the candidate components according to the preset spatial coordinate sorting rules based on the spatial coordinates and component type; generate a mapping relationship between relative serial numbers and component identifiers based on the sorting results; use the mapping relationship to map the relative serial numbers of entities to component identifiers; Step 6: Calculate the matching degree of each candidate component in multiple dimensions, and sum the matching degrees of each dimension according to the preset weights to obtain the total matching degree of each candidate component; select the candidate component with the highest total matching degree as the target BIM component. Step 7: Locate the corresponding component in the BIM model based on the component identifier of the target BIM component, associate the detection result with the component and store it, and update the attributes or status of the component.

[0008] Furthermore, step 1 specifically includes: Step 11: Use a directional acoustic acquisition device to collect the original voice signals of on-site personnel in the construction environment. The directional acoustic acquisition device is integrated with a micro vibration mechanism or airflow dust removal mechanism with self-cleaning function to remove dust in the construction environment. Step 12: Cut the acquired raw speech signal into segments according to the preset frame length and frame shift to obtain a series of speech frames. Apply a Hamming window to each speech frame for windowing processing. Step 13: Apply an adaptive filter based on the least mean square algorithm to each windowed speech frame for real-time filtering. The adaptive filter tracks noise features in real time through the least mean square algorithm and dynamically adjusts its step size factor to filter out power frequency interference and background mechanical noise, thus obtaining a denoised speech frame. Step 14: Apply a first-order high-pass filter to the denoised speech frame for high-frequency pre-emphasis processing to obtain the denoised speech signal; the transfer function H(z) of the first-order high-pass filter is: H(z) = 1 - μz -1 Where μ is the pre-emphasis coefficient, ranging from 0.9 to 0.98, to enhance speech details in the 3kHz to 8kHz frequency band; z represents a complex variable; Step 15: Calculate the signal-to-noise ratio (SNR) of the denoised speech signal. If the SNR is lower than the preset threshold, automatically adjust the gain of the directional acoustic acquisition device and reacquire the speech signal.

[0009] Furthermore, step 2 specifically includes: Step 21: Construct and train a speech recognition model. Input the denoised speech signal into the trained speech recognition model for processing to obtain text information. Step 22: Preprocess the text information, use the pre-trained entity recognition model to identify entities containing predefined categories from the preprocessed text information, and perform confidence verification and output on the entity recognition results.

[0010] Furthermore, step 21, which involves constructing and training a speech recognition model, specifically includes: Step 211: Obtain multiple speech samples from the building inspection corpus. Each speech sample is labeled with corresponding standardized text to form a speech-text pair. Divide the speech-text pairs into training set, validation set and test set according to a preset ratio for model training, parameter tuning and performance evaluation, respectively. Step 212: Use a pre-trained speech recognition model based on an encoder-decoder Transformer architecture as the base model; the encoder consists of multiple layers of Transformer blocks and is used to process the input speech features; the decoder consists of multiple layers of Transformer blocks and is used to generate the corresponding text sequence. Step 213: Freeze the parameters of the underlying Transformer block of the encoder so that it does not participate in the update during backpropagation, in order to preserve the speech recognition model's ability to extract general speech features. Step 214: Train the speech recognition model using the training set. During the training process, only the parameters of the high-level Transformer blocks of the encoder and the parameters of all Transformer blocks of the decoder are iteratively trained. The AdamW optimizer is used for optimization, and a mixed precision training strategy is used to speed up the training. The training loss function is the cross-entropy loss function, which is used to calculate the difference between the predicted text sequence output by the speech recognition model and the labeled text sequence. Step 215: Optimize the parameters of the speech recognition model using the validation set, and evaluate the performance of the speech recognition model using the test set, finally obtaining the trained speech recognition model; Step 22 specifically includes: Step 221: Define the entity category, including spatial location, component type, relative sequence number, detection parameters, and detection results; Step 222: Preprocess the text information, including text cleaning and normalization, word segmentation, removal of stop words and punctuation, and index construction and vectorization preparation; text cleaning and normalization refers to removing irrelevant symbols and redundant words, and unifying terminology format; word segmentation refers to dividing the text into words to obtain a series of words; removal of stop words and punctuation refers to removing meaningless stop words and punctuation marks, retaining valid words; index construction and vectorization preparation refers to converting each word into an index and preparing a vector representation corresponding to the vocabulary list; Step 223: The entity recognition model consists of a text vectorization representation layer, a bidirectional contextual feature fusion layer, and a conditional random field annotation and optimal decoding layer; Step 224: Input each preprocessed word into the text vectorization representation layer, and transform each word into a high-dimensional semantic vector through the pre-trained word vector model; Step 225: Input the high-dimensional semantic vector into the bidirectional context feature fusion layer, and process the high-dimensional semantic vector from left to right and from right to left respectively through the bidirectional long short-term memory network to capture the context information before and after each word and output the context feature vector. Step 226: Input the context feature vector into the conditional random field annotation and the optimal decoding layer, use the predefined BIO sequence annotation system to annotate each word, and use the optimal label Viterbi algorithm to decode to obtain the optimal entity label sequence; Step 227: Calculate the confidence level of each entity label in the conditional random field annotation and the output of the optimal decoding layer. If the confidence level is lower than the set threshold, trigger the manual confirmation process and output the entity recognition result after manual review. If the confidence level is not lower than the set threshold, output the entity recognition result. The entity recognition result includes the entity category, entity value and its corresponding confidence level.

[0011] Furthermore, step 22 is followed by: Step 23: Obtain construction inspection text, and extract action-target-result triples from the construction inspection text using a pre-trained semantic role labeling model and output them; specifically including: Step 231: Generate text samples based on speech-text pairs. Each text sample is labeled with an action-target-result triplet. The labeled text samples are then divided into training, validation, and test sets according to a preset ratio. The action represents the specific behavior performed in the detection scenario, the target represents the specific component affected by the action, and the result represents the state or conclusion obtained after performing the action. Step 232: Load the pre-trained semantic role labeling model and train it using the training set. During training, adjust the output layer and related parameters of the semantic role labeling model according to the labeled action-target-result triplets, so that the semantic role labeling model learns to recognize and associate the semantic roles of actions, targets, and results in the text samples. Step 233: During training, the performance of the semantic role labeling model is monitored using the validation set; Step 234: After training is complete, use the test set to evaluate the performance of the semantic role labeling model; Step 235: Input the construction inspection text to be analyzed into the trained semantic role labeling model; the semantic role labeling model analyzes the input construction inspection text and outputs the identified action entities, target entities and result entities, which are automatically combined into structured action-target-result triples.

[0012] Furthermore, step 3 specifically includes: Step 31: Parse the project IFC file using the IFC parser to extract the spatial coordinates and attribute information of each component in the project IFC file. The attribute information includes the component type, axis position, and floor to which it belongs. Step 32: Use the three-dimensional boundary cube of the entire building as the root node; Step 33: Recursively divide the root node into eight child nodes along the X, Y, and Z coordinate axes to form an octree node structure; each child node represents a cubic region within the building space. Step 34: Based on the spatial coordinates and attribute information of the components, determine the geometric center coordinates or outer box of each component and the spatial range of the child node, and register it to the child node containing the component; Step 35: Store the octree node structure and the list of components contained in each child node to form a three-dimensional spatial index.

[0013] Furthermore, step 4 specifically includes: Step 41: Map the spatial location of the entity to a specific spatial location range; wherein, the floor description is mapped to the corresponding Z coordinate range, and the planar area description is mapped to the X and Y coordinate range through the spatial object definition or grid coordinate derivation in the BIM model. Step 42: Using the spatial location range as the query range, perform a range query in the three-dimensional spatial index: starting from the root node, recursively determine whether the coordinate range of the child node intersects with the query range. If they intersect, enter the child node to continue the search and collect all components in the child nodes that intersect with the query range. Step 43: Based on the component type of the entity, perform type filtering on the components to obtain a list of candidate components containing spatial coordinates and axis positions.

[0014] Furthermore, step 5 specifically includes: Step 51: Extract the spatial coordinates of the candidate components; Step 52: Dynamically determine the spatial coordinate sorting rules based on the component type, and sort all candidate components according to the spatial coordinate sorting rules: If the component type is column, the candidate components are first sorted in ascending order of X-axis coordinates in the architectural coordinate system. If the difference between the X-axis coordinates of two adjacent candidate components is less than a preset threshold, they are further sorted in ascending order of Y-axis coordinates. If the component type is a beam or wall, the candidate components are first sorted in ascending order of Y-axis coordinates in the architectural coordinate system. If the difference between the Y-axis coordinates of two adjacent candidate components is less than a preset threshold, they are further sorted in ascending order of X-axis coordinates. When the spatial coordinates of different candidate components overlap, the axial positions of the candidate components are read and compared for sorting. Step 53: Assign a corresponding relative number to each candidate component based on the current component type sorting result, generate a unique code for each relative number, obtain the floor and area where the candidate component is located based on the spatial coordinates of the candidate component, and generate a unique component identifier for each candidate component based on the floor, area, component type and code, and generate a mapping table containing the mapping relationship between relative numbers and component identifiers under the current component type; Step 54: Determine the target component type based on the component type of the entity, find the corresponding mapping table based on the target component type, and then match the corresponding component identifier from the mapping table based on the relative sequence number of the entity.

[0015] Furthermore, in step 6, the matching degree of each candidate component is calculated across multiple dimensions, and the matching degrees of each dimension are weighted and summed according to preset weights to obtain the total matching degree of each candidate component; specifically including: Step 61: Calculate the matching degree of each candidate component in multiple dimensions, including space, type, serial number, historical records, and detection parameters. For the spatial matching degree S1, the spatial matching degree S1 is calculated based on the Euclidean distance d between the geometric center of the candidate component and the reference coordinate point resolved from the detection parameters. The formula is: S1 = 1 - (d / Dmax), where Dmax is the preset maximum span of the region. When d > Dmax, S1 = 0. For type matching degree S2, it is calculated based on the cosine similarity of word vectors between the type name of the candidate component and the component type text of the entity. The formula is: S2 = cos(θ), where θ is the angle between the word vectors of the type name of the candidate component and the component type text of the entity. For the sequence number matching degree S3, if the candidate component is successfully matched according to the mapping relationship, then S3=1.0; if it is a partial match, then S3=0.8; if it is not matched, then S3=0. For the historical matching degree S4, query the frequency with which the candidate component has been matched by the same relative sequence number in the historical record. If it has been matched, then S4 = 0.9 + 0.1 × (number of matches / total number of tests for the candidate component); otherwise, S4 = 0.5. For the detection parameter matching degree S5, it is determined whether the attribute parameters of the candidate component are consistent with the detection parameters of the entity. If they are completely consistent, then S5=1.0; if they are partially consistent, then S5=0.6; otherwise, S5=0. Step 62: Calculate the total matching degree S for each candidate component using the weighted summation formula: S = w1·S1+ w2·S2+ w3·S3+ w4·S4+ w5·S5; Where w1 represents the weight coefficient of the spatial dimension, w2 represents the weight coefficient of the type dimension, w3 represents the weight coefficient of the sequence dimension, w4 represents the weight coefficient of the historical record dimension, and w5 represents the weight coefficient of the detection parameter dimension, and w1 + w2 + w3 + w4 + w5 = 1; the specific values ​​of w1, w2, w3, w4 and w5 are set according to the component density of different projects.

[0016] Furthermore, step 7 specifically includes: Step 71: Encapsulate the spatial location of the entity, the component type of the entity, the relative sequence number of the entity, the semantic role labeling result, the original voice signal, the noise-reduced voice signal, the on-site photos, and the component identifier of the target BIM component into a detection record in JSON format. Step 72: Store the structured data in the detection records into a relational database, store the unstructured data into a non-relational database, and establish an association index; Step 73: Using the API interface of the BIM platform, locate the corresponding component in the BIM model based on the component identifier of the target BIM component, and write the detection result into the custom attribute field of the component. Step 74: Based on the status of the detection results, trigger the visualization update command of the BIM model to update the material color or display status of the target BIM component in the BIM model.

[0017] By adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art: This invention deconstructs the semantic elements of ambiguous voice descriptions at construction sites and combines them with the spatial layout rules of BIM models to automatically and accurately bind construction inspection behaviors with BIM components. It is applicable to construction quality inspection scenarios for various building projects such as residential buildings, public buildings, and municipal engineering projects, and can be directly integrated into existing BIM management platforms without modifying the basic attributes of the model.

[0018] 1. Significantly improved information association efficiency, completely freeing up manual labor: Achieving a fully automated process of "voice acquisition → component association," eliminating the need for manual component searching in BIM software. It also eliminates the manual step of converting "fuzzy serial number → BIM identifier," directly completing the matching through automatic mapping, further shortening process time.

[0019] 2. Significantly improved matching accuracy and reduced mismatch risk: The multi-dimensional weighted matching algorithm (space + type + sequence number + history + parameters) combined with a domain-adaptive speech recognition model, entity recognition model and semantic role labeling model results in a component matching error rate lower than existing technologies; the fuzzy sequence number mapping has high accuracy, and differentiated sorting rules are designed for different component types such as "column, beam and wall" to cover abnormal scenarios such as coordinate overlap and avoid mismatch caused by sequence number confusion; 3. Strong adaptability to construction scenarios and outstanding anti-interference ability: It integrates physical noise reduction with directional microphones and directional acoustic acquisition devices with self-cleaning function for algorithmic digital noise reduction, adapting to high noise environments of 70-90dB construction sites. 4. No need to modify the BIM model, strong compatibility: The fuzzy sequence mapping algorithm achieves association through "spatial coordinate sorting → mapping table generation", without modifying the original attributes of the BIM model (such as adding the "Xth root" sequence field); it supports mainstream BIM file formats such as IFC, is compatible with commonly used software such as Revit and BIMBase, and can be directly integrated into the existing BIM platform without custom development.

[0020] 5) Multi-component type adaptation, covering diverse testing scenarios: Differentiated sorting rules are designed for core components such as "columns, beams, and walls" (e.g., columns are sorted by X-axis → Y-axis, and beams by Y-axis → X-axis) to adapt to the testing needs of various components in building construction; it supports multiple fuzzy serial number descriptions such as "the Xth one" and "the Xth one from the left", which is compatible with the operating habits of on-site personnel and does not require standardized wording. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is an execution flowchart of a construction management and BIM component intelligent association method based on semantic space parsing provided in an embodiment of the present invention.

[0023] Figure 2 This is a schematic diagram of the framework of a construction management and BIM component intelligent association system based on semantic space parsing provided in an embodiment of the present invention. Detailed Implementation

[0024] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be particularly noted that the following embodiments are for illustrative purposes only and do not limit the scope of the invention. Similarly, the following embodiments are only some, not all, embodiments of the present invention, and all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] Please see Figure 1 and Figure 2 The present invention provides a method for intelligent association of construction management and BIM components based on semantic space parsing, comprising the following steps: Step 1: Collect the original voice signals of on-site personnel in the construction environment, and perform noise reduction processing on the original voice signals to obtain the noise-reduced voice signals; In this embodiment, step 1 specifically includes: Step 11: Use a directional acoustic acquisition device to collect the original voice signal of the on-site personnel in the construction environment at a preset sampling frequency (16kHz). The directional acoustic acquisition device is integrated with a micro vibration mechanism (micro ultrasonic generator) or an airflow dust removal mechanism with self-cleaning function to remove dust in the construction environment. Step 12: The acquired raw speech signal is segmented according to the preset frame length (the length of each segment of the analyzed signal, such as 20ms) and frame shift (the time interval from the start point of the current frame to the start point of the next frame, such as 10ms) to obtain a series of speech frames (each frame contains 320 sampling points). A Hamming window (window function coefficient α=0.54) is applied to each speech frame to reduce spectral leakage. Step 13: Apply an adaptive FIR (Finite Impulse Response) filter based on the Least Mean Square (LMS) algorithm to each windowed speech frame for real-time filtering. The adaptive filter tracks noise features in real time through the LMS algorithm and dynamically adjusts its step size factor to filter out power frequency interference and background mechanical noise, thus obtaining a denoised speech frame. Step 14: Apply a first-order high-pass filter to the denoised speech frame for high-frequency pre-emphasis processing to obtain the denoised speech signal; the transfer function H(z) of the first-order high-pass filter is: H(z) = 1 - μz -1 Where μ is the pre-emphasis coefficient, ranging from 0.9 to 0.98, to enhance speech details in the 3kHz to 8kHz frequency band (corresponding to the high-frequency components of consonants such as "pillar," "beam," and "wall" in human voice); z represents a complex variable; The process includes: a) Input: The denoised speech frame contains low-frequency vowel formants and high-frequency consonants and detail sounds; b) Filtering: The first-order high-pass filter suppresses low-frequency slow-change components (slowly changing parts) in the signal by calculating the weighted difference between the current sampling point and the previous sampling point. c) Output: The high-frequency components, especially the auxiliary frequency components in the range of 3kHz to 8kHz, are preserved and relatively highlighted after filtering; d) Effect: The processed speech signal has improved consonant clarity, more obvious speech details, and improved overall clarity and intelligibility, while the background low-frequency noise is relatively reduced; Step 15: Calculate the signal-to-noise ratio (SNR) of the denoised speech signal. If the SNR is lower than the preset threshold (15dB), automatically adjust the gain of the directional acoustic acquisition device (such as a microphone) (gain range 0-20dB) and reacquire the speech signal; trigger the microphone gain automatic adjustment mechanism, which performs the following closed-loop control: a) Detect the signal-to-noise ratio of the current speech signal; b) Calculate the difference between the signal-to-noise ratio value and the target signal-to-noise ratio threshold; c) Convert the difference into the desired gain adjustment factor; d) Limit the adjustment factor to the range of 0 to 20 dB; e) Adjust the microphone's hardware or software gain parameters according to the aforementioned adjustment factor; The system is triggered to reacquire the voice signal and perform signal-to-noise ratio verification again.

[0026] Step 2: Convert the denoised speech signal into text information; perform entity recognition on the text information to extract entities containing spatial location, component type, relative sequence number, detection parameters, and detection results; In this embodiment, step 2 specifically includes: Step 21: Construct and train a speech recognition model. Input the denoised speech signal into the trained speech recognition model for processing to obtain text information. Specifically, constructing and training the speech recognition model includes: Step 211: Obtain multiple (e.g., 100,000) speech samples from the building inspection corpus (including 30,000 recorded in noisy environments (construction machinery, wind noise, power tools, etc., to enhance the model's noise resistance). Each speech sample is labeled with corresponding standardized text (the text is standardized (removing redundant words and unifying terms such as "column", "beam", "wall")), forming a speech-text pair. The speech-text pairs are divided into a training set (80%, 80,000 samples), a validation set (10%, 10,000 samples), and a test set (10%, 10,000 samples) according to a preset ratio, which are used for model training, parameter tuning, and performance evaluation, respectively. Step 212: A pre-trained speech recognition model based on an encoder-decoder Transformer architecture is used as the base model; the encoder consists of multiple layers of Transformer blocks and is used to process the input speech features; the decoder consists of multiple layers of Transformer blocks and is used to generate the corresponding text sequence; the pre-trained speech recognition model has been trained on large-scale multilingual speech data and has strong feature extraction capabilities. Step 213: Freeze the parameters of the lower-level Transformer blocks of the encoder so that they do not participate in the update during backpropagation. This preserves the speech recognition model's ability to extract general speech features, does not destroy the original pre-trained speech representation, and reduces training computation while avoiding overfitting. Principle: The lower-level Transformer blocks of the encoder mainly learn "general features" (such as speech time-frequency patterns and basic phonemes). The higher-level Transformer blocks of the encoder and all Transformer blocks of the decoder are closer to specific tasks (such as terminology recognition and noise environment adaptation). Effect of freezing: The lower-level weights are preserved without update; only the parameters of the higher-level Transformer blocks of the encoder and all Transformer blocks of the decoder are trained to learn features relevant to the building detection domain. Step 214: Train the speech recognition model using the training set. During the training process, only the parameters of the high-level Transformer blocks of the encoder and the parameters of all Transformer blocks of the decoder are iteratively trained. The AdamW optimizer is used for optimization, and a mixed precision training strategy is used to speed up the training. The training loss function is the cross-entropy loss function, which is used to calculate the difference between the predicted text sequence output by the speech recognition model and the labeled text sequence. Training process design: Forward propagation: Input speech → Convert to Mel spectrum → Encoder extracts features. The weights of the lower-level encoder remain unchanged, and the output is directly fed to the partial-layer encoder + decoder for task adaptation.

[0027] Backpropagation: Gradients are propagated only through the last part of the encoder and decoder layers, while some lower-level gradients are masked and do not participate in the update at all.

[0028] Iterative training: Number of training rounds: 100 epochs.

[0029] Batch size: 32.

[0030] Optimizer: AdamW (weight decay 0.01).

[0031] Mixed precision training (FP16) improves speed and memory utilization.

[0032] Loss function: Cross-entropy loss is used to compare the difference between the predicted text and the labeled text.

[0033] Step 215: Optimize the parameters of the speech recognition model using the validation set, and evaluate the performance of the speech recognition model using the test set, finally obtaining the trained speech recognition model; Step 22: Preprocess the text information, use the pre-trained entity recognition model to identify entities containing predefined categories from the preprocessed text information, and perform confidence verification and output on the entity recognition results. Specifically, it includes: Step 221: Define the entity category, including spatial location, component type, relative sequence number, detection parameters, and detection results; Define five types of entities: "spatial location (e.g., third floor, area A), component type (e.g., column), relative sequence number (e.g., the third one), inspection parameters (e.g., verticality deviation of 5mm), and inspection result (e.g., qualified)". The BIO tagging system is used for annotation, and the specific annotation method is defined as follows: The “B-location” tag indicates the beginning word of a “spatial location” entity; The “I-location” tag indicates an internal term of a “spatial location” entity; The “B-component” tag indicates the start word of a “component type” entity; The “I-Component” tag represents an internal term of a “Component Type” entity; The “B-sequence” tag indicates the start word of a “relative sequence” entity; The “I-Sequence” tag indicates an internal word of a “relative sequence” entity; The “B-parameter” tag indicates the start word of a “detection parameter” entity; The “I-parameter” tag indicates an internal term of the “detection parameter” entity; The “B-Result” tag indicates the start word of a “detection result” entity; The “I-Results” tag indicates an internal term of the “Detection Result” entity; The “O” tag indicates a word that does not belong to any predefined entity.

[0034] The annotation results are: [("Third floor", "B-location"), ("Area A", "I-location"), ("Third root", "B-serial number"), ("column", "B-component"), ("of", "O"), ("verticality", "B-parameter"), ("deviation", "I-parameter"), ("for", "O"), ("5mm", "I-parameter"), (",", "O"), ("inspection", "O"), ("result", "B-result"), ("qualified", "I-result")].

[0035] Step 222: Preprocess the text information, including text cleaning and normalization, word segmentation, removal of stop words and punctuation, and index construction and vectorization preparation; text cleaning and normalization refers to removing irrelevant symbols and redundant words, and unifying terminology format; word segmentation refers to dividing the text into words to obtain a series of words; removal of stop words and punctuation refers to removing meaningless stop words and punctuation marks, retaining valid words; index construction and vectorization preparation refers to converting each word into an index and preparing a vector representation corresponding to the vocabulary list; Step 223: The entity recognition model consists of a text vectorization representation layer, a bidirectional contextual feature fusion layer, and a conditional random field annotation and optimal decoding layer; Step 224: Input each preprocessed word into the text vectorization representation layer, and transform each word into a high-dimensional semantic vector through the pre-trained word vector model; Step 225: Input the high-dimensional semantic vector into the bidirectional context feature fusion layer. Process the high-dimensional semantic vector from left to right and from right to left through the bidirectional long short-term memory network to capture the context information before and after each word and output the context feature vector. This provides rich context information for the conditional random field annotation and the optimal decoding layer, thereby enhancing the accuracy of entity recognition. Step 226: Input the context feature vector into the Conditional Random Field (CRF) annotation and optimal decoding layer, and use a predefined BIO sequence annotation system to annotate each word. The CRF annotation and optimal decoding layer not only consider the local feature scores of each word, but also the transition probabilities between different BIO tags to ensure that the annotation results conform to the BIO sequence annotation system; and use the optimal tag Viterbi algorithm to decode to obtain the optimal entity tag sequence, which conforms to the predefined BIO sequence annotation system. Step 227: Calculate the confidence level of each entity label in the conditional random field annotation and the output of the optimal decoding layer. If the confidence level is lower than a set threshold (e.g., 0.8), it is marked as low confidence, triggering a manual verification process to ensure the accuracy of the results. After manual review, the entity recognition result is output. If the confidence level is not lower than the set threshold, it is marked as high confidence, and the entity recognition result is output. The entity recognition result includes the entity category, entity value, and its corresponding confidence level.

[0036] Step 23: Obtain construction inspection text, and extract action-target-result triples from the construction inspection text using a pre-trained semantic role labeling model and output them; specifically including: Step 231: Generate text samples based on speech-text pairs. Each text sample is labeled with an action-target-result triplet. The labeled text samples are divided into training set, validation set, and test set according to a preset ratio (80% / 10% / 10%). The action represents the specific behavior performed in the detection scenario (such as "detection", "measurement", etc.), the target represents the specific component affected by the action (such as "pillar", "beam", etc.), and the result represents the state or conclusion obtained after performing the action (such as "qualified", "unqualified", etc.). Step 232: Load the pre-trained semantic role labeling model and train the semantic role labeling model using the training set. During the training process, adjust the output layer and related parameters of the semantic role labeling model according to the labeled action-target-result triplet, so that the semantic role labeling model learns to recognize and associate the semantic roles of actions, targets, and results in the text samples. Step 233: During training, the performance of the semantic role labeling model is monitored using the validation set; Step 234: After training is complete, use the test set to evaluate the performance of the semantic role labeling model; Step 235: Input the construction inspection text to be analyzed into the trained semantic role labeling model; the semantic role labeling model analyzes the input construction inspection text and outputs the identified action entities, target entities and result entities, which are automatically combined into structured action-target-result triples.

[0037] For example, if the input text is "The verticality deviation of the detected column is 3mm, and the detection result is qualified", the model output triple is: {action: detection, target: column, result: qualified}.

[0038] Step 3: Parse the target building's project IFC file to extract the spatial coordinates and attribute information of the components; construct a 3D spatial index based on the spatial coordinates and attribute information of the components; In this embodiment, step 3 specifically includes: Step 31: Parse the project IFC file using the IFC parser (IfcOpenShell library) to extract the spatial coordinates and attribute information of each component in the project IFC file. The attribute information includes component type, axis position and floor to which it belongs. Step 32: Use the three-dimensional boundary cube of the entire building as the root node; Step 33: Recursively divide the root node into eight child nodes along the X, Y, and Z coordinate axes to form an octree node structure; each child node represents a cubic region (X / Y / Z range) within the building space. Step 34: Based on the spatial coordinates and attribute information of the components, determine the geometric center coordinates or outer box of each component and the spatial range of the child node, and register it to the child node containing the component; Step 35: Store the octree node structure and the list of components contained in each child node to form a three-dimensional spatial index.

[0039] Step 4: Based on the spatial location of the entity, use the three-dimensional spatial index to retrieve all components within the spatial location range, and filter them according to the component type of the entity to obtain candidate components; In this embodiment, step 4 specifically includes: Step 41: Map the spatial location of the entity to a specific spatial location range; wherein, the floor description is mapped to the corresponding Z coordinate range, and the planar area description is mapped to the X and Y coordinate range through the spatial object definition or grid coordinate derivation in the BIM model. Floor Mapping: BIM models typically have a structured definition (IfcBuildingStorey) for each floor (e.g., the third floor). The system reads its elevation or Z-coordinate range (e.g., 10m–13m) to determine the vertical extent of the third floor in model space.

[0040] Zone (Area A) Mapping: In the BIM model, planar areas are often represented by room / zone objects (IfcSpace or a custom "Zone" identifier). If the spatial object "Area A" is explicitly defined in the BIM file, its X / Y boundary coordinates are obtained directly by name or ID. If the model does not have an explicit spatial object, the grid is derived. For example, Zone A may correspond to the intersection of "Axis A ~ Axis B" and "Axis 1 ~ Axis 3", and these axes all have coordinate values ​​in the BIM.

[0041] Step 42: Using the spatial location range as the query range, perform a range query in the three-dimensional spatial index: starting from the root node, recursively determine whether the coordinate range of the child node intersects with the query range. If they intersect, enter the child node to continue the search and collect all components in the child nodes that intersect with the query range. Step 43: Based on the component type of the entity (such as a column), filter the components by type (filter out components that are different from the component type of the entity and keep only the components that are the same as the component type of the entity) to obtain a list of candidate components containing spatial coordinates and axis positions (only the same component type is left).

[0042] Step 5: Sort the candidate components according to the preset spatial coordinate sorting rules based on the spatial coordinates and component type; generate a mapping relationship between relative serial numbers and component identifiers based on the sorting results; use the mapping relationship to map the relative serial numbers of entities to component identifiers; In this embodiment, step 5 specifically includes: Step 51: Extract the spatial coordinates of the candidate components; Step 52: Dynamically determine the spatial coordinate sorting rules based on the component type, and sort all candidate components according to the spatial coordinate sorting rules: If the component type is a column, the candidate components are first sorted in ascending order of their X-axis coordinates in the architectural coordinate system (X-axis is east-west, Y-axis is north-south, and Z-axis is height) (east → west). If the difference between the X-axis coordinates of two adjacent candidate components is less than a preset threshold (e.g., 0.5m), they are further sorted in ascending order of their Y-axis coordinates (south → north). If the component type is a beam or wall, the candidate components are first sorted in ascending order of Y-axis coordinates in the architectural coordinate system (South → North). If the difference between the Y-axis coordinates of two adjacent candidate components is less than a preset threshold (e.g., 0.5m), they are further sorted in ascending order of X-axis coordinates (East → West). When the spatial coordinates of different candidate components overlap (difference < 0.1m), the axial positions of the candidate components are read and compared (e.g., A-axis intersects A-axis 1) for sorting. Step 53: Based on the sorting results of the current component type, assign a corresponding relative serial number (first root, second root) to each candidate component, generate a unique code for each relative serial number (e.g., 001, 002), obtain the floor and area where the candidate component is located based on its spatial coordinates, and then generate a unique component identifier for each candidate component based on the floor, area, component type, and code (e.g., 3F-A-COL-001, 3F-A-COL-002, ..., 3F-A-COL-005). Generate a mapping table containing the mapping relationship between relative serial numbers and component identifiers under the current component type (one mapping table for each component type). Step 54: Determine the target component type according to the component type of the entity, find the corresponding mapping table based on the target component type, and then match the corresponding component identifier from the mapping table according to the relative serial number of the entity.

[0043] Step 6: Calculate the matching degrees of each candidate component in multiple dimensions, perform weighted summation on the matching degrees of each dimension according to the preset weights to obtain the total matching degree of each candidate component; take the candidate component with the highest total matching degree as the target BIM component. In this embodiment, in Step 6, calculating the matching degrees of each candidate component in multiple dimensions and performing weighted summation on the matching degrees of each dimension according to the preset weights to obtain the total matching degree of each candidate component specifically includes: Step 61: Calculate the matching degrees of each candidate component in multiple dimensions including space, type, serial number, historical record, and detection parameters respectively. For the space matching degree S1, the space matching degree S1 is calculated based on the Euclidean distance d between the geometric center of the candidate component and the reference coordinate point parsed from the detection parameters. The formula is: S1 = 1 - (d / Dmax), where Dmax is the preset maximum span of the area (default 10m). When d > Dmax, S1 = 0. For the type matching degree S2, the type matching degree S2 is calculated based on the cosine similarity of word vectors between the type name of the candidate component and the component type text of the entity. The formula is: S2 = cos(θ), where θ is the angle of the word vectors between the type name of the candidate component and the component type text of the entity; for example, the similarity between "column" and "pillar" is 0. 98, and the similarity between "column" and "beam" is 0.1. For the serial number matching degree S3, if the candidate component is successfully matched according to the mapping relationship, then S3 = 1.0; if it is partially matched, then S3 = 0.8; if it is not matched, then S3 = 0. For the historical matching degree S4, query the frequency of the candidate component being hit by the same relative serial number description in the historical record. If it has been hit, then S4 = 0.9 + 0.1× (number of hits / total number of detections of this candidate component); otherwise, S4 = 0.5. For the detection parameter matching degree S5, judge whether the attribute parameters (such as orientation, size) of the candidate component are consistent with the detection parameters of the entity. If they are completely consistent, then S5 = 1.0; if they are partially consistent, then S5 = 0.6; otherwise, S5 = 0. Step 62: Calculate the total matching degree S of each candidate component through the weighted summation formula: S = w 1·S1+ w2·S2+ w3·S3+ w4·S4+ w5·S5; Where w1 represents the weight coefficient of the spatial dimension, w2 represents the weight coefficient of the type dimension, w3 represents the weight coefficient of the sequence number dimension, w4 represents the weight coefficient of the historical record dimension, and w5 represents the weight coefficient of the detection parameter dimension, and w1 + w2 + w3 + w4 + w5 = 1; the specific values ​​of w1, w2, w3, w4 and w5 are set according to the component density of different projects. In this embodiment, w1 = 0.3, w2 = 0.25, w3 = 0.2, w4 = 0.15, and w5 = 0.1.

[0044] Step 7: Locate the corresponding component in the BIM model based on the component identifier of the target BIM component, associate the detection result with the component and store it, and update the attributes or status of the component.

[0045] In this embodiment, step 7 specifically includes: Step 71: Encapsulate the spatial location of the entity, the component type of the entity, the relative sequence number of the entity, the output of the semantic role labeling model, the original speech signal, the denoised speech signal, the on-site photos, and the component identifier of the target BIM component into a detection record in JSON format. Step 72: Store the structured data (spatial location of entities, component type of entities, relative sequence number of entities, component identifier of target BIM components) in the detection record into a relational database (RDBMS), and store the unstructured data (output results of semantic role labeling model, original speech signal, noise-reduced speech signal, and on-site photos) into a non-relational database (including NoSQL object storage database or document storage database), and establish an associated index; Step 73: Using the API interface (standard BIM open interface) of the BIM platform, locate the corresponding component in the BIM model according to the component identifier of the target BIM component, and write the detection result into the custom attribute field of the component. Step 74: Based on the status of the detection results, trigger the visualization update command of the BIM model to update the material color or display status of the target BIM component in the BIM model, thereby synchronizing the BIM model with the physical site.

[0046] Scenario: On the third floor of a residential project, in an environment with a noise level of 75dB, an inspector is using a mobile terminal to test the flatness of component “3F-A-COL-005”. The voice description is: “The flatness of the fifth column on the third floor, 1.4m from the west, has been tested and found to be satisfactory. A photo has been taken.”

[0047] S1: Voice Acquisition and Noise Reduction The inspector presses the "Record" button in the APP, and the directional microphone collects the voice signal (sampling rate: 16kHz, bit depth: 16bit); the automatic microphone cleaning function is activated, generating high-frequency vibrations through the built-in micro ultrasonic generator to strip the attached dust on the surface of the microphone dust-proof film, ensuring that the dust-proof film is dust-free; noise reduction module processing: Frame cutting: Cut with a frame length of 20ms and a frame shift of 10ms to generate 50-frame voice segments; Adaptive FIR filtering: Filter out mechanical noise in a 75dB environment; Pre-emphasis: Enhance the high-frequency components of consonants such as "column" and "qualified" through a coefficient of 0.97; Calculate SNR = 22dB (>15dB threshold), no need to re-collect, and output the noise-reduced voice.

[0048] S2: Voice to Text and Semantic Analysis The mobile terminal uploads the noise-reduced voice to the server's speech recognition service via 5G; Whisper model (speech recognition model) processing: Convert the voice into text "The flatness inspection of the fifth column in Area A on the third floor, 1.4m west, the result is qualified, and a photo is taken"; Entity recognition service processing: The BiLSTM-CRF model (entity recognition model) outputs entities: Spatial location: Area A on the third floor (confidence 0.94); Component type: column (confidence 0.99); Relative serial number: The fifth one (confidence 0.89); Detection parameters: 1.4m west, flatness (confidence 0.91); Detection result: qualified (confidence 0.98); Attachment: photo (confidence 0.95); Semantic role annotation service outputs a triple: Action (detection) - Target (column) - Result (qualified); The server returns the structured entity data to the mobile terminal (time-consuming 0.3 seconds).

[0049] S3: BIM Analysis and Component Matching The BIM workstation reads the project IFC file through the IFC parser and extracts the attributes of all column components in Area 3F-A: Component list: 3F-A-COL-001 (X = 10.2m, Y = 5.1m), 3F-A-COL-002 (X = 12.4m, Y = 5.1m),..., 3F-A-COL-008 (X = 24.8m, Y = 5.1m); Octree indexing: Retrieve Area 3F-A (X = 8 - 26m, Y = 4 - 8m, Z = 9 - 12m) to obtain 8 candidate column components; Fuzzy sequence mapping: Extract candidate component X-axis coordinates: [10.2,12.4,14.6,16.8,19.0,21.2,23.4,24.8]; Sort by X-axis in ascending order (East → West), generating a mapping table: 1st component → 001, 2nd component → 002, ..., 5th component → 005 (X=19.0m, Y=5.1m); Match “5th component” to 3F-A-COL-005; Multi-dimensional matching calculation: Spatial matching degree (S1): The coordinates corresponding to the detection parameter "1.4m west" (X=19.0m, Y=5.1m+1.4m=6.5m) are d=1.4m away from component 005, S1=1-(1.4 / 10)=0.86; Type matching degree (S2): The semantic similarity between "column" and component type is 0.99, S2=0.99; Sequence matching degree (S3): Mapping successful, S3=1.0; Historical matching degree (S4): This component has no historical records, S4=0.5; Detection parameter matching degree (S5): "Flatness" is consistent with the detectable parameters of the component, S5=1.0; Total matching degree S=0.3×0.86 + 0.25×0.99 + 0.2×1.0 + 0.15×0.5 + 0.1×1.0= 0.8805 (>0.7 threshold); Determine the optimal matching component: 3F-A-COL-005 (A axis intersects 5 axis column), and return to the mobile terminal.

[0050] S4: Data Linkage and BIM Updates The mobile terminal generates an inspection record (JSON format) and uploads a photo (2.5MB) to the MongoDB server. The server writes the structured data into a MySQL database, generating recordId=DET-20240520-001. The BIMBase server uses an API call to locate component 3F-A-COL-005, sets the "Inspection Status" parameter to "Qualified," and replaces the component's material with green (RGB:0,255,0). The data is then synchronized to the mobile terminal app via WebSocket, displaying "Association Successful" and previewing the green component in the BIM model. The inspection personnel confirm the results, completing the entire process.

[0051] The above description is only a part of the embodiments of the present invention and does not limit the scope of protection of the present invention. Any equivalent device or equivalent process transformation made based on the content of the present invention specification and drawings, or direct or indirect application in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for intelligent association of construction management and BIM components based on semantic space parsing, characterized in that, Includes the following steps: Step 1: Collect the original voice signals of on-site personnel in the construction environment, and perform noise reduction processing on the original voice signals to obtain the noise-reduced voice signals; Step 2: Convert the denoised speech signal into text information; perform entity recognition on the text information to extract entities containing spatial location, component type, relative sequence number, detection parameters, and detection results; Step 3: Parse the target building's project IFC file to extract the spatial coordinates and attribute information of the components; construct a 3D spatial index based on the spatial coordinates and attribute information of the components; Step 4: Based on the spatial location of the entity, use the three-dimensional spatial index to retrieve all components within the spatial location range, and filter them according to the component type of the entity to obtain candidate components; Step 5: Sort the candidate components according to the preset spatial coordinate sorting rules based on the spatial coordinates and component type; Generate a mapping relationship between relative serial numbers and component identifiers based on the sorting results; The relative sequence number of an entity is mapped to a component identifier using the mapping relationship. Step 6: Calculate the matching degree of each candidate component in multiple dimensions, and sum the matching degrees of each dimension according to the preset weights to obtain the total matching degree of each candidate component. The candidate component with the highest overall matching degree is selected as the target BIM component; Step 7: Locate the corresponding component in the BIM model based on the component identifier of the target BIM component, associate the detection result with the component and store it, and update the attributes or status of the component.

2. The construction management and BIM component intelligent association method based on semantic space parsing as described in claim 1, characterized in that, Step 1 specifically includes: Step 11: Use a directional acoustic acquisition device to collect the original voice signals of on-site personnel in the construction environment. The directional acoustic acquisition device is integrated with a micro vibration mechanism or airflow dust removal mechanism with self-cleaning function to remove dust in the construction environment. Step 12: Cut the acquired raw speech signal into segments according to the preset frame length and frame shift to obtain a series of speech frames. Apply a Hamming window to each speech frame for windowing processing. Step 13: Apply an adaptive filter based on the least mean square algorithm to each windowed speech frame for real-time filtering. The adaptive filter tracks noise features in real time through the least mean square algorithm and dynamically adjusts its step size factor to filter out power frequency interference and background mechanical noise, thus obtaining a denoised speech frame. Step 14: Apply a first-order high-pass filter to the denoised speech frame for high-frequency pre-emphasis processing to obtain the denoised speech signal; the transfer function H(z) of the first-order high-pass filter is: H(z) = 1 - μz -1 Where μ is the pre-emphasis coefficient, ranging from 0.9 to 0.98, to enhance speech details in the 3kHz to 8kHz frequency band; z represents a complex variable; Step 15: Calculate the signal-to-noise ratio (SNR) of the denoised speech signal. If the SNR is lower than the preset threshold, automatically adjust the gain of the directional acoustic acquisition device and reacquire the speech signal.

3. The construction management and BIM component intelligent association method based on semantic space parsing as described in claim 1, characterized in that, Step 2 specifically includes: Step 21: Construct and train a speech recognition model. Input the denoised speech signal into the trained speech recognition model for processing to obtain text information. Step 22: Preprocess the text information, use the pre-trained entity recognition model to identify entities containing predefined categories from the preprocessed text information, and perform confidence verification and output on the entity recognition results.

4. The construction management and BIM component intelligent association method based on semantic space parsing as described in claim 3, characterized in that, Step 21, which involves constructing and training a speech recognition model, specifically includes: Step 211: Obtain multiple speech samples from the building inspection corpus. Each speech sample is labeled with corresponding standardized text to form a speech-text pair. Divide the speech-text pairs into training set, validation set and test set according to a preset ratio for model training, parameter tuning and performance evaluation, respectively. Step 212: Use a pre-trained speech recognition model based on an encoder-decoder Transformer architecture as the base model; the encoder consists of multiple layers of Transformer blocks and is used to process the input speech features; the decoder consists of multiple layers of Transformer blocks and is used to generate the corresponding text sequence. Step 213: Freeze the parameters of the underlying Transformer block of the encoder so that it does not participate in the update during backpropagation, in order to preserve the speech recognition model's ability to extract general speech features. Step 214: Train the speech recognition model using the training set. During the training process, only the parameters of the high-level Transformer blocks of the encoder and the parameters of all Transformer blocks of the decoder are iteratively trained. The AdamW optimizer is used for optimization, and a mixed precision training strategy is used to speed up the training. The training loss function is the cross-entropy loss function, which is used to calculate the difference between the predicted text sequence output by the speech recognition model and the labeled text sequence. Step 215: Optimize the parameters of the speech recognition model using the validation set, and evaluate the performance of the speech recognition model using the test set, finally obtaining the trained speech recognition model; Step 22 specifically includes: Step 221: Define the entity category, including spatial location, component type, relative sequence number, detection parameters, and detection results; Step 222: Preprocess the text information, including text cleaning and normalization, word segmentation, removal of stop words and punctuation, and index construction and vectorization preparation; text cleaning and normalization refers to removing irrelevant symbols and redundant words, and unifying terminology format; word segmentation refers to dividing the text into words to obtain a series of words; removal of stop words and punctuation refers to removing meaningless stop words and punctuation marks, retaining valid words; index construction and vectorization preparation refers to converting each word into an index and preparing a vector representation corresponding to the vocabulary list; Step 223: The entity recognition model consists of a text vectorization representation layer, a bidirectional contextual feature fusion layer, and a conditional random field annotation and optimal decoding layer; Step 224: Input each preprocessed word into the text vectorization representation layer, and transform each word into a high-dimensional semantic vector through the pre-trained word vector model; Step 225: Input the high-dimensional semantic vector into the bidirectional context feature fusion layer, and process the high-dimensional semantic vector from left to right and from right to left respectively through the bidirectional long short-term memory network to capture the context information before and after each word and output the context feature vector. Step 226: Input the context feature vector into the conditional random field annotation and the optimal decoding layer, use the predefined BIO sequence annotation system to annotate each word, and use the optimal label Viterbi algorithm to decode to obtain the optimal entity label sequence; Step 227: Calculate the confidence level of each entity label in the conditional random field annotation and the output of the optimal decoding layer. If the confidence level is lower than the set threshold, trigger the manual confirmation process and output the entity recognition result after manual review. If the confidence level is not lower than the set threshold, output the entity recognition result. The entity recognition result includes the entity category, entity value and its corresponding confidence level.

5. The construction management and BIM component intelligent association method based on semantic space parsing as described in claim 3, characterized in that, Step 22 is followed by: Step 23: Obtain construction inspection text, and extract action-target-result triples from the construction inspection text using a pre-trained semantic role labeling model and output them; specifically including: Step 231: Generate text samples based on speech-text pairs. Each text sample is labeled with an action-target-result triplet. The labeled text samples are then divided into training, validation, and test sets according to a preset ratio. The action represents the specific behavior performed in the detection scenario, the target represents the specific component affected by the action, and the result represents the state or conclusion obtained after performing the action. Step 232: Load the pre-trained semantic role labeling model and train it using the training set. During training, adjust the output layer and related parameters of the semantic role labeling model according to the labeled action-target-result triplets, so that the semantic role labeling model learns to recognize and associate the semantic roles of actions, targets, and results in the text samples. Step 233: During training, the performance of the semantic role labeling model is monitored using the validation set; Step 234: After training is complete, use the test set to evaluate the performance of the semantic role labeling model; Step 235: Input the construction inspection text to be analyzed into the trained semantic role labeling model; the semantic role labeling model analyzes the input construction inspection text and outputs the identified action entities, target entities and result entities, which are automatically combined into structured action-target-result triples.

6. The construction management and BIM component intelligent association method based on semantic space parsing as described in claim 1, characterized in that, Step 3 specifically includes: Step 31: Parse the project IFC file using the IFC parser to extract the spatial coordinates and attribute information of each component in the project IFC file. The attribute information includes the component type, axis position, and floor to which it belongs. Step 32: Use the three-dimensional boundary cube of the entire building as the root node; Step 33: Recursively divide the root node into eight child nodes along the X, Y, and Z coordinate axes to form an octree node structure; each child node represents a cubic region within the building space. Step 34: Based on the spatial coordinates and attribute information of the components, determine the geometric center coordinates or outer box of each component and the spatial range of the child node, and register it to the child node containing the component; Step 35: Store the octree node structure and the list of components contained in each child node to form a three-dimensional spatial index.

7. The construction management and BIM component intelligent association method based on semantic space parsing as described in claim 1, characterized in that, Step 4 specifically includes: Step 41: Map the spatial location of the entity to a specific spatial location range; wherein, the floor description is mapped to the corresponding Z coordinate range, and the planar area description is mapped to the X and Y coordinate range through the spatial object definition or grid coordinate derivation in the BIM model. Step 42: Using the spatial location range as the query range, perform a range query in the three-dimensional spatial index: starting from the root node, recursively determine whether the coordinate range of the child node intersects with the query range. If they intersect, enter the child node to continue the search and collect all components in the child nodes that intersect with the query range. Step 43: Based on the component type of the entity, perform type filtering on the components to obtain a list of candidate components containing spatial coordinates and axis positions.

8. The construction management and BIM component intelligent association method based on semantic space parsing as described in claim 1, characterized in that, Step 5 specifically includes: Step 51: Extract the spatial coordinates of the candidate components; Step 52: Dynamically determine the spatial coordinate sorting rules based on the component type, and sort all candidate components according to the spatial coordinate sorting rules: If the component type is column, the candidate components are first sorted in ascending order of X-axis coordinates in the architectural coordinate system. If the difference between the X-axis coordinates of two adjacent candidate components is less than a preset threshold, they are further sorted in ascending order of Y-axis coordinates. If the component type is a beam or wall, the candidate components are first sorted in ascending order of Y-axis coordinates in the architectural coordinate system. If the difference between the Y-axis coordinates of two adjacent candidate components is less than a preset threshold, they are further sorted in ascending order of X-axis coordinates. When the spatial coordinates of different candidate components overlap, the axial positions of the candidate components are read and compared for sorting. Step 53: Assign a corresponding relative number to each candidate component based on the current component type sorting result, generate a unique code for each relative number, obtain the floor and area where the candidate component is located based on the spatial coordinates of the candidate component, and generate a unique component identifier for each candidate component based on the floor, area, component type and code, and generate a mapping table containing the mapping relationship between relative numbers and component identifiers under the current component type; Step 54: Determine the target component type based on the component type of the entity, find the corresponding mapping table based on the target component type, and then match the corresponding component identifier from the mapping table based on the relative sequence number of the entity.

9. The construction management and BIM component intelligent association method based on semantic space parsing as described in claim 1, characterized in that, Step 6 involves calculating the matching degree of each candidate component across multiple dimensions, and then weighting and summing the matching degrees of each dimension according to preset weights to obtain the total matching degree of each candidate component; specifically, this includes: Step 61: Calculate the matching degree of each candidate component in multiple dimensions, including space, type, serial number, historical records, and detection parameters. For the spatial matching degree S1, the spatial matching degree S1 is calculated based on the Euclidean distance d between the geometric center of the candidate component and the reference coordinate point resolved from the detection parameters. The formula is: S1 = 1 - (d / Dmax), where Dmax is the preset maximum span of the region. When d > Dmax, S1 = 0. For type matching degree S2, it is calculated based on the cosine similarity of word vectors between the type name of the candidate component and the component type text of the entity. The formula is: S2 = cos(θ), where θ is the angle between the word vectors of the type name of the candidate component and the component type text of the entity. For the sequence number matching degree S3, if the candidate component is successfully matched according to the mapping relationship, then S3=1.0; if it is a partial match, then S3=0.8; if it is not matched, then S3=0. For the historical matching degree S4, query the frequency with which the candidate component has been matched by the same relative sequence number in the historical record. If it has been matched, then S4 = 0.9 + 0.1 × (number of matches / total number of tests for the candidate component); otherwise, S4 = 0.

5. For the detection parameter matching degree S5, it is determined whether the attribute parameters of the candidate component are consistent with the detection parameters of the entity. If they are completely consistent, then S5=1.0; if they are partially consistent, then S5=0.6; otherwise, S5=0. Step 62: Calculate the total matching degree S for each candidate component using the weighted summation formula: S = w1·S1+ w2·S2+ w3·S3+ w4·S4+ w5·S5; Where w1 represents the weight coefficient of the spatial dimension, w2 represents the weight coefficient of the type dimension, w3 represents the weight coefficient of the sequence dimension, w4 represents the weight coefficient of the historical record dimension, and w5 represents the weight coefficient of the detection parameter dimension, and w1 + w2 + w3 + w4 + w5 = 1; the specific values ​​of w1, w2, w3, w4 and w5 are set according to the component density of different projects.

10. The construction management and BIM component intelligent association method based on semantic space parsing as described in claim 1, characterized in that, Step 7 specifically includes: Step 71: Encapsulate the spatial location of the entity, the component type of the entity, the relative sequence number of the entity, the semantic role labeling result, the original voice signal, the noise-reduced voice signal, the on-site photos, and the component identifier of the target BIM component into a detection record in JSON format. Step 72: Store the structured data in the detection records into a relational database, store the unstructured data into a non-relational database, and establish an association index; Step 73: Using the API interface of the BIM platform, locate the corresponding component in the BIM model based on the component identifier of the target BIM component, and write the detection result into the custom attribute field of the component. Step 74: Based on the status of the detection results, trigger the visualization update command of the BIM model to update the material color or display status of the target BIM component in the BIM model.