A building outer wall hollowing data acquisition method and system
Patent Information
- Application Number
- CN202511861042.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-12-11
AI Technical Summary
[0002]楼宇外墙空鼓是建筑运维中的常见隐患,若未及时检测易引发墙面开裂、脱落,威胁公共安全
[0017]本发明有效突破楼宇外墙阴角、挑檐下方、窗框边缘等特殊部位的检测盲区,通过在常规墙面构建通用嵌入提取器、在特殊部位以少样本支撑集适配原型网络生成类别原型,解决特殊部位样本稀缺导致的模型泛化难题;结合多模态数据增强与模态丢弃技术提升模型抗干扰能力,引入风险位置动态加权因子优化判断逻辑,既避免高风险部位空鼓漏判,又减少低风险部位过度检测;整体方法兼顾检测准确性与实用性,降低采集成本,为外墙空鼓全面、精准检测提供可靠方案。
Smart Images

Figure CN121298904B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data processing technology, specifically relating to a method and system for collecting data on hollow areas in building exterior walls. Background Technology
[0002] Hollow spots in building exterior walls are a common hidden danger in building operation and maintenance. If not detected in time, they can easily lead to wall cracks and peeling, threatening public safety. Currently, the detection of hollow spots in exterior walls mostly relies on traditional methods, such as manual tapping to locate the sound and ultrasonic testing. However, these methods have obvious limitations in special areas such as inside corners, under eaves, and window frame edges.
[0003] Due to structural obstructions and material differences, such as the different materials of window frames and walls, traditional equipment cannot obtain complete and effective signals in the aforementioned special areas. Manual tapping is easily affected by environmental noise, making it impossible to accurately distinguish between hollow areas and structural artifacts. The signals of ultrasonic and other equipment are easily blocked and reflected, resulting in blind spots in the detection, and the hidden dangers of hollow areas are often missed.
[0004] Meanwhile, existing technologies largely rely on training detection models with a large number of labeled samples. However, collecting samples from special locations is difficult and costly, making it difficult to meet the needs of large-sample training and limiting the model's generalization ability. In addition, environmental factors such as wind speed and lighting during the detection process can easily interfere with the signal, further reducing detection accuracy and making it difficult to achieve both comprehensive coverage and accurate judgment. Summary of the Invention
[0005] To address this issue, the present invention provides a method and system for collecting data on hollow areas in building exterior walls, thereby solving the aforementioned technical problems.
[0006] This invention provides a method for collecting data on hollow areas in building exterior walls, comprising the following steps: Multimodal detection signals are collected from the conventional wall surface of the building exterior. The multimodal detection signals include acoustic impact signals, body acceleration signals, and visual image signals. A universal embedding extractor is established based on the multimodal detection signals. A small number of labeled samples were collected from special parts of the building's exterior wall. Based on the labeled samples, a support set containing multimodal raw data and metadata information was constructed. The special parts included inside corners, under eaves, and window frame edges. Based on the general embedding vector output by the general embedding extractor, the support set is adapted through a few-shot learning model to generate the category prototype corresponding to the special part; wherein, the few-shot learning model adopts a prototype network, calculates the mean vector of each category in the embedding space as the category prototype, and performs classification based on Euclidean distance or cosine distance; During the external wall inspection process, the distance between the category prototype and the sample to be tested is used to determine whether the area corresponding to the sample to be tested is a hollow area.
[0007] Furthermore, the universal embedding extractor adopts a multimodal branch parallel architecture. The acoustic or vibration branch transforms the original signal into a time-frequency spectrum through short-time Fourier transform or wavelet transform, and then uses a lightweight convolutional network to extract acoustic or vibration features. The visual texture branch uses a lightweight visual backbone network to extract visual features. The metadata branch encodes the discrete or continuous attributes of the acquisition point category and device orientation to obtain metadata features. The acoustic or vibration features, visual features, and metadata features are concatenated, mapped to a unified embedding space, and then normalized to obtain a universal embedding vector.
[0008] Furthermore, during adaptation, the last layer of the general embedding extractor is fine-tuned iteratively with a small learning rate and an early stopping strategy based on the support set.
[0009] Furthermore, after fine-tuning, the general embedding extractor is called again to generate new general embedding vectors for the support set samples. Based on the mean vector, the special part category prototype library is updated to ensure that the category prototypes match the feature output of the fine-tuned extractor.
[0010] Furthermore, the support set includes a small number of samples for each special part, and each sample contains at least one modal signal and a corresponding annotation result, wherein the annotation result belongs to one of the following: normal, hollow, material splicing or seam, and structural boundary artifact.
[0011] Furthermore, the acoustic striking signal is subjected to time stretching or compression, environmental noise is added, simulated echoes are added, or random cropping is performed. The visual image is subjected to viewpoint transformation, illumination change, partial occlusion, or texture noise injection. In addition, some modal signals are randomly discarded during the training process.
[0012] Furthermore, during the training of the general embedding extractor, the weights of each modality feature are dynamically adjusted based on modal-level attention-weighted fusion. Specifically, based on the feature vector before fusion, the multilayer perceptron outputs three weight values, corresponding to acoustic or acceleration, visual, and metadata modalities, respectively, with the weights summed to 1. The feature vector of each modality is multiplied by the corresponding attention weight, and then summed to obtain the weighted fused feature vector, which is finally mapped to a unified embedding vector.
[0013] In another aspect, this application also provides a building exterior wall hollowness data acquisition system, comprising: A general embedding extractor construction module is used to collect multimodal detection signals on the conventional wall surface of building exteriors. The multimodal detection signals include acoustic impact signals, body acceleration signals, and visual image signals. A general embedding extractor is established based on the multimodal detection signals. The support set construction module is used to collect a small number of labeled samples at special locations on the exterior walls of buildings, and construct a support set containing multimodal raw data and metadata information based on the labeled samples. The special locations include inside corners, under eaves, and window frame edges. The category prototype generation module is used to adapt the support set to the general embedding vector output by the general embedding extractor and generate the category prototype corresponding to the special part by means of a few-shot learning model; wherein, the few-shot learning model adopts a prototype network, calculates the mean vector of each category in the embedding space as the category prototype, and performs classification based on Euclidean distance or cosine distance. The hollow area detection and judgment module is used to determine whether the area corresponding to the test sample is a hollow area by measuring the distance between the category prototype and the test sample during the external wall inspection process.
[0014] In another aspect, this application also provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method for collecting data on hollow exterior walls as described above.
[0015] In another aspect, this application provides a computer-readable storage medium having stored thereon computer program instructions that can be executed by a processor to implement a method for collecting data on hollow areas in building exterior walls as described above.
[0016] Another aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements a method for collecting data on hollow areas in building exterior walls as described above.
[0017] This invention effectively overcomes the detection blind spots in special areas such as the internal corners of building exterior walls, under eaves, and window frame edges. By constructing a universal embedding extractor on regular wall surfaces and using a small sample support set to adapt to a prototype network to generate category prototypes in special areas, it solves the model generalization problem caused by the scarcity of samples in special areas. Combining multimodal data augmentation and modality discarding techniques improves the model's anti-interference ability, and introducing a dynamic weighting factor for risk locations optimizes the judgment logic, avoiding missed detection of hollow areas in high-risk areas and reducing over-detection in low-risk areas. The overall method balances detection accuracy and practicality, reduces data acquisition costs, and provides a reliable solution for comprehensive and accurate detection of hollow areas in exterior walls. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0019] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 A flowchart of a method for collecting data on hollow areas in building exterior walls, provided as an embodiment of the present invention.
[0020] Figure 2 This is a schematic diagram of a general embedding extractor architecture provided for an embodiment of the present invention.
[0021] Figure 3 This is a schematic diagram of the category prototype generation process provided in an embodiment of the present invention.
[0022] Figure 4 This is a schematic diagram of a building exterior wall hollowness data acquisition system provided in an embodiment of the present invention.
[0023] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0025] The technical solutions of this application will be described in detail below with reference to various embodiments.
[0026] like Figure 1 As shown in the figure, an embodiment of the present invention discloses a method for collecting data on hollow areas in building exterior walls, including the following method steps: S101, Collect multimodal detection signals on the conventional wall surface of the building exterior. The multimodal detection signals include acoustic impact signals, body acceleration signals and visual image signals. Establish a universal embedding extractor based on the multimodal detection signals. S102, Collect a small number of labeled samples at special parts of the building exterior wall, and construct a support set containing multimodal raw data and metadata information based on the labeled samples, wherein the special parts include inside corners, under eaves and window frame edges; S103, based on the general embedding vector output by the general embedding extractor, the support set is adapted through a few-shot learning model to generate the category prototype corresponding to the special part; wherein, the few-shot learning model adopts a prototype network, calculates the mean vector of each category in the embedding space as the category prototype, and performs classification based on Euclidean distance or cosine distance. S104, during the external wall inspection process, the distance between the prototype of the category and the sample to be tested is measured to determine whether the area corresponding to the sample to be tested is a hollow area.
[0027] In some embodiments, for S101, a regular wall surface refers to a structurally regular area of the exterior wall of a building that has no special obstructions or shapes, such as a large flat wall surface, which is different from special parts such as inside corners, under eaves, and window frame edges. Signal acquisition in such areas has no blind spots, and a large amount of complete and interference-free multimodal data can be obtained, providing general representation learning samples for general embedding extractors.
[0028] In this embodiment, the collection area is evenly divided on the conventional wall surface, such as each 1m×1m is a collection unit, and 3-5 collection points are set in each unit to cover different orientations, including east, south, west or north, and conventional walls of different heights such as 1-3 floors and 4-6 floors, so as to avoid the data uniformity leading to insufficient generalization ability of the extractor.
[0029] Each collection point synchronously records corresponding metadata, which, for example, includes: Basic information includes: the type of collection point is marked as a regular wall, the orientation of the equipment is such as due south or 15° north of due east, and the collection height; Environmental information includes: collection time, weather conditions (sunny, cloudy, or rainy), wind speed, and light intensity; Operational information, including operator ID, facilitates the traceability of data sources and ensures that environmental features can be supplemented through metadata encoding during subsequent extractor training.
[0030] At the same time, the collected raw signals are preliminarily cleaned to remove invalid data. For example, for acoustic signals, segments containing obvious sudden noises such as horns from passing vehicles are removed, and valid signals are retained; for acceleration signals, abnormal jump data caused by sensor loosening is removed; for visual images, blurry, overexposed, or images with excessively large occlusion areas, such as >30%, are removed to ensure that textures are clear and distinguishable.
[0031] In one embodiment, such as Figure 2As shown, the general embedding extractor adopts a multimodal branch parallelism and feature fusion architecture, designs dedicated processing branches for three types of signals, and finally outputs a uniform-dimensional embedding vector.
[0032] Specifically, for the acoustic or acceleration signal branch, for example, the first step is time-frequency transformation, which converts the time-domain signal into a spectrum. For instance, short-time Fourier transform (STFT) or wavelet transform are performed on the acoustic time-domain signal and the acceleration time-domain signal, respectively. Among them, the STFT parameters can be window length 256-512, overlap rate 50%-75%, to obtain a two-dimensional time-frequency spectrum, where the horizontal axis is time, the vertical axis is frequency, and the pixel value is the signal amplitude, which can clearly show the frequency difference between hollow and normal wall surfaces. For example, the hollow signal has a higher amplitude in the 200-500Hz frequency band. Wavelet transform is performed using the db4 or sym8 wavelet basis, with 3-5 decomposition layers. Low-frequency approximation coefficients and high-frequency detail coefficients are extracted to preserve the transient characteristics of the vibration signal, such as the longer decay time of a hollow drum being struck.
[0033] The second step is feature extraction. This embodiment uses a lightweight network to improve the speed of feature extraction. Specifically, the time-frequency spectrum is input into a lightweight convolutional network, such as a lightweight CNN or a 1D convolutional network. For example... The network structure employs 3-5 convolutional layers, exemplarily with a 3×3 kernel size and a stride of 1, and 2-3 pooling layers, exemplarily with max pooling, with a 2×2 kernel, and 1 fully connected layer to avoid overfitting caused by complex networks. The output of the fully connected layer is mapped to a 64-dimensional feature vector, which is enabled for dimensional matching for subsequent fusion, i.e., acoustic or acceleration feature vectors.
[0034] For the visual image signal branch, for example, the first step is image preprocessing. This involves normalizing the high-resolution visual image, such as scaling pixel values to 0-1, and unifying the size, such as resizing to 224×224, to adapt it to the input of a lightweight visual network.
[0035] The second step is feature extraction, which, for example, employs a lightweight backbone network. For instance, a lightweight visual backbone network at the MobileNetV3 or ResNet-18 level can be used. MobileNetV3 reduces parameters through depthwise separable convolutions, making it suitable for deployment on edge devices; ResNet-18 avoids gradient vanishing through residual connections, making it suitable for extracting deep texture features. The specific model selection is not limited in this invention. Taking MobileNetV3 as an example, the last classification layer of the network is removed, and the output of the penultimate fully connected layer, such as the output of the fc1 layer of ResNet-18, is retained and mapped to a 64-dimensional feature vector, i.e., a visual feature vector.
[0036] For the metadata branch, for example, the first step is metadata encoding. Discrete and continuous attributes in the metadata are processed separately. Specifically, for discrete attributes such as the collection point category "regular wall" and the weather "sunny", one-hot encoding is used, such as regular wall being encoded as [1,0] and sunny being encoded as [1,0,0]; for continuous attributes such as collection height 2.5m and light intensity 5000 lux, Min-Max normalization is used, that is, scaling to the 0-1 range. The second step is feature concatenation and compression. The encoded discrete attribute vector is concatenated with the continuous attribute vector, where the total dimension is 16. This concatenation is then compressed into a 32-dimensional feature vector, i.e., the metadata feature vector, through a fully connected layer.
[0037] In one embodiment, to address the issue of dimensional differences between different branches of features, such as acoustic feature vectors having a numerical range of 0-10 and visual feature vectors having a numerical range of 0-1, this embodiment employs a unified processing method.
[0038] Specifically, dimensional unification: acoustic or acceleration feature vectors (64-dimensional for example), visual feature vectors (64-dimensional for example), and metadata feature vectors (32-dimensional for example) are concatenated to obtain a 160-dimensional feature vector before fusion. L2 normalization: L2 normalization is performed on the 160-dimensional feature vector. For example, the Euclidean norm of the vector is calculated, and each element of the vector is divided by the norm to make the norm of the vector 1. This eliminates numerical scale differences and ensures the fairness of subsequent distance metrics such as Euclidean distance or cosine distance. The unified embedding space mapping uses a fully connected layer to map the normalized 160-dimensional vector to a 128-dimensional or 256-dimensional unified embedding space. The 128-dimensional space is preferred to balance feature representation capability and computational efficiency, and finally a general embedding vector is obtained.
[0039] Optionally, during the training of the general embedding extractor, the weights of each modality feature are dynamically adjusted based on modality-level attention-weighted fusion to improve feature discriminativeness. Specifically, To generate attention weights, we design a short multilayer perceptron, or MLP, which contains one hidden layer and uses ReLU as the activation function. The input is the 160-dimensional feature vector before fusion, and the output is three weight values, corresponding to acoustic or acceleration, visual, and metadata modalities, respectively, with a weight sum of 1. Weighted fusion involves multiplying the feature vectors of each modality by their corresponding attention weights, then summing them to obtain a weighted fused feature vector, which is ultimately mapped to a unified embedding vector. This allows the model to automatically learn the modal importance in different scenarios during training, such as giving higher weight to the visual modality on sunny days and higher weight to the acoustic modality on rainy days, thus avoiding interference from invalid features caused by simple splicing.
[0040] In one embodiment, the general embedding extractor is optimized through a two-stage process of baseline training and episode training to ensure that it can learn the hollow-normal general representation of a regular wall surface.
[0041] Specifically, from the multimodal data collected from regular wall surfaces, the training set and validation set were divided in a 7:3 ratio, and each sample was labeled as normal or hollow.
[0042] Baseline training employs mean squared error loss (MSE), with the mapping error between the general embedding vector and the labeled vector (e.g., 0 for normal and 1 for hollow) as the optimization objective.
[0043] Episode training involves sampling one episode from the training set at a time. Each episode contains: Category sampling: randomly sample N categories, for example, N=2, i.e., normal and hollow. Sample sampling: s support samples are sampled for each class, for example s=5-10, to simulate a small sample size; q query samples are sampled, for example q=10-20, to verify the classification effect. To ensure that the embedding vectors output by the extractor satisfy the condition that samples of the same class are close to each other and samples of different classes are far apart, negative log-likelihood (the main loss) and triplet loss (the auxiliary loss) are used. Negative log-likelihood is calculated based on the Euclidean distance between the embedding vector and the class prototype, and the mean of the embedding vectors of each class's supporting samples. The loss is the negative log-likelihood of the probability and the labeled label. The acquisition of the class prototype will be explained in detail below and will not be repeated here.
[0044] The triplet loss sample takes a triplet of anchor sample-positive sample-negative sample from each episode. The loss function brings the anchor closer to the positive sample and widens the distance between the anchor and the negative sample, thereby enhancing the discriminative power of the features.
[0045] In some embodiments, for S102, the inside angle refers to the angle between the exterior wall surface and the wall surface, or between the wall surface and the roof, which is usually 90° or 135°. Because the angle obstructs the signal reflection of traditional detection equipment, it is easy to miss the hollow areas. The area below the eaves refers to the area below the eaves of the exterior wall that cantilever outwards, such as the edge of the roof or the area below the eaves of the balcony. This area has a structural dead angle where the vertical plane transitions to the horizontal plane, and signal acquisition is easily blocked. The window frame edge refers to the joint between the window frame and the wall surface and the surrounding 5-10cm area. Due to the difference between the window frame material, such as metal or plastic, and the wall material, signal artifacts are easily generated. Traditional methods make it difficult to distinguish between hollow areas and material boundary artifacts.
[0046] In this embodiment, the number of samples collected for each type of special area is controlled to 5-20 to avoid increasing collection costs due to excessive samples. Simultaneously, it ensures that the samples cover different hollow states of the area, including normal, slight hollowness, and obvious hollowness, and environmental conditions. Multimodal data from the same collection point, including acoustic, vibration, and visual data, are collected simultaneously to avoid feature misalignment due to time differences. Each sample is associated with complete metadata to ensure traceability of the collection scenario during subsequent model adaptation, improving the generalization ability with few samples. The specific collection process can be referred to in the previous embodiment for collecting data from a conventional wall surface, and will not be repeated here.
[0047] In this embodiment, each sample contains at least one modal signal and a corresponding annotation result, wherein the annotation result belongs to one of the following: normal, hollow, material splicing or seam, and structural boundary artifact.
[0048] Furthermore, the multimodal raw data, metadata, and annotation labels are categorized and integrated according to special parts to form support sets for three special parts.
[0049] Specifically, support sets for internal corners, under eaves, and window frame edges are constructed separately. Each support set contains all samples for that part, including multimodal raw data, metadata, and annotation labels. Each support set contains 5-20 samples, and the labels are evenly distributed, such as hollow samples accounting for 30%-40%, normal samples accounting for 30%-40%, and material splicing or seams and structural boundary artifacts accounting for 20%-40% in total, to avoid model bias caused by too few samples of a certain type of label.
[0050] In some embodiments, for S103, a sample identifier such as "Corner-01-Hollow" is added to the general embedding vector of each support set sample, and associated with the original sample's part-label group to form a part-label-embedded vector correspondence table, which facilitates subsequent calculation of the category prototype by group.
[0051] Based on the general embedding vector of the support set samples, the mean vector of each special part-label is calculated by the prototype network, which is the category prototype. The prototype network is naturally adapted to the k-shot (k small) scenario. The special part samples are difficult to collect and the number is small. For example, there are only 5-20 samples per class. The prototype network distinguishes the classes by calculating the class mean vector, which does not require a large number of samples to train complex parameters. It can achieve effective classification with few samples, which fits the characteristic of scarce special part data.
[0052] For example, such as Figure 3 The diagram shows the category prototype generation process, and its specific implementation is as follows: S301, grouping of embedded vectors of similar samples. Specifically, grouping according to location-label, such as the under-eaves-normal group containing 8 samples and the under-eaves-hollow group containing 6 samples. Extract all common embedded vectors of each group from the location-label-embedded vector correspondence table to form a set of similar embedded vectors, such as the 8 128-dimensional vectors of the under-eaves-normal group.
[0053] S302, calculate the category prototype, i.e., the mean vector. For example, for each set of embedding vectors of the same type, calculate the arithmetic mean of the vector set to obtain the category prototype of that group, as shown in the following formula:
[0054] in, Represents the category prototype of the k-th group, such as the hollow corner-hollow drum, exemplarily a 128 or 256-dimensional vector; m represents the number of samples in the k-th group, exemplarily 5-20; This represents the general embedding vector of the i-th sample in the k-th group; S303, construct a category prototype library, integrate all category prototypes according to the special part dimension, and construct a special part category prototype library. For example, the inside corner prototype library contains four category prototypes: inside corner - normal, inside corner - hollow, inside corner - material splicing, and inside corner - structural boundary artifact; the window frame edge prototype library contains category prototypes with corresponding four labels; each prototype is labeled with part - label - vector dimension - number of samples, such as inside corner - hollow - 128 dimensions - 7 samples, which facilitates subsequent traceability and updates.
[0055] Optionally, to further optimize the accuracy of the category prototype, during adaptation, the last layer of the general embedding extractor is fine-tuned with a small learning rate and an early stopping strategy based on the support set through a small number of iterations.
[0056] Specifically, a light-tuning is performed on the last fully connected layer of the general embedding extractor without changing the parameters of the preceding convolutional layers and feature fusion layers, so as to avoid destroying the general feature extraction capability. The learning rate is set to 1e-5 to 1e-4, a small learning rate is used here to prevent overfitting, the batch size is set to 4 to 8, which is suitable for a small number of samples, and the number of training rounds is set to 5 to 10, a small number of iterations are used here to avoid overfitting; for example, the class discrimination of the embedding vectors of the support set samples is used as the indicator that the Euclidean distance between the same sample and the prototype is ≤0.4 and the distance between the different sample and the prototype is ≥0.8. If the indicator does not improve for two consecutive rounds, fine-tuning is stopped.
[0057] After fine-tuning, the general embedding extractor is called again to generate new general embedding vectors for the support set samples. The special part category prototype library is updated according to the mean vector calculation logic to ensure that the prototype matches the feature output of the fine-tuned extractor.
[0058] Optionally, the acoustic striking signal may be time-stretched or compressed, environmental noise may be added, an echo may be simulated or randomly cropped, the visual image may be subjected to viewpoint transformation, illumination change, partial occlusion or texture noise injection, and some modal signals may be randomly discarded during training.
[0059] Specifically, for time stretching or compression, for example, by changing the time axis length of the acoustic signal, the time domain distribution of the signal is adjusted, without changing the frequency characteristics, to simulate the differences in signal duration caused by the speed of tapping and the device sampling response delay in actual testing.
[0060] To incorporate environmental noise, for example, common environmental noises in building exterior wall detection scenarios, such as wind noise and urban traffic noise, can be injected into the original acoustic signal to improve the model's robustness to noise interference and avoid missed or false judgments due to environmental differences.
[0061] For simulating echoes, the acoustic environment of reflecting light from eaves or corners is simulated. For example, special locations such as under eaves or corners are prone to echoes due to structural obstruction, i.e., the original knocking sound and the sound reflected from the wall are superimposed. By adding delayed reflection signals, the acoustic characteristics of this scene are simulated, allowing the model to distinguish between echo interference and hollow signals.
[0062] For random cropping, the effective segment of the simulated local signal is extracted. For example, in actual detection, due to accidental device touch or sudden environmental noise, only a part of the effective acoustic signal can be obtained, such as only the attenuation segment of 0.5-1.5s after the impact is extracted. By randomly cropping the original signal, the model can learn to extract hollow features from the local segment.
[0063] For perspective changes, simulate shooting angle deviations. For example, in actual detection, due to limited operating space, such as needing to look up to shoot under the eaves, the camera perspective may deviate from the vertical direction. By adjusting the image perspective through affine transformation, this scenario can be simulated to avoid the model missing the detection of hollow areas due to a fixed perspective.
[0064] To address lighting changes, simulate lighting conditions at different times of day. For example, special areas such as shadowed corners may be in backlit areas, and the edges of window frames may produce strong shadows due to direct sunlight. By adjusting the image brightness and contrast, simulate different lighting conditions, allowing the model to ignore lighting interference and focus on hollow textures.
[0065] For partial occlusion, simulate the occlusion interference of brackets or window frames. For example, special parts such as the edge of the window frame may be partially occluded by the window frame or external wall brackets such as air conditioner outdoor unit brackets. By adding occlusion objects, simulate this scenario and let the model learn to identify hollow features such as cracks in the unoccluded area under partial occlusion.
[0066] For texture noise injection, the texture interference of water stains or dust is simulated. For example, there may be water stains or dust on the surface of the exterior wall, which will cause visual texture interference. By injecting random texture noise, this scenario is simulated, allowing the model to distinguish between noisy textures and hollow textures, such as the irregular cracks of hollow textures vs. the uniform coverage of dust.
[0067] In one embodiment, randomly discarding some modal signals during training can improve the model's modal complementarity. Specifically, during the training phase of the general embedding extractor, which includes regular large-sample training and fine-tuning phases with few samples for special parts, one or two modalities are randomly masked each time multimodal data is input, such as masking acoustics and retaining visual and vibration, forcing the model to adjust feature weights and extract hollow features based on effective modalities, thus avoiding over-reliance on a single modality.
[0068] For example, for the acoustic, vibration, and visual modalities, the independent discard probability of each modality is set to 0.2, meaning there is a 20% probability of masking a certain modality during each training session; at least one modality is retained to avoid all three modalities being discarded, leaving no features to extract; the visual modality is preferentially retained, which can be understood as the visual modality can directly present features such as seams and textures, and its stability is higher than that of acoustic or vibration modalities. Retaining the visual modality can reduce the risk of having no effective features at all; when masking a modality, the feature vector of that modality is set to an all-zero vector, rather than being directly deleted, to ensure that the feature dimension of the input general embedding extractor remains unchanged and to avoid model errors.
[0069] In some embodiments, for S104, the distance between the general embedding vector of the sample to be tested and the prototype vector of the corresponding part category is calculated using Euclidean distance or cosine distance to quantify the feature similarity between the two.
[0070] Specifically, before detection, visual recognition such as image segmentation can be used to determine the type of the area to be tested, and mark it as a high-risk area, such as a corner, under an eave, or the edge of a window frame, or a low-risk area, such as a regular wall surface. The corresponding category prototype library is then associated with the area. For example, if the area to be tested is the edge of a window frame, the window frame edge category prototype library is called.
[0071] Among them, for the category prototype library corresponding to low-risk parts, the characteristics of low-risk parts are regular structure and no occlusion or reflection interference, such as large flat wall surfaces on the exterior of buildings. A large amount of complete and distortion-free multimodal data can be collected, including acoustic, vibration, visual and metadata data. There is no need to rely on small sample adaptation. The category prototype library can be generated directly through conventional large sample training. The same prototype network as the previous embodiment can be used to directly utilize the general embedding vector of conventional wall surface large samples to generate the low-risk part category prototype library. For the specific generation process, please refer to the previous embodiment, which will not be repeated here.
[0072] For the prototype library of the part to which the test area belongs, such as the prototype library of window frame edge, the distance between the embedding vector of the test sample and four types of prototypes, namely window frame edge-normal, window frame edge-hollow, window frame edge-material splicing, and window frame edge-structural boundary artifact, is calculated to obtain four distance values.
[0073] Optionally, in some embodiments, to further improve the accuracy of detecting hollow areas, risk location information is introduced as a dynamic weighting factor into the detection and judgment logic. For example, when the sampling point is located in a high-risk location such as a corner, under an eave, or a window frame, even if the distance between the sample and the prototype is at a critical point, the model is more inclined to classify it as a hollow area after combining the dynamic weighting factor. When the sampling point is located in a low-risk location, the model may classify it as normal at the same prototype distance, avoiding misjudgment caused by over-detection.
[0074] Specifically, the aforementioned embodiment relies on the distance between the sample and the prototype for detection. When the distance is within the critical range between normal and hollow areas, the model is prone to random misjudgments. Detecting high-risk areas as normal leads to missed potential problems, while detecting low-risk areas as hollow leads to over-repair. The solution, after introducing a dynamic weighting factor, is as follows: High-risk areas include inside corners, under eaves, and window frames. A positive weighting is applied to samples at critical distances, increasing the weight for classifying them as hollow. For example, if a sample is at a distance *d* from the prototype of a hollow area, it falls within the critical value. After being weighted at high-risk locations, the model calculates a higher probability of hollowness, thus classifying it as hollow and avoiding overlooking potential hazards in high-risk areas.
[0075] For low-risk areas, negative weighting is applied to samples at the critical distance, thus reducing the weight for classifying them as hollow. For samples at the same distance d, the probability of hollowness decreases after weighting at low-risk locations, making the model more likely to classify them as normal, reducing false positives caused by over-detection, and lowering unnecessary detection and repair costs.
[0076] Specifically, a critical distance threshold is preset. As an example, taking the distance between the normal prototype and the hollow prototype as the benchmark, if the distance d_hollow_prototype between the test sample and the hollow prototype satisfies d_(normal prototype - hollow prototype)×0.4≤d_hollow_prototype≤d_(normal prototype - hollow prototype)×0.6, it is determined to be in the critical range. For example, if the distance between the normal prototype and the hollow prototype is 1.2, the critical range is 0.48-0.72.
[0077] For example, the dynamic weighting rule includes the following: for high-risk areas such as inside corners, under eaves, or window frame edges, if the sample d_hollow is in the critical range, a positive weighting factor such as 1.2 is applied, and the weighted distance d'_hollow = d_hollow × 0.8. Here, reverse weighting is applied to reduce the distance value and increase the similarity with the hollow prototype; if d'_hollow < the critical lower limit such as 0.48, it is determined to be hollow.
[0078] For low-risk areas such as regular walls, if the sample d_hollow is in the critical range, a negative weighting factor such as 0.8 is applied. The weighted distance d'_hollow = d_hollow × 1.2. Increasing the distance value here reduces the similarity to the hollow prototype. If d'_hollow > the critical upper limit such as 0.72, it is judged as normal.
[0079] For non-critical regions, direct judgment is made. For example, if the test sample d_void < the lower critical limit such as 0.48, it is directly judged as void regardless of whether it is a high-risk area; if the test sample d_void > the upper critical limit such as 0.72, it is directly judged as normal regardless of whether it is a low-risk area; if the distance between the test sample and the material splicing or structural boundary artifact prototype is the smallest, such as d_3=0.3<d_1, d_2, d_4, it is judged as a non-void interference item, and the void risk is excluded.
[0080] Figure 4 A building exterior wall hollowness data acquisition system 400 is shown. The system embodiment is similar to... Figure 1 Corresponding to the illustrated method embodiments, this system can be specifically applied to various electronic devices. Specifically, it includes: A general embedding extractor construction module 401 is used to collect multimodal detection signals on the conventional wall surface of a building exterior. The multimodal detection signals include acoustic impact signals, body acceleration signals, and visual image signals. A general embedding extractor is established based on the multimodal detection signals. The support set construction module 402 is used to collect a small number of labeled samples at special parts of the building exterior wall, and construct a support set containing multimodal raw data and metadata information based on the labeled samples. The special parts include inside corners, under eaves and window frame edges. The category prototype generation module 403 is used to adapt the support set to the general embedding vector output by the general embedding extractor and generate the category prototype corresponding to the special part by means of a few-shot learning model; wherein, the few-shot learning model adopts a prototype network, calculates the mean vector of each category in the embedding space as the category prototype, and performs classification based on Euclidean distance or cosine distance. The hollow area detection and judgment module 404 is used to determine whether the area corresponding to the test sample is a hollow area by measuring the distance between the category prototype and the test sample during the external wall inspection process.
[0081] Based on the same inventive concept, this application also provides an electronic device. The method corresponding to the electronic device can be the method in the foregoing embodiments, and its problem-solving principle is similar to that method. The electronic device provided in this application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the methods and / or technical solutions of the foregoing embodiments of this application.
[0082] Figure 5 The diagram illustrates the structure of an apparatus suitable for implementing the methods and / or technical solutions in the embodiments of this application. The apparatus 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 502 or a program loaded from a storage portion 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for system operation. The CPU 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0083] The following components are connected to I / O interface 505: input section 506 including keyboard, mouse, touch screen, microphone, infrared sensor, etc.; output section 507 including cathode ray tube (CRT), liquid crystal display (LCD), LED display, OLED display, etc., and speakers, etc.; storage section 508 including one or more computer-readable media such as hard disk, optical disk, magnetic disk, semiconductor memory, etc.; and communication section 509 including network interface card such as LAN (local area network) card, modem, etc. Communication section 509 performs communication processing via a network such as the Internet.
[0084] In particular, the methods and / or embodiments in this application can be implemented as computer software programs. For example, the embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. When the computer program is executed by the central processing unit (CPU) 501, it performs the functions defined in the methods of this application.
[0085] Another embodiment of this application provides a computer-readable storage medium having computer program instructions stored thereon, which can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of this application described above.
[0086] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application.
[0087] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device through software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any specific order.
Claims
1. A method for collecting data on hollow areas in building exterior walls, characterized in that, Includes the following steps: Multimodal detection signals are collected from the conventional wall surface of the building exterior. The multimodal detection signals include acoustic impact signals, body acceleration signals, and visual image signals. A universal embedding extractor is established based on the multimodal detection signals. The general embedding extractor employs a multimodal branch parallel architecture. The acoustic or vibration branch transforms the original signal into a time-frequency spectrum using short-time Fourier transform or wavelet transform, and then extracts acoustic or vibration features using a lightweight convolutional network. The visual texture branch extracts visual features using a lightweight visual backbone network. The metadata branch encodes the discrete or continuous attributes of the acquisition point category and device orientation to obtain metadata features. The acoustic or vibration features, visual features, and metadata features are concatenated, normalized, and mapped to a unified embedding space through a fully connected layer to obtain a general embedding vector. A small number of labeled samples were collected from special parts of the building's exterior walls. Based on these labeled samples, a support set containing multimodal raw data and metadata information was constructed. The special parts include inside corners, under eaves, and window frame edges. The inside corners refer to the angles between exterior wall surfaces and between the wall surface and the roof. Under eaves refers to the area under the eaves of the exterior wall that cantilever outwards. Window frame edges refer to the joint between the window frame and the exterior wall surface and the surrounding area. Based on the general embedding vector output by the general embedding extractor, the support set is adapted using a few-shot learning model to generate category prototypes corresponding to the special parts. The few-shot learning model employs a prototype network, calculating the mean vectors of each category in the embedding space as category prototypes, and classifying based on Euclidean distance or cosine distance. During the exterior wall detection process, the distance between the category prototypes and the test sample is used to determine whether the area corresponding to the test sample is a hollow area. When the sampling point of the sample to be tested is in a high-risk area, and the distance between the sample to be tested and the prototype of the hollow type is within the critical range between normal and hollow, a positive weighting is applied to the sample to be tested, increasing the weight of being judged as hollow and improving the probability of hollow calculated by the model; when the sampling point of the sample to be tested is in a low-risk area, and the distance between the sample to be tested and the prototype of the hollow type is within the critical range between normal and hollow, a negative weighting is applied to the sample to be tested, decreasing the weight of being judged as hollow and decreasing the probability of hollow calculated by the model; wherein the special area is a high-risk area, and the regular wall surface is a low-risk area.
2. The method for collecting data on hollow areas in building exterior walls according to claim 1, characterized in that, It also includes, When adapting the support set using a few-shot learning model, the last layer of the general embedding extractor is fine-tuned iteratively with a small learning rate and an early stopping strategy based on the support set.
3. The method for collecting data on hollow areas in building exterior walls according to claim 2, characterized in that, Also includes: After fine-tuning, the general embedding extractor is called again to generate new general embedding vectors for the support set samples. Based on the mean vector, the special part category prototype library is updated to ensure that the category prototypes match the feature output of the fine-tuned general embedding extractor.
4. The method for collecting data on hollow areas in building exterior walls according to claim 1, characterized in that, Also includes: The support set includes a small number of samples for each special part. Each sample contains at least one modal signal and a corresponding annotation result. The annotation result belongs to one of the following: normal, hollow, material splicing or seam, and structural boundary artifact.
5. The method for collecting data on hollow areas in building exterior walls according to claim 1, characterized in that, The acoustic impact signal is subjected to time stretching or compression, environmental noise is added, simulated echo or random cropping, the visual image signal is subjected to viewpoint transformation, illumination change, local occlusion or texture noise injection, and some modal signals are randomly discarded during training.
6. The method for collecting data on hollow areas in building exterior walls according to claim 1, characterized in that, During the training of the general embedding extractor, the weights of each modality feature are dynamically adjusted based on modal-level attention-weighted fusion. Specifically, based on the feature vector before fusion, the multilayer perceptron outputs three weight values, corresponding to acoustic or acceleration, visual, and metadata modalities, respectively, with the weights summed to 1. The feature vector of each modality is multiplied by the corresponding attention weight, and then summed to obtain the weighted fused feature vector, which is finally mapped to a unified embedding vector.
7. A data acquisition system for hollow building exterior walls, characterized in that, include: A general embedding extractor construction module is used to collect multimodal detection signals on the conventional wall surface of building exteriors. The multimodal detection signals include acoustic impact signals, body acceleration signals, and visual image signals. A general embedding extractor is established based on the multimodal detection signals. The general embedding extractor employs a multimodal branch parallel architecture. The acoustic or vibration branch transforms the original signal into a time-frequency spectrum using short-time Fourier transform or wavelet transform, and then extracts acoustic or vibration features using a lightweight convolutional network. The visual texture branch extracts visual features using a lightweight visual backbone network. The metadata branch encodes the discrete or continuous attributes of the acquisition point category and device orientation to obtain metadata features. The acoustic or vibration features, visual features, and metadata features are concatenated, normalized, and mapped to a unified embedding space through a fully connected layer to obtain a general embedding vector. The support set construction module is used to collect a small number of labeled samples at special locations on the exterior walls of buildings, and construct a support set containing multimodal raw data and metadata information based on the labeled samples. The special locations include inside corners, under eaves, and window frame edges. The inside corners refer to the angles between exterior wall surfaces and between the wall surface and the roof. Under eaves refers to the area under the eaves of the exterior wall that cantilever outwards. Window frame edges refer to the joint between the window frame and the wall surface and the surrounding area. The category prototype generation module is used to adapt the support set to the general embedding vector output by the general embedding extractor and generate the category prototype corresponding to the special part by means of a few-shot learning model; wherein, the few-shot learning model adopts a prototype network, calculates the mean vector of each category in the embedding space as the category prototype, and performs classification based on Euclidean distance or cosine distance. The hollow area detection and judgment module is used to determine whether the area corresponding to the test sample is a hollow area by measuring the distance between the category prototype and the test sample during the external wall inspection process. When the sampling point of the sample to be tested is in a high-risk area, and the distance between the sample to be tested and the prototype of the hollow type is within the critical range between normal and hollow, a positive weighting is applied to the sample to be tested, increasing the weight of being judged as hollow and improving the probability of hollow calculated by the model; when the sampling point of the sample to be tested is in a low-risk area, and the distance between the sample to be tested and the prototype of the hollow type is within the critical range between normal and hollow, a negative weighting is applied to the sample to be tested, decreasing the weight of being judged as hollow and decreasing the probability of hollow calculated by the model; wherein the special area is a high-risk area, and the regular wall surface is a low-risk area.
8. A computer-readable medium having computer program instructions stored thereon, characterized in that, The computer program instructions can be executed by a processor to implement the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Tunnel guniting intelligent detection method and system based on multi-modal information fusion
CN118585963A
Full-automatic wall surface acquisition device and intelligent detection method for wall surface hollowing
CN120629370A