House acceptance method based on deep learning and medium

Through deep learning-based house acceptance methods and knowledge graph technology, the problem of low house acceptance efficiency is solved, and an efficient and accurate house acceptance process is achieved.

CN120013477AInactive Publication Date: 2025-05-16ZHUHAI HUAZHANG ENG MANAGEMENT CONSULTING CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510103539.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

House acceptance efficiency is low, and there are problems such as a large amount of manpower investment, long acceptance cycle, and difficult to guarantee the accuracy and reliability of results.

Method used

The house acceptance method based on deep learning is adopted to fusion of multi-source heterogeneous data and construct a knowledge graph in the field of house acceptance, and use the knowledge graph to perform semantic interpretation and correlation analysis to generate an acceptance report.

Benefits of technology

It improves the efficiency of house acceptance, realizes automatic identification and quantification of key visual and geometric characteristics of the house, and enhances the accuracy and reliability of acceptance results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013477A_ABST
    Figure CN120013477A_ABST
Patent Text Reader

Abstract

The invention discloses a house acceptance method based on deep learning, and relates to the field of house acceptance, and the method comprises the steps: collecting the multi-modal detection data of a to-be-accepted house; extracting visual features of the RGB image by adopting a pre-trained convolutional neural network ResNet; a PointNet network is adopted to extract three-dimensional point cloud features of the depth image; an attention fusion mechanism is adopted to obtain fused multi-modal house features; collecting structured data of the house acceptance field to obtain structured knowledge representation of the house acceptance field; an entity is identified by adopting a BiLSTM-CRF model; entity attributes are identified through rule matching; building a house acceptance field knowledge graph by adopting a TransE knowledge graph embedding model, taking acceptance entities as nodes and taking entity attributes and relationships between entities as edges; and inputting the fused multi-modal house features into the constructed house acceptance field knowledge graph to generate a house acceptance report. Aiming at low house acceptance efficiency in the prior art, the acceptance efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of house acceptance, and in particular to a house acceptance method and medium based on deep learning. Background Art

[0002] With the rapid advancement of urbanization, the real estate industry is booming, and the scale of residential construction is expanding. In order to ensure the quality and safety of residential projects, house acceptance, as a key link before the house is delivered for use, occupies an important position in project management. Traditional house acceptance mainly relies on manual work. The acceptance personnel conduct on-site inspections and measurements of various indicators of the house according to the acceptance specifications to determine whether they meet the requirements. However, with the increasing complexity of residential projects, house acceptance faces many challenges.

[0003] On the one hand, the workload of house acceptance is large and covers a wide range. Modern residential projects are large in scale and complex in structure. The acceptance content covers multiple aspects such as structure, appearance, equipment installation, and decoration. The acceptance standards for different professions and different parts are different, and the acceptance work is cumbersome and time-consuming. The acceptance process requires a lot of manpower input, resulting in low acceptance efficiency and long acceptance cycle, which affects the delivery and use of houses.

[0004] On the other hand, housing acceptance is highly professional and requires a high level of professionalism from the acceptance personnel. The traditional acceptance process relies on the experience and quality of the acceptance personnel, and has problems such as strong subjectivity and inconsistent standards. The accuracy and reliability of the acceptance results are difficult to guarantee, and different acceptance personnel may have different acceptance conclusions on the same project, affecting the fairness of the acceptance. In addition, the acceptance specifications are updated quickly, and the acceptance personnel need to continue to learn new standards and specifications, which increases the workload of the acceptance personnel.

[0005] At the same time, housing acceptance data is complex and multi-source, lacking effective management and utilization methods. The housing acceptance process generates a large amount of heterogeneous data such as pictures, videos, and texts, which contain rich information on the project status. However, these data are scattered in different systems and carriers, lacking systematic organization and association, making it difficult to achieve data sharing and in-depth mining. The acceptance data lacks association with knowledge in the fields of acceptance specifications and acceptance standards, making it difficult to achieve digital and intelligent management of the entire acceptance process. Summary of the invention

[0006] In response to the problem of low efficiency of house acceptance in the prior art, the present application provides a house acceptance method and system based on deep learning. By fusing multi-source heterogeneous data and constructing a knowledge graph in the field of house acceptance, the knowledge graph is used to perform semantic interpretation and association analysis on house acceptance, and an acceptance report is generated, thereby improving the acceptance efficiency.

[0007] The purpose of this application is achieved through the following technical solutions.

[0008] One aspect of the present application provides a house acceptance method based on deep learning, comprising: collecting multimodal detection data of a house to be accepted, including RGB images and depth images; wherein the depth image refers to image data obtained by a sensor and containing information about the distance from the object surface to the sensor; and preprocessing the collected detection data; based on the preprocessed RGB image, using a pre-trained convolutional neural network ResNet to extract RGB image visual features; based on the preprocessed depth image, using a PointNet network to extract three-dimensional point cloud features of the depth image; using an attention fusion mechanism to perform feature fusion on the extracted RGB image visual features and the depth image three-dimensional point cloud features to obtain fused multimodal house features; S3, collecting house The structured data in the acceptance field is extracted by regular expression matching and template filling methods to obtain a structured knowledge representation in the field of house acceptance; named entity recognition is performed on the structured knowledge representation in the field of house acceptance, and the BiLSTM-CRF model is used to identify the entities in the text containing acceptance items, acceptance objects, acceptance standards and acceptance methods; and the entity attributes are identified through rule matching; the TransE knowledge graph embedding model is used to construct a knowledge graph in the field of house acceptance with acceptance entities as nodes and entity attributes and relationships between entities as edges; among them, the nodes include acceptance item nodes, acceptance standard nodes and acceptance method nodes; the fused multimodal house features are input into the constructed knowledge graph in the field of house acceptance to generate a house acceptance report.

[0009] Furthermore, the collected detection data is preprocessed, including: using a bilateral filtering algorithm to smooth and reduce the noise of the RGB image to remove high-frequency noise in the image; using a K-nearest neighbor interpolation algorithm to fill in the missing values ​​in the depth image; performing geometric alignment on the preprocessed RGB image and depth image, using an image alignment algorithm based on feature point matching to extract the ORB feature points of each image, determining the geometric transformation relationship between images of different modalities through feature point matching, and aligning the two modal images to the same coordinate system.

[0010] Furthermore, according to the preprocessed RGB image, a pretrained convolutional neural network ResNet is used to extract the visual features of the RGB image, including: scaling the preprocessed RGB image to a specified input size, the input size includes the spatial dimension and the number of color channels of the image; using a pretrained ResNet network to extract features of the RGB image, the ResNet network includes a convolution layer, a pooling layer and multiple residual blocks; the convolution layer uses multiple convolution kernels to perform a convolution operation on the input image, and outputs a feature map with a preset step size; wherein the convolution operation slides the convolution kernel on the input image to extract local features; the pooling layer uses a pooling window to perform a convolution operation on the convolution layer The output feature map is downsampled and the downsampled feature map is output with a preset step size; the output of the pooling layer passes through multiple residual blocks in sequence, each residual block consists of at least two convolutional layers and at least one identity mapping: the first convolutional layer uses a convolution kernel of size N1xN1 to reduce the dimension of the input feature map; the second convolutional layer uses a convolution kernel of size N2xN2 to convolve the reduced feature map to extract features; the identity mapping directly adds the input feature map to the output of the second convolutional layer to form a residual connection; a global pooling layer is added after the last convolutional layer of the ResNet network to convert the output feature map into RGB image visual features.

[0011] Furthermore, according to the preprocessed depth image, the PointNet network is used to extract the three-dimensional point cloud features of the depth image, including: converting the preprocessed depth image into three-dimensional point cloud data, each point cloud data point is represented by a three-dimensional coordinate (x, y, z), wherein x and y represent the position of the data point in the image plane, and z represents the distance of the data point from the depth sensor in the depth direction; inputting the converted three-dimensional point cloud data into the PointNet network for feature extraction, and the PointNet network is composed of a local feature extraction module, an attention pooling module and a feature fusion module connected in sequence: the local feature extraction module uses edge convolution operations to perform feature aggregation on the local neighborhood points of each point in the three-dimensional point cloud data, and by constructing a connection relationship between a point and its neighborhood points, convolving the three-dimensional coordinates of the point through a convolution kernel in the local neighborhood to generate a feature representation of the point. The attention pooling module uses the attention mechanism to perform weighted fusion on the generated local feature vectors. By learning the importance weight of each point, the local feature vectors of different points are weighted summed to generate a global feature vector that represents the global shape of the entire point cloud. The feature fusion module uses a multi-layer perceptron to perform nonlinear transformation and cascade on the generated local feature vectors and global feature vectors. Through the fully connected layer and nonlinear activation function, the local feature vectors and the global feature vectors are mapped to the same feature space, and feature splicing is performed to generate a point cloud semantic feature vector. The point cloud semantic feature vector is dynamically selected and aggregated through the attention pooling operation. By learning the weights of the feature dimensions, the semantic features of different dimensions are weighted summed to obtain the three-dimensional point cloud feature vector of the depth image of the specified dimension as the output of the PointNet network.

[0012] Furthermore, structured data in the field of house acceptance is collected, and information is extracted using a regular expression matching method, including: defining a regular expression template for the field to be extracted; wherein the regular expression template describes the text features of the field through a regular expression syntax, and the text features include keywords, numerical ranges, and length limits of the field; a finite state machine is used to compile the defined regular expression template, convert the regular expression template into a state transition graph, and generate a state transition rule library for matching text; wherein the nodes in the state transition graph represent states in the matching process, and the edges represent transition conditions between states; the collected structured data is traversed, and each piece of data is matched through a finite state machine using the generated state transition rule library; the finite state machine starts from an initial state, and performs state transitions according to the current characters and state transition rules until the termination state, and identifies field text fragments that meet the defined regular expression template.

[0013] Furthermore, structured data in the field of house acceptance is collected, and information is extracted using a template filling method to obtain a structured knowledge representation in the field of house acceptance, which also includes: for the identified field text fragments, a pre-trained named entity recognition model is used to perform entity recognition, and structured information containing the acceptance item name, value and unit in the field text fragment is extracted; the extracted structured information is filled and combined according to a predefined structured knowledge representation template to generate a structured domain knowledge representation that meets the template definition; wherein the structured domain knowledge representation template defines the organization form and attribute fields of knowledge, including the acceptance item ID, acceptance item name, value and unit.

[0014] Furthermore, named entity recognition is performed on the structured knowledge representation of house acceptance, including: using the pre-trained BERT model to encode the text in the structured knowledge representation at the character level, mapping each character into a dense vector of fixed dimension through the character embedding layer, extracting the context information and dependency relationship between characters through the multi-layer Transformer encoder, and generating a character-level context representation; extracting local features from the character-level context representation through the convolutional neural network CNN, obtaining character combination patterns of different sizes through multi-scale convolution kernels, and generating a character-level local feature representation; splicing the character-level context representation and the local feature representation on the character dimension. The character representation containing context and local features is cascaded using conditional random fields (CRFs) to perform sequence annotation on the character representation containing context and local features, and the named entities are identified. According to the identified named entities, a rule-based entity attribute recognition method is used to extract the attribute information of the entity. Among them, the attribute recognition method defines the trigger words, part-of-speech patterns and position ranges of the attributes, matches them in the entities and entity contexts, identifies the values ​​and units, and generates a structured entity-attribute representation.

[0015] Furthermore, conditional random fields (CRFs) are used to perform sequence labeling on character representations containing context and local features to identify named entities, including: defining a feature template of CRF, extracting observation features and transfer features between characters; wherein the observation features reflect the semantic information of the characters, including the characters themselves, part-of-speech tags, and character types; the transfer features reflect the dependencies and constraints between adjacent labels, including the transfer probability of labels; taking the character representation containing context and local features as input, and mapping it to the input feature sequence of CRF through the feature template of CRF; taking the extracted input feature sequence of CRF as input, constructing the log-linear model of CRF, and defining The conditional probability of the annotation sequence is the weighted sum of the feature function; the feature function maps the feature representation of the character to a real value, and the weight parameter represents the contribution of the feature to the annotation structure; based on the training data, by maximizing the log-likelihood function of the annotation sequence, the feature weight parameters of the CRF model are learned using the gradient descent algorithm, and the learned feature weight parameters are used as the parameters of the CRF model; in the inference stage, the character representation sequence to be annotated is used as input, and the Viterbi algorithm is used to decode it on the CRF model to obtain the optimal annotation path; at the end of the sequence, the global optimal annotation path of the entire sequence is obtained by backtracking as the result output of named entity recognition; Furthermore, the TransE knowledge graph embedding model is adopted, with acceptance entities as nodes, entity attributes and relationships between entities as edges, to construct a knowledge graph in the field of house acceptance, including: mapping acceptance entities to nodes in the knowledge graph, and mapping entities of different types to different types of nodes; wherein acceptance item entities are mapped to acceptance item nodes, acceptance object entities are mapped to acceptance object nodes, acceptance standard entities are mapped to acceptance standard nodes, and acceptance method entities are mapped to acceptance method nodes; mapping entity attributes to attribute edges in the knowledge graph, connecting entity nodes and attribute value nodes; wherein the starting node of the attribute edge is an entity node, the ending node is an attribute value node, and the label of the edge is the attribute name; constructing relationship edges between entity nodes based on the semantic relationship between acceptance entities; the starting node and the ending node of the relationship edge are both entity nodes, and the label of the edge is the relationship type; adopting the TransE knowledge graph embedding model to perform representation learning on the constructed knowledge graph, embedding the nodes and edges in the knowledge graph into a continuous low-dimensional vector space, and obtaining a knowledge graph in the field of house acceptance; Furthermore, the fused multimodal house features are input into the constructed knowledge graph in the field of house acceptance to generate a house acceptance report, including: standardizing the fused multimodal house features obtained by the attention fusion mechanism, and mapping the features of different modes to the same scale space; wherein, the standardization is achieved by normalizing the maximum and minimum values ​​of each feature dimension, and mapping the feature values ​​to the interval [0, 1]; inputting the standardized multimodal house features into the constructed knowledge graph in the field of house acceptance, matching the acceptance item nodes in the knowledge graph with the multimodal features, and realizing the association between the multimodal features and the acceptance items; wherein, the matching process is achieved by calculating the cosine similarity between the acceptance item node embedding vector and the multimodal feature vector, and selecting the acceptance item node with the highest similarity as the matching result; according to the matching result, extracting the acceptance standard node and the acceptance method node connected to the matched acceptance item node in the knowledge graph, and obtaining the corresponding acceptance standard. The multimodal house features are compared with the extracted acceptance criteria to determine whether the multimodal features meet the acceptance criteria; wherein, the comparison process determines whether the value of the multimodal features falls within the threshold range by setting the threshold range of the acceptance criteria; according to the comparison results, combined with the extracted acceptance method, the house acceptance results are generated; wherein, the house acceptance results include the acceptance item name, whether it is qualified and the acceptance description; the acceptance item name is obtained from the matching acceptance item node, whether it is qualified is determined according to the comparison results, and the acceptance description is generated according to the text description of the acceptance method node; the generated house acceptance results are summarized, and according to the predefined house acceptance report template, the acceptance item name, whether it is qualified and the acceptance description are filled in to generate a structured house acceptance report; wherein, the house acceptance report template defines the organizational structure and format of the acceptance report, including chapters such as the acceptance item summary table and the acceptance result details; the generated house acceptance report is output to complete the house acceptance process.

[0016] Another aspect of the present application also provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed by a processor, the method of the present application is implemented.

[0017] Compared with the prior art, the advantages of this application are: This application uses pre-trained deep convolutional neural networks ResNet and PointNet to extract visual features and 3D point cloud features from RGB images and depth images respectively. ResNet learns local and global visual features of images through convolution, pooling and residual blocks, and PointNet encodes local and global geometric patterns of point cloud data through edge convolution and attention pooling. The extracted multimodal features are organically combined through the attention fusion mechanism to obtain a house acceptance feature representation containing rich semantic information. This enables the automatic identification and quantification of key visual and geometric characteristics of the house during the acceptance process.

[0018] This application extracts information from structured data in the field of housing acceptance, and uses regular expression matching, named entity recognition, and rule attribute matching technologies to automatically extract key acceptance entities such as acceptance items, acceptance objects, acceptance standards, and acceptance methods, and their attributes, and generates structured acceptance knowledge representation using a template filling method. This structured domain knowledge representation provides a high-quality data foundation for the subsequent construction of an acceptance knowledge graph.

[0019] This application constructs a knowledge graph in the acceptance field with acceptance entities as nodes and entity attributes and relationships as edges, and uses the TransE knowledge graph embedding model to learn the low-dimensional vector representation of nodes and edges in the graph. The multimodal acceptance features of the house are input into the knowledge graph, and the structured acceptance knowledge in the graph is used to perform semantic interpretation and association analysis on the house features. The knowledge graph can reveal the semantic relationship between acceptance entities, support the symbolic representation of acceptance rules, and make acceptance decisions explainable. Through semantic reasoning on the knowledge graph, the degree of compliance of the acceptance data with the acceptance criteria can be analyzed, and the acceptance conclusion can be automatically generated. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The present application will be further described in the form of exemplary embodiments, which will be described in detail by the accompanying drawings. These embodiments are not restrictive, and in these embodiments, the same number represents the same structure, wherein: Figure 1 is an exemplary flow chart of a housing acceptance method based on deep learning as shown in the present application; Figure 2 is an exemplary flow chart of obtaining a preprocessed data set according to the present application; Figure 3 is an exemplary flow chart for constructing a structured domain knowledge representation according to the present application; Figure 4 is an exemplary flow chart for constructing a structured entity-attribute representation according to the present application. DETAILED DESCRIPTION

[0021] The method and system provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0022] Figure 1This is an exemplary flow chart of a deep learning-based house acceptance method shown in the present application, including: collecting multimodal detection data of the house to be accepted, including RGB images and depth images; wherein the depth image refers to image data obtained by a sensor and containing information about the distance from the object surface to the sensor; and preprocessing the collected detection data; based on the preprocessed RGB image, using a pre-trained convolutional neural network ResNet to extract RGB image visual features; based on the preprocessed depth image, using a PointNet network to extract three-dimensional point cloud features of the depth image; using an attention fusion mechanism to fuse the extracted RGB image visual features and the depth image three-dimensional point cloud features to obtain fused multimodal house features; collecting The structured data in the field of house acceptance is extracted by regular expression matching and template filling methods to obtain structured knowledge representation in the field of house acceptance; named entity recognition is performed on the structured knowledge representation in the field of house acceptance, and the BiLSTM-CRF model is used to identify the entities containing acceptance items, acceptance objects, acceptance standards and acceptance methods in the text; and the entity attributes are identified through rule matching; the TransE knowledge graph embedding model is used to construct a knowledge graph in the field of house acceptance with acceptance entities as nodes and entity attributes and relationships between entities as edges; the nodes include acceptance item nodes, acceptance standard nodes and acceptance method nodes; the fused multimodal house features are input into the constructed knowledge graph in the field of house acceptance to generate a house acceptance report Figure 2 According to the exemplary flowchart of obtaining the preprocessed data set shown in the present application, multimodal detection data of the house to be inspected is collected, including RGB images and depth images; wherein the depth image refers to image data obtained by the sensor containing information about the distance from the object surface to the sensor; and the collected detection data is preprocessed; the RGB image is smoothed and denoised by using a bilateral filtering algorithm to remove high-frequency noise in the image; specifically, the RGB image is smoothed and denoised by using a bilateral filtering algorithm to remove high-frequency noise in the image, including for each pixel point in the RGB image , define its neighborhood window Ω, the window size is , where k is a positive integer; for each pixel point q (m, n) in the neighborhood window Ω, calculate its spatial distance with the central pixel point p and color distance : ; ;in, and Represent the RGB color values ​​of pixel points p and q respectively; calculate the spatial weight coefficient and color weight coefficient : ; ;in, and are the standard deviations of spatial distance and color distance, respectively, which control the smoothness of the filter; calculate the bilateral filtering output value of pixel point p : , where Σ represents the sum of all pixels in the neighborhood window Ω; for all pixels of the RGB image, the above steps are performed to obtain a smoothed image after bilateral filtering.

[0023] The K nearest neighbor interpolation algorithm is used to fill in the missing values ​​in the depth image; in the process of acquiring the depth image, due to the influence of factors such as the measurement error of the sensor and the reflective characteristics of the surface of the object, there are often missing values ​​in the depth image, that is, the depth values ​​at some pixels are zero or undefined. For the needs of subsequent processing, it is necessary to fill and repair the missing values ​​in the depth image. In this embodiment, the K nearest neighbor interpolation algorithm is used to fill in the missing values ​​of the depth image. Specifically, let the depth image be D, the width and height of the image be W and H respectively, then D can be expressed as a W×H matrix. For any missing value pixel point p (x, y) in the depth image D, where x and y represent the horizontal and vertical coordinates of the pixel point in the image, respectively, the K nearest neighbors of pixel p are defined as the set of K non-missing value pixels closest to p, expressed as ,in is the i-th neighbor of p. The basic idea of ​​K-nearest neighbor interpolation is that for a missing pixel p, search for the K nearest non-missing pixels in its neighborhood, and then use the weighted average of the depth values ​​of these K neighboring pixels to estimate the depth value at p. Assume that the depth values ​​of the K nearest neighbor pixels of p are , then the estimated depth value at p It can be calculated by the following formula: ,in, is the i-th neighbor pixel The interpolation weight can be calculated based on The weight is calculated based on the spatial distance between p and p, color similarity and other factors. A commonly used weight calculation method is Gaussian weight, that is: , where σ is the standard deviation of the Gaussian kernel function, which controls the influence of neighborhood pixels on the interpolation result.

[0024] The preprocessed RGB image and depth image are geometrically registered. The image registration algorithm based on feature point matching is used to extract the ORB feature points of each image. The geometric transformation relationship between different modal images is determined by feature point matching, and the two modal images are registered to the same coordinate system. Extract the ORB feature points of the RGB image and the depth image. ORB feature points are extracted from the preprocessed RGB image and the depth image respectively. ORB feature points are a scale- and rotation-invariant feature point based on the oriented FAST and Rotated BRIEF algorithm. The FAST corner detection algorithm is used to detect corners in the image. By comparing the gray value difference between the pixel point and its surrounding neighborhood pixels, it is determined whether the point is a corner point. The grayscale centroid method is used to assign direction information to the detected corner point, and the grayscale centroid of the pixels in the corner point neighborhood is calculated. The main direction of the corner point is determined with the centroid as the center. The BRIEF feature descriptor with rotation invariance is extracted in the corner point area, and a binary feature vector is generated by binary comparison of pixel pairs in the corner point neighborhood. The corner point position, direction information and BRIEF feature descriptor are combined to form the final ORB feature point.

[0025] The geometric transformation relationship between the RGB image and the depth image is determined by matching ORB feature points. The scale and rotation invariance of ORB features are used to achieve robust matching between RGB images and depth images at different viewing angles and scales. The RANSAC algorithm based on random sampling consistency is used for ORB feature point matching. A certain number of ORB feature point pairs are randomly selected, and the geometric transformation relationship between the feature point pairs, including the rotation matrix and the translation vector, is estimated by the least squares method. All feature points are transformed using the estimated transformation relationship, and the error between the transformed feature points and the actual corresponding points is calculated. Feature point pairs with errors less than the set threshold are regarded as inliers, and the rest are outliers. The random sampling and transformation estimation process is iterated multiple times, and the transformation relationship with the largest number of inliers is selected as the optimal estimate. The optimal transformation relationship is used to perform least squares estimation on all inliers to solve the optimal geometric transformation matrix from the RGB image to the depth image. The optimal geometric transformation matrix contains two parts: the rotation matrix and the translation vector, which respectively characterize the rotation and displacement transformation relationship of the RGB image relative to the depth image.

[0026] According to the geometric transformation relationship, the coordinate system of the depth image is transformed to the coordinate system of the RGB image. The depth image is spatially transformed using the optimal geometric transformation matrix. A rotation matrix and a translation vector are applied to each pixel coordinate of the depth image to map it to the RGB image coordinate system. Assume that the pixel coordinates on the depth image are , the corresponding RGB image coordinates are , then the mapping relationship is: ; Where T is the optimal geometric transformation matrix, including the rotation matrix R and the translation vector t: ; Transform all pixel coordinates of the depth image to obtain the transformed depth image, and realize the mapping of the depth image to the RGB image coordinate system.

[0027] Achieve geometric alignment of the RGB image and the depth image in the same coordinate system. Through coordinate transformation, align the coordinate system of the depth image with the coordinate system of the RGB image. The aligned RGB image and the depth image completely overlap in the same coordinate system, and each RGB pixel corresponds to a unique depth value. For each pixel coordinate on the RGB image , the corresponding depth value d can be found on the transformed depth image. Through the coordinate correspondence, the RGB image and the depth image can be fused to generate an RGBD image, in which each pixel contains RGB color information and corresponding depth information.

[0028] According to the preprocessed RGB image, the pretrained convolutional neural network ResNet is used to extract the visual features of the RGB image; including: scaling the preprocessed RGB image to an input size of 224×224×3 as the input of the ResNet network; where 224×224 represents the spatial dimension of the image, and 3 represents the three RGB color channels; using the pretrained ResNet-50 network to extract features from the RGB image, the ResNet-50 network contains 1 convolution layer, 1 maximum pooling layer and 16 residual blocks: the convolution layer uses 64 3×3×3 convolution kernels to perform convolution operations on the input image, with a step size of 2, and the output size is a feature map of 112×112×64 ; The maximum pooling layer uses a 3×3 pooling window to downsample the convolution output with a step size of 2, and the output feature map is 56×56×64; the output of the maximum pooling layer passes through 16 residual blocks in sequence, each residual block consists of 2 convolutional layers and 1 identity mapping, and the residual connection method is used to alleviate the network degradation problem; the number of channels of the residual blocks is 64, 128, 256, and 512 respectively, and the residual blocks of each scale are repeated 4 times, gradually reducing the feature map size to 28×28, 14×14, and 7×7; a global average pooling layer is added after the last convolution layer of the ResNet-50 network to convert the 7×7×512 output feature map into a 512-dimensional RGB image visual feature vector.

[0029] After obtaining the preprocessed depth image, in order to extract the 3D geometric information contained in it, the depth image must first be converted into 3D point cloud data. Specifically, for each pixel point p (u, v) in the depth image, the coordinates (x, y, z) of the pixel point in 3D space can be calculated based on its coordinates (u, v) in the image and the corresponding depth value d. Assume that the width of the depth image is W, the height is H, and the focal length of the depth sensor is and , the principal point coordinates are , then the three-dimensional coordinates of pixel point p can be calculated by the following formula: ; ; z = d; By calculating the three-dimensional coordinates of all pixels in the depth image, a point cloud dataset containing N three-dimensional points can be obtained , where each point By three-dimensional coordinates express.

[0030] The converted 3D point cloud data P is input into the PointNet network for feature extraction. The PointNet network is a deep learning model specifically used to process 3D point cloud data, which can directly learn high-level semantic features from the original point cloud data. The PointNet network mainly consists of three modules: local feature extraction module, attention pooling module and feature fusion module. The purpose of the local feature extraction module is to extract the local neighborhood features of each point in the point cloud.

[0031] Specifically, the converted three-dimensional point cloud data P is input into the PointNet network. The three-dimensional point cloud data P converted in S12 is used as the input of the PointNet network. The point cloud data P is represented as an N×3 matrix, where N is the number of points in the point cloud and each row represents the three-dimensional coordinates (x, y, z) of a point. S2: Use the local feature extraction module of the PointNet network to extract the local neighborhood features of each point in the point cloud. For each point in the point cloud , find the k nearest neighbor points around it through the k nearest neighbor algorithm to form a local neighborhood set . Use edge convolution to aggregate features of points in the local neighborhood. The edge convolution operation steps are as follows: and its neighboring points , calculate the relative position vector between them . The coordinates and relative position vectors are spliced ​​into a 6-dimensional vector . Input the concatenated vector into a multi-layer perceptron network In the above example, we transform the feature vector . , the multi-layer perceptron network h_Θ is obtained through learning, which can encode the relative positions between points and extract local geometric features. All neighboring points of Perform feature transformation to obtain a set of feature vectors Perform the maximum pooling operation on all feature vectors to obtain the point The local eigenvector of . , the maximum pooling operation can extract the most significant features in the local neighborhood and enhance the robustness of the features. Repeat the above steps to extract the local feature vector for each point in the point cloud to obtain a set of local feature vectors .

[0032] The purpose of the attention pooling module is to perform weighted fusion of local feature vectors to generate a feature vector that represents the global shape of the entire point cloud. Specifically, the attention mechanism is used to learn the importance weight of each point and perform weighted summation of local feature vectors of different points. Let the attention weight of the i-th point be , then the global eigenvector g can be calculated by the following formula: ; Among them, the attention weight It can be adaptively calculated based on the local feature vector of the point through the softmax function: ; where w is a learnable parameter vector used to measure the importance of different points.

[0033] The purpose of the feature fusion module is to fuse the local feature vector with the global feature vector to generate the final point cloud semantic feature vector. Specifically, the local feature vector is fused using a multi-layer perceptron. And the global feature vector g is transformed and concatenated nonlinearly. First, the fully connected layer transforms and g are mapped to the same feature space: ; ; Then, the transformed local feature vector and the global feature vector are concatenated to obtain the point The semantic feature vector of : ; Finally, the semantic feature vectors of all points are dynamically selected and aggregated through the attention pooling operation. Similar to the attention pooling module, by learning the attention weights of the feature dimensions, the semantic features of different dimensions are weighted and summed to obtain the point cloud feature vector of the specified dimension: ;in, is the attention weight of the i-th point in the feature dimension, which can be adaptively calculated by the softmax function: ; where v is a learnable parameter vector used to measure the importance of different feature dimensions. Through the above steps, the PointNet network can be used to extract point cloud feature vectors representing three-dimensional geometric shapes from depth images, providing a basis for subsequent house acceptance feature fusion and acceptance report generation.

[0034] This implementation adopts the attention fusion mechanism to perform adaptive weighted fusion of the RGB image visual features and the depth image 3D point cloud features. Specifically, let the RGB image visual feature vector be , the feature vector of the 3D point cloud of the depth image is ,in and are the dimensions of the two features respectively. The purpose of the attention fusion mechanism is to learn an attention weight matrix , which is used to measure the correlation and importance between the features of the two modalities and perform weighted fusion of the features according to the attention weights.

[0035] First, the RGB image visual features v and the depth image 3D point cloud features s are transformed and dimensionally aligned through matrix multiplication: ; ;in, and are two learnable linear transformation matrices used to map v and s to the same feature space and unify the feature dimension to d.

[0036] Then, the attention mechanism is used to calculate the visual features of the RGB image and depth image 3D point cloud features The attention weight matrix between Specifically, the attention weight matrix is ​​calculated by the following formula: ,in, is a learnable query matrix that measures the correlation between the features of two modalities. and With the query matrix By performing matrix multiplication, we can obtain a d×d correlation matrix, which indicates the degree of correlation between the two modal features in different dimensions. Then, the correlation matrix is ​​nonlinearly transformed by the tanh activation function, and each row is normalized by the softmax function to obtain the final attention weight matrix. .

[0037] Finally, using the attention weight matrix Visual features of RGB images and depth image 3D point cloud features Perform weighted fusion to obtain the fused multimodal house feature vector f: , by and Performing matrix multiplication, we can get a The weighted feature matrix of represents the importance distribution of the RGB image visual features in different dimensions of the depth image 3D point cloud features. Then, the weighted feature matrix is ​​combined with Perform matrix multiplication to obtain the final d-dimensional multimodal house feature vector f. In the attention fusion mechanism, the attention weight matrix It can be trained and optimized together with the entire network in an end-to-end manner. During the training process, the network will adaptively learn the degree of correlation and importance between different modal features, and dynamically adjust the attention weights so that the fused features can better represent the multimodal information of the house.

[0038] Figure 3 It is an exemplary flow chart for constructing structured domain knowledge representation according to the application. When obtaining the domain knowledge required for the automatic generation of the house acceptance report, it is necessary to extract key information from the structured data in the house acceptance field and organize it into a structured knowledge representation form. This embodiment adopts a method of regular expression matching and template filling. By defining the text feature template of the field to be extracted, a finite state machine is used to perform text matching and information extraction, and the extracted structured information is organized according to the predefined knowledge representation template to obtain a structured knowledge representation of the house acceptance field. Specifically, it is necessary to define the regular expression template of the field to be extracted first. Regular expression is a formal language for describing text patterns. By defining a series of character rules and matching conditions, text fragments that meet specific patterns can be matched. In this embodiment, the regular expression template is used to describe the text features of the key fields that need to be extracted in the house acceptance report, such as the name of the acceptance item, the numerical range, the unit, etc. The regular expression template defines the text features of the field such as keywords, numerical ranges and length limits through regular expression syntax. For example, for the acceptance item name field, you can define a regular expression template that contains keywords such as "construction", "quality", and "acceptance", and limit the field length to within 10 characters.

[0039] After the regular expression template is defined, it needs to be compiled into a state transition rule that can be used for text matching, so that the text fragments that meet the template definition can be efficiently matched and identified in the subsequent information extraction process. This embodiment uses a finite state machine to compile the regular expression template and convert the regular expression template into the form of a state transition diagram. A finite state machine is a computing model consisting of a set of states, an input alphabet, a transition function, an initial state, and a set of acceptance states. The finite state machine can perform state transitions according to the rules defined by the transition function based on the current state and the input characters until the acceptance state is reached or the transition cannot continue. Finite state machines are often used for tasks such as string matching and lexical analysis. First, a state node is created for each matching rule of the regular expression template. For example, for the regular expression template "Construction.*Quality.*Acceptance" that matches the name of the acceptance item, three state nodes can be created, representing the matching of the three keywords "Construction", "Quality" and "Acceptance". Then, according to the matching rules in the regular expression template, state transition edges are added between the state nodes. The transition edge is marked with the matching condition that the current character needs to meet, that is, the transition rule. For example, between the two state nodes "Construction" and "Quality", a transfer edge marked with "*" can be added, indicating that after matching "Construction", any character can be matched until "Quality" is matched. Next, the starting position and the ending position of the regular expression template are represented by special state nodes, called the initial state and the accepting state, respectively. The initial state represents the starting point of the matching process, and the accepting state represents the end point of a successful match. Finally, according to the matching rules of the regular expression template, the state nodes and transfer edges are connected to form a complete state transition graph. Any path from the initial state to the accepting state in the graph corresponds to a matching result that satisfies the regular expression template.

[0040] After completing the compilation of the regular expression template and the generation of the state transfer rule library, the finite state machine can be used to perform text matching and field extraction on the structured data in the field of house acceptance. This implementation method traverses the collected structured data, matches and transfers the state of each data character by character, and identifies the field text fragments that meet the definition of the regular expression template. Traverse the structured data in the field of house acceptance and process each data. The structured data can be data organized in the form of tables, XML, JSON, etc., which contains various information on house acceptance, such as acceptance items, acceptance results, acceptance date, etc. For each structured data, a finite state machine is initialized using the generated state transfer rule library. The initial state of the finite state machine is the initial state node defined when compiling the regular expression template.

[0041] According to the current character and the transition conditions in the state transition rule base, the finite state machine performs state transition. The finite state machine searches for the transition conditions in the state transition rule base based on the current state and the input character. Assume that the finite state machine is currently in state s and the input character is c. The finite state machine searches for all transition rules in the state transition rule base with state s as the starting state. The form of the transition rule is (starting state, input character condition, target state). Specifically, the finite state machine searches for all transition edges in the current state and finds the transition edge that matches the current character. For each transition rule (s, condition, target) of state s, the finite state machine checks whether the input character c satisfies the condition. The condition can be a single character, a set of characters, a wildcard, etc., which is used to match the input character. If the input character c satisfies the condition of the transition rule, a matching transition edge is found.

[0042] If a matching transition edge is found, the finite state machine performs state transfer according to the transfer condition on the transition edge. If a matching transition edge (s, condition, target) is found, the finite state machine transfers the current state from s to the target state target. The finite state machine updates the current state to the target state target and prepares to process the next input character. S4.4: If there is no matching transition edge, it means that the current character does not meet the matching condition of the regular expression template, and the finite state machine stops matching. If no matching transition edge is found, it means that the current input character c does not meet any transfer conditions. This means that the current character does not meet the matching rules defined by the regular expression template, and the finite state machine cannot continue to transfer states. The finite state machine stops matching, indicating that the current matching process has failed and needs to be restarted or ended.

[0043] The finite state machine continuously reads the next character and performs state transfer according to the state transfer rules until it reaches the terminal state or cannot continue to transfer. The finite state machine continues to read the next character in the structured data and repeats the state transfer process. After completing the state transfer of the current character, the finite state machine reads the next character in the structured data. The read character is used as a new input and the state transfer process is repeated. During the transfer process, the finite state machine will record the state nodes passed and the matched character sequence. When the finite state machine performs state transfer, it will record the state nodes passed for subsequent matching result analysis. At the same time, the finite state machine will record the current matching character sequence as a potential field text fragment.

[0044] The finite state machine continuously performs state transitions until it reaches the terminal state or fails to find a matching transition edge. The finite state machine repeats execution, continuously reads the next character, and performs state transitions. If the finite state machine reaches the terminal state defined by the regular expression template, it indicates that the match is successful, and the matching character sequence can be extracted as a field text fragment. If the finite state machine cannot find a matching transition edge, it indicates that the match has failed, and the matching process needs to be restarted or ended.

[0045] If the finite state machine reaches the accepting state defined by the regular expression template, the matched character sequence is identified as the corresponding field text fragment. When the finite state machine reaches the accepting state defined by the regular expression template, it means that the currently matched character sequence meets the matching conditions of the regular expression template. During the state transition process, if the finite state machine reaches the accepting state defined in the regular expression template, it means that the currently matched character sequence meets the matching rules of the regular expression. The accepting state is a special state predefined in the regular expression template, which is used to indicate the termination condition of a successful match. When the finite state machine reaches the accepting state, it means that the currently matched character sequence is a complete text fragment that meets the matching conditions.

[0046] Identify the matched character sequence as the corresponding field text fragment. When the finite state machine reaches the accept state, identify the currently matched character sequence as the text fragment of the corresponding field. According to the definition of the regular expression template, the accept state corresponds to a specific field name or type. Associate the matched character sequence with the corresponding field to indicate that the character sequence is the value of the field. For example, if the finite state machine matches a text fragment such as "the qualified rate of construction quality acceptance is 95%" and reaches the accept state of the acceptance item name field, "construction quality acceptance" can be identified as the value of the acceptance item name field. Assume that a sub-expression for matching the acceptance item name is defined in the regular expression template, and its corresponding accept state is marked as the "acceptance item name" field. When the finite state machine matches the text fragment "the qualified rate of construction quality acceptance is 95%", it will reach the accept state of the "acceptance item name" field. At this time, the matched character sequence "construction quality acceptance" can be identified as the value of the "acceptance item name" field.

[0047] Extract the identified field text fragment and mark it according to the field name defined by the regular expression template. Extract the field text fragment identified in S6 as a field value of the structured data. The finite state machine identifies the field text fragment corresponding to the matching character sequence. Extract the text fragment as the value of the corresponding field in the structured data. Mark the extracted field value according to the field name defined by the regular expression template. In the regular expression template, each field has a corresponding field name or type. According to the definition of the regular expression template, associate the extracted field value with the corresponding field name. By marking the field name, the field to which each extracted field value belongs can be clearly identified. For example, "Construction Quality Acceptance" can be marked as the "Acceptance Item Name" field. In the example, "Construction Quality Acceptance" is identified as the value of the "Acceptance Item Name" field. Mark the extracted field value "Construction Quality Acceptance" and clearly specify that the field to which it belongs is "Acceptance Item Name". The marked result can be expressed as ("Acceptance Item Name" "Construction Quality Acceptance"), indicating that "Construction Quality Acceptance" is the value of the "Acceptance Item Name" field.

[0048] Continue to read the next character in the data, and repeat until all characters of the complete data are processed. The finite state machine continues to read the next character in the structured data and repeats. After completing the matching and field extraction of the current character sequence, the finite state machine continues to read the next character in the structured data. Use the read character as a new input and repeat the steps of state transfer (S4), recording the state and matching character sequence (S5), identifying field text fragments (S6), and marking field names (S7). S8.2: Until all characters of the current structured data are processed, the text matching and field extraction of the data are completed. The finite state machine repeats until all characters of the current structured data are read and processed. When all characters are processed, it means that the text matching and field extraction of the structured data have been completed. The extracted field values ​​and corresponding field names constitute the structured information of the structured data.

[0049] Repeat for each structured data in the house acceptance field until all data is processed. Traverse each structured data in the house acceptance field and repeat for each data. After completing the processing of the current structured data, continue to traverse the next structured data in the house acceptance field. For each structured data, repeat the steps of initializing the finite state machine (S2), reading characters (S3), state transition (S4), recording states and matching character sequences (S5), identifying field text fragments (S6), marking field names (S7) and processing the next character (S8).

[0050] Until all structured data is processed, text matching and field extraction of the entire housing acceptance field data are completed. Until all structured data in the housing acceptance field are traversed and processed. When all structured data are processed, it means that text matching and field extraction of the entire housing acceptance field data have been completed. The extracted field values ​​and corresponding field names constitute the structured information set of the housing acceptance field.

[0051] For the identified field text fragments, it is necessary to further extract the structured information contained therein, such as the acceptance item name, value and unit, etc. This embodiment uses a pre-trained named entity recognition model to perform entity recognition on the field text fragments. Named entity recognition is a basic task in the field of natural language processing, which aims to identify entities with specific meanings from text, such as names of people, places, and institutions. In the field of house acceptance, it is necessary to identify key entities such as acceptance item names, values, and units. The pre-trained named entity recognition model can automatically learn the feature representations of different entity categories by training on large-scale annotated data, and perform entity recognition on new text based on contextual information. By inputting the field text fragment into the pre-trained named entity recognition model, the structured information such as the acceptance item name, value, and unit contained therein can be automatically identified.

[0052] Finally, the extracted structured information needs to be organized according to the predefined structured knowledge representation template. The structured knowledge representation template defines the organizational form and attribute fields of the knowledge in the field of house acceptance, such as acceptance item ID, acceptance item name, value, unit, etc. In this embodiment, the extracted structured information is filled and combined according to the fields defined in the template to generate a structured knowledge representation that conforms to the template definition. For example, for the extracted acceptance item name "Construction Quality Acceptance", the value "95" and the unit "%", they can be filled into the corresponding fields in the structured knowledge representation template to generate a complete structured acceptance knowledge, such as "{Acceptance Item ID: 1, Acceptance Item Name: Construction Quality Acceptance, Value: 95, Unit: %}".

[0053] After obtaining the structured knowledge representation of the housing acceptance domain, it is necessary to further perform named entity recognition on it to identify key entities such as acceptance items, acceptance objects, acceptance criteria and acceptance methods contained in the text. This embodiment uses the BiLSTM-CRF model for named entity recognition, and identifies the attribute information of the entity by rule matching. Including: using the pre-trained BERT model to encode the text in the structured knowledge representation at the character level. BERT is a pre-trained language model based on the Transformer architecture. By performing self-supervised learning on large-scale unlabeled text data, a contextual representation rich in semantic information can be obtained. In this embodiment, a character embedding layer is used to map each character to a dense vector of a fixed dimension, and then the contextual information and dependencies between characters are extracted through a multi-layer Transformer encoder to generate a character-level contextual representation.

[0054] The local features of the context representation at the character level are extracted through the convolutional neural network (CNN). CNN slides on the character sequence through multi-scale convolution kernels to obtain character combination patterns of different sizes and generate local feature representations at the character level. This can capture the local dependencies and pattern information between characters and enrich the feature representation of characters. The character-level context representation and local feature representation are spliced ​​in the character dimension as the input of the bidirectional long short-term memory network BiLSTM. BiLSTM is a recurrent neural network that can perform forward and backward calculations on the character sequence at the same time to capture the bidirectional long-distance dependencies between characters. At each character position, BiLSTM cascades the forward and backward hidden states to generate a character representation that contains context and local features.

[0055] Conditional random field (CRF) is used to perform sequence annotation on character representations containing context and local features to identify named entities. Define the feature template of CRF to extract observation features and transfer features between characters. Observation features reflect the semantic information of characters and include the following aspects: Character itself: Extract the vector representation of the current character as the observation feature. Let , in Represents the i-th character. Part-of-speech tag: Tag the character sequence with parts of speech and extract the part-of-speech tag of the current character as the observation feature. , in Represents the part-of-speech tag of the i-th character. Character type: Extract character type features based on the attributes of the character (such as numbers, English, punctuation, etc.). Let , in represents the type of the i-th character. The transition feature reflects the dependency and constraints between adjacent labels, and mainly includes the transition probability of the label. , in represents the transition probability from label u to label v.

[0056] The character representation containing context and local features is taken as input and mapped to the input feature sequence of CRF through the feature template of CRF. , extract its observation features , and obtain the observed feature vector For adjacent characters and , extract its transfer features , and obtain the transfer eigenvalue. and the transfer eigenvalues Combined into the feature vector of the i-th position The feature vectors of all positions are concatenated in order to form the input feature sequence f(y, x) of CRF.

[0057] Construct a log-linear model of CRF and define the conditional probability of the labeled sequence as the weighted sum of feature functions. Mapped to the conditional probability of the labeled sequence, the form is: ; Where Z(x) is the normalization factor; n is the sequence length; K is the number of feature functions; is the kth characteristic function; is the weight parameter corresponding to the kth feature function. Feature function Map the feature representation of the character to a real value to represent the impact of the feature on the annotation result. Specifically include: Observation feature function , maps the observed feature vector to a real value. Transfer feature function , maps the transfer feature value to a real value. Weight parameter Characterization function The contribution to the annotation structure. The value of the weight parameter reflects the importance of the feature function to the annotation result. Based on the training data, the feature weight parameters of the CRF model are learned using the gradient descent algorithm by maximizing the log-likelihood function of the annotation sequence.

[0058] For the training set , maximize the log-likelihood function of the labeled sequence: ; Use the gradient descent algorithm to optimize the weight parameter λ and iteratively update the weight parameter: ; Where α is the learning rate, which controls the step size of each update. Repeat the iteration until the log-likelihood function converges or the maximum number of iterations is reached. The learned weight parameter λ is used as the parameter of the CRF model for sequence labeling and named entity recognition. CRF extracts observation features and transfer features between characters through feature templates, constructs a log-linear model, and learns model parameters by maximizing the log-likelihood function. This method makes full use of the semantic information of characters and the dependencies between adjacent labels, and can effectively perform named entity recognition. In the inference stage, the character sequence to be labeled is used as input, and the learned CRF model parameters are used for decoding to obtain the optimal labeling path, thereby identifying the named entity. Commonly used decoding algorithms include the Viterbi algorithm and the forward-backward algorithm.

[0059] In the inference stage, the Viterbi algorithm is used to decode the CRF model to obtain the specific implementation of the optimal annotation path. In the inference stage, given the character sequence to be annotated , the goal is to find the optimal annotation path , so that the conditional probability P(y|x) is maximized. The Viterbi algorithm is a dynamic programming algorithm that recursively calculates the optimal subpath score for each label at each position and finds the optimal labeling path for the entire sequence at the end of the sequence. Define state variables and recursion: Define state variables , indicating that a label is marked at position i The score of the optimal subpath of . Define the recursive formula: , where w represents the feature weight vector of the CRF model and f represents the feature function vector. The recursive formula describes the transition relationship between the current state and the previous state, that is, labeling the position i The optimal subpath score depends on all possible labels of the previous position i-1 The optimal subpath score and the characteristic function value of the current position.

[0060] Initialize the starting state of the Viterbi algorithm: Initialize the state variable at the starting position of the sequence to 0, that is, Initialize the state variables at the remaining positions to negative infinity, that is, , indicating an unreachable state. Recursively calculate the optimal subpath score for each label at each position: for each position i (i=1, 2, ..., n) in the sequence and each possible label , calculated by recursion During the calculation process, the optimal precursor state of each state is recorded , that is, label at position i The previous state of the optimal subpath. ; By recording the optimal predecessor state, the optimal annotation path of the entire sequence can be found in the subsequent backtracking process. At the end of the sequence, the end point of the optimal annotation path of the entire sequence is found based on the state variable at the final position: ; That is, at the last position n, find the state variable The label with the largest value , as the end point of the optimal annotation path. Starting from the end of the sequence, according to the optimal predecessor state of each position, reverse backtracking to find the optimal annotation path of the entire sequence: from the end point of the optimal annotation path Start by going through the optimal front-end state Backtrack and find Finally, the optimal annotation path for the entire sequence is obtained The optimal annotation path is used as the output result of named entity recognition. Using the Viterbi algorithm to decode on the CRF model, the optimal annotation path for a given character sequence can be obtained to achieve named entity recognition. The Viterbi algorithm uses the idea of ​​dynamic programming to find the optimal solution for the entire sequence at the end of the sequence by recursively calculating and recording the optimal predecessor state, avoiding the high complexity of exhausting all possible annotation paths.

[0061] Figure 4It is an exemplary flowchart for constructing a structured entity-attribute representation according to the present application. According to the identified named entity, a rule-based entity attribute recognition method is used to extract the attribute information of the entity. Define the trigger words of the attribute: For each type of attribute, define a set of trigger words, which represent the keywords or phrases of the attribute. For example, for the "length" attribute, the trigger words may include "long", "length", "length", etc. The trigger words may be manually defined or automatically extracted from the training data. Define the part-of-speech pattern of the attribute: For each type of attribute, define a set of part-of-speech patterns, which represent the part-of-speech combination of the attribute value. For example, for the "length" attribute, the part-of-speech pattern may be "numeral + quantifier" or "numeral + unit word". The part-of-speech pattern may be manually defined or automatically learned from the training data. Define the position range of the attribute: For each type of attribute, define the position range of the attribute value relative to the entity. For example, the attribute value may appear in front of, behind, or inside the entity. The position range may be manually defined or automatically learned from the training data. Match in the entity and entity context: For each identified named entity, search for the trigger words of the attribute in its context. If a trigger word is found, the location of the potential attribute value is determined based on the location of the trigger word and the location range of the attribute. The potential attribute value is tagged with parts of speech and matched with the part-of-speech pattern of the attribute. If the match is successful, the matched text is used as the attribute value. Identify the value and unit: For the matched attribute value, further identify the value and unit. You can use regular expressions or machine learning models to identify the value and unit. For example, for the "length" attribute, the value "10" and the unit "meter" can be identified. Generate a structured entity-attribute representation: Combine the identified attribute value and the corresponding value and unit into a structured entity-attribute representation.

[0062] The TransE knowledge graph embedding model is adopted to construct a knowledge graph in the field of house acceptance, with acceptance entities as nodes and entity attributes and relationships between entities as edges. Acceptance entities are mapped to nodes in the knowledge graph, and different types of entities are mapped to different types of nodes. Acceptance item entities are mapped to acceptance item nodes, such as "door and window acceptance", "water and electricity acceptance", etc. Acceptance object entities are mapped to acceptance object nodes, such as "door", "window", "wire", etc. Acceptance standard entities are mapped to acceptance standard nodes, such as "flatness", "sealing", "insulation", etc. Acceptance method entities are mapped to acceptance method nodes, such as "visual inspection", "measurement with ruler", "instrument testing", etc. Entity attributes are mapped to attribute edges in the knowledge graph, connecting entity nodes and attribute value nodes. The starting node of the attribute edge is the entity node, and the ending node is the attribute value node. The label of the edge is the attribute name, such as "material", "size", "color", etc. Examples: "(door, material, solid wood)" "(window, size, 1.5m x 1.2m)", etc. According to the semantic relationship between acceptance entities, the relationship edge between entity nodes is constructed. The starting node and the ending node of the relationship edge are both entity nodes. The label of the edge is the relationship type, such as "belongs to", "includes", "corresponds to", etc. Examples: "(Door and Window Acceptance, Includes, Door)" "(Door, Corresponds, Sealing)" and so on.

[0063] The TransE knowledge graph embedding model is used to learn the representation of the constructed knowledge graph, and the embedding vectors of all nodes and relationships in the knowledge graph are randomly initialized. The dimension of the node embedding vector is set to , the dimension of the relation embedding vector is . Create a shape of The node embedding matrix E has a shape of The relation embedding matrix R of is the number of nodes, is the number of relations. Use Xavier initialization method or Gaussian distribution initialization method to randomly initialize the element values ​​of matrices E and R.

[0064] Define the score function f(h, r, t) of the TransE model. The score function f(h, r, t) is used to evaluate the rationality of the triple (h, r, t). , that is, the Euclidean distance between the sum of the head entity embedding vector h and the relation embedding vector r and the tail entity embedding vector t. The smaller the score function value, the more reasonable the triple is, that is, the head entity h should be as close to the tail entity t as possible after being transformed by the relation r.

[0065] Define the loss function L of the TransE model: The definitions of the parameters are as follows: f(h, r, t) represents the score function of the TransE model, which is used to calculate the representation error of the triple (h, r, t). h, r, t represent the embedding vectors of the head entity, relation, and tail entity, respectively. and is a hyperparameter, which indicates the interval between positive and negative triplets. Introducing two different interval parameters can more flexibly control the distance between positive and negative samples. is the regularization coefficient, corresponding to the L2 regularization term of the head entity, relation, and tail entity embedding vectors. By setting different regularization strengths for different types of embedding vectors, the model's representation ability and overfitting risk can be better balanced. and is an additional regularization coefficient used to control the norm of the embedding vector. The corresponding upper bound constraint on the embedding vector norm is, The corresponding lower bound constraint on the norm of the embedding vector. is the upper threshold of the embedding vector norm, corresponding to the head entity, relationship and tail entity respectively. By setting a suitable upper threshold, the norm of the embedding vector can be prevented from being too large, thus improving the robustness of the model. is the lower bound threshold of the embedding vector norm, corresponding to the head entity, relation, and tail entity respectively. By setting an appropriate lower bound threshold, the norm of the embedding vector can be prevented from being too small, ensuring that the embedding vector has sufficient representation power. Two different hyperparameters are introduced for the intervals of positive and negative triples The flexibility of the model is improved. Different regularization coefficients are set for the embedding vectors of the head entity, relation, and tail entity, respectively, which can control the regularization strength of different types of embedding vectors in a more fine-grained manner. The upper and lower bound constraints of the embedding vector norm are introduced, and the strength of the constraint is controlled by μ1 and μ2. and Setting the threshold of the norm can effectively prevent the norm of the embedded vector from being too large or too small, thereby improving the robustness and representation ability of the model.

[0066] Mini-batch stochastic gradient descent algorithm is used to optimize the loss function L. Set the batch size to B, and randomly sample B positive triples from the knowledge graph as a batch. For each positive triple, construct the corresponding number of negative triples by replacing the head entity or the tail entity. Calculate the score function values ​​of the positive and negative triples respectively, and calculate the loss value of the current batch according to the loss function L defined in S543. Calculate the gradient of the loss function L to the node embedding matrix E and the relationship embedding matrix R through the back-propagation algorithm. Use the gradient descent algorithm to update the parameters of the matrices E and R: ; ; where α is the learning rate. Repeat the iteration until the preset number of training rounds is reached or the loss function value converges.

[0067] Get the optimized node embedding matrix E and relationship embedding matrix R, that is, the TransE knowledge graph embedding model. Each row of the matrix E represents the low-dimensional embedding vector of the corresponding node. Each row of the matrix R represents the low-dimensional embedding vector of the corresponding relationship. Calculate the semantic similarity between the acceptance entity node and the multimodal house features. Extract feature vectors from the house image through a pre-trained convolutional neural network (such as VGG, ResNet, etc.). Convert the house text into word vectors through a pre-trained word embedding model (such as Word2Vec, GloVe, etc.), and obtain the text feature vector through average pooling and other methods. Map the image feature vector and the text feature vector to the same embedding space as the TransE model through linear transformation or nonlinear mapping to obtain the embedded representation of the house features. Define the cosine similarity function as the similarity measurement function: ;in, and are two vectors. The cosine similarity between the embedding vector of each acceptance entity node and the embedding vector of the house feature is calculated to obtain a similarity score. The house features are associated with the acceptance entity according to the similarity score to obtain a house acceptance knowledge graph that integrates multimodal information.

[0068] The fused multimodal house features are input into the constructed knowledge graph of the house acceptance field to generate a house acceptance report, and the fused multimodal house features are standardized. The fused multimodal house features obtained by the attention fusion mechanism in S2 are standardized. The maximum and minimum normalization method is used to normalize each feature dimension separately, and the feature value is mapped to the interval [0, 1]. The normalization formula is: ; where x is the original feature value, min and max are the minimum and maximum values ​​of the feature of this dimension, respectively. is the normalized eigenvalue. After normalization, the features of different modalities will be mapped to the same scale space, which is convenient for subsequent feature matching and comparison.

[0069] Input the standardized multimodal house features into the knowledge graph of house acceptance to match the features with the acceptance items. Input the standardized multimodal house features into the constructed knowledge graph of house acceptance. The matching of multimodal features and acceptance items is achieved by calculating the cosine similarity between the embedding vector of the acceptance item node and the multimodal feature vector. The cosine similarity formula is: similarity=cos(θ)=(A·B) / (||A||*||B||); where A and B are two vectors, ||A|| and ||B|| represent the L2 norm of the vector. For each multimodal feature vector, calculate its cosine similarity with the embedding vectors of all acceptance item nodes, and select the acceptance item node with the highest similarity as the matching result.

[0070] Extract the acceptance criteria and acceptance methods connected to the matching acceptance item nodes in the knowledge graph. According to the matching results of S62, find the acceptance criteria nodes and acceptance method nodes connected to the matching acceptance item nodes in the knowledge graph. Extract the association information between the acceptance item nodes and the acceptance criteria nodes and the acceptance method nodes through the relationship edges in the knowledge graph. Obtain the specific acceptance criteria value or text description in the acceptance criteria node, and the acceptance method text description in the acceptance method node. Compare the multimodal house features with the extracted acceptance criteria to determine whether the acceptance criteria are met. According to the extracted acceptance criteria, set the corresponding threshold range. Compare the standardized multimodal house features with the threshold range of the acceptance criteria. Determine whether the value of the multimodal feature falls within the threshold range of the acceptance criteria to determine whether the feature meets the acceptance criteria. For numerical features, directly compare the feature value with the threshold range; for text features, comparison can be performed by keyword matching, semantic similarity and other methods.

[0071] Generate the house acceptance result based on the comparison results and the acceptance method. Generate the house acceptance result based on the comparison results and the extracted acceptance method. The house acceptance result contains the following content: Acceptance item name: Get the name of the acceptance item from the matching acceptance item node. Whether it is qualified: According to the comparison results, determine whether the acceptance item meets the acceptance criteria, so as to determine whether it is qualified. Acceptance description: According to the text description of the acceptance method node, generate a detailed description of the acceptance process and results of the acceptance item. For each matching acceptance item, generate the corresponding acceptance result to form a complete set of house acceptance results.

[0072] Summarize the acceptance results and generate a structured house acceptance report. Summarize and organize the generated house acceptance results. Organize and fill in the acceptance result information according to the predefined house acceptance report template. The house acceptance report template defines the structure and format of the acceptance report. Acceptance project summary table: List the names of all acceptance items, whether they are qualified, and other summary information in a table. Acceptance result details: According to the acceptance project classification, describe the acceptance criteria, acceptance methods, and acceptance results of each acceptance item in detail. Conclusion: Give the overall conclusion of the house acceptance based on the acceptance results to determine whether the house has passed the acceptance. Fill the generated acceptance results into the corresponding chapters according to the template requirements to generate a structured house acceptance report. Output the house acceptance report to complete the house acceptance process. Output and save the generated house acceptance report. The acceptance report can be generated in PDF, Word and other formats for subsequent viewing and delivery. The output house acceptance report is the final result of the house acceptance process for reference and archiving by relevant parties.

Claims

1. A house acceptance method based on deep learning, characterized in that: include: Collect multimodal detection data of the house to be inspected, including RGB images and depth images; where the depth image refers to image data obtained by the sensor containing the distance information from the object surface to the sensor; and pre-process the collected detection data; According to the preprocessed RGB image, the pre-trained convolutional neural network ResNet is used to extract the visual features of the RGB image; According to the preprocessed depth image, the PointNet network is used to extract the three-dimensional point cloud features of the depth image; The attention fusion mechanism is used to fuse the extracted RGB image visual features and the 3D point cloud features of the depth image to obtain the fused multi-modal house features. Collect structured data in the field of house acceptance, and use regular expression matching and template filling methods to extract information to obtain structured knowledge representation in the field of house acceptance; Perform named entity recognition on the structured knowledge representation of housing acceptance, and use the BiLSTM-CRF model to identify the entities containing acceptance items, acceptance objects, acceptance criteria, and acceptance methods in the text; and identify entity attributes through rule matching; The TransE knowledge graph embedding model is adopted to construct a knowledge graph in the field of house acceptance, with acceptance entities as nodes and entity attributes and relationships between entities as edges. The nodes include acceptance item nodes, acceptance standard nodes, and acceptance method nodes. The fused multimodal house features are input into the constructed house acceptance domain knowledge graph to generate a house acceptance report.

2. The house acceptance method based on deep learning according to claim 1, characterized in that: Preprocess the collected test data, including: The bilateral filtering algorithm is used to smooth and reduce noise of RGB images to remove high-frequency noise in the image. The K nearest neighbor interpolation algorithm is used to fill in the missing values ​​in the depth image; The preprocessed RGB image and depth image are geometrically registered. The image registration algorithm based on feature point matching is used to extract the ORB feature points of each image. The geometric transformation relationship between images of different modalities is determined by feature point matching, and the two modal images are registered to the same coordinate system.

3. The house acceptance method based on deep learning according to claim 2, characterized in that: The pre-trained convolutional neural network ResNet is used to extract the visual features of RGB images, including: Scale the preprocessed RGB image to the specified input size, which includes the spatial dimension and number of color channels of the image; A pre-trained ResNet network is used to extract features from RGB images. The ResNet network contains convolutional layers, pooling layers, and multiple residual blocks. The convolution layer uses multiple convolution kernels to perform convolution operations on the input image and outputs a feature map with a preset step size. The convolution operation slides the convolution kernel on the input image to extract local features. The pooling layer uses a pooling window to downsample the feature map output by the convolutional layer, and outputs the downsampled feature map with a preset step size; The output of the pooling layer is passed through multiple residual blocks in sequence. Each residual block consists of at least two convolutional layers and at least one identity mapping: The first convolutional layer uses a convolution kernel of size N1xN1 to reduce the dimension of the input feature map; The second convolutional layer uses a convolution kernel of size N2xN2 to convolve the reduced feature map to extract features; The identity mapping adds the input feature map directly to the output of the second convolutional layer, forming a residual connection; A global pooling layer is added after the last convolutional layer of the ResNet network to convert the output feature map into RGB image visual features.

4. The house acceptance method based on deep learning according to claim 3 is characterized in that: The PointNet network is used to extract the three-dimensional point cloud features of the depth image, including: The preprocessed depth image is converted into three-dimensional point cloud data, and each point cloud data point is represented by a three-dimensional coordinate (x, y, z), where x and y represent the position of the data point in the image plane, and z represents the distance of the data point from the depth sensor in the depth direction; The converted 3D point cloud data is input into the PointNet network for feature extraction. The PointNet network consists of a local feature extraction module, an attention pooling module, and a feature fusion module connected in sequence: The local feature extraction module uses edge convolution operation to aggregate features of local neighborhood points of each point in the 3D point cloud data. By building a connection relationship between a point and its neighborhood points, the 3D coordinates of the point are convolved by the convolution kernel in the local neighborhood to generate a local feature vector that represents the local geometric shape of the point. The attention pooling module uses the attention mechanism to perform weighted fusion on the generated local feature vectors. By learning the importance weight of each point, the local feature vectors of different points are weighted summed to generate a global feature vector that represents the global shape of the entire point cloud. The feature fusion module uses a multi-layer perceptron to perform nonlinear transformation and cascade on the generated local feature vectors and global feature vectors. Through a fully connected layer and a nonlinear activation function, the local feature vectors and the global feature vectors are mapped to the same feature space, and feature splicing is performed to generate a point cloud semantic feature vector. The point cloud semantic feature vector is dynamically selected and aggregated through the attention pooling operation. By learning the weights of the feature dimensions, the semantic features of different dimensions are weighted summed to obtain the three-dimensional point cloud feature vector of the depth image of the specified dimension as the output of the PointNet network.

5. The house acceptance method based on deep learning according to any one of claims 2 to 4, characterized in that: Collect structured data in the field of housing acceptance and use regular expression matching to extract information, including: Define a regular expression template for the field to be extracted; the regular expression template describes the text features of the field using regular expression syntax, and the text features include the keywords, value range, and length limit of the field; The defined regular expression template is compiled by using a finite state machine, and the regular expression template is converted into a state transition graph to generate a state transition rule base for matching text; wherein the nodes in the state transition graph represent the states in the matching process, and the edges represent the transition conditions between the states; Traverse the collected structured data, use the generated state transition rule library, and match each data through a finite state machine; the finite state machine starts from the initial state, and performs state transition according to the current characters and state transition rules until the terminal state, identifying the field text fragments that meet the defined regular expression template.

6. The housing inspection method based on deep learning according to claim 5, characterized in that: The structured knowledge representation of housing acceptance domain is obtained, including: For the identified field text fragments, a pre-trained named entity recognition model is used for entity recognition to extract the structured information including the acceptance item name, value and unit in the field text fragments; The extracted structured information is filled and combined according to the predefined structured knowledge representation template to generate a structured domain knowledge representation that meets the template definition; wherein the structured domain knowledge representation template defines the organizational form and attribute fields of the knowledge, including the acceptance item ID, acceptance item name, value and unit.

7. The house acceptance method based on deep learning according to claim 6, characterized in that: Perform named entity recognition on the structured knowledge representation of housing acceptance, including: Use the pre-trained BERT model to encode the text in the structured knowledge representation at the character level, map each character into a dense vector of fixed dimension through the character embedding layer, extract the context information and dependencies between characters through the multi-layer Transformer encoder, and generate character-level context representation; The local features of the context representation at the character level are extracted through the convolutional neural network (CNN), and the character combination patterns of different sizes are obtained through multi-scale convolution kernels to generate local feature representation at the character level. The context representation and local feature representation at the character level are concatenated in the character dimension as the input of the bidirectional long short-term memory network BiLSTM. BiLSTM performs forward and backward calculations on the character sequence at the same time, cascading the hidden states in both directions at each character position to generate a character representation that contains context and local features. Conditional random field (CRF) is used to perform sequence labeling on character representations containing context and local features to identify named entities. According to the identified named entities, a rule-based entity attribute recognition method is used to extract the attribute information of the entity. Among them, the attribute recognition method defines the trigger words, part-of-speech patterns and position range of the attributes, matches them in the entities and entity contexts, recognizes the values ​​and units, and generates a structured entity-attribute representation.

8. The housing inspection method based on deep learning according to claim 7, characterized in that: Conditional random field (CRF) is used to perform sequence annotation on character representations containing context and local features, and to identify named entities, including: Define the feature template of CRF to extract the observation features and transfer features between characters; the observation features reflect the semantic information of the characters, including the characters themselves, part-of-speech tags, and character types; the transfer features reflect the dependencies and constraints between adjacent labels, including the transfer probability of labels; The character representation containing context and local features is taken as input and mapped to the input feature sequence of CRF through the feature template of CRF; The extracted input feature sequence of CRF is used as input to construct the log-linear model of CRF, and the conditional probability of the annotation sequence is defined as the weighted sum of the feature function; wherein the feature function maps the feature representation of the character to a real value, and the weight parameter represents the contribution of the feature to the annotation structure; based on the training data, the feature weight parameters of the CRF model are learned by maximizing the log-likelihood function of the annotation sequence using the gradient descent algorithm, and the learned feature weight parameters are used as the parameters of the CRF model; In the inference stage, the character representation sequence to be annotated is taken as input, and the Viterbi algorithm is used to decode it on the CRF model to obtain the optimal annotation path; at the end of the sequence, the global optimal annotation path of the entire sequence is obtained by reverse backtracking as the result output of named entity recognition.

9. The housing inspection method based on deep learning according to claim 7, characterized in that: The TransE knowledge graph embedding model is used to construct a knowledge graph in the housing acceptance field, with acceptance entities as nodes and entity attributes and relationships between entities as edges, including: Map the acceptance entities to nodes in the knowledge graph, and map different types of entities to different types of nodes; among them, the acceptance item entity is mapped to the acceptance item node, the acceptance object entity is mapped to the acceptance object node, the acceptance standard entity is mapped to the acceptance standard node, and the acceptance method entity is mapped to the acceptance method node; Map entity attributes to attribute edges in the knowledge graph, connecting entity nodes and attribute value nodes; the starting node of the attribute edge is the entity node, the ending node is the attribute value node, and the label of the edge is the attribute name; According to the semantic relationship between the acceptance entities, the relationship edge between the entity nodes is constructed; the starting node and the ending node of the relationship edge are both entity nodes, and the label of the edge is the relationship type; The TransE knowledge graph embedding model is used to perform representation learning on the constructed knowledge graph, and the nodes and edges in the knowledge graph are embedded into a continuous low-dimensional vector space to obtain a knowledge graph in the field of house acceptance.

10. A computer-readable storage medium storing computer instructions, which implement the method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Cited By

  • Intelligent guest room checking system based on image recognition

    CN121033564A