Element fusion and accurate matching method and system based on multi-dimensional attribute classification
By classifying and processing real estate data and integrating factor, combining long-term and short-term memory networks and sorting learning models, the problem of difficulty in mining multi-dimensional attribute information in the existing technology is solved, and the refined classification, efficient integration and precise matching of real estate data is achieved, and the level of intelligent data processing and business application value is improved.
Patent Information
- Application Number
- CN202510084568.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
AI Technical Summary
It is difficult for existing technology to effectively explore multi-dimensional attribute information to achieve refined classification, efficient integration and accurate matching of real estate data, resulting in low level of intelligent data processing and business application value.
By classifying and fusion of the property description data set, an improved long-term and short-term memory network and sorting learning model are built, and a house rent evaluation result and accurate matching result are generated by combining historical rent matching data and user behavior data.
It realizes the mining of multi-dimensional attribute information of real estate data, improves the intelligence level of data processing and the value of business applications, and can more accurately match user needs.
Smart Images

Figure CN120013639A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multidimensional attribute classification, and in particular relates to a method and system for element fusion and precise matching based on multidimensional attribute classification. Background Art
[0003] There is an urgent need for a factor fusion and precise matching method based on multidimensional attribute classification, which can fully mine the multidimensional attribute information of the data, realize refined classification, efficient fusion and precise matching, thereby improving the intelligent level of data processing and the value of business applications. Summary of the invention
[0004] In order to solve the above technical problems, the present invention provides a method for factor fusion and precise matching based on multidimensional attribute classification, comprising:
[0005] Process the structured field feature approximate data of the real estate to obtain a real estate description data set;
[0006] After classifying the property description dataset, perform element fusion to obtain a fused dataset;
[0007] Constructing an improved long short-term memory network, training the long short-term memory network based on a historical rent matching data set, and obtaining a house rent evaluation model;
[0008] Inputting the fused data set into the house rent evaluation model to generate a house rent evaluation result;
[0009] A ranking learning model is constructed to obtain user behavior data, and the house rental evaluation result and the user behavior data are input into the ranking learning model for matching to obtain a house matching result.
[0010] Preferably, the process of processing the structured field feature approximate data of the real estate to obtain the real estate description data set includes:
[0011] Obtain metadata of structured field feature approximate data of the real estate, and construct a real estate data description based on the metadata;
[0012] Normalizing the real estate data description to obtain normalized data, and performing correlation analysis on the normalized data to obtain a correlation data set;
[0013] The associated data set is indexed in XML format and then stored to obtain the property description data set.
[0014] Preferably, the process of obtaining the fused data set includes:
[0015] Classifying the property description dataset based on multiple dimensional features to obtain a classified dataset; wherein the dimensional features include data content, data structure, and data semantics;
[0016] The interactive features of the classification data set are constructed, polynomial feature combination is performed on the interactive features to generate higher-order features, and then the discrete features are converted into dense vector representations, which are then combined with numerical features to generate the fused data set.
[0017] Preferably, converting the discrete features into a dense vector representation comprises:
[0018] Obtain the data set to be classified, and extract the content, structure and semantic feature vector of each data;
[0019] Based on the extracted data feature vectors, the K-means clustering algorithm is used to divide the data into several clusters. The data in each cluster has similarities in content, structure and semantics.
[0020] In each cluster, the Apriori association analysis algorithm is used to mine the association rules between data and obtain the frequently co-occurring data combination patterns;
[0021] According to the clustering results and association rules, the corresponding data classification labels are automatically generated to form a multi-dimensional and fine-grained data classification system;
[0022] A dense vector representation is constructed based on the multi-dimensional and fine-grained data classification system.
[0023] Preferably, the process of constructing an improved long short-term memory network includes:
[0024] Setting the range of the number of hidden layer nodes, learning rate and batch size of the initial long short-term memory network hyperparameters, training the initial long short-term memory network, optimizing the parameters based on the mean square error of the training results as the fitness function of the ant colony algorithm, and obtaining the optimal model parameters;
[0025] The optimal model parameters are introduced into the initial long short-term memory network to generate the long short-term memory network.
[0026] Preferably, the process of obtaining the house matching result includes:
[0027] The house rental evaluation result and user behavior are taken as a pair of candidate items, the candidate items are embedded in a low-dimensional space, and the low-dimensional data is input into the ranking learning model to generate a house recommendation list for each user. The recommendation list is sorted according to the user's behavior history and preferences, as well as the house rental evaluation results.
[0028] On the other hand, the present invention also provides an element fusion and precise matching system based on multi-dimensional attribute classification, comprising:
[0029] A data processing module is used to process the structured field feature approximate data of the real estate to obtain a real estate description data set;
[0030] An element fusion module is used to classify the property description dataset and then fuse the elements to obtain a fused dataset;
[0031] A training module, used for constructing an improved long short-term memory network, training the long short-term memory network based on a historical rent matching data set, and obtaining a house rent evaluation model;
[0032] An evaluation module, used for inputting the fused data set into the house rent evaluation model to generate a house rent evaluation result;
[0033] The matching module is used to build a ranking learning model, obtain user behavior data, input the house rental evaluation result and the user behavior data into the ranking learning model for matching, and obtain a house matching result.
[0034] On the other hand, the present invention also provides an electronic device, comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein the processor implements the method for element fusion and precise matching based on multidimensional attribute classification when executing the computer program.
[0035] On the other hand, the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the factor fusion and precise matching method based on multidimensional attribute classification.
[0036] Compared with the prior art, the present invention has the following advantages and technical effects:
[0037] The present invention can classify house features into different categories (such as geographic location, apartment type, price, etc.) by classifying and processing the property description data, which helps the model extract key information from structured and unstructured data and enhances the adaptability of the model to different types of properties. LSTM is particularly suitable for processing data with time series relationships, such as historical rental data. By using LSTM, the model can effectively capture the patterns, periodic fluctuations and potential trends of rent changes over time. By introducing ranking learning models (such as RankNet, LambdaMART, etc.) and combining user behavior data, the model can match the house rent evaluation results with the user's interests and behaviors (such as browsing, clicking, collecting, etc.), thereby generating a personalized house recommendation list. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0039] Figure 1 The figure is a schematic diagram of a method flow of an embodiment of the present invention. DETAILED DESCRIPTION
[0040] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0041] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0042] Embodiment 1
[0043] like Figure 1 As shown, this embodiment provides a method for element fusion and precise matching based on multi-dimensional attribute classification, including:
[0044] Process the structured field feature approximate data of the real estate to obtain a real estate description data set;
[0045] After classifying the property description dataset, perform element fusion to obtain a fused dataset;
[0046] Constructing an improved long short-term memory network, training the long short-term memory network based on a historical rent matching data set, and obtaining a house rent evaluation model;
[0047] Inputting the fused data set into the house rent evaluation model to generate a house rent evaluation result;
[0048] A ranking learning model is constructed to obtain user behavior data, and the house rental evaluation result and the user behavior data are input into the ranking learning model for matching to obtain a house matching result.
[0049] The "multi-dimensional attributes" of real estate and rent include but are not limited to the following aspects: Geographic location: cities, districts, streets, public facilities near real estate (schools, hospitals, shopping malls, transportation, etc.); basic information of the house: apartment type (area, number of bedrooms, number of bathrooms, etc.); floor, orientation, building type (high-rise, low-rise, villa, etc.); house status and decoration: the age of the house, decoration style (fine decoration, rough, etc.); whether there are supporting facilities such as furniture and electrical appliances; rent and payment method: monthly rent, deposit, rental payment method (monthly, quarterly, etc.); market dynamics and economic factors: surrounding housing price trends, rent changes, inflation rate and other market supply and demand conditions; landlord and tenant attributes: landlord and tenant credit scores; landlord's rental history, tenant's rental history;
[0050] To further optimize the solution, the process of processing the structured field feature approximate data of the real estate to obtain the real estate description data set includes:
[0051] Obtain metadata of structured field feature approximate data of the real estate, and construct a real estate data description based on the metadata;
[0052] Normalizing the real estate data description to obtain normalized data, and performing correlation analysis on the normalized data to obtain a correlation data set;
[0053] The associated data set is indexed in XML format and then stored to obtain the property description data set.
[0054] Based on the metadata information such as the format, standard, and attributes of the structured field characteristics, a unified data description model is constructed to standardize the representation and storage of data from different sources and eliminate data heterogeneity.
[0055] Obtain metadata information of structured field feature approximate data, including data format, data standard, and data attributes, and build a unified data description model based on the metadata information. For heterogeneous data from different sources, the constructed unified data description model is used to normalize them. The heterogeneous data after normalization is stored in a normalized manner to eliminate the heterogeneity of the data. Through data mining algorithms, association analysis is performed on the normalized stored data to mine the association rules and patterns between the data. Using machine learning algorithms, the normalized stored data is classified and clustered to discover the inherent structure and laws of the data. Based on the results of association analysis and machine learning, a semantic association network of the data is constructed to reveal the semantic relationship between the data. Integrate association rules, clustering results, and semantic association networks to form a unified understanding and application of structured field feature approximate data, and realize the effective integration and utilization of heterogeneous data.
[0056] When dealing with structured field feature approximation data, effective preprocessing and feature engineering are required for different types of data.
[0057] (1) Processing of time series data:
[0058] Normalization and standardization: Normalize or standardize time series data such as price data and trading volume to avoid the impact of data of different dimensions on model training.
[0059] Time series window division: Divide the time series data into time windows of fixed length according to business needs and model requirements. This helps capture the dependencies of historical information.
[0060] (2) Processing of non-time series data:
[0061] Standardization of numerical features: For numerical data such as regional population density and housing area, standardization (such as Z-score standardization) or normalization is performed.
[0062] Categorical data processing: Categorical data (such as region, property type, policy information, etc.) can be encoded using One-Hot Encoding or Embeddings.
[0063] (3) Data fusion:
[0064] Data alignment: Since the timestamps of data from different sources may be inconsistent, the time series data needs to be interpolated and aligned to ensure that the timestamps of all data are consistent for easy input into LSTM.
[0065] Multimodal data fusion: For multi-source data in the real estate field (such as price, features, policies, etc.), they can be fused through different data processing methods. For example, time series data and non-time series data can be merged and used as input to the LSTM model through feature splicing.
[0066] Further optimizing the scheme, the process of obtaining the fused data set includes:
[0067] Classifying the property description dataset based on multiple dimensional features to obtain a classified dataset; wherein the dimensional features include data content, data structure, and data semantics;
[0068] First, we can select the potential features according to different property dimensions (such as geographic information, house characteristics, market conditions, etc.). We can use correlation analysis, LASSO regression and other methods to select features. We can combine numerical features such as region and house area with categorical features such as house type and process them through One-Hot encoding or embedding vectors.
[0069] The interactive features of the classification data set are constructed, polynomial feature combination is performed on the interactive features to generate higher-order features, and then the discrete features are converted into dense vector representations, which are then combined with numerical features to generate the fused data set.
[0070] Clean up duplicate records in the data and fill or remove missing values.
[0071] Standardize or normalize numerical features (such as area, rent, etc.) to ensure that the data is on the same scale and avoid large numerical differences that affect neural network training.
[0072] For categorical variables (such as decoration style, geographic location, etc.), one-hot encoding or label encoding can be used for conversion.
[0073] For text data such as house descriptions, natural language processing (NLP) techniques such as bag-of-words, TF-IDF, Word2Vec, or BERT can be used for vectorization.
[0074] The one-hot encoded category features can be concatenated, or an embedding layer can be used to embed the category features into a low-dimensional vector to further reduce the dimension and capture the potential relationship between categories.
[0075] For multidimensional data, feature fusion can be performed by weighted summation. For example, by combining certain features of the geographic location (such as transportation convenience) with the features of the house (such as area and number of rooms), different weights can be set to determine the influence of each feature based on historical data or domain knowledge.
[0076] For numerical features, higher-order features can be generated through polynomial feature combination to capture nonlinear relationships. For example, house prices are not only related to area, but may have a stronger correlation with the square of the area (i.e., the quadratic term of the area). This combination can be automatically generated through PolynomialFeatures.
[0077] Further optimizing the scheme, the conversion of discrete features into dense vector representation includes:
[0078] Obtain the data set to be classified, and extract the content, structure and semantic feature vector of each data;
[0079] Based on the extracted data feature vectors, the K-means clustering algorithm is used to divide the data into several clusters. The data in each cluster has similarities in content, structure and semantics.
[0080] In each cluster, the Apriori association analysis algorithm is used to mine the association rules between data and obtain the frequently co-occurring data combination patterns;
[0081] According to the clustering results and association rules, the corresponding data classification labels are automatically generated to form a multi-dimensional and fine-grained data classification system;
[0082] A dense vector representation is constructed based on the multi-dimensional and fine-grained data classification system.
[0083] According to the pre-established unified data description model, the data set to be classified is obtained, and the content, structure and semantic feature vectors of each data are extracted. According to the extracted data feature vectors, the K-means clustering algorithm is used to divide the data into several clusters. The data in each cluster has a high degree of similarity in content, structure and semantics. In each cluster, the Apriori association analysis algorithm is used to mine the association rules between the data to obtain the frequently co-occurring data combination pattern. According to the clustering results and association rules, the corresponding data classification labels are automatically generated to form a multi-dimensional and fine-grained data classification system. For the newly added data, its feature vector is extracted, and its classification is determined by calculating the similarity with the existing cluster center, and the corresponding association rules are updated. According to the user's data retrieval needs, the relevant data subsets can be quickly located through the combination and screening of classification labels to improve the efficiency of data search and access. The clustering results and association rules are regularly evaluated, and the classification labels and rule thresholds are dynamically adjusted according to the changes in the data to maintain the accuracy and timeliness of the classification system.
[0084] Design unified data fusion rules and strategies for classified similar data, and realize automatic alignment and fusion of data from different sources through data cleaning, conversion, association and other operations to form a consistent data view. Data fusion rules include attribute mapping rules, conflict resolution rules and data merging rules, which are used to handle attribute differences, data conflicts and data merging between data from different sources.
[0085] Preprocess the fused high-quality data, extract key fields and attribute information, and build an attribute index; use natural language processing technology to perform word segmentation and semantic analysis on text data, extract keywords and subject information, and build a full-text index; for spatial data, use spatial partitioning and encoding technology to build a spatial index to support spatial range queries; select appropriate index structures such as B+ trees, hash tables, inverted indexes, etc. according to data characteristics and query patterns to balance query efficiency and storage overhead; design a multi-dimensional combined query algorithm, dynamically generate query plans based on multiple query conditions entered by the user, and use indexes to quickly locate candidate results; sort and filter candidate results by relevance, combine user preferences and query context, and optimize the quality and diversity of query results; implement an incremental update mechanism to dynamically update the index when data is updated to ensure the synchronization and consistency of the index and data and improve the real-time performance of queries.
[0086] Further optimizing the scheme, the process of constructing the improved long short-term memory network includes:
[0087] Setting the range of the number of hidden layer nodes, learning rate and batch size of the initial long short-term memory network hyperparameters, training the initial long short-term memory network, optimizing the parameters based on the mean square error of the training results as the fitness function of the ant colony algorithm, and obtaining the optimal model parameters;
[0088] The optimal model parameters are introduced into the initial long short-term memory network to generate the long short-term memory network.
[0089] The probability of an ant choosing a path is affected by the pheromone concentration and the heuristic information of the path (such as distance, cost, etc.). The probability of choosing path (i, j) can be expressed by the following formula:
[0090]
[0091] P ij (t) is the probability of the ant choosing from node i to node j at time t. τ ij (t) is the pheromone concentration on path (i, j). η ij is the heuristic information of the path, usually the inverse length of the path (i.e., η ij =1 / d ij η, where d ij is the distance from node i to node j). α is the influence factor of pheromone, which controls the weight of pheromone. β is the influence factor of heuristic information, which controls the weight of heuristic information. i is the set of neighbor nodes of node i.
[0092] Ranking learning relies on labeled datasets. Typically, the data contains a query and a set of candidate items (such as web pages in search results, products in recommendation systems, etc.), and the relevance of each set of candidate items has been annotated.
[0093] Data labels are divided into the following forms:
[0094] Pairwise Labels: For each pair of candidates, label their relative order (e.g., candidate A is more relevant than candidate B).
[0095] Listwise Labels: Score or sort the entire list of candidates and mark the priority order of the entire list.
[0096] The selection of the ranking learning model can be based on the specific task requirements. Common ranking learning methods include:
[0097] Pairwise Ranking: Such as RankNet, SVMRank, etc. This type of method is trained by comparing the relative order of two instances.
[0098] Listwise Ranking: Such as LambdaRank, LambdaMART, etc. These models optimize the entire ranking list and are usually used in information retrieval tasks.
[0099] Learn variants of sorting models (such as neural network models): For example, use neural network models (such as DeepRank, DeepMatching) for sorting tasks.
[0100] Use historical data to train the sorting model, and the model will learn how to sort based on the relationship between house features and user behavior data. Appropriate loss functions (such as logarithmic loss, sorting loss function, etc.) can be used during training.
[0101] To further optimize the solution, the process of obtaining the house matching result includes:
[0102] The house rental evaluation result and user behavior are taken as a pair of candidate items, the candidate items are embedded in a low-dimensional space, and the low-dimensional data is input into the ranking learning model to generate a house recommendation list for each user. The recommendation list is sorted according to the user's behavior history and preferences, as well as the house rental evaluation results.
[0103] In actual applications, the trained ranking model is applied to the real-time input of user behavior data and house feature data to generate a house recommendation list for each user. The recommendation list is sorted according to the user's behavior history and preferences, as well as the house's rental evaluation results.
[0104] Based on the output of the model, the final matching house is given. The recommendation system will prioritize the eligible houses according to the user's preferences, budget, location and other factors and display them to the user.
[0105] As users interact with the recommended results (such as clicking, collecting, renting, etc.), new behavior data can be continuously collected as feedback for model training. The model can be updated regularly to improve the recommendation effect.
[0106] The recommendation model is adjusted and optimized in real time based on changes in user behavior or market changes (such as rental fluctuations, changes in housing supply, etc.).
[0107] On the other hand, this embodiment also provides a factor fusion and precise matching system based on multi-dimensional attribute classification, including:
[0108] A data processing module is used to process the structured field feature approximate data of the real estate to obtain a real estate description data set;
[0109] An element fusion module is used to classify the property description dataset and then fuse the elements to obtain a fused dataset;
[0110] A training module, used for constructing an improved long short-term memory network, training the long short-term memory network based on a historical rent matching data set, and obtaining a house rent evaluation model;
[0111] An evaluation module, used for inputting the fused data set into the house rent evaluation model to generate a house rent evaluation result;
[0112] The matching module is used to build a ranking learning model, obtain user behavior data, input the house rental evaluation result and the user behavior data into the ranking learning model for matching, and obtain a house matching result.
[0113] On the other hand, this embodiment also provides an electronic device, including a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein the processor implements the factor fusion and precise matching method based on multidimensional attribute classification when executing the computer program.
[0114] On the other hand, this embodiment further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the factor fusion and precise matching method based on multi-dimensional attribute classification.
[0115] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A method for factor fusion and precise matching based on multidimensional attribute classification, characterized in that: include: Process the structured field feature approximate data of the real estate to obtain a real estate description data set; After classifying the property description dataset, perform element fusion to obtain a fused dataset; Constructing an improved long short-term memory network, training the long short-term memory network based on a historical rent matching data set, and obtaining a house rent evaluation model; Inputting the fused data set into the house rent evaluation model to generate a house rent evaluation result; A ranking learning model is constructed to obtain user behavior data, and the house rental evaluation result and the user behavior data are input into the ranking learning model for matching to obtain a house matching result.
2. The method according to claim 1, characterized in that The process of processing the structured field feature approximate data of the real estate to obtain the real estate description data set includes: Obtain metadata of structured field feature approximate data of real estate, and construct real estate data description based on the metadata; Normalizing the real estate data description to obtain normalized data, and performing correlation analysis on the normalized data to obtain a correlation data set; The associated data set is indexed in XML format and then stored to obtain the property description data set.
3. The method according to claim 1, characterized in that The process of obtaining the fused data set includes: Classifying the property description dataset based on multiple dimensional features to obtain a classified dataset; wherein the dimensional features include data content, data structure, and data semantics; The interactive features of the classification data set are constructed, polynomial feature combination is performed on the interactive features to generate higher-order features, and then the discrete features are converted into dense vector representations, which are then combined with numerical features to generate the fused data set.
4. The method according to claim 3, characterized in that The step of converting discrete features into dense vector representations includes: Obtain the data set to be classified, and extract the content, structure and semantic feature vector of each data; Based on the extracted data feature vectors, the K-means clustering algorithm is used to divide the data into several clusters. The data in each cluster has similarities in content, structure and semantics. In each cluster, the Apriori association analysis algorithm is used to mine the association rules between data and obtain the frequently co-occurring data combination patterns; According to the clustering results and association rules, the corresponding data classification labels are automatically generated to form a multi-dimensional and fine-grained data classification system; A dense vector representation is constructed based on the multi-dimensional and fine-grained data classification system.
5. The method according to claim 1, characterized in that The process of constructing the improved long short-term memory network includes: Setting the range of the number of hidden layer nodes, learning rate and batch size of the initial long short-term memory network hyperparameters, training the initial long short-term memory network, optimizing the parameters based on the mean square error of the training results as the fitness function of the ant colony algorithm, and obtaining the optimal model parameters; The optimal model parameters are introduced into the initial long short-term memory network to generate the long short-term memory network.
6. The method according to claim 1, characterized in that The process of obtaining the house matching result includes: The house rental evaluation result and user behavior are taken as a pair of candidate items, the candidate items are embedded in a low-dimensional space, and the low-dimensional data is input into the ranking learning model to generate a house recommendation list for each user. The recommendation list is sorted according to the user's behavior history and preferences, as well as the house rental evaluation results.
7. A factor fusion and precise matching system based on multi-dimensional attribute classification, characterized in that: include: A data processing module is used to process the structured field feature approximate data of the real estate to obtain a real estate description data set; An element fusion module is used to classify the property description dataset and then fuse the elements to obtain a fused dataset; A training module, used for constructing an improved long short-term memory network, training the long short-term memory network based on a historical rent matching data set, and obtaining a house rent evaluation model; An evaluation module, used for inputting the fused data set into the house rent evaluation model to generate a house rent evaluation result; The matching module is used to build a ranking learning model, obtain user behavior data, input the house rental evaluation result and the user behavior data into the ranking learning model for matching, and obtain a house matching result.
8. An electronic device comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method described in any one of claims 1 to 7 is implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.