A house source processing method based on user portrait and deep reinforcement learning

CN122817541APending Publication Date: 2026-09-25HUBEI PROVINCIAL HOUSING SECURITY CONSTRUCTION MANAGEMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610709295.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-21
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

现有方法在处理该类数据时,多采用简单拼接或独立建模的方式,缺乏对多源特征之间内在关联性的统一刻画机制,导致不同类型房源特征之间的信息无法有效融合,进而难以形成一致的高维表达空间

Benefits of technology

本发明通过对多源用户数据构建分层用户画像体系以形成具备稳定约束、动态偏好与潜在需求统一表达的用户特征向量,并对多源房源数据执行统一特征提取与增强处理以形成高维一致表达的房源特征向量,从而在同一特征空间中建立用户特征与房源特征之间的深层语义对应关系;进一步引入深度强化学习对房源匹配推荐模型执行联合建模,使匹配度量结果不仅依赖静态特征相似性,还能够结合用户实时交互反馈对匹配策略进行动态优化,同时通过多源特征融合与非线性交互建模有效刻画复杂关联关系,避免浅层匹配带来的信息损失,从而提升用户特征与房源特征匹配度分析的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817541A_ABST
    Figure CN122817541A_ABST
Patent Text Reader

Abstract

The application provides a house source processing method based on user portrait and deep reinforcement learning, and relates to the technical field of data processing.The method comprises the following steps: obtaining multi-source user data corresponding to a target user and performing preprocessing to construct a user portrait system, thereby generating a user feature vector; obtaining multi-source house source data through the user feature vector; performing feature extraction and enhancement processing on the multi-source house source data to generate a house source feature vector; constructing a house source matching recommendation model according to the user feature vector and the house source feature vector, and performing joint modeling through deep reinforcement learning to output a matching degree measurement result between the user feature vector and the house source feature vector; and generating a house source recommendation sequence based on the matching degree measurement result.The application can improve the accuracy of user feature and house source feature matching degree analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and specifically to a method for processing housing information based on user profiling and deep reinforcement learning. Background Technology

[0002] With the development of internet-based real estate information platforms, property recommendation has gradually become an important technical means to improve user decision-making efficiency and platform service capabilities. Existing technologies typically involve constructing user profiles and combining them with basic property attribute information, then using collaborative filtering, content-based recommendation methods, or shallow machine learning models to recommend properties to users. These methods generally analyze users' historical browsing behavior, click behavior, and basic attribute information to form user interest characteristics, and then match these characteristics with structured features such as property price, area, and location to output recommendation results. Some improved solutions also introduce deep learning models to model user behavior sequences to improve recommendation accuracy; however, overall, these methods still rely heavily on static modeling or offline training, and the recommendation strategy lacks dynamic feedback adjustment capabilities.

[0003] Multi-source housing data typically originates from different platforms and data structure systems, including structured attribute data, semi-structured tag data, and unstructured text or image data. These different data types exhibit significant differences in expression, feature dimensions, and semantic levels, while also displaying complex non-linear relationships between various features. Existing methods for processing this type of data often employ simple concatenation or independent modeling, lacking a unified mechanism for characterizing the inherent relationships between multi-source features. This results in the ineffective fusion of information from different types of housing features, making it difficult to form a consistent high-dimensional representation space. In this situation, the interaction between user features and housing features remains at a shallow matching level, failing to reflect potential preference drivers and implicit semantic connections, thus reducing matching accuracy. Summary of the Invention

[0004] This invention provides a housing resource processing method based on user profiling and deep reinforcement learning, which can improve the accuracy of the matching degree analysis between user features and housing resource features.

[0005] In a first aspect, the present invention provides a method for processing housing listings based on user profiling and deep reinforcement learning, the method comprising: Acquire multi-source user data corresponding to the target user and perform preprocessing to build a user profile system, thereby generating user feature vectors; Multi-source housing data is obtained through the user feature vector; Feature extraction and enhancement processing are performed on the multi-source housing data to generate housing feature vectors; A property matching recommendation model is constructed based on the user feature vector and the property feature vector, and joint modeling is performed through deep reinforcement learning to output the matching metric result between the user feature vector and the property feature vector; A property recommendation sequence is generated based on the matching metric results.

[0006] In a second aspect of the invention, a housing resource processing apparatus based on user profiling and deep reinforcement learning is provided. The apparatus is used to execute a housing resource processing method based on user profiling and deep reinforcement learning as described above. The apparatus includes an acquisition module, a processing module, and an output module, wherein: The acquisition module is used to acquire multi-source user data corresponding to the target user and perform preprocessing to construct a user profile system, thereby generating user feature vectors; The processing module is used to obtain multi-source housing data through the user feature vector; The processing module is used to perform feature extraction and enhancement processing on the multi-source housing data to generate housing feature vectors; The processing module is used to construct a housing matching recommendation model based on the user feature vector and the housing feature vector, and to perform joint modeling through deep reinforcement learning to output the matching metric result between the user feature vector and the housing feature vector; The output module is used to generate a property recommendation sequence based on the matching metric results.

[0007] In a third aspect of the invention, an electronic device is provided, including a processor, a memory, a user interface, and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any of the preceding embodiments.

[0008] In a fourth aspect of the invention, a non-transitory computer-readable storage medium is provided, the computer-readable storage medium storing instructions that, when executed, perform the method as described in any of the preceding claims.

[0009] In summary, one or more technical solutions provided in the embodiments of the present invention have at least the following technical effects or advantages: This invention constructs a hierarchical user profile system from multi-source user data to form user feature vectors with stable constraints, dynamic preferences, and a unified expression of potential needs. It then performs unified feature extraction and enhancement processing on multi-source housing data to form high-dimensional, consistent housing feature vectors, thereby establishing a deep semantic correspondence between user features and housing features in the same feature space. Furthermore, it introduces deep reinforcement learning to perform joint modeling on the housing matching and recommendation model, enabling the matching measurement results to not only rely on static feature similarity but also dynamically optimize the matching strategy based on real-time user interaction feedback. Simultaneously, it effectively characterizes complex relationships through multi-source feature fusion and nonlinear interaction modeling, avoiding information loss caused by shallow matching, thus improving the accuracy of user feature and housing feature matching degree analysis. Attached Figure Description

[0010] Figure 1 This is a flowchart illustrating a housing resource processing method based on user profiling and deep reinforcement learning disclosed in an embodiment of the present invention. Figure 2 This is a schematic diagram of a housing resource processing device based on user profiling and deep reinforcement learning disclosed in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present invention.

[0011] Explanation of reference numerals in the attached drawings: 201, acquisition module; 202, processing module; 203, output module; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. Detailed Implementation

[0012] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0013] In the description of the embodiments of the present invention, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0014] In the description of the embodiments of the present invention, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0015] Existing housing recommendation methods mainly rely on static matching mechanisms based on user profiles and basic housing attributes. Although deep learning is introduced to model user behavior in some scenarios, the overall approach is still dominated by offline training and shallow feature interaction, lacking the ability to dynamically adjust to user feedback. Furthermore, in the process of processing multi-source housing data, due to the differences in structural form, feature dimension, and semantic level of different data, existing methods are unable to uniformly model the nonlinear correlation between multi-source features, resulting in insufficient housing feature fusion. User features and housing features can only achieve surface matching, failing to characterize deep preference-driven mechanisms, thus limiting the accuracy and adaptability of housing matching and recommendation.

[0016] This invention discloses a housing resource processing method based on user profiling and deep reinforcement learning, which is applied to a server. The server includes, but is not limited to, electronic devices such as mobile phones, tablets, wearable devices, and PCs (Personal Computers), and can also be a backend server running a housing resource processing method based on user profiling and deep reinforcement learning. The server can be implemented using a standalone server or a server cluster composed of multiple servers.

[0017] This embodiment discloses a housing resource processing method based on user profiling and deep reinforcement learning, referring to... Figure 1 It includes the following steps: S110: Obtain multi-source user data corresponding to the target user and perform preprocessing to build a user profile system, thereby generating user feature vectors.

[0018] S120 obtains multi-source housing data through user feature vectors.

[0019] S130 performs feature extraction and enhancement processing on multi-source housing data to generate housing feature vectors.

[0020] S140: Construct a housing matching recommendation model based on user feature vectors and housing feature vectors, and perform joint modeling through deep reinforcement learning to output the matching metric results between user feature vectors and housing feature vectors.

[0021] S150 generates a property recommendation sequence based on the matching metric results.

[0022] In practice, the first step is to establish a multi-source user data collection scope centered around the target users, ensuring that the subsequently generated user feature vectors have a complete behavioral and semantic foundation. Multi-source user data refers to a collection of user-related data from different sources, with different structures and reflecting different perspectives. It typically includes basic attribute data, platform interaction data, consultation text data, search condition data, device environment data, and location trajectory data. Basic attribute data is used to characterize the relatively stable characteristics of the target users, such as age range, occupation type, budget range, family structure, rental / purchase preferences, and move-in timeframe. Platform interaction data is used to characterize the dynamic behavior of the target users on the real estate information platform, such as browsing duration, dwell time, page turning frequency, click frequency, favorites behavior, unfavorites behavior, consultation triggering behavior, and sharing behavior. Consultation text data is used to characterize the explicit and implicit needs expressed by the target users through the search box, online consultation window, or voice transcription entry. Search condition data is used to characterize the filtering conditions actively set by the target users, such as target area, price range, apartment type range, area range, decoration requirements, and commuting requirements. Device environment data is used to characterize the access terminal, access time period, network environment, and access frequency, thereby assisting in identifying the user's current decision-making activity level. Location trajectory data is used to characterize workplace, residence, frequently used activity areas, and cross-regional migration tendencies. In this embodiment, the target user refers to the user object currently entering the recommendation service link and requiring the generation of personalized housing recommendation results. Multi-source means that the data does not only come from a single behavior log table, but covers multiple business entry points and multiple information dimensions. User feature vector refers to the numerical expression result of mapping multi-source user data to the same feature space after unified processing, used as the standard input for subsequent housing matching and recommendation models.

[0023] After collecting multi-source user data, identity alignment and time-series alignment are performed on the multi-source user data to ensure that data from different sources can be attributed to the same target user and form a continuous behavioral chain. Identity alignment is used to solve the problem of different identifiers for the same target user in different data sources, such as inconsistent account identifiers, device identifiers, session identifiers, and consultation identifiers. During processing, a unified user identifier can be established based on login relationships, device association relationships, temporal adjacency relationships, and behavioral continuity relationships. Time-series alignment is used to solve the problem of different sampling frequencies and recording time granularities from different data sources, so that browsing behavior, collection behavior, consultation behavior, and trajectory behavior can be arranged according to a unified time axis. In this embodiment, identity alignment refers to the unified merging of data records pointing to the same target user from different sources. Time-series alignment refers to mapping data with different time granularities to a continuous time series. A behavioral chain refers to the continuous interaction process formed by a target user within one or more decision-making cycles. Through identity alignment and time-series alignment, the problem of the same user being split or the same behavior being misinterpreted can be avoided in the subsequent user profile system construction process, thereby improving the authenticity and usability of user feature vectors.

[0024] After alignment, preprocessing is performed on multi-source user data to eliminate interference from dirty data, outliers, and heterogeneous representations in subsequent modeling. Preprocessing typically includes missing value completion, outlier cleaning, duplicate record removal, field standardization, category coding, numerical normalization, and text normalization. Missing value completion fills in gaps in budget ranges, family structures, or regional preferences, and can be based on historical stable behavior or statistical results from similar user groups. Outlier cleaning removes records of behavior that clearly does not conform to real decision-making logic, such as a large number of consecutive clicks on multiple unrelated property listings within a very short period. Duplicate record removal eliminates redundant behavioral data caused by multi-terminal synchronization or repeated writes to interfaces. Field standardization unifies synonymous fields from different systems into a consistent representation, such as unifying total price, rent, and rental price into a price field, and unifying business district, area, and region into a region field. Category coding converts discrete fields such as occupation type, family structure, and decoration preferences into computable representations. Numerical normalization compresses fields with different numerical ranges, such as price, area, dwell time, and browsing frequency, to a uniform scale. Text normalization is used to perform word segmentation, stop item cleanup, synonym merging, and noise word removal on consultation texts, search terms, and message texts. In this embodiment, preprocessing refers to the cleaning, correction, unification, and transformation of the raw data before modeling. Field standardization refers to adopting unified naming and value rules for fields from different sources. Numerical normalization refers to mapping data at different numerical scales to a comparable range. Through preprocessing, multi-source user data can be transformed from raw, discrete records into a standardized dataset that can support the construction of a user profile system.

[0025] After preprocessing, a user profile system is constructed around the target users, enabling a hierarchical expression structure of their stable attributes, dynamic preferences, and potential needs. The user profile system is a structured expression system that organizes and models user information according to business semantics, transforming raw, multi-source user data into an interpretable set of profile features. To ensure the consistency of subsequent technical features, the user profile system can be divided into a basic attribute layer, a behavioral feature layer, and a potential need layer. The basic attribute layer represents relatively stable and clearly constrained characteristics of the target users, such as budget range, rental / purchase preferences, family structure, occupation type, move-in timeframe, and regional restrictions. The behavioral feature layer represents the dynamic interaction patterns of target users on the real estate information platform, such as browsing intensity, click bias, collection tendency, consultation depth, frequency of filter adjustment, and frequency of regional jumps. The potential need layer represents preferences that are not explicitly set by the target users but can be inferred from text semantics, behavioral migration, and location trajectory, such as a preference for quiet living environments, preference for school district resources, preference for low commuting burdens, preference for a complete social circle, and preference for improved living conditions. In this embodiment, the user profiling system is not a simple data table concatenation result, but a hierarchical feature expression structure organized around the recommendation goal. The basic attribute layer emphasizes stable constraints, the behavioral feature layer emphasizes dynamic preferences, and the latent demand layer emphasizes implicit intentions. By constructing a user profiling system, the subsequently generated user feature vectors can not only reflect users' surface preferences, but also their deeper decision-driving factors.

[0026] When constructing the basic attribute layer, stable fields that have a rigid constraint effect on property listing selection are first extracted from the preprocessed multi-source user data, and then regularized and mapped onto them. For budget range, budget features can be determined based on the user's input budget, the price ceiling mentioned in the consultation, and the historical price distribution of viewed properties. For family structure, family structure features can be determined based on user registration information, the semantics of family members in the consultation text, and the preference for apartment types. For renting and buying preferences, they can be determined based on the business entry point accessed by the user, the transaction type of the viewed properties, and the consultation intent. For move-in time limit, it can be extracted based on the moving time, signing time, or semester start time mentioned by the user in the consultation text. For regional restrictions, they can be determined based on the regional filtering conditions directly set by the user and the long-term browsing regional distribution. In this embodiment, rigid constraints refer to user conditions that are prioritized in the property listing selection process, and deviations from these conditions will significantly reduce the matching degree. Regularized mapping refers to converting the original fields into standardized features according to preset business rules. By constructing the basic attribute layer, clear boundary conditions can be provided for subsequent property listing acquisition and matching, ensuring that the recommendation results do not deviate from the basic residential constraints of the target user.

[0027] When constructing the behavioral feature layer, the focus is on extracting dynamic interaction features that reflect the target user's true interest intensity and interest migration trends. During processing, browsing behavior can be used to extract browsing frequency, continuous browsing duration, detail page dwell time, image switching frequency, return visits, and page-turning behavior. Click behavior can be used to extract click density, regional click concentration, price range click distribution, and apartment type click distribution. Collection behavior can be used to extract the percentage of collected properties, uncollection frequency, collection regional concentration, and collection price range preference. Inquiry behavior can be used to extract inquiry trigger frequency, inquiry round depth, inquiry topic distribution, and inquiry target change trajectory. Filtering behavior can be used to extract the stability of filter conditions, frequency of filter condition adjustments, and trend of filter direction changes. In this embodiment, the behavioral feature layer refers to a profile layer specifically designed to express dynamic interest patterns. Interest intensity refers to the degree of attention a target user pays to a certain type of property or a certain region. Interest migration trend refers to the direction and magnitude of changes in a target user's preferences over time. By constructing the behavioral feature layer, seemingly scattered browsing, clicking, collection, and inquiry behaviors can be transformed into measurable preference expressions, providing a basis for subsequent models to identify current key needs.

[0028] When constructing a layer around latent needs, joint analysis of consultation text, search terms, regional migration trajectories, and behavioral sequences is required to identify implicit needs that are not explicitly set by the target user but actually exist. For consultation text, word segmentation, named entity recognition, intent recognition, sentiment semantic recognition, and contextual analysis can be performed to extract latent need signals from the text, such as proximity to schools, convenient commuting, quieter environment, convenient nearby amenities, future value retention, and suitability for elderly living. For search terms, demand-indicating words such as school district, subway, improvement, move-in ready, and low total price can be identified. For location trajectories, workplace location, resident living area, weekend activity area, and regional migration direction can be identified. For behavioral sequences, potential preference transfer patterns such as convergence from distant to nearby areas, migration from small apartments to three-bedroom apartments, and migration from low-priced to high-end properties can be identified. In this embodiment, the latent needs layer refers to the implicit preference expression layer obtained through semantic understanding and behavioral inference. Named entity recognition refers to identifying objects with clear semantic orientations, such as area names, transportation nodes, and school names, from the text. Intent recognition refers to determining the underlying needs expressed in text, such as whether it emphasizes commuting, school districts, improvement, or price control. By constructing a layer of latent needs, we can overcome the limitations of relying solely on explicit filtering conditions, enabling user profiling systems to express implicit decision-making factors.

[0029] After the basic attribute layer, behavioral feature layer, and latent demand layer are constructed, unified mapping and fusion processing are performed on the features of each layer to generate user feature vectors. Unified mapping refers to mapping features of different types, dimensions, and sources to the same feature space through methods such as category embedding, continuous value projection, text semantic encoding, and sequence representation encoding. Fusion processing refers to combining the three layers of features into a unified representation result according to preset weights or learnable weights. For the basic attribute layer, stable constraint representations can be obtained using category embedding and numerical projection. For the behavioral feature layer, dynamic preference representations can be obtained using statistical aggregation representations and temporal encoding representations. For the latent demand layer, implicit intent representations can be obtained using text semantic encoding and trajectory encoding. Then, the above representations are concatenated, weighted, or gated fusion to form the final user feature vector. In this embodiment, unified mapping refers to transforming heterogeneous data into comparable representations in the same space. Feature space refers to a vector space that carries various features and allows subsequent models to perform unified calculations. Gated fusion refers to a fusion method that dynamically adjusts the contribution level of different features based on the current feature importance. Through unified mapping and fusion processing, the multidimensional information of the target user can be output in the form of a single continuous vector, providing a standardized input for subsequent acquisition of multi-source housing data through user feature vectors and construction of housing matching and recommendation models.

[0030] After generating user feature vectors, it is necessary to perform profile consistency verification and vector usability verification to ensure that the user feature vectors can accurately reflect the current state of the target user and directly enter the subsequent recommendation process. Profile consistency verification is used to check whether there are obvious conflicts between the basic attribute layer, behavioral feature layer, and potential demand layer. For example, if the budget range is low but the user continues to browse high-end improved housing listings and the consultation text emphasizes the low total price, it is necessary to further determine whether the budget has changed or the browsing behavior is exploratory. Vector usability verification is used to check whether there is a high proportion of missing features, excessive short-term abnormal behavior disturbances, or excessive semantic feature noise. If an anomaly is detected, the process returns to the preprocessing stage to re-perform missing correction, anomaly cleaning, or feature recalculation. If no anomaly is detected, the user feature vector is written to the recommendation cache and enters the subsequent housing acquisition and matching process. In this embodiment, consistency verification refers to checking whether different profile layers support each other in terms of business semantics. Usability verification refers to checking whether the generated user feature vectors meet the input conditions of the subsequent model. Through this process, it can be ensured that the user profile system is not a static storage result, but a valid input result that can directly support dynamic recommendation decisions.

[0031] In one possible implementation, after constructing a property matching and recommendation model based on user feature vectors and property feature vectors, the method further includes: constructing user feature vectors and property feature vectors as states in a Markov decision process, and outputting action values ​​through a deep Q-network. Specifically, a user profile system is constructed based on a basic attribute layer, a behavioral feature layer, and a potential demand layer to generate user feature vectors; property image data, text description data, structural attribute data, and spatial location data are encoded and weighted and fused using a convolutional neural network and an attention mechanism to generate property feature vectors; candidate properties are ranked based on action values ​​to generate a property recommendation sequence; a multi-dimensional reward function is constructed based on user interaction feedback, and the parameters of the property matching and recommendation model are updated using an experience replay mechanism and a target network update mechanism; and the property recommendation sequence is adjusted based on real-time interaction data and offline training using an ε-greedy strategy based on historical data.

[0032] Specifically, user feature vectors and property feature vectors are first organized into a state representation that can be used for reinforcement learning decision-making. This allows the property matching recommendation model to make action selections at each recommendation time based on the current user demand state and the candidate property attribute state. In this embodiment, a Markov decision process refers to a mathematical model that describes a continuous decision-making process using states, actions, rewards, and state transitions. The current decision depends only on the current state and action, and does not directly depend on the original observation records from earlier times. A state refers to the comprehensive information used to represent the recommendation environment at the current recommendation time. An action refers to the recommendation selection behavior performed by the model on candidate properties in the current state. Action value refers to the long-term benefit estimate obtained after choosing a certain action in a certain state. To construct the state representation, a user profile system needs to be built based on the basic attribute layer, behavioral feature layer, and potential demand layer. The user profile system is then mapped to user feature vectors. Then, unified encoding is performed on property image data, text description data, structural attribute data, and spatial location data to generate property feature vectors. Then, the user feature vector, the property feature vector in the candidate property set, the current session context features, and the historical feedback trajectory are combined into the state input to ensure that the state not only reflects the user's stable preferences, but also reflects the real-time interest changes in the current session and the attribute distribution of the candidate properties.

[0033] The generation of user feature vectors relies on the hierarchical modeling results of the user profile system. The basic attribute layer expresses stable constraint features such as budget range, rental / purchase preferences, family structure, occupation type, and move-in timeframe. The behavioral feature layer expresses dynamic preference features such as browsing intensity, click density, collection trends, consultation depth, and frequency of filtering adjustments. The latent demand layer expresses implicit demand features parsed from consultation text, search expressions, and location trajectories. In this embodiment, the user profile system refers to a hierarchical user information expression structure constructed for recommendation tasks. The basic attribute layer emphasizes relatively stable constraints, the behavioral feature layer emphasizes recent behavior-induced preference changes, and the latent demand layer emphasizes deep-seated needs not yet fully expressed through explicit filtering conditions. To enable the three layers of features to participate in subsequent state construction, unified mapping and fusion processing is required for different types of features to form a unified-dimensional user feature vector. The expression for the user feature vector is:

[0034]

[0035] Where u represents the user feature vector; This represents the feature vector of the basic attribute layer, which is used to carry stable constraint features such as budget range, family structure, rental and purchase preferences, occupation type and move-in time limit; This represents the feature vector of the behavioral feature layer, which is used to carry dynamic preference features such as browsing behavior, click behavior, collection behavior, consultation behavior, and filtering behavior; This represents the feature vector of the potential demand layer, which carries potential demand features such as text semantic parsing results, location trajectory relationships, and implicit demand labels. , and These represent the mapping weight matrices corresponding to the three layers of features, used to map features from different sources to a unified feature space; This represents the bias vector, used to correct the offset of different user groups in the overall representation. The principle behind this formula is to map user information at different semantic levels to the same representation space, and then fuse them through learnable weights, so that the user feature vector simultaneously retains stable constraints, dynamic preferences, and potential needs.

[0036] The generation of property feature vectors relies on the joint encoding results of property image data, text description data, structural attribute data, and spatial location data. In this embodiment, property image data refers to visual information such as actual property photos, floor plans, and community maps. Text description data refers to textual information such as titles, selling point descriptions, amenities descriptions, and surrounding information. Structural attribute data refers to structured fields such as price, area, unit type, floor, orientation, decoration status, and building age. Spatial location data refers to latitude and longitude location, administrative division, business district affiliation, proximity to rail transit, and surrounding educational, medical, and commercial facilities. To ensure that the above different types of property data can be uniformly incorporated into the property matching and recommendation model, image encoding, text encoding, attribute mapping, and location embedding are performed separately, and then weighted and fused using a convolutional neural network and attention mechanism. A convolutional neural network is a deep network structure adept at extracting local spatial patterns, suitable for extracting visual semantics such as room layout, lighting characteristics, decoration style, and spatial cleanliness from property images. Attention mechanisms are computational mechanisms that assign different weights to different feature components, and can be used to highlight property attributes that are more important to the current user. The expression for the property feature vector is:

[0037]

[0038] Where v represents the property feature vector; The i-th type of housing feature component can correspond to image feature component, text semantic component, structural attribute component or spatial location component; represents the attention weight of the i-th type of property feature component, used to characterize the importance of this feature component to the current matching task; n represents the total number of property feature components participating in the fusion. The principle of this formula is to control the contribution of various property features to the final property feature vector through attention weights, so that property features that are more relevant to the current user's needs occupy a larger proportion in the final expression, thereby enhancing the personalized adaptation capability of the property feature vector.

[0039] After obtaining the user feature vector and the property feature vector, they are used as the state input in a Markov decision process, and the action value is output through a deep Q-network. A deep Q-network is a reinforcement learning model that uses a deep neural network to approximate the action value function; its core function is to estimate the long-term payoff of each candidate action given a state. Specifically, a joint mapping is first performed on the user feature vector and the candidate property feature vector, explicitly incorporating the correspondence between user preferences and property attributes into the state representation, which is then input into the deep Q-network. To enhance the state's ability to describe the recommendation context, recently viewed sequences, recently clicked results, and recent changes in filtering conditions in the current session can be introduced, thus avoiding the state only reflecting static user profiles and static property attributes. The state representation can be written as:

[0040]

[0041] in, This represents the state vector corresponding to time t; This represents the user feature vector at time t; Represents the set of candidate housing features at time t; It represents the contextual characteristics of the current session and is used to carry real-time behavioral information such as recent browsing, recent clicks, recent favorites, and changes in filter conditions; It represents the characteristics of historical feedback trajectories and is used to carry long-term feedback patterns formed over multiple recommendation periods. This represents the state construction function, used to map information from different sources into a unified state representation. The principle behind this formula is to map user state, candidate property state, immediate context, and historical feedback together into an environmental description at the current decision-making moment, enabling deep Q-networks to learn recommendation strategies based on more complete information.

[0042] After receiving the state vector, the deep Q-network outputs the action value corresponding to each candidate action. In this embodiment, a candidate action can be defined as recommending a specific property, recommending a combination of properties, or determining a display order among the candidate properties. The higher the action value, the higher the expected long-term benefit of performing the action in the current state. The expression for the action value output by the deep Q-network is:

[0043]

[0044] in, This indicates that the network parameters are At that time, state Select action The corresponding action value; This represents the expectation operation, used to characterize the average long-term return under future stochastic feedback conditions; This represents the discount factor, which typically ranges from zero to one and is used to balance the importance of current returns with future returns. This represents the reward value obtained at the k-th step starting from time t. The principle behind this formula is that the action value does not only measure the immediate effect of the current action, but also considers the cumulative impact of the action on the recommendation effect in multiple future steps. Therefore, it is suitable for continuous interactive decision-making scenarios such as property recommendation.

[0045] Based on the action value output by the deep Q-network, candidate properties are ranked and a property recommendation sequence is generated. In this embodiment, candidate properties refer to the set of properties to be ranked after basic constraint filtering, elastic preference screening, and implicit demand compensation recall at the current recommendation time; the property recommendation sequence refers to the property display results arranged according to recommendation priority. Specifically, each property in the candidate property set is first combined with the current user state to form a candidate action. Then, the action value corresponding to each candidate action is calculated using the deep Q-network. Finally, the properties are ranked from highest to lowest action value to generate the property recommendation sequence. To improve the diversity and robustness of the recommendation results, regional dispersion constraints, price range coverage constraints, and appropriate exposure constraints for newly listed properties can be introduced during the ranking process to avoid the recommendation results being too concentrated on a certain type of property. The expression for the ranking score can be written as:

[0046]

[0047] in, This represents the ranking score corresponding to the i-th candidate property. This represents the action value corresponding to recommending the i-th candidate property in the current state; This represents the matching supplement score of the i-th candidate property, which reflects the degree to which the property matches the user's constraints, implicit demand tags, and current session interests. This represents the penalty score for the i-th candidate property, used to reflect adverse factors such as repeated exposure, excessive concentration in a particular area, or decreased timeliness. , and This represents the weight parameters for different scoring items. The principle behind this formula is to prioritize the ranking based on action value, while simultaneously adjusting the ranking results by incorporating matching supplementary items and penalty items, so that the final property recommendation sequence takes into account long-term benefits, personalized matching, and result quality.

[0048] After the property recommendation sequence is generated and delivered to the user, a multi-dimensional reward function is constructed based on user interaction feedback to quantify the actual effect of the recommendation results. In this embodiment, user interaction feedback refers to user behaviors such as clicking, staying, saving, inquiring, scheduling viewings, completing transactions, swiping away, quickly exiting, blocking, or reporting on each property in the property recommendation sequence. The multi-dimensional reward function is a function that maps various feedback behaviors to a unified reward value according to their different business values. Since the value of different feedback behaviors varies significantly in the property recommendation scenario, different weights need to be assigned to positive feedback, weak negative feedback, and strong negative feedback so that the model can learn optimization directions that better align with real business objectives. The expression for the multi-dimensional reward function is:

[0049]

[0050] in, This represents the overall reward value at time t; Indicates the intensity of click feedback; Indicates the intensity of dwell time feedback; Indicates the intensity of feedback regarding collection; Indicates the intensity of consultation feedback; This indicates the level of feedback regarding appointment bookings and viewings; Indicates the strength of transaction feedback; Indicates the intensity of the feedback after the swipe away; Indicates the intensity of rapid exit feedback; Indicates the strength of the shielding feedback; Indicates the intensity of the whistleblower feedback; to This represents the weight parameters of each feedback item. The principle behind this formula is to uniformly map user feedback at different levels into continuous reward signals, enabling the property matching and recommendation model to learn not only surface-level click behavior but also higher-value inquiries, viewings, and transactions. At the same time, it penalizes negative feedback, thereby forming an optimization direction that is more in line with the platform's goals.

[0051] After the reward value is generated, the parameters of the housing matching recommendation model are updated by combining the experience replay mechanism and the target network update mechanism. The experience replay mechanism refers to storing the states, actions, rewards, and next-state samples formed during the model's historical recommendation process in an experience sample pool, and randomly sampling these samples during training to learn from them, thereby breaking the strong correlation between adjacent samples and improving training stability. The target network update mechanism refers to using a target network with a low update frequency to calculate the target action value during training, thereby reducing the instability caused by rapid fluctuations in the training target. Specifically, the current recommendation process's state, actions, reward, and next state are first constructed into state transition samples, then written into the experience sample pool. Training batches are then extracted from the experience sample pool, and the target action value is calculated through the target network. The expression for the target action value is:

[0052]

[0053] in, This represents the value of the target action at time t. This represents the overall reward value at time t; Indicates the discount factor; Indicates the target network in the next state Next candidate action Estimation of the value of the output action; Indicates the target network parameters; This means selecting the action with the highest value from all candidate actions. The principle behind this formula is that it uses both the current actual reward and the optimal future reward in the next state as training objectives, enabling the model to learn the long-term optimal recommendation strategy.

[0054] After obtaining the target action value, the deviation between the current network output and the target action value is calculated using a loss function, and the parameters of the housing matching recommendation model are updated accordingly. The loss function can be written as:

[0055]

[0056] in, This represents the loss value in the current training batch; N represents the number of samples in the training batch. This represents the target action value of the t-th sample; Indicates the current state of the network. Next action Estimation of the value of the output action; This represents the current network parameters. The principle behind this formula is that by minimizing the mean squared deviation between the current estimate and the training objective, the network output gradually approaches the true long-term return, thereby continuously improving the decision-making accuracy of the housing matching recommendation model.

[0057] After the model parameter update pipeline is established, offline training is performed based on historical data using an ε-greedy strategy, and the property recommendation sequence is adjusted based on real-time interactive data. In this embodiment, the ε-greedy strategy refers to using probability during the action selection process... Perform exploration actions, with probability The action with the highest current value is selected; offline training refers to training the initial model using historical interaction data before official launch, enabling the model to have basic decision-making capabilities; real-time interaction data refers to the new feedback samples continuously received during the actual recommendation process after the model goes live; property recommendation sequence adjustment refers to dynamically correcting the ranking structure, regional quotas, and exploration ratios of the current recommendation results based on the latest action value and the latest feedback status. The expression for the ε-greedy action selection rule is:

[0058]

[0059] in, Indicates the action chosen at time t; This represents the probability of exploration and is used to control the frequency of random exploration actions. This indicates choosing the action with the highest value in the current state. The principle behind this formula is that while utilizing the existing optimal strategy, a certain proportion of random exploration is retained, thereby avoiding the model getting trapped in local optima and providing exposure opportunities for new properties, new areas, or new preference patterns.

[0060] The offline training phase primarily utilizes historical browsing, click, favorites, inquiry, transaction, and property attribute data to construct training samples. This allows the model to learn the basic mapping relationship between user profiles and property attributes, as well as the long-term benefit relationship between recommendation actions and feedback. The real-time adjustment phase incrementally refines the property recommendation sequence based on the latest interaction data. For example, when a user continuously clicks on properties in a certain area during the current session, the ranking weight of candidate properties in that area is increased in the property recommendation sequence; when a user continuously swipes away from a certain type of property, the display ratio of that type of property is reduced; when newly listed properties lack sufficient behavioral samples in historical data, an ε-greedy strategy is used to allocate appropriate exploration opportunities for them. After this processing, the property recommendation sequence no longer simply relies on the static output of offline training results but can continuously adjust with real-time interaction feedback, thereby improving the timeliness, adaptability, and long-term benefit performance of the recommendation results.

[0061] By linking the above processes together, a complete technical chain is formed, encompassing user profile construction, property feature extraction, state construction, action value estimation, property ranking, reward construction, experience replay training, target network updates, and online recommendation adjustments. In this chain, user feature vectors provide the foundation for expressing user needs, property feature vectors provide the foundation for expressing property attributes, Markov decision processes provide a continuous decision-making framework, deep Q-networks provide action value estimation capabilities, multidimensional reward functions provide quantitative feedback, experience replay mechanisms and target network update mechanisms provide training stability guarantees, and ε-greedy strategies provide a balance between exploration and utilization.

[0062] In one possible implementation, after generating a property recommendation sequence based on the matching metric results, the method further includes: constructing a recommendation snapshot record based on the property recommendation sequence and establishing a mapping relationship between recommendation identifiers and user interaction feedback; performing event standardization and time decay aggregation on the user interaction feedback to generate a target feedback vector; performing weighted mapping on the target feedback vector to generate a comprehensive reward value; constructing state transition samples based on the comprehensive reward value and performing incremental updates on the user feature vector in conjunction with user interaction feedback to generate the next state; writing the state transition samples into an experience sample pool and performing hierarchical storage and priority management; extracting training batches based on the experience sample pool and calculating the target action value in conjunction with the target network to update the parameters of the property matching recommendation model; and performing backpropagation based on the loss between the target action value and the current action value to optimize the property matching recommendation model.

[0063] Specifically, after the property recommendation sequence is generated, a recommendation snapshot record is first established around this recommendation behavior, so that any subsequent user interaction feedback can be accurately traced back to the corresponding recommendation context. The recommendation snapshot record is a structured record that completely freezes the recommendation output result at the time of generation, used to save the environmental state, candidate objects, ranking results, and decision basis of this recommendation behavior. When constructing the recommendation snapshot record, a unique recommendation identifier is assigned to each recommended property in the property recommendation sequence, and this recommendation identifier, along with the target user identifier, recommendation time, display page position, display order, session identifier, user feature vector summary, property feature vector summary, matching metric result, action selection result, candidate property set summary, and the exploration probability used at that time, are written into the recommendation log. In this embodiment, the recommendation identifier is a unique index mark used to uniquely correspond to a specific recommended property being delivered to a specific target user at a specific recommendation time; the recommendation snapshot record is not a simple exposure record, but a frozen record containing state, action, and ranking context. This process allows subsequent user interactions such as clicks, favorites, inquiries, viewings, transactions, swipes, blocking, and reports to be accurately mapped back to the decision state at the time the recommendation result was generated, thus providing a reliable basis for subsequent reward construction, state transition construction, and model updates.

[0064] After establishing recommendation snapshot records, user interaction feedback is continuously collected, and a one-to-one mapping relationship is established between recommendation identifiers and user interaction feedback. User interaction feedback refers to the explicit or implicit behaviors of target users towards recommended properties in the property recommendation sequence, including clicking after exposure, staying on the details page, switching images, adding to favorites, unadding to favorites, inquiring, scheduling viewings, visiting properties in person, completing a transaction, swiping away, quickly exiting, blocking, and reporting. When establishing the mapping relationship, feedback events are first extracted from the front-end behavior logs, inquiry system logs, viewing system logs, transaction system logs, and risk control system logs. Then, each feedback event is linked back to the corresponding recommendation snapshot record through the recommendation identifier. If the same recommended property generates multiple behavior events on different terminals, the session identifier, behavior timestamp, and device identifier are further combined to merge multi-source feedback. The core function of this process is to enable the model to clearly know which recommendation generated which type of feedback, avoiding subsequent misattribution of irrelevant feedback to incorrect recommendation actions, thereby ensuring a strict causal correspondence between states, actions, and rewards in reinforcement learning.

[0065] After the feedback loop is completed, event standardization processing is performed on user interaction feedback, enabling behavior records from different business subsystems to be converted into a unified format of feedback event sequences. Event standardization processing includes unified encoding of event types, unified format of event times, unified quantification of behavior intensity, merging of duplicate events, removal of abnormal events, and cross-terminal feedback association. Unified encoding of event types is used to map synonymous behavior expressions from different systems to unified feedback categories. For example, viewing details, opening property listings, and clicking cards are uniformly mapped to click events; initiating consultation IM, triggering telephone consultations, and submitting message consultations are uniformly mapped to consultation events. Unified quantification of behavior intensity is used to transform originally discrete or textual business records into continuous value representations. For example, the intensity value of dwell behavior can be determined based on dwell time, number of image switching, and detail expansion depth; the intensity value of consultation behavior can be determined based on whether a consultation occurred, consultation rounds, and consultation topic coverage. In this embodiment, event standardization refers to converting heterogeneous feedback records into feedback expressions with unified definitions, unified formats, and unified granularity, so that subsequent reward functions can directly process them. After event standardization is completed, user interaction feedback is no longer a scattered log, but a computable, comparable, and aggregable sequence of feedback events.

[0066] After event standardization, time-decay aggregation is performed on user interaction feedback to generate a target feedback vector. The purpose of time-decay aggregation is that, in property recommendation scenarios, recent user feedback typically reflects the current demand state better than older feedback; therefore, different weights need to be assigned to feedback at different time points. The target feedback vector is a unified vector representation formed by time-weighted aggregation of a set of standardized feedback events corresponding to the same recommendation identifier, where each dimension corresponds to the aggregation strength of a certain type of feedback. The expression for time-decay aggregation is:

[0067]

[0068] in, This represents the target feedback vector at time t; K represents the number of feedback events associated with the current recommendation identifier. This represents the k-th standardized feedback event vector, used to characterize the type and intensity of an event; This represents the time interval between the k-th feedback event and the recommendation time. This represents the time decay coefficient, used to control the rate at which the impact of feedback decays over time. The principle behind this formula is to assign higher weights to feedback closer to the recommendation time and lower weights to feedback occurring later, thus making the target feedback vector more biased towards reflecting the direct impact of the current recommendation action. This process distinguishes between immediate inquiries after a click and weak interactions that occur a considerable time after exposure, facilitating a more accurate characterization of the recommendation effect.

[0069] After obtaining the target feedback vector, a weighted mapping process is performed on the target feedback vector to generate a comprehensive reward value. The comprehensive reward value is the core supervision signal in the reinforcement learning chain, used to quantify the effectiveness of a recommendation action in the current state. Because user feedback in property recommendation scenarios has multiple levels of value differences, clicks and non-clicks cannot be simply used as the sole reward; instead, a layered weighting based on behavioral value is required. Clicks and dwell time are typically considered primary interest feedback, favorites and inquiries are considered intermediate intention feedback, appointments for viewings and transactions are considered high-value conversion feedback, and swiping away, quick exit, blocking, and reporting are considered negative feedback. The expression for the comprehensive reward value is:

[0070]

[0071] in, This represents the comprehensive reward value corresponding to time t; This represents the click feedback component in the target feedback vector; Indicates the dwell feedback component; This indicates the amount of feedback received for adding the item to the collection. Indicates the weight of consultation feedback; This indicates the amount of feedback received regarding the scheduled viewing; Indicates the weight of transaction feedback; This indicates that the feedback component has been swiped away; Indicates a quick exit from the feedback component; Indicates that the feedback component is masked; Indicates the weight of the report / feedback; to These represent the weight parameters for different feedback components. The principle behind this formula is that by assigning different positive and negative weights to different feedback types, the model's learning objective expands from simply increasing click-through rate to increasing consultation rate, viewing rate, and conversion rate, while simultaneously suppressing negative experiences. The weight parameters can be set based on business objectives or calibrated using historical data. For example, in conversion-focused scenarios, the weights corresponding to appointment viewing and conversion feedback are higher than those for click and dwell time feedback.

[0072] After the comprehensive reward value is generated, a state transition sample is constructed around the current recommendation action, and the user feature vector is incrementally updated based on user interaction feedback to generate the next state. The state transition sample is a standard training unit required for reinforcement learning training, typically composed of the current state, current action, current reward, and next state. The current state is taken from the frozen decision-making state expression in the recommendation snapshot record; the current action is taken from the specific selection result of candidate properties in the recommendation snapshot record; the current reward is the aforementioned comprehensive reward value; and the next state needs to be reconstructed after updating the user state based on the current feedback. In this embodiment, the next state is not simply a backward shift of the current state, but rather a reflection of the user's changing interest after receiving the recommendation and generating feedback. The expression for the incremental update of the user feature vector is:

[0073]

[0074] in, This represents the user feature vector after the feedback effect; This represents the user feature vector before the feedback effect; This represents the incremental update coefficient, which is usually set between zero and one to control the intensity of the impact of this feedback on the original user feature vector. Represents the target feedback vector; This indicates the current session context features, used to characterize recently browsed areas, recently consulted topics, recently clicked property types, and recent changes in filter criteria; This represents a feedback-driven state update function used to generate new interest expressions based on the original user state, current feedback, and current context. The principle behind this formula is to smoothly integrate the original long-term preferences with the current short-term feedback, avoiding excessive manipulation of the user state by a single, occasional action, while allowing continuous and consistent feedback to gradually change the user's feature vector. The updated user feature vector is then recombine with a new set of candidate properties, contextual information, and historical feedback trajectories to form the next state, thus completing the construction of the state transition sample.

[0075] After the state transition samples are constructed, they are written into the experience sample pool, and hierarchical storage and priority management are implemented in the experience sample pool. The experience sample pool is a cache structure used to store historical state transition samples. Its purpose is to provide a multi-time period, multi-type, and multi-distribution sample source for model training, thereby avoiding overfitting of recent data caused by updating the model only based on the latest samples. Hierarchical storage refers to dividing samples into different sample layers based on user type, region type, room type, price range, feedback category, and time window; priority management refers to assigning different sampling probabilities to different samples based on their importance, scarcity, and prediction error. High-value samples typically include appointment viewing samples, transaction samples, report samples, inquiries after continuous favorites, and samples with significant model prediction bias. The expression for priority scoring is:

[0076]

[0077] in, This represents the priority score of the i-th state transition sample; This represents the temporal difference error corresponding to the sample, which reflects the magnitude of the model's current prediction bias for that sample. This indicates the business value intensity of the sample, reflecting whether the sample corresponds to high-value events such as consultation, viewing, transaction, or strong negative feedback; This is an index representing the freshness of the sample, used to reflect how far the sample is from the current time. , and These represent the weight parameters for the prediction error term, business value term, and freshness term, respectively. The principle behind this formula is to allow the model to learn more frequently from samples that are currently difficult to predict, have high business value, and are more recent, while retaining ordinary samples to maintain overall distribution stability. Through this process, the experience sample pool becomes not just a simple cache, but a training resource management structure that balances distribution coverage and focused learning.

[0078] After establishing and maintaining the experience sample pool, training batches are extracted from the experience sample pool and combined with the target network to calculate the target action value, thereby updating the parameters of the housing matching recommendation model. A training batch refers to a set of samples selected from the experience sample pool according to certain sampling rules, used for a single parameter update. The target network is an auxiliary network used for stable training in deep Q-learning; its parameter update frequency is lower than that of the current network. It is used to calculate the training objective, thereby reducing training oscillations caused by rapid changes in the current network parameters. The core structure of the housing matching recommendation model can be designed as five parts: a user encoding subnetwork, a housing encoding subnetwork, a state fusion subnetwork, an action value estimation subnetwork, and a target network. The user encoding subnetwork receives inputs from the basic attribute layer, behavioral feature layer, and potential demand layer, and outputs user feature vectors. The property encoding subnetwork receives property image data, text description data, structural attribute data, and spatial location data, and outputs property feature vectors. The state fusion subnetwork jointly maps user feature vectors, candidate property feature vectors, current session context, and historical feedback trajectories to form a current state representation. The action value estimation subnetwork outputs action values ​​for different candidate actions based on the current state. The target network has the same structure as the action value estimation subnetwork, but its parameters are updated using a delayed synchronization method. If a candidate action is defined as recommending a single candidate property, the state fusion subnetwork can construct a state-action pair for each candidate property. If a candidate action is defined as ranking the top few properties in the candidate set, the action value estimation subnetwork can output the value of the corresponding combined action. The expression for the target action value is:

[0079]

[0080] in, This represents the value of the target action at time t. This represents the overall reward value; This represents the discount factor, used to balance current rewards with future rewards; Indicates the target network in the next state Next candidate action Estimation of the value of the output action; This represents the target network parameters. The principle behind this formula is to use both the immediate benefit of the current action and the maximum future benefit in the next state as training objectives, enabling the model to learn a recommendation strategy that maximizes long-term benefits, rather than optimizing only one click.

[0081] After the target action value is calculated, backpropagation is performed based on the loss between the target action value and the current action value to optimize the property matching recommendation model. The current action value is output by the current network, while the target action value is determined jointly by the target network and the overall reward. The deviation between the two is the learning error that the model needs to reduce. The expression for the loss function is:

[0082]

[0083] in, This represents the loss value corresponding to the current training batch; N represents the number of samples in the training batch. This represents the target action value of the t-th sample; Indicates the current state of the network. The following are the actual actions to be performed. Estimation of the value of the output action; This represents the current network parameters. The principle behind this formula is to measure the current network's prediction bias of action value using mean squared error, and then adjust the network parameters along the error-reducing direction using backpropagation. During backpropagation, the gradient of the action value estimation subnetwork is forwarded to the state fusion subnetwork, and then to the user encoding subnetwork and the property encoding subnetwork, thereby achieving end-to-end joint optimization from value error to user representation, property representation, and state representation. If there are abnormally large reward samples or abnormally large error samples during training, gradient clipping constraints can be added outside the loss function to prevent drastic oscillations in parameter updates.

[0084] The complete training and application process of the property matching and recommendation model includes three stages: offline pre-training, online updating, and online inference. In the offline pre-training stage, large-scale state transition samples are constructed using historical multi-source user data, historical property data, and historical feedback data. The basic parameters of the user encoding subnetwork, property encoding subnetwork, and action value estimation subnetwork are trained first, enabling the model to possess basic user demand understanding and property value judgment capabilities. In the online updating stage, the model continuously receives recommendation snapshot records and user interaction feedback during actual recommendation processes, constructs state transition samples in real time and writes them into the experience sample pool, and then incrementally updates the current network by sampling training batches at a set frequency. Simultaneously, the current network parameters are synchronized to the target network according to a preset update cycle. In the online inference stage, when a new recommendation request arrives, the user encoding subnetwork first generates a user feature vector, then the property encoding subnetwork generates candidate property feature vectors, followed by the state fusion subnetwork constructing the current state, and finally the action value estimation subnetwork outputs the action value of each candidate action, generating a property recommendation sequence based on the action value. In this way, recommendation, feedback, training and re-recommendation form a closed loop, enabling the housing matching recommendation model to continuously absorb new interactive information and gradually adjust its decision-making strategy.

[0085] To further enhance the model's adaptability in housing recommendation scenarios, a dual-stream structure can be introduced into the action value estimation subnetwork. This involves modeling the user-side interest intensity branch and the housing-side accessibility branch separately, then fusing them at the end. The user-side interest intensity branch focuses on characterizing user preferences for different regions, price ranges, and apartment types; the housing-side accessibility branch focuses on characterizing the housing's availability, repetition rate, historical interaction frequency, and regional supply scarcity. The fused action value output from these two branches is more suitable for housing recommendation scenarios that simultaneously emphasize preference matching and supply availability. To further enhance the model's ability to perceive multi-step recommendation chains, the recent recommendation results and feedback results can be input into the state fusion subnetwork using temporal encoding. This ensures that the current action value depends not only on the current single-point state but also on recent recommendation trajectories. This better addresses the strategy adjustment issues when users continuously browse, reject, or focus on a particular region.

[0086] In the entire model chain, recommendation snapshots freeze the decision-making process; user interaction feedback provides evidence of actual effectiveness; event standardization and time decay aggregation transform complex behaviors into unified feedback expressions; the comprehensive reward value converts feedback expressions into learnable signals; state transition samples transform a single recommendation action into a reinforcement learning training unit; the experience sample pool accumulates multi-time-period and multi-type training resources; the target network provides stable training objectives; and backpropagation drives model parameter iteration. The property matching recommendation model continuously refines its understanding of user needs, property features, and recommendation strategies through these mechanisms. Through this chain, the model can evolve from an initial recommendation method relying on static matching metrics to a dynamic recommendation method capable of continuous optimization based on long-term benefits, thereby improving the accuracy of the property recommendation sequence.

[0087] In one possible implementation, a property matching recommendation model is constructed based on user feature vectors and property feature vectors, and joint modeling is performed through deep reinforcement learning to output a matching metric between the user feature vectors and property feature vectors. Specifically, this includes: performing uniform dimension alignment and component normalization on the user feature vectors and property feature vectors, and partitioning the space to establish feature associations; performing bilinear interaction mapping on the user feature vectors and property feature vectors to generate coupled matching representations; performing dimension-wise difference calculation and weighted mapping on the user feature vectors and property feature vectors based on feature associations and coupled matching representations to generate difference penalty representations; generating gate vectors based on the user feature vectors and property feature vectors to construct dynamic feature activation relationships; performing gated fusion processing on dimension-wise consistency based on the gate vectors to generate consistency representations; and performing weighted fusion on the coupled matching representations, difference penalty representations, and consistency representations to generate a matching metric.

[0088] Specifically, the user feature vector and property feature vector are first subjected to unified dimension alignment and component normalization, and then subspace partitioning is performed to establish a common representation basis required for subsequent matching calculations. Unified dimension alignment means mapping the user feature vector and property feature vector to a feature space with the same dimension scale and the same semantic order, so that the components at corresponding positions can form a comparable relationship. Component normalization means compressing feature components from different sources and of different magnitudes into a unified numerical range, avoiding the unbalanced dominant role of components such as price, area, commuting distance, and semantic strength in subsequent matching calculations due to excessive differences in their original numerical ranges. During processing, the user feature vector is first divided into constraint preference subspace, behavioral preference subspace, and potential demand subspace based on the semantic structure of the basic attribute layer, behavioral feature layer, and potential demand layer in the user profiling system. Then, based on the semantic structure of the property image data, text description data, structural attribute data, and spatial location data, the property feature vector is divided into structural attribute subspace, location attribute subspace, supporting attribute subspace, and semantic representation subspace. Mapping relationships between subspaces are then established. For example, a primary association is established between the constraint preference subspace and the structural attribute subspace; a dynamic association is established between the behavioral preference subspace and the semantic representation subspace; and an implicit association is established between the potential demand subspace and the location attribute subspace and supporting attribute subspace. In this embodiment, subspace division refers to dividing the original high-dimensional vector into multiple local representation regions with clear meanings according to business semantics; feature association refers to the predefined or learned correspondence between a certain type of user demand feature and a certain type of property attribute feature. This process avoids performing coarse-grained comparisons at the overall vector level in subsequent similarity calculations, instead enabling fine-grained matching within local subspaces with clear business significance. The dimension-aligned vector can be represented as:

[0089]

[0090] Where u represents the original user feature vector; v represents the original property feature vector; This represents the dimension-aligned and normalized user feature vector; This represents the dimensionally aligned and normalized feature vector of the property; and These represent the dimensionality mapping and normalization functions for the user side and the property side, respectively, used to transform heterogeneous features into a unified feature space representation. The purpose of this formula is to provide a unified input basis for subsequent bilinear interaction, difference penalty, and gating fusion.

[0091] After completing the unified dimension alignment and subspace partitioning, bilinear interaction mapping is performed based on user feature vectors and property feature vectors to generate a coupled matching representation. Bilinear interaction mapping means that instead of treating user feature components and property feature components as simple, independent correspondences, it explicitly characterizes the cross-coupling strength between any user component and any property component through a learnable interaction matrix. Since there are usually non-linear coupling relationships between budget preference and price attributes, family structure and apartment size, commuting tolerance and location accessibility, and amenities preference and surrounding facility strength in property recommendation scenarios, it is difficult to reflect these interaction semantics using only cosine similarity or Euclidean distance. During processing, bilinear interaction matrices are first constructed in each partitioned subspace, and then the interaction results of each subspace are concatenated or weighted to form an overall coupled matching representation. In this embodiment, the coupled matching representation refers to the joint response result of user needs and property attributes in the cross-feature space; the higher the value, the more consistent the two are in the key semantic correspondence. The expression for bilinear interaction mapping is:

[0092]

[0093] in, This represents the coupling matching characteristic; This represents the dimension-aligned and normalized user feature vector; This represents the dimensionally aligned and normalized feature vector of the property; This represents the bilinear interaction weight matrix, used to describe the pairwise interaction strength between user feature components and property feature components. In this formula, if a user demand component and a property attribute component have a strong positive matching relationship, the corresponding interaction weight will increase during training, thereby enhancing the contribution of this dimension's cross-interaction to the overall coupled matching representation; if the interaction is weak, the corresponding weight will be reduced. The coupled matching representation obtained in this way can more accurately reflect the deep correspondence between user needs and property attributes than a simple inner product.

[0094] After obtaining the coupling matching representation, based on the feature association and coupling matching representation, a dimension-by-dimensional difference calculation is performed on the user feature vector and the property feature vector, and a weighted mapping is applied to generate a difference penalty representation. Dimension-by-dimensional difference calculation refers to calculating the degree of deviation for each corresponding or associated dimension component of the user feature vector and the property feature vector under the constraints of the established feature association, thereby identifying conflicting dimensions that significantly reduce the acceptability of the recommendation. Because property recommendations do not necessarily achieve high quality simply by having some dimensions highly matched, if the budget is severely exceeded, commuting time significantly exceeds the tolerance threshold, the apartment size is significantly insufficient, or the school district requirements are clearly not met, then even if other dimensions have high coupling, a strong penalty should be applied to the final matching measurement result. In processing, a difference selection mask is first generated based on the feature association, then the absolute difference or squared difference between the user feature vector and the property feature vector in the corresponding dimensions is extracted, and then the key conflicting dimensions are highlighted through a weighted mapping matrix. In this embodiment, the difference penalty representation refers to the centralized quantitative result of the degree of inconsistency between user needs and property attributes. The expression for the difference penalty representation is:

[0095]

[0096] in, represents the difference penalty representation; q represents the penalty aggregation vector, used to weight and summarize different difference components; This represents the difference mapping matrix, used to amplify the penalty contribution of key conflict dimensions; This represents the one-dimensional absolute difference vector between the user feature vector and the property feature vector after feature association mapping; This represents the feature association mapping matrix, used to rearrange or project property feature vectors into the association dimension space corresponding to user feature vectors. The purpose of this formula is to explicitly extract key deviations that may be masked in the coupled matching representation and transform them into suppression terms in the subsequent fusion process, ensuring that the final matching measurement result does not ignore significant conflict dimensions due to a few highly coupled dimensions.

[0097] After the difference penalty representation is constructed, a gating vector is generated based on the user feature vector and the property feature vector to construct a dynamic feature activation relationship. The gating vector is a control vector that adaptively determines whether each feature dimension should be strengthened or weakened in subsequent consistency calculations based on the current user's needs and the current property's attributes. A gating vector is needed because the same user's priorities change at different stages. For example, during a budget-constrained period, price and commuting are prioritized; during family changes, apartment layout and school district are prioritized; and during improved living conditions, lighting, quality, and community amenities are prioritized. Therefore, the influence of consistency or difference in the same dimension on the final matching result is not fixed under different states. During processing, the user feature vector and the property feature vector are input into the gating generation network, which can output the activation intensity corresponding to each dimension using a linear mapping combined with an activation function. In this embodiment, the dynamic feature activation relationship refers to the real-time participation intensity relationship of each feature dimension in the final matching calculation under the current user and property combination state. The expression for the gating vector is:

[0098]

[0099] Where g represents the gate vector; This represents the activation function, used to compress the result of a linear mapping to a range between zero and one. This represents the user feature mapping matrix, used to extract the user-side impact on gating; This represents the property feature mapping matrix, used to extract the impact of property-side access control on gate control; This represents the gating bias vector. The principle behind this formula is to map user needs and property attributes together into a set of dimensional weights. When the gating value of a certain dimension is close to one, it indicates that this dimension should be given priority consideration in the current matching scenario; when the gating value is close to zero, it indicates that even if there are similarities or differences in this dimension, it should not have a significant impact on the final matching result. This mechanism allows the matching calculation to have scenario-adaptive capabilities.

[0100] After obtaining the gating vector, gating fusion processing is performed on the dimension-wise consistency based on the gating vector to generate a consistency representation. Dimension-wise consistency refers to the closeness between the user feature vector and the property feature vector in corresponding or related dimensions, reflecting their matching level in each local semantic dimension. Gating fusion processing means that the consistency contribution of different dimensions is dynamically adjusted by the gating vector when calculating consistency. During processing, the dimension-wise consistency strength between the user feature vector and the property feature vector is first calculated, for example, by subtracting the absolute difference, using local similarity mapping, or local kernel functions to obtain basic consistency. Then, the gating vector is applied to the basic consistency dimension by dimension. Finally, the results of each dimension are aggregated to obtain the overall consistency representation. In this embodiment, the consistency representation refers to the quantitative result of the overall consistency between user needs and property attributes under the current dynamic feature activation relationship. Its expression is:

[0101]

[0102] in, This represents a consistent representation; represents the transpose of a single vector, used to sum and aggregate dimension-wise gated consistency results; g represents the gate vector; This represents element-wise multiplication. This represents the dimension-wise difference between the user feature vector and the property feature vector in the correlation dimension; This represents the strength of dimension-wise consistency. The principle behind this formula is that if a dimension has a small difference and a large gate value, then the contribution of that dimension to the consistent representation will be significantly increased; if a dimension has a small difference but a low gate value, then that dimension will only produce a limited contribution; if a dimension has a large difference, then its consistency itself is low, and even a large gate value will hardly improve its contribution. This allows the consistent representation to reflect both the degree of dimension-wise proximity and the importance of dimensions in the current scenario.

[0103] After the coupling matching representation, difference penalty representation, and consistency representation are all generated, a weighted fusion is performed on the three to generate the final matching measurement result. Weighted fusion refers to unifying the coupling matching representation (reflecting positive cross-coupling relationships), the difference penalty representation (reflecting the degree of key conflict), and the consistency representation (reflecting the degree of scene adaptation consistency) into the same scoring framework, thereby forming a matching measurement result that combines deep semantic coupling capabilities, key conflict suppression capabilities, and dynamic attention adaptation capabilities. In this embodiment, the matching measurement result refers to the comprehensive matching score of the user feature vector and the property feature vector under the current model structure. The higher the value, the more suitable the current property is as a recommendation for the user. Its expression is:

[0104]

[0105] Where S represents the matching metric result; This represents the coupling matching characteristic; This represents the differential punishment characteristic; This represents a consistent representation; The fusion weight represents the coupling matching representation and is used to control the contribution of deep interaction matching relationships to the final result. The fusion weights representing the difference penalty representation are used to control the inhibitory strength of the key conflict dimensions on the final result; The fusion weight, representing the consistency representation, is used to control the degree to which dynamic consistency enhances the final result. The principle behind this formula is to simultaneously incorporate the positive correspondence, negative conflict, and dynamic consistency relationships between user needs and property attributes into the calculation, so that the final matching measurement result is no longer a single-directional similarity measure, but a composite scoring result with multi-layered semantic interpretation capabilities.

[0106] To facilitate the direct integration of this matching metric into the ranking, thresholding, and reinforcement learning decision-making processes of the subsequent property matching and recommendation model, the matching metric can be normalized and compressed to stably map to a finite interval. The normalized matching metric can be represented as:

[0107]

[0108] in, The expression represents the normalized matching metric result; S represents the original matching metric result after weighted fusion. The purpose of this formula is to compress the originally unbounded matching score to between zero and one, making the matching results between different users, different properties, and different recommendation times more numerically stable and comparable. Subsequently, in the property matching recommendation model, the normalized matching metric result, along with action value, candidate property timeliness status, property freshness, and exploration probability, can be input into the ranking chain, thus enabling the matching metric result to truly participate in the generation process of the final property recommendation sequence. After this processing, from unified dimension alignment, subspace partitioning, bilinear interaction, difference penalty, gating generation, consistency fusion to the final matching metric output, the entire process is streamlined.

[0109] Furthermore, the cross-regional candidate housing dynamic migration recommendation scenario addresses situations where target users have phased migration needs across different cities or regions. By jointly modeling the trajectory features, behavioral evolution features, and potential demand features in the user feature vector, it identifies the dynamic migration process of users from regional exploration to regional convergence and then to stable migration, and constructs a regional migration preference distribution based on this. At the same time, it performs unified encoding and regional correction processing on housing feature vectors from multiple regions, and introduces cross-regional distance constraints to characterize commuting accessibility and migration costs, thereby establishing a comparable unified expression space among housing sets in different regions. Furthermore, it constructs a recommendation state expression based on user feature vectors, regional migration preference distribution, housing feature vectors, and cross-regional distance constraints, and performs hierarchical decision-making at the regional and housing layers through deep reinforcement learning, enabling housing recommendation results to dynamically adjust with user migration stages and real-time behavioral feedback. This achieves continuous evolution of cross-regional housing recommendations from multi-regional exploration to target region-centric recommendations, improving the accuracy and adaptability of recommendation results in complex migration scenarios.

[0110] In one possible implementation, generating a property recommendation sequence based on the matching metric results further includes: performing migration intensity modeling on the cross-regional behavioral evolution sequence based on user feature vectors to generate a regional migration preference distribution, wherein multi-source user data from multiple regions is acquired to construct a cross-regional migration user profile system, thereby generating user feature vectors; introducing cross-regional distance constraints and correcting the property feature vectors to form a cross-regional candidate property set, wherein multi-source property data from multiple regions is acquired and uniformly standardized to construct a cross-regional property set, and uniform feature encoding and regional correction processing are performed on the cross-regional property set to generate property feature vectors; constructing a recommendation state expression based on user feature vectors, regional migration preference distribution, property feature vectors, and cross-regional distance constraints; mapping the recommendation state expression to a Markov decision process, and performing hierarchical decision-making based on deep reinforcement learning to generate corrected matching metric results; and generating a cross-regional property recommendation sequence based on the corrected matching metric results.

[0111] Specifically, a cross-regional migration user profile system is first constructed based on cross-regional multi-source user data. This system enables user feature vectors to express not only the static residential preferences of target users within a single region but also their phased migration trends across multiple regions. The cross-regional multi-source user data includes basic attribute data, historical browsing data, consultation text data, regional filtering records, candidate work location data, stay area data, location trajectory data, budget change data, and family constraint change data. Basic attribute data characterizes stable constraints such as budget range, rental / purchase preferences, family structure, and move-in timeframe. Historical browsing data characterizes the distribution and intensity of target users' attention across different regions. Consultation text data characterizes the migration intentions and regional preferences actively expressed by target users. Candidate work location data and location trajectory data characterize cross-regional commuting relationships and spatial activity patterns. Budget change data and family constraint change data characterize the dynamic changes in constraints during the migration process. The cross-regional migration user profile system is a hierarchical user expression structure constructed for cross-regional housing selection scenarios. It not only describes who the user is but also which region the user might migrate to, when they might migrate, and which regional attributes they focus on at different migration stages. By performing identity unification, time unification, regional coding unification, abnormal trajectory cleaning, semantic normalization, and feature fusion processing on the above multi-source data, user feature vectors with cross-regional migration semantics can be generated, providing a unified expression basis for subsequent regional migration preference distribution modeling.

[0112] After obtaining the user feature vector, migration intensity modeling is performed on the cross-regional behavior evolution sequence based on the user feature vector to generate a regional migration preference distribution. The cross-regional behavior evolution sequence refers to the time-ordered sequence of browsing, collection, consultation, filtering, and dwell behaviors of target users around different regions within multiple time windows. This sequence reflects the migration process of target users from initial regional attention to candidate region convergence and then to target region focus. Migration intensity modeling refers to the comprehensive quantification of the target user's attention strength, dwell depth, consultation intention, commuting suitability, and budget suitability in different candidate regions, thereby obtaining the migration probability distribution of users towards multiple candidate regions. If a target user continuously browses listings in a certain region, frequently collects listings in that region, and the name of that region appears multiple times in consultation text, and the commuting relationship between that region and the candidate workplace is relatively good, then the weight of that region in the regional migration preference distribution will significantly increase; if a region is only briefly browsed and no sustained attention is subsequently formed, the migration weight of that region will gradually decrease. The expression for the regional migration preference distribution is:

[0113]

[0114] in, This represents the probability of a target user's migration preference for the j-th candidate region, with a value ranging from zero to one, and the sum of the probabilities of all candidate regions is one. This represents the browsing attention intensity corresponding to the j-th candidate area, which is determined based on the number of views, dwell time, and return frequency. This represents the collection intensity corresponding to the j-th candidate region, which is determined based on the number of collections, the frequency of uncollections, and the duration of collection. This represents the intensity of consultation intention corresponding to the j-th candidate region, which is determined based on whether consultation has occurred, the consultation round, and the relevance of the consultation topic. This represents the spatial correlation strength between the j-th candidate region and the target user's location trajectory, which is determined based on the proximity of the workplace, the degree of overlap of the user's frequently used activity areas, and the path repetition. This indicates the fit strength between the price range of the j-th candidate region and the budget range of the target user; This represents the overall migration cost for a target user to migrate to the j-th candidate region, which is determined based on geographical distance, commuting time, and the complexity of switching living circles. to The weight parameters for each factor are denoted by M, which is trained using historical migration samples or set by business rules. M represents the total number of candidate regions. This formula obtains the regional migration preference distribution that can be used for subsequent cross-regional recommendations by exponentially normalizing the target user's attention behavior, intention behavior, trajectory relationship, and constraint adaptation relationship in multiple regions.

[0115] After generating the regional migration preference distribution, multi-source housing data from multiple regions is acquired and subjected to unified standardization processing to construct a cross-regional housing data set. This multi-source housing data includes structural attribute data, semi-structured tag data, text description data, image data, location data, transportation data, educational and medical facility data, and housing status data from different cities, districts, or functional areas. Unified standardization processing maps housing data from different regions, sources, and expression standards to a unified field system and value system, thereby eliminating differences in price units, area definitions, regional naming methods, granularity of supporting tags, and housing type expression rules between cross-regional data. For example, price fields in different regions may use different expressions such as rent, total price, monthly payment, and unit price; after unified standardization, these need to be mapped to comparable price expressions. Similarly, school tags, commercial tags, and rail transit tags from different regions also need to be mapped to a unified supporting tag structure. The cross-regional housing data set refers to the collection of housing data constructed after the above unified standardization processing, which can be processed together in the same retrieval and modeling process. This process ensures that the generation of subsequent housing feature vectors is no longer affected by differences in data structure between regions, thus providing a data foundation for cross-regional comparison and ranking.

[0116] After constructing the cross-regional housing data set, unified feature encoding and regional correction are performed on the set to generate housing feature vectors. Unified feature encoding involves performing visual semantic extraction on housing image data, textual semantic encoding on textual description data, category embedding and continuous value mapping on structural attribute data, and location embedding and supporting facility embedding on spatial location and supporting facility data. This transforms information from different types of housing into vector representations within a unified feature space. Regional correction considers the inherent differences between different regions in price centers, housing types, supply and demand density, transportation conditions, and living costs. It applies regional offset correction to the original housing features, making similar housing comparable across different regions. For example, housing of the same size and type may have drastically different price levels and transportation value in core urban areas and peripheral urban areas; regional correction can explicitly separate these regional baseline differences from the housing features. The expression for the housing feature vector is:

[0117]

[0118] in, This represents the feature vector of the i-th property. This represents the original property representation of the i-th property after image encoding, text encoding, structural mapping, and location embedding. This represents the global feature mapping matrix, used to map the original property representations to a unified feature space; This represents the regional correction vector of the area to which the i-th property belongs, used to characterize the regional price baseline, supply and demand density, accessibility level and supporting facilities baseline; This represents the area correction mapping matrix, used to map area correction vectors to the property feature space; This represents the bias vector, used to correct the overall feature distribution. The purpose of this formula is to ensure that each property feature vector simultaneously carries the property's own attribute information and the baseline information of its region, thereby supporting the subsequent construction and comparison of cross-regional candidate property sets.

[0119] After generating the property feature vectors, cross-regional distance constraints are introduced and the feature vectors are corrected to form a cross-regional candidate property set. Cross-regional distance constraints refer to a mechanism that jointly constrains the physical distance, commuting time, migration stage deviation, and living circle switching complexity between candidate properties and the target user's current spatial state in cross-regional recommendation scenarios. The reason for introducing cross-regional distance constraints is that matching solely based on user preferences and property attributes may lead to the erroneous promotion of properties that, while highly similar in attributes, are extremely far away, have unacceptable commutes, or are at unsuitable migration times. In processing, a cross-regional distance constraint value is first calculated for each property based on the target user's current main activity area, candidate workplace, current migration stage, and regional migration preference distribution. This constraint value is then used to correct the accessibility of the original property feature vectors, suppressing properties with high inaccessibility in subsequent rankings. The expression for the cross-regional distance constraint is:

[0120]

[0121] in, This represents the cross-regional distance constraint value corresponding to the i-th property. This represents the geographical distance from the target user's current main activity area to the area where the i-th property is located; The maximum reference value for geographic distance, used for normalization; This represents the commuting time for the target user from their current primary activity area or candidate workplace to the i-th property. Indicates the maximum tolerable commute time; This indicates the degree of stage deviation between the area where the i-th property is located and the current migration stage of the target user; This represents the complexity of a target user switching from their existing living area to the living area of ​​the i-th property. to This represents the weight parameters of each constraint term. The property feature vector can then be modified as follows:

[0122]

[0123] in, This represents the modified feature vector of the i-th property. This represents the original feature vector of the i-th property. This represents the constraint correction mapping coefficient, used to map cross-regional distance constraint values ​​to the property feature space. The purpose of this process is to explicitly integrate cross-regional accessibility and migration feasibility into the property representation, ensuring that the final cross-regional candidate property set not only meets attribute matching requirements but also satisfies practical migration conditions.

[0124] After forming a cross-regional candidate housing set, a recommendation state representation is constructed based on user feature vectors, regional migration preference distributions, housing feature vectors, and cross-regional distance constraints. The recommendation state representation is a unified state vector used to describe the entire cross-regional recommendation environment at the current recommendation time. It simultaneously represents the target user's current demand state, the priority structure of candidate regions, the distribution of candidate housing attributes, and cross-regional accessibility conditions. During construction, migration-stage-related components are first extracted from the user feature vectors to generate migration-stage state vectors. Then, the regional migration preference distribution is expanded to form a regional state matrix. Subsequently, the modified housing feature vectors are grouped, sorted, and organized by region to form regional housing state tensors. Finally, the cross-regional distance constraints are organized into a constraint state matrix and mapped to a unified state space along with the aforementioned state components. The expression for the recommendation state representation is:

[0125]

[0126] Where s represents the recommendation state representation; u represents the user feature vector; P represents the expanded vector of the regional migration preference distribution, which carries the migration preference probability and regional priority structure of each candidate region; V^* represents the expanded representation of the modified housing feature vector set, which carries the attribute distribution of each candidate housing in multiple regions. This represents the expanded vector of the cross-regional distance constraint matrix, used to carry the accessibility constraints of each candidate property. , , and These represent the mapping matrices corresponding to each state component; This represents the state bias vector. The purpose of this formula is to encode user needs, regional migration tendencies, candidate property attributes, and cross-regional constraints into a single state representation, enabling subsequent deep reinforcement learning models to select actions within the same decision space.

[0127] After the recommendation state representation is constructed, it is mapped to a Markov Decision Process (MDP) and hierarchical decision-making is performed based on deep reinforcement learning to generate corrected matching metrics. A Markov Decision Process is a continuous decision-making model consisting of states, actions, rewards, and state transitions. The decision at the current moment depends only on the current state and not directly on the original history from earlier moments. Deep reinforcement learning uses deep neural networks to approximate the action-value function or policy function, thereby learning the optimal decision-making strategy in a high-dimensional state space. Hierarchical decision-making involves breaking down cross-regional recommendation actions into two levels: regional layer decision-making and property layer decision-making. First, it determines which region to recommend first, and then it determines which properties to recommend first within that region. This design is because cross-regional recommendation is not simply a uniform ranking of all properties; rather, it first determines which regions the current user is more likely to migrate to, and then performs fine-grained property selection within the candidate regions. In specific processing, two sub-networks can be set up: a regional decision network and a housing decision network. The regional decision network outputs the regional action value of each candidate region based on the recommendation state expression. The housing decision network outputs the housing action value within the selected region based on the user feature vector, regional state, and modified housing feature vector. Then, the regional action value and the housing action value are fused to obtain the modified matching metric result. The expression for the modified matching metric result is:

[0128]

[0129] in, The corrected matching metric result is represented by u; u represents the user feature vector. This represents the modified feature vector of the i-th property. This represents the inner product of the two in a unified feature space, used to characterize the degree of attribute matching; and These represent the magnitudes of the corresponding vectors, used to eliminate the influence of amplitude; This represents the probability of a target user's migration preference for the area where the i-th property is located; This represents the cross-regional distance constraint value corresponding to the i-th property. , and This represents the weight parameters for attribute matching, regional migration preference, and constraint penalty. This formula incorporates property attribute matching, regional migration tendency, and actual accessibility into a unified scoring framework, thereby obtaining a modified matching metric suitable for cross-regional scenarios.

[0130] After the matching metric results are corrected, a cross-regional property recommendation sequence is generated based on these results. This sequence refers to the property display results presented to the target user at the current recommendation time, sorted by cross-regional migration priority and property suitability. Specifically, candidate regions are first sorted based on regional migration preference distribution and regional action value to determine core recommendation regions, extended recommendation regions, and exploratory recommendation regions. Then, within each region, candidate properties are finely sorted based on the corrected matching metric results, and a quota for each region in the property recommendation sequence is set according to the current migration stage. For example, in the regional exploration stage, the coverage ratio of multiple high-migration-preference regions is increased to help target users compare regions; in the regional convergence stage, the property density in core recommendation regions is increased, and the proportion of properties in irrelevant regions is reduced; in the stable migration stage, highly matched properties in the main target region and adjacent regions are concentratedly output. The final cross-regional property recommendation sequence not only reflects the user's current attribute matching needs but also their cross-regional migration tendency, migration stage, and actual accessibility. Therefore, it is more suitable for cross-regional phased migration scenarios than traditional single-region static sorting results. The entire process involves constructing a cross-regional user profile system, modeling regional migration preferences, organizing cross-regional housing resources, generating housing feature vectors, correcting cross-regional distance constraints, constructing recommendation state representations, modeling Markov decision processes, and outputting hierarchical decision-making and recommendation sequences.

[0131] In one possible implementation, a recommendation state representation is constructed based on user feature vectors, regional migration preference distributions, property feature vectors, and cross-regional distance constraints. Specifically, this includes: performing phase decoupling based on user feature vectors to generate migration phase state vectors; performing expansion and feature supplementation based on regional migration preference distributions to generate regional state matrices; grouping property feature vectors by region and performing local sorting and fixed-length construction to generate regional property state tensors; performing multi-component mapping and regional organization based on cross-regional distance constraints to generate constraint state matrices; performing phase weighting on the regional state matrix based on the migration phase state vectors, and performing regional gating on the regional property state tensors based on the weighting results, combined with constraint decay in the constraint state matrix, to generate an initial recommendation state representation; performing context enhancement on the initial recommendation state representation using session context features and historical feedback trajectory features to generate a final recommendation state representation; and performing state compression and state verification on the recommendation state representation to form a standard state input, thereby completing the construction of the recommendation state representation.

[0132] Specifically, when performing phase decoupling on the user feature vector, the phase discrimination component reflecting the cross-regional migration process is first extracted from the user feature vector, and this component is separated from the general preference component to form an independent migration phase state vector. In this embodiment, phase decoupling refers to separating the stable constraint information, dynamic preference information, and migration process information that were originally mixed in the same user feature vector, so that the migration process information can participate independently in subsequent state construction; the migration phase state vector is a vector expression used to characterize the current cross-regional migration phase of the target user. During processing, phase discrimination features are first constructed around the regional browsing dispersion, regional collection concentration, migration deterministic semantics in the consultation text, candidate work location change stability, budget fluctuation amplitude, and cross-regional stay frequency in the user feature vector, and then the phase discrimination features are mapped to the phase space corresponding to the regional trial state, regional convergence state, and stable migration state. If the target user frequently switches between multiple regions and their consultation topics are scattered, the region exploration state component in the migration phase state vector is high; if the target user's browsing areas begin to concentrate on a few candidate regions and their consultation topics revolve around specific areas, the region convergence state component increases; if the target user continuously browses, favorites, and consults around a single main region and their commuting relationship tends to stabilize, the stable migration state component increases. The expression for the migration phase state vector is:

[0133]

[0134] Where m represents the migration phase state vector; u represents the user feature vector; The stage mapping matrix is ​​used to project user feature vectors onto the migration stage discriminant space; This represents the stage bias vector, used to correct the overall offset of different user groups during the migration stage; This represents the normalization mapping function, used to convert the stage discrimination result into a probability distribution corresponding to multiple migration stages. The function of this formula is to explicitly extract the information reflecting the migration process from the user feature vector and form the standard state input required for weighted calculation in subsequent stages.

[0135] When performing expansion and feature supplementation on the regional migration preference distribution, the original regional migration preference distribution, which exists in the form of regional probabilities, is first converted into a regional state matrix with the ability to describe regional attributes. In this embodiment, expansion refers to organizing the migration preference probabilities of each candidate region into a region-by-region expression according to regional order; feature supplementation refers to adding information such as regional category, commuting correlation strength between the region and the candidate work location, regional cost of living suitability, regional supporting facilities completeness, regional popularity fluctuation degree, and regional historical feedback strength on the basis of the original regional migration preference probabilities; the regional state matrix is ​​a structured representation of the comprehensive state of multiple candidate regions in matrix form. During processing, for each candidate region, the regional migration preference probability, regional level identifier, regional price center suitability, regional transportation accessibility, regional education, medical and commercial supporting facilities completeness, and regional historical interaction strength are extracted. These components are then arranged in the same row, and all candidate regions are stacked sequentially to form a matrix. After this processing, each row corresponds to a candidate region, and each column corresponds to a regional state attribute, preserving both the magnitude of regional preferences and the characteristics of the regions themselves. The expression for the regional state matrix is:

[0136]

[0137] Where R represents the region state matrix; This represents the probability of a target user's migration preference for the j-th candidate region; This represents the cost-of-living adaptation feature of the j-th candidate region; This represents the commuting association feature between the j-th candidate region and the candidate workplace; This represents the completeness feature of the matching elements in the j-th candidate region; This represents the heat fluctuation characteristics of the j-th candidate region; Let represent the feedback intensity feature of the j-th candidate region in historical recommendations; M represents the number of candidate regions. The purpose of this formula is to enable subsequent decision-making links to perceive migration preferences and regional environments on a region-by-region basis, rather than relying solely on a single probability value to complete region judgment.

[0138] When grouping property feature vectors by region and performing local sorting and fixed-length construction, the property feature vectors in the candidate property set are first divided using the region identifier as the key. Then, within each region, local sorting is performed based on local priority. Finally, the properties within each region are organized into a unified-length regional property state tensor. In this embodiment, local sorting refers to sorting properties within a single region based on the degree of local fit between the property and the current user's needs. Fixed-length construction refers to unifying the property sets with inconsistent numbers within different regions into a state subset of the same length for subsequent unified calculation. The regional property state tensor is a three-dimensional state representation organized by region hierarchy and containing the features of candidate properties within each region. During processing, the local priority of property feature vectors belonging to the same region is first calculated based on price fit, unit type fit, commuting fit, supporting facilities fit, property availability, and historical interaction heat. Then, within each region, they are sorted from high to low local priority, and the top-ranked properties are selected. Each property constitutes a subset of the property status for a given area; if the actual number of candidate properties in an area is less than K, placeholder vectors are used to fill the gap. The expression for local priority is:

[0139]

[0140] in, This indicates the local priority of the i-th property within its respective area; Indicates price-matching characteristics; Indicates the characteristics of apartment type compatibility; Indicates commuting adaptation features; Indicates matching and compatibility features; Indicates the time-sensitive status of the property listing; Indicates the characteristics of historical interaction popularity; to This represents the weight parameters of each local ranking factor. After sorting and truncation, the fixed-length housing subsets within all regions are stacked in regional order to form a regional housing state tensor. This ensures that both the regional structure and the housing structure within each region are preserved, and meets the input requirements for subsequent unified state calculations.

[0141] When performing multi-component mapping and regional organization on cross-regional distance constraints, a single constraint value is first decomposed into multiple constraint components with clear business semantics, and then organized into a constraint state matrix according to the correspondence between regions and properties. In this embodiment, multi-component mapping refers to decomposing cross-regional distance constraints into geographical distance constraint components, commuting time constraint components, migration stage deviation constraint components, and living circle switching complexity constraint components; regional organization refers to categorizing and arranging the multi-component constraints corresponding to each property according to the region to which the property belongs; the constraint state matrix is ​​a structured representation of the cross-regional accessibility constraints of candidate properties in matrix form. During processing, for each candidate property, the geographical distance from the target user's current main activity area to the region to which the property belongs, the commuting time from the candidate work location to the property, the stage deviation between the property's region and the current migration stage, and the living circle switching complexity are calculated. These results are then written as constraint vectors into the constraint sub-blocks under the corresponding region. The expression for the constraint state matrix is:

[0142]

[0143] Where C represents the constraint state matrix; This represents the geographic distance constraint component corresponding to the k-th property within the j-th region; Indicates the commuting time constraint component; This indicates the deviation from the constraint components during the migration phase; The matrix represents the complexity constraint component for switching living areas; M represents the number of candidate areas; and K represents the number of properties retained in each area. This matrix can explicitly provide accessibility information for each area and each property during subsequent state fusion, avoiding the recommendation decision relying solely on preference information while ignoring actual migration conditions.

[0144] When performing stage weighting on the region state matrix based on the migration stage state vector, the effective contribution of each candidate region in the current stage is adjusted by first utilizing the magnitude of the different stage components in the migration stage state vector. In this embodiment, stage weighting refers to changing the weight distribution of each row in the region state matrix according to the current migration stage state, so that the region state matrix can dynamically change with the stage the target user is in. If the migration stage state vector shows that the target user is in the region exploration stage, multiple candidate regions are allowed to maintain high activity; if the migration stage state vector shows that the target user is in the region convergence stage, a few high migration preference regions are strengthened and weakly correlated regions are suppressed; if the migration stage state vector shows that the target user is in the stable migration stage, the contribution of non-major regions is further compressed. The expression for stage weighting is:

[0145]

[0146] in, This represents the region state matrix after stage weighting; m represents the migration stage state vector. Let R represent the stage weight matrix generated from the state vector of the migration stage; R represents the original region state matrix. The purpose of this formula is to make the region state matrix no longer a fixed expression, but to form a direct coupling relationship with the current migration stage.

[0147] After the regional state matrix is ​​weighted at each stage, regional gating is applied to the sub-regional property state tensor based on the weighting results, and constraint decay is performed in conjunction with the constraint state matrix to generate an initial recommendation state representation. In this embodiment, regional gating refers to generating regional-level gating weights using the stage-weighted regional state matrix, and then applying these gating weights to the property state components within the corresponding region. Constraint decay refers to compressing the property state intensity using the multi-component constraint values ​​in the constraint state matrix, reducing the contribution of properties with poor accessibility to the state representation. The initial recommendation state representation refers to a representation that integrates information from user migration stages, regional preferences, property attributes, and cross-regional constraints before incorporating session context and historical feedback. During processing, the regional gating coefficient for each region is first calculated based on the stage-weighted regional state matrix, and then multiplied by the regional gating coefficient into the sub-regional property state tensor. Afterward, a constraint decay factor is introduced for each property to weaken property state components with excessive geographical distance, high commuting time, significant stage deviation, or excessive complexity in switching living areas. The expressions after regional gating and constraint decay are as follows:

[0148]

[0149] in, This represents the state vector of the k-th property in the j-th region after gating and constraint attenuation. This represents the region gating coefficient for the j-th region; This represents the original property status vector of the k-th property within the j-th region; This represents the constraint attenuation coefficient, used to control the intensity of the constraint value's suppression of the property's condition. This represents the comprehensive constraint value of the k-th property within the j-th region. Then, the state vectors of all gated and attenuated properties across all regions are concatenated or mapped with the state vectors from the migration phase and the phase-weighted region state matrix to generate the initial recommendation state representation. This ensures that the initial recommendation state representation includes both positive preference information and real-world constraint information.

[0150] After the initial recommendation state expression is generated, context enhancement is performed on the initial recommendation state expression using session context features and historical feedback trajectory features to generate a final recommendation state expression. In this embodiment, session context features refer to real-time behavioral information occurring in the current recommendation session, including recently browsed areas, recently clicked property types, recently visited property characteristics, recent changes in filter conditions, and recent inquiry topics. Historical feedback trajectory features refer to the continuous feedback patterns generated by the target user on properties in different regions and property types over multiple recommendation periods, including continuous click trends, continuous collection trends, cross-regional interest migration trends, and negative rejection trends. Context enhancement refers to further injecting real-time behavioral and historical evolution information into the initial recommendation state expression, so that the final recommendation state expression not only describes static preferences and static candidates but also describes changes in the decision-making environment at the current point in time. During processing, the session context features and historical feedback trajectory features are first encoded into context vectors, then mapped to the same feature space as the initial recommendation state expression, and finally fused. The expression for the recommendation state expression is:

[0151]

[0152] Where s represents the recommendation status expression; The initial recommendation state is represented by x; the session context feature vector is represented by f; and the historical feedback trajectory feature vector is represented by f. This represents the initial recommendation state mapping matrix; Represents the session context mapping matrix; Represents the historical feedback trajectory mapping matrix; This represents the state bias vector. The purpose of this formula is to inject both short-term context and long-term feedback into the state representation, enabling subsequent decision-making models to simultaneously perceive both current interest shifts and historical interest evolution.

[0153] When performing state compression and state verification on the recommended state representation, the highly redundant and low-information-density components in the recommended state representation are first compressed. Then, the completeness, consistency, and usability of the compressed state result are checked to form a standard state input. In this embodiment, state compression refers to reducing the dimensionality of the recommended state representation using linear dimensionality reduction, nonlinear projection, or gating screening methods to avoid excessive expansion of the state space when the number of candidate regions and candidate properties increases. State verification refers to checking whether there are problems such as missing regions, misaligned region order, abnormal property occupancy ratio, constraint dimension mismatch, context feature mismatch, or numerical out-of-bounds errors in the recommended state representation. Standard state input refers to the state vector that meets the requirements of subsequent Markov decision process modeling and deep reinforcement learning inference. During processing, the recommended state representation can be compressed and mapped first:

[0154]

[0155] in, s represents the compressed standard state input; s represents the recommended state expression before compression. This represents the state compression matrix, used to map high-dimensional recommendation state representations to a compact state space; This represents the compressed bias vector. After compression, the result is checked to see if it still retains migration stage information, regional priority information, property attribute distribution information, and constraint information. If any missing or abnormal information is detected, the process returns to the preceding steps to re-execute regional state matrix correction, property state tensor completion, or constraint state matrix recalculation. If the detection result is normal, the compressed result is written to the state cache as the standard state input for subsequent property matching and recommendation models. This process creates a continuous technical chain for the entire recommendation state representation construction process, from migration stage state vector generation, regional state matrix construction, regional property state tensor generation, constraint state matrix construction, stage weighting, regional gating, constraint decay, context enhancement to state compression and verification. This chain can stably support subsequent decision calculations in cross-regional property recommendation scenarios.

[0156] In one possible implementation, multi-source housing data is obtained through user feature vectors, specifically including: performing semantic decomposition on user feature vectors and constructing rigid acquisition constraints, flexible acquisition constraints, and implicit acquisition constraints; performing regional intent parsing based on user feature vectors and fusing multi-source regional features to generate a target region set and regional acquisition weights; constructing housing acquisition requests based on structured constraint expressions and the target region set, and mapping them to multiple housing data sources; performing rigid constraint filtering on the data returned by multiple housing data sources to form an initial candidate housing set; performing flexible constraint screening on the initial candidate housing set and calculating the flexible constraint deviation to form a flexible candidate housing set; performing implicit demand compensation recall based on the implicit acquisition constraints in user feature vectors to generate an extended candidate housing set; performing cross-source deduplication and consistency correction on the extended candidate housing set to form a unified candidate housing set; and performing dynamic quantity control and regional quota control on the unified candidate housing set based on user feature vectors to generate multi-source housing data.

[0157] Specifically, when performing semantic decomposition on user feature vectors, the stable constraint information, dynamic preference information, and implicit demand information that coexist in the user feature vectors are first separated and mapped to rigid acquisition constraints, flexible acquisition constraints, and implicit acquisition constraints, respectively. In this embodiment, semantic decomposition refers to performing hierarchical inversion and rule reconstruction on user feature vectors according to different roles in the housing acquisition chain, so that the vector expression is transformed back into a constraint expression with clear business meaning. Rigid acquisition constraints refer to constraints that must be prioritized in the housing acquisition process, and if they are not met, the corresponding housing will be directly eliminated. These typically include budget range, rental / purchase type, move-in time limit, housing type range, and core area restrictions. Flexible acquisition constraints refer to constraints that allow candidate opportunities to be retained within a certain deviation range. These typically include floor preference, decoration preference, orientation preference, commuting tolerance, and supporting facilities preference intensity. Implicit acquisition constraints refer to constraints that are not directly expressed through explicit screening conditions, but are indirectly inferred through consultation text, browsing path, collection behavior, and trajectory relationships. These typically include school district preference, quiet living preference, complete living circle preference, low commuting burden preference, and improved living preference. During processing, stable components are first extracted from the basic attribute layer of the user profile system to generate rigid acquisition constraints; then, fluctuating preference components are extracted from the behavioral feature layer to generate flexible acquisition constraints; finally, semantic intent components are extracted from the latent demand layer to generate implicit acquisition constraints. The structured constraint expression can be represented as follows:

[0158]

[0159] Where Q represents the structured constraint expression; This represents a rigid set of acquisition constraints, used to bear inviolable constraints such as budget, rental / purchase type, housing type range, occupancy period, and core area restrictions; This represents a set of flexible acquisition constraints, used to accommodate deviable constraints such as commuting tolerance, acceptable floor level, acceptable decoration, acceptable orientation, and preferred amenities. This represents a set of implicit acquisition constraints, used to carry implicit demand conditions inferred from text semantics, behavioral paths, and location trajectories. The purpose of this formula is to transform abstract user feature vectors into a structured set of constraints that can directly participate in the housing acquisition process.

[0160] When performing region intent parsing based on user feature vectors, explicit region components, behavioral region components, semantic region components, and trajectory region components that can represent region preferences are first extracted from the user feature vectors. These region components are then fused into a target region set and a region acquisition weight. In this embodiment, region intent parsing refers to identifying the set of regions that the target user is most likely to be interested in, compare, or migrate to; the target region set refers to the set of regions allowed to enter the candidate recall range within the current property acquisition cycle; and the region acquisition weight is a quantitative coefficient representing the priority of different target regions in the property acquisition process. During processing, the user's explicitly set region filtering conditions are first extracted to form explicit region components; then, the distribution of browsing history, favorites history, and consultation history in each region is statistically analyzed to form behavioral region components; next, region entity words, transportation node words, and community-related words are identified from consultation text and search expressions to form semantic region components; finally, trajectory region components are formed based on the workplace, frequently visited activity areas, and frequently stayed areas in the location trajectory. After fusing the above region components, a target region set is generated, and a region acquisition weight is calculated for each region. The expression for the region acquisition weight is:

[0161]

[0162] in, This indicates the weight to be acquired for region r. This indicates the strength of explicit region preference, determined based on the region filtering conditions directly set by the user. This indicates the intensity of attention to the area corresponding to the browsing behavior, determined based on the number of views, dwell time, and frequency of return visits; It represents the intensity of regional attention in text semantics, and is determined based on the frequency of regional entity occurrence, the degree of contextual relevance, and the intensity of intention expression; This indicates the spatial correlation strength between the location trajectory and region r, and is determined based on workplace proximity, dwell frequency, and path repetition. This indicates the intensity of regional preference corresponding to the collection behavior, which is determined based on the proportion of collected properties in region r and the duration of the collection. to This parameter represents the weighting of information from different regions. Its purpose is to unify the explicit region representation, behavioral region representation, semantic region representation, and trajectory region representation of the target user into a unified region acquisition priority, thereby supporting the construction of subsequent property acquisition requests.

[0163] When constructing a property acquisition request based on structured constraints and a target region set, rigid acquisition constraints, flexible acquisition constraints, implicit acquisition constraints, and the target region set are first transformed into a unified set of request fields. This unified set of request fields is then mapped to the local field systems of multiple property data sources. In this embodiment, a property acquisition request refers to a standardized request object that initiates retrieval, filtering, and supplementary recall against property data sources. Multiple property data sources refer to the basic property database, real-time property listing stream, external cooperative property interfaces, historical transaction property index, surrounding amenities database, transportation database, education and medical amenities database, and property image and text resource database. During processing, required fields, deviable fields, and compensation recall fields are first determined based on the structured constraints. Then, the main recall region, extended recall region, and exploratory recall region are determined based on the target region set. Finally, a unified property acquisition request is constructed. Because different property data sources have different field naming methods, field granularity, and encoding rules, a mapping rule from unified fields to local fields needs to be established. For example, the price range in the unified field needs to be mapped to the rent or total price field in different systems; the area code in the unified field needs to be mapped to the city, district, or business district field in different systems; and the room type range in the unified field needs to be mapped to the room or apartment template field in different systems. This process ensures that the same set of user constraints can trigger a consistent recall across multiple housing data sources.

[0164] When performing rigid constraint filtering on data returned from multiple housing data sources, a first round of strong filtering is performed on all returned housings based on rigid acquisition constraints to form an initial candidate housing set. In this embodiment, rigid constraint filtering refers to directly judging the inviolable conditions such as budget range, rental / purchase type, room type range, occupancy period, core area restrictions, and housing availability status. Housings that do not meet the conditions will not enter the subsequent processing chain. The initial candidate housing set refers to the first-level candidate housing set retained after passing the rigid constraint filtering. During processing, each housing is checked item by item to see if the price falls within the budget range, if the transaction type matches the rental / purchase preference, if the room type falls within the allowed range, if the occupancy period meets the occupancy period limit, if the housing area is in the main recall area or allowed area of ​​the target area set, and if the housing status is still in a valid listing state. Housings with missing key fields, significantly out-of-bounds prices, mismatched areas, or invalid status are directly eliminated. The rigid constraint filtering function can be expressed as:

[0165]

[0166] in, This indicates the result of whether the i-th property has passed the rigid constraint filtering; This indicates an indicator function that takes the value of one when the condition is true, and zero otherwise. This indicates that the i-th property is in the i-th position. The condition is a satisfaction flag under rigid constraints; K represents the total number of rigid constraint terms. If the i-th property does not meet the condition under any rigid constraint term, the corresponding product result is zero, and the property is eliminated. The purpose of this formula is to transform multiple inviolable constraints into a binary judgment result of pass or elimination, thereby quickly narrowing down the range of candidate properties.

[0167] When performing flexible constraint screening and calculating the flexible constraint deviation on the initial candidate housing set, the housing units that passed the rigid constraint filtering are first subjected to item-by-item deviation measurement on the flexible acquisition constraint dimension. Then, the degree of deviation determines whether the housing unit enters the next level of the candidate set. In this embodiment, flexible constraint screening refers to a secondary screening based on the proximity between the housing unit and the user's flexible preferences without violating rigid constraints. The flexible constraint deviation refers to the comprehensive degree of deviation of the housing unit from the user's flexible preferences in dimensions such as commuting time, floor level, decoration, orientation, and amenities. During processing, for each housing unit, the differences between commuting time and the user's tolerance threshold, the floor level and the user's acceptable range, the decoration status and the user's preference level, the orientation attribute and the user's orientation preference, and the surrounding amenities and the user's amenities preference are calculated separately. These differences are then weighted and summarized. The expression for the flexible constraint deviation is:

[0168]

[0169] in, This represents the deviation of the elastic constraint corresponding to the i-th property. This represents the commuting deviation component, determined based on the difference between the commuting time from the property to the workplace or high-frequency activity area and the user's tolerance threshold. This indicates the floor deviation component, determined based on the difference between the available floor and the user's acceptable floor range; This indicates the deviation from the standard decoration rating, determined based on the difference between the property's current decoration status and the user's decoration preferences. This indicates the orientation deviation component, determined based on the difference between the property's orientation and the user's orientation preference; This indicates the deviation component, determined based on the difference between surrounding transportation, commercial, educational, and medical facilities and user preferences; to This represents the weighting parameter of each flexible constraint component. The purpose of this formula is to quantify the deviation conditions, preventing the premature elimination of high-quality boundary properties due to the use of only hard thresholds. Subsequently, properties with smaller deviations can be retained based on preset deviation thresholds or sorting truncation rules, forming a flexible candidate property set.

[0170] When performing implicit demand compensation recall based on implicit acquisition constraints in user feature vectors, the process first extracts demand directions that are not covered by explicit screening conditions but persist in behavior and semantics from the user feature vectors. Then, compensation recall is performed on multiple housing data sources around these demand directions to generate an expanded candidate housing set. In this embodiment, implicit demand compensation recall refers to additionally recalling housings that, while not necessarily appearing in the explicit screening conditions, are actually highly matched with the user's deep-seated needs, based on the explicit screening results. The expanded candidate housing set refers to a wider range of candidate sets composed of the flexible candidate housing set and the compensation recall housings. During processing, implicit demand tags are first identified based on consultation text, search expression, browsing dwell patterns, collection patterns, and trajectory relationships. Examples include tags for quiet living, low commuting burden, convenient school district, improved housing type, complete living circle, and value preservation tendency. Subsequently, housings with high correlation to these implicit demand tags are retrieved from the surrounding amenities information database, historical transaction housing index, text resource database, and external cooperation interfaces. The compensation recall results are then incorporated into the flexible candidate housing set. The implicit demand correlation score can be represented as:

[0171]

[0172] in, This represents the association score between the i-th property and the implicit acquisition constraint; This indicates the strength of the semantic association between the property listing text description and the implicit demand tags; This indicates the strength of the correlation between property attributes and historical patterns of high dwell time and high collection frequency; This indicates the strength of the correlation between the location of a property and its location trajectory; This indicates the strength of a property's ability to support implicit needs in specific living scenarios; to This represents the weighting parameter of each implicit demand component. The purpose of this formula is to provide a quantitative basis for compensation recall, enabling implicit demands to directly participate in the housing acquisition process, rather than only being indirectly reflected in the subsequent ranking stage.

[0173] When performing cross-source deduplication and consistency correction on the expanded candidate property list set, duplicate records from different property data sources that actually point to the same property entity are first identified. Then, conflicting fields in the duplicate records are corrected for consistency, forming a unified candidate property list set. In this embodiment, cross-source deduplication refers to merging the same property record that may exist simultaneously in the basic property database, the real-time listing stream, and the external cooperative property interface, based on information such as location, building, unit type, area, image fingerprint, and text similarity. Consistency correction refers to uniformly selecting or correcting conflicting fields based on source credibility, update time, and field completeness when multiple sources provide different prices, areas, decoration status, building age, listing time, or amenity descriptions for the same property. The unified candidate property list set is a standard property list set that can directly enter the subsequent feature extraction and matching chain after deduplication and correction. During processing, property entity fingerprints can be constructed first, then clustered and merged based on the entity fingerprints. Subsequently, confidence scores are performed on the fields within each cluster, retaining the field value with the highest confidence, and supplementing missing fields from other sources. The expression for field confidence is:

[0174]

[0175] in, This represents the field confidence score of the i-th property in field j; This represents the source credibility component, used to reflect the reliability of the data source to which the field belongs; This represents the update time component, used to reflect the freshness of the field data; The completeness component reflects whether the field and related fields exist together in completeness. This represents the consistency component, used to reflect the degree of consistency between this field and other source fields; to This represents the weighting parameter of each confidence component. The purpose of this formula is to provide a unified basis for adjudicating multi-source conflict fields, ensuring that the output unified candidate housing set possesses stable, complete, and comparable data quality.

[0176] When performing dynamic quantity control and regional quota control on a unified candidate housing set based on user feature vectors, the total number of candidate housings is first determined according to the stability of current user demand, regional migration tendency, budget rigidity, and exploration needs. Then, the quota ratio of different regions in the output results is determined according to the target region set and region acquisition weight, ultimately generating multi-source housing data. In this embodiment, dynamic quantity control means that the number of candidate housings is not fixed and output uniformly, but rather the candidate scale is dynamically determined based on user status and recall quality. Regional quota control means that different numbers of candidate seats are allocated to different regions according to their importance, exploration degree, and accessibility constraints. Multi-source housing data refers to the standardized housing data set output after constraint parsing, region parsing, multi-source recall, rigid filtering, elastic screening, compensated recall, cross-source deduplication, and quota control. During processing, if user regional intent is highly concentrated, budget range is narrow, and behavioral preferences are stable, the total number of candidates is reduced and the quota for core regions is increased; if user regional intent is dispersed, migration trend is obvious, and recent behavior fluctuates greatly, the total number of candidates is increased and the quotas for extended and explored regions are increased. The expression for regional quotas is:

[0177]

[0178] in, This represents the number of candidate properties allocated to region r; This indicates the total number of candidate properties within the current acquisition period; This indicates the weight to be acquired for region r. This represents the sum of weights obtained from all target regions; This represents the regional adjustment coefficient, used to adjust quotas based on the region's role. Its value increases when the region is a core region, and decreases moderately or increases at specific stages when the region is an exploration region. This represents the number of target areas. The purpose of this formula is to ensure that the quantity and regional distribution of candidate properties simultaneously conform to the user's current state, rather than using a static acquisition method with a fixed number and fixed regional proportions. Through this processing, the final multi-source property data maintains consistency with the user's feature vector in content, with the target area set and regional acquisition weights in regional distribution, and with the complexity of the current recommendation task in scale. This allows it to directly enter the subsequent property feature extraction and property matching recommendation model chain.

[0179] This embodiment also discloses a housing resource processing device based on user profiling and deep reinforcement learning, referring to... Figure 2 The device includes an acquisition module 201, a processing module 202, and an output module 203. It is used to execute any of the above-described methods for processing housing resources based on user profiles and deep reinforcement learning, wherein: The acquisition module 201 is used to acquire multi-source user data corresponding to the target user and perform preprocessing to build a user profile system, thereby generating user feature vectors; Processing module 202 is used to obtain multi-source housing data through user feature vectors; The processing module 202 is used to perform feature extraction and enhancement processing on multi-source housing data to generate housing feature vectors; The processing module 202 is used to construct a housing matching recommendation model based on user feature vectors and housing feature vectors, and to perform joint modeling through deep reinforcement learning to output the matching metric results between user feature vectors and housing feature vectors; Output module 203 is used to generate a housing recommendation sequence based on the matching metric results.

[0180] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0181] This embodiment also discloses an electronic device, as shown in the reference. Figure 3 The electronic device may include: at least one processor 301, at least one communication bus 302, user interface 303, network interface 304, and at least one memory 305.

[0182] The communication bus 302 is used to enable communication between these components.

[0183] The user interface 303 may include a display screen and a camera. Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.

[0184] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0185] The processor 301 may include one or more processing cores. The processor 301 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 305, and by calling data stored in memory 305. Optionally, the processor 301 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 301 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 301 and may be implemented as a separate chip.

[0186] The memory 305 may include random access memory (RAM) or read-only memory. Optionally, the memory may include a non-transitory computer-readable storage medium. The memory 305 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 305 may also be at least one storage device located remotely from the aforementioned processor 301. As a computer storage medium, the memory 305 may include an operating system, a network communication module, a user interface 303 module, and an application program for a housing resource processing method based on user profiling and deep reinforcement learning.

[0187] exist Figure 3In the electronic device shown, the user interface 303 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 301 can be used to call an application stored in the memory 305 that is a housing resource processing method based on user profile and deep reinforcement learning. When executed by one or more processors 301, the electronic device executes one or more methods as described in the above embodiments.

[0188] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, as some steps can be performed in other orders or simultaneously according to the present invention. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0189] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0190] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between apparatuses or units may be electrical or other forms.

[0191] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0192] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0193] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 305 and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned memory 305 includes various media capable of storing program code, such as a USB flash drive, external hard drive, magnetic disk, or optical disk.

[0194] The present invention also discloses a non-transitory computer-readable storage medium storing instructions. When executed by one or more processors 301, these instructions cause an electronic device to perform one or more methods as described in the above embodiments.

[0195] The above are merely exemplary embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of other embodiments of this disclosure upon considering the specification and the disclosure of practical truths. This invention is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.

Claims

1. A method for processing housing listings based on user profiling and deep reinforcement learning, characterized in that, The method includes: Acquire multi-source user data corresponding to the target user and perform preprocessing to build a user profile system, thereby generating user feature vectors; Multi-source housing data is obtained through the user feature vector; Feature extraction and enhancement processing are performed on the multi-source housing data to generate housing feature vectors; A property matching recommendation model is constructed based on the user feature vector and the property feature vector, and joint modeling is performed through deep reinforcement learning to output the matching metric result between the user feature vector and the property feature vector; A property recommendation sequence is generated based on the matching metric results.

2. The housing resource processing method based on user profiling and deep reinforcement learning according to claim 1, characterized in that, After constructing the property matching recommendation model based on the user feature vector and the property feature vector, the method further includes: The user feature vector and the property feature vector are constructed as the state in a Markov decision process, and the action value is output through a deep Q network. The user profile system is constructed based on the basic attribute layer, behavioral feature layer and potential demand layer to generate the user feature vector. The property image data, text description data, structural attribute data and spatial location data are encoded and weighted and fused through a convolutional neural network and attention mechanism to generate the property feature vector. The candidate properties are ranked based on the value of the actions, and a property recommendation sequence is generated. A multi-dimensional reward function is constructed based on user interaction feedback, and the parameters of the housing matching recommendation model are updated by combining the experience replay mechanism and the target network update mechanism. The property recommendation sequence is adjusted by performing offline training based on historical data and adjusting based on real-time interactive data using an ε-greedy strategy.

3. The housing resource processing method based on user profiling and deep reinforcement learning according to claim 1, characterized in that, After generating the property recommendation sequence based on the matching metric results, the method further includes: Based on the property recommendation sequence, a recommendation snapshot record is constructed, and a mapping relationship between recommendation identifiers and user interaction feedback is established; The user interaction feedback is subjected to event standardization and time decay aggregation processing to generate a target feedback vector; Perform a weighted mapping process on the target feedback vector to generate a comprehensive reward value; Based on the comprehensive reward value, a state transition sample is constructed, and the user feature vector is incrementally updated in combination with the user interaction feedback to generate the next state; The state transition samples are written into the experience sample pool and hierarchical storage and priority management are performed. Training batches are extracted from the experience sample pool and the target action value is calculated in combination with the target network to update the parameters of the housing matching recommendation model; Backpropagation is performed based on the loss between the target action value and the current action value to optimize the property matching recommendation model.

4. The housing resource processing method based on user profiling and deep reinforcement learning according to claim 1, characterized in that, The step of constructing a property matching recommendation model based on the user feature vector and the property feature vector, and performing joint modeling through deep reinforcement learning to output a matching metric result between the user feature vector and the property feature vector, specifically includes: For the user feature vector and the property feature vector, perform unified dimension alignment and component normalization processing, and perform subspace partitioning to establish feature association relationships; Bilinear interactive mapping is performed based on the user feature vector and the property feature vector to generate a coupled matching representation; Based on the feature association relationship and the coupling matching representation, a dimension-wise difference calculation and weighted mapping are performed on the user feature vector and the housing feature vector to generate a difference penalty representation. Based on the user feature vector and the property feature vector, a gating vector is generated to construct a dynamic feature activation relationship; Based on the gated vector, gating fusion processing is performed on the dimension-wise consistency to generate a consistency representation; A weighted fusion is performed on the coupling matching representation, the difference penalty representation, and the consistency representation to generate the matching metric result.

5. The housing resource processing method based on user profiling and deep reinforcement learning according to claim 1, characterized in that, The process of generating a property recommendation sequence based on the matching metric results further includes: Based on the user feature vector, migration intensity modeling is performed on the cross-regional behavior evolution sequence to generate regional migration preference distribution. In this process, cross-regional multi-source user data is acquired to construct a cross-regional migration user profile system, thereby generating the user feature vector. A cross-regional distance constraint is introduced and the property feature vector is corrected to form a cross-regional candidate property set. In this process, multi-source property data from multiple regions are acquired and uniform standardization processing is performed to construct the cross-regional property set. The cross-regional property set is then subjected to uniform feature encoding and regional correction processing to generate the property feature vector. Based on the user feature vector, the regional migration preference distribution, the property feature vector, and the cross-regional distance constraint, a recommendation state representation is constructed; The recommended state representation is mapped to a Markov decision process, and hierarchical decision-making is performed based on deep reinforcement learning to generate a modified matching metric result; Based on the corrected matching metric results, a cross-regional housing recommendation sequence is generated.

6. The housing resource processing method based on user profiling and deep reinforcement learning according to claim 5, characterized in that, The step of constructing a recommendation state representation based on the user feature vector, the regional migration preference distribution, the property feature vector, and the cross-regional distance constraint specifically includes: Based on the user feature vector, the execution phase decoupling is performed to generate the migration phase state vector; Based on the aforementioned regional migration preference distribution, expansion and feature supplementation are performed to generate a regional state matrix; Based on the property feature vectors, the properties are grouped by region and local sorting and fixed-length construction are performed to generate regional property state tensors. Based on the cross-regional distance constraint, multi-component mapping and regional organization are performed to generate a constraint state matrix; Based on the migration stage state vector, stage weighting is performed on the regional state matrix, and regional gating is performed on the sub-regional housing state tensor based on the weighting result. In combination with the constraint state matrix, constraint decay is performed to generate an initial recommendation state expression. Context enhancement is performed on the initial recommendation state expression using session context features and historical feedback trajectory features to generate a recommendation state expression; The recommended state expression is then compressed and validated to form a standard state input, thereby completing the construction of the recommended state expression.

7. The housing resource processing method based on user profiling and deep reinforcement learning according to claim 1, characterized in that, The process of obtaining multi-source housing data through the user feature vector specifically includes: Semantic decomposition is performed on the user feature vector, and rigid acquisition constraints, flexible acquisition constraints, and implicit acquisition constraints are constructed. Based on the user feature vector, perform region intent parsing and fuse multi-source region features to generate a target region set and region acquisition weights; A housing retrieval request is constructed based on the structured constraint expression and the target region set, and mapped to multiple housing data sources; Rigid constraint filtering is performed on the data returned by multiple housing data sources to form an initial candidate housing set; The initial candidate housing set is subjected to flexible constraint screening and the flexible constraint deviation is calculated to form a flexible candidate housing set; Implicit demand compensation recall is performed based on the implicit acquisition constraints in the user feature vector to generate an expanded candidate housing set; Cross-source deduplication and consistency correction are performed on the expanded candidate housing set to form a unified candidate housing set; Based on the user feature vector, dynamic quantity control and regional quota control are performed on the unified candidate housing set to generate the multi-source housing data.

8. A housing information processing device based on user profiling and deep reinforcement learning, characterized in that, The apparatus is used to execute a housing resource processing method based on user profiling and deep reinforcement learning as described in any one of claims 1-7. The apparatus includes an acquisition module, a processing module, and an output module, wherein: The acquisition module is used to acquire multi-source user data corresponding to the target user and perform preprocessing to construct a user profile system, thereby generating user feature vectors; The processing module is used to obtain multi-source housing data through the user feature vector; The processing module is used to perform feature extraction and enhancement processing on the multi-source housing data to generate housing feature vectors; The processing module is used to construct a housing matching recommendation model based on the user feature vector and the housing feature vector, and to perform joint modeling through deep reinforcement learning to output the matching metric result between the user feature vector and the housing feature vector; The output module is used to generate a property recommendation sequence based on the matching metric results.

9. An electronic device, characterized in that, The device includes a processor, a communication bus, a user interface, a network interface, and a memory. The memory is used to store instructions. The user interface and the network interface are both used to communicate with other devices. The communication bus is used to enable communication between the various components within the electronic device. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method described in any one of claims 1-7.

10. A non-transitory computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1-7.