House renting recommendation method and system based on multi-modal data fusion analysis

Through multimodal data fusion analysis, a new user's rental demand matrix is ​​constructed and the housing matching degree is calculated, which solves the problem of low correlation between recommendation results of new users when the rental recommendation system is cold-started, and realizes personalized and efficient rental recommendations.

CN119919211AActive Publication Date: 2025-05-02YOU DISTRICT LIFE (SHENZHEN) NETWORK TECHNOLOGY CO LTD

Patent Information

Application Number
CN202411976865.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-02
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

When facing new users, the existing rental recommendation system lacks user historical behavior data, resulting in low correlation of recommendation results and cold start problems.

Method used

A rental recommendation method based on multimodal data fusion analysis is adopted. By obtaining the delivery channel information of new users, the user's geographical location and delivery material attribute data, a rental demand matrix is ​​constructed, and the matching degree is calculated based on the housing description data and travel distance, and a personalized housing recommendation list is generated.

Benefits of technology

It effectively alleviates the cold start problem of new users, improves the relevance and personalization of recommendation results, ensures that the recommendation results are highly consistent with user needs, and optimizes the user experience of the rental platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919211A_ABST
    Figure CN119919211A_ABST
Patent Text Reader

Abstract

The invention discloses a house renting recommendation method and system based on multi-modal data fusion analysis, and relates to the technical field of big data analysis. Obtaining delivery channel information, user geographic positions and delivery material attribute data corresponding to the new drainage users; inputting the delivery channel information and the delivery material attribute data into a house renting demand prediction model to construct a house renting demand matrix corresponding to drainage new users; screening a candidate house resource set from a house resource library according to the geographic position of the user; for each candidate house resource, calculating a matching degree relative to the house renting demand matrix according to the house resource description data of the candidate house resource and the corresponding travel distance; and generating a house resource recommendation list for the drainage new users according to a preset number of candidate house resources with the corresponding matching degree ranked in the front. Therefore, through a multi-modal data fusion mode, accurate rental house resource recommendation under the cold start condition is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of big data analysis, and in particular to a rental house recommendation method and system based on multimodal data fusion analysis. Background Art

[0002] With the continuous development of Internet technology and big data applications, rental recommendation systems have gradually become an indispensable tool in modern urban life, improving the rental experience and efficiency by providing users with housing information that meets their needs.

[0003] The current housing rental recommendation system mainly adopts a recommendation method based on user portraits. By collecting users' personal information (such as requirements for houses, such as house type, environment preferences, rental preferences, etc.) and users' historical behavior data, a user "portrait" model is constructed, and personalized recommendations are made based on the portrait.

[0004] However, the cold start problem is one of the major drawbacks of such rental recommendation systems. Specifically, when the system lacks sufficient user behavior data, it cannot effectively make personalized recommendations. In particular, in rental recommendations, the rental preferences of newly registered users are not yet fully understood, resulting in the recommendation system being unable to quickly adapt to the needs of new users. Similarly, when new listings are added to the system, due to the lack of historical user interaction data, the system often cannot accurately predict the potential user groups of the new listings, resulting in poor recommendation results.

[0005] In response to the above problems, the industry has not yet proposed a better technical solution. Summary of the invention

[0006] The present application provides a rental house recommendation method, system, storage medium, computer program product and electronic device based on multimodal data fusion analysis, which is used to at least solve the cold start problem of traditional rental house recommendation systems when facing new users, due to the lack of users' historical behavior data, resulting in low relevance of recommendation results.

[0007] In a first aspect, an embodiment of the present application provides a rental recommendation method based on multimodal data fusion analysis, comprising: when a user trigger operation for placing materials on a rental platform is detected, obtaining the placement channel information, user geographic location and placement material attribute data of the corresponding new user; the attraction material attribute data includes material preference group type and material house description information; inputting the placement channel information and the placement material attribute data into a rental demand prediction model to output multiple rental demand types and corresponding prediction confidences, thereby constructing a rental demand matrix corresponding to the new user; each row of the rental demand matrix represents a rental demand type, and each column represents a corresponding prediction confidence; selecting a candidate house set from a house library according to the user's geographic location, wherein the travel distance between the house geographic location indicated by the candidate house and the user's geographic location is less than a preset threshold; for each of the candidate houses, calculating the matching degree relative to the rental demand matrix according to the house description data and the corresponding travel distance of the candidate house; and generating a house recommendation list for the new user according to a preset number of candidate houses ranked top according to the corresponding matching degree.

[0008] In a second aspect, an embodiment of the present application provides a rental recommendation system based on multimodal data fusion analysis, including: a traffic data acquisition unit, which is used to obtain the delivery channel information, user geographic location and delivery material attribute data of the corresponding new traffic users when a user trigger operation for delivering materials to the rental platform is detected; the delivery material attribute data includes the material preference group type and material house description information; a rental demand prediction unit, which is used to input the delivery channel information and the delivery material attribute data into a rental demand prediction model to output multiple rental demand types and corresponding prediction confidences, thereby constructing a rental demand matrix corresponding to the new traffic users; the Each row of the rental demand matrix represents a type of rental demand, and each column represents a corresponding prediction confidence; a candidate housing source screening unit is used to screen a set of candidate housing sources from a housing source library according to the user's geographical location, wherein the travel distance between the housing source geographical location indicated by the candidate housing source and the user's geographical location is less than a preset threshold; a housing source matching calculation unit is used to calculate the matching degree relative to the rental demand matrix for each of the candidate housing sources according to the housing source description data of the candidate housing source and the corresponding travel distance; a drainage housing source recommendation unit is used to generate a housing source recommendation list for the drainage new user according to a preset number of candidate housing sources ranked top according to the corresponding matching degree.

[0009] According to a third aspect, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the steps of the method for recommending a house to be rented based on multimodal data fusion analysis of any embodiment of the present application.

[0010] In a fourth aspect, an embodiment of the present application provides a storage medium on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the rental recommendation method based on multimodal data fusion analysis of any embodiment of the present application are implemented.

[0011] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the rental recommendation method based on multimodal data fusion analysis of any embodiment of the present application.

[0012] The housing rental recommendation method and system based on multimodal data fusion analysis provided by this application can produce at least the following technical effects:

[0013] (1) By introducing multimodal data such as delivery channel information and lead material attributes (including material preference group type and house description information), an initial rental demand matrix is ​​constructed for newly registered users. By combining the attribute information of delivery materials with the delivery channel information of lead users and applying model prediction technology, the system can preliminarily predict the type of rental demand and its confidence of new users based on the preferences of the target group. Even without the user's historical interaction data, a recommendation list that meets their potential preferences can be generated, thereby effectively alleviating the cold start problem of new users.

[0014] (2) The candidate housing set is filtered according to the user's geographic location to ensure that the recommended results meet the user's travel convenience needs. In addition, the system matches the housing description data of each candidate housing source with the user demand matrix and sorts the candidate housing sources according to the matching calculation results, thereby ensuring that the recommended housing sources are highly consistent with user needs, which can significantly improve the relevance and personalization of the recommended housing sources.

[0015] Through this technical solution, the new user's drainage information is analyzed, and the new user's rental demand is predicted based on the delivery channel information and delivery material attributes of the corresponding drainage material, and quickly matched with each candidate house with a short travel distance, and more accurate recommended houses can be generated without the user's historical interaction data. Therefore, through the multimodal data fusion method, accurate rental recommendations are achieved in cold start conditions, which effectively improves the cold start performance of the recommendation system, ensures that the recommendation results are highly consistent with the needs of new users, and optimizes the user experience of the rental platform. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 A flowchart of an example of a rental recommendation method based on multimodal data fusion analysis according to an embodiment of the present application is shown;

[0018] Figure 2 A schematic diagram showing a structural connection of an example of a generator in a CGAN architecture of a rental housing demand prediction model according to an embodiment of the present application is shown;

[0019] Figure 3 An operation flow chart of an example of matching degree calculation according to an embodiment of the present application is shown;

[0020] Figure 4 A structural block diagram of an example of a rental recommendation system based on multimodal data fusion analysis according to an embodiment of the present application is shown;

[0021] Figure 5 It is a schematic structural diagram of an embodiment of an electronic device of the present application. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0023] In the technical solution of this application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved shall comply with the provisions of relevant laws and regulations and shall not violate public order and good morals.

[0024] Figure 1 A flowchart of an example of a rental recommendation method based on multimodal data fusion analysis according to an embodiment of the present application is shown.

[0025] Regarding the executor of the method of the embodiment of the present application, it can be any controller or processor with computing or processing capabilities. Specifically, it can be implemented by a housing recommendation analysis system or a rental platform server. By performing data mining and analysis on the drainage information of new users, integrating and analyzing multiple modal data sources such as delivery channel information, material attribute data, geographic location and housing description data, the recommendation system no longer relies solely on the user's personal portrait and historical behavior data, but can infer the user's potential needs based on multi-level information. Through the fusion analysis of multimodal data, the user's rental needs and preferences can be more comprehensively portrayed, thereby improving the accuracy and applicability of the recommendation.

[0026] In some examples, it may be integrated and configured in an electronic device or terminal by means of software, hardware, or a combination of software and hardware, and the type of the terminal or electronic device may be diverse, such as a mobile phone, a tablet computer, or a desktop computer, etc.

[0027] like Figure 1 As shown, in step S110, when a user triggering operation to release materials on the rental platform is detected, the corresponding release channel information, user geographic location and release material attribute data of the new user are obtained, and the material attribute data includes the material preference group type and material house description information.

[0028] In some implementations, the material listing description information includes at least one of the following: the type, decoration style, price range, and listing highlights of the material listing.

[0029] For example, a house rental platform promotes itself by placing traffic-generating materials on various advertising platforms (for example, short video platforms, search engine platforms, etc.). When a user is interested in a certain traffic-generating material, such as clicking, following, liking, collecting or downloading, and entering the house rental platform system, the relevant information of the traffic-generating material can be analyzed, including the delivery channel information (for example, which advertising channel the traffic is generated from, such as social media, search engines, etc.), the user's geographic location information, and the attribute data of the traffic-generating material.

[0030] It should be noted that each traffic-generating material on the rental platform is generally carefully designed, and often has relatively distinct house characteristics, and can simultaneously record information of multiple dimensions through house labels. For example, house type (such as single room, apartment, whole house, etc.), decoration style (such as modern, simple, retro, etc.), price range (such as 1,000-1,500 yuan / month), and house highlights (such as "direct access to the subway", "nearby supermarkets"). These attribute data provide an initial basis for predicting user needs.

[0031] In step S120, the delivery channel information and delivery material attribute data are input into the rental demand prediction model to output multiple rental demand types and corresponding prediction confidences, thereby constructing a rental demand matrix corresponding to attracting new users.

[0032] Specifically, the system uses the collected delivery channel information and delivery material attribute data as input features and inputs them into the rental demand prediction model. The model can be a multi-classification model based on machine learning or deep learning, which can be trained to identify the possible rental demand types of users.

[0033] It should be understood that the types of rental demand prediction models can be diverse and are not limited in this embodiment. In one example of an embodiment of the present application, the rental demand prediction model can adopt a rule-based model to infer user needs through a series of predefined rules. These rules can be set based on platform historical data and expert experience. For example, if the user comes from a specific delivery channel (such as the "Fashion Home" page of a social platform), he may prefer houses with fashionable decoration styles. In another example of an embodiment of the present application, the rental demand prediction model can also adopt a deep learning model (for example, a convolutional neural network CNN, a long short-term memory network LSTM, etc.) to process large amounts of heterogeneous data and automatically extract high-dimensional features.

[0034] In some embodiments, the rental demand type includes at least one of the following: high cost performance, fine decoration, convenient transportation, complete living facilities, single residence or family residence. Furthermore, the model will predict different rental demand types (such as "high cost performance", "fine decoration", "convenient transportation", etc.), output the prediction confidence of different demand types, and form a rental demand matrix. Each row of the rental demand matrix represents the rental demand type, and each column represents the corresponding prediction confidence. For example, for users who prefer "high cost performance" housing sources, the model may assign a higher confidence to this demand type.

[0035] Here, the rental demand matrix can be regarded as a preliminary portrait of user needs, so that the system can quickly capture new user preferences in the cold start phase. Using the drainage material information that users are exposed to, the system can infer the type of user rental needs. This demand prediction method based on delivery channels and material attributes breaks through the reliance on historical user behavior data and realizes user demand identification in the cold start state. By constructing a demand matrix, the system can accurately express the user's various potential needs in a quantitative way, providing a basis for personalized and accurate recommendations of housing sources.

[0036] In step S130, a candidate housing set is screened from the housing database according to the user's geographical location, wherein the travel distance between the geographical location of the housing indicated by the candidate housing and the geographical location of the user is less than a preset threshold.

[0037] In some embodiments, based on the user's geographic location, houses that meet the user's travel distance requirements are screened out from the housing database to form a candidate housing set. The travel distance can be calculated in real time through a map API (such as Baidu Maps, Amap, etc.) to screen out houses with travel distances less than a threshold. The travel distance threshold can be adjusted according to the platform's usage scenarios. For example, the city center area can be set to 3 kilometers, and the suburbs can be set to 5 kilometers or more, so as to flexibly adapt to the needs of different users.

[0038] In this way, in rental recommendations, the distance factor is the core concern of users, and the screening mechanism of the candidate housing set significantly improves the relevance of recommended housing. By performing preliminary screening of candidate housing based on the user's geographic location synchronized by the traffic platform, it ensures that the location of the recommended housing is consistent with the user's living radius, and can also reduce the system resource consumption for comparing and analyzing massive housing sources.

[0039] In step S140, for each candidate house, the matching degree with respect to the housing rental demand matrix is ​​calculated according to the house description data of the candidate house and the corresponding travel distance.

[0040] It should be noted that the matching degree can be calculated in a variety of ways. For example, based on the cosine similarity calculation results or the dot product calculation results, the feature vector of the property description and the demand type characteristics in the user's demand matrix are calculated to evaluate the correlation between the property and the user's demand. No restriction is made here for the time being.

[0041] In some embodiments, for the selected candidate housing sources, the matching degree of the housing source relative to the user's needs is calculated based on the description data of the housing source (such as housing source type, decoration style, price, etc.) and the travel distance from the user's geographical location, combined with the confidence value of each demand type in the rental demand matrix. The matching degree calculation can be implemented by weighted scoring, multiplying each attribute of the housing source with the corresponding demand confidence value of the demand matrix, and superimposing them to obtain the total matching degree score. For example, for users who prefer "high cost performance" and "convenient transportation", housing sources that meet these conditions will be preferred and given a higher matching degree.

[0042] In step S150, a list of recommended properties for attracting new users is generated based on a preset number of candidate properties ranked high in corresponding matching degrees.

[0043] In some embodiments, the candidate housing set is sorted in descending order of matching degree, and a preset number of housings ranked at the top are selected to generate a final housing recommendation list. In addition, the display order of the list can also be determined based on the matching degree score, so as to ensure that the user can see the housing that best meets his needs first when browsing the recommendation list.

[0044] It should be understood that the recommendation list can be presented in different forms, such as list form, card form, map view, etc., to help users intuitively understand the housing information. At the same time, the matching highlights of each housing source and user needs (such as "close distance" and "high cost performance") can be marked in the recommendation list to enhance users' understanding and trust in the recommendation results.

[0045] In the embodiment of the present application, the delivery channel and material attribute information of the user-triggered operation are input into the rental demand prediction model, multiple demand types and their prediction confidence are output, and a demand matrix is ​​constructed to reflect the diverse demand tendencies of new users. This prediction of multiple types of demand not only increases the flexibility of the recommendation, but also improves the reliability of the recommendation, allowing the system to provide more comprehensive recommendation options when the user first contacts it.

[0046] Then, based on the user's geographic location, we filter out candidate properties with travel distances within a preset threshold to ensure that the recommended properties meet the user's commuting or living area requirements. Furthermore, by calculating and sorting the matching degree between the description data of the candidate properties and the demand matrix, we can prioritize the properties that best meet the user's needs. This not only improves the relevance of the recommendation results, but also significantly improves the personalization of the recommendation list, improving the user experience.

[0047] Through the embodiments of the present application, multimodal data is used to construct a preliminary demand portrait of a new user, and the calculation of the geographic location and demand matching degree is introduced into the house recommendation analysis system, which can effectively solve the cold start problem, improve the accuracy of the recommendation results and user satisfaction, and optimize the house recommendation effect of new users in the cold start situation of the rental platform.

[0048] In some examples of the embodiments of the present application, the rental demand prediction model adopts a conditional generative adversarial network (CGAN), which includes a generator and a discriminator, and uses the adversarial training mechanism of the generator and the discriminator to generate high-quality demand prediction results.

[0049] Specifically, the discriminator is used to distinguish the difference between real demand (for example, historical user demand data or expert-annotated demand forecast) and the generated demand type. It helps the generator adjust its generation process through feedback information to make the generated demand forecast more real and accurate. The discriminator generally only participates in the optimization of the generator in the model training phase, but does not operate in the reasoning phase. The structure of the discriminator in the current related technology can be referenced or borrowed, so it will not be described here.

[0050] As the core part of the rental demand prediction model CGAN, the generator is used to generate conditional outputs of rental demand predictions. Specifically, the generator generates a rental demand prediction that matches the given conditional input (such as delivery material attributes, delivery channel information, etc.). Based on these conditional information, the generator generates a possible demand distribution of users, namely the rental demand matrix, such as the user's rental demand type (such as "high cost performance", "fine decoration", etc.) and the prediction confidence of each demand type.

[0051] During the training process of CGAN, the generator and the discriminator conduct adversarial training and optimize each other. The training goal of the generator is to make the demand forecast it generates as close to the real demand data as possible, so as to deceive the discriminator and avoid being identified as false data by the discriminator. During the training process, the generator continuously adjusts its parameters until the difference between the generated demand forecast and the real demand data is minimized. The training goal of the discriminator is to accurately distinguish the generated false data from the real data, helping the generator to gradually improve its generation strategy.

[0052] Through the embodiments of the present application, the CGAN architecture is used to construct a rental demand prediction model. When the user lacks historical behavior data, the conditional generative adversarial network can generate preliminary demand forecasts based on existing external information (such as delivery channels, material attributes, etc.), thereby alleviating the cold start problem.

[0053] Figure 2 A schematic diagram of the structural connection of an example of a generator in the CGAN architecture of the rental demand prediction model according to an embodiment of the present application is shown.

[0054] like Figure 2 As shown, the generator 200 includes an input layer 210, a fusion layer 220, a feature decomposition layer 230 and an output layer 240.

[0055] The input layer 210 is used to embed the input delivery channel information and delivery material attribute data into a unified vector space through multi-modal feature encoding to generate corresponding encoding features of each modality.

[0056] Here, the main function of the input layer 210 is to convert various input data (such as delivery channel information, delivery material attribute data, etc.) into a unified vector space representation through multimodal feature encoding. Exemplarily, discretization encoding (such as one-hot encoding) can be used to convert channel information (such as social media, search engines, advertising platforms, etc.) into vector representation. In addition, for delivery material attribute data, such as material preference group type (such as "students", "office workers") and material house description information (such as "decoration style"), it can be vectorized through the embedding layer, and for numerical data (such as price range), it is standardized; if it is category data (such as house type), one-hot encoding or embedding representation is used, and so on.

[0057] As a result, the encoded features of each modality will be embedded in a unified vector space, which can effectively eliminate the dimensionality differences between different features and enable subsequent processing to be performed based on a unified representation.

[0058] The fusion layer 220 is used to fuse the encoding features of each modality to determine the corresponding multimodal fusion features.

[0059] Here, the fusion layer is used to fuse features from different modes (delivery channel information, material attributes, etc.) to generate multimodal fusion features. It should be understood that the fusion method can be diverse, such as feature splicing, weighted fusion or other methods, which are not limited here.

[0060] The feature decomposition layer 230 is used to perform hierarchical processing on the multimodal fusion features according to the condition information to determine the personalized demand matrix of attracting new users at the coarse-grained and refined demand levels.

[0061] Here, the conditional information is defined based on the delivery channel information, delivery material attribute data, and the preset rental demand type set. The delivery channel information reflects the user's initial interest or demand, and the delivery material attribute data includes the type of material house, price range, and other information, as well as the high cost performance, fine decoration, and convenient transportation in the rental demand type set, to help the model understand the context and level of user needs.

[0062] In some embodiments, first, in the coarse-grained demand hierarchy analysis, a coarse-grained prediction of user demand is generated based on the user's delivery channel information and material attributes. For example, based on the user's click behavior from the "cost-effective housing" advertisement, it is predicted that the user may be more interested in "budget-friendly housing". Then, in the refined demand hierarchy analysis, the detailed attribute information of the delivery material (such as the decoration style of the housing source, the specific price range, etc.) is combined for refined decomposition to obtain the user's demand prediction on certain details. For example, the user's preference for "convenient transportation" may be further refined into the demand for "close to the subway station". Afterwards, the feature decomposition layer outputs a personalized demand matrix, in which each element represents the user's interest level or demand confidence in a certain demand type.

[0063] Through the feature decomposition layer and hierarchical demand decomposition, we can capture the different dimensions of user needs and avoid over-simplification of user needs. Hierarchical processing can provide more fine-grained personalized demand prediction, achieve accurate modeling of user needs, and enhance the depth of understanding of user needs.

[0064] The output layer 240 is used to output multiple rental demand types and corresponding prediction confidences according to the personalized demand matrix.

[0065] In some embodiments, the output layer will output multiple rental demand types and their corresponding prediction confidences based on the decomposed personalized demand matrix, which indicates the user's interest in the demand type. In addition, the output layer can also calculate the probability of each demand type through forward propagation and use it as the corresponding confidence, or calibrate the personalized demand matrix to output a high-precision matching degree for each demand type. A demand type with a high confidence value indicates that the type is more relevant to the user.

[0066] It should be noted that in cGAN, the output of the generator is not necessarily the actual demand type distribution, but through adversarial training with the discriminator, it can generate demand prediction results that are close to the actual distribution.

[0067] Through the generator structure provided in the embodiment of the present application, relying on multimodal data fusion, conditional constraints and information hierarchical processing, the demand prediction distribution for attracting new users is effectively generated, and the authenticity and diversity of the generated effect are improved by utilizing the adversarial training in the CGAN architecture, ultimately achieving a higher demand prediction accuracy when user historical interaction data is scarce.

[0068] In some examples of the embodiments of the present application, the fusion layer 220 adopts a fusion layer based on an adaptive attention mechanism. Specifically, the fusion layer performs an adaptive attention mechanism on each modal feature to identify the different effects of each modality on the user's rental demand. Based on adaptive attention, the model can highlight the features that are more critical to demand prediction in multimodal information.

[0069] Specifically, the fusion layer 220 is used to perform the following operations:

[0070] Calculate the attention weights corresponding to the encoding features of each modality:

[0071]

[0072]

[0073] In the formula, α i represents the normalized attention weight of modality i, is the embedding vector of the i-th modality, N represents the number of modalities, represents the attention weight of modality i, σ att is the activation function used for attention weight calculation; is the projection bias term, is the projection weight matrix, which transforms the feature dimension d of mode i into i Convert to uniform dimension d att ; is the sum of attention weights of all modalities.

[0074] Based on the attention weights, the encoding features of each modality are weighted and fused to obtain weighted multimodal features:

[0075]

[0076] In the formula, Z fusion Represents weighted multimodal features.

[0077] Through the adaptive attention mechanism, the fusion layer can dynamically assign weights to multimodal information according to the specific characteristics of different users, thereby ensuring that during the fusion process, modal features that have a more significant impact on demand prediction are given higher weights.

[0078] In addition, in order to further strengthen the interactive relationship between the features of each modality, a modal interaction mechanism is introduced in the fusion layer so that information from different modalities can influence each other, thereby improving the accuracy of demand forecasting.

[0079] Specifically, a bilinear interaction method is used to generate modal interaction terms to capture the correlation characteristics between different modalities:

[0080]

[0081] In the formula, I ij represents the modal interaction term between modality i and modality j, for The transposed vector of int is the activation function of modal interaction; is the weight matrix of modal interaction, indicating the strength of association between modalities i and j; d int Represents the feature dimension of modal interaction, d fusion Represents Z fusion The characteristic dimension of is the bias term of modal interaction, which is used to adjust the linear combination of inter-modal interaction features.

[0082] By calculating the interaction terms between modalities through the modal interaction mechanism and capturing the correlation information between different modal features, the complementarity of multimodal features can be effectively utilized, so that each modal feature is not only a simple weighted superposition, but also can interact with each other, thereby providing a more comprehensive expression of user needs. For example, user needs may be affected by the attributes of the delivered material and the geographical location at the same time, and the modal interaction mechanism can capture these associations and generate more refined demand forecasts for users.

[0083] The modal interaction terms and weighted multimodal features are concatenated to form the final fused feature representation containing the interaction information:

[0084] Z final =[Z fusion ;I ij ] , Formula (5)

[0085] In the formula, Z final Represents multimodal fusion features.

[0086] Through the weighted fusion of multimodal features and the generation of interaction terms, the multimodal fusion feature Z output by the fusion layer is final It contains personalized information about user preferences, allowing the model to more accurately predict the user's rental demand type and its corresponding preference intensity.

[0087] Through the embodiments of the present application, adaptive attention weight allocation and modal interaction mechanism are adopted in the fusion layer, which improves the expressive ability of the rental demand prediction model in multimodal feature processing, enhances the model's adaptability to personalized needs and generalization ability to diversified user needs, and reduces information redundancy, which helps to improve the prediction accuracy of rental demand.

[0088] In some examples of the embodiments of the present application, the feature decomposition layer 230 is used to perform multi-level computation processing on the multimodal fusion features from coarse granularity to fine granularity. Specifically, it performs the following operations to perform hierarchical processing on the multimodal fusion features:

[0089] Calculate the demand category weight vector for multimodal fusion features:

[0090] γ=softmax(W γ ·Z final +b γ ), Formula (6)

[0091] Where γ represents the demand category weight vector, which indicates the user's preference for each broad demand category. The softmax function is used to normalize each weight to ensure that the sum of the weights of each category is 1. γ and b γ They represent the weight matrix and bias term used to generate the demand category weight vector respectively.

[0092] The demand category weight vector γ is generated based on the user's multimodal fusion features, which can adaptively assign weights to different demand categories, so that the model can dynamically adjust the attention paid to different demand categories based on the user's performance in the delivery channel information and material attribute data. As a result, the model can more accurately capture the user's broad rental demand tendencies (such as high cost performance, convenient transportation, etc.), laying the foundation for personalized recommendations.

[0093] The multimodal fusion features are weighted and aggregated according to the demand category weight vector to generate a coarse-grained demand matrix:

[0094] D coarse =γ⊙(W coarse ·Z final +b coarse ), Formula (7)

[0095] Where ⊙ represents element-wise multiplication, D coarse represents the coarse-grained demand matrix, W coarse and b coarse They represent the weight matrix and bias term used to generate coarse-grained features respectively.

[0096] Here, we use the multimodal features of users to generate a personalized demand weight vector γ to represent the user's preference for each broad demand category, and then generate a coarse-grained demand matrix through dynamic aggregation of user multimodal features rather than a fixed weight matrix. Through the dynamic aggregation mechanism of demand categories, even when there is less user information, we can still extract the initial matching degree of demand from the delivery channels and material attributes, which improves the adaptability of the model to cold-start users.

[0097] Calculate the preliminary matching degree of the delivery channel information and delivery material attribute data with respect to each rental demand type in the rental demand type set, and construct condition information based on each preliminary matching degree:

[0098]

[0099] In the formula, E C is the embedded representation of the delivery channel information, T q represents the embedding representation of the qth requirement type, ||E C || is E C The L2 norm of ||T q || is T q The L2 norm of M C,q is the delivery channel information and the qth demand type T q The initial matching degree indicates the similarity between the delivery channel information and the qth demand type; M M,q To place material attributes and T q The initial matching degree of E represents the similarity between the attributes of the delivered material and the qth demand type; M is the embedded representation of the delivery material attributes, ||E M || is the L2 norm of the delivery material attribute; C cond It is the embedding representation of conditional information, which represents the user's personalized demand preferences; q represents the weighting coefficient of the qth rental demand type, and Q represents the total number of rental demand types in the rental demand type set.

[0100] Here, the cosine similarity is used to calculate the initial matching degree between the delivery channel information and the delivery material attributes and each demand type, thereby generating the conditional information embedding C that reflects the user's preference. cond , ensuring that the model can extract the user's real demand preferences from the delivery information and material attributes, and convert them into conditional information vectors.

[0101] The high-dimensional relationship between coarse-grained features and conditional information is captured through multi-dimensional convolutional layers to generate a personalized demand matrix:

[0102] D input =[D coarse ; C cond ] , Formula (11)

[0103]

[0104] In the formula, represents the personalized demand matrix, which is used to represent the fine-grained demand preferences for attracting new users. h′ and w′ are the height and width of the personalized demand matrix, d fine The characteristic dimension that represents the refined demand characteristics; D input Represents the combined input features of the multi-dimensional convolutional layer; represents a multidimensional convolution kernel, which is used to capture the complex relationship between coarse-grained features and conditional information in the spatial dimension and feature dimension.conv The size in the spatial dimension is k h ×k w , * indicates convolution operation, b conv represents the convolution bias term, σ conv represents the convolutional layer activation function, d input Represents the feature dimension of the combined input features; u and v are the convolution kernels W conv The spatial position index of represents the index of the convolution kernel in the height direction and width direction respectively; Represents the weight matrix of the convolution kernel at position (u,v).

[0105] Through multi-dimensional conditional convolution, the feature decomposition layer can capture the high-dimensional relationship between coarse-grained demand features and condition information, and can convolve coarse-grained features and condition information in spatial and feature dimensions, extract their local patterns, generate more refined user demand expressions, and realize the refined demand matrix D fine In this way, the model can capture the user's preferences for detailed needs such as rental range, decoration style, and transportation convenience at a fine-grained level, and effectively reflect the user's personalized needs in specific scenarios.

[0106] Therefore, the feature decomposition layer realizes the fusion expression of broad demand and fine-grained demand in the rental demand prediction model, which not only improves the accuracy of the model for user needs, but also enhances the adaptability of the model.

[0107] In some examples of the embodiments of the present application, the output layer 240 is used to perform the following operations:

[0108] Based on the global weighted pooling mechanism, the features of each spatial position in the personalized demand matrix are weighted pooled to obtain the weighted pooled demand features:

[0109]

[0110] Where D pool is the weighted pooling demand feature; α u,v is the pooling weight, which represents the feature weight at the spatial position (u, v); is the feature representation of the personalized demand matrix at the spatial position (u, v), W α and b α Respectively represent the weight matrix and bias term used to calculate the pooling weight, exp represents the exponential function; is a normalization term to ensure that all α u,v The sum of is 1.

[0111] Through the global weighted pooling mechanism, the output layer can refine the demand matrix D fineAssign different weights to the features of each spatial position and generate adaptive pooling weights α u,v , adaptively weighted sum the features at different spatial positions, suppress irrelevant information, and enhance the influence of key feature positions.

[0112] The weighted pooling demand features are processed based on the softmax function to obtain the preliminary predicted probability distribution of the demand type:

[0113]

[0114] In the formula, Represents the preliminary predicted probability distribution of the demand type, where the predicted probability value is used to quantify the confidence of the corresponding rental demand type; and Represent the classification weight matrix and classification bias term respectively.

[0115] Here, the predicted probability of the model is directly used as the confidence, and the predicted probability output by softmax naturally reflects the confidence of the model.

[0116] The preliminary predicted probability distribution of the demand type is calibrated based on the conditional information to output multiple rental demand types and corresponding prediction confidences:

[0117]

[0118] In the formula, represents the corrected probability distribution of demand type, and Represent the conditional correction weight matrix and conditional correction bias term respectively; σ sig Represents the sigmoid activation function to ensure that the correction coefficient of the conditional information is in the interval [0,1].

[0119] Here, after obtaining the prediction probability, a dynamic correction mechanism based on conditional information is introduced to further improve the personalized ability of prediction. cond Dynamically adjust the predicted probability of the demand type, and supervise and adjust the predicted probability based on the user's delivery channel information and material attributes, so that the final output of the model is more in line with the user's actual needs.

[0120] It should be noted that in the cold start scenario, the user's historical behavior data may be relatively small, and it is difficult for the model to accurately grasp the user's preferences. Through the global weighted pooling mechanism and conditional information correction, the existing distribution channel information and material attributes can be reused for supervision. The pooling mechanism ensures the effective aggregation of key features, and the conditional correction can infer potential preferences through information such as the user's registration path and the advertising materials seen.

[0121] Therefore, the global weighted pooling mechanism and the supervised correction mechanism based on conditional information are combined to make the prediction results output by the model more in line with the personalized needs of users. Whether in the cold start user scenario or the diversified user preference scenario, the model can analyze the needs and preferences of new users based on the drainage information.

[0122] It should be noted that the generator and discriminator of the rental demand prediction model based on the cGAN architecture are trained alternately according to the data sample set. Specifically, during the training process, the generator is first fixed and the discriminator is trained to enable it to better distinguish between real data and generated data and reduce the discriminator loss; then the discriminator is fixed and the generator is trained to enable the data it generates to be more able to "cheat" the discriminator and reduce the generator loss.

[0123] In the adversarial network of the embodiment of the present application, the task of the discriminator is to judge the difference between the generated data (demand forecast distribution generated based on conditional information) and the real data. To achieve this goal, the discriminator can receive both the generated data and the conditional information C at the input. cond , so as to more effectively identify whether the generated data is reasonable.

[0124] Specifically, the loss function of the discriminator is expressed as follows:

[0125]

[0126] In the formula, S is the number of samples in the data sample set, s is the sample index; x (s) represents the input feature of the sth sample, Is a generator based on x (s) and condition information The generated demand forecast distribution, represents the conditional information of the sth sample, Indicates the unique hot encoding of the real rental demand type label of the sth sample; is the discriminator's response to the true demand label y (s) and condition information The output probability of , with an expected value of 1, indicates that the requirement type is true, so that the discriminator can learn to establish the correct association between real data and conditional information; is the output of the discriminator to the generator and condition information The output probability of , with an expected value of 0, indicates that the demand type is generated data, so that the discriminator can learn to distinguish generated data (under the same conditional information) from real data.

[0127] Here, by introducing the conditional information C cond,The discriminator can detect whether the generated demand distribution meets the requirements of the conditional information, and then provide effective feedback to the generator, helping the generator to better utilize the conditional information to generate personalized solutions that meet user needs in subsequent iterations.

[0128] In cGAN, the input of the generator G contains not only the user’s input data x (i.e., the user’s feature data), but also the conditional information C cond In this way, the generator controls the generated demand type prediction distribution through conditional information, so that the generated results are more in line with the actual needs of users.

[0129] The loss function of the generator is expressed as follows:

[0130]

[0131] In the formula, represents the loss function of the generator, It represents the loss term of the generator's classification of the rental demand type of the sth sample in the data sample set, which is used to measure the difference between the generated rental demand type prediction and the true label; represents the adversarial loss term for the sth sample in the data sample set, indicating the possibility that the generated data is judged as real by the discriminator; λ 1 and λ 2 Respectively represent the corresponding loss term hyperparameters; Represents a generator based on x (s) and condition information The generated predicted probability for rental demand type l.

[0132] Here, the generator loss jointly optimizes the classification loss and the adversarial loss to ensure that the demand distribution generated by the generator not only conforms to the real distribution of user demand, but also confuses the discriminator, making it difficult to distinguish the generated data. In this framework, the conditional information C cond Embedded in the input of the generator, it directly affects the category distribution and feature expression of the generated results, making the generated demand distribution meet the specific conditional requirements of the user. In addition, the generator generates personalized demand prediction distribution under the guidance of conditional information through classification loss and adversarial loss.

[0133] Specifically, in each round, the generator is fixed and the discriminator loss is optimized using the real data and the data generated by the generator. Maximizing this loss improves the discriminator’s ability to distinguish between real and generated data. In addition, in each round, the discriminator is fixed and the generator loss is minimized. The data generated by the generator can "fool" the discriminator, while making the demand type prediction as close to the true label as possible.

[0134] Through the adversarial loss of the generator and the discriminator, the model can continuously optimize the generation results during the adversarial training process, so that the generated data distribution gradually approaches the real data distribution. Ultimately, the generator is prompted to learn the real demand distribution characteristics, thereby reflecting higher authenticity in the generated rental demand forecast. By embedding conditional information into the input of the generator and the discriminator, and using the conditional information for guidance in the adversarial training of the generator and the discriminator, the model can dynamically adjust the generated demand forecast according to the user's specific background (such as delivery channels, material attributes, etc.), and generate a more personalized demand forecast distribution.

[0135] Figure 3 An operational flow chart of an example of matching degree calculation according to an embodiment of the present application is shown.

[0136] like Figure 3 As shown, in step S310, the house description data and the corresponding travel distance are concatenated to generate a corresponding house feature vector.

[0137] Specifically, when constructing the property feature vector, travel distance is used as a specific feature dimension and integrated into the same vector together with the property description feature. Travel distance is a numerical value that can be converted to the same scale through normalization or other data standardization methods so that it can be combined with other features.

[0138] F prop =[f 1 ,f 2 ,…,f G-1 ,d distance ] , Formula (21)

[0139] In the formula, f 1 ,f 2 ,…,f G-1 Represents the information characteristics of each dimension in the property description data, d distance represents the normalized feature of travel distance, F prop represents the property feature vector;

[0140] In step S320, a bilinear pooling matching operation is performed on the house source feature vector and the rental demand matrix to obtain a matching result matrix through interactive feature analysis.

[0141] Through bilinear pooling, we can establish multi-level second-order interaction relationships between rental demand characteristics and housing characteristics, and capture the complex associations between different demand types and housing characteristics. Compared with simple linear weighting, this method can more accurately reflect the user's focus on specific needs. In addition, by giving each feature interaction a different weight, we can more intelligently handle the correlation between different demand types and housing characteristics, and improve the personalized matching effect.

[0142] Here, travel distance is an important component of the property feature vector, and we can give it a higher weight to highlight its influence in the matching calculation. This can be achieved by adjusting the weights in the bilinear pooling. For example, we can adjust the last dimension of the feature vector containing the travel distance (i.e., d distance ) are given higher weight.

[0143]

[0144] In the formula, represents the prediction confidence of the qth rental demand type in the rental demand matrix; F prop is the property feature vector, containing G-dimensional features; represents the g-th eigenvalue in the property feature vector, where the eigenvalue of the last dimension represents the travel distance feature; W (q,g) represents the weight parameter of bilinear pooling, which is used to control the interaction between the rental demand characteristics and the housing source feature vector; d distance represents the normalized travel distance feature; W (q,G) Represents the weight parameter of the travel distance feature, which is used to control the importance of the travel distance feature in the matching calculation; μ is the additional weight coefficient of the travel distance feature, and η is the hyperparameter for adjusting the change rate of the travel distance weight; P match Represents the matching result matrix; C qg Score the correlation between demand type q and the g-th eigenvalue in the property feature vector; is a normalized term, which represents the sum of the correlation scores of all demand and house feature vectors in the rental demand matrix.

[0145] In the above formula, by normalizing the correlation scores of demand characteristics and property characteristics, the contributions of different demand types and property characteristics can be better balanced, making the matching degree calculation more stable and avoiding the excessive influence of individual characteristics on the results.

[0146] It should be noted that by weighting the distance between the user's location and the listing location as an independent feature, the weight of the travel distance is dynamically adjusted according to the distance, so that listings with shorter distances have higher matching degrees. This can achieve a sensitive response to the user's travel needs and ensure that the system gives priority to the user's geographical location needs when recommending listings.

[0147] In addition, the above two types of hyperparameters for travel distance can be customized according to the needs of the scene on the one hand, and the weight of the travel distance feature can be automatically optimized during model training on the other hand, so that the model can adaptively adjust the importance of travel distance in the matching calculation according to the needs and preferences of different users, thereby achieving personalized recommendations.

[0148] In step S330, the matching result matrix is ​​projected by means of feature compression and linear projection to obtain the matching degree.

[0149] S match =σ mat (W proj vec(P match )+b proj ), Formula (26)

[0150] In the formula, S match represents the matching degree of candidate housing, σ mat represents the nonlinear activation function, W proj is the projection matrix, which is used to map the high-dimensional matching matrix to the low-dimensional space; vec(·) represents the operation of expanding the matrix into a vector; b proj Represents the bias vector.

[0151] Here, by adopting feature compression and linear projection, the high-dimensional matching matrix can be mapped to a low-dimensional space, thereby greatly reducing the amount of calculation and effectively retaining the core information between multi-dimensional features. At the same time, the speed of matching calculation is improved, ensuring real-time responsiveness in large-scale recommendation scenarios and meeting the real-time requirements of high-concurrency recommendation systems.

[0152] In the embodiment of the present application, the travel distance is explicitly included in the property feature vector. By assigning a higher weight, the travel distance feature is highlighted in the matching calculation, making the recommendation result more in line with the user's geographical needs. Furthermore, bilinear pooling is used to capture all second-order interactions between user needs and property features. As part of the property features, the travel distance feature interacts with the demand matrix in multiple ways, which can more carefully reflect the correlation between the user's travel preferences and other needs. Projecting on the high-dimensional matching matrix generated by bilinear pooling maps the high-dimensional information to the low-dimensional matching score, which not only retains rich interaction information but also maintains computational efficiency.

[0153] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of actions combined, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application. In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0154] Figure 4A structural block diagram of an example of a house rental recommendation system based on multimodal data fusion analysis according to an embodiment of the present application is shown.

[0155] like Figure 4 As shown, the house rental recommendation system 400 based on multimodal data fusion analysis includes a traffic data acquisition unit 410, a house rental demand prediction unit 420, a candidate house source screening unit 430, a house source matching calculation unit 440 and a traffic house source recommendation unit 450.

[0156] The traffic data acquisition unit 410 is used to obtain the distribution channel information, user geographic location and distribution material attribute data of the corresponding new traffic users when a user trigger operation to distribute materials on the rental platform is detected; the traffic material attribute data includes the material preference group type and material house description information.

[0157] The rental demand prediction unit 420 is used to input the delivery channel information and the delivery material attribute data into the rental demand prediction model to output multiple rental demand types and corresponding prediction confidences, thereby constructing a rental demand matrix corresponding to the new users; each row of the rental demand matrix represents a rental demand type, and each column represents a corresponding prediction confidence.

[0158] The candidate listing screening unit 430 is configured to screen a set of candidate listings from a listing database according to the user's geographic location, wherein the travel distance between the listing geographic location indicated by the candidate listing and the user's geographic location is less than a preset threshold.

[0159] The housing source matching calculation unit 440 is used to calculate the matching degree of each candidate housing source relative to the housing rental demand matrix according to the housing source description data of the candidate housing source and the corresponding travel distance.

[0160] The house recommendation unit 450 is used to generate a house recommendation list for the new user to be attracted according to a preset number of candidate houses ranked high in corresponding matching degrees.

[0161] In some embodiments, an embodiment of the present application provides a non-volatile computer-readable storage medium, which stores one or more programs including execution instructions, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute the steps of any of the above-mentioned methods for renting recommendations based on multimodal data fusion analysis in the present application.

[0162] In some embodiments, the embodiments of the present application also provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes any step of the above-mentioned rental recommendation method based on multimodal data fusion analysis.

[0163] In some embodiments, the embodiments of the present application also provide an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of a rental recommendation method based on multimodal data fusion analysis.

[0164] Figure 5 is a schematic diagram of the hardware structure of an electronic device for executing a rental recommendation method based on multimodal data fusion analysis provided by another embodiment of the present application, such as Figure 5 As shown, the device includes:

[0165] One or more processors 510 and memory 520, Figure 5 A processor 510 is taken as an example.

[0166] The device for executing the housing rental recommendation method based on multimodal data fusion analysis may further include: an input device 530 and an output device 540 .

[0167] The processor 510, the memory 520, the input device 530 and the output device 540 may be connected via a bus or other means. Figure 5 The example of connecting through bus is taken in the following.

[0168] The memory 520 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions / modules corresponding to the rental recommendation method based on multimodal data fusion analysis in the embodiment of the present application. The processor 510 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions and modules stored in the memory 520, that is, realizing the rental recommendation method based on multimodal data fusion analysis in the above method embodiment.

[0169] The memory 520 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 520 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 520 may optionally include a memory remotely arranged relative to the processor 510, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0170] The input device 530 may receive input digital or character information and generate signals related to user settings and function control of the electronic device. The output device 540 may include a display device such as a display screen.

[0171] The one or more modules are stored in the memory 520, and when executed by the one or more processors 510, perform the rental recommendation method based on multimodal data fusion analysis in any of the above method embodiments.

[0172] The above-mentioned product can execute the method provided in the embodiment of the present application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of the present application.

[0173] The electronic devices of the embodiments of the present application exist in various forms, including but not limited to:

[0174] (1) Mobile communication equipment: This type of equipment is characterized by having mobile communication functions and its main purpose is to provide voice and data communications. This type of terminal includes: smart phones, multimedia phones, functional phones, and low-end phones.

[0175] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have mobile Internet access features. These terminals include: PDA, MID and UMPC devices, etc.

[0176] (3) Portable entertainment devices: These devices can display and play multimedia content. They include audio and video players, handheld game consoles, e-books, smart toys, and portable car navigation devices.

[0177] (4) Other onboard electronic devices with data interaction functions, such as on-board devices installed in vehicles.

[0178] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0179] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, and of course, by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0180] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A rental recommendation method based on multimodal data fusion analysis, comprising: When a user triggering operation for placing materials on the rental platform is detected, the placement channel information, user geographic location and placement material attribute data of the corresponding new users are obtained; the material attribute data for the new users includes the material preference group type and material house description information; Inputting the delivery channel information and the delivery material attribute data into a rental demand prediction model to output a plurality of rental demand types and corresponding prediction confidences, thereby constructing a rental demand matrix corresponding to the new users attracted; each row of the rental demand matrix represents a rental demand type, and each column represents a corresponding prediction confidence; Filtering a candidate housing set from a housing database according to the user's geographic location, wherein the travel distance between the housing location indicated by the candidate housing and the user's geographic location is less than a preset threshold; For each of the candidate housing sources, calculating a matching degree relative to the housing rental demand matrix according to the housing source description data of the candidate housing source and the corresponding travel distance; A list of recommended houses for attracting new users is generated based on a preset number of candidate houses ranked high in corresponding matching degrees.

2. The method according to claim 1, wherein: The material house description information includes at least one of the following: the type, decoration style, price range and highlights of the material house; the rental demand type includes at least one of the following: high cost performance, fine decoration, convenient transportation, complete living facilities, single residence or family residence.

3. The method according to claim 1, wherein: The rental demand prediction model adopts a conditional generative adversarial network, which includes a generator and a discriminator; wherein the generator includes an input layer, a fusion layer, a feature decomposition layer and an output layer: The input layer is used to embed the input delivery channel information and delivery material attribute data into a unified vector space through multi-modal feature encoding to generate corresponding coding features of each modality; The fusion layer is used to fuse the encoding features of each modality to determine the corresponding multimodal fusion features; The feature decomposition layer is used to perform hierarchical processing on the multimodal fusion features according to condition information to determine the personalized demand matrix of new users at the coarse-grained and refined demand levels; the condition information is defined according to the delivery channel information, the delivery material attribute data and the preset rental demand type set; The output layer is used to output multiple rental demand types and corresponding prediction confidences according to the personalized demand matrix.

4. The method according to claim 3, wherein: The fusion layer adopts a fusion layer based on an adaptive attention mechanism and is used to perform the following operations: Calculate the attention weights corresponding to the encoding features of each modality: In the formula, α i represents the normalized attention weight of modality i, is the embedding vector of the i-th modality, N represents the number of modalities, represents the attention weight of modality i, σ att is the activation function used for attention weight calculation; is the projection bias term, is the projection weight matrix, which transforms the feature dimension d of mode i into i Convert to uniform dimension d att ; is the sum of attention weights of all modalities; Based on the attention weights, the encoding features of each modality are weighted and fused to obtain weighted multimodal features: In the formula, Z fusion represents weighted multimodal features; The bilinear interaction method is used to generate modal interaction terms to capture the correlation characteristics between different modalities: In the formula, I ij represents the modal interaction term between modality i and modality j, for The transposed vector of int is the activation function of modal interaction; is the weight matrix of modal interaction, indicating the strength of association between modalities i and j; d int Represents the feature dimension of modal interaction, d fusion Represents Z fusion The characteristic dimension of is the bias term of modal interaction, which is used to adjust the linear combination of inter-modal interaction features; The modal interaction terms and weighted multimodal features are concatenated to form a fused feature representation containing interactive information: WITH final =[From fusion ;AND ij ], In the formula, Z final Represents multimodal fusion features.

5. The method according to claim 4, wherein: The feature decomposition layer performs hierarchical processing on the multimodal fusion features by performing the following operations: Calculate the demand category weight vector for the multimodal fusion feature: γ=softmax(W γ ·WITH final +b γ ), Where γ represents the demand category weight vector, which indicates the user's preference for each broad demand category. The softmax function is used to normalize each weight to ensure that the sum of the weights of each category is 1. γ and b γ They represent the weight matrix and bias term used to generate the demand category weight vector respectively; The multimodal fusion features are weighted and aggregated according to the demand category weight vector to generate a coarse-grained demand matrix: D coarse =γ⊙(W coarse ·Z final +b coarse ), Where ⊙ represents element-wise multiplication, D coarse represents the coarse-grained demand matrix, W coarse and b coarse Respectively represent the weight matrix and bias term used to generate coarse-grained features; Calculate the preliminary matching degree of the delivery channel information and the delivery material attribute data with respect to each rental demand type in the rental demand type set, and construct condition information according to each preliminary matching degree: In the formula, E C is the embedded representation of the delivery channel information, T q represents the embedding representation of the qth requirement type, ||E C || is E C The L2 norm of ||T q || is T q The L2 norm of M C,q is the delivery channel information and the qth demand type T q The initial matching degree indicates the similarity between the delivery channel information and the qth demand type; M M,q To place material attributes and T q The initial matching degree of E represents the similarity between the attributes of the delivered material and the qth demand type; M is the embedded representation of the delivery material attributes, ||E M || is the L2 norm of the delivery material attribute; C cond Embedding the conditional information to represent the user's personalized demand preferences; β q represents the weighting coefficient of the qth rental demand type, Q represents the total number of rental demand types in the rental demand type set; The high-dimensional relationship between coarse-grained features and conditional information is captured through multi-dimensional convolutional layers to generate a personalized demand matrix: D input =[D coarse ;C cond ], In the formula, represents the personalized demand matrix, which is used to represent the fine-grained demand preferences for attracting new users. h′ and w′ are the height and width of the personalized demand matrix, d fine Characteristic dimensions that represent refined demand characteristics; D input Represents the combined input features of the multi-dimensional convolutional layer; represents a multidimensional convolution kernel, which is used to capture the complex relationship between coarse-grained features and conditional information in the spatial dimension and feature dimension. conv The size in the spatial dimension is k h ×k w , * indicates convolution operation, b conv represents the convolution bias term, σ conv represents the convolutional layer activation function, d input Represents the feature dimension of the combined input features; u and v are the convolution kernels W conv The spatial position index of represents the index of the convolution kernel in the height direction and width direction respectively; Represents the weight matrix of the convolution kernel at position (u,v).

6. The method according to claim 5, wherein: The output layer is used to perform the following operations: Based on the global weighted pooling mechanism, the features of each spatial position in the personalized demand matrix are weighted pooled to obtain the weighted pooled demand features: Where D pool is the weighted pooling demand feature; α u,v is the pooling weight, which represents the feature weight at the spatial position (u, v); is the feature representation of the personalized demand matrix at the spatial position (u, v), W α and b α Respectively represent the weight matrix and bias term used to calculate the pooling weight, exp represents the exponential function; is a normalization term to ensure that all α u,v The sum of is 1; The weighted pooled demand features are processed based on the softmax function to obtain a preliminary predicted probability distribution of the demand type: In the formula, Represents the preliminary predicted probability distribution of the demand type, where the predicted probability value is used to quantify the confidence of the corresponding rental demand type; and Represent the classification weight matrix and classification bias term respectively; The preliminary predicted probability distribution of the demand type is calibrated based on the condition information to output multiple rental demand types and corresponding prediction confidences: In the formula, represents the corrected probability distribution of demand type, and Represent the conditional correction weight matrix and conditional correction bias term respectively; σ sig Represents the sigmoid activation function to ensure that the correction coefficient of the conditional information is in the interval [0,1].

7. The method according to claim 6, wherein: The generator and the discriminator of the rental demand prediction model are alternately trained according to the data sample set; The loss function of the discriminator is expressed as follows: In the formula, S is the number of samples in the data sample set, s is the sample index; x (s) represents the input feature of the sth sample, Is a generator based on x (s) and condition information The generated demand forecast distribution, represents the conditional information of the sth sample, Indicates the unique hot encoding of the real rental demand type label of the sth sample; is the discriminator's response to the true demand label y (s) and condition information The output probability of is 1, indicating that the demand type is true; is the output of the discriminator to the generator and condition information The output probability of is 0, indicating that the demand type is generated data; The loss function of the generator is expressed as follows: In the formula, represents the loss function of the generator, It represents the loss term of the generator's classification of the rental demand type of the sth sample in the data sample set, which is used to measure the difference between the generated rental demand type prediction and the true label; represents the adversarial loss term for the sth sample in the data sample set, indicating the possibility that the generated data is judged as real by the discriminator; λ1 and λ2 represent the corresponding loss term hyperparameters; Represents a generator based on x (s) and condition information The generated predicted probability for rental demand type l.

8. The method according to any one of claims 4 to 7, wherein: Calculating the matching degree with respect to the housing rental demand matrix according to the housing description data of the candidate housing sources and the corresponding travel distances includes: Concatenate the listing description data and the corresponding travel distance to generate the corresponding listing feature vector: F prop =[f1,f2,…,f G-1 ,d distance ], Where f1, f2, …, f G-1 Represents the information characteristics of each dimension in the property description data, d distance represents the normalized feature of travel distance, F prop represents the property feature vector; A bilinear pooling matching operation is performed on the house source feature vector and the rental demand matrix to obtain a matching result matrix through interactive feature analysis: In the formula, represents the prediction confidence of the qth rental demand type in the rental demand matrix; F prop is the property feature vector, containing G-dimensional features; represents the g-th eigenvalue in the property feature vector, where the eigenvalue of the last dimension represents the travel distance feature; W (q,g) represents the weight parameter of bilinear pooling, which is used to control the interaction between the rental demand characteristics and the housing source feature vector; d distance represents the normalized travel distance feature; W (q,G) Represents the weight parameter of the travel distance feature, which is used to control the importance of the travel distance feature in the matching calculation; μ is the additional weight coefficient of the travel distance feature, and η is the hyperparameter for adjusting the change rate of the travel distance weight; P match Represents the matching result matrix; C qg Score the correlation between demand type q and the g-th eigenvalue in the property feature vector; is a normalized term, which represents the sum of the correlation scores of all demand and house feature vectors in the rental demand matrix; The matching result matrix is ​​projected by feature compression and linear projection to obtain the matching degree: S match =s mat (W proj ·vec(P match )+b proj ), In the formula, S match represents the matching degree of candidate housing, σ mat represents the nonlinear activation function, W proj is the projection matrix, which is used to map the high-dimensional matching matrix to the low-dimensional space; vec(·) represents the operation of expanding the matrix into a vector; b proj Represents the bias vector.

9. A house rental recommendation system based on multimodal data fusion analysis, comprising: A traffic data acquisition unit is used to acquire the delivery channel information, user geographic location and delivery material attribute data of the corresponding new traffic users when a user triggering operation of delivering materials to the rental platform is detected; the delivery material attribute data includes the material preference group type and material house description information; A rental demand prediction unit, used for inputting the delivery channel information and the delivery material attribute data into a rental demand prediction model to output a plurality of rental demand types and corresponding prediction confidences, thereby constructing a rental demand matrix corresponding to the new users attracted; each row of the rental demand matrix represents a rental demand type, and each column represents a corresponding prediction confidence; A candidate housing source screening unit, configured to screen a candidate housing source set from a housing source library according to the user's geographical location, wherein the travel distance between the housing source geographical location indicated by the candidate housing source and the user's geographical location is less than a preset threshold; A housing source matching calculation unit, configured to calculate, for each of the candidate housing sources, a matching degree relative to the housing rental demand matrix according to the housing source description data of the candidate housing source and the corresponding travel distance; The house recommendation unit is used to generate a house recommendation list for the new user to be attracted according to a preset number of candidate houses ranked high in corresponding matching degrees.

Citation Information

Patent Citations

  • A device and a method for automatically extracting house resource tags based on word segmentation and multi-mode matching

    CN109739955A

  • Method for constructing recommendation engine based on LBS house renting scene

    CN111161034A

  • House resource recommendation method and device

    CN111383042A

  • Real estate oriented marketing method

    CN112633943A

  • Shared rental place house resource recommendation method, system and device and a storage medium

    CN113034225A

Cited By

  • Multidimensional data intelligent analysis and evaluation system for buildings

    CN121120134A