A method and device for recommending hotel room types

By extracting features from user historical data and hotel room type data and using a multi-task learning model to recommend hotel room type, the problem that the recommendation results in the existing recommendation methods do not meet user expectations, and a recommendation effect that is more in line with user needs is achieved.

CN120011646BActive Publication Date: 2025-06-17SHENZHEN HUOLI TIAN HUI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510480138.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-06-17
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The existing hotel room type recommendation method relies on user browsing and booking history, and lacks in-depth correlation between user-side data and hotel-side data, resulting in the recommendation results not meeting user expectations.

Method used

The user-side text features are obtained by extracting voice data and evaluation text data from user historical data, image data and introduction text data are extracted from hotel room type data to obtain product-side fusion features, and a multi-task learning model is used to train based on these features to generate a recommendation model for hotel room type recommendation.

Benefits of technology

By deeply linking users and hotel room types, the recommendation results are more in line with user needs and enhance the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011646B_ABST
    Figure CN120011646B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and device for recommending hotel room types, belonging to the technical field of intelligent recommendation. The method includes: obtaining user-side text features based on voice data and evaluation text data in user historical data, obtaining product-side fusion features based on hotel room type image data and hotel room type introduction text data, obtaining multiple groups of feature relationship pairs based on the user-side text features and the product-side fusion features, using the multiple groups of feature relationship pairs as a training set to train a multi-task learning model to obtain a recommendation model, and completing the recommendation of hotel room types for users based on the recommendation model. Through the user-side text features and the product-side fusion features, this recommendation method deeply associates users with hotel room types, making the recommended hotel room types more in line with user needs and effectively improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent recommendation, and particularly relates to a method and device for recommending hotel room types. Background Art

[0002] In recent years, with the improvement of living standards and the change of personal consumption concepts, people have paid more and more attention to personalized experiences during travel. Especially the younger generation is more inclined to choose accommodation methods that can provide unique feelings and have social sharing value. Therefore, the hotel room type recommendation system needs to fully consider market demands and provide room type recommendations that meet the needs of different customer groups.

[0003] The existing hotel room type recommendation methods mainly recommend room types that meet the user's preferences by analyzing the user's browsing and reservation history. For example, users who often book double rooms may be more inclined to recommend double room types. However, this existing hotel room type recommendation method only relies on the user's browsing and reservation history, lacks in-depth association between user-side data and hotel-side data, and is prone to the problem that the recommended hotel room types do not meet the user's expectations. Summary of the Invention

[0004] Therefore, the present invention provides a method and device for recommending hotel room types to solve the problem that the existing hotel room type recommendation methods are prone to the problem that the recommended hotel room types do not meet the user's expectations.

[0005] To achieve the above objectives, the present invention adopts the following technical solutions:

[0006] In a first aspect, the present invention provides a method for recommending hotel room types, including:

[0007] Obtaining user-side text features based on user historical data; the user historical data includes voice data and evaluation text data;

[0008] Obtaining product-side fusion features based on hotel room type data; the hotel room type data includes hotel room type image data and hotel room type introduction text data;

[0009] Obtaining multiple groups of feature relationship pairs between the user-side text features and the product-side fusion features based on the user-side text features and the product-side fusion features;

[0010] Using the multiple groups of feature relationship pairs as a training set to train a multi-task learning model to obtain a recommendation model;

[0011] Completing the recommendation of hotel room types for users based on the recommendation model.

[0012] Further, the obtaining product-side fusion features based on hotel room type data includes:

[0013] Extract the image of the ROI feature region of the hotel room type from the hotel room type image data to obtain the product-side image feature;

[0014] Based on the product-side image feature, obtain the product-side image feature to be fused based on the ResNet50 model; the product-side image feature to be fused is a vector feature of the first preset dimension;

[0015] Based on the hotel room type introduction text data, obtain the product-side text feature to be fused based on the XL-Net model; the product-side text feature to be fused is a vector feature of the second preset dimension;

[0016] Through the linear layer processing of the DNN model, convert the product-side image feature to be fused and the product-side text feature to be fused into the same dimension;

[0017] After aligning the dimensions of the converted product-side image feature to be fused and the product-side text feature to be fused, add the corresponding dimension values to obtain the first fusion feature;

[0018] Perform downsampling processing on the first fusion feature using the average pooling layer to obtain the second fusion feature;

[0019] Based on the product-side image feature to be fused, perform downsampling processing through the max pooling layer to obtain the third fusion feature;

[0020] Based on the second fusion feature and the third fusion feature, obtain the product-side fusion feature through the decision layer.

[0021] Further, the hotel room type image data includes the key frame images in the hotel room type introduction video and the promotional images of the hotel room type. The extracting the image of the ROI feature region of the hotel room type from the hotel room type image data to obtain the product-side image feature includes:

[0022] Extract the image of the ROI feature region from the key frame images and the promotional images through faster-RCNN as the product-side image feature.

[0023] Further, the obtaining the user-side text feature based on the user historical data includes:

[0024] Convert the voice data to obtain the first text data;

[0025] Based on the first text data and the evaluation text data, obtain the user-side text feature through the WordPiece method.

[0026] Further, the obtaining multiple groups of feature relationship pairs of the user-side text feature and the product-side fusion feature based on the user-side text feature and the product-side fusion feature includes:

[0027] Input the user-side text features and the product-side fusion features into the ImageBERT model to obtain multiple groups of feature relationship pairs between the user-side text features and the product-side fusion features.

[0028] Further, the multi-task learning model is an MMoE model.

[0029] Further, the hotel room type recommendation for the user based on the recommendation model includes:

[0030] When a hotel room type recommendation needs to be made for the current user, obtain the user's historical data of the current user;

[0031] Obtain the current user-side text features based on the user's historical data of the current user;

[0032] Input the current user-side text features into the recommendation model to obtain the target hotel room type.

[0033] Further, the method further includes:

[0034] Complete the online recall function of hotel room types based on FAISS.

[0035] In a second aspect, the present invention provides a hotel room type recommendation device, including:

[0036] A user-side preprocessing module, configured to obtain user-side text features based on user historical data; the user historical data includes voice data and evaluation text data;

[0037] A product-side preprocessing module, configured to obtain product-side fusion features based on hotel room type data; the hotel room type data includes hotel room type image data and hotel room type introduction text data;

[0038] A feature relationship pair module, configured to obtain multiple groups of feature relationship pairs between the user-side text features and the product-side fusion features based on the user-side text features and the product-side fusion features;

[0039] A training module, configured to use multiple groups of the feature relationship pairs as a training set to train a multi-task learning model to obtain a recommendation model;

[0040] A recommendation module, configured to complete the hotel room type recommendation for the user based on the recommendation model.

[0041] The present invention adopts the above technical solutions and has at least the following beneficial effects:

[0042] A method and device for recommending hotel room types are provided. User-side text features are obtained based on voice data and evaluation text data in user historical data. Product-side fusion features are obtained based on hotel room type image data and hotel room type introduction text data. Multiple groups of feature relationship pairs are obtained based on the user-side text features and the product-side fusion features. The multiple groups of feature relationship pairs are used as a training set to train a multi-task learning model to obtain a recommendation model. Based on the recommendation model, hotel room type recommendations for users are completed. Through the user-side text features and the product-side fusion features, this recommendation method deeply associates users with hotel room types, making the hotel room type recommendations more in line with user needs and effectively improving the user experience.

[0043] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. Brief Description of the Drawings

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 It is a flowchart of a method for recommending hotel room types shown in an exemplary embodiment of the present invention;

[0046] Figure 2 It is a schematic block diagram of a device for recommending hotel room types shown in an exemplary embodiment of the present invention.

[0047] The present invention will be further described below in conjunction with the drawings and specific embodiments. Detailed Embodiments

[0048] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be described in detail below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other implementation manners obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0049] Existing hotel room type recommendation methods mainly recommend room types that match user preferences by analyzing users' browsing and reservation histories. Many potential high-quality user-side, product-side images, and visual features that are not utilized are available. Comprehensively using the information of text and visual patterns can significantly improve the model metrics in the recommendation scenario.

[0050] The technologies represented by the English abbreviations used in this application are as follows: ROI stands for Region of Interest, which is the target region in image segmentation and extraction; the ResNet50 model is a residual convolutional neural network containing 50 residual blocks; the XL-Net model is Transformer with Extra Long-range Net, a long-sequence large language model based on Transformer-XL; the DNN model is Deep Neural Network, a deep neural network; faster-RCNN is Faster Region-based Convolutional Neural Network, a region-based fast object detection and segmentation convolutional neural network model; WordPiece is word piece segmentation, a word segmentation algorithm in natural language processing; the ImageBERT model is a model that can fuse visual and language information based on the BERT large language model; the MMoE model is Multi-gate Mixture-of-Experts, a multi-task recommendation model based on the multi-gate mechanism; FAISS is Facebook AI Similarity Search, a high-efficiency similarity retrieval tool for Facebook AI; Whisper is a whisperer, an automatic speech recognition system based on the Transformer architecture; OpenCV is Open Source Computer Vision Library, an open-source computer vision and machine learning software tool library; Feature Map is a feature map, an intermediate product of the forward calculation of image sample data in a neural network; DecoderLayer is a decoding layer; Input is an input layer; Embedding is an embedded feature; Focal Loss is a focal loss.

[0051] The methods and devices in the present invention will be described below through specific embodiments.

[0052] Please refer to Figure 1 , Figure 1 which is a flowchart of a hotel room type recommendation method shown in an exemplary embodiment of the present invention. Refer to Figure 1 The method includes:

[0053] Step S11: Obtain user-side text features based on user historical data; the user historical data includes voice data and evaluation text data;

[0054] Step S12: Obtain product-side fusion features based on hotel room type data; the hotel room type data includes hotel room type image data and hotel room type introduction text data;

[0055] Step S13: Obtain multiple pairs of feature relationships between the user-side text features and the product-side fusion features based on the user-side text features and the product-side fusion features;

[0056] Step S14: Use the multiple pairs of feature relationships as a training set to train a multi-task learning model to obtain a recommendation model;

[0057] Step S15: Complete the hotel room type recommendation for the user based on the recommendation model.

[0058] It should be noted that the technical solution provided in this embodiment can be loaded and used in an existing hotel system or application in the form of a small program or a plug-in, or in the form of a separate application, and the hotel room type recommendation function can be implemented through an external interface. Applicable scenarios include, but are not limited to: hotel room type recommendation.

[0059] It can be understood that for the method provided in this embodiment, the user-side text features are obtained based on the voice data and evaluation text data in the user's historical data, the product-side fusion features are obtained based on the hotel room type image data and hotel room type introduction text data, multiple pairs of feature relationships are obtained based on the user-side text features and the product-side fusion features, the multiple pairs of feature relationships are used as a training set to train a multi-task learning model to obtain a recommendation model, and the hotel room type recommendation for the user is completed based on the recommendation model. This recommendation method deeply associates the user with the hotel room type through the user-side text features and the product-side fusion features, making the hotel room type recommendation more in line with the user's needs and effectively improving the user's experience.

[0060] In specific practice, in step S11, "obtain the user-side text features based on the user's historical data" includes: converting the voice data to obtain the first text data, and obtaining the user-side text features through the WordPiece method based on the first text data and the evaluation text data.

[0061] Specifically, use Whisper to convert the voice data into the first text data, and combine the user's evaluation text data, and use the WordPiece method to obtain the user-side text features.

[0062] It should be noted that the user's historical data generally collects the user-side data in the past year.

[0063] In specific practice, "obtaining the fused feature on the product side according to the hotel room type data" in step S12 includes: extracting the image of the ROI feature region of the hotel room type from the hotel room type image data to obtain the product side image feature; obtaining the to-be-fused product side image feature based on the ResNet50 model according to the product side image feature; the to-be-fused product side image feature is a vector feature of the first preset dimension; obtaining the to-be-fused product side text feature based on the XL-Net model according to the hotel room type introduction text data; the to-be-fused product side text feature is a vector feature of the second preset dimension; converting the to-be-fused product side image feature and the to-be-fused product side text feature into the same dimension through the linear layer of the DNN model; adding the values of the corresponding dimensions after aligning the dimensions of the converted to-be-fused product side image feature and the to-be-fused product side text feature to obtain the first fused feature; performing downsampling processing on the first fused feature using the average pooling layer to obtain the second fused feature; performing downsampling processing on the to-be-fused product side image feature through the max pooling layer to obtain the third fused feature; obtaining the fused feature on the product side through the decision layer according to the second fused feature and the third fused feature.

[0064] It should be noted that the hotel room type data generally collects the product side data within the past year.

[0065] Specifically, the hotel room type image data includes the key frame images in the hotel room type introduction video and the promotional images of the hotel room type. Extracting the image of the ROI feature region of the hotel room type from the hotel room type image data to obtain the product side image feature includes: extracting the image of the ROI feature region from the key frame images and the promotional images through faster-RCNN as the product side image feature.

[0066] It should be noted that the key frame images in the hotel room type introduction video are obtained using OpenCV, with the dimension represented as (224, 224, 3). Adding the existing promotional images of the hotel room type to obtain the hotel room type image data; using faster-RCNN to extract all the ROI feature regions (including the FeatureMaps of the windows and beds) in the hotel room type image data of each hotel room type, and the dimension of the feature map is also (224, 224, 3).

[0067] Specifically, the to-be-fused product-side image features are obtained based on the product-side image features using the ResNet50 model; the to-be-fused product-side image features are vector features of the first preset dimension, including: using the ResNet50 pre-trained model, taking the product-side image features as Input data, and the whole process is forward calculation to obtain the best representation of the visual features as the to-be-fused product-side image features, where the best representation is the feature map of the ResNet50 average pooling layer (AvgPooling Layer) with a dimension of (6, 6, 2048), and this dimension is the first preset dimension.

[0068] Specifically, the to-be-fused product-side text features are obtained based on the hotel room type introduction text data using the XL-Net model; the to-be-fused product-side text features are vector features of the second preset dimension, including: using the XL-Net pre-trained model, taking the hotel room type introduction text data as Input data, and similarly, the whole process is forward calculation to obtain the best representation of the text features as the to-be-fused product-side text features, where the best representation is the text Embedding in the Decoder Layer before the output layer of XL-Net, and the dimension representation is (192, 384), and this dimension is the second preset dimension.

[0069] Specifically, the to-be-fused product-side image features and the to-be-fused product-side text features are transformed into the same dimension through the linear layer of the DNN model, including: based on the linear layer of the DNN model, in order to make the dimension of the to-be-fused product-side text features consistent with the dimension of the to-be-fused product-side image features, here is actually through the parameter matrix W of the linear layer T to perform a linear transformation of dimension alignment:

[0070] ; where, V represents the feature vector, W is the custom LinearLayer weight parameter matrix, b is the bias, and the dimension change of the feature vector is as follows:

[0071]

[0072] , where, D represents the number of eigenvalues of each dimension in the feature vector V, D input represents the input layer dimension, and D output represents the output layer dimension.

[0073] The to-be-fused product-side text features output by this linear layer are multi-dimensional feature vectors with the same dimension as the to-be-fused product-side image features.

[0074] Specifically, the converted image features on the product side to be fused and the text features on the product side to be fused are aligned in dimension, and the values ​​of the corresponding dimensions are added to obtain the first fused feature, including: based on the addition layer of the DNN model, the feature vectors of the image features on the product side to be fused and the text features on the product side to be fused after the dimension alignment are numerically added according to the corresponding positions, that is, Pointwise addition. After this operation, the dimension of the feature vector remains unchanged:

[0075] .

[0076] Specifically, the first fusion feature is downsampled using an average pooling layer to obtain a second fusion feature, including: an average pooling layer based on a DNN model: after the addition layer, observing the corresponding heat map of the first fusion feature, it can be found that the first fusion feature is too sharp. In order to retain more image background information and reduce the increase in the variance of the final estimate caused by the large change in the neighborhood size, an average pooling window with a window size of 2×2 is used for downsampling, which reduces subsequent training parameters and calculations, and avoids overfitting of the model:

[0077] , where V (x,y) It indicates that the output FeatureMap position is (x, y) and the feature value is output after average pooling calculation; p and q respectively limit the representation range of the pooling window; the fused feature obtained at this time is the second fused feature.

[0078] Specifically, the third fusion feature is obtained by downsampling the image features of the product side to be fused through the maximum pooling layer, including: based on the maximum pooling layer, in order to retain more texture information and reduce the deviation of the estimated value mean caused by the convolution kernel parameter error, we downsample the image features of the product side to be fused obtained by forward calculation from the ResNet50 pre-training model using a maximum pooling window with a window size of 2×2 to obtain the third fusion feature:

[0079] ; Among them, V (x,y) It indicates the feature value output after the maximum pooling calculation at the output FeatureMap location (x, y); p and q respectively limit the representation range of the pooling window.

[0080] Specifically, the commodity-side fusion feature is obtained through the decision layer according to the second fusion feature and the third fusion feature, including: the decision layer learns the following decision function by inputting the obtained feature value (second fusion feature) and the target value (third fusion feature): ; Where F(x) is the average value of all channel values ​​at the corresponding position in the target FeatureMap, here is: ; X is the average value of all channel values at the corresponding position in the Feature Map, which is: .

[0081] Here, to achieve the goal of enabling the model to converge quickly, the parameter matrix W of the decision layer is initialized T as weight values that conform to the standard normal distribution X∽N(0, 1).

[0082] It should be noted that for the optimal solution of the decision function, we adopt the following loss function here:

[0083] ; where F(X) p and X p are the target value and feature value at the corresponding position p respectively; D1 and D2 are the number of feature values of each dimension except the channel C dimension; ||W|| is the norm of the parameter matrix, and λ1 and λ2 are adjustable penalty term parameters respectively. To make the second fusion feature most similar to the features in the original feature space, Feature Map represents the feature map. We solve the weight coefficient matrix and the adjustable penalty parameter value that satisfy the following relationship to minimize the value of the loss function in the current training data: , we use the optimal model obtained from the existing training, input the text features and visual features on the commodity side into the model for forward calculation, and the obtained second fusion feature can be used as the commodity-side fusion feature.

[0084] In specific practice, in step S13, "obtaining multiple groups of feature relationship pairs of the user-side text features and the commodity-side fusion features based on the user-side text features and the commodity-side fusion features" includes: inputting the user-side text features and the commodity-side fusion features into the ImageBERT model to obtain multiple groups of feature relationship pairs of the user-side text features and the commodity-side fusion features.

[0085] It should be noted that we screen the feature relation pairs with relatively high matching degrees through the image-text matching task of ImageBERT. Using the pre-trained ImageBERT model, the above-mentioned user-side text features are processed into 64-dimensional embedding word vectors, and at the same time, the product-side fusion features are also processed into embedding vectors of the same dimension. Finally, they are concatenated into Linguistic Embedding. Before inputting into the Transformer of ImageBERT, it is also necessary to add Segment Embedding that can distinguish feature modalities and Sequense Position Embedding that can distinguish the positional relationships between tokens. Additionally, in the current recommendation scenario, in order to reflect the geographical location information of the samples, the geographical location generated by the GWR model is added. At the same time, in order to better represent the global features of the images, we add the global visual features of the key-frame images and promotional images. Experiments have shown that the model can achieve a better screening and pairing effect after adding the global visual features. Feature extraction: After fine-tuning the pre-trained ImageBERT model, the threshold is set to 0.72, and the feature pairs of all feature relation pairs whose final output results of the model exceed this threshold are cached, so as to obtain the offline recall pool of the calculated feature pairs of the target user and the hotel room types of interest.

[0086] In specific practice, in step S14, "training the multi-task learning model with multiple groups of feature relation pairs to obtain a recommendation model", the "multi-task learning model" is the MMoE model.

[0087] It should be noted that the training process of the recommendation model is as follows: Based on each feature relation pair obtained in the above feature extraction step as the model input, use MMoE as the recommendation model for training and tuning. Here, we introduce an additional objective (order placement rate) as a secondary indicator for online evaluation. At the same time, in order to alleviate the performance impact caused by the imbalance of positive and negative sample ratios in the current scenario, we use Focal Loss to further optimize the loss function. The expression of the Focal Loss function is as follows:

[0088] , where:

[0089] p t represents the probability that the sample is predicted by the model as the target of the t-th class. p t The larger it is, the higher the confidence of the classification, indicating that the sample is easier to classify; p tThe smaller it is, the lower the confidence of classification, indicating that the sample is more difficult to classify. Therefore, focal loss is equivalent to increasing the weight of difficult-to-classify samples in the loss function, making the loss function tend to difficult-to-classify samples, which helps to improve the accuracy of difficult-to-classify samples; γ is an adjustable factor, γ > 0. From the experiments in the current scenario, when γ takes 0.82, the recommended model can achieve the optimal effect.

[0090] In specific practice, "completing the hotel room type recommendation for the user based on the recommended model" in step S15 includes: when it is necessary to recommend a hotel room type to the current user, obtaining the user's historical data; obtaining the text features on the current user side based on the user's historical data; inputting the text features on the current user side into the recommended model to obtain the target hotel room type, and recommending the target hotel room type to the current user.

[0091] In specific practice, the method further includes: completing the online recall function of hotel room types based on FAISS.

[0092] It should be noted that in order to improve the real-time performance of the recommendation results, FAISS is used to complete the online recall, which improves the online recall efficiency and the real-time performance of the recommendation results.

[0093] Please refer to Figure 2 , Figure 2 is a schematic block diagram of a hotel room type recommendation device shown in an exemplary embodiment of the present invention. Refer to Figure 2 , the hotel room type recommendation device 100 includes:

[0094] The user-side preprocessing module 101 is used to obtain the text features on the user side based on the user's historical data; the user's historical data includes voice data and evaluation text data;

[0095] The product-side preprocessing module 102 is used to obtain the product-side fusion features based on the hotel room type data; the hotel room type data includes hotel room type image data and hotel room type introduction text data;

[0096] The feature relationship pair module 103 is used to obtain multiple groups of feature relationship pairs between the text features on the user side and the product-side fusion features based on the text features on the user side and the product-side fusion features;

[0097] The training module 104 is used to use multiple groups of feature relationship pairs as a training set to train a multi-task learning model to obtain a recommended model;

[0098] The recommendation module 105 is used to complete the hotel room type recommendation for the user based on the recommended model.

[0099] It should be noted that the applicable scenarios of the device provided in this embodiment include, but are not limited to: hotel room type recommendation.

[0100] It can be understood that the device provided in this embodiment obtains the user-side text features based on the voice data and evaluation text data in the user's historical data, obtains the product-side fusion features based on the hotel room type image data and hotel room type introduction text data, obtains multiple groups of feature relationship pairs based on the user-side text features and the product-side fusion features, uses the multiple groups of feature relationship pairs as a training set to train a multi-task learning model to obtain a recommendation model, and completes the hotel room type recommendation for the user based on the recommendation model. Through the user-side text features and the product-side fusion features, this recommendation method deeply associates the user with the hotel room type, making the hotel room type recommendation more in line with the user's needs and effectively improving the user experience.

[0101] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0102] It should be noted that, in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.

[0103] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0104] Each embodiment in this specification is described in a related manner. For the same or similar parts between each embodiment, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the relevant part of the method embodiment for the relevant content.

[0105] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.

[0106] The above-described embodiments only represent several implementation manners of this application, and the description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application patent shall be subject to the appended claims.

Claims

1. A method for recommending hotel room types, characterized in that: The method comprises: Obtaining user-side text features based on user history data; the user history data includes voice data and evaluation text data; Obtaining commodity-side fusion features based on hotel room type data; the hotel room type data includes hotel room type image data and hotel room type introduction text data; Obtaining multiple groups of feature relationship pairs between the user-side text features and the product-side fusion features according to the user-side text features and the product-side fusion features; Using the multiple groups of feature relationship pairs as training sets to train the multi-task learning model to obtain a recommendation model; Complete hotel room type recommendations for users based on the recommendation model; The product-side fusion features obtained based on the hotel room type data include: Extracting an image of a ROI feature area of ​​a hotel room type from the hotel room type image data to obtain a product side image feature; Obtaining the product side image features to be fused based on the ResNet50 model according to the product side image features; the product side image features to be fused are vector features of a first preset dimension; Obtaining the product-side text features to be integrated based on the XL-Net model according to the hotel room type introduction text data; the product-side text features to be integrated are vector features of the second preset dimension; The image features of the product to be fused and the text features of the product to be fused are converted into the same dimension through the linear layer processing of the DNN model; Aligning the dimensions of the converted image feature on the product side to be fused with the converted text feature on the product side to be fused, and adding the values ​​of the corresponding dimensions to obtain a first fusion feature; The first fusion feature is downsampled using an average pooling layer to obtain a second fusion feature; According to the product side image features to be fused, downsampling is performed through a maximum pooling layer to obtain a third fusion feature; Obtaining the product-side fusion feature through a decision layer according to the second fusion feature and the third fusion feature; The obtaining, based on the user-side text features and the commodity-side fusion features, a plurality of feature relationship pairs of the user-side text features and the commodity-side fusion features comprises: The user-side text features and the product-side fusion features are input into the ImageBERT model to obtain multiple groups of feature relationship pairs between the user-side text features and the product-side fusion features.

2. The recommendation method according to claim 1, characterized in that: The hotel room type image data includes key frame images in the hotel room type introduction video and promotional images of the hotel room type, and the image of the ROI feature area of ​​the hotel room type is extracted from the hotel room type image data to obtain the product side image features, including: The image of the ROI feature area is extracted from the key frame image and the promotional image by using Faster-RCNN as the product side image feature.

3. The recommendation method according to claim 1, characterized in that: The obtaining of user-side text features based on user historical data includes: Convert the voice data to obtain first text data; The user-side text features are obtained by using the WordPiece method according to the first text data and the evaluation text data.

4. The recommendation method according to claim 1, characterized in that: The multi-task learning model is an MMoE model.

5. The recommendation method according to claim 3, characterized in that: The method of completing the hotel room type recommendation for the user based on the recommendation model includes: When it is necessary to recommend a hotel room type to the current user, obtain the user history data of the current user; Obtaining a current user-side text feature according to the user history data of the current user; The current user-side text features are input into the recommendation model to obtain the target hotel room type.

6. The recommendation method according to claim 5, characterized in that: The method further comprises: Complete the online recall function of hotel room types based on FAISS.

7. A device for recommending hotel room types, characterized in that: The device comprises: A user-side preprocessing module, used to obtain user-side text features based on user history data; the user history data includes voice data and evaluation text data; The product side preprocessing module is used to obtain product side fusion features based on hotel room type data; the hotel room type data includes hotel room type image data and hotel room type introduction text data; the image of the ROI feature area of ​​the hotel room type is extracted from the hotel room type image data to obtain product side image features; the product side image features to be fused are obtained based on the ResNet50 model according to the product side image features; the product side image features to be fused are vector features of the first preset dimension; the product side text features to be fused are obtained based on the XL-Net model according to the hotel room type introduction text data; the product side text features to be fused are vector features of the second preset dimension dimensional vector features; convert the image features of the product to be fused and the text features of the product to be fused into the same dimension through the linear layer processing of the DNN model; align the dimensions of the converted image features of the product to be fused and the converted text features of the product to be fused, and then add the values ​​of the corresponding dimensions to obtain a first fusion feature; downsample the first fusion feature using an average pooling layer to obtain a second fusion feature; downsample the image features of the product to be fused through a maximum pooling layer to obtain a third fusion feature; obtain the product side fusion feature through a decision layer based on the second fusion feature and the third fusion feature; A feature relationship pair module is used to obtain multiple sets of feature relationship pairs of the user-side text features and the product-side fusion features based on the user-side text features and the product-side fusion features; the user-side text features and the product-side fusion features are input into the ImageBERT model to obtain multiple sets of feature relationship pairs of the user-side text features and the product-side fusion features; A training module, used for training a multi-task learning model using the plurality of feature relationship pairs as training sets to obtain a recommendation model; The recommendation module is used to recommend hotel room types to users based on the recommendation model.

Citation Information

Patent Citations

  • Commodity recommendation method and device

    CN115293853A

  • Hotel guest room recommendation method, system and device based on business travel and storage medium

    CN116911947A