Hotel room type recommendation method and device

By extracting features from user historical data and hotel room type data and using a multi-task learning model to recommend hotel room type, the problem that the recommendation results in the existing technology do not meet user expectations, and the recommendation effect is achieved that is more in line with user needs.

CN120011646AActive Publication Date: 2025-05-16SHENZHEN HUOLI TIAN HUI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510480138.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-16
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The existing hotel room type recommendation method relies on user browsing and booking history, and lacks in-depth correlation between user-side data and hotel-side data, resulting in the recommendation results not meeting user expectations.

Method used

The user-side text features are obtained by extracting voice data and evaluation text data from user historical data, image data and introduction text data are extracted from hotel room type data to obtain product-side fusion features, and a multi-task learning model is used to train based on these features to generate a recommendation model for hotel room type recommendation.

Benefits of technology

By deeply linking users and hotel room types, the recommendation results are more in line with user needs and effectively improve the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011646A_ABST
    Figure CN120011646A_ABST
Patent Text Reader

Abstract

The invention relates to a hotel room type recommendation method and device, and belongs to the technical field of intelligent recommendation, and the method comprises the steps: obtaining a user side text feature according to voice data and evaluation text data in user historical data, obtaining a commodity side fusion feature according to hotel room type image data and hotel room type introduction text data, the method comprises the steps of obtaining user side text features and commodity side fusion features, obtaining multiple groups of feature relationship pairs according to the user side text features and the commodity side fusion features, training a multi-task learning model by taking the multiple groups of feature relationship pairs as a training set to obtain a recommendation model, and completing hotel room type recommendation for a user based on the recommendation model. The user and the hotel room type are deeply associated, so that the hotel room type recommendation is more in line with the user demand, and the user experience is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent recommendation, and in particular relates to a method and device for recommending hotel room types. Background Art

[0002] In recent years, with the improvement of living standards and changes in personal consumption concepts, people pay more and more attention to personalized experience during travel. Especially the younger generation, they are more inclined to choose accommodation that can provide unique feelings and has social sharing value. Therefore, the hotel room type recommendation system needs to fully consider market demand and provide room type recommendations that meet the needs of different customer groups.

[0003] The existing hotel room type recommendation method mainly recommends room types that meet user preferences by analyzing user browsing and booking history. For example, users who frequently book king-size rooms may be more inclined to recommend king-size rooms. However, this existing hotel room type recommendation method only relies on user browsing and booking history, lacks deep association between user-side data and hotel-side data, and is prone to problems where hotel room type recommendations do not meet user expectations. Summary of the invention

[0004] To this end, the present invention provides a method and device for recommending hotel room types, so as to solve the problem that the existing hotel room type recommendation method is prone to the hotel room type recommendation not meeting the user's expectations.

[0005] To achieve the above objectives, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a method for recommending hotel room types, comprising: Obtaining user-side text features based on user history data; the user history data includes voice data and evaluation text data; Obtaining commodity-side fusion features based on hotel room type data; the hotel room type data includes hotel room type image data and hotel room type introduction text data; Obtaining multiple groups of feature relationship pairs between the user-side text features and the product-side fusion features according to the user-side text features and the product-side fusion features; Using the multiple groups of feature relationship pairs as training sets to train the multi-task learning model to obtain a recommendation model; Complete hotel room type recommendations for users based on the recommendation model.

[0006] Furthermore, obtaining product-side fusion features based on hotel room type data includes: Extracting an image of a ROI feature area of ​​a hotel room type from the hotel room type image data to obtain a product side image feature; Obtaining the product side image features to be fused based on the ResNet50 model according to the product side image features; the product side image features to be fused are vector features of a first preset dimension; Obtaining the product-side text features to be integrated based on the XL-Net model according to the hotel room type introduction text data; the product-side text features to be integrated are vector features of the second preset dimension; The image features of the product to be fused and the text features of the product to be fused are converted into the same dimension through the linear layer processing of the DNN model; Aligning the dimensions of the converted image feature on the product side to be fused with the text feature on the product side to be fused, and adding the values ​​of the corresponding dimensions to obtain a first fused feature; The first fusion feature is downsampled using an average pooling layer to obtain a second fusion feature; According to the product side image features to be fused, downsampling is performed through a maximum pooling layer to obtain a third fusion feature; The commodity-side fusion feature is obtained through a decision layer according to the second fusion feature and the third fusion feature.

[0007] Further, the hotel room type image data includes key frame images in the hotel room type introduction video and promotional images of the hotel room type, and the image of the ROI feature area of ​​the hotel room type is extracted from the hotel room type image data to obtain the product side image features, including: The image of the ROI feature area is extracted from the key frame image and the promotional image by using Faster-RCNN as the product side image feature.

[0008] Furthermore, obtaining user-side text features based on user historical data includes: Convert the voice data to obtain first text data; The user-side text features are obtained by using the WordPiece method according to the first text data and the evaluation text data.

[0009] Furthermore, obtaining a plurality of feature relationship pairs of the user-side text features and the product-side fusion features based on the user-side text features and the product-side fusion features includes: The user-side text features and the product-side fusion features are input into the ImageBERT model to obtain multiple groups of feature relationship pairs between the user-side text features and the product-side fusion features.

[0010] Furthermore, the multi-task learning model is a MMoE model.

[0011] Furthermore, the step of completing the hotel room type recommendation for the user based on the recommendation model includes: When it is necessary to recommend a hotel room type to the current user, obtain the user history data of the current user; Obtaining a current user-side text feature according to the user history data of the current user; The current user-side text features are input into the recommendation model to obtain the target hotel room type.

[0012] Furthermore, the method further comprises: Complete the online recall function of hotel room types based on FAISS.

[0013] In a second aspect, the present invention provides a device for recommending hotel room types, comprising: A user-side preprocessing module, used to obtain user-side text features based on user history data; the user history data includes voice data and evaluation text data; A product-side preprocessing module, used to obtain product-side fusion features based on hotel room type data; the hotel room type data includes hotel room type image data and hotel room type introduction text data; A feature relationship pair module, used to obtain multiple sets of feature relationship pairs of the user-side text features and the product-side fusion features according to the user-side text features and the product-side fusion features; A training module, used for training a multi-task learning model using the plurality of feature relationship pairs as training sets to obtain a recommendation model; The recommendation module is used to recommend hotel room types to users based on the recommendation model.

[0014] The present invention adopts the above technical solution and has at least the following beneficial effects: A method and device for recommending hotel room types are provided. User-side text features are obtained based on voice data and evaluation text data in user historical data. Product-side fusion features are obtained based on hotel room image data and hotel room introduction text data. Multiple groups of feature relationship pairs are obtained based on user-side text features and product-side fusion features. A multi-task learning model is trained using the multiple groups of feature relationship pairs as training sets to obtain a recommendation model. Hotel room types are recommended to users based on the recommendation model. The recommendation method deeply associates users and hotel room types through user-side text features and product-side fusion features, so that hotel room type recommendations are more in line with user needs and the user experience is effectively improved.

[0015] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1 is a flow chart of a method for recommending hotel room types shown in an exemplary embodiment of the present invention; Figure 2 It is a schematic block diagram of a device for recommending hotel room types according to an exemplary embodiment of the present invention.

[0018] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be described in detail below. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other implementation methods obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention.

[0020] The existing hotel room type recommendation method mainly recommends room types that meet user preferences by analyzing user browsing and booking history. When making recommendations, many potential untapped high-quality user-side and product-side images and visual features are used in combination with text and visual pattern information to significantly improve model indicators in recommendation scenarios.

[0021] The technologies represented by the English abbreviations used in this application are as follows: ROI is Region of Interest, which is the target region in image segmentation and extraction; ResNet50 model is a residual convolutional neural network containing 50 residual blocks; XL-Net model is Transformer with Extra Long-range Net, a long sequence large language model based on Transformer-XL; DNN model is Deep Neural Network, a deep neural network; faster-RCNN is Faster Region-based Convolutional Neural Network, a region-based fast target detection and segmentation convolutional neural network model; WordPiece is word piece segmentation, a word segmentation algorithm in natural language processing; ImageBERT model is a model based on BERT large language model that can integrate visual and language information; MMoE model is Multi-gate Mixture-of-Experts, a mixed expert multi-task recommendation model based on multi-gate mechanism; FAISS is Facebook AI Similarity Search, an efficient similarity retrieval tool for Facebook AI; Whisper is a whisperer, an automatic speech recognition system based on Transformer architecture; OpenCV is Open Source Computer Vision Library, an open source computer vision and machine learning software tool library; Feature Map is the feature map, the intermediate product of the forward calculation of image sample data in the neural network; DecoderLayer is the decoding layer; Input is the input layer; Embedding is the embedded feature; Focal Loss is the focal loss.

[0022] The method and device of the present invention are described below through specific embodiments.

[0023] See also Figure 1 , Figure 1 is a flowchart of a method for recommending hotel room types according to an exemplary embodiment of the present invention. Figure 1 , the method comprising: Step S11, obtaining user-side text features based on user history data; user history data includes voice data and evaluation text data; Step S12: obtaining product-side fusion features based on hotel room type data; the hotel room type data includes hotel room type image data and hotel room type introduction text data; Step S13, obtaining multiple groups of feature relationship pairs of user-side text features and product-side fusion features based on the user-side text features and the product-side fusion features; Step S14: using multiple groups of feature relationship pairs as training sets to train the multi-task learning model to obtain a recommendation model; Step S15: Recommend hotel room types to the user based on the recommendation model.

[0024] It should be noted that the technical solution provided in this embodiment can be loaded into an existing hotel system or application in the form of a small program or a plug-in in actual practice, or in the form of a separate application to implement the hotel room type recommendation function through an external interface. Applicable scenarios include but are not limited to: hotel room type recommendation.

[0025] It can be understood that the method provided in this embodiment obtains user-side text features based on voice data and evaluation text data in user historical data, obtains product-side fusion features based on hotel room image data and hotel room introduction text data, obtains multiple groups of feature relationship pairs based on user-side text features and product-side fusion features, uses the multiple groups of feature relationship pairs as training sets to train a multi-task learning model to obtain a recommendation model, and completes hotel room recommendations for users based on the recommendation model. This recommendation method deeply associates users and hotel room types through user-side text features and product-side fusion features, so that hotel room recommendations are more in line with user needs and effectively improve user experience.

[0026] In specific practice, in step S11, "obtaining user-side text features based on user historical data" includes: converting the voice data to obtain first text data, and obtaining user-side text features through the WordPiece method based on the first text data and the evaluation text data.

[0027] Specifically, Whisper is used to convert the voice data into first text data, and combined with the user's evaluation text data, the WordPiece method is used to obtain the user-side text features.

[0028] It should be noted that user historical data generally collects user-side data within the past year.

[0029] In specific practice, "obtaining product-side fusion features based on hotel room type data" in step S12 includes: extracting images of the ROI feature area of ​​the hotel room type from the hotel room type image data to obtain product-side image features; obtaining product-side image features to be fused based on the ResNet50 model based on the product-side image features; the product-side image features to be fused are vector features of a first preset dimension; obtaining product-side text features to be fused based on the XL-Net model based on the hotel room type introduction text data; the product-side text features to be fused are vector features of a second preset dimension; converting the product-side image features to be fused and the product-side text features to be fused to the same dimension through the linear layer processing of the DNN model; aligning the dimensions of the converted product-side image features to be fused with the product-side text features to be fused, and then adding the values ​​of the corresponding dimensions to obtain the first fusion feature; downsampling the first fusion feature using the average pooling layer to obtain the second fusion feature; downsampling the product-side image features to be fused through the maximum pooling layer to obtain the third fusion feature; obtaining the product-side fusion feature through the decision layer based on the second fusion feature and the third fusion feature.

[0030] It should be noted that hotel room type data generally collects product-side data within the past year.

[0031] Specifically, the hotel room image data includes key frame images in the hotel room introduction video and promotional images of the hotel room. The image of the ROI feature area of ​​the hotel room is extracted from the hotel room image data to obtain the product side image features, including: extracting the image of the ROI feature area from the key frame images and the promotional images as the product side image features through faster-RCNN.

[0032] It should be noted that OpenCV is used to obtain the key frame images in the hotel room introduction video, with the dimension represented as (224, 224, 3), and the hotel room image data is obtained by adding the promotional images of the existing hotel room types. Faster-RCNN is used to extract all ROI feature areas (including FeatureMaps of windows and beds) in the hotel room image data of each hotel room type, and the feature map dimension is also (224, 224, 3).

[0033] Specifically, the product side image features to be fused are obtained based on the ResNet50 model according to the product side image features; the product side image features to be fused are vector features of a first preset dimension, including: using the ResNet50 pre-trained model, taking the product side image features as input data, the whole process is a forward calculation, and the best representation of the visual features is obtained as the product side image features to be fused, wherein the best representation, i.e., the feature map (Feature Map) of the ResNet50 average pooling layer (AvgPooling Layer), has a dimension of (6, 6, 2048), which is the first preset dimension.

[0034] Specifically, the text features of the product side to be fused are obtained based on the XL-Net model according to the text data of the hotel room type introduction. The text features of the product side to be fused are vector features of the second preset dimension, including: using the XL-Net pre-trained model and taking the text data of the hotel room type introduction as input data. Similarly, the whole process is a forward calculation to obtain the best representation of the text features as the text features of the product side to be fused, wherein the best representation is the text Embedding in the Decoder Layer before the XL-Net output layer, and the dimension is represented as (192, 384), which is the second preset dimension.

[0035] Specifically, the image features of the product to be fused and the text features of the product to be fused are converted into the same dimension through the linear layer processing of the DNN model, including: based on the linear layer of the DNN model, in order to make the dimension of the text features of the product to be fused consistent with the dimension of the image features of the product to be fused, here is actually through the parameter matrix W of the linear layer T A linear transformation is performed to align the dimensions: ; Where V represents the feature vector, W is the custom LinearLayer weight parameter matrix, b is the bias, and the feature vector dimension changes as follows:

[0036] , where D represents the number of eigenvalues ​​of each dimension in the eigenvector V, D input Denotes the input layer dimension, D output Represents the output layer dimension.

[0037] The text features of the product to be fused output by this linear layer are multi-dimensional feature vectors with the same dimensions as the image features of the product to be fused.

[0038] Specifically, the converted image features on the product side to be fused and the text features on the product side to be fused are aligned in dimension, and the values ​​of the corresponding dimensions are added to obtain the first fused feature, including: based on the addition layer of the DNN model, the feature vectors of the image features on the product side to be fused and the text features on the product side to be fused after the dimension alignment are numerically added according to the corresponding positions, that is, Pointwise addition. After this operation, the dimension of the feature vector remains unchanged: .

[0039] Specifically, the first fusion feature is downsampled using an average pooling layer to obtain a second fusion feature, including: an average pooling layer based on a DNN model: after the addition layer, observing the corresponding heat map of the first fusion feature, it can be found that the first fusion feature is too sharp. In order to retain more image background information and reduce the increase in the variance of the final estimate caused by the large change in the neighborhood size, an average pooling window with a window size of 2×2 is used for downsampling, which reduces subsequent training parameters and calculations, and avoids overfitting of the model: , where V (x,y) It indicates that the output FeatureMap position is (x, y) and the feature value is output after average pooling calculation; p and q respectively limit the representation range of the pooling window; the fused feature obtained at this time is the second fused feature.

[0040] Specifically, the third fusion feature is obtained by downsampling the image features of the product side to be fused through the maximum pooling layer, including: based on the maximum pooling layer, in order to retain more texture information and reduce the deviation of the estimated value mean caused by the convolution kernel parameter error, we downsample the image features of the product side to be fused obtained by forward calculation from the ResNet50 pre-training model using a maximum pooling window with a window size of 2×2 to obtain the third fusion feature: ; Among them, V (x,y) It indicates the feature value output after the maximum pooling calculation at the output FeatureMap location (x, y); p and q respectively limit the representation range of the pooling window.

[0041] Specifically, the commodity-side fusion feature is obtained through the decision layer according to the second fusion feature and the third fusion feature, including: the decision layer learns the following decision function by inputting the obtained feature value (second fusion feature) and the target value (third fusion feature): ; Where F(x) is the average value of all channel values ​​at the corresponding position in the target FeatureMap, here is: ; X is the average value of all channel values ​​at the corresponding position in the feature FeatureMap, here is: .

[0042] In order to achieve the purpose of rapid convergence of the model, the parameter matrix W of the decision layer is initialized. T is the weight value that conforms to the standard normal distribution of X∽N(0, 1).

[0043] It should be noted that for the optimization solution of the decision function, we use the following loss function: ; Among them, F(X) p and X p are the target value and eigenvalue of the corresponding position p respectively; D1 and D2 are the number of eigenvalues ​​of each dimension except the channel C dimension respectively; ||W|| is the modulus of the parameter matrix, and λ1 and λ2 are the adjustable penalty term parameters respectively. In order to make the second fusion feature most similar to the feature of the original feature space, Feature Map represents the feature map, we solve the weight coefficient matrix and the adjustable penalty parameter value that satisfy the following relationship, so that the value of the loss function is minimized in the current training data: We use the optimal model obtained through existing training, input the text features and visual features of each product side into the model for forward calculation, and the obtained second fusion features can be used as the product side fusion features.

[0044] In specific practice, step S13 of "obtaining multiple sets of feature relationship pairs of user-side text features and product-side fusion features based on user-side text features and product-side fusion features" includes: inputting user-side text features and product-side fusion features into the ImageBERT model to obtain multiple sets of feature relationship pairs of user-side text features and product-side fusion features.

[0045] It should be noted that we screen the feature relationship pairs with high matching degree through the image-text matching task of ImageBERT. Using the pre-trained ImageBERT model, the above user-side text features are processed into 64-dimensional embedding word vectors, and the product-side fusion features are also processed into embedding vectors of the same dimension, and finally spliced ​​into Linguistic Embedding. Before inputting into the ImageBERT Transformer, it is also necessary to add Segment Embedding that can distinguish feature modes and Sequense Position Embedding that distinguishes the position relationship between tokens. In addition, in order to reflect the geographic location information of the sample in the current recommendation scenario, the geographic location generated by the GWR model is added. At the same time, in order to better represent the global features of the image, we add the global visual features of the keyframe image and the promotional image. Experiments have shown that the model after adding global visual features can achieve better screening and matching effects. Feature extraction: After fine-tuning the ImageBERT pre-trained model, the threshold is set to 0.72, and the feature pairs of all feature relationship pairs whose final output results of the model exceed this threshold are cached. In this way, an offline recall pool of calculated feature pairs of target users and hotel room types of interest is obtained.

[0046] In specific practice, in step S14, "using multiple groups of feature relationship pairs as training sets to train the multi-task learning model to obtain a recommendation model", the "multi-task learning model" is the MMoE model.

[0047] It should be noted that the training process of the recommendation model is as follows: based on each feature relationship pair obtained in the above feature extraction step as the model input, MMoE is used as the recommendation model for training and tuning. Here, we introduce an additional target (order rate) as a secondary indicator for online inspection. At the same time, in order to alleviate the performance impact caused by the imbalance of positive and negative sample ratios in the current scenario, we use Focal Loss to further optimize the loss function. The expression of the Focal Loss loss function is as follows: ,in: p t It represents the probability that the sample is predicted by the model as the t-th type target. t The larger the value, the higher the confidence of the classification, which means the samples are easier to classify; tThe smaller it is, the lower the confidence of the classification is, which means the sample is more difficult to distinguish. Therefore, focal loss is equivalent to increasing the weight of difficult samples in the loss function, making the loss function tend to be difficult samples, which helps to improve the accuracy of difficult samples; γ is an adjustable factor, γ>0, and the current scenario experiment shows that when γ is 0.82, the recommended model effect can reach the best.

[0048] In specific practice, step S15 of "completing hotel room type recommendation for users based on recommendation model" includes: when it is necessary to recommend a hotel room type to the current user, obtaining the user history data of the current user; obtaining the current user side text features based on the user history data of the current user; inputting the current user side text features into the recommendation model to obtain the target hotel room type, and recommending the target hotel room type to the current user.

[0049] In specific practice, the method also includes: completing the online recall function of hotel room types based on FAISS.

[0050] It should be noted that in order to improve the real-time performance of recommendation results, FAISS is used to complete online recall, which improves the online recall efficiency and the real-time performance of recommendation results.

[0051] See also Figure 2 , Figure 2 is a schematic block diagram of a hotel room type recommendation device shown in an exemplary embodiment of the present invention, see Figure 2 , the hotel room type recommendation device 100 includes: The user-side preprocessing module 101 is used to obtain user-side text features based on user history data; the user history data includes voice data and evaluation text data; The product side preprocessing module 102 is used to obtain product side fusion features according to the hotel room type data; the hotel room type data includes hotel room type image data and hotel room type introduction text data; A feature relationship pair module 103 is used to obtain a plurality of feature relationship pairs of user-side text features and product-side fusion features based on user-side text features and product-side fusion features; A training module 104 is used to train a multi-task learning model using multiple sets of feature relationship pairs as training sets to obtain a recommendation model; The recommendation module 105 is used to recommend hotel room types to users based on the recommendation model.

[0052] It should be noted that the device provided in this embodiment is applicable to scenarios including but not limited to: hotel room type recommendation.

[0053] It can be understood that the device provided in this embodiment obtains user-side text features based on voice data and evaluation text data in user historical data, obtains product-side fusion features based on hotel room image data and hotel room introduction text data, obtains multiple groups of feature relationship pairs based on user-side text features and product-side fusion features, uses the multiple groups of feature relationship pairs as training sets to train a multi-task learning model to obtain a recommendation model, and completes hotel room type recommendations for users based on the recommendation model. This recommendation method deeply associates users and hotel room types through user-side text features and product-side fusion features, so that hotel room type recommendations are more in line with user needs and effectively improve user experience.

[0054] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0055] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0056] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data for analysis, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0057] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0058] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0059] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A method for recommending hotel room types, characterized in that: The method comprises: Obtaining user-side text features based on user history data; the user history data includes voice data and evaluation text data; Obtaining commodity-side fusion features based on hotel room type data; the hotel room type data includes hotel room type image data and hotel room type introduction text data; Obtaining multiple groups of feature relationship pairs between the user-side text features and the product-side fusion features according to the user-side text features and the product-side fusion features; Using the multiple groups of feature relationship pairs as training sets to train the multi-task learning model to obtain a recommendation model; Complete hotel room type recommendations for users based on the recommendation model.

2. The recommendation method according to claim 1, characterized in that: The product-side fusion features obtained based on the hotel room type data include: Extracting an image of a ROI feature area of ​​a hotel room type from the hotel room type image data to obtain a product side image feature; Obtaining the product side image features to be fused based on the ResNet50 model according to the product side image features; the product side image features to be fused are vector features of a first preset dimension; Obtaining the product-side text features to be integrated based on the XL-Net model according to the hotel room type introduction text data; the product-side text features to be integrated are vector features of the second preset dimension; The image features of the product to be fused and the text features of the product to be fused are converted into the same dimension through the linear layer processing of the DNN model; Aligning the dimensions of the converted image feature on the product side to be fused with the converted text feature on the product side to be fused, and adding the values ​​of the corresponding dimensions to obtain a first fusion feature; The first fusion feature is downsampled using an average pooling layer to obtain a second fusion feature; According to the product side image features to be fused, downsampling is performed through a maximum pooling layer to obtain a third fusion feature; The commodity-side fusion feature is obtained through a decision layer according to the second fusion feature and the third fusion feature.

3. The recommendation method according to claim 2, characterized in that: The hotel room type image data includes key frame images in the hotel room type introduction video and promotional images of the hotel room type, and the image of the ROI feature area of ​​the hotel room type is extracted from the hotel room type image data to obtain the product side image features, including: The image of the ROI feature area is extracted from the key frame image and the promotional image by using Faster-RCNN as the product side image feature.

4. The recommendation method according to claim 1, characterized in that: The obtaining of user-side text features based on user historical data includes: Convert the voice data to obtain first text data; The user-side text features are obtained by using the WordPiece method according to the first text data and the evaluation text data.

5. The recommendation method according to claim 1, characterized in that: The obtaining, based on the user-side text features and the commodity-side fusion features, a plurality of feature relationship pairs of the user-side text features and the commodity-side fusion features comprises: The user-side text features and the product-side fusion features are input into the ImageBERT model to obtain multiple groups of feature relationship pairs between the user-side text features and the product-side fusion features.

6. The recommendation method according to claim 1, characterized in that: The multi-task learning model is an MMoE model.

7. The recommendation method according to claim 4, characterized in that: The method of completing the hotel room type recommendation for the user based on the recommendation model includes: When it is necessary to recommend a hotel room type to the current user, obtain the user history data of the current user; Obtaining a current user-side text feature according to the user history data of the current user; The current user-side text features are input into the recommendation model to obtain the target hotel room type.

8. The recommendation method according to claim 7, characterized in that: The method further comprises: Complete the online recall function of hotel room types based on FAISS.

9. A device for recommending hotel room types, characterized in that: The device comprises: A user-side preprocessing module, used to obtain user-side text features based on user history data; the user history data includes voice data and evaluation text data; A product-side preprocessing module, used to obtain product-side fusion features based on hotel room type data; the hotel room type data includes hotel room type image data and hotel room type introduction text data; A feature relationship pair module, used to obtain multiple sets of feature relationship pairs of the user-side text features and the product-side fusion features according to the user-side text features and the product-side fusion features; A training module, used for training a multi-task learning model using the plurality of feature relationship pairs as training sets to obtain a recommendation model; The recommendation module is used to recommend hotel room types to users based on the recommendation model.

Citation Information

Patent Citations

  • Commodity recommendation method and device

    CN115293853A

  • Hotel guest room recommendation method, system and device based on business travel and storage medium

    CN116911947A

  • Vector shift-based recommendation method, apparatus, computer device, and non-volatile readable storage medium

    WO2021051515A1