Hotel personalized recommendation method and system, storage medium and computer
By generating unique semantic IDs for hotels and using Transformer model to predict user historical choices, the problem of inefficiency of existing hotel recommendation systems is solved, and personalized and automated hotel recommendations are achieved.
Patent Information
- Application Number
- CN202411700315.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-11-26
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing hotel recommendation system relies on manual customization rules, is inefficient and cannot track changes in user preferences in real time, and the recommendation results are not highly personalized.
Generate a unique hotel semantic ID based on the hotel's relevant information, and predict the hotel data selected by the user's historically through the Transformer generative model to achieve personalized recommendations.
实现了快速、准确、个性化的酒店推荐,自动学习用户偏好,无需人工制定复杂规则,推荐结果更加符合用户需求。
Smart Images

Figure CN120298065A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hotel data processing, and in particular to a hotel personalized recommendation method, system, storage medium and computer. Background Art
[0002] With the continuous improvement of people's living standards, more and more people begin to choose to travel to enrich their cultural life and experience the local customs and practices. Accommodation is an indispensable part of the travel process, and a good accommodation experience will add happiness to the whole journey. Therefore, it is very meaningful to find a hotel that meets one's own expectations.
[0003] At present, major tourism companies and platforms have launched hotel search and recommendation functions. However, these systems mainly rely on manually customized rules for matching, with high rule maintenance costs. They need to rely on expert experience to formulate and update rules, resulting in problems such as low efficiency, inability to track changes in user preferences in real time, narrow coverage, and low personalization degree of recommendation results. Summary of the Invention
[0004] Based on this, the purpose of the present invention is to provide a hotel personalized recommendation method, system, storage medium and computer to solve the technical problems existing in the prior art.
[0005] The present invention proposes a hotel personalized recommendation method, including:
[0006] Generating a unique hotel semantic ID for each hotel based on the relevant information of the hotel to form a hotel semantic ID list;
[0007] Obtaining the hotel data selected by the user within a period of time, constructing a historical hotel semantic ID sequence according to the hotel data, and inputting the historical hotel semantic ID into a Transformer model for prediction processing to obtain a predicted hotel semantic ID list;
[0008] Traversing and matching each predicted hotel semantic ID in the predicted hotel semantic ID list with the hotel semantic ID list, and obtaining a corresponding actual hotel list according to the matching result;
[0009] Recommending the actual hotel list to the user for selection to complete the personalized recommendation of the hotel.
[0010] The beneficial effects of the present invention are as follows: The hotel personalized recommendation method provided by the present invention first generates a unique hotel semantic ID for each hotel based on the relevant information of the hotel to form a hotel semantic ID list for convenient model processing. Then, it obtains the hotel data selected by the user within a certain period of time, constructs a historical hotel semantic ID sequence according to the hotel data, and inputs the historical hotel semantic ID into the Transformer model for prediction processing to obtain a predicted hotel semantic ID list. Based on the hotel data selected by the user historically and combined with the method of the Transformer generative model, personalized hotel recommendation is realized. Compared with manual rule-based recommendation, the Transformer generative model also has stronger generalization ability. This method can automatically learn user preferences without the need for manual formulation of complex rules, and the recommendation results are more personalized, achieving fast, accurate, and personalized hotel recommendation results, providing a better experience for users and having broad application value.
[0011] Preferably, a unique hotel semantic ID is generated for each hotel based on the relevant information of the hotel to form a hotel semantic ID list;
[0012] Obtain the hotel data selected by the user within a certain period of time, construct a historical hotel semantic ID sequence according to the hotel data, and input the historical hotel semantic ID sequence into the Transformer model for prediction processing to obtain a predicted hotel semantic ID list;
[0013] Traverse and match each predicted hotel semantic ID in the predicted hotel semantic ID list with the hotel semantic ID list, and obtain the corresponding actual hotel list according to the matching result;
[0014] Recommend the actual hotel list to the user for selection to complete the personalized recommendation of the hotel.
[0015] Preferably, the step of generating a unique hotel semantic ID for each hotel based on the relevant information of the hotel includes:
[0016] Obtain the relevant information of each hotel, and construct the corresponding text according to the relevant information of the hotel, where the relevant information of the hotel at least includes information such as the location, room type, price, and policy of the hotel;
[0017] The constructed text is successively subjected to text encoding, dimensionality reduction, and vectorization processing through a hotel semantic ID generation model to obtain an ordered tuple containing multiple code words, and the ordered tuple is used as the semantic ID of the hotel.
[0018] Preferably, the step of successively subjecting the constructed text to text encoding, dimensionality reduction, and vectorization processing through a hotel semantic ID generation model to obtain an ordered tuple containing multiple code words and using the ordered tuple as the semantic ID of the hotel includes:
[0019] Input the constructed text data into the text encoder to obtain the semantic embedding vectors corresponding to each hotel;
[0020] Input the semantic embedding vectors into a dimensionality reduction encoder for dimensionality reduction processing to reduce the number of elements in the semantic embedding vectors;
[0021] Construct a quantization module including several quantization layers, divide the elements in the dimensionality-reduced semantic embedding vectors into several codebooks corresponding to the quantization layers in the quantization module, use the codebook corresponding to the first quantization layer as the input to the first quantization layer of the quantization module for vector quantization processing, and obtain the vector closest to the semantic embedding vector in the codebook and the corresponding codeword as the output;
[0022] Calculate the residual between the output and the input of the first quantization layer in the quantization module, and use the residual and the codebook corresponding to the second quantization layer as the input to the second quantization layer of the quantization module for processing;
[0023] And so on, obtain the codewords corresponding to the semantic embedding vectors in all codebooks to obtain an ordered tuple containing multiple codewords, and use the ordered tuple as the semantic ID of the hotel.
[0024] Preferably, before performing text encoding, dimensionality reduction, and vector quantization processing on the constructed text through the hotel semantic ID generation model, the hotel personalized recommendation method further includes: training the hotel semantic ID generation model, and the steps for training the hotel semantic ID generation model are as follows:
[0025] Obtain the relevant information of several specific hotels, and batch the several specific hotels as the training set;
[0026] Select a batch of hotels in the training set as the first training hotel, freeze the parameters of the text encoder, and input the relevant information of the first training hotel into the original hotel semantic ID generation model for processing;
[0027] Obtain the comprehensive difference loss between the output and the input of the original hotel semantic ID generation model, and update the hotel semantic ID generation model according to the comprehensive difference loss;
[0028] Obtain another batch of hotels in the training set as the second training hotel, and input the relevant information of the second training hotel into the hotel semantic ID generation model updated once for processing;
[0029] Iteratively update in turn until the comprehensive difference loss between the output and the input of the hotel semantic ID generation model meets the preset requirements and then stop the iteration, and perform dimensionality increase processing through a dimensionality increase decoder to complete the training of the hotel semantic ID generation model.
[0030] Preferably, the similarity between different hotels is positively correlated with the similarity of the hotel semantic IDs corresponding to different hotels.
[0031] Preferably, the steps of obtaining the hotel data selected by the user within a period of time, constructing a historical hotel semantic ID sequence according to the hotel data, and inputting the historical hotel semantic ID into a Transformer model for prediction processing to obtain a list of predicted hotel semantic IDs include:
[0032] Obtain the hotels that the user has stayed in or browsed within a period of time, and organize the corresponding hotel semantic IDs into a historical hotel semantic ID sequence;
[0033] Convert the historical hotel semantic ID into a token sequence of code words, and input the token sequence into the Transformer model for a first prediction to obtain a first predicted hotel semantic ID;
[0034] Use the first predicted hotel semantic ID and the original historical hotel semantic ID sequence as a new historical hotel semantic ID, convert it into a new token sequence, and then input it back into the Transformer model for prediction to obtain a second predicted hotel semantic ID;
[0035] Iterate through a preset cycle in sequence, sort the obtained several predicted hotel semantic IDs in chronological order of prediction, and obtain the list of predicted hotel semantic IDs.
[0036] Preferably, before inputting the historical hotel semantic ID into the Transformer model for a first prediction, the hotel personalized recommendation method further includes: training the Transformer model, and the steps of training the Transformer model are:
[0037] Obtain the hotels that several users have stayed in or browsed, and organize the hotel semantic IDs corresponding to all hotels in chronological order to form several training hotel semantic ID sequences corresponding to the users;
[0038] Select the training hotel semantic ID sequence corresponding to one of the users and input it into the original Transformer model for processing, and perform an iterative operation. The steps of the iterative operation are: match the predicted hotel semantic ID output by the original Transformer model with the hotel semantic ID actually selected by the user, obtain the matching difference between the predicted hotel semantic ID output by the original Transformer model and the hotel semantic ID of the hotel where the user actually stays, and update the original Transformer model according to the matching difference to complete one iteration;
[0039] Select the sequence of the training hotel semantic IDs corresponding to other users in sequence, and input them into the iterated Transformer model for processing in sequence, and repeat the iteration operation;
[0040] Stop the iteration until the predicted hotel semantic ID output by the Transformer model matches the hotel semantic ID actually selected by the corresponding user to meet the preset requirements, and complete the training of the Transformer model.
[0041] The present invention also provides a hotel personalized recommendation system, including:
[0042] A forming module, configured to generate a unique hotel semantic ID for each hotel based on the relevant information of the hotel to form a hotel semantic ID list;
[0043] A prediction module, configured to obtain the hotel data selected by a user within a period of time, construct a historical hotel semantic ID sequence according to the hotel data, input the historical hotel semantic ID into the Transformer model for prediction processing, and obtain a predicted hotel semantic ID list;
[0044] A matching module, configured to traverse and match each predicted hotel semantic ID in the predicted hotel semantic ID list with the hotel semantic ID list, and obtain a corresponding actual hotel list according to the matching result;
[0045] A recommendation module, configured to recommend the actual hotel list to the user for selection to complete the personalized recommendation of the hotel.
[0046] The present invention also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned hotel personalized recommendation method is implemented.
[0047] The present invention also provides a computer, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the above-mentioned hotel personalized recommendation method is implemented.
[0048] Additional aspects and advantages of the present invention will be given in part in the following description, will become apparent in part from the following description, or will be understood through the practice of the present invention. Description of the Drawings
[0049] Figure 1 It is a flowchart of the hotel personalized recommendation method in the first embodiment of the present invention;
[0050] Figure 2 It is a structural block diagram of the hotel personalized recommendation system in the third embodiment of the present invention;
[0051] Figure 3 This is a structural block diagram of a computer in the seventh embodiment of the present invention.
[0052] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. Specific Embodiments
[0053] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0055] Embodiment 1
[0056] Please refer to Figure 1 , which shows a hotel personalized recommendation method in the first embodiment of the present invention. The hotel personalized recommendation method specifically includes steps S10 to S40:
[0057] S10. Generate a unique hotel semantic ID for each hotel based on the relevant information of the hotel to form a hotel semantic ID list;
[0058] In specific implementation, first generate a unique hotel semantic ID for each hotel based on the relevant information of the hotel. The IDs of similar hotels can share some code words, that is, the more the same code words there are for similar hotels, the corresponding hotel can be directly found through the hotel semantic ID. The relevant information of the hotel includes, but is not limited to, information such as the location, room type, price, and policy of the hotel;
[0059] Optionally, in this embodiment, the step of generating a unique hotel semantic ID for each hotel based on the relevant information of the hotel is:
[0060] Obtain the relevant information of each hotel and construct the corresponding text according to the relevant information of the hotel; first construct the corresponding text according to the relevant information of the hotel to facilitate subsequent processing.
[0061] Input the constructed text data into the Sentence-T5 text encoder to obtain the semantic embedding vectors corresponding to each hotel; the Sentence-T5 text encoder is an existing encoder on the market and can be pre-trained before use, which will not be elaborated here. The semantic embedding vectors generated after the Sentence-T5 text encoder can be understood as an array, which contains many elements. After being processed by the Sentence-T5 text encoder, each element is represented in decimal form, and the value of each element is related to the relevant information of the hotel.
[0062] Input the semantic embedding vectors into the RQ-VAE model for dimensionality reduction and vector quantization processing in sequence to obtain an ordered tuple containing multiple codewords, and use the ordered tuple as the semantic ID of the hotel.
[0063] Specifically, the RQ-VAE model includes at least one dimensionality reduction encoder and a quantization module. The dimensionality reduction encoder can perform dimensionality reduction processing on the semantic embedding vectors to reduce the number of elements in the array. The quantization module performs vector quantization processing on each element in the dimensionality-reduced semantic embedding vectors to find the closest codeword for each element, that is, perform rounding processing on each element represented in decimal form to obtain the ordered tuple. And use the obtained ordered tuple as the hotel semantic ID. Among them, the similarity between different hotels is positively correlated with the similarity of the hotel semantic IDs corresponding to different hotels, that is, a higher similarity of hotel semantic IDs means a higher similarity between hotels.
[0064] Optionally, in this embodiment, before inputting the semantic embedding vectors into the RQ-VAE model for dimensionality reduction and vector quantization processing in sequence, the hotel personalized recommendation method further includes: training the RQ-VAE model. The steps for training the RQ-VAE model are as follows:
[0065] Obtain the relevant information of several specific hotels, and batch the several specific hotels as the training set; select a batch of hotels in the training set as the first training hotel, and after processing the relevant information of the first training hotel through the Sentence-T5 text encoder, obtain the corresponding one-dimensional array x (x1, x2...), and use x as the input and input it into the original RQ-VAE model for processing to obtain the corresponding one-dimensional array output The RQ-VAE model also includes a reconstruction decoder. Use the RQ-VAE model to Reconstruct the output, obtain the comprehensive difference loss between the output and the input of the original RQ-VAE model, and update the original RQ-VAE model according to the comprehensive difference loss; specifically, the update algorithm can use the existing gradient descent algorithm for updating, which is equivalent to performing an iteration on the original RQ-VAE model.
[0066] Among them, the comprehensive difference loss includes the reconstruction loss L recon and the quantization loss L rqvae The reconstruction loss L recon and the quantization loss L rqvae and the functional expressions of the comprehensive difference loss L are respectively:
[0067]
[0068]
[0069] L = L recon + L rqvae
[0070] In the formula, x is the input parameter of the RQ-VAE model; is the output parameter of the RQ-VAE model; sg[r i means that r i does not provide gradients; r i is the residual; is the embedding vector in the custom codebook; β is the custom weight parameter; d is the d-th iteration during training; m is the total number of iterations during training.
[0071] After one iteration, another batch of hotels in the training set is obtained as the second training hotel, and the relevant information of the second training hotel is input into the RQ-VAE model updated once for processing; then the comprehensive difference loss between the output and the input of the RQ-VAE model is obtained, and the RQ-VAE model is updated according to the second comprehensive difference loss; this is equivalent to performing a second iteration on the original RQ-VAE model. Iterations are carried out in this way until the comprehensive difference loss between the output and the input of the RQ-VAE model meets the preset requirements and then the iteration stops, completing the training of the RQ-VAE model. The specific comprehensive difference loss can be specifically set according to the time requirements and accuracy requirements of the RQ-VAE model. For example, it can be set that the comprehensive difference loss is less than 5% as the preset requirement, and no specific restrictions are made here.
[0072] S20. Obtain the hotel data selected by the user within a period of time, construct a historical hotel semantic ID sequence according to the hotel data, and input the historical hotel semantic ID sequence into the Transformer model for prediction processing to obtain a predicted hotel semantic ID list.
[0073] In specific implementation, the steps of the predicted hotel semantic ID list include:
[0074] Retrieve the hotels that the user has stayed in or browsed within a certain period of time, and organize the corresponding hotel semantic IDs into a historical hotel semantic ID sequence. Specifically, hotel data that the user has stayed in within the most recent year can be retrieved, and the hotel semantic IDs corresponding to the hotels that the user has stayed in can be organized into a historical hotel semantic ID sequence. In some alternative embodiments, hotels that the user has browsed can also be selected, and the hotel semantic IDs corresponding to the hotels that the user has browsed can be organized into a historical hotel semantic ID sequence for subsequent operations.
[0075] In this embodiment, eight hotels that the user has stayed in are selected, and the hotel semantic IDs corresponding to the eight hotels are organized into a historical hotel semantic ID sequence. The historical hotel semantic ID is converted into a sequence of codeword tags, and the sequence of tags is input into the Transformer model for a first prediction to obtain a first predicted hotel semantic ID. The encoder in the Transformer model mainly processes the sequence of codeword tags. Therefore, before processing, the historical hotel semantic ID needs to be converted into a sequence of codeword tags and then input into the Transformer model for prediction. The process of obtaining the predicted hotels that have been browsed is similar and will not be elaborated here.
[0076] Furthermore, in this embodiment, before inputting the historical hotel semantic ID into the Transformer model for a first prediction, the hotel personalized recommendation method further includes: training the Transformer model. The steps for training the Transformer model are as follows:
[0077] Retrieve several hotels that the user has stayed in or browsed, and organize the hotel semantic IDs corresponding to all the hotels in chronological order to form several training hotel semantic ID sequences corresponding to the user.
[0078] Specifically, the several training hotel semantic ID sequences will be denoted as A1 / A2 / A3…, B1 / B2 / B3…, C1 / C2 / C3… respectively, where A, B, and C can be understood as different users. Optionally, the number of users may be more, which will not be elaborated here. A1 / A2 / A3… represents the sequence of hotel semantic IDs of several hotels that user A has stayed in, B1 / B2 / B3… represents the sequence of hotel semantic IDs of several hotels that user B has stayed in, and C1 / C2 / C3… represents the sequence of hotel semantic IDs of several hotels that user C has stayed in. The number of hotels can be the same or different. A1 / A2 / A3… are sorted in chronological order of the stay time, that is, the hotel corresponding to A1 is the first hotel that user A stayed in within a certain period of time, followed by the hotel corresponding to A2, the hotel corresponding to A3, and so on.
[0079] Select the training hotel semantic ID sequence corresponding to one of the users and input it into the original Transformer model for processing, and perform iterative operations: match the output hotel semantic ID in the original Transformer model with the hotel semantic ID actually selected by the user, obtain the matching difference between the input and output of the original Transformer model, and update the original Transformer model according to the matching difference to complete one iteration;
[0080] Specifically, in this embodiment, first select the hotels A1 / A2 / A3... where user A has stayed and input them into the original Transformer model for processing. Match the predicted semantic ID of the original Transformer model with the hotel semantic ID actually selected by the user, find the difference between them and perform the matching, and perform iterative operations. The specific steps of the iterative operations can be: take A1 as the test sample, A2 as the validation sample, input A1 into the original Transformer model for processing, match the obtained predicted hotel semantic ID with A2, and obtain the matching difference 1; take A1 and A2 as the test samples, A3 as the validation sample, input A1 and A2 into the original Transformer model for processing, match the obtained predicted hotel semantic ID with A3, and obtain the matching difference 2; take A1, A2, and A3 as the test samples, A4 as the validation sample, input A1, A2, and A3 into the original Transformer model for processing, match the obtained predicted hotel semantic ID with A4, and obtain the matching difference 3. Take A1, A2, A3, and A4 as the test samples, A5 as the validation sample, input A1, A2, A3, and A4 into the original Transformer model for processing, match the obtained predicted hotel semantic ID with A5, and obtain the matching difference 4, and so on, to obtain several matching differences. Take the average of the obtained several matching differences, that is, obtain the matching difference of inputting the hotel data where user A has stayed into the original Transformer model, and update the original Transformer model according to the matching difference of user A to complete one iteration. That is, the hotel data where user A has stayed is sampled to train the original Transformer model once. In this embodiment, by traversing each user sequence, each hotel can appear as a validation or test sample, achieving the effect of cross-validation.
[0081] Select the corresponding training hotel semantic ID sequences of other users in turn and input them into the Transformer model after one iteration for processing, and repeat the iterative operations in turn;
[0082] Specifically, in this embodiment, after training the original Transformer model with the hotel data where User A stayed, the above iterative operation is performed again with the hotel data where User B stayed. The hotel data where User B stayed is input into the Transformer model after the first iteration for the second iteration, and the hotel data where User C stayed is input into the Transformer model after the second iteration for the third iteration, and so on.
[0083] Iteration stops until the predicted hotel semantic ID output by the Transformer model matches the hotel semantic ID actually selected by the corresponding user and meets the preset requirements, and the training of the Transformer model is completed.
[0084] Specifically, in the training stage, the maximum length of each user sequence is 20 to reduce the computational complexity; optionally, in this embodiment, when the Transformer model goes through several iterations, if the predicted hotel semantic ID output by the Transformer model is the same as the hotel semantic ID actually selected by the user, the iteration stops, and the training of the Transformer model is completed. The trained Transformer model is used as the model for subsequent hotel ID prediction.
[0085] Optionally, in this embodiment, after inputting the token sequence of the code words corresponding to eight hotels into the Transformer model for a first prediction, the decoder in the Transformer model can automatically regress to generate a predicted hotel semantic ID, denoted as the first predicted hotel semantic ID.
[0086] The first predicted hotel semantic ID and the original historical hotel semantic ID sequence are used as the new historical hotel semantic ID. After being converted into a new token sequence, they are re-input into the Transformer model for prediction to obtain the second predicted hotel semantic ID; specifically, in this embodiment, the semantic IDs of the original eight hotels where the user stayed and the first predicted hotel semantic ID just predicted, a total of nine hotels, are used as the new historical hotel semantic ID and re-input into the Transformer model for a second prediction, and then a predicted hotel semantic ID is generated again, denoted as the second predicted hotel semantic ID. The above prediction operations are performed in sequence to obtain the third predicted hotel semantic ID, the fourth predicted hotel semantic ID, and so on.
[0087] Iterate through the preset cycle in sequence, sort the obtained several predicted hotel semantic IDs in order of time of prediction, and obtain a list of predicted hotel semantic IDs. Specifically, the iteration cycle can be determined according to accuracy and experience. The later the prediction result, the farther it is from the user's expectation. Therefore, in this embodiment, the number of iterations is five. The accuracies of the first predicted hotel semantic ID, the second predicted hotel semantic ID, the third predicted hotel semantic ID, the fourth predicted hotel semantic ID, and the fifth predicted hotel semantic ID obtained in sequence are worse. According to the five predicted hotel semantic IDs, a list of predicted hotel semantic IDs is obtained, that is, the list of predicted hotel semantic IDs is a list of hotel semantic IDs that the user may be satisfied with; among them, the first predicted hotel semantic ID has the highest possibility of satisfying the user's needs, and the fifth predicted hotel semantic ID has the lowest possibility of satisfying the user's needs.
[0088] S30. Traverse and match each predicted hotel semantic ID in the list of predicted hotel semantic IDs with the list of hotel semantic IDs, and obtain the corresponding actual hotel list according to the matching result.
[0089] After generating the list of predicted hotel semantic IDs in step S20, match the hotel semantic IDs in the generated list of predicted hotel semantic IDs back to specific hotels, that is, match them with the list of hotel semantic IDs in step S10. Match the hotel semantic IDs in the list of predicted hotel semantic IDs with the list of specific hotel semantic IDs respectively, find the hotel semantic IDs that match the predicted hotel semantic IDs, and then obtain the matching actual hotel list. In this way, the predicted semantic IDs can be very efficiently matched to specific hotels by directly looking up the table, completing the retrieval process within the Transformer model. Specifically, the actual hotel list is sorted according to the possibility of meeting the user's expectation.
[0090] Further, since the number of elements in the generated hotel semantic ID and the predicted hotel semantic ID may be different. In this embodiment, the hotel semantic ID predicted by the Transformer model has one less element than the hotel semantic ID generated by the RQ-VAE model. For example, the hotel semantic ID predicted by the Transformer model is generally a one-dimensional array, and the one-dimensional array contains 3 elements, denoted as [25.1.6] for example. The hotel semantic ID generated by the RQ-VAE model may include [25.1.6.0], [25.1.6.1], [25.1.6.2]... Then when the hotel semantic ID predicted by the Transformer model is [25.1.6], the system will automatically identify the hotels corresponding to [25.1.6.0], [25.1.6.1], [25.1.6.2]... It can be understood that when the hotel semantic ID generated by the RQ-VAE model only includes [25.1.6.0], it is simply denoted as [25.1.6].
[0091] S40, recommend the actual hotel list to the user for selection to complete the personalized recommendation of hotels.
[0092] Embodiment 2
[0093] This embodiment also provides a method for personalized hotel recommendation. The difference between this embodiment and the method for personalized hotel recommendation in Embodiment 1 is that:
[0094] The step of generating a unique hotel semantic ID for each hotel based on the relevant information of the hotel includes:
[0095] Obtain the relevant information of each hotel, and construct the corresponding text according to the relevant information of the hotel. Among them, the relevant information of the hotel at least includes the information of the location, room type, price, and policy of the hotel;
[0096] Perform text encoding, dimensionality reduction, and vector quantization processing on the constructed text in sequence through the hotel semantic ID generation model to obtain an ordered tuple containing multiple code words, and use the ordered tuple as the hotel semantic ID.
[0097] In this embodiment, the hotel semantic ID generation model includes a text encoder, a dimensionality reduction encoder, and a quantization module. The quantization layer of the quantization module has three layers.
[0098] The step of performing text encoding, dimensionality reduction, and vector quantization processing on the constructed text in sequence through the hotel semantic ID generation model to obtain an ordered tuple containing multiple code words, and using the ordered tuple as the hotel semantic ID includes:
[0099] Input the constructed text data into the text encoder to obtain the semantic embedding vector corresponding to each hotel;
[0100] Input the semantic embedding vector into a dimensionality reduction encoder for dimensionality reduction processing to reduce the number of elements in the semantic embedding vector;
[0101] Construct a quantization module including several quantization layers, divide the elements in the dimensionality-reduced semantic embedding vector into several codebooks corresponding to the quantization layers in the quantization module, use the codebook corresponding to the first quantization layer as the input of the first quantization layer of the quantization module, perform vector quantization processing, and obtain the vector closest to the semantic embedding vector in the codebook and the corresponding codeword as the output;
[0102] Calculate the residual between the output and the input of the first quantization layer in the quantization module, and use the residual and the codebook corresponding to the second quantization layer as the input of the second quantization layer of the quantization module for processing;
[0103] And so on, obtain the codewords corresponding to the semantic embedding vector in all codebooks to obtain an ordered tuple containing multiple codewords, and use the ordered tuple as the semantic ID of the hotel.
[0104] Among them, the text encoder uses Sentence-T5, and the dimensionality reduction encoder can use a deep neural network model (DNN).
[0105] Before performing text encoding, dimensionality reduction, and vector quantization processing on the constructed text through the hotel semantic ID generation model, the hotel personalized recommendation method further includes: training the hotel semantic ID generation model, and the steps for training the hotel semantic ID generation model are:
[0106] Obtain the relevant information of several specific hotels, and batch the several specific hotels as the training set;
[0107] Select a batch of hotels in the training set as the first training hotel, freeze the parameters of the text encoder, and input the relevant information of the first training hotel into the original hotel semantic ID generation model for processing;
[0108] Obtain the comprehensive difference loss between the output and the input of the original hotel semantic ID generation model, and update the hotel semantic ID generation model according to the comprehensive difference loss;
[0109] Obtain another batch of hotels in the training set as the second training hotel, and input the relevant information of the second training hotel into the hotel semantic ID generation model updated once for processing;
[0110] Perform iterative updates in turn until the comprehensive difference loss between the output and the input of the hotel semantic ID generation model meets the preset requirements and then stop the iteration, and perform dimensionality increase processing through a dimensionality increase decoder to complete the training of the hotel semantic ID generation model.
[0111] In this embodiment, the hotel semantic ID generation model further includes an upsampling decoder. During the training iteration of the hotel semantic ID generation model, the parameters of the text encoder can be frozen, and only the parameters of the encoder, quantization module, and decoder are updated. The upsampling decoder can also adopt a deep neural network model (DNN); the training process is similar to the training process of the RQ-VAE model in Embodiment 1 and will not be elaborated here. Optionally, in this embodiment, when training the Transformer model, the sequences of all hotels clicked and stayed by each user since registration / for one year / half a year / three months / one month can be input into the model for training; 10-20 hotels can be inferred for final display.
[0112] Embodiment 3
[0113] This embodiment provides a hotel personalized recommendation system, as Figure 2 shown, the hotel personalized recommendation system includes:
[0114] A formation module 101, configured to generate a unique hotel semantic ID for each hotel based on relevant information of the hotel to form a hotel semantic ID list;
[0115] The formation module 101 includes:
[0116] A construction unit, configured to obtain relevant information of each hotel and construct a corresponding text according to the relevant information of the hotel, where the hotel relevant information at least includes information on the location, room type, price, and policy of the hotel;
[0117] A first input unit, configured to input the constructed text data into a Sentence-T5 text encoder to obtain a semantic embedding vector corresponding to each hotel;
[0118] A processing unit, configured to input the semantic embedding vector into an RQ-VAE model to perform dimensionality reduction and vector quantization processing in sequence to obtain an ordered tuple containing multiple codewords, and use the ordered tuple as the hotel semantic ID, where the similarity between different hotels is positively correlated with the similarity of the hotel semantic IDs corresponding to different hotels.
[0119] The RQ-VAE model at least includes a dimensionality reduction encoder and a quantization module;
[0120] The processing unit includes:
[0121] A dimensionality reduction subunit, configured to perform dimensionality reduction processing on the semantic embedding vector through the dimensionality reduction encoder to reduce the number of elements in the semantic embedding vector;
[0122] A quantization subunit, configured to perform vector quantization processing on each element in the dimension-reduced semantic embedding vector through the quantization module, find the closest codeword for each element, so as to obtain the ordered tuple.
[0123] A prediction module 201, configured to obtain hotel data selected by a user within a period of time, construct a historical hotel semantic ID sequence according to the hotel data, input the historical hotel semantic ID sequence into a Transformer model for prediction processing, and obtain a predicted hotel semantic ID list;
[0124] The prediction module 201 includes:
[0125] An arrangement unit, configured to obtain hotels that the user has stayed in or browsed within a period of time, and arrange the corresponding hotel semantic IDs into a historical hotel semantic ID sequence;
[0126] A second input unit, configured to convert the historical hotel semantic ID into a sequence of codeword tags, input the tag sequence into the Transformer model for a first prediction, and obtain a first predicted hotel semantic ID;
[0127] A third input unit, configured to use the first predicted hotel semantic ID and the original historical hotel semantic ID sequence as a new historical hotel semantic ID, convert them into a new tag sequence, and then re-input them into the Transformer model for prediction to obtain a second predicted hotel semantic ID;
[0128] A sorting unit, iterating a preset period in sequence, sorting the obtained several predicted hotel semantic IDs in chronological order of prediction to obtain the predicted hotel semantic ID list.
[0129] A matching module 301, configured to traverse and match each predicted hotel semantic ID in the predicted hotel semantic ID list with the hotel semantic ID list, and obtain a corresponding actual hotel list according to the matching result;
[0130] A recommendation module 401, configured to recommend the actual hotel list to the user for selection to complete personalized hotel recommendation.
[0131] Embodiment 4
[0132] This embodiment also provides a hotel personalized recommendation system. The difference between this embodiment and the hotel personalized recommendation system in Embodiment 3 is that it includes
[0133] A first training module, configured to train the RQ-VAE model before inputting the semantic embedding vector into the RQ-VAE model for dimension reduction and vector quantization processing in sequence.
[0134] The steps for training the RQ-VAE model are as follows:
[0135] Obtain the relevant information of several specific hotels, and batch the several specific hotels as the training set;
[0136] Select a batch of hotels in the training set as the first training hotel, and input the relevant information of the first training hotel into the original RQ-VAE model for processing;
[0137] Obtain the comprehensive difference loss between the output and the input of the original RQ-VAE model, and update the original RQ-VAE model according to the comprehensive difference loss;
[0138] Obtain another batch of hotels in the training set as the second training hotel, and input the relevant information of the second training hotel into the RQ-VAE model updated once for processing;
[0139] Iterate in sequence until the comprehensive difference loss between the output and the input of the RQ-VAE model meets the preset requirements, then stop the iteration and complete the training of the RQ-VAE model.
[0140] The second training module is used to train the Transformer model before inputting the historical hotel semantic ID into the Transformer model for a first prediction. The steps for training the Transformer model are as follows:
[0141] Obtain several hotels that the user has ever stayed in or browsed, and sort the hotel semantic IDs corresponding to all hotels in chronological order to form several training hotel semantic ID sequences corresponding to the user;
[0142] Select the training hotel semantic ID sequence corresponding to one of the users and input it into the original Transformer model for processing, and perform an iterative operation: match the predicted hotel semantic ID output in the original Transformer model with the hotel semantic ID actually selected by the user, obtain the matching difference between the predicted hotel semantic ID output in the original Transformer model and the hotel semantic ID of the hotel where the user actually stayed, and update the original Transformer model according to the matching difference to complete one iteration;
[0143] Select the training hotel semantic ID sequences corresponding to other users in sequence, and input them into the iterated Transformer model for processing in sequence, and repeat the iterative operation;
[0144] Stop iterating until the predicted hotel semantic ID output by the Transformer model matches the actual hotel semantic ID selected by the corresponding user and meets the preset requirements, and complete the training of the Transformer model.
[0145] Embodiment 5
[0146] This embodiment also provides a hotel personalized recommendation system. The difference between the hotel personalized recommendation system in this embodiment and that in Embodiment 3 is that it includes:
[0147] A generation module, configured to generate a unique hotel semantic ID for each hotel based on the relevant information of the hotel to form a hotel semantic ID list;
[0148] Specifically, the generation module is used for:
[0149] Obtain the relevant information of each hotel, and construct the corresponding text according to the relevant information of the hotel. Among them, the relevant information of the hotel at least includes the information of the location, room type, price, and policy of the hotel;
[0150] Perform text encoding, dimensionality reduction, and vectorization processing on the constructed text in sequence through a hotel semantic ID generation model to obtain an ordered tuple containing multiple code words, and use the ordered tuple as the semantic ID of the hotel.
[0151] The step of performing text encoding, dimensionality reduction, and vectorization processing on the constructed text in sequence through a hotel semantic ID generation model to obtain an ordered tuple containing multiple code words, and using the ordered tuple as the semantic ID of the hotel includes:
[0152] Input the constructed text data into a text encoder to obtain a semantic embedding vector corresponding to each hotel;
[0153] Input the semantic embedding vector into a dimensionality reduction encoder for dimensionality reduction processing to reduce the number of elements in the semantic embedding vector;
[0154] Construct a quantization module including several quantization layers, divide the elements in the dimensionality-reduced semantic embedding vector into several codebooks corresponding to the quantization layers in the quantization module, use the codebook corresponding to the first quantization layer as the input of the first quantization layer of the quantization module, perform vectorization processing, and obtain the vector closest to the semantic embedding vector in the codebook and the corresponding code word as the output;
[0155] Calculate the residual between the output and the input of the first quantization layer in the quantization module, and use the residual and the codebook corresponding to the second quantization layer as the input of the second quantization layer of the quantization module for processing;
[0156] By analogy, obtain the codewords corresponding to the semantic embedding vectors in all codebooks to obtain an ordered tuple containing multiple codewords, and use the ordered tuple as the semantic ID of the hotel.
[0157] A training and learning module, used before the hotel semantic ID generation model performs text encoding, dimensionality reduction, and vectorization processing on the constructed text in sequence. The hotel personalized recommendation method further includes: training the hotel semantic ID generation model. The steps for training the hotel semantic ID generation model are as follows:
[0158] Obtain the relevant information of several specific hotels, and batch the several specific hotels as the training set;
[0159] Select a batch of hotels in the training set as the first training hotel, freeze the parameters of the text encoder, and input the relevant information of the first training hotel into the original hotel semantic ID generation model for processing;
[0160] Obtain the comprehensive difference loss between the output and input of the original hotel semantic ID generation model, and update the hotel semantic ID generation model according to the comprehensive difference loss;
[0161] Obtain another batch of hotels in the training set as the second training hotel, and input the relevant information of the second training hotel into the hotel semantic ID generation model updated once for processing;
[0162] Perform iterative updates in sequence until the comprehensive difference loss between the output and input of the hotel semantic ID generation model meets the preset requirements and then stop the iteration. Perform dimensionality increase processing through a dimensionality increase decoder to complete the training of the hotel semantic ID generation model.
[0163] Embodiment Six
[0164] An embodiment of the present invention proposes a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the hotel personalized recommendation method as described in Embodiment One or Embodiment Two above.
[0165] Embodiment Seven
[0166] The present invention also proposes a computer. Please refer to Figure 3 , as shown in the computer in this embodiment, including a memory 10, a processor 20, and a computer program 30 stored on the memory 10 and executable on the processor 20. When the processor 20 executes the computer program 30, it implements the hotel personalized recommendation method as described in Embodiment One or Embodiment Two above.
[0167] Among them, the memory 10 includes at least one type of storage medium, and the storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. The memory 10 can be an internal storage unit of a computer in some embodiments, such as the hard disk of the computer. The memory 10 can also be an external storage device in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 10 can also include both an internal storage unit of a computer and an external storage device. The memory 10 can be used not only to store application software installed on the computer and various types of data, but also to temporarily store data that has been output or will be output.
[0168] Among them, the processor 20 can be an Electronic Control Unit (ECU, also known as a vehicle computer), a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor or other data processing chips in some embodiments, and is used to run the program code stored in the memory 10 or process data, such as executing an access restriction program, etc.
[0169] It should be noted that Figure 3 The structure shown does not constitute a limitation on the computer. In other embodiments, the computer may include fewer or more components than shown in the figure, or combine certain components, or have different component arrangements.
[0170] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus or device and execute the instructions), or used in combination with these instruction execution systems, apparatus or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by or in combination with an instruction execution system, apparatus or device.
[0171] More specific examples (nonexhaustive list) of computer-readable media include the following: electrical connections (electronic devices) having one or more wirings, portable computer diskettes (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.
[0172] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0173] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as falling within the scope described in this specification.
[0174] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A hotel personalized recommendation method, characterized in that, including; generating a unique hotel semantic ID for each hotel based on the relevant information of the hotel to form a hotel semantic ID list; obtaining the hotel data selected by the user within a period of time, constructing a historical hotel semantic ID sequence according to the hotel data, and inputting the historical hotel semantic ID sequence into a Transformer model for prediction processing to obtain a predicted hotel semantic ID list; traversing and matching each predicted hotel semantic ID in the predicted hotel semantic ID list with the hotel semantic ID list, and obtaining a corresponding actual hotel list according to the matching result; recommending the actual hotel list to the user for selection to complete the personalized recommendation of hotels.
2. The hotel personalized recommendation method according to claim 1, wherein The step of generating a unique hotel semantic ID for each hotel based on the relevant information of the hotel includes: obtaining the relevant information of each hotel, and constructing a corresponding text according to the relevant information of the hotel, where the relevant information of the hotel at least includes information such as the location, room type, price, and policy of the hotel; performing text encoding, dimensionality reduction, and vector quantization processing on the constructed text in sequence through a hotel semantic ID generation model to obtain an ordered tuple containing multiple code words, and using the ordered tuple as the semantic ID of the hotel.
3. The hotel personalized recommendation method according to claim 2, wherein The step of performing text encoding, dimensionality reduction, and vector quantization processing on the constructed text in sequence through a hotel semantic ID generation model to obtain an ordered tuple containing multiple code words, and using the ordered tuple as the semantic ID of the hotel includes: inputting the constructed text data into a text encoder to obtain a semantic embedding vector corresponding to each hotel; inputting the semantic embedding vector into a dimensionality reduction encoder for dimensionality reduction processing to reduce the number of elements in the semantic embedding vector; constructing a quantization module including several quantization layers, dividing the elements in the dimensionality-reduced semantic embedding vector into several codebooks corresponding to the quantization layers in the quantization module, using the codebook corresponding to the first quantization layer as the input of the first quantization layer of the quantization module, performing vector quantization processing, and obtaining the vector closest to the semantic embedding vector in the codebook and the corresponding code word as the output; calculating the residual between the output and the input of the first quantization layer in the quantization module, and using the residual and the codebook corresponding to the second quantization layer as the input of the second quantization layer of the quantization module for processing; and so on, obtaining the code words corresponding to the semantic embedding vector in all codebooks to obtain an ordered tuple containing multiple code words, and using the ordered tuple as the semantic ID of the hotel.
4. The hotel personalized recommendation method according to claim 3, wherein Before performing text encoding, dimensionality reduction, and vector quantization processing on the constructed text in sequence through a hotel semantic ID generation model, the hotel personalized recommendation method further includes: training the hotel semantic ID generation model, and the steps of training the hotel semantic ID generation model are: obtaining the relevant information of several specific hotels, and dividing the several specific hotels into batches as a training set; selecting a batch of hotels in the training set as the first training hotel, freezing the parameters of the text encoder, and inputting the relevant information of the first training hotel into the original hotel semantic ID generation model for processing; Obtain the comprehensive difference loss between the output and the input of the original hotel semantic ID generation model, and update the hotel semantic ID generation model according to the comprehensive difference loss; Obtain another batch of hotels in the training set as the second training hotels, and input the relevant information of the second training hotels into the hotel semantic ID generation model updated once for processing; Perform iterative updates in sequence until the comprehensive difference loss between the output and the input of the hotel semantic ID generation model meets the preset requirements, then stop the iteration, and perform dimensionality increase processing through a dimensionality increase decoder to complete the training of the hotel semantic ID generation model.
5. The hotel personalized recommendation method according to claim 1, wherein The similarity between different hotels is positively correlated with the similarity of the hotel semantic IDs corresponding to different hotels.
6. The hotel personalized recommendation method according to claim 1, characterized in that The steps of obtaining the hotel data selected by the user within a period of time, constructing a historical hotel semantic ID sequence according to the hotel data, and inputting the historical hotel semantic ID into the Transformer model for prediction processing to obtain a list of predicted hotel semantic IDs include: Obtain the hotels that the user has stayed in or browsed within a period of time, and organize the corresponding hotel semantic IDs into a historical hotel semantic ID sequence; Convert the historical hotel semantic ID into a token sequence of code words, and input the token sequence into the Transformer model for a first prediction to obtain a first predicted hotel semantic ID; Use the first predicted hotel semantic ID and the original historical hotel semantic ID sequence as the new historical hotel semantic ID, convert them into a new token sequence, and then re-enter them into the Transformer model for prediction to obtain a second predicted hotel semantic ID; Iterate for a preset period in sequence, and sort the obtained several predicted hotel semantic IDs in chronological order of prediction to obtain the list of predicted hotel semantic IDs.
7. The hotel personalized recommendation method according to claim 6, wherein Before inputting the historical hotel semantic ID into the Transformer model for a first prediction, the hotel personalized recommendation method further includes: performing model training on the Transformer model, and the steps of performing model training on the Transformer model are: Obtain several hotels that the user has stayed in or browsed, and organize the hotel semantic IDs corresponding to all hotels in chronological order to form several training hotel semantic ID sequences corresponding to the user; Select the training hotel semantic ID sequence corresponding to one user and input it into the original Transformer model for processing, and perform iterative operations. The steps of the iterative operations are: match the predicted hotel semantic ID output by the original Transformer model with the hotel semantic ID actually selected by the user, obtain the matching difference between the predicted hotel semantic ID output by the original Transformer model and the hotel semantic ID of the hotel where the user actually stays, and update the original Transformer model according to the matching difference to complete one iteration; Select the sequence of the training hotel semantic IDs corresponding to other users in turn, and input them into the iterated Transformer model for processing in turn, and repeat the iterated operation; Stop the iteration until the predicted hotel semantic ID output by the Transformer model matches the hotel semantic ID actually selected by the corresponding user and meets the preset requirements, and complete the training of the Transformer model.
8. A hotel personalized recommendation system, characterized in that, Including: A forming module, configured to generate a unique hotel semantic ID for each hotel based on relevant information of the hotel to form a hotel semantic ID list; A prediction module, configured to obtain hotel data selected by a user within a period of time, construct a historical hotel semantic ID sequence according to the hotel data, input the historical hotel semantic ID sequence into a Transformer model for prediction processing, and obtain a predicted hotel semantic ID list; A matching module, configured to traverse and match each predicted hotel semantic ID in the predicted hotel semantic ID list with the hotel semantic ID list, and obtain a corresponding actual hotel list according to the matching result; A recommendation module, configured to recommend the actual hotel list to the user for selection to complete personalized recommendation of hotels.
9. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the hotel personalized recommendation method according to any one of claims 1 to 7.
10. A computer, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the hotel personalized recommendation method according to any one of claims 1 to 7.
Citation Information
Cited By
Hotel design method and system based on user portrait vector construction
CN121328316A
Generation method and recommendation method of commodity reasoning identifier and related device
CN121544353A
Model training method, recall link construction method and device, and storage medium
CN121615717A