Artificial Intelligence-Based Hotel Demand Model

JP2024523377A5Pending Publication Date: 2025-06-10ORACLE INT CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023577727
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-08-11
Filing Date
2022-06-09
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Hoteliers face challenges in accurately estimating demand for hotel rooms due to the heterogeneity of customer preferences, leading to ineffective one-size-fits-all revenue management policies, and existing demand forecasting tools often fail to account for the ordering of room and rate code pairs, resulting in suboptimal pricing and marketing strategies.

Method used

A dynamic iterative reconfigurable clustering algorithm is employed to form multiple customer clusters based on characteristics, using a semi-parametric mixture of multinomial logit models to predict room category and rate code combinations, incorporating unsupervised learning techniques like random forests to iteratively update cluster probabilities and weights, thereby generating a demand model that adapts to customer preferences.

Benefits of technology

The solution enables highly accurate prediction of room and service combinations, improving personalized pricing and marketing strategies by up to 4% compared to static clustering, and provides a practical approach for hoteliers to profile guests and optimize revenue through personalized recommendations and pricing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

An embodiment generates a demand model of potential hotel customers for hotel rooms. An embodiment forms a plurality of clusters based on features of the potential hotel customers, each cluster including a corresponding weight and cluster probability. An embodiment generates an initial estimated mixture of multinomial logit ("MNL") models corresponding to each of the plurality of clusters, the mixture of the MNL models including a weighted likelihood function based on the features and the weights. An embodiment determines revised cluster probabilities and updates the weights. An embodiment estimates an updated estimated mixture of the MNL models and maximizes the weighted likelihood function based on the revised cluster probabilities and the updated weights. Based on the updated weights and the updated estimated mixture of the MNL models, an embodiment generates a demand model adapted to predict a selection probability of a room category and rate code combination for a potential hotel customer.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 215,688, filed June 28, 2021, the disclosure of which is incorporated herein by reference.

[0002] Field One embodiment relates generally to computer systems, and more particularly to a computer system that generates an artificial intelligence based hotel demand model. [Background technology]

[0003] Background information Increasing competition in the hotel industry has led hoteliers to explore more innovative revenue management policies such as personalized pricing and recommendations. Over the last few years, hoteliers have come to understand that not all guests are equal and traditional one-size-fits-all policies may not be effective. Hence, hotels need to profile their guests and present them with the right product / service at the right price with the goal of maximizing profits. Summary of the Invention [Means for solving the problem]

[0004] overview An embodiment generates a demand model of potential hotel customers for hotel rooms. An embodiment forms a plurality of clusters based on features of the potential hotel customers, each cluster including a corresponding weight and cluster probability. An embodiment generates an initial estimated mixture of multinomial logit ("MNL") models corresponding to each of the plurality of clusters, the mixture of the MNL models including a weighted likelihood function based on the features and the weights. An embodiment determines revised cluster probabilities and updates the weights. An embodiment estimates an updated estimated mixture of the MNL models and maximizes the weighted likelihood function based on the revised cluster probabilities and the updated weights. Based on the update weights and the updated estimated mixture of the MNL models, an embodiment generates a demand model adapted to predict a selection probability of a room category and rate code combination for a potential hotel customer. [Brief description of the drawings]

[0005] [Figure 1] FIG. 1 is a schematic block diagram of a hotel reservation system according to an embodiment of the present invention. [Diagram 2] FIG. 2 is a block diagram of a computer server / system according to one embodiment of the present invention. [Diagram 3] 3 is a flow diagram of the functionality of the room demand model module of FIG. 2 for generating a room demand model, according to one embodiment. [Figure 4] FIG. 1 illustrates an example of initial clustering according to an embodiment. [Diagram 5] 1 is an example showing various asking prices, room categories, and rate codes. [Figure 6] FIG. 1 illustrates guest cluster selection modeling according to an example embodiment. [Figure 7] FIG. 1 illustrates the initial assignment of MNL models to each cluster, according to an embodiment. [Figure 8] FIG. 1 illustrates a proposed likelihood function to be used with the EM function according to an embodiment. [Figure 9] FIG. 2 illustrates a portion of an EM function according to an embodiment. [Figure 10] FIG. 2 illustrates a portion of an EM function according to an embodiment. [Figure 11] FIG. 1 shows an example of an embodiment of the present invention for three clusters. [Figure 12] FIG. 1 shows an example of an embodiment of the present invention for three clusters. [Figure 13] FIG. 1 shows an example of an embodiment of the present invention for three clusters. [Figure 14] FIG. 1 shows an example of an embodiment of the present invention for three clusters. [Figure 15] FIG. 1 shows an example of an embodiment of the present invention for three clusters. [Figure 16] FIG. 1 shows an example of an embodiment of the present invention for three clusters. [Figure 17] FIG. 13 illustrates a comparison of prediction accuracy by iteration between CCR and MSE according to an embodiment of the present invention. [Figure 18] FIG. 13 illustrates how cluster properties change with iterations given two clusters, according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0006] Detailed Description An embodiment predicts a customer's selection of a hotel room category and associated service type based on estimating parameters of a discrete choice model built on clusters of dynamically determined observations. Each observation corresponds to a choice made by a customer to book a hotel room and select an associated service type from an ordered set of room category and service type pairs offered at a particular price. Each room category and service type is described by a set of features that determine the value or utility of the customer's choice. In addition, each customer is characterized by a unique set of attributes that determine the cluster to which the customer belongs, also known as a "persona type." It is envisioned that each persona type may have its own utility of booking choice.

[0007] The choice probability is modeled as a multinomial logit function based on the utility of room-service pairs for each persona type. The embodiments build the basis for prescriptive analytics applications to optimize personalized presentations by improving prediction accuracy and maximizing expected revenue. The embodiments can be used as a standalone system or as the core part of a personalized hotel room price optimization system and a room category and rate code sequence display optimization system. Instead of using static clustering traditionally used for this purpose, the embodiments utilize iteratively reconfigurable dynamic clustering based on semi-parametric mixtures of discrete choice models to fully reflect customer choice behavior.

[0008] In general, in the hotel industry as well as other comparable industries, increasing competition is driving more innovative revenue management practices such as personalized offers and pricing. Not all customers are the same and traditional one-size-fits-all policies may prove ineffective. Accurate estimation of demand as input to personalized recommendation systems is crucial.

[0009] Embodiments address the need to more accurately estimate demand for hotel rooms by modeling demand taking into account heterogeneous customers who differ in (1) willingness to pay (as indicated by selected price range), (2) rate plan choice (corporate discount, breakfast included, etc.), (3) travel attributes, (4) booking channel, (5) booking window, (6) length of stay, (7) arrival date, and / or (8) group / family size, number of children, etc. Factors influencing choice may include room features, rate plan features, price, and the order in which offers are presented.

[0010] Reference will now be made in detail to the embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. Wherever possible, like reference numbers are used for like elements.

[0011] Figure 1 is a schematic block diagram of a hotel reservation system 100 according to an embodiment of the present invention. Figure 1 includes reservation channels 102 with which potential hotel customers may interact to reserve a hotel room. The channels include Global Distribution Systems ("GDS") 111, including "Amadeus", "Sabre", "Travel Port", etc., Online Travel Agencies ("OTAs") 112, including "Booking.com", "Expedia", etc., metasearch sites 113, and other means through which customers may reserve a hotel room, including websites maintained by hotel chains or individual hotels.

[0012] Each hotel chain operation 104 is accessed by an application programming interface ("API") 140 as a web service, such as "WebLogic Server" by Oracle Corp. The hotel chain operations 104 include a hotel property management system ("PMS") 121, such as "OPERA Cloud Property Management" by Oracle Corp., a hotel central reservation system ("CRS") 122, and a demand modeling module 150 that interfaces with the systems 121, 122 to provide optimized demand modeling as disclosed herein.

[0013] A hotel customer or potential hotel customer using the system 100 to acquire a hotel room typically engages in a three-step reservation process. First, an area availability search is performed. Multiple hotel chains are represented and the hotel CRS 122 provides static data. The static data may include minimum / maximum rates, availability dates, etc.

[0014] If the booking guest selects a hotel, the booking guest proceeds to the next step of performing a property search that includes a single hotel property, multiple rooms, and rate plans. For a single hotel property, the information may include descriptive data for room categories, rate plan descriptions, and room prices, each of which are presented in a particular order. The property search includes real-time availability data that allows the booking guest to select a room. Once a room is selected, the final step is the final reservation, where the reservation is guaranteed by credit card or other form of payment.

[0015] FIG. 2 is a block diagram of a computer server / system 10 according to one embodiment of the present invention. Although shown as a single system, the functionality of system 10 may be implemented as a distributed system. Additionally, the functionality disclosed herein may be implemented on separate servers or devices that may be coupled to each other via a network. Additionally, one or more components of system 10 may not be included. For example, when implemented as a web server or cloud-based functionality, system 10 may be implemented as one or more servers and no user interface such as a display, mouse, etc. is required. In an embodiment, system 10 may be used to implement any of the elements shown in FIG. 1.

[0016] The system 10 includes a bus 12 or other communication mechanism for communicating information, and a processor 22 coupled to the bus 12 for processing information. The processor 22 may be any type of general-purpose or special-purpose processor. The system 10 further includes a memory 14 for storing information and instructions executed by the processor 22. The memory 14 may be comprised of any combination of random access memory ("RAM"), read-only memory ("ROM"), static storage devices such as magnetic or optical disks, or any other type of computer-readable medium. The system 10 further includes a communication device 20, such as a network interface card, to provide access to a network. Thus, a user may interface with the system 10 directly, remotely through a network, or otherwise.

[0017] Computer-readable media may be any available media that can be accessed by processor 22 and includes both volatile and nonvolatile media, removable and non-removable media, and communication media. Communication media may include computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism, and includes any information delivery media.

[0018] The processor 22 is further coupled via the bus 12 to a display 24, such as a liquid crystal display ("LCD"), A keyboard 26 and a cursor control device 28, such as a computer mouse, are further coupled to the bus 12 to allow a user to interface with the system 10.

[0019] In one embodiment, the memory 14 stores software modules that provide functionality when executed by the processor 22. The modules include an operating system 15 that provides operating system functionality for the system 10. The modules further include a room demand model module 16 that generates a room demand model to maximize expected hotel room revenue, and all other functionality disclosed herein. Because variable operating costs for a hotel are relatively small, expected revenue (i.e., the product of room reservation probability and room price) is the primary optimization goal in the embodiment. The system 10 can be part of a larger system. Thus, the system 10 can include one or more additional functional modules 18 to include additional functionality, such as functionality of a property management system ("PMS") (e.g., "Oracle Hospitality OPERA Property" or "Oracle Hospitality OPERA Cloud Services") or an enterprise resource planning ("ERP") system. A database 17 is coupled to the bus 12 and provides centralized storage for the modules 16, 18, storing guest data, hotel data, transaction data, and the like. In one embodiment, database 17 is a relational database management system ("RDBMS") capable of using Structured Query Language ("SQL") to manage stored data.

[0020] In one embodiment, especially with a large number of hotel locations, a large number of guests, and a large amount of historical data, database 17 is implemented as an in-memory database ("IMDB"). An IMDB is a database management system that relies primarily on main memory for computer data storage. This is in contrast to database management systems that use disk storage mechanisms. Main memory databases are faster than disk-optimized databases because disk access is slower than memory access, and the internal optimization algorithms are simpler and execute fewer CPU instructions. Accessing data in memory eliminates seek times when querying data, resulting in faster and more predictable performance than disk.

[0021] In one embodiment, the database 17, when implemented as an IMDB, is implemented based on a distributed data grid. A distributed data grid is a system in which a collection of computer servers cooperate in one or more clusters to manage information and related operations, such as computation, in a distributed or clustered environment. A distributed data grid can be used to manage application objects and data shared among the servers. Distributed data grids provide low response times, high throughput, predictable scalability, continuous availability, and reliability of information. In a particular example, distributed data grids, such as the "Oracle Coherence" data grid by Oracle Corp., store information in-memory to achieve higher performance and employ redundancy in keeping copies of that information synchronized across multiple servers, ensuring resilience of the system and continued availability of data even in the event of server failure.

[0022] In one embodiment, system 10 is a computing / data processing system that includes an application or collection of distributed applications for an enterprise organization, and may also implement logistics, manufacturing, and inventory management functions. Application and computing system 10 may be configured to operate with or be implemented as a cloud-based networking system, a software as a service ("SaaS") architecture, or other type of computing solution.

[0023] The embodiments solve the problem of forecasting the demand for multiple hotel room category and service type combinations based on hotel customer attributes, room category and service type features, offered prices, and the order in which room-rate pairs are offered to customers. Rather than assuming that customer characteristics are homogeneous (i.e., expected demand should be the same when offered the same price), the embodiments assume that the customer population includes several clusters to allow customer characteristics and choice patterns to be heterogeneous across clusters. In addition to forecasting the demand of these heterogeneous customers (i.e., expected demand may differ even when offered the same price), the embodiments estimate the dynamic size of each cluster and the centroid of each cluster, which are iteratively recalculated to reflect new allocations. The main output of this problem is the probability that each individual customer will book a room in a particular room category and service type combination.

[0024] In an embodiment, a dynamic clustering approach is utilized to enable the prediction of room and service combinations by booking customers with high accuracy. The embodiment starts with an initial clustering that divides customers into clusters such that the characteristics of customers in each cluster are more likely to be homogenous than those from other clusters, and assumes a personalized selection model within each cluster. Since the cluster membership of customers (i.e., which cluster each customer belongs to) is unobservable, the embodiment employs a soft clustering approach that captures the "mix" through the probability of customers belonging to each cluster.

[0025] To do so, embodiments implement unsupervised clustering using a random forest clustering algorithm with a certain number of clusters based on the characteristics of potential hotel customers, including the order of room-service pairs, and the offered prices. Next, embodiments derive a weighted likelihood function from the observed customers based on a discrete choice multinomial logit ("MNL") model that corresponds to the clusters with weights set on the cluster probabilities obtained from the initial clustering. Then, embodiments maximize the weighted likelihood function to obtain coefficient and intercept values ​​for each covariate in the MNL model. Selection probabilities for multiple hotel room category and service type combinations for each customer are calculated from these values. The number of clusters is selected to the value that results in the highest prediction accuracy.

[0026] In an embodiment, the initial clustering is based on customer characteristics, not customer choices. To incorporate customer choice behavior into the clustering, an embodiment updates the weights as the initial clustering probability multiplied by the choice probability calculated in the previous step, which can be seen as the E step of the Expected-Maximization ("EM") algorithm. Then, an embodiment refits the model with the newly formed clusters performing a dynamic clustering step by maximizing the updated weighted likelihood function, which constitutes the M step of the EM algorithm. Finally, an embodiment repeats the E and M steps until a convergence criterion is met.

[0027] After convergence, the embodiment obtains the final estimates of the model parameters. For a new customer with the characteristics of the room category including the customer's own characteristics, the order of the room-service pairs, and the offered price, the embodiment can predict the selection probability of the new customer after estimating the association with each cluster by solving the classification problem with a supervised random forest classifier.

[0028] Dynamically iterative reconfigurable clustering algorithms / functions In general, the embodiments implement a dynamic iterative reconfigurable clustering algorithm / function for forecasting demand to generate a hotel room demand model. Assume that the target customer population is composed of multiple clusters G (G>1), and the room booking patterns among customers within each cluster are relatively homogeneous, but the booking patterns among the clusters are heterogeneous. Under this assumption, it is intuitive to consider G selection models that are different among the clusters, i.e., selection models adapted to each cluster individually. However, in practice, the cluster membership indicating which cluster each customer belongs to is unobservable. In contrast, the embodiments implement a novel algorithm / function to address the problem of estimating the heterogeneous booking patterns of customers among clusters when the cluster membership is unknown.

[0029] In particular, customer i (i = 1,...,n) has observable covariates

[0030]

number

[0031] Let J be the number of products considered in the market, and S i The customer i The set of products available to i Let ⊂{1,...,J}. y i Let y denote the product selection made by customer i, where y i ∈S i The products j=1,...,J are the observable variables.

[0032]

number

[0033] Then, the MNL room selection probability in cluster g can be expressed as follows: i Let be the cluster membership indicator for customer i,

[0034]

number

[0035] where:

[0036]

number

[0037] is for identifiability and B∈{1,...,J} is the baseline product. In the embodiment, B=J is set to demonstrate the embodiment of the present invention. The cluster membership indicator l i Since is unobservable, it is considered a latent variable and a model is required to explain the different probabilities of belonging to a cluster across different customer characteristics. Specifically, we use a mixture distribution called i Assume the model is as follows:

[0038]

number

[0039] Where:

[0040]

number

[0041] is unknown

[0042]

number

[0043] is the general notation for a probability mass function that depends on l i One common approach to modeling is to assume an MNL (also known as logit) model, which assumes that customers have probability

[0044]

number

[0045] where, for discriminability, an embodiment may:

[0046]

number

[0047] Set the vector

[0048]

number

[0049] specifies how customer characteristics affect clustering, i.e., which cluster a customer belongs to. However, product selection y i Unlike the cluster membership indicator l i Since is unobservable, the true structure of the mixture distribution is actually unknown, and it is difficult to verify whether the specified model is correct. Even if a parametric family is pre-specified for the mixture distribution as in equation (1), it may not match the true mixture distribution, which is called the model misspecification problem, leading to biased parameter estimates or poor goodness-of-fit measures, affecting the prediction accuracy.

[0050] To avoid such model misspecification and improve predictive performance, embodiments implement a semiparametric mixture of discrete choice models by assuming Equation (1) and Equation (2) instead of Equation (1) and the MNL model (3). The model parameters are

[0051]

number

[0052] This can be expressed as:

[0053]

number

[0054] The likelihood function of

[0055]

number

[0056] where:

[0057]

number

[0058] There is no pre-specified parametric model form imposed on , which can be estimated using non-parametric clustering techniques such as random forests. Other embodiments can use other unsupervised machine learning techniques for clustering. Next, the embodiment uses the same idea of ​​the EM algorithm as follows: The potential clustering membership indicators l i Assume that is known. Then the complete likelihood function is

[0059]

number

[0060] and the full log-likelihood function is

[0061]

number

[0062] It becomes. In the EM algorithm, a maximizer for the objective function in equation (4) can be found by using the following iterative method:

[0063]

number

[0064] Specifically, in the embodiment, the following E step and M step are repeated as follows. E-step Observation data

[0065]

number

[0066] Given l i Calculate the conditional expectation of M-step Parameters

[0067]

number

[0068] , the formula

[0069]

number

[0070] Update by solving As disclosed, embodiments employ unsupervised clustering techniques, such as random forests, based on customer features. Thus, the EM algorithm can be tailored to the context of what is called iterative reconfigurable clustering, as follows:

[0071] Initial Clustering Customer-level covariates

[0072]

number

[0073] Run unsupervised soft clustering (e.g., Random Forest, K-means) on G clusters based on

[0074]

number

[0075] As a result,

[0076]

number

[0077] Here,

[0078]

number

[0079] teeth

[0080]

number

[0081] is the initial estimate of

[0082]

number

[0083] Using the formula

[0084]

number

[0085] The initial parameter values ​​are obtained by solving E-step An embodiment determines the conditional cluster probabilities by using the observed choices and the fitted discrete choice model as follows: y i If =j,

[0086]

number

[0087] As a result,

[0088]

number

[0089] It is. M-step The selection model parameters are calculated by solving the following equation:

[0090]

number

[0091] and

[0092]

number

[0093] For each g, first, for j=1,...,J-1,

[0094]

number

[0095] Find the solution to:

[0096]

number

[0097] next,

[0098]

number

[0099] of,

[0100]

number

[0101] Regarding the formula

[0102]

number

[0103] where g=2,...,G. For any ε>0, the convergence criterion

[0104]

number

[0105] Repeat (E step) and (M step) until is satisfied. The above can be considered as a variation of the EM algorithm. In an embodiment, the Dempster et al. (1977) theorem can be applied to the proposed iterative algorithm, which states that

[0106]

number

[0107] but

[0108]

number

[0109] where,

[0110]

number

[0111] is our objective function

[0112]

number

[0113] It is a maximizer of. Predicting room category and rate code combinations After convergence, embodiments may provide final estimates of the model parameters

[0114]

number

[0115] get.

[0116]

number

[0117] For a new customer characterized by * Let be the available product, j∈S * About

[0118]

number

[0119] Where:

[0120]

number

[0121] is the predicted probability of belonging to cluster g by soft clustering,

[0122]

number

[0123] is the feature vector of room j available for new customers. Figure 3 is a flow diagram of the functionality of the room demand model module 16 of Figure 2 for generating a room demand model, according to one embodiment. In one embodiment, the functionality of the flow diagram of Figure 3 is implemented by software stored in memory or other computer readable or tangible medium and executed by a processor. In other embodiments, the functionality may be performed by hardware (e.g., through the use of application specific integrated circuits ("ASICs"), programmable gate arrays ("PGAs"), field programmable gate arrays ("FPGAs"), etc.), or any combination of hardware and software.

[0124] At 302, an initial unsupervised soft clustering is developed to cluster customers based on multiple attributes / characteristics assigned to each customer. In an embodiment, the attributes may include one or more of the following: (1) global distribution system used (e.g., Amadeus, SABRE, etc.), (2) booking channel, (3) length of stay, (4) number of arriving guests, (5) advance booking date, (6) weekend vs. weekday, (7) corporate bookings.

[0125] The initial clustering at 302 is based on customer features, not customer selection. Customer features are features known at the time of the room request and include data such as arrival date and time, number of people, and booking channel. In addition, customer feature data includes other inferred features such as booking window (i.e., time from booking to arrival date).

[0126] Both the initial clustering and the dynamic clustering described below, in which the initial clustering and subsequent clustering are dynamically updated, incorporate machine learning. Specifically, the initial clustering in 302 can incorporate any unsupervised machine learning technique for clustering, such as a random forest or a soft clustering algorithm using a Gaussian mixture model. Unlike customer selection, cluster membership is unobservable, so it is more difficult to assume a pre-specified parametric model for how clusters are formed based on customer characteristics, and it is difficult to verify whether the pre-specified parametric model is correct. Failure to specify the correct model will result in biased parameter estimates or low goodness-of-fit measures, affecting prediction accuracy. The embodiment does not require a pre-specified parametric model form for the clustering structure, and therefore can avoid possible bias due to model misspecification.

[0127] FIG. 4 shows an example of initial clustering according to an embodiment. As shown in FIG. 4, three clusters are formed based on guest characteristics, external factors, and travel attributes. In an embodiment, the number of clusters is a predefined parameter based on the interpretability of the clustering, typically limiting the number of clusters to a single digit. In various embodiments, 2-4 clusters are used.

[0128] At 304, embodiments estimate an initial mixture of multinomial logit ("MNL") models for demand for hotel room category and rate code combinations based on parameters related to the hotel room offerings, including: (1) the offer price, (2) the position of the room category and rate plan in the offer, and (3) room and rate features, such as view, room size, whether breakfast is included, free cancellation, etc. For each cluster formed at 302, a separate MNL model is constructed at 304. FIG. 5 is an example showing various offer prices (e.g., $335), room categories (e.g., deluxe or superior, king or queen bed), and rate codes (e.g., "rate with breakfast"). At 304, the MNL model is calculated based on the parameters related to equation (4) above by solving equation (5) above.

[0129]

number

[0130] To estimate the parameters, historical booking data stored in a database (eg, database 17 in FIG. 2) is used.

[0131] Figure 6 illustrates guest cluster choice modeling according to an example embodiment. As shown in Figure 6, each cluster uses a unique discrete choice model to predict each customer's hotel room and rate code combination choice. Figure 7 illustrates the initial assignment of MNL models to each cluster according to an embodiment.

[0132] 306, 308, and 310 collectively and iteratively form an expectation maximization ("EM") function. The EM function includes 306, 308, and 310, and also includes a soft clustering that is updated in the E-step of 306. The soft clustering in 302 is an initial clustering that is not repeated. In 306, for the expectation "E-step", the cluster probabilities are updated by incorporating customer selection probabilities evaluated at the parameter values ​​of the current iteration.

[0133] FIG. 8 illustrates a proposed likelihood function used with the EM function according to an embodiment. As shown, the proposed likelihood function includes both the cluster model generated at 302 and shown in FIG. 4 and the selection model generated at 304 and shown in FIG. 6. The proposed likelihood function is the objective function of the EM function. The embodiment finds a maximizer for this objective function to estimate the model parameters by using the EM function.

[0134] 9 shows a portion of the EM function according to an embodiment. After an "initial step", which is a soft clustering performed in 302, the expectation E step is determined in 306.

[0135] For the maximization "M-step", at 308, embodiments estimate an updated MNL model mixture, whose mixture probabilities are the cluster probabilities updated in the E-step. At 310, 306 and 308 are repeated until the convergence criterion |new prediction error-old prediction error|<0.0001 is met.

[0136] At 312, a demand model is generated that predicts the new customer's probability of selection of the room category and rate code combination using the estimated parameters from 306, 308. At 314, the function ends.

[0137] The functionality in Figure 3 combines discrete choice modeling estimation with data-driven identification of customer segments to capture the various preferences of a heterogeneous customer population and provide an interpretable model output. The demand model generated in 312 provides a practical approach that can help hoteliers profile their customers / guests based on their preferences, which serves as a valuable input to: (1) develop more efficient marketing policies and provide personalized recommendations that are more likely to be accepted, and (2) generate optimal personalized prices and display locations for each room type (e.g., suites with water views and queen beds).

[0138] Figure 10 illustrates a portion of the EM function according to an embodiment. Figure 10 illustrates the M-step at 308 and the iterations until convergence at 310.

[0139] 11-16 show an example of an embodiment of the invention for three clusters. FIG. 11 shows soft clustering at 1101 (302 in FIG. 3) and selection modeling at 1102 (304 in FIG. 3), where a different MNL model is generated for each of the clusters from the soft clustering. The number of clusters is predetermined before the EM function is used. To select the best number of clusters, prediction accuracy measures are compared across several different numbers of clusters, and the best number that achieves the most accurate prediction is selected. There is a different MNL for each cluster, but all model parameters are estimated together. Initial data 1103 for each guest is used as input, and initial cluster probabilities 1104 for each guest are generated.

[0140] FIG. 12 shows the first iteration (306 and 308 in FIG. 3) using an E-step to reassign the conditional cluster probabilities at 1201.

[0141] FIG. 13 shows the first iteration (306 and 308 in FIG. 3) using the M-step to update the selection model at 1301, which modifies the conditional cluster probabilities.

[0142] Figure 14 shows the second iteration using the E-step, which is the updated conditional cluster probabilities at 1301, and Figure 15 shows the second iteration using the M-step. For illustration purposes, only two iterations are assumed.

[0143] FIG. 16 next illustrates the generation of a demand model using the estimated model parameters to form a prediction of the selection probability of new customers.

[0144] Evaluation Indicators To investigate the performance of the iterative reconfigurable clustering according to the embodiment, the embodiment splits the data set into a training data set and a test data set. After estimating the model parameters and the initial clustering from the training data, the embodiment obtains a prediction of product selection among customers in the test data. For prediction accuracy measurements, the embodiment uses the correct classification ratio ("CCR") and the mean squared error ("MSE").

[0145] The CCR is calculated as the proportion of observations in which the option with the highest predicted probability matches the observed choice.

[0146] The MSE is calculated as follows:

[0147]

number

[0148] where y i and

[0149]

number

[0150] are the true and predicted choices of customer i for room type j, and φ te is the set of customer indices in the test data, n teis the number of customers in the test data. This metric, also known as the Brier score, is commonly used in evaluating probabilistic forecasts.

[0151] In the experiments, the embodiment was applied using a real hotel dataset by utilizing a dedicated dataset of multiple hotels in multiple cities and countries. The data includes reservation information and corresponding customer characteristics. In addition, log data (i.e., real-time customer reservation requests and corresponding responses by the reservation server system) is included. From the log, information about the display order of rooms and rate codes was extracted. This order was strategically entered by each hotelier, and each customer will have a different display order. The final dataset included 9,173 reservations for 18 different rooms and 15 different rate codes from July 2, 2019 to July 19, 2019.

[0152] The embodiment first finds the optimal number of clusters for the customer population, which is usually unknown. The embodiment adopts a predictive criteria approach to select the optimal number of clusters. Specifically, the embodiment adopts MSE to select the number of clusters with the highest predictive accuracy among 2, 3, 4, and 5 clusters. Experiments confirm that 2 clusters has the best performance among the four choices.

[0153] Given two clusters, the embodiment implements an embodiment of the present invention using a real hotel booking dataset. Specifically, the prediction accuracy of the embodiment is compared to a single cluster benchmark. The embodiment splits each dataset into a training dataset (80%) and a test dataset (20%). The results are shown in Table 1 below, and the iteration was stopped at 17 because the MSE based criterion was met (i.e., close to 0.0001).

[0154] [Table 1]

[0155] FIG. 17 shows a comparison of prediction accuracy by iterations of CCR and MSE according to an embodiment of the present invention. In FIG. 11, lines 1701 and 1702 are for two clusters, and lines 1703 and 1704 are for a single cluster. As shown, prediction performance is improved by iterations for two clusters using an embodiment of the present invention. Specifically, the value of MSE decreases by iterations while the value of CCR increases. Also, it is observed that the results of a single cluster (i.e., a known solution) are worse than the two cluster case.

[0156] FIG. 18 illustrates how cluster characteristics change with iterations given two clusters, according to an embodiment. Specifically, FIG. 18 illustrates how centroid values ​​move with iterations, with curve 1850 for CCR and curve 1860 for MSE. Seven attributes are used to cluster customers to address heterogeneous customer populations: global delivery system (1803), booking channel (1802), length of stay (1806), number of guests arriving (1807), advance booking date (1801), whether the customer arrived on a weekend (1805), and whether the customer booked via a company code (1804). FIG. 18 illustrates how each attribute moves for each cluster with iterations.

[0157] As disclosed, embodiments incorporate novel techniques for predicting customer choices and estimating the relative value of room category and service type features in the hotel industry based on booking guest attributes, the order of room-service pairs in the presentation, and the offered price. Specifically, most of the demand forecasting tools currently used by the hotel industry are aimed at providing total bookings based on time series analysis assuming a single cluster (i.e., homogeneous customer population), thus ignoring heterogeneous customer populations. These demand modeling tools are often ineffective when there are heterogeneous customers with significantly different willingness to pay and behavioral patterns. Even if some tools do consider heterogeneous customer populations, they use standard cluster algorithms and may not reflect customer choice behavior during the clustering process. Furthermore, generally, no demand forecasting tool addresses the order of room category and rate code pairs. The order of presentation on a website affects customer choice behavior in addition to the offered price.

[0158] The embodiments enable highly accurate prediction of room and service combinations by booking guests. Through computational experiments, the embodiments show that the predicted rates using the dynamic clustering method are about 4% higher than the static clustering method. Furthermore, the embodiments can input information about the order of room categories and rate codes into the display optimization system, which can help hoteliers develop more appropriate marketing strategies and suggest personalized recommendations that are more likely to be accepted.

[0159] In addition, embodiments can incorporate any unsupervised machine learning technique for clustering, such as a random forest or a soft clustering algorithm using Gaussian mixture models, into the first step of the algorithm. Unlike customer selection, cluster membership is unobservable, so it is more difficult to assume a pre-specified parametric model for how clusters are formed based on customer characteristics, and it is difficult to verify whether the pre-specified parametric model is correct. Failure to specify the correct model will result in biased parameter estimates or poor goodness-of-fit measures, affecting prediction accuracy. Since embodiments do not require a pre-specified parametric model format for the clustering structure, they can avoid possible bias due to model misspecification.

[0160] The embodiments implement dynamic clustering as a form of machine learning, especially when it involves training as in the embodiments of the present invention. The embodiments use unsupervised learning that takes in a dataset containing only the input and finds structures in the data, such as groupings or clusterings of data points. Cluster analysis is the assignment of a set of observations to subsets called clusters, such that observations in the same cluster are similar according to one or more pre-specified criteria, while observations extracted from different clusters are dissimilar. Different clustering techniques are often defined by some similarity metric and make different assumptions about the structure of the data, for example, assessed by internal compactness, i.e., similarity between members of the same cluster, and separation, i.e., differences between clusters. Dynamic clustering as a form of unsupervised online / incremental machine learning considers two concepts: (1) incrementality of the learning method to devise a clustering model, and (2) self-adaptation of the learned model (parameters and structure).

[0161] The features, structures, or characteristics of the present disclosure described throughout this specification can be combined in any suitable manner in one or more embodiments. For example, the use of "one embodiment," "some embodiments," "an embodiment," "particular embodiment," or other similar language throughout this specification refers to the fact that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the present disclosure. Thus, the appearance of the phrases "one embodiment," "some embodiments," "an embodiment," "particular embodiment," or other similar language throughout this specification does not necessarily all refer to the same group of embodiments, and the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.

[0162] Those skilled in the art will readily appreciate that the above-described embodiments may be implemented with steps in a different order and / or with elements in configurations different from those disclosed. Thus, while this disclosure contemplates embodiments outlined, it will be apparent to those skilled in the art that certain modifications, variations, and alternative constructions will be apparent while remaining within the spirit and scope of the disclosure. Accordingly, reference should be made to the appended claims to determine the metes and bounds of the present disclosure.

Claims

**Claim 1** A method for generating a demand model for potential hotel customers in a hotel room, comprising: forming a plurality of clusters based on the characteristics of the potential hotel customers, each cluster including a corresponding weight and a cluster probability; the method further comprises: generating an initial estimated mixture of multinomial logit (MNL) models corresponding to each of the plurality of clusters, the estimated mixture of the MNL models including a weighted likelihood function based on the characteristics and the weights; the method further comprises: determining modified cluster probabilities and updating the weights; estimating an updated estimated MNL model and maximizing the weighted likelihood function based on the modified cluster probabilities and the updated weights; generating the demand model adapted to predict the selection probability of a combination of a room category and a rate code of the potential hotel customers based on the updated weights and the updated estimated mixture of the MNL models. **Claim 2** The method according to claim 1, wherein the characteristics of the potential hotel customers are known when the potential hotel customers request the hotel room. **Claim 3** The method according to claim 1 or 2, wherein the step of generating the estimated mixture of the MNL models is based on a presented price, the position of the room category and the rate plan in the presentation, and the characteristics of the room and the rate. **Claim 4** The method according to claim 1 or 2, wherein the step of forming the plurality of clusters includes unsupervised machine learning. **Claim 5** The method according to claim 4, wherein the unsupervised machine learning includes one of dynamic clustering or soft clustering using a Gaussian mixture model. **Claim 6** The method according to claim 1 or 2, wherein the determining and the estimating are repeated until a convergence criterion is reached, and the demand model is generated after the convergence criterion is reached. **Claim 7** The method according to claim 1 or 2, wherein the characteristics include at least one of an arrival time, the number of people, a reservation channel, or a reservation window. **Claim 8** The method according to claim 1 or 2, wherein the demand model is adapted to maximize the revenue of the hotel room. **Claim 9** A computer-readable program storing instructions that, when executed by one or more processors, cause the one or more processors to generate a demand model for potential hotel customers in a hotel room, wherein generating the demand model comprises forming a plurality of clusters based on the characteristics of the potential hotel customers, each cluster including a corresponding weight and a cluster probability, generating an initial estimated mixture of multinomial logit (MNL) models corresponding to each of the plurality of clusters, the estimated mixture of the MNL models including a weighted likelihood function based on the characteristics and the weights, determining revised cluster probabilities and updating the weights, estimating an updated estimated mixture of the MNL model and maximizing the weighted likelihood function based on the revised cluster probabilities and the updated weights, generating the demand model adapted to predict the selection probabilities of combinations of room categories and rate codes of the potential hotel customers based on the updated weights and the updated estimated mixture of the MNL model.

10. The computer-readable program according to claim 9, wherein the characteristics of the potential hotel customers are known when the potential hotel customers request the hotel room.

11. The computer-readable program according to claim 9 or 10, wherein generating the estimated mixture of the MNL model is based on the presented price, the position of the room category and rate plan in the presentation, and the characteristics of the room and rate.

12. The computer-readable program according to claim 9 or 10, wherein forming the plurality of clusters includes unsupervised machine learning.

13. The computer-readable program according to claim 12, wherein the unsupervised machine learning includes one of dynamic clustering or soft clustering using a Gaussian mixture model.

14. The computer-readable program according to claim 9 or 10, wherein the determining and the estimating are repeated until a convergence criterion is reached, and the demand model is generated after the convergence criterion is reached.

15. The computer-readable program according to claim 9 or 10, wherein the characteristics include at least one of arrival date and time, number of people, reservation channel, or reservation window.

16. The computer-readable program according to claim 9 or 10, wherein the demand model is adapted to maximize the revenue of the hotel room.

17. A hotel reservation system for generating a demand model of potential hotel customers for hotel rooms, comprising one or more processors coupled to stored instructions, and a database for storing past reservation data, wherein the one or more processors are configured to form a plurality of clusters based on the characteristics of the potential hotel customers, each cluster including a corresponding weight and cluster probability, are configured to generate an initial estimated mixture of multinomial logit (MNL) models corresponding to each of the plurality of clusters, the estimated mixture of the MNL model including a weighted likelihood function based on the characteristics and the weights, determine modified cluster probabilities and update the weights, estimate an updated estimated mixture of the MNL model and maximize the weighted likelihood function based on the modified cluster probabilities and the updated weights, and generate the demand model adapted to predict the selection probability of a combination of a room category and a rate code of the potential hotel customer based on the updated weights and the updated estimated mixture of the MNL model.

18. The hotel reservation system according to claim 17, wherein the characteristics of the potential hotel customer are known when the potential hotel customer requests the hotel room.

19. Generating the estimated MNL model is based on the presented price, the position of the room category and the rate plan in the presentation, and the characteristics of the room and the rate, according to the hotel reservation system of claim 17 or 18.

20. Forming the plurality of clusters includes unsupervised machine learning, according to the hotel reservation system of claim 17 or 18.