Intelligent air travel payment discount recommendation system
By introducing dynamic user status space and reinforcement learning strategies, combined with real-time update mechanism, the problem of insufficient personalization and real-time performance of the existing intelligent air travel payment discount recommendation system is solved, and personalized and flexible discount recommendations are achieved to adapt to user needs and market changes.
Patent Information
- Application Number
- CN202510579733.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-15
AI Technical Summary
The existing intelligent air travel payment discount recommendation system has shortcomings in personalized recommendations, real-time updates and adapting to changes in user needs, especially for new users, and cannot respond in a timely manner to real-time factors such as airline promotions and flight delays.
Introduce dynamic user status space, reinforcement learning strategies and real-time update mechanisms, and by building user status space, policy execution module, update module and dynamic optimization module, personalized and flexible recommendation functions are realized, predict users' future payment needs and provide timely and accurate discount recommendations.
It realizes accurate and real-time discount recommendations in a dynamically changing environment, adapts to user needs and market changes, improves the personalization and flexibility of the recommendation system, and meets the diverse needs of users.
Smart Images

Figure CN120494827A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent air travel, and more specifically, to an intelligent air travel payment discount recommendation system. Background Art
[0002] Existing intelligent travel payment discount recommendation systems typically rely on traditional content-based or collaborative filtering algorithms to recommend travel-related payment discounts to users. These systems generate recommendations by analyzing users' historical behavior data, preferences, and the behavior of similar users. Most systems use static rules and fixed recommendation models to provide discount information, for example, based on a user's past purchase history and flight browsing history. While these recommendation systems can enhance the user shopping experience to a certain extent, they have limitations in terms of personalized recommendations, real-time updates, and adaptability to changing user needs.
[0003] However, existing intelligent travel payment discount recommendation systems also have several problems. First, traditional recommendation systems are limited in their effectiveness in personalized recommendations, especially for new users. Due to the lack of sufficient user historical data, recommendation results are poor, and the recommended content may deviate from the user's actual needs. Second, existing recommendation systems are often unable to adapt to changes in user needs in real time. For example, real-time factors such as airline promotions and flight delays may affect user decisions, and existing systems are often unable to adjust recommended content immediately. In addition, the system lacks sufficient flexibility when dealing with behavioral differences between different users, which may result in overly simple recommendations and difficulty meeting the diverse needs of users. Therefore, how to provide accurate and personalized recommendations in a dynamically changing environment has become an urgent problem that needs to be solved. Summary of the Invention
[0004] The present invention addresses the technical problems existing in the prior art and provides an intelligent travel payment discount recommendation system. By introducing a dynamic user state space, a reinforcement learning strategy, and a real-time update mechanism, the system achieves personalized and flexible recommendation functions. The system can predict users' future payment needs based on their historical payment behavior, real-time flight dynamics, and the correlation between cross-platform itineraries, and provide timely and accurate discount recommendations when user needs change, such as flight delays or itinerary changes. In addition, through a real-time feedback mechanism and a dynamic optimization module, the system can continuously adjust and optimize the recommendation strategy to adapt to changes in user needs and the market environment, thereby solving the problems raised in the above-mentioned background technology.
[0005] The technical solution of the present invention to solve the above technical problems is as follows: specifically comprising: a state generation module, a strategy execution module, an update module and a dynamic optimization module;
[0006] State generation module: Constructs a user state space S, which includes: historical user payment behavior data, real-time flight dynamics information, available payment channels, and associated discount rules. In addition, the user state space also includes dynamic travel plan features: the user's subsequent flight booking status, cross-platform itinerary correlation (such as the matching degree of hotel booking and flight time), which is used to predict the user's long-term payment needs. This design is designed to address the user's urgent payment needs due to itinerary changes (such as hotel cancellations due to flight delays) and the need to reserve discount resources in advance. Based on this state space, a recommended action space A is generated, which includes: payment channel discount combinations, airline co-branded card discounts, third-party discount coupons, and points deduction strategies;
[0007] Policy execution module: According to the current user status S t , through reinforcement learning strategy π(a|S t ) Select the optimal payment preferential action A t , and push it to the user terminal in real time;
[0008] Update module: collect user feedback data on recommended offers and calculate instant rewards R t , the rewards are based on the user's actual payment conversion rate, preferential utilization rate and user satisfaction index;
[0009] Dynamic Optimization Module: This module updates the policy parameters of the reinforcement learning model and integrates multi-source data with dynamic flight information and discount rules to optimize the recommended action at the next moment.
[0010] In a preferred embodiment, in the status generation module, the user's historical payment behavior data specifically includes: user preferred payment channels, historical discount usage records, and flight class selection preferences.
[0011] In a preferred embodiment, the real-time flight dynamic information specifically includes: flight delay status, number of remaining seats, and current cabin price fluctuations.
[0012] In a preferred embodiment, the available payment channels and associated preferential rules are specifically: obtaining dynamic preferential strategies of banks, third-party payment platforms and airlines in real time through an API interface.
[0013] In a preferred embodiment, in the policy execution module, the reinforcement learning strategy is a deep deterministic policy gradient algorithm, and its action selection process satisfies:
[0014] A t =μ(S t |θ μ )+Y t ;
[0015] Among them, μ(·) represents the policy network, θμ represents the network parameters, Y t Represents exploration noise, which is used to balance the exploration and exploitation of the strategy.
[0016] In a preferred embodiment, in the update module, the instant reward R t The calculation formula is:
[0017] R t =α·s+β·f+γ·h;
[0018] Among them, α, β, γ represent preset weight coefficients, and α+β+γ=1, s represents the payment success mark, f represents the preferential usage amount, and h represents the user rating.
[0019] In a preferred embodiment, in the dynamic optimization module, the dynamic optimization process of the policy update is:
[0020] S1. When the flight delay time exceeds the threshold, the emergency discount recommendation strategy is triggered, and airline compensation coupons or free rebooking benefits are given priority;
[0021] S2. When the user's payment channel balance is insufficient, it will automatically switch to the backup payment method and match the corresponding discount;
[0022] S3. If it is detected that the user has not completed check-in N hours before the flight departure, the "Fast Track Package" recommendation (including priority check-in coupon + delay insurance combination) will be automatically triggered;
[0023] S4. When the user's GPS location shows that they are in the airport security area, exclusive discounts for the airport business district (such as instant discount coupons for duty-free shops) are recommended, and scene perception is enhanced through LBS data.
[0024] In a preferred embodiment, the specific steps of multi-source data fusion are:
[0025] S1. Align the features of the real-time data of the flight information system, the operation logs of the user payment terminal, and the external discount rule library;
[0026] S2. Use the attention mechanism to weightedly fuse multi-source data and generate a comprehensive state vector S t .
[0027] In a preferred embodiment, the dynamic optimization module includes an anti-fraud verification unit for verifying the availability of the preferential rules and the legitimacy of the user identity before the recommendation is executed.
[0028] In a preferred embodiment, the anti-fraud verification unit specifically includes:
[0029] S1. Build a discount abuse detection model based on user historical behavior to intercept abnormally high-frequency discount requests;
[0030] S2. Use blockchain technology to record preferential issuance and usage records to prevent duplicate cancellations.
[0031] In a preferred embodiment, the method specifically includes the following steps:
[0032] The beneficial effects of the present invention are as follows: the intelligent travel payment discount recommendation system realizes personalized and flexible recommendation functions by introducing a dynamic user state space, reinforcement learning strategy and real-time update mechanism; it can predict the user's future payment needs based on the user's historical payment behavior, real-time flight dynamics, and the correlation of cross-platform itineraries, and provide timely and accurate discount recommendations when user needs change, such as flight delays and itinerary changes; in addition, through the real-time feedback mechanism and dynamic optimization module, the system can continuously adjust and optimize the recommendation strategy to adapt to changes in user needs and market environment. Therefore, the present application can effectively solve the challenges faced by traditional recommendation systems such as inaccurate personalized recommendations, insufficient real-time updates and differentiated user behaviors, and provide more accurate and real-time discount recommendations. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 Flow chart of the method of the present invention;
[0034] Figure 2 This is a system structure diagram of the present invention. DETAILED DESCRIPTION
[0035] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0036] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the described features. In the description of this application, "plurality" means two or more, unless otherwise specifically specified.
[0037] In the description of this application, the term "for example" is used to mean "used as an example, illustration or explanation". Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art will recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes will not be elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in this application.
[0038] Example 1
[0039] This embodiment provides Figure 1-2 The intelligent travel payment discount recommendation system shown in the figure specifically includes: a state generation module, a strategy execution module, an update module and a dynamic optimization module;
[0040] State generation module: Constructs a user state space S, which includes historical user payment behavior data, real-time flight dynamics, available payment channels, and associated discount rules. Furthermore, the user state space also includes dynamic travel plan features: the user's subsequent flight booking status and cross-platform itinerary correlations (such as the matching degree between hotel bookings and flight times). This is used to predict the user's future payment needs. This design is designed to address the possibility that users may have urgent payment needs due to itinerary changes (such as hotel cancellations due to flight delays) and need to reserve discount resources in advance. Based on the state space, a recommended action space A is generated. The action space includes payment channel discount combinations, airline co-branded card discounts, third-party discount coupons, and points deduction strategies.
[0041] Policy execution module: According to the current user status S t , through reinforcement learning strategy π(a|S t ) Select the optimal payment preferential action A t , and push it to the user terminal in real time;
[0042] Update module: collect user feedback data on recommended offers and calculate instant rewards R t ,rewards are based on users’ actual payment conversion rate, preferential utilization rate and user satisfaction indicators;
[0043] Dynamic Optimization Module: Updates the strategy parameters of the reinforcement learning model and integrates multi-source data with dynamic flight information and discount rules to optimize the recommended action at the next moment.
[0044] In this embodiment, the state generation module specifically needs to be explained. The user's historical payment behavior data specifically includes: the user's preferred payment channel, historical discount usage records, and flight class selection preferences;
[0045] Real-time flight dynamic information includes: flight delay status, remaining seats, and current cabin price fluctuations;
[0046] The available payment channels and associated discount rules are as follows: real-time access to dynamic discount strategies of banks, third-party payment platforms and airlines through the API interface.
[0047] In this embodiment, the specific requirements for the strategy execution module are as follows: the reinforcement learning strategy is a deep deterministic policy gradient algorithm, and its action selection process satisfies:
[0048] A t =μ(S t |θ μ )+Y t ;
[0049] Among them, μ(·) represents the policy network, θ μ represents the network parameters, Y t Represents exploration noise, which is used to balance the exploration and exploitation of the strategy.
[0050] In this embodiment, it is necessary to explain the update module, the instant reward R t The calculation formula is:
[0051] R t =α·s+β·f+γ·h;
[0052] Among them, α, β, γ represent preset weight coefficients, and α+β+γ=1, s represents a payment success flag, f represents the preferential usage amount, and h represents the user rating.
[0053] In this embodiment, the dynamic optimization module specifically needs to be explained. The dynamic optimization process of the strategy update is as follows:
[0054] S1. When the flight delay time exceeds the threshold, the emergency discount recommendation strategy is triggered, and airline compensation coupons or free rebooking benefits are given priority;
[0055] S2. When the user's payment channel balance is insufficient, it will automatically switch to the backup payment method and match the corresponding discount;
[0056] S3. If it is detected that the user has not completed check-in N hours before the flight departure, the "Fast Track Package" recommendation (including priority check-in coupon + delay insurance combination) will be automatically triggered;
[0057] S4. When the user's GPS location indicates they are in the airport security area, we recommend exclusive discounts in the airport shopping district (such as instant discount coupons at duty-free shops), enhancing scene perception through LBS data.
[0058] The specific steps of multi-source data fusion are:
[0059] S1. Align the features of the real-time data of the flight information system, the operation logs of the user payment terminal, and the external discount rule library;
[0060] S2. Use the attention mechanism to weightedly fuse multi-source data and generate a comprehensive state vector S t ;
[0061] The dynamic optimization module includes an anti-fraud verification unit, which is used to verify the availability of preferential rules and the legitimacy of user identity before recommendation execution. The anti-fraud verification unit specifically includes:
[0062] S1. Build a discount abuse detection model based on user historical behavior to intercept abnormally high-frequency discount requests;
[0063] S2. Use blockchain technology to record preferential issuance and usage records to prevent duplicate cancellations.
[0064] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0065] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0066] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0067] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0068] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0069] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0070] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. An intelligent travel payment discount recommendation system, characterized in that: Specifically include: State generation module, strategy execution module, update module and dynamic optimization module; State generation module: Constructs a user state space S, which includes historical user payment behavior data, real-time flight dynamics, available payment channels, and associated discount rules. Based on this state space, it generates a recommended action space A, which includes payment channel discount combinations, airline co-branded card discounts, third-party instant discount coupons, and points deduction strategies. Policy execution module: According to the current user status S t , through reinforcement learning strategy π(a|S t ) Select the optimal payment preferential action A t , and push it to the user terminal in real time; Update module: collect user feedback data on recommended offers and calculate instant rewards R t , the rewards are based on the user's actual payment conversion rate, preferential utilization rate and user satisfaction index; Dynamic Optimization Module: Updates the strategy parameters of the reinforcement learning model and integrates multi-source data with dynamic flight information and discount rules to optimize the recommended action at the next moment.
2. The intelligent travel payment discount recommendation system according to claim 1, characterized in that: In the state generation module, the user's historical payment behavior data specifically includes: user preferred payment channels, historical discount usage records, and flight class selection preferences.
3. The intelligent travel payment discount recommendation system according to claim 2, characterized in that: The real-time flight dynamic information specifically includes: flight delay status, remaining seats, and current cabin price fluctuations.
4. The intelligent travel payment discount recommendation system according to claim 3, characterized in that: The available payment channels and associated preferential rules are specifically: obtaining dynamic preferential strategies of banks, third-party payment platforms and airlines in real time through API interfaces.
5. The intelligent travel payment discount recommendation system according to claim 4, characterized in that: In the policy execution module, the reinforcement learning strategy is a deep deterministic policy gradient algorithm, and its action selection process satisfies: A t =μ(S t |θ μ )+Y t ; Among them, μ(·) represents the policy network, θ μ represents the network parameters, Y t represents the exploration noise.
6. The intelligent travel payment discount recommendation system according to claim 5, characterized in that: In the update module, the immediate reward R t The calculation formula is: R t =α·s+β·f+γ·h; Among them, α, β, γ represent preset weight coefficients, and α+β+γ=1, s represents a payment success flag, f represents the preferential usage amount, and h represents the user rating.
7. The intelligent air travel payment discount recommendation system according to claim 6, characterized in that: In the dynamic optimization module, the dynamic optimization process of strategy update is: S1. When the flight delay time exceeds the threshold, the emergency discount recommendation strategy is triggered, and airline compensation coupons or free rebooking benefits are given priority; S2. When the balance of the user's payment channel is insufficient, it will automatically switch to the backup payment method and match the corresponding discount.
8. The intelligent air travel payment discount recommendation system according to claim 7, characterized in that: The specific steps of multi-source data fusion are: S1. Align the features of the real-time data of the flight information system, the operation logs of the user payment terminal, and the external discount rule library; S2. Use the attention mechanism to weightedly fuse multi-source data and generate a comprehensive state vector S t .
9. The intelligent air travel payment discount recommendation system according to claim 8, characterized in that: The dynamic optimization module includes an anti-fraud verification unit for verifying the availability of preferential rules and the legitimacy of user identity before recommendation execution.
10. The intelligent air travel payment discount recommendation system according to claim 9, characterized in that: The anti-fraud verification unit specifically includes: S1. Build a discount abuse detection model based on user historical behavior to intercept abnormally high-frequency discount requests; S2. Use blockchain technology to record preferential issuance and usage records to prevent duplicate cancellations.
Citation Information
Cited By
Public payment method and system for multi-mode identity modeling and intelligent right and interest linkage
CN121146773A