Financial product recommendation method, device and equipment, medium and program product
By combining deep learning and reinforcement learning in financial product recommendation methods, and utilizing long and short distance data for feature interaction and strategy optimization, this approach solves the problem that existing recommendation systems cannot fully understand the demand side, achieving higher recommendation accuracy and relevance.
Patent Information
- Application Number
- CN202511130120.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-18
AI Technical Summary
Financial product recommendation systems often fail to fully understand the actual needs of consumers, resulting in a low degree of alignment between recommendations and expectations.
By acquiring long-range and short-range data from the demand side of financial products, deep learning network models are used to interact explicit and implicit features, and reinforcement learning network models are combined to update the recommendation list and optimize the recommendation strategy.
It improves the accuracy and diversity of financial product recommendations, enabling a better understanding of the actual needs of the demanders and enhancing the alignment between recommendations and expectations.
Smart Images

Figure CN120975883A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of financial technology, and in particular to a method, apparatus, device, medium and program product for recommending financial products. Background Technology
[0002] With the development of the financial technology field, the types and structures of financial products are constantly being enriched. It is difficult for financial product demanders to quickly identify financial products that meet their needs from a large number of financial products. Therefore, personalized recommendations for financial products are necessary.
[0003] Currently, financial product recommendations can be made by sales personnel combining their experience with thorough communication with clients seeking financial products, or by using online channels to make recommendations based on a single type of authorization information from the client. However, these methods fail to provide a comprehensive understanding of the client's actual needs, resulting in a low degree of alignment between the recommended and desired financial products. Summary of the Invention
[0004] This invention provides a method, apparatus, device, medium, and program product for recommending financial products, which can improve the matching degree between recommended financial products and desired financial products.
[0005] In a first aspect, embodiments of the present invention provide a method for recommending financial products, including:
[0006] Acquire long-distance and short-distance data related to demanders of financial products. The long-distance data includes at least long-distance basic financial data, interest preference data, and financial product data. The short-distance data includes at least short-distance basic financial data and operational feedback data.
[0007] By using a deep learning network model, explicit and implicit feature interactions are performed based on the long-distance data to determine a candidate product recommendation list.
[0008] The candidate product recommendation list is updated based on the short-range data using a reinforcement learning network model to determine the target product recommendation list.
[0009] Secondly, embodiments of the present invention provide a financial product recommendation device, comprising:
[0010] The acquisition module is used to acquire long-distance and short-distance data related to the demand side of financial products. The long-distance data includes at least long-distance basic financial data, interest preference data, and financial product data, and the short-distance data includes at least short-distance basic financial data and operation feedback data.
[0011] The deep learning module is used to determine a candidate product recommendation list by performing explicit and implicit feature interactions based on the long-distance data through a deep learning network model.
[0012] The reinforcement learning module is used to update the candidate product recommendation list based on the short-range data using a reinforcement learning network model, and to determine the target product recommendation list.
[0013] Thirdly, embodiments of the present invention provide an electronic device, including:
[0014] At least one processor; and
[0015] A memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the method as described in the first aspect.
[0017] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer instructions that cause a processor to execute the method described in the first aspect.
[0018] Fifthly, embodiments of the present invention provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the method described in the first aspect.
[0019] The technical solution of this invention acquires long-distance and short-distance data related to the demand side of financial products. The long-distance data includes at least long-distance basic financial data, interest preference data, and financial product data, while the short-distance data includes at least short-distance basic financial data and operational feedback data. A deep learning network model is used to determine a candidate product recommendation list based on explicit and implicit feature interactions with the long-distance data. A reinforcement learning network model is then used to update the candidate product recommendation list based on the short-distance data to determine a target product recommendation list. This solution expands the types and time dimensions of data used for financial product recommendations, providing more comprehensive data support. It combines deep learning and reinforcement learning, possessing the dual advantages of deep learning's ability to handle complex multi-source information and reinforcement learning's emphasis on dynamic changes. This allows for a better understanding of the actual needs of the demand side of financial products, improving the fit between recommended and desired financial products.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of a financial product recommendation method provided in Embodiment 1 of the present invention;
[0023] Figure 2 This is a schematic diagram illustrating the working principle of a financial product recommendation method according to Embodiment 1 of the present invention;
[0024] Figure 3 This is a flowchart of a financial product recommendation method provided in Embodiment 2 of the present invention;
[0025] Figure 4 This is a schematic diagram of the structure of a deep learning network model provided in Embodiment 2 of the present invention;
[0026] Figure 5 This is a schematic diagram illustrating the working principle of a reinforcement learning network model according to Embodiment 2 of the present invention;
[0027] Figure 6 This is a schematic diagram of the structure of a financial product recommendation device according to Embodiment 3 of the present invention;
[0028] Figure 7 This is a schematic diagram of the structure of an electronic device that implements an embodiment of the present invention. Detailed Implementation
[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0030] It should be noted that the terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0031] It should be noted that the acquisition, storage, use, and processing of data involved in this invention are all authorized by the user and comply with relevant laws and regulations. The data may only be used for recommending financial products.
[0032] Example 1
[0033] Figure 1 This is a flowchart of a financial product recommendation method according to Embodiment 1 of the present invention. This embodiment is applicable to situations involving financial product recommendations. The method can be executed by a financial product recommendation device, which can be implemented in software and / or hardware and integrated into an electronic device. Furthermore, the electronic device includes, but is not limited to, computers, laptops, smartphones, etc.
[0034] like Figure 1 As shown, the method includes:
[0035] S110. Obtain long-distance and short-distance data related to the demand side of financial products. The long-distance data includes at least long-distance basic financial data, interest preference data, and financial product data. The short-distance data includes at least short-distance basic financial data and operation feedback data.
[0036] The demand side for financial products can be understood as users who need financial product recommendations, but this is not limited here.
[0037] Long-range data can be data recorded over a long period of time related to the financial operations of those demanding financial products. Short-range data can be data recorded over a short period of time related to the financial operations of those demanding financial products. The definitions of long and short time spans are not limited; for example, a long time span can be several years (e.g., data accumulated several years prior to a certain point in time), and a short time span can be several months (e.g., data accumulated several months prior to a certain point in time). The specific definitions can be determined based on the actual application requirements.
[0038] In this embodiment of the invention, data related to financial operations of financial product demanders can be divided into four categories: basic financial data, interest and preference data, financial product data, and operational feedback data. Both long-distance and short-distance basic financial data are essentially basic financial data, differing only in their time span. All of the aforementioned data has been authorized for access and use by the financial product demanders, and this data can originate from databases built within institutions based on user dimensions. The following provides illustrative examples of these data categories, including but not limited to the data content described below, which can be fine-tuned according to actual circumstances and application scenarios.
[0039] Basic financial data can be fundamental data related to the financial operations of those seeking financial products, such as customer level, customer risk assessment results, records of large asset changes, and records of acquiring financial products. Some information in the records of large asset changes is unstructured data, which can be represented as a dictionary based on common types and then converted into structured data. For example, transaction types can be converted by income / expense categories: 1-Income, 2-Expense. Records of acquiring financial products can be further refined into multiple indicators such as the product name, product period, product type, product risk level, number of acquisitions, holding time, and return information. Product type and product risk level can also be represented as a dictionary based on common types. The basic financial data is recorded in the form of a set, which includes subsets corresponding to customer level, customer risk assessment results, records of large asset changes, and records of acquiring financial products.
[0040] Interest preference data can be data related to the financial operations of those seeking financial products, reflecting their interests and preferences. This data may include: the industry of the seeker, the industry corresponding to large asset change records, the temporal relationship between large asset change records and financial product acquisition operations, the dwell time of the seeker on the financial product recommendation platform, the relevant fields and keywords on the browsing interface, interaction on the browsing interface (e.g., whether to add "follow"), and the investment areas of the acquired financial products. The temporal relationship between large asset change records and financial product acquisition operations can be quantified and structured into data based on time intervals. For example, if the time interval between a financial product acquisition operation and a large asset change is less than 5 days, it is considered a strong correlation, corresponding to dictionary key 5. Similarly, 4 represents a moderately strong correlation, 3 a moderately strong correlation, 2 a moderately weak correlation, and 1 a weak correlation. The interest preference data is recorded as a set, which includes multiple subsets of the data within the interest preference data set.
[0041] Financial product data can be data related to financial products that support recommendations to those seeking such products. This data may include, for example, the product name, product period, product type, product risk level, return information, and other relevant indicators. Financial product data can be recorded as a set, with each financial product-related data point within the set being recorded as a subset.
[0042] Operational feedback data can be the operational feedback from financial product demanders after the financial product is recommended to them. This feedback may include: browsing time, browsing frequency, adding "follow" status, understanding of product details, rejection of recommendation, and product acquisition status.
[0043] S120. Using a deep learning network model, based on the long-distance data, explicit feature interaction and implicit feature interaction are performed to determine a candidate product recommendation list.
[0044] A deep learning network model can be understood as a model capable of implementing deep learning. Deep learning is a branch of machine learning, an algorithm that uses artificial neural networks as its architecture to learn representations of data. Machine learning is a sub-branch of artificial intelligence, an algorithm that automatically analyzes data to obtain patterns and uses these patterns to predict unknown data.
[0045] Deep learning network models can employ an architecture combining explicit and implicit feature interaction components. The explicit feature interaction component inputs long-range data for high-order explicit feature interactions at the vector level, yielding explicit feature interaction results. This part of the architecture enhances the model's memorization ability, enabling it to learn direct, high-frequency co-occurring feature associations within the long-range data. The implicit feature interaction component inputs long-range data for high-order implicit feature interactions, yielding implicit feature interaction results. This part of the architecture enhances the model's generalization ability, uncovering potential, indirect patterns by learning implicit associations among features in the long-range data. The results of both explicit and implicit feature interactions are combined to determine a candidate product recommendation list. This candidate product recommendation list is a list of potential products for recommendation and requires further updates and optimization.
[0046] S130. The candidate product recommendation list is updated based on the short-range data using a reinforcement learning network model to determine the target product recommendation list.
[0047] A reinforcement learning network model can be understood as a model capable of implementing reinforcement learning. Reinforcement learning is a field of machine learning that emphasizes how to act based on the environment, seeking a balance between prior information and unknown posterior information to achieve the maximum expected reward.
[0048] Reinforcement learning schemes in reinforcement learning network models can be based on multilayer perceptrons, using deep neural networks to approximate the state-action reward value, and making decisions based on this value. The aim is to achieve low computational cost and high online training efficiency. Essentially, it involves learning the optimal decision-making strategy through interaction with the environment; that is, choosing the action that yields the best long-term cumulative reward when facing a given state. The state-action reward value is the reward value obtainable by a specific action in a given state.
[0049] In this embodiment of the invention, short-range data encoding can be used as the aforementioned state, and a product recommendation list can be used as the aforementioned action (the initial action can be the candidate product recommendation list obtained in S120). A reinforcement learning network model is used to continuously learn and make decisions. That is, through the reinforcement learning network model, the state is continuously updated. During the state update process, the possible actions in each state and the execution reward after the action is executed are determined (which can be the operation feedback data obtained after recommending the product recommendation list corresponding to the action to the financial product demander). The corresponding state action reward value is further determined, and the optimal decision strategy is determined from the determined state action reward values. That is, it is determined which action in which state can obtain the optimal state action reward value. Finally, the selected action is the target product recommendation list to be recommended to the financial product demander.
[0050] Once a target product recommendation list is determined, it can be displayed on the browsing interface of financial product demanders on the financial product recommendation platform to achieve financial product recommendations to financial product demanders.
[0051] The following content summarizes and explains S110-S130 above. Figure 2 This is a schematic diagram illustrating the working principle of a financial product recommendation method according to Embodiment 1 of the present invention. Figure 2 As shown, long-range data (long-range basic financial data, interest preference data, and financial product data) are input into a deep learning network model, which outputs a candidate product recommendation list. Short-range data (short-range basic financial data and operational feedback data) are input into a reinforcement learning network model, which updates the candidate product recommendation list output by the deep learning network model to obtain the target product recommendation list.
[0052] The technical solution of this invention acquires long-distance and short-distance data related to the demand side of financial products. The long-distance data includes at least long-distance basic financial data, interest preference data, and financial product data, while the short-distance data includes at least short-distance basic financial data and operational feedback data. A deep learning network model is used to determine a candidate product recommendation list based on explicit and implicit feature interactions with the long-distance data. A reinforcement learning network model is then used to update the candidate product recommendation list based on the short-distance data to determine a target product recommendation list. This solution expands the types and time dimensions of data used for financial product recommendations, providing more comprehensive data support. It combines deep learning and reinforcement learning, possessing the dual advantages of deep learning's ability to handle complex multi-source information and reinforcement learning's emphasis on dynamic changes. This allows for a better understanding of the actual needs of the demand side of financial products, improving the fit between recommended and desired financial products.
[0053] Example 2
[0054] Figure 3 This is a flowchart of a financial product recommendation method according to Embodiment 2 of the present invention. This embodiment is based on Embodiment 1 above, further refining the determination of the candidate product recommendation list by using a deep learning network model to perform explicit feature interaction and implicit feature interaction based on the long-distance data; and further refining the determination of the target product recommendation list by using a reinforcement learning network model to update the candidate product recommendation list based on the short-distance data.
[0055] like Figure 3 As shown, the method includes:
[0056] S110. Obtain long-distance and short-distance data related to the demand side of financial products. The long-distance data includes at least long-distance basic financial data, interest preference data, and financial product data. The short-distance data includes at least short-distance basic financial data and operation feedback data.
[0057] S121. The long-distance data is input into the embedding layer of the deep learning network model for preprocessing to obtain preprocessed long-distance data.
[0058] S122. Input the preprocessed long-distance data into the compressed interaction network of the deep learning network model to perform explicit feature interaction and obtain explicit feature interaction results.
[0059] S123. Input the preprocessed long-distance data into the attention mechanism network of the deep learning network model to perform implicit feature interaction and obtain the implicit feature interaction result.
[0060] S124. The candidate product recommendation list is determined based on the explicit feature interaction results and the implicit feature interaction results through the output unit of the deep learning network model.
[0061] Figure 4 This is a schematic diagram of the structure of a deep learning network model according to Embodiment 2 of the present invention, combined with... Figure 4 The structure shown provides the following explanation for S121-S124:
[0062] Long-range data is preprocessed by inputting it into the embedding layer of a deep learning network model. This preprocessing maps the long-range data into a low-dimensional dense representation, resulting in preprocessed long-range data that facilitates subsequent capture of feature relationships. The long-range data input into the embedding layer can be understood as multi-dimensional discrete data.
[0063] Preprocessed long-range data is input into a compressed interaction network of a deep learning network model. High-order explicit feature interactions are performed at the vector level within this compressed interaction network to obtain the explicit feature interaction results. During the high-order explicit feature interaction process, each layer of the neural network calculates based on the previous layer and the original feature vectors. The order of the ultimately learned feature interaction is determined by the number of layers in the network. Each hidden layer is connected to the output unit through a pooling operation, ensuring that the output unit can see feature interaction patterns of different orders. For each feature vector, it sequentially interacts with multiple other feature vectors, and finally, it is compressed into a vector using a weight matrix.
[0064] Preprocessed long-range data is input into the attention mechanism network of a deep learning network model, where high-order latent feature interactions are performed to obtain the latent feature interaction results. The attention mechanism network consists of an encoder and decoder composed of modules based on multiple self-attention mechanisms, which can effectively capture the implicit and complex feature relationships between long-range basic financial data, interest preference data, and financial product data, and is highly sensitive to long-range dependencies.
[0065] The execution order of explicit and implicit feature interactions is not limited, as long as the results of both explicit and implicit feature interactions are obtained. Once the results of both interactions are obtained, the output unit can determine the candidate product recommendation list based on both.
[0066] The advantage of this setup is that, through deep learning network models, based on long-distance data of corresponding multiple data types, a high-order explicit and implicit dual feature interaction can be constructed. This can uncover the explicit and implicit information included in the long-distance data, thereby improving the accuracy and diversity of financial product recommendations.
[0067] In one embodiment, determining a candidate product recommendation list based on the explicit feature interaction results and the implicit feature interaction results includes:
[0068] The results of the interaction between the explicit features and the interaction between the implicit features are weighted, and the weighted result is added to the bias term to obtain the summation result.
[0069] The summation result is processed by the execution function of the binary classification task to obtain the score results of each financial product included in the financial product data;
[0070] Based on the scoring results of each financial product, a list of recommended candidate products is determined.
[0071] The output y of a deep learning network model can be expressed as: y = σ(w1x + w2p + b). Here, x and p represent the results of the explicit and implicit feature interactions, respectively; w1 and w2 represent the weights of the explicit and implicit feature interactions, respectively; b is a bias term used to adjust the overall baseline probability; and σ represents the execution function for the binary classification task.
[0072] The execution function for a binary classification task can be understood as a function that performs the binary classification task, mapping the linear output to the interval [0,1]. The value in this interval represents the probability that a sample (i.e., a financial product) belongs to the positive class (i.e., y=1). The closer the value is to 1, the more recommended the sample should be; the closer the value is to 0, the greater the probability that the sample belongs to the negative class.
[0073] It should be noted that the optimal values of the weights w1 and w2 mentioned above can be obtained by continuously learning the difference between the predicted and actual values of the binary classification task through the cross-entropy loss function and minimizing the log-likelihood.
[0074] Based on the output y above, the score results of each financial product included in the financial product data can be obtained. The score result is the probability that each financial product belongs to the positive class, which is the value in the interval [0,1] above. Based on the score results of each financial product, a candidate product recommendation list can be formed by selecting some financial products from each financial product. The selection criteria can be set according to the actual application needs.
[0075] The advantage of this setup is that it can combine the results of explicit feature interactions with those of implicit feature interactions to determine the probability that each financial product included in the financial product data should be recommended, and then select some financial products to form a candidate product recommendation list.
[0076] In one embodiment, determining a candidate product recommendation list based on the rating results of each financial product includes: sorting each financial product from largest to smallest according to the rating results, and selecting the top predetermined number of financial products to form a candidate product recommendation list.
[0077] In the candidate product recommendation list, a set number of financial products are arranged in descending order of their ratings. There is no limit to the set number; it can be determined based on the total number of financial products included in the financial product data.
[0078] The advantage of this setup is that it allows for the selection of financial products that may meet the needs of those seeking financial products from all the financial products included in the financial product data, facilitating rapid product recommendations.
[0079] S131. Using a reinforcement learning network model, the combination of the short-range basic financial data and the operational feedback data is encoded into a state to be updated.
[0080] S132. Based on the state to be updated, a greedy algorithm is used to obtain the action to be updated, the execution reward of the action to be updated, and the update state of the state to be updated; wherein, the action to be updated is initially the candidate product recommendation list, and the execution reward is the operation feedback data obtained after the action to be updated is executed.
[0081] S133. Add the pending update status, the pending update action, the execution reward, and the update status to the experience pool.
[0082] S134. Determine the state-action prediction reward value of the network to be trained and the state-action target reward value of the target network through the experience pool, and update the parameters of the network to be trained by combining the state-action prediction reward value and the state-action target reward value using a loss function.
[0083] S135. The updated state is used as a new state to be updated for a new round of learning until the learning termination condition is met, and the termination state of the state to be updated is obtained.
[0084] S136. Based on the termination action in the termination state, determine the target product recommendation list.
[0085] Figure 5 This is a schematic diagram illustrating the working principle of a reinforcement learning network model according to Embodiment 2 of the present invention, combined with... Figure 5 The following explanation is given for S131-S136:
[0086] Short-range basic financial data and operational feedback data are used as inputs to a reinforcement learning network model, and several combinations of the two are encoded as states to be updated.
[0087] In each iteration, a greedy algorithm is used to obtain a list of recommended products based on the state to be updated s as the action to be updated a (a can initially be a list of candidate recommended products). Executing the action to be updated a means displaying the list of recommended products corresponding to a to the financial product demander. The execution reward r after the execution of a is obtained, which is the feedback data of the financial product demander's operation on a. After determining the execution reward r, the update state s' of the state to be updated s is obtained.
[0088] Add the pending update status s, pending update action a, execution reward r, and update status s' to the experience pool.
[0089] Based on the experience pool, the predicted state-action reward value of the network to be trained and the target state-action reward value of the target network are determined. The network to be trained can take the state to be updated s as input and output the predicted state-action reward value of selecting the action to be updated 'a' under the state to be updated. The target network can be a network with the same structure as the network to be trained, but with lagging parameter updates. It consists of the execution reward r of the action to be updated, plus a discount factor, and the product of the optimal state-action reward value of selecting the action to be updated 'a' under the updated state ', and outputs the target state-action reward value. The difference between the predicted state-action reward value and the target state-action reward value is calculated using a loss function. The parameters of the network to be trained are updated based on this difference, and the parameter updates make the predicted state-action reward value continuously approach the target state-action reward value.
[0090] The updated state is treated as the new state to be updated for a new round of learning until the learning termination condition is met, such as reaching a set number of iterations, to obtain the termination state of the state to be updated. The product recommendation list corresponding to the termination action in the termination state is determined as the target product recommendation list. The termination action in the termination state can be understood as selecting the termination action that yields the optimal state action reward value.
[0091] The advantage of this setup is that by using a reinforcement learning network model to update the candidate product recommendation list based on short-range basic financial data and operational feedback data, the target product recommendation list can be obtained. This enables dynamic optimization of the recommendation strategy to adapt to changes in the needs of financial product customers and improve user experience.
[0092] In one embodiment, the reinforcement learning network model includes an input layer, three hidden layers, and an output layer; wherein the input layer has an embedding layer for preprocessing the input data, and the hidden layers consist of fully connected layers and activation functions.
[0093] The reinforcement learning network model is configured with the following five layers: the first layer is the input layer, which includes an embedding layer to perform dimensionality reduction and dense mapping on the input data (short-distance basic financial data and operational feedback data); the last layer is the output layer, which outputs a list of recommended target products; and the middle three layers are hidden layers, which consist of fully connected layers and activation functions. The activation functions are selected based on both computational efficiency and gradient characteristics, and no specific restrictions are imposed here.
[0094] The advantage of this setup is that it allows for lower computational cost and higher online training efficiency in reinforcement learning network models by using multilayer perceptrons.
[0095] The technical solution of this invention enhances the system performance of financial product recommendation through deep learning, enabling it to process complex and unstructured information, deeply mine the relationship between implicit and explicit information, and improve the accuracy and diversity of recommendations. By focusing on interactive operation feedback data through reinforcement learning, it continuously optimizes recommendation strategies to adapt to the dynamic needs of financial product demanders and improve user experience.
[0096] Example 3
[0097] Figure 6 This is a schematic diagram of a financial product recommendation device according to Embodiment 3 of the present invention. This embodiment is applicable to situations where financial product recommendations are implemented, such as... Figure 6 As shown, the specific structure of the device includes:
[0098] The acquisition module 61 is used to acquire long-distance and short-distance data related to the demand side of financial products. The long-distance data includes at least long-distance basic financial data, interest preference data and financial product data, and the short-distance data includes at least short-distance basic financial data and operation feedback data.
[0099] The deep learning module 62 is used to determine a candidate product recommendation list by performing explicit feature interaction and implicit feature interaction based on the long-distance data through a deep learning network model.
[0100] The reinforcement learning module 63 is used to update the candidate product recommendation list based on the short-range data through a reinforcement learning network model, and to determine the target product recommendation list.
[0101] The financial product recommendation device provided in this embodiment acquires long-distance and short-distance data related to financial product demanders through an acquisition module. The long-distance data includes at least long-distance basic financial data, interest preference data, and financial product data, while the short-distance data includes at least short-distance basic financial data and operational feedback data. A deep learning module uses a deep learning network model to perform explicit and implicit feature interactions based on the long-distance data to determine a candidate product recommendation list. A reinforcement learning module uses a reinforcement learning network model to update the candidate product recommendation list based on the short-distance data to determine a target product recommendation list. This solution expands the types and time dimensions of data used for financial product recommendations, providing more comprehensive data support. It combines deep learning and reinforcement learning, possessing the dual advantages of deep learning's ability to handle complex multi-source information and reinforcement learning's emphasis on dynamic changes. This allows for a better understanding of the actual needs of financial product demanders, improving the match between recommended and desired financial products.
[0102] Furthermore, the deep learning module 62 is specifically used for:
[0103] The long-distance data is input into the embedding layer of the deep learning network model for preprocessing to obtain preprocessed long-distance data.
[0104] The preprocessed long-distance data is input into the compressed interaction network of the deep learning network model to perform explicit feature interaction, and the explicit feature interaction result is obtained.
[0105] The preprocessed long-distance data is input into the attention mechanism network of the deep learning network model to perform implicit feature interaction, and the implicit feature interaction result is obtained.
[0106] The candidate product recommendation list is determined based on the explicit feature interaction results and the implicit feature interaction results through the output unit of the deep learning network model.
[0107] Furthermore, the deep learning module 62 is specifically used for:
[0108] The results of the interaction between the explicit features and the interaction between the implicit features are weighted, and the weighted result is added to the bias term to obtain the summation result.
[0109] The summation result is processed by the execution function of the binary classification task to obtain the score results of each financial product included in the financial product data;
[0110] Based on the scoring results of each financial product, a list of recommended candidate products is determined.
[0111] Furthermore, the deep learning module 62 is specifically used for:
[0112] Sort the financial products by their scores from highest to lowest, and select the top 100 financial products to form a candidate product recommendation list.
[0113] Furthermore, reinforcement learning module 63 is specifically used for:
[0114] The combination of the short-range basic financial data and the operational feedback data is encoded into a state to be updated using a reinforcement learning network model.
[0115] Based on the state to be updated, a greedy algorithm is used to obtain the action to be updated, the execution reward of the action to be updated, and the update status of the state to be updated; wherein, the action to be updated is initially the candidate product recommendation list, and the execution reward is the operation feedback data obtained after the action to be updated is executed;
[0116] Add the pending update status, the pending update action, the execution reward, and the update status to the experience pool;
[0117] The state-action prediction reward value of the network to be trained and the state-action target reward value of the target network are determined by the experience pool, and the parameters of the network to be trained are updated by combining the state-action prediction reward value and the state-action target reward value using a loss function.
[0118] The updated state is used as the new state to be updated for a new round of learning until the learning termination condition is met, and the terminated state of the state to be updated is obtained.
[0119] Based on the termination action in the terminated state, a target product recommendation list is determined.
[0120] Furthermore, the reinforcement learning network model has a model structure including an input layer, three hidden layers, and an output layer; wherein, the input layer is equipped with an embedding layer to preprocess the input data, and the hidden layer consists of a fully connected layer and an activation function.
[0121] The financial product recommendation device provided in this embodiment of the invention can execute the financial product recommendation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0122] Example 4
[0123] Figure 7This is a schematic diagram of the structure of an electronic device implementing embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0124] like Figure 7 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory 12 or a random access memory 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the read-only memory 12 or loaded from storage unit 18 into the random access memory 13. The random access memory 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, read-only memory 12, and random access memory 13 are interconnected via a bus 14. An input / output interface 15 is also connected to the bus 14.
[0125] Multiple components in electronic device 10 are connected to input / output interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of monitors, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0126] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, central processing units, graphics processing units, various special-purpose artificial intelligence computing chips, various processors running machine learning model algorithms, digital signal processors, and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as financial product recommendation methods.
[0127] In some embodiments, the financial product recommendation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via read-only memory 12 and / or communication unit 19. When the computer program is loaded into random access memory 13 and executed by processor 11, one or more steps of the financial product recommendation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the financial product recommendation method by any other suitable means (e.g., by means of firmware).
[0128] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays, application-specific integrated circuits (ASICs), application-specific standard products (ASICs), systems-on-a-chip (SoCs), payload programmable logic devices, computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0129] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0130] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0131] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a cathode ray tube or liquid crystal display) for displaying information to a user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with a user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0132] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0133] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product within the cloud computing service system. This addresses the shortcomings of traditional physical hosts and virtual private servers, such as high management difficulty and weak business scalability.
[0134] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0135] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for recommending financial products, characterized in that, include: Acquire long-distance and short-distance data related to demanders of financial products. The long-distance data includes at least long-distance basic financial data, interest preference data, and financial product data. The short-distance data includes at least short-distance basic financial data and operational feedback data. By using a deep learning network model, explicit and implicit feature interactions are performed based on the long-distance data to determine a candidate product recommendation list. The candidate product recommendation list is updated based on the short-range data using a reinforcement learning network model to determine the target product recommendation list.
2. The method according to claim 1, characterized in that, Using a deep learning network model, based on the long-distance data, explicit and implicit feature interactions are performed to determine a candidate product recommendation list, including: The long-distance data is input into the embedding layer of the deep learning network model for preprocessing to obtain preprocessed long-distance data. The preprocessed long-distance data is input into the compressed interaction network of the deep learning network model to perform explicit feature interaction, and the explicit feature interaction result is obtained. The preprocessed long-distance data is input into the attention mechanism network of the deep learning network model to perform implicit feature interaction, and the implicit feature interaction result is obtained. The candidate product recommendation list is determined based on the explicit feature interaction results and the implicit feature interaction results through the output unit of the deep learning network model.
3. The method according to claim 2, characterized in that, A candidate product recommendation list is determined based on the explicit feature interaction results and the implicit feature interaction results, including: The results of the interaction between the explicit features and the interaction between the implicit features are weighted, and the weighted result is added to the bias term to obtain the summation result. The summation result is processed by the execution function of the binary classification task to obtain the score results of each financial product included in the financial product data; Based on the scoring results of each financial product, a list of recommended candidate products is determined.
4. The method according to claim 3, characterized in that, Based on the rating results of each financial product, a list of recommended candidate products is determined, including: Sort the financial products by their scores from highest to lowest, and select the top 100 financial products to form a candidate product recommendation list.
5. The method according to claim 1, characterized in that, By using a reinforcement learning network model, the candidate product recommendation list is updated based on the short-range data to determine the target product recommendation list, including: The combination of the short-range basic financial data and the operational feedback data is encoded into a state to be updated using a reinforcement learning network model. Based on the state to be updated, a greedy algorithm is used to obtain the action to be updated, the execution reward of the action to be updated, and the update status of the state to be updated; wherein, the action to be updated is initially the candidate product recommendation list, and the execution reward is the operation feedback data obtained after the action to be updated is executed; Add the pending update status, the pending update action, the execution reward, and the update status to the experience pool; The state-action prediction reward value of the network to be trained and the state-action target reward value of the target network are determined by the experience pool, and the parameters of the network to be trained are updated by combining the state-action prediction reward value and the state-action target reward value using a loss function. The updated state is used as the new state to be updated for a new round of learning until the learning termination condition is met, and the terminated state of the state to be updated is obtained. Based on the termination action in the terminated state, a target product recommendation list is determined.
6. The method according to claim 1, characterized in that, The reinforcement learning network model has an input layer, three hidden layers, and an output layer. The input layer has an embedding layer for preprocessing the input data, and the hidden layers consist of fully connected layers and activation functions.
7. A financial product recommendation device, characterized in that, include: The acquisition module is used to acquire long-distance and short-distance data related to the demand side of financial products. The long-distance data includes at least long-distance basic financial data, interest preference data, and financial product data, and the short-distance data includes at least short-distance basic financial data and operation feedback data. The deep learning module is used to determine a candidate product recommendation list by performing explicit and implicit feature interactions based on the long-distance data through a deep learning network model. The reinforcement learning module is used to update the candidate product recommendation list based on the short-range data using a reinforcement learning network model, and to determine the target product recommendation list.
8. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the method as described in any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method as described in any one of claims 1-6.