Model training method and device
By preprocessing and fine-tuning the service recommendation data of credit services through reinforcement learning, a sample generation model is generated and a second language model is trained, which solves the problem of insufficient service strategy matching score in the existing technology and achieves more efficient credit service recommendation.
Patent Information
- Application Number
- CN202510936836.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-11-07
AI Technical Summary
As competition intensifies in large language model services, existing technologies struggle to effectively improve the accuracy and efficiency of service recommendations, especially in credit service scenarios where service strategy matching and scoring processes are inadequate.
By preprocessing the service recommendation data of credit services, a pre-trained first language model is used for policy matching scoring, and a reinforcement learning method is used to fine-tune the model to generate a sample generation model. The second language model is then trained through the sample generation model to obtain a scoring model, thereby realizing policy matching scoring of service recommendation data.
This improved the accuracy and efficiency of service recommendations, enhanced the training efficiency of the model, ensured that the scoring model could more accurately perform policy matching and scoring, and improved the service recommendation effect of credit services.
Smart Images

Figure CN120911545A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present document relates to the technical field of data processing, and particularly relates to a model training method and device. BACKGROUND
[0002] With the continuous development and popularization of the Internet and artificial intelligence, many services can be automatically processed through large language models, such as in the scene of user service access, services can be processed to users through calling large language models, specifically, services can be recommended to users through calling large language models, but with more and more services using large language models, the competition between services is becoming more and more fierce, under this condition, higher requirements are put forward for service processing through large language models. SUMMARY
[0003] One or more embodiments of the present specification provide a model training method, comprising: preprocessing service recommendation data of a credit service to obtain preprocessed data. The preprocessed data and the prompt text are input into the first language model to perform matching processing of the preprocessed data and the service policy, and obtain the policy matching score of the preprocessed data. Based on the score label corresponding to the preprocessed data and the policy matching score, the first language model is fine-tuned, and a sample generation model is obtained after fine-tuning. Sample service data and sample labels are generated through the sample generation model, and a second language model is trained according to the generated sample service data and sample labels, to obtain a scoring model after training.
[0004] One or more embodiments of the present specification provide a service recommendation processing method, comprising: preprocessing service recommendation data of a credit service to obtain preprocessed data. The preprocessed data and the prompt text are input into the scoring model to perform matching score processing of the service policy, and obtain the policy matching score. Based on the policy matching score, service recommendation processing is performed on the service user corresponding to the service recommendation data. The scoring model is obtained by training a second language model based on sample service data and sample labels generated by a sample generation model, and the sample generation model is obtained by fine-tuning a first language model in a reinforcement learning manner.
[0005] One or more embodiments of the present specification provide a model training apparatus, comprising: a preprocessing module configured to preprocess service recommendation data of a credit service to obtain preprocessed data. A matching processing module is configured to perform matching processing of the preprocessed data with a prompt text input first language model and service policy, to obtain a policy matching score of the preprocessed data. A model fine-tuning module is configured to fine-tune the first language model based on the score label corresponding to the preprocessed data and the policy matching score, and obtain a sample generation model after fine-tuning is completed. A model training module is configured to generate sample service data and sample labels through the sample generation model, and train a second language model according to the generated sample service data and sample labels, to obtain a scoring model after training is completed.
[0006] One or more embodiments of the present specification provide a service recommendation processing apparatus, comprising: a preprocessing module configured to preprocess service recommendation data of a credit service to obtain preprocessed data. A matching score module is configured to perform matching score processing of the preprocessed data with a prompt text input scoring model and service policy, to obtain a policy matching score. A service recommendation module is configured to perform service recommendation processing on a service user corresponding to the service recommendation data based on the policy matching score. Wherein, the scoring model is obtained by training a second language model based on sample service data and sample labels generated by a sample generation model, and the sample generation model is obtained by fine-tuning a first language model in a reinforcement learning manner.
[0007] One or more embodiments of the present specification provide a model training device, comprising: a processor; and a memory configured to store computer executable instructions that, when executed, cause the processor to: preprocess service recommendation data of a credit service to obtain preprocessed data. Perform matching processing of the preprocessed data with a prompt text input first language model and service policy, to obtain a policy matching score of the preprocessed data. Fine-tune the first language model based on the score label corresponding to the preprocessed data and the policy matching score, and obtain a sample generation model after fine-tuning is completed. Generate sample service data and sample labels through the sample generation model, and train a second language model according to the generated sample service data and sample labels, to obtain a scoring model after training is completed.
[0008] The one or more embodiments of the specification provide a service recommendation processing device, comprising: a processor; and a memory configured to store computer executable instructions which, when executed, cause the processor to: preprocess service recommendation data of a credit service to obtain preprocessed data. Perform matching score processing on the preprocessed data and prompt text input to a service policy to obtain a policy matching score. Perform service recommendation processing on a service user corresponding to the service recommendation data based on the policy matching score. The score model is obtained by training a second language model based on sample service data and sample labels generated by a sample generation model, and the sample generation model is obtained by fine-tuning a first language model in a reinforcement learning manner.
[0009] The one or more embodiments of the specification provide a computer readable storage medium for storing computer executable instructions, which, when executed, implement the following processes: preprocess service recommendation data of a credit service to obtain preprocessed data. Perform matching processing on the preprocessed data and prompt text input to a service policy to obtain a policy matching score of the preprocessed data. Fine-tune the first language model based on the score label corresponding to the preprocessed data and the policy matching score, and obtain a sample generation model after fine-tuning is completed. Generate sample service data and sample labels through the sample generation model, and train a second language model according to the generated sample service data and sample labels to obtain a score model after training is completed.
[0010] The one or more embodiments of the specification provide another computer readable storage medium for storing computer executable instructions, which, when executed, implement the following processes: preprocess service recommendation data of a credit service to obtain preprocessed data. Perform matching score processing on the preprocessed data and prompt text input to a score model to obtain a policy matching score. Perform service recommendation processing on a service user corresponding to the service recommendation data based on the policy matching score. The score model is obtained by training a second language model based on sample service data and sample labels generated by a sample generation model, and the sample generation model is obtained by fine-tuning a first language model in a reinforcement learning manner. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the one or more embodiments of the specification or the prior art, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments described in the specification, and those skilled in the art can also obtain other drawings according to these drawings without creative labor. Figure 1 A schematic diagram of a model training method implementation environment provided for one or more embodiments of the present specification; Figure 2 A model training method processing flowchart provided for one or more embodiments of the present specification; Figure 3 A model training method processing flowchart applied to a credit service scenario provided for one or more embodiments of the present specification; Figure 4 A service recommendation processing method processing flowchart provided for one or more embodiments of the present specification; Figure 5 A schematic diagram of a model training device embodiment provided for one or more embodiments of the present specification; Figure 6 A schematic diagram of a service recommendation processing device embodiment provided for one or more embodiments of the present specification; Figure 7 A structural schematic diagram of a model training device provided for one or more embodiments of the present specification Figure 8 A structural schematic diagram of a service recommendation processing device provided for one or more embodiments of the present specification. DETAILED DESCRIPTION
[0012] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of the present specification, the technical solutions in one or more embodiments of the present specification will be described clearly and completely below in conjunction with the drawings in one or more embodiments of the present specification. Obviously, the described embodiments are only a part of the embodiments of the present specification, not all embodiments. Based on one or more embodiments of the present specification, all other embodiments obtained by those skilled in the art without creative labor should fall within the protection scope of the present document.
[0013] The model training method provided by one or more embodiments of the present specification can be applied to a model training system implementation environment, referring to Figure 1 The implementation environment at least includes: The first language model 101, the model fine-tuning component 102, the second language model 103 and the model training component 104; The first language model 101 refers to a pre-acquired base language model, the model fine-tuning component 102 is used for model fine-tuning of the first language model 101, and the sample generation model 105 is obtained after fine-tuning. The obtained sample generation model 105 is used to generate training samples and sample labels for model training of the second language model 103; The second language model 103 refers to the constructed natural language model. The model training component 104 is used to train the second language model 103 based on the training samples and sample labels generated by the sample generation model 105. After the model training is completed, the scoring model 106 is obtained. The generated scoring model 106 is used to perform corresponding strategy matching and scoring processing on the service recommendation data of resource services.
[0014] In this implementation environment, during model training, the service recommendation data for credit services is first preprocessed using the model fine-tuning component 102 to obtain preprocessed data. The preprocessed data and prompt text are then input into the first language model 101 for matching the preprocessed data with the service strategy to obtain a strategy matching score. Based on the score labels corresponding to the preprocessed data and the strategy matching score, the first language model 101 is fine-tuned, thus obtaining a sample generation model 105 after fine-tuning. Further, based on the sample generation model 105 obtained in the above training stage, sample service data and sample labels are generated by the sample generation model 105. The model training component 104 then trains the second language model 103 based on the generated sample service data and sample labels, thus obtaining a scoring model 106 capable of performing strategy matching scoring on the service recommendation data for credit services after training. In this way, a scoring model is obtained through the combination of fine-tuning and training of the natural language model in two stages.
[0015] One or more embodiments of a model training method provided in this specification are as follows: Reference Figure 2 The model training method provided in this embodiment specifically includes steps S202 to S208.
[0016] Step S202: Preprocess the service recommendation data of the credit service to obtain preprocessed data.
[0017] The credit service described in this embodiment refers to credit-related services provided by the service provider, or services provided by the service provider that perform corresponding processing based on credit. Optionally, the credit service includes credit payment services for order payments based on credit limits; for example, credit payment services provided to service users for order payments or order installment payments based on credit limits.
[0018] The service recommendation data refers to data related to service marketing of the credit service or service recommendation for a service user, or data used for service marketing or service recommendation of the credit service. Specifically, the service recommendation data can be user data related to the service user, such as user contact information, user intention, and / or user demand of the service user, and the like. The service recommendation data can also be data related to a service field in which the service party is located, such as market trend data, competitor dynamic data, or hot event data of the service field, and the like.
[0019] In a specific implementation process, the service recommendation data can also be lead data or business lead data related to marketing or recommendation of the credit service. In this case, in order to improve the collection efficiency of the service recommendation data, the data collection model can be used to collect the service recommendation data. Specifically, the lead data, business lead data, or service lead data obtained by the data collection model for collecting the lead data of the credit service can be used as the service recommendation data, that is, the service recommendation data includes the lead data, business lead data, or service lead data obtained by the data collection model for collecting the lead data of the credit service.
[0020] In a specific implementation, in the process of model training based on the obtained service recommendation data of the credit service, in order to improve the training accuracy of the model training and further improve the model training efficiency, the service recommendation data can be improved by data padding, so that the service recommendation data after data padding can carry more comprehensive and complete data information, thereby helping to improve the efficiency of model training based on the service recommendation data.
[0021] Specifically, in an optional implementation provided by the embodiment, the service recommendation data of the credit service is preprocessed to obtain preprocessed data, including: The service recommendation data is filled in a preset field for the credit service to obtain filled service data; The filled service data is verified, and the filled service data that passes the verification is used as the preprocessed data.
[0022] In a specific implementation process, the service recommendation data is filled in a preset field for the credit service, which means that the service recommendation data is filled in a preset field related to the service recommendation of the credit service, such as a transaction total amount field (or a transaction amount field), a user number field, and / or a product number field. Further, after the preset field filling of the service recommendation data is performed to obtain the filled service data, the filled fields may have abnormal, inaccurate or missing fields, etc. In order to avoid the influence of these abnormal situations on subsequent model training, the filled service data is subjected to data verification. The data verification can be to verify whether the filled service data has abnormal fields, such as verifying whether the filled service data has abnormal fields and / or missing fields. If there are abnormal fields, the filled service data with abnormal fields is removed. If there are no abnormal fields, the filled service data that passes the verification is used as preprocessed data.
[0023] In step S204, the preprocessed data is input into the first language model with the prompt text to perform matching processing of the preprocessed data and the service policy, and obtain a policy matching score of the preprocessed data.
[0024] The first language model refers to a pre-trained natural language model, which can be a pre-trained base large language model or a base language model. Specifically, a network architecture containing a large number of parameters can be used, such as an open-source base large language model using a neural network architecture.
[0025] Optionally, the first language model is configured with a policy knowledge base, and the policy knowledge base contains a plurality of service policies of the credit service. Specifically, in the case of using a base language model or a base large language model for the first language model, the policy knowledge base can be configured for the base language model or the base large language model.
[0026] The service policy refers to a service recommendation policy for recommending services to the service user of the credit service. Optionally, the service policy includes an increase policy for increasing the credit limit and / or a recommendation policy for recommending the service user to use the payment service to pay for the order.
[0027] In the implementation, on the basis that the first language model is configured with the policy knowledge base, the preprocessed data and the prompt text are input into the first language model. The prompt text is used to prompt the first language model to perform corresponding processing according to the content contained in the prompt text. Here, the prompt text is used to prompt the first language model to perform matching processing of the preprocessed data and the service policy. Specifically, the prompt text is used to prompt the first language model to perform scoring processing on the matching relationship between the preprocessed data and each service policy in the policy knowledge base.
[0028] Further, in order to make the first language model more accurate and comprehensive in scoring the matching relationship between the preprocessed data and each service policy in the policy knowledge base, the matching score dimension for scoring the matching relationship between the preprocessed data and each service policy in the policy knowledge base can be specified in the prompt text, so that the first language model scores the matching relationship between the preprocessed data and each service policy in the policy knowledge base according to the matching score dimension specified in the prompt text.
[0029] Optionally, the prompt text is used to prompt the first language model to score the matching relationship between the preprocessed data and each service policy according to the matching score dimension contained in the prompt text to obtain a sub-score; accordingly, the obtained policy matching score includes the sub-score of each matching score dimension.
[0030] The sub-score can be a score level or a score score corresponding to the score level, such as setting 5 score levels, and setting a score score for each score level. The first language model scores the matching relationship between the preprocessed data and each service policy according to the matching score dimension contained in the prompt text, specifically determines which score level the matching relationship between the preprocessed data and each service policy belongs to, and then takes the score score of the score level as the sub-score.
[0031] Further, the first language model can be prompted by the prompt text to weight the sub-score of each matching score dimension, so as to improve the calculation amount and processing efficiency of subsequent fine-tuning of the first language model; optionally, the prompt text is also used to prompt the first language model to weight and calculate the sub-score of each matching score dimension to obtain a weighted score; accordingly, the obtained policy matching score also includes the weighted score.
[0032] The weight of the matching score dimension can be specified in the prompt text, in addition to which the weight of the matching score dimension can also be determined by the first language model to determine the matching score dimension, specifically the weight of each matching score dimension can be determined by the first language model analyzing the importance of the matching score dimension, so that the weighted calculation can be performed on the basis of the weight of the matching score dimension.
[0033] In addition, the number of output sub-scores can also be set in the prompt text, such as prompting the first language model to sort the sub-score of the matching relationship between the preprocessed data and each service policy under each matching score dimension in descending order, and outputting the sub-score of the last N positions; similarly, the number of output weighted scores can also be set in the prompt text.
[0034] Optionally, the matching score dimension comprises at least one of the following: a data expression dimension for evaluating data density, redundancy and / or significance of the preprocessed data, a strategy matching dimension for evaluating matching degree of the preprocessed data with the service strategy, and a service value dimension for evaluating predicted service value of the service strategy.
[0035] The data expression dimension for evaluating data density, redundancy and / or significance of the preprocessed data refers to a dimension for evaluating data expression capability of the preprocessed data in at least one of data density, redundancy and significance. Specifically, the data expression capability of the preprocessed data in at least one of data density, redundancy and significance can be represented by a sub-score of the data expression dimension. Specifically, in the process of scoring the matching relationship between the preprocessed data and each service strategy by the first voice model in the data expression dimension, the data expression capability of the preprocessed data in at least one of data density, redundancy and significance is evaluated, a score level is determined according to the data expression capability, and a score of the score level is taken as a sub-score in the data expression dimension. Similarly, the strategy matching dimension for evaluating matching degree of the preprocessed data with the service strategy refers to a dimension for evaluating the matching relationship between the preprocessed data and the service strategy, and specifically can be a dimension for evaluating matching accuracy and / or matching completeness of the matching relationship between the preprocessed data and the service strategy. In the process of scoring the matching relationship between the preprocessed data and each service strategy by the first voice model in the strategy matching dimension, the matching degree of the preprocessed data with the service strategy is evaluated, a score level is determined according to the matching degree, and a score of the score level is taken as a sub-score in the strategy matching dimension. The service value dimension for evaluating predicted service value of the service strategy refers to a dimension for evaluating expected service value of the service strategy matched by the preprocessed data, i.e., a dimension for evaluating expected service value that can be generated by service recommendation according to the service strategy matched by the preprocessed data on the basis of the matching relationship between the preprocessed data and each service strategy in the strategy knowledge base. In the process of scoring the matching relationship between the preprocessed data and each service strategy by the first voice model in the service value dimension, the expected service value of the service strategy matched by the preprocessed data is evaluated, a score level is determined according to the service value, and a score of the score level is taken as a sub-score in the service value dimension.
[0036] In step S206, the first language model is fine-tuned based on the score label corresponding to the preprocessed data and the strategy matching score, and a sample generation model is obtained after fine-tuning is completed.
[0037] In this embodiment, the first language model is fine-tuned in a reinforcement learning manner, and the fine-tuned first language model is obtained after fine-tuning. The fine-tuned first language model is used to generate training samples and sample labels in the subsequent model training process. Therefore, the first language model obtained after fine-tuning is referred to as a sample generation model.
[0038] In actual implementation, in the process of fine-tuning the first language model in a reinforcement learning manner, the preprocessed data and the prompt text are input into the first language model to obtain the policy matching score of the preprocessed data through the matching processing of the preprocessed data and the service policy. Based on this, the first language model is fine-tuned based on the score label corresponding to the preprocessed data and the policy matching score. In this way, reinforcement learning of the first language model is realized, and finally a sample generation model is obtained after fine-tuning.
[0039] Specifically, in an optional implementation provided by the embodiment, fine-tuning the first language model based on the score label corresponding to the preprocessed data and the policy matching score includes: calculating the reward score of each policy matching score of the preprocessed data through a reward function; calculating the loss according to the reward score and the corresponding score label, and adjusting the parameters of the first language model according to the loss.
[0040] In the actual execution process, the loss can be calculated through a pre-set loss function in the process of calculating the loss according to the reward score and the corresponding score label. The above process of fine-tuning the first language model is repeated until the loss function converges or the parameters of the first language model converge. At this time, the first language model is the sample generation model obtained after fine-tuning. That is, the first language model in the case where the loss function converges or the parameters of the first language model converge is used as the sample generation model.
[0041] Step S208: generating sample service data and sample labels through the sample generation model, and training the second language model according to the generated sample service data and sample labels to obtain a scoring model after training.
[0042] In the first stage of model training, the first language model is fine-tuned in a reinforcement learning manner to obtain a sample generation model. On the basis of the obtained sample generation model, sample service data and sample labels for training the second language model are generated through the sample generation model, and the second language model is trained according to the sample service data and sample labels to obtain a scoring model after training. That is, a scoring model for performing policy matching score processing on the service recommendation data of the credit service is obtained.
[0043] The second language model refers to a pre-constructed natural language model, which can specifically adopt a network architecture containing preset parameters. The second language model can be a lightweight natural language model (lightweight language model), such as a base language model adopting a neural network architecture, and a scoring model obtained after training.
[0044] Here, the purpose of training the second language model is to enable the second language model to learn the matching relationship knowledge of the sample service data and the sample label, and to map the learned knowledge to the network architecture of the second language model. Specifically, the sample service data includes service recommendation data of credit services, and the sample label includes a policy matching score, that is, a score of the matching relationship between the service recommendation data and the service policy in the policy knowledge base. Based on this, the second language model can learn the ability to score the matching relationship between the service recommendation data and the service policy through model training, and after mapping the learned knowledge to the network architecture of the second language model, the scoring model obtained after training can have the ability to score the matching relationship between the service recommendation data and the service policy, that is, the ability of policy matching score processing of the service recommendation data.
[0045] Optionally, the second language model includes a plurality of matching score networks corresponding to the matching score dimensions and a score weighting module. The matching score network is used to score the matching relationship of the input service recommendation data corresponding to the matching score dimension to obtain a sub-score. The score weighting module is used to weight the sub-scores of the matching score networks to obtain a weighted score.
[0046] In specific implementation, in the process of training the second language model according to the generated sample service data and sample label, the sample service data and the prompt text are input into the second language model for policy matching score processing, the loss is calculated according to the output policy matching score and the sample label, and the parameters of the second language model are adjusted according to the obtained training loss. The training process is repeated until the convergence condition or the iteration termination condition is met, and the second language model meeting the convergence condition or the iteration termination condition is taken as the scoring model.
[0047] Specifically, in an optional implementation provided by the embodiment, the second language model is trained according to the generated sample service data and sample label, including: The sample service data and the prompt text are input into the second language model for policy matching score processing to obtain the service policy matched by the sample service data and the policy matching score. The loss is calculated based on the sample label and the policy matching score, and the parameters of the second language model are adjusted according to the obtained training loss.
[0048] In actual application, after the training of the second language model is completed to obtain the scoring model, the scoring model can be deployed to the credit service for corresponding policy matching scoring processing. Further, in order to improve the service recommendation effect of the credit service, feedback of the scoring model deployed to the credit service for actual policy matching scoring processing can be used to further optimize the scoring model, so as to improve the accuracy of the scoring model deployed to the credit service for policy matching scoring processing. Specifically, in an optional implementation provided by the embodiment, after the scoring model is obtained after the training is completed, the scoring result of the scoring model deployed to the credit service for policy matching scoring processing is obtained, and the scoring model is parameter-optimized according to the scoring result.
[0049] To sum up, the model training method provided by the embodiment improves the data comprehensiveness and integrity of the service recommendation data by preprocessing the service recommendation data in the model training process. Further, the first language model is fine-tuned to obtain a sample generation model by using a reinforcement learning method. Specifically, the preprocessed data obtained by preprocessing and the prompt text are input into the first language model to obtain a policy matching score of the preprocessed data through matching processing of the preprocessed data and the service policy. The first language model is fine-tuned based on the score label corresponding to the preprocessed data and the policy matching score. In this way, the sample generation model is obtained after the fine-tuning is completed. Further, based on the obtained sample generation model, the sample service data and the sample label for training the second language model are generated by the sample generation model, and the second language model is trained according to the sample service data and the sample label. After the training is completed, the scoring model is obtained, that is, the scoring model for policy matching scoring processing of the service recommendation data of the credit service is obtained. In this way, the two-stage model training is used to realize automatic scoring model training, and the training efficiency of the model training is improved.
[0050] The following takes the application of the model training method provided by the embodiment in the credit service scenario as an example, and combines the above description of the model training method provided by the embodiment. Figure 3 The model training method applied in the credit service scenario is further described with reference to Figure 3 The model training method applied in the credit service scenario includes the following steps.
[0051] In step S302, the service recommendation data of the credit service is pre-filled in a preset field to obtain filled service data.
[0052] In step S304, the filled service data is subjected to data verification, and the filled service data that passes the verification is used as preprocessed data.
[0053] In step S306, the preprocessed data and the prompt text are input into the first language model to perform matching processing of the preprocessed data and the service policy, and a policy matching score of the preprocessed data is obtained.
[0054] In step S308, the reward score of each policy matching score of the preprocessed data is calculated by a reward function.
[0055] In step S310, the loss is calculated according to the reward score and the corresponding score label, and the first language model is adjusted in parameters according to the loss, and the sample generation model is obtained after fine-tuning is completed.
[0056] In step S312, the sample service data and the sample label are generated by the sample generation model.
[0057] In step S314, the sample service data and the prompt text are input into the second language model for policy matching score processing, and the policy matching score of the sample service data is obtained.
[0058] In step S316, the loss is calculated based on the sample label and the policy matching score, and the second language model is adjusted in parameters according to the training loss, so as to obtain the scoring model after training is completed.
[0059] In step S318, the scoring result of the scoring model deployed to the credit service for policy matching score processing is obtained, and the scoring model is optimized in parameters according to the scoring result.
[0060] It should be noted that any one step or combination of any multiple steps in steps S302 to S318 can be combined with any one step or multiple steps in steps S202 to S208 to form a new implementation mode according to the needs of implementation and deployment. In addition, any one or more technical features in steps S302 to S318 can be combined with any one or more technical features provided in steps S202 to S208 to form a new implementation mode according to the actual needs of deployment. Alternatively, any one or more technical features in steps S302 to S318 can be replaced by any one or more technical features provided in steps S202 to S208 to form a new implementation mode according to the actual needs of deployment, which will not be described one by one here.
[0061] The service recommendation processing method provided in the specification is implemented as follows: Referring to Figure 4 The service recommendation processing method provided in the embodiment specifically includes steps S402 to S406.
[0062] In step S402, the service recommendation data of the credit service is preprocessed to obtain preprocessed data.
[0063] The credit service described in the embodiment refers to a credit-related service provided by the service party or a service provided by the service party based on credit for corresponding processing; optionally, the credit service includes a credit payment service based on a credit limit for order payment; for example, a credit payment service based on a credit limit for order payment or order installment payment provided to the service user.
[0064] The service recommendation data refers to data related to service marketing or service recommendation of the credit service for service users, or data used for service marketing or service recommendation of the credit service; specifically, the service recommendation data can be user data related to the service user, such as user contact information, user intention, and / or user demand of the service user, and the like. The service recommendation data can also be data related to the service field of the service party, such as market trend data, competitor dynamic data, or hot event data of the service field, and the like.
[0065] In the specific implementation process, the service recommendation data can also be lead data or business lead data related to the marketing or recommendation of the credit service; in this case, in order to improve the collection efficiency of the service recommendation data, the data collection model can be used to collect the service recommendation data; specifically, the lead data, business lead data, or service lead data obtained by the data collection model for the lead data acquisition of the credit service can be used as the service recommendation data, that is, the service recommendation data includes the lead data, business lead data, or service lead data obtained by the data collection model for the lead data collection of the credit service.
[0066] In the specific implementation, in the process of performing service recommendation based on the obtained service recommendation data of the credit service, in order to improve the accuracy of the service recommendation, the service recommendation data can be improved by data filling, so that the service recommendation data after data filling can carry more comprehensive and complete data information, thereby helping to improve the efficiency of service recommendation based on the service recommendation data.
[0067] Specifically, in an optional implementation provided by the embodiment, the service recommendation data of the credit service is preprocessed to obtain preprocessed data, including: The service recommendation data is filled in a preset field of the credit service to obtain filled service data; The filled service data is verified, and the filled service data that passes the verification is used as the preprocessed data.
[0068] In a specific implementation process, the preset field filling of the credit service on the service recommendation data refers to filling the preset fields related to the service recommendation of the credit service in the service recommendation data, such as filling the transaction total field (or the transaction amount field), the user number field, and / or the commodity number field in the service recommendation data. Further, after obtaining the filled service data by performing the preset field filling of the credit service on the service recommendation data, the filled fields may have abnormal, inaccurate, or missing fields. In order to avoid the influence of these abnormal situations on subsequent model training, the filled service data is subjected to data verification. The data verification can be to verify whether the filled service data has abnormal fields, such as verifying whether the filled service data has abnormal fields and / or missing fields. If there are abnormal fields, the filled service data with abnormal fields is removed. If there are no abnormal fields, the filled service data that passes the verification is used as the preprocessed data.
[0069] In step S404, the preprocessed data is input into the scoring model together with the prompt text to perform matching scoring processing with the service policy, and a policy matching score is obtained.
[0070] Optionally, the scoring model is obtained by training the second language model based on the sample service data and the sample label generated by the sample generation model. The sample generation model is obtained by fine-tuning the first language model in a reinforcement learning manner.
[0071] In this embodiment, in the process of fine-tuning the first language model in a reinforcement learning manner, the preprocessed data obtained by preprocessing the service recommendation data is input into the first language model together with the prompt text to perform matching processing of the preprocessed data with the service policy, and a policy matching score of the preprocessed data is obtained. Based on this, the first language model is fine-tuned based on the score label corresponding to the preprocessed data and the policy matching score of the preprocessed data. In this way, the reinforcement learning of the first language model is realized, and finally the sample generation model is obtained after the fine-tuning is completed, that is, the sample generation model is obtained by fine-tuning the first language model in a reinforcement learning manner.
[0072] The first language model refers to a natural language model that has been pre-trained, which can be a pre-trained base large language model or a base language model. Specifically, a network architecture containing a large number of parameters can be used, such as an open-source base large language model using a neural network architecture.
[0073] Optionally, the first language model is configured with a policy knowledge base, and the policy knowledge base contains multiple service policies of the credit service. Specifically, in the case where the first language model uses a base language model or a base large language model, the policy knowledge base can be configured for the base language model or the base large language model.
[0074] The service strategy refers to a service recommendation strategy used for service recommendation to a service user of a credit service; optionally, the service strategy includes an increase strategy for increasing a credit limit and / or a recommendation strategy for recommending the service user to use a payment service for order payment.
[0075] In a specific implementation, on the basis that the first language model is configured with the strategy knowledge base, the first language model is prompted to perform corresponding processing according to the content contained in the prompt text by inputting the preprocessed data and the prompt text into the first language model, and the prompt text is specifically used to prompt the first language model to perform matching processing of the preprocessed data and the service strategy. Specifically, the prompt text is used to prompt the first language model to perform scoring processing on the matching relationship between the preprocessed data and each service strategy in the strategy knowledge base.
[0076] Further, in order to make the scoring processing of the first language model on the matching relationship between the preprocessed data and each service strategy in the strategy knowledge base more accurate and more comprehensive, the matching scoring dimensions for scoring processing on the matching relationship between the preprocessed data and each service strategy in the strategy knowledge base can also be specified in the prompt text, so that the first language model performs scoring processing on the matching relationship between the preprocessed data and each service strategy in the strategy knowledge base according to the matching scoring dimensions specified in the prompt text.
[0077] Optionally, the prompt text is used to prompt the first language model to score the matching relationship between the preprocessed data and each service strategy according to the matching scoring dimensions contained in the prompt text to obtain sub-scores; correspondingly, the obtained strategy matching score includes sub-scores of each matching scoring dimension.
[0078] The sub-score can be a scoring level or a scoring score corresponding to the scoring level, such as setting 5 scoring levels, and setting a scoring score for each scoring level. The first language model scores the matching relationship between the preprocessed data and each service strategy according to the matching scoring dimensions contained in the prompt text, specifically determines which scoring level the matching relationship between the preprocessed data and each service strategy belongs to, and then takes the scoring score of the scoring level as the sub-score.
[0079] Further, the first language model can also be prompted by the prompt text to perform weighting processing on the sub-scores of each matching scoring dimension, so as to improve the calculation amount and processing efficiency of subsequent fine-tuning of the first language model; optionally, the prompt text is also used to prompt the first language model to perform weighted calculation on the sub-scores of each matching scoring dimension to obtain a weighted score; correspondingly, the obtained strategy matching score also includes the weighted score.
[0080] The weight of the matching score dimension can be specified in the prompt text, and in addition, the weight of the matching score dimension can be determined by the first language model to determine the matching score dimension, and specifically, the weight of each matching score dimension can be determined by the first language model analyzing the importance of the matching score dimension, so that weighted calculation can be performed on the basis of the weight of the matching score dimension.
[0081] In addition, the number of output sub-scores can also be set in the prompt text, such as for each matching score dimension, prompting the first language model to sort the sub-scores of the matching relationship between the preprocessed data and each service policy under each matching score dimension in descending order, and outputting the sub-scores of the last N positions; similarly, the number of output weighted scores can also be set in the prompt text.
[0082] Optionally, the matching score dimension includes at least one of the following: a data expression dimension for evaluating the data density, redundancy and / or significance of the preprocessed data, a policy matching dimension for evaluating the matching degree of the preprocessed data and the service policy, and a service value dimension for evaluating the predicted service value of the service policy.
[0083] The data expression dimension for evaluating the data density, redundancy and / or significance of the preprocessed data refers to a dimension for evaluating the data expression capability of the preprocessed data in at least one of the data density, redundancy and significance, and specifically, the data expression capability of the preprocessed data in at least one of the data density, redundancy and significance can be represented by a sub-score of the data expression dimension; specifically, in the process of scoring the matching relationship between the preprocessed data and each service policy by the first voice model in the data expression dimension, the data expression capability of the preprocessed data in at least one of the data density, redundancy and significance is evaluated, the score level is determined according to the data expression capability, and the score of the score level is taken as the sub-score in the data expression dimension; Similarly, the policy matching dimension for evaluating the matching degree of the preprocessed data and the service policy refers to a dimension for evaluating the matching relationship between the preprocessed data and the service policy, and specifically, it can be a dimension for evaluating the matching accuracy and / or matching integrity of the matching relationship between the preprocessed data and the service policy; in the process of scoring the matching relationship between the preprocessed data and each service policy by the first voice model in the policy matching dimension, the matching degree of the preprocessed data and the service policy is evaluated, the score level is determined according to the matching degree, and the score of the score level is taken as the sub-score in the policy matching dimension; The service value dimension for evaluating the predicted service value of the service strategy is a dimension for evaluating the expected service value of the service strategy matched with the preprocessed data, that is, a dimension for evaluating the expected service value of the service recommendation performed by the service strategy matched with the preprocessed data on the basis of the matching relationship between the preprocessed data and each service strategy in the strategy knowledge base; in the process of scoring the matching relationship between the preprocessed data and each service strategy in the service value dimension by the first voice model, the expected service value of the service strategy matched with the preprocessed data is specifically evaluated, the scoring level is determined according to the service value, and the score of the scoring level is taken as the sub-score in the service value dimension.
[0084] In an optional implementation of the embodiment, the first language model is fine-tuned based on the scoring label corresponding to the preprocessed data and the strategy matching score, including: The reward score of each strategy matching score of the preprocessed data is calculated by the reward function; The loss is calculated according to the reward score and the corresponding scoring label, and the parameters of the first language model are adjusted according to the loss.
[0085] In the specific execution process, the loss can be calculated by a pre-set loss function in the process of calculating the loss according to the reward score and the corresponding scoring label; the above fine-tuning process of the first language model is repeated until the loss function converges or the parameters of the first language model converge, and the first language model at this time is the sample generation model obtained after fine-tuning, that is, the first language model in the case of converging loss function or converging parameters of the first language model is taken as the sample generation model.
[0086] In the embodiment, on the basis of fine-tuning the first language model by the reinforcement learning method to obtain the sample generation model, the sample service data and the sample label for training the second language model are generated by the sample generation model, and the second language model is trained according to the sample service data and the sample label to obtain the scoring model after training, that is, to obtain the scoring model for performing the strategy matching scoring process on the service recommendation data of the credit service.
[0087] The second language model refers to a pre-constructed natural language model, which can specifically adopt a network architecture containing preset parameters, and can be a lightweight natural language model (lightweight language model), such as a base language model adopting a neural network architecture, to obtain the scoring model after training; as a natural language model, the scoring model can perform the strategy matching scoring process on the service data according to the input prompt text.
[0088] Here, the purpose of training the second language model is to enable the second language model to learn the knowledge of the matching relationship between the sample service data and the sample label, and map the learned knowledge into its own network architecture. Specifically, the sample service data includes service recommendation data of a credit service, and the sample label includes a policy matching score, i.e., a score of the matching relationship between the service recommendation data and a service policy in the policy knowledge base. Based on this, the model training can enable the second language model to learn the ability to score the matching relationship between the service recommendation data and the service policy, and after mapping the learned knowledge into the network architecture of the second language model, the scoring model obtained after the training can have the ability to score the matching relationship between the service recommendation data and the service policy, i.e., have the ability of policy matching score processing of the service recommendation data.
[0089] Optionally, the second language model includes a plurality of matching score networks corresponding to the matching score dimensions and a score weighting module; wherein the matching score network is used to score the matching relationship of the corresponding matching score dimension of the input service recommendation data to obtain a sub-score; and the score weighting module is used to weight the sub-scores of the matching score networks to obtain a weighted score.
[0090] In specific implementation, in the process of training the second language model according to the generated sample service data and sample label, the sample service data and the prompt text are input into the second language model for policy matching score processing, the loss is calculated according to the output policy matching score and the sample label, and the parameters of the second language model are adjusted according to the obtained training loss. The training process is repeated until the convergence condition or the iteration termination condition is met, and the second language model meeting the convergence condition or the iteration termination condition is taken as the scoring model.
[0091] Specifically, in an optional implementation provided by the embodiment, training the second language model according to the generated sample service data and sample label includes: inputting the sample service data and the prompt text into the second language model for policy matching score processing to obtain the service policy matched by the sample service data and the policy matching score; calculating the loss based on the sample label and the policy matching score, and adjusting the parameters of the second language model according to the obtained training loss.
[0092] Step S406, performing service recommendation processing on the service user corresponding to the service recommendation data based on the policy matching score.
[0093] In a specific implementation, in a process of performing service recommendation processing on a service user corresponding to the service recommendation data based on the policy matching score, a target service policy can be determined according to the policy matching score, and the service recommendation processing is performed on the service user corresponding to the service recommendation data according to the target service policy.
[0094] The service policy is a service recommendation policy used for performing service recommendation on a service user of a credit service. Optionally, the service policy includes an increase policy for increasing a credit limit and / or a recommendation policy for recommending the service user to use a payment service to perform order payment.
[0095] In a process of determining the target service policy according to the policy matching score, the policy matching scores can be sorted in descending order, and a service policy corresponding to a policy matching score at a top of the sorting is taken as the target service policy. The service policy corresponding to the policy matching score can be a service policy corresponding to the policy matching score output by the scoring model.
[0096] The model training apparatus provided in the specification is implemented, for example, as follows. In the above-described embodiments, a model training method is provided, and a model training apparatus corresponding to the model training method is also provided, which will be described below with reference to the accompanying drawings.
[0097] Reference is made to Figure 5 which shows a schematic diagram of an embodiment of the model training apparatus provided in the present embodiment.
[0098] Since the apparatus embodiment corresponds to the method embodiment, the description is relatively simple, and the relevant part can be referred to the above-described corresponding description of the method embodiment. The apparatus embodiment described below is merely illustrative.
[0099] The present embodiment provides a model training apparatus, which comprises: The preprocessing module 502 is configured to preprocess service recommendation data of a credit service to obtain preprocessed data. The matching processing module 504 is configured to perform matching processing of the preprocessed data and a service policy by inputting the preprocessed data and a prompt text into a first language model, to obtain a policy matching score of the preprocessed data. The model fine-tuning module 506 is configured to fine-tune the first language model based on a scoring label corresponding to the preprocessed data and the policy matching score, and obtain a sample generation model after the fine-tuning is completed. The model training module 508 is configured to generate sample service data and sample labels by using the sample generation model, and train a second language model according to the generated sample service data and sample labels, to obtain a scoring model after the training is completed.
[0100] The service recommendation processing apparatus provided in the specification implements, for example, the following: In the above-described embodiment, a service recommendation processing method is provided, and a service recommendation processing apparatus corresponding thereto is also provided, which will be described below with reference to the accompanying drawings.
[0101] Reference Figure 6 which shows a schematic diagram of an embodiment of a service recommendation processing apparatus provided in the present embodiment.
[0102] Since the apparatus embodiment corresponds to the method embodiment, it is described more simply, and the relevant part can be seen from the above-described corresponding description of the method embodiment. The apparatus embodiment described below is only schematic.
[0103] The present embodiment provides a service recommendation processing apparatus, which comprises: The preprocessing module 602 is configured to preprocess the service recommendation data of the credit service to obtain preprocessed data. The matching score module 604 is configured to perform matching score processing on the preprocessed data and the prompt text input score model according to the service policy to obtain a policy matching score. The service recommendation module 606 is configured to perform service recommendation processing on the service user corresponding to the service recommendation data based on the policy matching score. The score model is obtained by training a second language model based on sample service data and sample labels generated by a sample generation model, and the sample generation model is obtained by fine-tuning a first language model in a reinforcement learning manner.
[0104] The model training device provided in the specification implements, for example, the following: Corresponding to the above-described model training method, based on the same technical concept, one or more embodiments of the present specification also provide a model training device for executing the above-described model training method, Figure 7 The structure diagram of the model training device provided in one or more embodiments of the present specification.
[0105] The model training device provided in the present embodiment comprises: As Figure 7As shown, the model training device can have a large difference due to different configurations or performances, and can include one or more processors 701 and memories 702, and one or more applications or data can be stored in the memories 702. The memories 702 can be temporary storage or persistent storage. The applications stored in the memories 702 can include one or more modules (not shown in the figure), and each module can include a series of computer executable instructions in the model training device. Further, the processor 701 can be configured to communicate with the memory 702 and execute a series of computer executable instructions in the memory 702 on the model training device. The model training device can also include one or more power supplies 703, one or more wired or wireless network interfaces 704, one or more input / output interfaces 705, one or more keyboards 706, and the like.
[0106] In one specific embodiment, the model training device includes a memory and one or more programs, wherein one or more programs are stored in the memory, and the one or more programs can include one or more modules, and each module can include a series of computer executable instructions in the model training device, and the one or more programs configured to be executed by one or more processors include computer executable instructions for: preprocessing the service recommendation data of the credit service to obtain preprocessed data; performing matching processing of the preprocessed data and the service policy by inputting the preprocessed data and the prompt text into the first language model, to obtain a policy matching score of the preprocessed data; based on the score label corresponding to the preprocessed data and the policy matching score, fine-tuning the first language model, and obtaining a sample generation model after fine-tuning is completed; generating sample service data and sample labels through the sample generation model, and training a second language model according to the generated sample service data and sample labels, to obtain a scoring model after training is completed.
[0107] The service recommendation processing device provided in the specification implements, for example: According to the same technical concept as described above, one or more embodiments of the specification also provide a service recommendation processing device for executing the model training method provided above, Figure 8 The structure diagram of the service recommendation processing device provided in one or more embodiments of the specification.
[0108] The model training device provided in the embodiment includes: As Figure 8As shown, the model training device can have a large difference due to different configurations or performances, and can include one or more processors 801 and memories 802, and one or more applications or data can be stored in the memories 802. The memories 802 can be temporary storage or persistent storage. The applications stored in the memories 802 can include one or more modules (not shown in the figure), and each module can include a series of computer executable instructions in the model training device. Further, the processor 801 can be configured to communicate with the memory 802 and execute the series of computer executable instructions in the memory 802 on the model training device. The model training device can also include one or more power supplies 803, one or more wired or wireless network interfaces 804, one or more input / output interfaces 805, one or more keyboards 806, and the like.
[0109] In one specific embodiment, the model training device includes a memory and one or more programs, wherein one or more programs are stored in the memory, and the one or more programs can include one or more modules, and each module can include a series of computer executable instructions in the model training device, and the one or more programs configured to be executed by the one or more processors include computer executable instructions for: preprocessing service recommendation data of a credit service to obtain preprocessed data; performing matching score processing of the preprocessed data with prompt text input scoring model and service policy to obtain policy matching score; performing service recommendation processing based on the policy matching score to a service user corresponding to the service recommendation data; The scoring model is obtained by training a second language model based on sample service data and sample labels generated by a sample generation model, and the sample generation model is obtained by fine-tuning a first language model in a reinforcement learning manner.
[0110] The computer readable storage medium provided in the specification implements, for example: Based on the same technical concept, the one or more embodiments of the specification also provide a computer readable storage medium corresponding to the model training method described above.
[0111] The computer readable storage medium provided in the embodiment is used to store computer executable instructions, and the computer executable instructions are executed to implement the following processes: preprocessing service recommendation data of a credit service to obtain preprocessed data; The preprocessed data and prompt text are input into the first language model for matching processing of the preprocessed data and the service policy, to obtain a policy matching score of the preprocessed data; The first language model is fine-tuned based on the score label corresponding to the preprocessed data and the policy matching score, and after fine-tuning is completed, a sample generation model is obtained; Sample service data and sample labels are generated through the sample generation model, and a second language model is trained according to the generated sample service data and sample labels, to obtain a score model after training is completed.
[0112] It should be noted that the embodiment of the computer-readable storage medium in the present specification is based on the same inventive concept as the embodiment of the model training method in the present specification, and therefore the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described herein.
[0113] Another computer-readable storage medium embodiment provided by the present specification is as follows: Corresponding to the service recommendation processing method described above, based on the same technical concept, one or more embodiments of the present specification also provide another computer-readable storage medium.
[0114] The computer-readable storage medium provided in the present embodiment is used to store computer executable instructions, and the computer executable instructions, when executed, implement the following flow: The service recommendation data of the credit service is preprocessed to obtain preprocessed data; The preprocessed data and prompt text are input into the score model for matching score processing with the service policy, to obtain a policy matching score; The service recommendation processing is performed on the service user corresponding to the service recommendation data based on the policy matching score; The score model is obtained by training a second language model based on sample service data and sample labels generated by a sample generation model, and the sample generation model is obtained by fine-tuning a first language model in a reinforcement learning manner.
[0115] It should be noted that the embodiment of another computer-readable storage medium in the present specification is based on the same inventive concept as the embodiment of the service recommendation processing method in the present specification, and therefore the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described herein.
[0116] A computer program product embodiment provided by the present specification is as follows: Corresponding to the model training method described above, based on the same technical concept, one or more embodiments of the present specification also provide a computer program product.
[0117] A computer program product comprises computer programs / instructions which, when executed by a processor, implement the following steps: Preprocessing service recommendation data of a credit service to obtain preprocessing data; Matching the preprocessing data with prompt text input into a first language model to obtain a policy matching score of the preprocessing data; Fine-tuning the first language model based on the score label corresponding to the preprocessing data and the policy matching score, and obtaining a sample generation model after fine-tuning is completed; Generating sample service data and sample labels through the sample generation model, and training a second language model according to the generated sample service data and sample labels to obtain a scoring model after training is completed.
[0118] It should be noted that the embodiments of the computer program product in the present specification and the embodiments of the model training method in the present specification are based on the same inventive concept, and therefore the specific implementation of the embodiments can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0119] Another computer program product provided by the present specification is as follows: Corresponding to the service recommendation processing method described above, based on the same technical concept, one or more embodiments of the present specification also provide another computer program product.
[0120] A computer program product comprises computer programs / instructions which, when executed by a processor, implement the following steps: Preprocessing service recommendation data of a credit service to obtain preprocessing data; Matching the preprocessing data with prompt text input into a first language model to obtain a policy matching score of the preprocessing data; Performing service recommendation processing on a service user corresponding to the service recommendation data based on the policy matching score; The scoring model is obtained by training a second language model based on sample service data and sample labels generated by a sample generation model, and the sample generation model is obtained by fine-tuning a first language model in a reinforcement learning manner.
[0121] It should be noted that the embodiments of the computer program product in the present specification and the embodiments of the model training method in the present specification are based on the same inventive concept, and therefore the specific implementation of the embodiments can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0122] The various embodiments in the specification are described in progressive manner, and the same or similar parts among the various embodiments can be referred to each other, and each embodiment focuses on the difference from other embodiments, such as the device embodiment, the equipment embodiment, the computer readable storage medium embodiment, the computer program product embodiment, which are similar to the method embodiment, so the description is relatively simple, and the related content in the device embodiment, the equipment embodiment, the computer readable storage medium embodiment and the computer program product embodiment can be referred to the part of the description of the method embodiment.
[0123] The above describes specific embodiments of the specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than the order in which they are recited and still achieve desirable results. In addition, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous or necessary.
[0124] In the 1930s, it was clear to distinguish whether an improvement in a technology was in hardware (e.g., improvement in circuit structures of diodes, transistors, switches, etc.) or in software (e.g., improvement in method flow). However, as technology has evolved, many improvements in method flow today can be considered as direct improvements in hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structures by programming the improved method flow into hardware circuits. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using hardware entity modules. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is an integrated circuit whose logic function is determined by user programming of the device. A digital system is "integrated" on a piece of PLD by the designer programming it by himself, without having to ask a chip manufacturer to design and manufacture a special integrated circuit chip. Moreover, instead of manually fabricating integrated circuit chips, this programming is now mostly implemented using "logic compiler" software, which is similar to the software compiler used when developing programs, and the original code before compilation also has to be written in a specific programming language, which is called a hardware description language (HDL), and there are many such languages, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc., and the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that it is easy to obtain hardware circuits that implement the logic method flow by simply logically programming the method flow in the above-mentioned hardware description languages and programming it into an integrated circuit.
[0125] The controller can be implemented in any suitable way, for example, the controller can take the form of a microprocessor or processor and a computer readable medium storing computer readable program code, such as software or firmware, executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20 and Silicone Labs C8051F320, the memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that, in addition to implementing the controller in pure computer readable program code, it is possible to implement the same functionality in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers and embedded microcontrollers by logically programming the method steps. Such a controller can therefore be considered to be a hardware component, and the means included therein for implementing the various functions can also be considered to be structures within the hardware component. Alternatively, or even additionally, the means for implementing the various functions can be considered to be both a software module implementing the method and a structure within the hardware component.
[0126] The systems, apparatuses, modules or units illustrated by the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0127] For the sake of description, the above apparatuses are described in various units with functions respectively. Of course, the functions of the units can be implemented in one or more software and / or hardware in implementing the embodiments of the present specification.
[0128] Those skilled in the art will understand that one or more embodiments of the present specification can be provided as a method, a system or a computer program product. Therefore, one or more embodiments of the present specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present specification can take the form of a computer program product implemented on one or more computer readable storage media (including but not limited to disk storage, CD-ROMs, optical storage devices, etc.) containing computer usable program code.
[0129] The specification is presented with reference to flow diagrams and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the specification. It will be understood that each block of the flow diagrams and / or block diagrams, and combinations of blocks in the flow diagrams and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing unit or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flow diagrams and / or block diagrams block or blocks. Figure 1 The flow diagrams and / or block diagrams in the specification can present a method, apparatus (system) and / or computer program product according to embodiments of the specification. It will be understood that each block of the flow diagrams and / or block diagrams, and combinations of blocks in the flow diagrams and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing unit or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flow diagrams and / or block diagrams block or blocks. Figure 1 The flow diagrams and / or block diagrams in the specification can present a method, apparatus (system) and / or computer program product according to embodiments of the specification. It will be understood that each block of the flow diagrams and / or block diagrams, and combinations of blocks in the flow diagrams and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing unit or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flow diagrams and / or block diagrams block or blocks.
[0130] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flow diagrams and / or block diagrams block or blocks. Figure 1 The flow diagrams and / or block diagrams in the specification can present a method, apparatus (system) and / or computer program product according to embodiments of the specification. It will be understood that each block of the flow diagrams and / or block diagrams, and combinations of blocks in the flow diagrams and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing unit or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flow diagrams and / or block diagrams block or blocks. Figure 1 The flow diagrams and / or block diagrams in the specification can present a method, apparatus (system) and / or computer program product according to embodiments of the specification. It will be understood that each block of the flow diagrams and / or block diagrams, and combinations of blocks in the flow diagrams and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing unit or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flow diagrams and / or block diagrams block or blocks.
[0131] These computer program instructions can also be loaded into a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flow diagrams and / or block diagrams block or blocks. Figure 1 The flow diagrams and / or block diagrams in the specification can present a method, apparatus (system) and / or computer program product according to embodiments of the specification. It will be understood that each block of the flow diagrams and / or block diagrams, and combinations of blocks in the flow diagrams and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing unit or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flow diagrams and / or block diagrams block or blocks. Figure 1 The flow diagrams and / or block diagrams in the specification can present a method, apparatus (system) and / or computer program product according to embodiments of the specification. It will be understood that each block of the flow diagrams and / or block diagrams, and combinations of blocks in the flow diagrams and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing unit or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flow diagrams and / or block diagrams block or blocks.
[0132] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0133] The memory can include non-persistent memory, Random Access Memory (RAM), and / or non-volatile memory, such as Read Only Memory (ROM) or flash memory, among others in a computer readable medium. The memory is an example of computer readable media.
[0134] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer-readable storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.
[0135] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover non-exclusive inclusions, such that a process, method, article or apparatus that comprises a list of elements does not only include those elements, but also other elements not explicitly listed or inherent to such a process, method, article or apparatus. Without more limitations, the element defined by the phrase "comprising at least one" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.
[0136] One or more embodiments of the present specification can be described in the general context of computer-executable instructions being executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform particular tasks or implement particular abstract data types. One or more embodiments of the present specification can also be practiced in a distributed computing environment, in which tasks are performed by remote processing devices that are connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0137] The above only describes the embodiments of the present document and is not intended to limit the present document. For those skilled in the art, the present document can have various changes and variations. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present document shall be included in the scope of claims of the present document.
Claims
1. A model training method, comprising: preprocessing service recommendation data of a credit service to obtain preprocessed data; performing matching processing of the preprocessed data and prompt text on a first language model to obtain a policy matching score of the preprocessed data; fine-tuning the first language model based on a score label corresponding to the preprocessed data and the policy matching score, and obtaining a sample generation model after fine-tuning is completed; generating sample service data and sample labels by the sample generation model, and training a second language model according to the generated sample service data and sample labels to obtain a scoring model after training is completed.
2. The model training method of claim 1, wherein the first language model is configured with a policy knowledge base; the policy knowledge base contains a plurality of service policies of the credit service, and the service policies are used for service recommendation to service users of the credit service.
3. The model training method of claim 2, wherein the prompt text is used to prompt the first language model to score a matching relationship between preprocessed data and each service policy according to a matching score dimension contained in the prompt text to obtain a sub-score; and the prompt text is also used to prompt the first language model to perform weighted calculation on the sub-score of each matching score dimension to obtain a weighted score.
4. The model training method of claim 3, wherein the matching score dimension comprises at least one of the following: a data expression dimension for evaluating data density, redundancy and / or significance of preprocessed data, a policy matching dimension for evaluating a matching degree between preprocessed data and a service policy, and a service value dimension for evaluating a predicted service value of a service policy.
5. The model training method of claim 1, wherein fine-tuning the first language model based on the score label corresponding to the preprocessed data and the policy matching score comprises: calculating a reward score of each policy matching score of the preprocessed data by a reward function; calculating a loss according to the reward score and the corresponding score label, and adjusting parameters of the first language model according to the loss.
6. The model training method of claim 1, wherein the second language model comprises a plurality of matching score networks corresponding to matching score dimensions and a score weighting module; wherein the matching score network is used to score a matching relationship of a corresponding matching score dimension of input service recommendation data to obtain a sub-score; and the score weighting module is used to weight the sub-score of each matching score network to obtain a weighted score.
7. The model training method of claim 1, wherein training the second language model according to the generated sample service data and sample labels comprises: inputting the sample service data and prompt text into the second language model to perform policy matching score processing, and obtaining a service policy matched by the sample service data and a policy matching score; performing loss calculation based on the sample label and the policy matching score, and adjusting parameters of the second language model according to a training loss obtained by calculation.
8. The model training method of claim 1, wherein the credit service comprises a credit payment service for order payment based on a credit limit. The service policy comprises a limit increase policy for increasing the credit limit and / or a recommendation policy for recommending the payment service to the service user for order payment.
9. The model training method of claim 8, wherein the pre-processing of the service recommendation data of the credit service to obtain pre-processed data comprises: performing pre-set field filling of the credit service on the service recommendation data to obtain filled service data; performing data verification on the filled service data, and taking the filled service data that passes the verification as the pre-processed data. The service recommendation data is obtained by a data collection model through lead data acquisition of the credit service.
10. The model training method of claim 1, further comprising: obtaining a score result of the scoring model after being deployed to the credit service for policy matching score processing; performing parameter optimization on the scoring model according to the score result.
11. The model training method of claim 1, wherein the scoring model is trained in an ensemble training manner on the second language model.
12. A service recommendation processing method, comprising: pre-processing service recommendation data of a credit service to obtain pre-processed data; inputting the pre-processed data and a prompt text into a scoring model for matching score processing with a service policy to obtain a policy matching score; performing service recommendation processing on a service user corresponding to the service recommendation data based on the policy matching score; The scoring model is obtained by training a second language model based on sample service data and sample labels generated by a sample generation model, and the sample generation model is obtained by fine-tuning a first language model in a reinforcement learning manner.
13. A model training apparatus, comprising: a pre-processing module configured to pre-process service recommendation data of a credit service to obtain pre-processed data; a matching processing module configured to input the pre-processed data and a prompt text into a first language model for matching processing of the pre-processed data and a service policy to obtain a policy matching score of the pre-processed data; a model fine-tuning module configured to fine-tune the first language model based on a score label corresponding to the pre-processed data and the policy matching score, and obtain a sample generation model after the fine-tuning is completed; a model training module configured to generate sample service data and sample labels by the sample generation model, and train a second language model according to the generated sample service data and sample labels to obtain a scoring model after the training is completed.
14. A service recommendation processing apparatus, comprising: a pre-processing module configured to pre-process service recommendation data of a credit service to obtain pre-processed data; a matching score module configured to input the pre-processed data and a prompt text into a scoring model for matching score processing with a service policy to obtain a policy matching score; a service recommendation module configured to perform service recommendation processing on a service user corresponding to the service recommendation data based on the policy matching score. The scoring model is obtained by training a second language model based on sample service data and sample labels generated by a sample generation model, and the sample generation model is obtained by fine-tuning a first language model in a reinforcement learning manner.
15. A model training device comprising: a processor; and a memory configured to store computer executable instructions that, when executed, cause the processor to: preprocess service recommendation data of a credit service to obtain preprocessed data; input the preprocessed data and prompt text into a first language model to perform matching processing of the preprocessed data and service policies, and obtain a policy matching score of the preprocessed data; fine-tune the first language model based on a scoring label corresponding to the preprocessed data and the policy matching score, and obtain a sample generation model after fine-tuning is completed; generate sample service data and sample labels through the sample generation model, and train a second language model based on the generated sample service data and sample labels to obtain a scoring model after training is completed.
16. A service recommendation processing device comprising: a processor; and a memory configured to store computer executable instructions that, when executed, cause the processor to: preprocess service recommendation data of a credit service to obtain preprocessed data; input the preprocessed data and prompt text into a scoring model to perform matching score processing with service policies, and obtain a policy matching score; perform service recommendation processing on a service user corresponding to the service recommendation data based on the policy matching score; wherein the scoring model is obtained by training a second language model based on sample service data and sample labels generated by a sample generation model, and the sample generation model is obtained by fine-tuning a first language model in a reinforcement learning manner.
17. A computer readable storage medium for storing computer executable instructions that, when executed, implement the steps of the method of claim 1 or 12.