Self-learning method of credible dialogue model

By constructing seed samples and conducting multiple rounds of quality assessment, the self-optimizing credit review dialogue model solves the domain adaptation problem of large language models in financial credit review scenarios, and improves the model's logical reasoning ability and decision transparency.

CN120806128APending Publication Date: 2025-10-17CHERY HUIYIN MOTOR FINANCE SERVICE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510843457.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In the context of financial credit review, existing large language models face domain adaptation issues when applied in vertical fields, especially given the high requirements for data quality and scale, and the high cost of manual annotation, making it difficult to achieve data scaling.

Method used

By constructing seed samples, using a credit-based dialogue model to generate prediction questions and reasoning processes, and combining a large language model and a validator model for multiple rounds of quality assessment, high-quality sample data is selected for self-optimization training, gradually improving logical reasoning ability.

Benefits of technology

The credit review dialogue model has achieved progressive self-optimization, which has improved logical reasoning ability and decision-making transparency, reduced the black box phenomenon, and improved the applicability and efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806128A_ABST
    Figure CN120806128A_ABST
Patent Text Reader

Abstract

The invention discloses a self-learning method of a credit review dialogue model, and the method comprises the steps: constructing a seed sample which is used for marking a reasoning process and dialogue data of a question corresponding to the reasoning process; inputting the dialogue data in the seed sample into a current credit examination dialogue model, outputting a prediction problem and a reasoning process corresponding to a prediction question by the credit examination dialogue model, and forming a series of sample data by the dialogue data, the reasoning process and the corresponding prediction question; evaluating the quality of currently generated sample data, screening out high-quality sample data as seed samples, and putting the seed samples into a seed sample database; and training the trust dialogue model based on the seed samples in the seed sample database. By iterating a small amount of dialogue data of questions marked with the reasoning process, progressive self-optimization of the credit review dialogue model is realized, and the logical reasoning ability of the credit review dialogue model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of credit investigation, and more particularly, relates to a self-learning method of a credit investigation dialogue model. BACKGROUND

[0002] With the rapid development of artificial intelligence, it is reshaping the global socio-economic pattern at an astonishing speed and becoming the core engine driving the progress of human civilization. In key fields such as finance, manufacturing, and medicine, artificial intelligence is upgrading from single-point technology application to systematic productivity revolution. For the financial scene, it has become a key hub connecting financial institutions and the real economy, and is experiencing a paradigm shift from "experience-dependent decision-making" to "digital intelligent governance".

[0003] As an actual and effective application scenario of artificial intelligence technology in the financial field, the credit investigation dialogue model embodies its significant advantages in task automation processing. The credit investigation dialogue model can replace traditional artificial customer service and complete the credit investigation process through interactive dialogue with users, greatly saving manpower and improving work efficiency.

[0004] As a comprehensive representative of artificial intelligence development, large language models exhibit unparalleled technical advantages and have broad application scenarios in intelligent dialogue systems. Such models rely on massive multi-source text data for unsupervised learning, with data sources covering news information, academic literature, social media texts, literary works, and other diverse carriers, enabling large language models to have strong language understanding capabilities and accurately grasp the meanings behind various natural language expressions. At the same time, in terms of intelligent dialogue, large language models integrate natural language understanding and dialogue generation capabilities, and can generate logically consistent, semantically smooth, and scenario-adapted reply content based on given contexts and user inputs, effectively expanding the functional boundaries and application depth of intelligent dialogue systems.

[0005] Although current large language models have shown strong dialogue capabilities in general fields, in vertical fields such as financial credit review, direct application of general models has significant domain adaptation problems due to strong domain constraints of professional terms and compliance specifications of dialogue logic. Fine-tuning training of specific vertical scene data based on pre-trained general large language models can effectively solve this problem, but it puts strict requirements on the quality, size and diversity of training data. Existing training data construction mainly relies on manual annotation and rule template matching to clean, filter and annotate data. Through the construction of a data processing rule library, data outliers are filtered and semantic annotations are realized, and the quality of data is further improved relying on manual annotation. Through template processing, the construction efficiency of structured data is improved, but there are two key problems, one is that the rule coverage is not complete in complex semantic scenarios, and the second is that there is no corresponding reasoning basis data in the database. If manual filtering and writing are used, the cost is very high, and it is difficult to realize data scaling. SUMMARY

[0006] The present application provides a self-learning method of a credit review dialogue model, aiming to improve at least one of the above problems.

[0007] The present application is implemented as follows: a self-learning method of a credit review dialogue model, the method comprising:

[0008] (1) constructing seed samples, the seed samples being dialogue data annotated with reasoning processes and corresponding questions in the reasoning processes;

[0009] (2) inputting the dialogue data in the seed samples into the current credit review dialogue model, the credit review dialogue model outputting predicted questions and corresponding reasoning processes of the predicted questions, and the dialogue data, the reasoning processes and the corresponding predicted questions forming a series of sample data;

[0010] (3) evaluating the quality of the generated sample data, selecting high-quality sample data as seed samples and putting them into a seed sample database;

[0011] (4) training the credit review dialogue model based on the seed samples in the seed sample database.

[0012] Further, the quality evaluation process of the sample data is as follows:

[0013] (31) first round of quality evaluation: detecting whether the reasoning processes and predicted questions in the sample data conform to the standard format, and retaining the sample data conforming to the standard format;

[0014] (32) second round of quality evaluation: calculating the cosine similarity between the predicted questions in the sample data and the annotated questions in the corresponding seed samples, and retaining the sample data corresponding to the predicted questions with high similarity;

[0015] (33) Third round of quality assessment: input sample data into a large language model, and the large language model discriminates whether the sample meets the evaluation standard.

[0016] Further, after each update of the seed sample database, the training of the credit review dialogue model is performed based on the updated seed sample.

[0017] Further, the credit review dialogue model training process is as follows:

[0018] The seed sample is divided into test samples and training samples. The credit review dialogue model is trained based on the training samples, and the credit review dialogue model is verified based on the test samples. If the prediction accuracy of the credit review dialogue model reaches the set standard, the training of the credit review dialogue model is completed. If the prediction accuracy of the credit review dialogue model is lower than the set standard, return to step (1).

[0019] Further, a plurality of sample data generated based on the same dialogue data are divided into high-quality and low-quality sample data. The dialogue data, the predicted questions of the high-quality samples based on the dialogue data, and the predicted questions of the low-quality samples form a sample. The verifier model is trained.

[0020] Further, the verifier model adopts a Qwen2.5 model, a Llama3.1 model, or a DeepSeek model.

[0021] Further, after the training of the credit review dialogue model is completed, the dialogue data is input into the trained credit review dialogue model. The credit review dialogue model outputs the predicted questions corresponding to the dialogue data and the reasoning process of forming the predicted questions. The output predicted questions are input into the verifier model. The verifier model outputs the quality evaluation of each predicted question. The predicted question with the best quality evaluation is output.

[0022] Further, the credit review dialogue model outputs a plurality of predicted questions corresponding to a dialogue data and the reasoning process corresponding to the predicted questions.

[0023] Further, the credit review dialogue model adopts a Qwen2.5 model, a Llama3.1 model, or a DeepSeek model.

[0024] The present application gradually optimizes the credit review dialogue model through iterative labeling of a small amount of dialogue data with reasoning process questions, and improves the logical reasoning ability of the credit review dialogue model. At the same time, the credit review dialogue model can generate detailed logical chains of reasoning processes, enhance the explainability and transparency of decision-making, and effectively reduce the black-box technology defects of traditional large language models. BRIEF DESCRIPTION OF DRAWINGS

[0025] Figure 1 The self-learning method flow chart of the credit review dialogue model provided by the embodiment of the present application. DETAILED DESCRIPTION

[0026] The specific implementation methods of the present invention will be further explained in detail below by describing the embodiments with reference to the accompanying drawings, so as to help those skilled in the art to have a more complete, accurate and in-depth understanding of the inventive concept and technical solution of the present invention.

[0027] Figure 1 This is a flow chart of a self-learning method for a credit review dialogue model provided by an embodiment of the present invention. The method is specifically as follows:

[0028] (1) Constructing seed samples, which are dialogue data with labeled reasoning processes and questions corresponding to the reasoning processes;

[0029] In an embodiment of the present invention, the conversation data between the credit reviewer and the credit review user is collected, and the conversation data is annotated. The questions in the conversation data and the reasoning process of forming the questions are annotated to form a seed sample. Since the annotation of the seed sample is relatively time-consuming and labor-intensive, only a small amount of seed samples can be formed.

[0030] (2) Input the dialogue data in the seed sample into the current credit review dialogue model, which outputs the predicted question and the reasoning process corresponding to the predicted question. The dialogue data, reasoning process, and corresponding predicted question form a series of sample data.

[0031] In the embodiments of the present invention, the credit review dialogue model uses the Qwen 2.5 model, the Llama 3.1 model, or the DeepSeek model. The above models are fine-tuned using seed samples to make them suitable for the credit review dialogue field, thereby forming a credit review dialogue model. The number of seed samples generated by manual annotation is relatively small. The present invention uses the current credit review dialogue model to expand sample data. The expansion process is as follows:

[0032] By inputting the dialogue data in the seed sample into the current credit review dialogue model, the credit review dialogue model will generate a new dialogue data x for each input dialogue data x. i Output several predicted questions y i1 、y i2 ,…,y in And the reasoning process T corresponding to each prediction problem i1 、T i2 ,…,T in , the conversation data x i , the corresponding prediction question y i1 、y i2 ,…,y in and the reasoning process that forms the predictive question T i1 、T i2 ,…,T in Composed of new sample data (xi , y i1 , T i1 ), (x i , y i2 , T i2 ), …, (x i , y in , T in ) are generated by the above process.

[0033] (3) Evaluate the quality of the currently generated sample data, select high-quality sample data as seed samples, and put them into the seed sample database;

[0034] Because the number of seed samples is relatively small, based on the seed samples to expand the sample data, the quality of the newly generated sample data is uneven, in order to ensure the quality of the sample data, the quality of the sample data generated based on the current credit review dialogue model needs to be selected, and high-quality sample data is selected as seed sample and put into the seed sample database, realizing the automatic expansion of the seed sample.

[0035] In the embodiment of the application, the quality evaluation process of the sample data is as follows:

[0036] (31) First round of quality evaluation: detect whether the reasoning process and the prediction question in the sample data conform to the standard format, retain the sample data conforming to the standard format, and perform the second round of quality evaluation, enter step (32);

[0037] In the embodiment of the application, in order to guide the credit review dialogue model to output the prediction question in the standard format and the reasoning process corresponding to the prediction question, the reasoning process and the question corresponding to the reasoning process in the seed sample are labeled in the set standard format, and the reasoning process and the question in the standard format are more conducive and more accurate for the next round of sample data quality evaluation, the application adopts a code matching mechanism based on regular expressions, and performs full string matching through a pre-set regular expression engine to verify whether it conforms to the set template structure.

[0038] (32) Second round of quality evaluation: calculate the cosine similarity between the prediction question in the sample data and the labeled question corresponding to the seed sample, retain the sample data corresponding to the prediction question with high similarity, and execute step (33);

[0039] In the embodiment of the application, because the prediction question in the sample data is obtained by predicting the dialogue data in the seed sample, the similarity between the prediction question in the current sample and the labeled question corresponding to the seed sample is calculated, the higher the similarity of the labeled question, the better the quality of the prediction question, therefore, the sample data with a similarity greater than a similarity threshold is retained, and the third round of quality evaluation is performed, and step (33) is entered;

[0040] In the embodiment of the present application, before calculating the cosine similarity between the predicted question in the sample data and the labeled question of the corresponding seed sample, the length of the predicted question needs to be filtered, and the sample data in the question process is discarded.

[0041] (33) The third round of quality assessment: input the sample data into the large language model, and the large language model discriminant module outputs whether the sample meets the evaluation standard.

[0042] In the embodiment of the present application, the large language model adopts an existing general large language model, such as the Qwen2.5 model, the Llama3.1 model, or the DeepSeek model, and the sample data with good logical coherence is screened out through the large language model.

[0043] (4) Training the credit review dialogue model based on the seed samples in the seed sample database.

[0044] In the embodiment of the present application, after generating new seed samples based on steps (1) to (3), the generated seed samples are put into the seed sample database. After each update of the seed sample database, the credit review dialogue model is trained based on the updated seed samples, and the training process is as follows:

[0045] The seed samples are divided into test samples and training samples. The credit review dialogue model is trained based on the training samples, and the credit review dialogue model is verified based on the test samples. If the prediction accuracy of the credit review dialogue model reaches the set standard, the training of the credit review dialogue model is completed. If the prediction accuracy of the credit review dialogue model is lower than the set standard, return to step (1).

[0046] Through step (3), the several sample data generated based on the same dialogue data are divided into high-quality and low-quality sample data. The dialogue data, the predicted question of the high-quality sample based on the dialogue data, and the predicted question of the low-quality sample form a sample. The verifier model is trained based on the above-mentioned sample. The verifier model adopts the Qwen2.5 model, the Llama3.1 model, or the DeepSeek model. After the above-mentioned model is trained through the sample, the verifier model has the function of evaluating the quality of the predicted question.

[0047] In the embodiment of the present application, after the training of the credit review dialogue model is completed, the dialogue data is input into the current credit review dialogue model, and the credit review dialogue model outputs the predicted question corresponding to the dialogue data and the reasoning process of forming the predicted question. The output predicted question is input into the verifier model, the verifier model outputs the quality evaluation of each predicted question, and the predicted question with the best quality evaluation is output and output to the credit review user.

[0048] The application realizes progressive self-optimization of the credit review dialogue model and improves the logical reasoning capability of the credit review dialogue model by iteratively labeling a small amount of dialogue data of the question with the reasoning process; meanwhile, the credit review dialogue model can generate a detailed logical chain of the reasoning process, enhance the decision explainability and transparency, and effectively reduce the black-box technology defects of the traditional large language model.

[0049] The application is described exemplarily, and it is obvious that the specific implementation of the application is not limited by the above manner, as long as various non-essential improvements are made by adopting the method concept and technical solution of the application, or the concept and technical solution of the application is directly applied to other occasions without improvement, which is within the protection scope of the application.

Claims

1. A self-learning method for a credit review dialogue model, characterized in that: The method comprises: (1) Constructing seed samples, which are dialogue data with labeled reasoning processes and questions corresponding to the reasoning processes; (2) Input the dialogue data in the seed sample into the current credit review dialogue model, which outputs the predicted question and the reasoning process corresponding to the predicted question. The dialogue data, reasoning process, and corresponding predicted question form a series of sample data. (3) Evaluate the quality of the currently generated sample data, select high-quality sample data as seed samples, and put them into the seed sample database; (4) Train the credit review dialogue model based on the seed samples in the seed sample database.

2. The self-learning method of the credit review dialogue model according to claim 1, characterized in that: The quality assessment process of sample data is as follows: (31) The first round of quality assessment: Check whether the reasoning process and prediction questions in the sample data conform to the standard format, and retain the sample data that conform to the standard format; (32) Second round of quality assessment: Calculate the cosine similarity between the predicted questions in the sample data and the labeled questions of the corresponding seed samples, and retain the sample data corresponding to the predicted questions with high similarity; (33) The third round of quality assessment: The sample data is input into the large language model, and the large language model discriminant module outputs whether the sample meets the assessment criteria.

3. The self-learning method of the credit review dialogue model according to claim 1, characterized in that: After each update of the seed sample database, the credit review dialogue model is trained based on the updated seed samples.

4. The self-learning method of the credit review dialogue model according to claim 3, characterized in that: The training process of the credit review dialogue model is as follows: The seed samples are divided into test samples and training samples. The credit review dialogue model is trained based on the training samples and verified based on the test samples. If the prediction accuracy of the credit review dialogue model reaches the set standard, the training of the credit review dialogue model is completed. If the prediction accuracy of the credit review dialogue model is lower than the set standard, return to step (1).

5. The self-learning method of the credit review dialogue model according to claim 3 is characterized in that: Several sample data generated based on the same conversation data are divided into high-quality and low-quality sample data. The conversation data, the predicted questions of the high-quality samples formed based on the conversation data, and the predicted questions of the low-quality samples are composed of samples to train the verifier model.

6. The self-learning method of the credit review dialogue model according to claim 5, characterized in that: The verifier model uses the Qwen2.5 model, Llama3.1 model or DeepSeek model.

7. The self-learning method of the credit review dialogue model according to claim 5, characterized in that: After the training of the credit review dialogue model is completed, the dialogue data is input into the trained credit review dialogue model. The credit review dialogue model outputs the predicted questions corresponding to the dialogue data and the reasoning process for forming the predicted questions. The output predicted questions are input into the verifier model. The verifier model outputs the quality evaluation of each predicted question and outputs the predicted question with the best quality evaluation.

8. The self-learning method of the credit review dialogue model according to claim 1, characterized in that: The credit review dialogue model outputs multiple predicted questions corresponding to a dialogue data and the reasoning process corresponding to the predicted questions.

9. The self-learning method of the credit review dialogue model according to claim 5 is characterized in that: The credit review dialogue model adopts the Qwen2.5 model, Llama3.1 model or DeepSeek model.

Citation Information

Cited By

  • Training method and system of medical dialogue model

    CN121351970A