Sample enhancement-based credit review settlement optimization method
By using high-quality sample data and knowledge graph technology in the credit review summary generation, the big model is trained, and the traditional credit review summary generation method is solved, and the content is not personalized and the big model lacks domain knowledge is achieved, and high-quality, personalized and credible credit review summary generation is achieved.
Patent Information
- Application Number
- CN202510291384.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
The traditional method of generating a summary of the credit review relies on fixed templates and cannot dynamically adjust the content, resulting in insufficient personalization and accuracy, and may cause information loss and redundancy; the large model lacks domain knowledge when generating the summary of the credit review, which is prone to "illusions" and affects credibility.
By collecting the speech of the credential review dialogue into text, inputting the trained big model to generate a credential review summary, and using high-quality sample data to train the big model, filtering out sample data consistent with the credential review dialogue text, using knowledge graph technology to filter low-quality samples, and improving the model's performance ability in the credential review field.
It improves the quality and efficiency of the credit review summary, reduces redundant information, enhances the personalization and accuracy of the summary, reduces the risk of hallucination in large models, and improves the credibility of the credit review summary.
Smart Images

Figure CN120218331A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of credit review, and provides a method for optimizing credit review summaries based on sample enhancement. Background Art
[0002] In the field of credit approval, a credit review summary is a summary by a credit reviewer of the customer's application materials and the results of telephone verification, and is an important basis for credit approval decisions. However, there are many technical problems in traditional credit review summary generation methods, making it difficult for the quality and efficiency of the summary to meet actual needs.
[0003] Traditional templatized methods rely on pre-set fixed templates and cannot dynamically adjust the content according to the specific situation of the customer, resulting in insufficient personalization and accuracy of the credit review summary, and there may also be problems of loss of relevant information. In addition, the template may contain a large amount of redundant information, making the credit review summary long and inaccurate.
[0004] When generating a credit review summary based on a general large model, the general large model often lacks in-depth understanding of the knowledge in the corresponding field when generating text, and the model is prone to "hallucinations" during the generation process, that is, generating inaccurate or irrelevant content, affecting the credibility of the credit review summary. Summary of the Invention
[0005] In view of this, the present application provides a method for optimizing credit review summaries based on sample enhancement, aiming to improve at least one of the above problems.
[0006] Specifically, it includes the following technical solutions:
[0007] On the one hand, an embodiment of the present application provides a method for optimizing credit review summaries based on sample enhancement, and the method is as follows:
[0008] Collect the credit review dialogue voices during the credit review process, convert the credit review dialogue voices into credit review dialogue texts, input the credit review dialogue texts into a trained large model, and the large model outputs a credit review summary corresponding to the credit review dialogue text;
[0009] The large model is trained with high-quality sample data, and the high-quality sample data is sample data where the credit review summary and the credit review dialogue text expressions tend to be consistent.
[0010] In some embodiments of the present invention, the credit review dialogue voices include: the first dialogue voice between the credit review robot and the loan application user and the second dialogue voice between the credit reviewer and the loan application user;
[0011] The first dialogue voice is that the credit review robot sequentially issues questions based on a set question list, and the loan application user answers the robot's questions in sequence;
[0012] The second dialogue voice is the question given by the credit reviewer based on the data of the first dialogue, and the loan application user answers the questions of the credit reviewer in turn.
[0013] In some embodiments of the present invention, the screening process of high-quality sample data is specifically as follows:
[0014] Construct a knowledge graph of the loan application user based on the credit review dialogue data corresponding to the sample features as the standard knowledge graph; construct a knowledge graph of the loan application user based on the credit review summary corresponding to the sample label, and calculate the similarity between this knowledge graph and the corresponding standard knowledge graph; use the sample data with high similarity as high-quality sample data;
[0015] Among them, the knowledge graph includes the identity identification of the loan application user and the credit review data of the loan application user.
[0016] In some embodiments of the present invention, the credit review data of the loan application user includes: first credit review data and second credit review data. The first credit review data is the credit review data that must be included in the knowledge graph, and the second credit review data is the credit review data that may be included in the knowledge graph.
[0017] In some embodiments of the present invention, the first credit review data includes: credit investigation data and work data. The credit investigation data includes: loan data and credit data. The work data includes: income data, work unit name and work years. The second credit review data includes: vehicle data and basic family information.
[0018] In some embodiments of the present invention, the process of constructing the knowledge graph of the loan application user is specifically as follows:
[0019] Input the credit review dialogue text corresponding to the sample features into the trained knowledge graph model TransE, and the knowledge graph model TransE outputs the standard knowledge graph of the loan application user;
[0020] Input the credit review summary corresponding to the sample label into the trained knowledge graph model TransE, and the knowledge graph model TransE outputs the knowledge graph of the loan application user.
[0021] In some embodiments of the present invention, the loss of the knowledge graph model TransE The calculation formula is specifically as follows:
[0022]
[0023] Among them, represents the hinge loss, represents the high-order contrast loss, represents the graph smoothing constraint loss, and α, λ, γ respectively represent the hinge loss high-order contrast loss Graph smoothing constraint loss , where α+λ+γ=1.
[0024] In some embodiments of the present invention, the high-order contrast loss calculation formula is as follows:
[0025]
[0026] Among them, V represents the node set consisting of all nodes in the knowledge graph, and the k-hop neighbor set N of node v is generated by random walk or graph diffusion. k (v), put the nodes within k hops of node v into the set N k (v), k is 2, set N k The nodes in (v) constitute the positive samples n of node v v + , the remaining nodes constitute the negative samples n of node v v - , τ is the adjustment coefficient.
[0027] In some embodiments of the present invention, the graph smoothing constraint loss The calculation formula is as follows:
[0028]
[0029] Where u is a node within one hop range of node v.
[0030] In some embodiments of the present invention, the large model adopts the Qwen2-14B model.
[0031] The present invention filters the sample data to select high-quality samples that are consistent with the knowledge graph expression of the credit review dialogue text, mines the implicit information in the data through the knowledge graph technology, filters out the low-quality samples that may cause knowledge hallucinations, and trains the large model with the high-quality samples to improve the performance of the large model in the credit review field and greatly reduce the risk of hallucinations caused by the large model. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0033] Figure 1 A flow chart of a credit review summary optimization method based on sample enhancement provided in an embodiment of the present invention;
[0034] Figure 2 A schematic diagram of a knowledge graph of a loan application user provided by an embodiment of the present invention;
[0035] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0036] The technical scheme in the embodiment of the present application will be clearly and completely described below in conjunction with the drawings in the embodiment of the present application. Obviously, the described embodiment is a part of the embodiment of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application. Unless otherwise defined, all technical terms used in the embodiment of the present application have the same meaning as those generally understood by ordinary technicians in the field.
[0037] Figure 1 A flow chart of a method for optimizing a credit review summary based on sample enhancement provided by an embodiment of the present invention, the method is specifically as follows:
[0038] Collect the voice of the credit review dialogue during the credit review process, convert the voice of the credit review dialogue into credit review dialogue text, input the credit review dialogue text into the trained big model, and the big model outputs the credit review summary corresponding to the credit review dialogue text.
[0039] In an embodiment of the present invention, the credit review dialogue voice includes: a first dialogue voice between a credit review robot and a loan applicant and a second dialogue voice between a credit review officer and a loan applicant. The first dialogue voice is the robot asking questions in sequence based on a set question list, and the loan applicant answers the questions of the credit review robot in sequence. The second dialogue voice is the questions asked by the credit review officer based on the first dialogue voice, and the loan applicant answers the questions of the credit review officer in sequence.
[0040] When selecting the large model, we comprehensively considered factors such as the model's scale, performance, and training cost. We ultimately selected Qwen2-14B as the base model, used the selected high-quality samples to train the large model, and updated the model parameters through back propagation.
[0041] In order to improve the quality of the credit review summary output by the big model, the present invention screens the collected sample data and uses high-quality sample data to train the big model. The high-quality sample data is sample data whose credit review summary is consistent with the text expression of the credit review dialogue. After training the big model with the high-quality sample data, the trained big model outputs a credit review summary that is consistent with the text expression of the credit review dialogue.
[0042] In an embodiment of the present invention, credit review dialogue data is collected. The credit review dialogue data can be credit review dialogue voice or credit review dialogue text. If it is credit review dialogue voice, it is converted into credit review dialogue text. At the same time, a credit review summary formed based on the credit review dialogue text is obtained. The credit review dialogue text is used as a feature, and the credit review summary corresponding to the credit review dialogue text is used as a label to form sample data. The credit review summary in the current sample data may be manually formed by a credit reviewer or formed by other existing methods. There may be redundant information or unobjective information in the credit review summary. Therefore, the present invention aims to screen the sample data to obtain high-quality sample data.
[0043] In an embodiment of the present invention, the screening process of the sample data is specifically as follows:
[0044] A knowledge graph of the loan application user is constructed based on the credit review dialogue text corresponding to the sample feature and used as a standard knowledge graph; a knowledge graph of the loan application user is constructed based on the credit review summary corresponding to the sample label, and the similarity between this knowledge graph and the corresponding standard knowledge graph is calculated; the sample data with high similarity is used as high-quality sample data to participate in the training of the large model. Among them, the knowledge graph includes the identity identifier of the loan application user and the credit review data of the loan application user.
[0045] In an embodiment of the present invention, the credit review data of the loan application user includes: first credit review data and second credit review data. The first credit review data is the credit review data that must be included in the knowledge graph, and the second credit review data is the credit review data that may be included in the knowledge graph. The first credit review data includes: credit investigation data and work data. The credit investigation data includes: loan data and credit data. The work data includes: income data, work unit name, and work years. The second credit review data includes: vehicle data and basic family information. The vehicle data includes: basic vehicle information and vehicle use. The basic family information includes: marital information and other family member information, as Figure 2 shown.
[0046] In an embodiment of the present invention, the process of constructing the knowledge graph of the loan application user is specifically as follows:
[0047] The credit review dialogue text corresponding to the sample feature is input into the trained knowledge graph model TransE, and the knowledge graph model TransE outputs the standard knowledge graph of the loan application user. The credit review summary corresponding to the sample label is input into the trained knowledge graph model TransE, and the knowledge graph model TransE outputs the standard knowledge graph of the loan application user;
[0048] In the embodiment of the present invention, before the knowledge graph model TransE is used, it needs to be trained with samples, including: credit review dialogue texts and annotated knowledge graphs, as well as credit review summaries and their annotated knowledge graphs. During the training process, the loss value of the knowledge graph model TransE is calculated. When the loss value is lower than the set loss threshold, the training stops, and the knowledge graph model TransE is trained. Among them, the calculation formula of the loss value is specifically as follows:
[0049]
[0050] Among them, represents the hinge loss, represents the high-order contrast loss, represents the graph smoothing constraint loss, and α, λ, and γ respectively represent the weights of the hinge loss high-order contrast loss graph smoothing constraint loss where α + λ + γ = 1.
[0051] In the embodiment of the present invention, the calculation formula of the high-order contrast loss is specifically as follows:
[0052]
[0053] Among them, V represents the set of nodes composed of all nodes in the knowledge graph. The knowledge graph is a tree structure, the identity information of the loan application user is the root node, the data types of the first credit review data and the second credit review data are the child nodes of the root node, which are two nodes, and the data information corresponding to the first credit review data and the second credit review data is the child node of the secondary node, that is, the tertiary node. The k-hop neighbor set N k (v) of node v is generated by random walk or graph diffusion, and the nodes within the k-hop range of node v are put into the set N k (v). k takes the value of 2, and the nodes in the set N k (v) form the positive sample n v + of node v, and the remaining nodes form the negative sample n v - of node v, and τ is the adjustment coefficient.
[0054] In the embodiment of the present invention, the calculation formula of the graph smoothing constraint loss is specifically as follows:
[0055]
[0056] Among them, u is the node within the one-hop range of node v.
[0057] The present invention filters the sample data to select high-quality samples that are consistent with the knowledge graph expression of the credit review dialogue text, mines the implicit information in the data through the knowledge graph technology, filters out the low-quality samples that may cause knowledge hallucinations, and trains the large model with the high-quality samples to improve the performance of the large model in the credit review field and greatly reduce the risk of hallucinations caused by the large model.
[0058] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the present application disclosed herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only.
[0059] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A credit review summary optimization method based on sample enhancement, characterized in that: The method is specifically as follows: Collect the voice of the credit review dialogue during the credit review process, convert the voice of the credit review dialogue into credit review dialogue text, input the credit review dialogue text into the trained big model, and the big model outputs the credit review summary corresponding to the credit review dialogue text; The large model is trained with high-quality sample data, which refers to sample data whose expression in the credit review summary is consistent with that in the credit review dialogue text.
2. The method for optimizing the credit review summary based on sample enhancement as claimed in claim 1, characterized in that: The credit review dialogue voice includes: the first dialogue voice between the credit review robot and the loan application user and the second dialogue voice between the credit reviewer and the loan application user; The first conversation is that the credit review robot asks questions based on a set list of questions, and the loan applicant answers the robot's questions in turn; The second conversation voice is questions asked by the credit reviewer based on the data from the first conversation, and the loan applicant answers the credit reviewer's questions in turn.
3. The method for optimizing the credit review summary based on sample enhancement as claimed in claim 1, characterized in that: The screening process of high-quality sample data is as follows: Based on the credit review dialogue data corresponding to the sample features, a knowledge graph of loan applicants is constructed and used as the standard knowledge graph; based on the credit review substructure corresponding to the sample labels, a knowledge graph of loan applicants is constructed and the similarity between the knowledge graph and the corresponding standard knowledge graph is calculated; sample data with high similarity is used as high-quality sample data; Among them, the knowledge graph includes the identity identification of the loan applicant and the credit review data of the loan applicant.
4. The method for optimizing the credit review summary based on sample enhancement as claimed in claim 3, characterized in that: The credit review data of the loan applicant includes: first credit review data and second credit review data. The first credit review data is the credit review data that must be included in the knowledge graph, and the second credit review data is the credit review data that may be included in the knowledge graph.
5. The method for optimizing the credit review summary based on sample enhancement as claimed in claim 4, characterized in that: The first credit review data includes: credit investigation data and work data. The credit investigation data includes: loan data and credit data. The work data includes: income data, name of the work unit and years of work. The second credit review data includes: vehicle data and basic family information.
6. The method for optimizing the credit review summary based on sample enhancement as claimed in claim 3, characterized in that: The knowledge graph construction process of loan application users is as follows: The credit review dialogue text corresponding to the sample features is input into the trained knowledge graph model TransE, and the knowledge graph model TransE outputs the standard knowledge graph of the loan application user; The credit review summary corresponding to the sample label is input into the trained knowledge graph model TransE, and the knowledge graph model TransE outputs the knowledge graph of the loan application user.
7. The method for optimizing the credit review summary based on sample enhancement as claimed in claim 6, characterized in that: The loss of the knowledge graph model TransE The calculation formula is as follows: in, Indicates hinge loss, represents the high-order contrast loss, represents the graph smoothing constraint loss, α, λ, and γ represent the hinge loss respectively. High-order contrast loss Graph smoothing constraint loss The weight of , where α+λ+γ=1.
8. The method for optimizing the credit review summary based on sample enhancement as claimed in claim 7, characterized in that: The high-order contrast loss calculation formula is as follows: Among them, V represents the node set consisting of all nodes in the knowledge graph, and the k-hop neighbor set N of node v is generated by random walk or graph diffusion. k (v), put the nodes within k hops of node v into the set N k (v), k is 2, set N k The nodes in (v) constitute the positive samples n of node v v + , the remaining nodes constitute the negative samples n of node v v - , τ is the adjustment coefficient.
9. The method for optimizing the credit review summary based on sample enhancement as claimed in claim 8, characterized in that: Graph smoothing constraint loss The calculation formula is as follows: Where u is a node within one hop range of node v.
10. The method for optimizing the credit review summary based on sample enhancement according to claim 1, characterized in that: The large model uses the Qwen2-14B model.