Automobile financial risk control question answering method based on social network enhanced retrieval
By constructing a knowledge graph and public opinion information graph that combine people and vehicles, and using graph retrieval enhancement technology, a large-scale question-and-answer model for auto finance risk control was created. This solved the problems of insufficient interpretability and poor timeliness of traditional auto finance risk control methods, and achieved efficient financial risk identification and improved question-and-answer accuracy.
Patent Information
- Application Number
- CN202510983833.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-10-31
AI Technical Summary
Traditional auto finance risk control methods rely on the experience of business experts and blacklist matching, which are insufficient in interpretability and timeliness. Insufficient understanding of the business by professional modelers also limits the effectiveness of risk identification and response.
The auto finance risk control question-and-answer method based on social network-enhanced retrieval constructs a knowledge graph and public opinion information graph that combine human and vehicle information, and uses graph retrieval enhancement technology to create a large-scale auto finance risk control question-and-answer model, achieving efficient multi-dimensional information retrieval and natural language question answering.
It improves the accuracy and robustness of financial risk control Q&A, enabling business personnel without professional modeling capabilities to effectively identify potential financial risks of loan applicants.
Smart Images

Figure CN120873138A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of automotive finance risk control technology, and in particular relates to an automotive finance risk control question-and-answer method based on social network enhanced retrieval. Background Technology
[0002] With the booming development of the automobile sales market, the auto finance industry has experienced rapid rise and expansion. However, this has also brought with it increasing and more complex financial risks. Effective control of these diverse risks has become a major challenge for the industry. Traditional risk control methods mainly rely on rules set by experts or on professional modelers to build predictive models using statistical methods, machine learning, and deep learning techniques. However, the effectiveness of this approach is highly dependent on the experience level of the business experts who formulate the rules and the depth of understanding of the business by the modelers. In reality, staff directly involved in the business often lack professional modeling skills, while professional modelers often do not have in-depth exposure to specific business operations. This disconnect leads to limitations in understanding and responding to financial risks. To address these issues, this invention proposes an auto finance risk control question-answering method based on social network-enhanced retrieval. This method constructs a social network based on combined human and vehicle data and uses graph retrieval enhancement to build a large-scale auto finance risk control question-answering model. This enables efficient retrieval of multi-dimensional information about applicants, effectively improving the accuracy and robustness of financial risk control questions and answers. Based on this large-scale financial risk control question-answering model, business staff without professional modeling capabilities can discover the financial risks of loan applicants through natural language question answering.
[0003] The existing technology still has the following shortcomings:
[0004] Traditional risk control strategies primarily rely on business experts to formulate risk control rules based on their own experience, or on identifying potential risky behaviors through historical blacklist matching. This approach, due to its dependence on past experience and blacklists, possesses a degree of interpretability. However, it is limited by the time frame of business experts' experience, and blacklists suffer from issues of timeliness and privacy protection, making it inadequate in addressing emerging financial risks. With advancements in artificial intelligence, machine learning and deep learning have been increasingly introduced into the field of financial risk control to more effectively identify potential risks. While these advanced methods can improve the timeliness of risk control, their effectiveness largely depends on the modelers' understanding of the business. In reality, professional modelers often do not directly participate in relevant business operations, resulting in insufficient depth of understanding. Conversely, those directly involved in relevant business operations are usually not professional modelers. This disconnect limits the effectiveness of financial risk identification and response.
[0005] To address the aforementioned issues, this invention proposes a social network-based augmented retrieval method for auto finance risk control question answering. This method constructs a social network by integrating relevant data on people and vehicles, and employs graph retrieval enhancement techniques to create a large-scale model for auto finance risk control question answering. This approach efficiently retrieves multi-dimensional information about applicants, significantly improving the accuracy and stability of financial risk control question answering. With this large-scale financial risk control question answering model, even business personnel without professional modeling capabilities can effectively identify potential financial risks of loan applicants through natural language question answering. Summary of the Invention
[0006] This invention provides a question-and-answer method for automotive finance risk control based on social network-enhanced retrieval, aiming to solve the above-mentioned problems.
[0007] This invention is implemented as follows: a question-and-answer method for auto finance risk control based on social network enhanced retrieval, comprising the following steps:
[0008] Step 1: Building a human-vehicle integrated social network, including the construction of a human-vehicle integrated knowledge graph and the construction of a knowledge graph based on public opinion information;
[0009] Step 2: Graph-guided retrieval based on social networks, including query enhancement, graph retrieval based on extended paths, and knowledge enhancement;
[0010] Step 3: Automotive finance risk control question answering based on a large model, including rewriting search results and generating answers based on knowledge enhancement.
[0011] Preferably, the knowledge graph construction combining people and vehicles includes: given a borrower P and the vehicle he / she purchases, constructing a basic relationship graph based on P's communication information, residential address, kinship, company information, etc., and constructing a person-vehicle association graph based on information such as the vehicle type and purchase channel of the vehicle he / she purchases.
[0012] Specifically, for borrower P's communication information, their communication records from the past three years are selected, based on the number of communications F. i c and communication duration Perform a weighted calculation to determine the association weight W between the borrower and their corresponding contact i. i r The calculation formula is as follows:
[0013]
[0014] Among them, F c and D c These represent the total number of communications and the total duration of communications for borrower P, respectively. and Represent the weights of communication frequency and communication duration respectively, and satisfy the following conditions: The relationship weights of all contacts of borrower P are calculated using the above formula, and then a communication relationship graph is constructed based on the relationship weights. The specific calculation formula is as follows:
[0015]
[0016] Among them, E i Indicates whether communicator i and borrower P are related (1 indicates related, 0 indicates not related), T c This represents the communication threshold (typically the mean of the borrower P's associated weights).
[0017] For borrower P's kinship and company information, relationships are typically calculated using close relatives within three generations, employees of the same company, or employees of the same department to construct kinship and company association graphs. The communication relationship graph, kinship graph, and company association graph are then merged (i.e., merging related individuals with the same ID number or mobile phone number) to ultimately form the basic relationship graph.
[0018] Regarding the type of vehicle and purchase channel information for the borrower's vehicle, firstly, statistics on borrower P... i The number of times fraud or deception has occurred in the past three months for the corresponding vehicle type. Number of times fraud or deception occurred with the vehicle's purchase channel Secondly, calculate borrower P i Other borrowers P j Vehicle type association weight The specific formula is as follows:
[0019]
[0020] Where, N VT This represents the total number of write-offs or frauds involving vehicle type VT. Then, the number of times borrower P... i Other borrowers P j The weight of the relationship between purchasing channels The specific formula is as follows:
[0021]
[0022] Where, N DC This represents the total number of times fraud or write-offs occurred at the vehicle purchase channel (DC). Finally, it assigns a weight to the association between vehicle types. Weight of the relationship with the purchase channel Weighted fusion is performed to obtain the final vehicle relationship weights, and a human-vehicle relationship graph is constructed by combining the weight thresholds. The calculation formula is as follows:
[0023]
[0024] in, and Represent the weights of vehicle type and purchase channel respectively, and satisfy the following conditions:
[0025] E i,j Indicates borrower P i With borrower P j Whether it is related (1 indicates related, 0 indicates not related), T V This represents the vehicle relationship threshold (usually the average of borrower P's vehicle relationships).
[0026] After completing the construction of the basic relationship graph and the human-vehicle association graph, the two association graphs are merged according to the borrower to finally construct a knowledge graph combining human and vehicle.
[0027] Preferably, the knowledge graph construction based on public opinion information includes discovering entities, attributes, and relationships contained in social network public opinion to construct triples in the form of <entity, relationship, entity>, including entity extraction, attribute extraction, and relationship extraction.
[0028] Preferably, the entity extraction specifically includes:
[0029] Step 1: Collect topic data from social network platforms using web crawling technology or API services;
[0030] Step 2: Preprocess the collected topic data, removing redundant information and sensitive information, etc.
[0031] Step 3: Use entity extraction techniques to obtain relevant entities from the processed topic data. There are two main ways to extract entities: rule-based and dictionary-based methods and machine learning-based methods.
[0032] Preferably, the attribute extraction specifically includes:
[0033] (1) Use web crawlers and API services to obtain relevant data and perform preprocessing;
[0034] (2) Extract attribute information from preprocessed data using attribute extraction techniques (including rule-based extraction methods and statistical extraction methods);
[0035] (3) Associate the extracted attribute information with the extracted entities.
[0036] Preferably, the relation extraction specifically includes:
[0037] Relevant data was obtained using web crawlers and API services, and the data was preprocessed.
[0038] The preprocessed data is analyzed to extract the relationships between entities (this patent uses template-based relationship extraction and machine learning-based relationship extraction).
[0039] Finally, all extracted relationships are associated with entities to obtain all entity relationship pairs in the public opinion knowledge graph.
[0040] Preferably, the query expansion specifically includes: terminology recognition, keyword extraction, context recognition, query terminology, and question enhancement. Specifically, terminology recognition involves identifying terms and abbreviations in the query question and generating accurate explanations for them. Keyword extraction involves extracting keywords from the query question and generating explanations and synonyms for them. Context recognition involves identifying the contextual information of terms, abbreviations, and keywords in the query question. Query terminology involves using the identified terms to query a financial terminology dictionary to obtain extended definitions, descriptions, and annotations. Question enhancement involves combining the original query question, keywords, contextual information, and detailed terminology definitions, and rewriting and modifying it using a large language model to form an enhanced query question, providing clear context and resolving ambiguities. The query decomposition method involves using query decomposition techniques to break down the original query question and the enhanced query question into smaller, more specific subqueries, ensuring that each subquery focuses on a specific aspect of the original query. Both the original and enhanced query questions are decomposed into clauses, each representing a specific association, and relevant triples are retrieved for each clause sequentially.
[0041] Preferably, the graph retrieval based on the extended path specifically includes:
[0042] For a given query, decompose the clauses and extract key topic entities from the sentences. Starting with the extracted topic entities, construct an expanded path by following a sequence prediction process. Assume the retrieval path at step t is p. t ,
[0043] p t =(r1,r2,...,r t )
[0044] Where, r i This represents the relation retrieved in step i. It is constructed by filling entities into the path relation to build a relation based on p. t Tree structure T t ,
[0045] T t =(e s ,r1,E1,r2,E2,...,r t E t )
[0046] Among them, e s E represents the subject entity of the sentence. i This indicates that relation r is retrieved through step i. i The set of related entities. The retrieval relationship in step t+1 starts from E. t The adjacency relationships are obtained by focusing on them, that is, by calculating the relationship r. t+1 Extended conditional probability p(r|p t To determine.
[0047] To calculate the conditional probability p(r|p t To determine the relevance between relation r and query clause q, this patent uses the cosine similarity between the embedding vectors of relation r and query clause q as the relevance calculation result sim(q,r).
[0048]
[0049] Here, `emb(·)` represents vector embedding. Since different retrieval relation time steps should focus on specific parts of the input query clause when expanding the relation, the original question is concatenated with historical expanded relations as input to update the question's embedding, i.e.
[0050] emb(q t )=f([q;T t ])
[0051] Among them, f is usually encoded using vector encoding methods such as Word2Vec and BERT.
[0052] Relation r t+1 Extended conditional probability p(r|q) t The calculation formula for ) is as follows:
[0053]
[0054] If p(r|q) t If the probability is greater than or equal to 0.5, then the relation with the highest probability is selected as the retrieval relation r for step t+1. t+1 If p(r|q) t If the probability is less than 0.5, path expansion stops. Based on the probability calculation formula for relation expansion mentioned above, the probability of the path corresponding to the query clause q is calculated using the joint distribution of all relations in the path.
[0055]
[0056] Here, n represents the number of relations in path p. Since the path with the highest probability cannot be guaranteed to be correct, a bundle search can be used to obtain multiple paths to construct a tree structure, which can then be used as the query result for the clause.
[0057] The multiple paths of the query are merged to generate the final subgraph of the clause query. Finally, after the expanded path retrieval calculations for all query question decomposition clauses are completed, graph retrieval results based on the expanded paths can be generated.
[0058] The optimization of search results includes, after obtaining the initial search results, using knowledge enhancement strategies to optimize and improve these results, specifically knowledge merging and knowledge pruning. Knowledge merging specifically involves: merging the retrieved results to achieve information compression and aggregation, helping to obtain a more comprehensive perspective by integrating relevant details from multiple sources; knowledge pruning specifically involves: filtering out irrelevant or redundant search information to optimize the results, calculating the similarity between the retrieved triples and the original query question and the enhanced query question; then, performing a comprehensive ranking based on the two similarity calculation results; finally, selecting the top n search results from the comprehensive ranking, concatenating them with their respective questions, and re-ranking them using the Sentence-BERT model, selecting the top k triples as the final search results.
[0059] Preferably, the rewriting of the search results specifically includes:
[0060] The core of rewriting search results is to convert structured triples into free-form text based on the KG-to-Text model. The specific process is as follows:
[0061] Given a knowledge graph G, the triples in the knowledge graph G are verbalized into ternary texts x by connecting head entities, relations, and tail entities. That is, the knowledge graph is transformed into a text set of "(head entity, relation, tail entity)".
[0062] The prompt word project constructs a knowledge graph-to-text (KG-to-Text) prompt word prompt1 based on the ternary form text x. The prompt word template is as follows:
[0063] "Your task is to transform the knowledge graph into one or more sentences."
[0064] The knowledge graph is: {ternary text x}.
[0065] The sentence is:
[0066] The prompt word prompt1 is input into a large model fine-tuned based on KG-to-Text labeled data for knowledge graph to text conversion to obtain free-format text y, thereby rewriting the search results.
[0067] When answering questions, each reasoning path is converted into a prompt word to obtain the corresponding free text from the input to the large model. The free text is then merged into the question-related knowledge.
[0068] Preferably, the knowledge-enhanced answer generation includes, in order to integrate the knowledge and questions generated based on KG-to-Text, this patent designs a large-model-based automotive finance risk control question-answering prompt word prompt2, the template of which is as follows:
[0069] "Your task is to answer the question based on the facts relevant to it."
[0070] Facts related to the question: {free-form text y}
[0071] Question: {Auto Finance Risk Control Issues Q}
[0072] Answer:".
[0073] By using the prompt word prompt2 template, the free-form text y and the auto finance risk control question Q are mapped to the knowledge graph-enhanced prompt word and input into the financial risk control big model to output the predicted answer.
[0074] Compared with the prior art, the embodiments of this application have the following main advantages:
[0075] This invention constructs a social network based on borrower information and embeds vehicle-related information into the social network to build a human-vehicle integrated social network. Secondly, it performs graph augmentation retrieval on the constructed human-vehicle integrated social network to output the social network information most closely related to the vehicle risk control issue. Finally, it combines the vehicle risk control issue with the retrieval results and inputs them into a large financial model for processing to generate the vehicle risk control issue result.
[0076] Compared with existing methods, this invention proposes a social network-based augmented retrieval method for auto finance risk control. By integrating relevant data on people and vehicles to construct a social network, and using graph retrieval enhancement technology to create a large model for auto finance risk control questions and answers, business personnel without professional modeling capabilities can effectively identify the potential financial risks of loan applicants through natural language question answering. Attached Figure Description
[0077] Figure 1 This is a schematic diagram of the process of a question-and-answer method for automotive finance risk control based on social network enhanced retrieval according to the present invention;
[0078] Figure 2 This is a schematic diagram of the human-vehicle integrated social network construction of the present invention;
[0079] Figure 3 This is a schematic diagram of the graph-guided retrieval based on social networks according to the present invention;
[0080] Figure 4This is a schematic diagram of the auto finance risk control question and answer based on a large model of the present invention; Detailed Implementation
[0081] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0082] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0083] This invention provides a question-and-answer method for auto finance risk control based on social network-enhanced retrieval, such as... Figure 1-4 As shown, it includes the following steps:
[0084] Step 1: Building a human-vehicle integrated social network, including the construction of a human-vehicle integrated knowledge graph and the construction of a knowledge graph based on public opinion information;
[0085] Step 2: Graph-guided retrieval based on social networks, including query enhancement, graph retrieval based on extended paths, and knowledge enhancement;
[0086] Step 3: Automotive finance risk control question answering based on a large model, including rewriting search results and generating answers based on knowledge enhancement.
[0087] The present invention constructs a knowledge graph combining people and vehicles. When constructing a knowledge graph combining people and vehicles, it is necessary to process the association information between borrowers and purchased vehicles, and merge the association graphs based on borrowers and purchased vehicles to construct a knowledge graph combining people and vehicles. This includes giving a borrower P and his / her purchased vehicle, constructing a basic relationship graph based on P's communication information, residential address, kinship, company information, etc., and constructing a human-vehicle association graph based on information such as the vehicle type and purchase channel of the purchased vehicle.
[0088] Specifically, for borrower P's communication information, their communication records from the past three years are selected, based on the number of communications F. ic and communication duration Perform a weighted calculation to determine the association weight W between the borrower and their corresponding contact i. i r The calculation formula is as follows:
[0089]
[0090] Among them, F c and D c These represent the total number of communications and the total duration of communications for borrower P, respectively. and Represent the weights of communication frequency and communication duration respectively, and satisfy the following conditions: The relationship weights of all contacts of borrower P are calculated using the above formula, and then a communication relationship graph is constructed based on the relationship weights. The specific calculation formula is as follows:
[0091]
[0092] Among them, E i Indicates whether communicator i and borrower P are related (1 indicates related, 0 indicates not related), T c This represents the communication threshold (typically the mean of the borrower P's associated weights).
[0093] For borrower P's kinship and company information, relationships are typically calculated using close relatives within three generations, employees of the same company, or employees of the same department to construct kinship and company association graphs. The communication relationship graph, kinship graph, and company association graph are then merged (i.e., merging related individuals with the same ID number or mobile phone number) to ultimately form the basic relationship graph.
[0094] Regarding the type of vehicle and purchase channel information for the borrower's vehicle, firstly, statistics on borrower P... i The number of times fraud or deception has occurred in the past three months for the corresponding vehicle type. Number of times fraud or deception occurred with the vehicle's purchase channel Secondly, calculate borrower P i Other borrowers P j Vehicle type association weight The specific formula is as follows:
[0095]
[0096] Where, N VT This represents the total number of write-offs or frauds involving vehicle type VT. Then, the number of times borrower P... i Other borrowers P j The weight of the relationship between purchasing channels The specific formula is as follows:
[0097]
[0098] Where, N DC This represents the total number of times fraud or write-offs occurred at the vehicle purchase channel (DC). Finally, it assigns a weight to the association between vehicle types. Weight of the relationship with the purchase channel Weighted fusion is performed to obtain the final vehicle relationship weights, and a human-vehicle relationship graph is constructed by combining the weight thresholds. The calculation formula is as follows:
[0099]
[0100] in, and Represent the weights of vehicle type and purchase channel respectively, and satisfy the following conditions:
[0101] E i,j Indicates borrower P i With borrower P j Whether it is related (1 indicates related, 0 indicates not related), T V This represents the vehicle relationship threshold (usually the average of borrower P's vehicle relationships).
[0102] After completing the construction of the basic relationship graph and the human-vehicle association graph, the two association graphs are merged according to the borrower to finally construct a knowledge graph combining human and vehicle.
[0103] The construction of a knowledge graph based on public opinion information involves discovering entities, attributes, and relationships contained in public opinion on social networks to construct triples in the form of <entity, relation, entity>. This includes entity extraction, attribute extraction, and relation extraction. In constructing a knowledge graph of public opinion information related to people and vehicles, this invention requires processing relevant information from social platforms to discover entities, attributes, and relationships contained in public opinion on social networks, constructing triples in the form of <entity, relation, entity>, and ultimately achieving the construction of a knowledge graph based on public opinion information. Entities, as the basic elements of a knowledge graph, represent a collection of objects of a certain type in the real world. The core task of entity extraction is to identify, detect, and classify these entities from text. In the analysis of public opinion on social networks, it is necessary to extract multiple types of entities in a complex environment, including but not limited to users participating in discussions, vehicles, and the spatiotemporal context of multi-platform interactions. To construct a knowledge graph of public opinion on social networks, this patent focuses on several key entities: platform entities, i.e., the online environment in which netizens discuss public opinion topics; individual entities, covering all individuals participating in topic discussions, such as ordinary netizens and self-media personalities; vehicle entities, involving vehicle-related topics, such as vehicle problems and vehicle reviews; organizational entities, involving institutions or departments related to the topic, such as government departments or non-profit organizations; and geographic location entities, which not only refer to the location where the event occurs, but also the location where information is published and the location of the participants. Accurately extracting and classifying these entities is crucial for understanding the dynamics of public opinion on social networks and provides solid data support for subsequent in-depth analysis.
[0104] Entity extraction specifically includes:
[0105] Step 1: Collect topic data from social network platforms using web crawling technology or API services;
[0106] Step 2: Preprocess the collected topic data, removing redundant information and sensitive information, etc.
[0107] Step 3: Use entity extraction techniques to obtain relevant entities from the processed topic data. There are two main ways to extract entities: rule-based and dictionary-based methods and machine learning-based methods.
[0108] Attributes are inherent characteristics of entities, arising with their existence and possessing objective reality. They not only describe the characteristics of the entity itself but also express the relationships between entities. Through attributes, we can gain a deeper understanding and reveal the concept and uniqueness of entities, and provide a more detailed description of the relationships between them. In the knowledge graph of social network public opinion, attribute extraction aims to capture key information reflecting the characteristics of public opinion, helping to comprehensively understand its features and deeper meanings. Attributes can be presented in two ways: one is in the form of <entity-relationship> pairs, showing the connections between entities; the other is by directly describing specific aspects of an entity, such as a user's nickname or ID. These two forms of representation provide valuable perspectives for analyzing and interpreting public opinion data. Through precise attribute extraction, we can not only enrich the profile of entities but also enhance our understanding of complex public opinion dynamics. This is relevant to attribute extraction in the knowledge graph of social network public opinion.
[0109] Attribute extraction specifically includes:
[0110] (1) Use web crawlers and API services to obtain relevant data and perform preprocessing;
[0111] (2) Extract attribute information from preprocessed data using attribute extraction techniques (including rule-based extraction methods and statistical extraction methods);
[0112] (3) Associate the extracted attribute information with the extracted entities.
[0113] Relationships refer to the connections between entities and between entities and attributes, and are an indispensable component of knowledge graph construction. In the knowledge graph of social network public opinion, relationship extraction aims to identify and extract the connections between all categories of entities, connecting previously isolated entities and weaving them into a complex public opinion network. Users establish various relationships between entities in the social network public opinion field through information behaviors such as following, commenting, forwarding, and liking, such as the inclusion relationship between public opinion topics and platforms, the publishing relationship between platforms and users, and the interaction relationship between users (including forwarding, commenting, and liking). Furthermore, by analyzing the text content expressed by users, the directional and semantic connections between texts can be discovered; and from the platform's perspective, parallel relationships also exist between different platforms. These diverse types of connections collectively constitute the foundation of the social network public opinion knowledge graph, enabling a more comprehensive understanding and analysis of public opinion dynamics.
[0114] Relation extraction specifically includes:
[0115] Relevant data was obtained using web crawlers and API services, and the data was preprocessed.
[0116] The preprocessed data is analyzed to extract the relationships between entities (this patent uses template-based relationship extraction and machine learning-based relationship extraction).
[0117] Finally, all extracted relationships are associated with entities to obtain all entity relationship pairs in the public opinion knowledge graph.
[0118] The construction of a knowledge graph based on public opinion information is essentially the process of representing public opinion information occurring on social network platforms using a knowledge graph. Through entity extraction, attribute extraction, and relation extraction, the nodes, edges, attributes, and other elements of public opinion information on social networks are extracted. Finally, the corresponding entities, attributes, and relations are merged to form a knowledge graph based on public opinion information.
[0119] By integrating knowledge graphs that combine people and vehicles with knowledge graphs based on public opinion information, a social network that combines people and vehicles can be formed.
[0120] Query Enhancement: To ensure high graph retrieval quality, this invention enhances the query information before retrieving data from human-vehicle integrated social networks. This is achieved through query expansion, query decomposition, and other query strategies to enrich the query information and enable better retrieval.
[0121] In the field of auto finance risk control, risk analysis is usually conducted on individual loan recipients and sales channels. Therefore, queries typically contain only short question texts such as individual loan recipients, sales channels, financial terms, and keywords, with limited information content. Query expansion improves search results by supplementing or refining the original query and adding additional relevant terms or concepts.
[0122] Query expansion specifically includes: terminology recognition, keyword extraction, context recognition, query terminology, and question enhancement. Terminology recognition specifically involves identifying terms and abbreviations in the query question and generating accurate explanations for them. Keyword extraction specifically involves extracting keywords from the query question and generating explanations and synonyms for them. Context recognition specifically involves identifying the contextual information of terms, abbreviations, and keywords in the query question. Query terminology specifically involves using the identified terms to query a financial terminology dictionary to obtain extended definitions, descriptions, and annotations. Question enhancement specifically involves combining the original query question, keywords, contextual information, and detailed terminology definitions, and using a large language model to rewrite and reshape the query question to provide clear context and resolve ambiguities. Query decomposition specifically involves using query decomposition techniques to break down the original and enhanced query questions into smaller, more specific subqueries, allowing each subquery to focus on a specific aspect of the original query. Both the original and enhanced query questions are decomposed into clauses, each representing a specific association, and relevant triples are retrieved sequentially for each clause.
[0123] Graph retrieval based on extended paths specifically refers to the following: In order to obtain retrieval results that are highly relevant to the query question, this invention uses an extended path approach to perform graph retrieval, so as to obtain high-quality and highly relevant graph retrieval results.
[0124] For a given query, decompose the clauses and extract key topic entities from the sentences. Starting with the extracted topic entities, construct an expanded path by following a sequence prediction process. Assume the retrieval path at step t is p. t ,
[0125] p t =(r1,r2,...,r t )
[0126] Where, r i This represents the relation retrieved in step i. It is constructed by filling entities into the path relation to build a relation based on p. t Tree structure T t ,
[0127] T t =(e s ,r1,E1,r2,E2,...,r t E t )
[0128] Among them, e s E represents the subject entity of the sentence. i This indicates that relation r is retrieved through step i. i The set of related entities. The retrieval relationship in step t+1 starts from E. t The adjacency relationships are obtained by focusing on them, that is, by calculating the relationship r. t+1 Extended conditional probability p(r|p t To determine.
[0129] To calculate the conditional probability p(r|p t To determine the relevance between relation r and query clause q, this patent uses the cosine similarity between the embedding vectors of relation r and query clause q as the relevance calculation result sim(q,r).
[0130]
[0131] Here, `emb(·)` represents vector embedding. Since different retrieval relation time steps should focus on specific parts of the input query clause when expanding the relation, the original question is concatenated with historical expanded relations as input to update the question's embedding, i.e.
[0132] emb(q t )=f([q;T t ])
[0133] Among them, f is usually encoded using vector encoding methods such as Word2Vec and BERT.
[0134] Relation r t+1 Extended conditional probability p(r|q) t The calculation formula for ) is as follows:
[0135]
[0136] If p(r|q) t If the probability is greater than or equal to 0.5, then the relation with the highest probability is selected as the retrieval relation r for step t+1. t+1 If p(r|q) t If the probability is less than 0.5, path expansion stops. Based on the probability calculation formula for relation expansion mentioned above, the probability of the path corresponding to the query clause q is calculated using the joint distribution of all relations in the path.
[0137]
[0138] Here, n represents the number of relations in path p. Since the path with the highest probability cannot be guaranteed to be correct, a bundle search can be used to obtain multiple paths to construct a tree structure, which can then be used as the query result for the clause.
[0139] The multiple paths of the query are merged to generate the final subgraph of the clause query. Finally, after the expanded path retrieval calculations for all query question decomposition clauses are completed, graph retrieval results based on the expanded paths can be generated.
[0140] Search result optimization includes optimizing and improving the initial search results using knowledge enhancement strategies. Knowledge merging involves integrating relevant knowledge from different sources to enrich the content of the search results, while knowledge pruning removes redundant or irrelevant information to ensure the results are concise and focused. The application of these techniques ensures that the final search results not only comprehensively cover content that users may be interested in but also highly match the user's specific information needs. Specifically, this involves knowledge merging and knowledge pruning. Knowledge merging specifically involves merging the retrieved results to achieve information compression and aggregation, helping to obtain a more comprehensive perspective by integrating relevant details from multiple sources. This method not only enhances the completeness and coherence of information but also alleviates problems related to input length limitations in the model. This patent improves reasoning efficiency by fusing entities and relations in the retrieved triples (e.g., merging entities and relations that are identical, similar in meaning, or have consistent referents) and merging corresponding nodes and edges to compress the retrieved subgraph. Knowledge pruning specifically involves filtering out less relevant or redundant search information to optimize the results. This patent first calculates the similarity between the retrieved triples and the original query question and the enhanced query question; then it performs a comprehensive ranking based on the two similarity calculation results; finally, it selects the top n search results from the comprehensive ranking, connects them with their respective questions, and uses the Sentence-BERT model to re-rank them, selecting the top k triples as the final search results.
[0141] After completing the graph-guided retrieval based on social networks, in order to better convert the knowledge representation of the knowledge graph into the corresponding text representation, this patent rewrites the graph retrieval results and then combines them with the question text to construct prompts which are then input into the large model. This can significantly improve the accuracy and robustness of the large model's answer output.
[0142] Rewriting search results specifically includes:
[0143] The core of rewriting search results is to convert structured triples into free-form text based on the KG-to-Text model. The specific process is as follows:
[0144] Given a knowledge graph G, the triples in the knowledge graph G are verbalized into ternary texts x by connecting head entities, relations, and tail entities. That is, the knowledge graph is transformed into a text set of "(head entity, relation, tail entity)".
[0145] The prompt word project constructs a knowledge graph-to-text (KG-to-Text) prompt word prompt1 based on the ternary form text x. The prompt word template is as follows:
[0146] "Your task is to transform the knowledge graph into one or more sentences."
[0147] The knowledge graph is: {ternary text x}.
[0148] The sentence is:
[0149] The prompt word prompt1 is input into a large model fine-tuned based on KG-to-Text labeled data for knowledge graph to text conversion to obtain free-format text y, thereby rewriting the search results.
[0150] When answering questions, each reasoning path is converted into a prompt word to obtain the corresponding free text from the input to the large model. The free text is then merged into the question-related knowledge.
[0151] Knowledge-enhanced answer generation includes a large-model-based prompt word prompt2 for automotive finance risk control, designed to integrate knowledge generated from KG-to-Text with questions. The template is as follows:
[0152] "Your task is to answer the question based on the facts relevant to it."
[0153] Facts related to the question: {free-form text y}
[0154] Question: {Auto Finance Risk Control Issues Q}
[0155] Answer:".
[0156] By using the prompt word prompt2 template, the free-form text y and the auto finance risk control question Q are mapped to the knowledge graph-enhanced prompt word and input into the financial risk control big model to output the predicted answer.
[0157] All electrical components mentioned in this article are electrically connected to the controller and power supply. The control method of this invention is controlled by the controller. The control circuit of the controller can be implemented by simple programming by those skilled in the art. The power supply is also common knowledge in the art. Furthermore, this invention is mainly used to protect mechanical devices, so the control method and circuit connection will not be explained in detail.
[0158] It should be noted that, for the sake of simplicity, the foregoing embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to the present invention. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0159] It should be understood that the disclosed apparatus can be implemented in other ways, given the several embodiments provided in this application. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units described above may be implemented in other ways in practice. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or communication connections shown or discussed may be through some interfaces; indirect coupling or communication connections between devices or units may be telecommunications or other forms.
[0160] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0161] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit the scope of protection of the invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on these embodiments, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art can still combine, add, delete, or otherwise adjust the features of the various embodiments of the present invention according to the circumstances without conflict or creative effort, thereby obtaining different technical solutions that do not fundamentally depart from the concept of the present invention. These technical solutions also fall within the scope of protection of the present invention.
Claims
1. A question-and-answer method for auto finance risk control based on social network enhanced retrieval, characterized in that, Includes the following steps: Step 1: Building a human-vehicle integrated social network, including the construction of a human-vehicle integrated knowledge graph and the construction of a knowledge graph based on public opinion information; Step 2: Graph-guided retrieval based on social networks, including query enhancement, graph retrieval based on extended paths, and knowledge enhancement; Step 3: Automotive finance risk control question answering based on a large model, including rewriting search results and generating answers based on knowledge enhancement.
2. The question-and-answer method for auto finance risk control based on social network enhanced retrieval as described in claim 1, characterized in that, The knowledge graph construction combining people and vehicles includes, given a borrower P and the vehicle he / she purchases, constructing a basic relationship graph based on P's communication information, residential address, kinship, company information, etc., and constructing a person-vehicle association graph based on the vehicle type, purchase channel, and other information of the vehicle he / she purchases. Specifically, for borrower P's communication information, their communication records from the past three years are selected, based on the number of communications F. i c and communication duration Perform a weighted calculation to determine the association weight W between the borrower and their corresponding contact i. i r The calculation formula is as follows: Among them, F c and D c These represent the total number of communications and the total duration of communications for borrower P, respectively. and Represent the weights of communication frequency and communication duration respectively, and satisfy the following conditions: The relationship weights of all contacts of borrower P are calculated using the above formula, and then a communication relationship graph is constructed based on the relationship weights. The specific calculation formula is as follows: Among them, E i Indicates whether communicator i and borrower P are related (1 indicates related, 0 indicates not related), T c This represents the communication threshold (typically the mean of the borrower P's associated weights). For borrower P's kinship and company information, relationships are typically calculated using close relatives within three generations, employees of the same company, or employees of the same department to construct kinship and company association graphs. The communication relationship graph, kinship graph, and company association graph are then merged (i.e., merging related individuals with the same ID number or mobile phone number) to ultimately form the basic relationship graph. Regarding the type of vehicle and purchase channel information for the borrower's vehicle, firstly, statistics on borrower P... i The number of times fraud or deception has occurred in the past three months for the corresponding vehicle type. Number of times fraud or deception occurred with the vehicle's purchase channel Secondly, calculate borrower P i Other borrowers P j Vehicle type association weight The specific formula is as follows: Where, N VT This represents the total number of write-offs or frauds involving vehicle type VT. Then, the number of times borrower P... i Other borrowers P j The weight of the relationship between purchasing channels The specific formula is as follows: Where, N DC This represents the total number of times fraud or write-offs occurred at the vehicle purchase channel (DC). Finally, it assigns a weight to the association between vehicle types. Weight of the relationship with the purchase channel Weighted fusion is performed to obtain the final vehicle relationship weights, and a human-vehicle relationship graph is constructed by combining the weight thresholds. The calculation formula is as follows: in, and Represent the weights of vehicle type and purchase channel respectively, and satisfy the following conditions: E i,j Indicates borrower P i With borrower P j Whether it is related (1 indicates related, 0 indicates not related), T V This represents the vehicle relationship threshold (usually the average of borrower P's vehicle relationships). After completing the construction of the basic relationship graph and the human-vehicle association graph, the two association graphs are merged according to the borrower to finally construct a knowledge graph combining human and vehicle.
3. The question-and-answer method for auto finance risk control based on social network enhanced retrieval as described in claim 2, characterized in that, The knowledge graph construction based on public opinion information includes discovering entities, attributes, and relationships contained in social network public opinion to construct triples in the form of <entity, relation, entity>, which includes entity extraction, attribute extraction, and relation extraction.
4. The question-and-answer method for auto finance risk control based on social network enhanced retrieval as described in claim 3, characterized in that, The entity extraction specifically includes: Step 1: Collect topic data from social network platforms using web crawling technology or API services; Step 2: Preprocess the collected topic data, removing redundant information and sensitive information, etc. Step 3: Use entity extraction techniques to obtain relevant entities from the processed topic data. There are two main ways to extract entities: rule-based and dictionary-based methods and machine learning-based methods.
5. The question-and-answer method for auto finance risk control based on social network enhanced retrieval as described in claim 3, characterized in that, The attribute extraction specifically includes: (1) Use web crawlers and API services to obtain relevant data and perform preprocessing; (2) Extract attribute information from preprocessed data using attribute extraction techniques (including rule-based extraction methods and statistical extraction methods); (3) Associate the extracted attribute information with the extracted entities.
6. The question-and-answer method for auto finance risk control based on social network enhanced retrieval as described in claim 3, characterized in that, The relation extraction specifically includes: Relevant data was obtained using web crawlers and API services, and the data was preprocessed. The preprocessed data is analyzed to extract the relationships between entities (this patent uses template-based relationship extraction and machine learning-based relationship extraction). Finally, all extracted relationships are associated with entities to obtain all entity relationship pairs in the public opinion knowledge graph.
7. The question-and-answer method for auto finance risk control based on social network enhanced retrieval as described in claim 1, characterized in that, The query expansion specifically includes: terminology recognition, keyword extraction, context recognition, query terminology, and question enhancement. Specifically, terminology recognition involves identifying terms and abbreviations in the query question and generating accurate explanations for them. Keyword extraction involves extracting keywords from the query question and generating explanations and synonyms for them. Context recognition involves identifying the contextual information of terms, abbreviations, and keywords in the query question. Query terminology involves using the identified terms to query a financial terminology dictionary to obtain extended definitions, descriptions, and annotations. Question enhancement involves combining the original query question, keywords, contextual information, and detailed terminology definitions, and rewriting and modifying it using a large language model to form an enhanced query question, providing clear context and resolving ambiguities. The query decomposition method involves breaking down the original and enhanced query questions into smaller, more specific subqueries, ensuring each subquery focuses on a specific aspect of the original query. Both the original and enhanced query questions are decomposed into clauses, each representing a specific association, and relevant triples are retrieved for each clause sequentially.
8. The question-and-answer method for auto finance risk control based on social network enhanced retrieval as described in claim 7, characterized in that, The graph retrieval based on the extended path specifically refers to: For a given query, decompose the clauses and extract key topic entities from the sentences. Starting with the extracted topic entities, construct an expanded path by following a sequence prediction process. Assume the retrieval path at step t is p. t , p t =(r1,r2,...,r t ) Where, r i This represents the relation retrieved in step i. It is constructed by filling entities into the path relation to build a relation based on p. t Tree structure T t , T t =(e s ,r1,E1,r2,E2,...,r t ,AND t ) Among them, e s E represents the subject entity of the sentence. i This indicates that relation r is retrieved through step i. i The set of related entities. The retrieval relationship in step t+1 starts from E. t The adjacency relationships are obtained by focusing on them, that is, by calculating the relationship r. t+1 Extended conditional probability p(r|p t To determine. To calculate the conditional probability p(r|p t To determine the relevance between relation r and query clause q, this patent uses the cosine similarity between the embedding vectors of relation r and query clause q as the relevance calculation result sim(q,r). Here, `emb(·)` represents vector embedding. Since different retrieval relation time steps should focus on specific parts of the input query clause when expanding the relation, the original question is concatenated with historical expanded relations as input to update the question's embedding, i.e. emb(q t )=f([q;T t ]) Among them, f is usually encoded using vector encoding methods such as Word2Vec and BERT. Relation r t+1 Extended conditional probability p(r|q) t The calculation formula for ) is as follows: If p(r|q) t If the probability is greater than or equal to 0.5, then the relation with the highest probability is selected as the retrieval relation r for step t+1. t+1 If p(r|q) t If the probability is less than 0.5, path expansion stops. Based on the probability calculation formula for relation expansion mentioned above, the probability of the path corresponding to the query clause q is calculated using the joint distribution of all relations in the path. Here, n represents the number of relations in path p. Since the path with the highest probability cannot be guaranteed to be correct, a bundle search can be used to obtain multiple paths to construct a tree structure, which can then be used as the query result for the clause. The multiple paths of the query are merged to generate the final subgraph of the clause query. Finally, after the expanded path retrieval calculations for all query question decomposition clauses are completed, graph retrieval results based on the expanded paths can be generated. The optimization of search results includes, after obtaining the initial search results, optimizing and improving these results by adopting knowledge enhancement strategies, specifically knowledge merging and knowledge pruning. Knowledge merging specifically involves: merging the retrieved results to achieve information compression and aggregation, which helps to obtain a more comprehensive perspective by integrating relevant details from multiple sources. Knowledge pruning specifically involves: filtering out irrelevant or redundant search information to optimize the results; calculating the similarity between the retrieved triples and the original query question and the enhanced query question; then, performing a comprehensive ranking based on the two similarity calculation results; finally, selecting the top n search results from the comprehensive ranking and connecting them with their respective questions, and re-ranking them using the Sentence-BERT model, selecting the top k triples as the final search results.
9. The question-and-answer method for auto finance risk control based on social network enhanced retrieval as described in claim 1, characterized in that, The rewriting of the search results specifically includes: The core of rewriting search results is to convert structured triples into free-form text based on the KG-to-Text model. The specific process is as follows: Given a knowledge graph G, the triples in the knowledge graph G are verbalized into ternary texts x by connecting head entities, relations, and tail entities. That is, the knowledge graph is transformed into a text set of "(head entity, relation, tail entity)". The prompt word project constructs a knowledge graph-to-text (KG-to-Text) prompt word prompt1 based on the ternary form text x. The prompt word template is as follows: Your task is to convert the knowledge graph into one or more sentences. The knowledge graph is: {ternary text x}. The sentence is: The prompt word prompt1 is input into a large model fine-tuned based on KG-to-Text labeled data for knowledge graph to text conversion to obtain free-format text y, thereby rewriting the search results. When answering questions, each reasoning path is converted into a prompt word to obtain the corresponding free text from the input to the large model. The free text is then merged into the question-related knowledge.
10. The question-and-answer method for auto finance risk control based on social network enhanced retrieval as described in claim 9, characterized in that, The knowledge-enhanced answer generation includes, in order to integrate the knowledge and questions generated based on KG-to-Text, the patent designed a large-model-based automotive finance risk control question-answering prompt word prompt2, the template of which is as follows: Your task is to answer the question based on the facts relevant to it. Facts related to the question: {free-form text y} Question: {Auto Finance Risk Control Issues Q} Answer:". By using the prompt word prompt2 template, the free-form text y and the auto finance risk control question Q are mapped to the knowledge graph-enhanced prompt word and input into the financial risk control big model to output the predicted answer.