Information decision-making method and system based on collection session big data analysis

By acquiring and enhancing key semantic nodes in big data collection sessions, generating sample training data, and learning decision networks, the problems of low efficiency and insufficient risk identification in traditional collection methods are solved, and efficient risk fraud detection is achieved.

CN119598284BActive Publication Date: 2025-12-09JIANGXI ZHIWEN ZIAN DIGITAL TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411642812.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-18
Publication Date
2025-12-09
Estimated Expiration
2044-11-18

AI Technical Summary

Technical Problem

Traditional debt collection methods rely on manual judgment, which is inefficient and makes it difficult to accurately identify and handle risk and fraud events during the debt collection process. Existing debt collection conversation analysis methods are also unable to capture deep semantic information and risk and fraud characteristics.

Method used

By acquiring big data from collection sessions, enhancing key semantic nodes, generating sample collection session training data, and performing parameter learning on the collection session decision network, a collection session decision network with completed parameter learning is generated, which is used to determine whether there are risky or fraudulent events in the collection session data.

Benefits of technology

It improves the accuracy and efficiency of debt collection session data analysis, enabling timely detection and handling of potential risk and fraud incidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0005138923990000011
    Figure HDA0005138923990000011
  • Figure HDA0005138923990000021
    Figure HDA0005138923990000021
Patent Text Reader

Abstract

The application provides an information decision-making method and system based on collection session big data analysis. First, collection session big data is acquired, then, for a session data stream in the collection session big data, a sample collection session training data is generated by enhancing a key semantic node of a collection target user in the session data stream. In this process, if a semantic representation vector associated with the key semantic node of the collection target user exists a risk fraud event, the semantic representation vector is included in the sample collection session training data. Finally, the sample collection session training data is used to learn parameters of a collection session decision network, thereby generating a completed parameter learning collection session decision network capable of accurately deciding whether the collection session data exists a risk fraud event. Thus, by deeply mining key information in the collection session big data, efficient identification of the risk fraud event is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to an information decision-making method and system based on collection conversation big data analysis. BACKGROUND

[0002] In the financial, credit and other industries, collection is an important link to ensure the collection of funds and maintain the rights and interests of creditors. However, the traditional collection method often relies on manual judgment and experience operation, which is not only inefficient, but also difficult to accurately identify and handle the risk fraud events that may occur in the collection process. Especially when facing a large amount of collection conversation data, manual analysis is not enough and is prone to omissions and misjudgments.

[0003] With the rapid development of big data technology, it is possible to use big data analysis methods to deeply mine and analyze collection conversation data. By collecting and organizing collection conversation data, a collection conversation big data set containing rich information can be formed. These data sets contain information about the behavior patterns, repayment willingness, fraud risk, and other aspects of the collection target users, providing strong data support for collection decision-making.

[0004] However, how to effectively extract key information from collection conversation big data and make accurate decisions based on it is still a problem to be solved. Existing collection conversation analysis methods mostly focus on simple data statistics and rule matching, and are difficult to capture deep semantic information and risk fraud features in the conversation. SUMMARY

[0005] In view of the above-mentioned problems, in combination with the first aspect of the present application, the embodiments of the present application provide an information decision-making method based on collection conversation big data analysis, which comprises:

[0006] Obtaining collection conversation big data, the collection conversation big data comprising a plurality of conversation data streams associated with a collection target user arranged according to the order of collection conversation nodes;

[0007] For a conversation data stream in the collection conversation big data, enhancing the key semantic nodes of the collection target user in the conversation data stream, generating sample collection conversation training data, the semantic representation vector associated with the key semantic nodes of the collection target user in the sample collection conversation training data exists a risk fraud event, a is a positive integer;

[0008] According to the sample collection conversation training data, the parameters of the collection conversation decision network are learned to generate a collection conversation decision network with completed parameter learning, which is used to decide whether the collection conversation data exists a risk fraud event.

[0009] In a possible implementation manner of the first aspect, the enhancing the key semantic nodes of the collection target user in the conversation data stream, and generating the sample collection conversation training data, comprises:

[0010] The each conversation semantic unit in the enhanced conversation data stream is enhanced, and an enhanced conversation data stream is generated. The conversation semantic units in the enhanced conversation data stream exhibit inconsistency with the conversation semantic units in the conversation data stream.

[0011] The enhanced conversation data stream and the conversation data stream are aggregated to generate a fictitious conversation data stream. The key semantic nodes of the collection target user in the fictitious conversation data stream exhibit inconsistency with the key semantic nodes of the collection target user in the conversation data stream. The non-key semantic nodes of the collection target user in the fictitious conversation data stream exhibit consistency with the non-key semantic nodes of the collection target user in the conversation data stream.

[0012] The a conversation data streams are covered as various corresponding fictitious conversation data streams of the a conversation data streams to generate the sample collection conversation training data.

[0013] In a possible implementation manner of the first aspect, the aggregating the enhanced conversation data stream and the conversation data stream to generate a fictitious conversation data stream comprises:

[0014] A plurality of key semantic node sets corresponding to the collection target user in the conversation data stream are acquired. Each key semantic node set comprises a plurality of key semantic nodes.

[0015] For any one key semantic node set in the plurality of key semantic node sets, a semantic feature field corresponding to the conversation data stream is generated according to the key semantic node set. The semantic feature field represents a penetration paragraph of the key semantic node set in the conversation data stream.

[0016] Under the limitation of the various corresponding semantic feature fields of the plurality of key semantic node sets, the enhanced conversation data stream and the conversation data stream are aggregated to generate the fictitious conversation data stream.

[0017] In a possible implementation manner of the first aspect, the generating the semantic feature field corresponding to the conversation data stream according to the key semantic node set comprises:

[0018] For any one conversation semantic unit in the conversation data stream, a first correlation degree between the conversation semantic unit and the penetration paragraph in which the key semantic node set is located is acquired.

[0019] determine a semantic feature value of the conversation semantic unit based on the first correlation degree, wherein the semantic feature value of the conversation semantic unit and the first correlation degree have a positive correlation relationship;

[0020] generate the semantic feature field according to the semantic feature values of the conversation semantic units in the conversation data stream.

[0021] In a possible implementation of the first aspect, the aggregating the enhanced conversation data stream and the conversation data stream under the restrictions of the semantic feature fields corresponding to the plurality of key semantic nodes to generate the fictitious conversation data stream comprises:

[0022] for any one of the semantic feature fields corresponding to the plurality of key semantic nodes, optimizing the enhanced conversation data stream according to the semantic feature field to generate a first composite conversation data stream, and optimizing the conversation data stream according to an inverse semantic feature field corresponding to the semantic feature field to generate a second composite conversation data stream, wherein the sum of the semantic feature values of the conversation semantic units of the same penetration paragraph in the semantic feature field and the inverse semantic feature field is 1;

[0023] generating a composite conversation data stream corresponding to the semantic feature field according to the first composite conversation data stream and the second composite conversation data stream;

[0024] aggregating the composite conversation data streams corresponding to the plurality of semantic feature fields to generate the fictitious conversation data stream.

[0025] In a possible implementation of the first aspect, the optimizing the enhanced conversation data stream according to the semantic feature field to generate the first composite conversation data stream, and optimizing the conversation data stream according to the inverse semantic feature field corresponding to the semantic feature field to generate the second composite conversation data stream comprises:

[0026] for any one of the conversation semantic units in the enhanced conversation data stream, multiplying the conversation semantic unit by the semantic feature value corresponding to the conversation semantic unit in the semantic feature field to generate the first composite conversation data stream;

[0027] for any one of the conversation semantic units in the conversation data stream, fusing the conversation semantic unit with the semantic feature value corresponding to the conversation semantic unit in the inverse semantic feature field to generate the second composite conversation data stream.

[0028] In a possible implementation of the first aspect, the enhancing the conversation semantic units in the conversation data stream to generate the enhanced conversation data stream comprises:

[0029] The session data stream is preliminarily segmented according to semantic integrity to generate a plurality of independent session semantic units, the types of words in each session semantic unit are identified, and key entities and semantic roles in each session semantic unit are marked based on the types of words to generate marking information of each session semantic unit, the key entities include debt information related to collection, financial products involved, identity information of a collection target user, and domain words related to risk fraud, and the semantic roles include an agent and a patient;

[0030] Based on the marking information of each session semantic unit, a corresponding semantic association knowledge base is constructed, the semantic association knowledge base is obtained by analyzing a collection-related corpus and a general corpus, and specifically, for a noun in each session semantic unit, an association relationship with other related nouns is established in the semantic association knowledge base, and for a verb in each session semantic unit, semantic associations under different combinations of subjects and objects are established, and for an adjective and an adverb, semantic associations with modified objects and semantic change associations under different contexts are established;

[0031] For each session semantic unit, a core semantic of the session semantic unit is extracted according to the semantic association knowledge base to obtain a session semantic unit set containing the core semantic;

[0032] For each session semantic unit containing the core semantic, semantic expansion is performed according to the semantic association knowledge base, wherein for a noun part in the core semantic, other nouns, adjectives, and verbs related to the noun part are searched from the semantic association knowledge base and then expanded into the session semantic unit, and for a verb part in the core semantic, synonyms, near-synonyms, and different tenses and voices of the verb part are searched and then expanded into the session semantic unit to obtain a session semantic unit set after semantic expansion;

[0033] Based on the session semantic unit set after semantic expansion, a semantic conversion operation is performed, part of semantics in the session semantic unit is converted according to the semantic association knowledge base, and a session semantic unit set after semantic conversion is obtained;

[0034] Based on prior context information of a collection session, context adaptability adjustment is performed on the session semantic unit set after semantic conversion to obtain a session semantic unit set after context adaptation;

[0035] Logical relationships between each session semantic unit in the session semantic unit set after context adaptation are analyzed, and it is determined whether there is semantic jump and semantic contradiction between adjacent session semantic units, if there is, the related session semantic units are adjusted to obtain a logically coherent session semantic unit set;

[0036] The logically coherent conversation semantic units are recombined in the order in the original implementation conversation data stream to form an enhanced conversation data stream.

[0037] In a possible implementation of the first aspect, the parameter learning of the collection conversation decision network according to the sample collection conversation training data to generate the collection conversation decision network after the parameter learning includes:

[0038] The first mapping knowledge vector of the sample collection conversation training data is obtained by using the collection conversation decision network, and the first mapping knowledge vector represents semantic transition information of a corresponding semantic representation vector of the collection target user in a time domain in the sample collection conversation training data;

[0039] The second mapping knowledge vector is generated by using the collection conversation decision network according to the first mapping knowledge vector, and the second mapping knowledge vector represents semantic transition information of the corresponding semantic representation vector of the collection target user in a space domain in the sample collection conversation training data;

[0040] The decision result of the sample collection conversation training data is generated by using the collection conversation decision network according to the second mapping knowledge vector, and the decision result of the sample collection conversation training data represents whether a semantic change of a semantic representation vector associated with a key semantic node of the collection target user in the sample collection conversation training data exists a risk fraud event;

[0041] The parameter learning of the collection conversation decision network is performed based on the decision result to generate the collection conversation decision network after the parameter learning.

[0042] In a possible implementation of the first aspect, the generating of the second mapping knowledge vector by using the collection conversation decision network according to the first mapping knowledge vector includes:

[0043] The first mapping knowledge vector is subjected to attention weight distribution processing to generate a first mapping knowledge vector after attention weight distribution processing;

[0044] The second mapping knowledge vector is generated by using the collection conversation decision network according to the first mapping knowledge vector after the attention weight distribution processing;

[0045] The knowledge segments in the first mapping knowledge vector are arranged in a time domain, and the knowledge segments in the first mapping knowledge vector after the attention weight distribution processing are arranged in a space domain;

[0046] In a possible implementation of the first aspect, the collection session decision network comprises a time domain knowledge encoder and a space domain knowledge encoder, the time domain knowledge encoder is generated according to a first initialization encoder and a first recurrent neural network, the space domain knowledge encoder is generated according to a second initialization encoder and a second recurrent neural network, the first recurrent neural network is configured to obtain the first mapping knowledge vector according to the first initialization encoder, and the second recurrent neural network is configured to obtain the second mapping knowledge vector according to the second initialization encoder.

[0047] The parameter learning of the collection session decision network based on the decision result comprises:

[0048] Based on the decision result, a training error parameter of the collection session decision network is determined.

[0049] The weight parameter information of the first initialization encoder and the weight parameter information of the second initialization encoder are locked, the weight parameter information of the first recurrent neural network and the weight parameter information of the second recurrent neural network are optimized based on the training error parameter, and the collection session decision network after the parameter learning is generated.

[0050] In another aspect, the embodiment of the present application further provides an information decision system based on collection session big data analysis, comprising a processor, a machine readable storage medium, the machine readable storage medium is connected with the processor, the machine readable storage medium is used for storing programs, instructions or codes, and the processor is used for executing the programs, instructions or codes in the machine readable storage medium to realize the above method.

[0051] Based on the above aspect, the embodiment of the present application enhances the key semantic nodes to generate sample collection session training data by acquiring and sorting collection session big data, and then performs parameter learning on the collection session decision network, and finally generates the collection session decision network after the parameter learning which can accurately decide whether the collection session data exists risk fraud event, effectively improves the accuracy and efficiency of the collection session data analysis, and helps to discover and handle potential risk fraud events in time. BRIEF DESCRIPTION OF DRAWINGS

[0052] Figure 1 is an execution flow diagram of the information decision method based on collection session big data analysis provided by the embodiment of the present application.

[0053] Figure 2 is a hardware architecture diagram of the information decision system based on collection session big data analysis provided by the embodiment of the present application. DETAILED DESCRIPTION

[0054] The application will be specifically described below in combination with the drawings of the specification, Figure 1 is a flowchart of an information decision-making method based on collection session big data analysis provided by an embodiment of the application. The information decision-making method based on collection session big data analysis will be described in detail below.

[0055] In step S110, collection session big data is acquired, which includes a plurality of session data streams associated with a collection target user arranged according to a collection session node sequence.

[0056] In this embodiment, the collection session big data is a collection of a series of session information related to the target users who need to be collected. For example, a financial institution provides a variety of financial products, such as personal loans, credit card consumption, etc. When the user fails to repay on time, the collection process will be entered.

[0057] The server will collect these session data from multiple systems. For example, the telephone communication record system of the customer service department and the collection target user. When the customer service personnel makes a collection call, the telephone system will automatically record the entire call process, including the words of the customer service and the response of the user. These call records are arranged in the order of the call occurrence, that is, the collection session node sequence. The call record of each call forms a session data stream. For another example, the chat record system of the online customer service and the collection target user. When the user has a conversation with the online customer service about the repayment collection through the official website or the mobile APP of the financial institution, these chat records will also be arranged in the order of the conversation and become a session data stream.

[0058] There are also some possible channels, such as the communication between the offline collection personnel and the user face to face, and the communication content is input into a special system. The server will also acquire these data and arrange them into session data streams according to the collection session node sequence. These numerous session data streams associated with the collection target user converge together to form the collection session big data acquired by the server.

[0059] Taking a specific credit card collection scenario as an example, the credit card of user A is overdue and not repaid. The customer service personnel first makes a call on the first day after the overdue period, and in the call, the customer service introduces his own identity, informs user A that the credit card has been overdue, and inquires about the repayment plan. User A says that the funds are tight recently and may need a few days to repay the money. This call record is part of a session data stream. On the third day, the customer service makes another call to remind user A that the overdue period may affect the credit record, and user A says that he is trying to raise money. The two call records in order complete a session data stream associated with user A, the collection target user. And the server will acquire a plurality of session data streams of different users to form the collection session big data.

[0060] In step S120, for a number of a of the collection session big data, the key semantic nodes of the collection target user in the session data stream are enhanced to generate sample collection session training data, and the semantic representation vector associated with the key semantic nodes of the collection target user exists in the sample collection session training data. Risk fraud events, and a is a positive integer.

[0061] In this embodiment, after obtaining the collection session big data, the a session data streams in it are processed. It is assumed that the server selects 100 (i.e. a = 100) from a large number of session data streams as processing objects.

[0062] Take one of the session data streams related to user B as an example. This user B is a personal loan overdue customer. In the original session data stream, it contains multiple dialogue records between the customer service and user B. Some key semantic nodes may include user B's mention of repayment ability, source of funds, repayment time and other related expressions.

[0063] First, the server enhances each session semantic unit in this session data stream to generate an enhanced session data stream. For the processing of the session semantic unit, the server will perform a series of complex operations. For example, the session data stream is preliminarily segmented according to semantic integrity to generate multiple independent session semantic units. It is assumed that in a dialogue, user B says "I have been unemployed recently and have no income source, so I can't repay the loan". The server will identify the word types in this session semantic unit, analyze the words such as "unemployment", "income source" and "can't repay the loan", and based on these word types, mark the key entities (such as the debt information related to collection, such as "loan") and semantic roles (such as user B as the agent, indicating that he is in the state of unemployment and no income, so he cannot perform the action of the patient, i.e. repaying the loan), to generate the marking information of the session semantic unit.

[0064] Then, the server constructs a corresponding semantic association knowledge base based on the marking information of each session semantic unit. For the noun "unemployment" in this session semantic unit, the server will establish an association relationship with other related nouns in the semantic association knowledge base, such as "unemployment" may be associated with "economic situation", "job market" and other nouns; for the verb "can't repay", establish its semantic association under different subject and object combinations, such as the cause-and-effect relationship of "because there is no income, so can't repay the loan", and for adjectives and adverbs (if any), establish the semantic association between them and the modified objects and the semantic change association in different contexts.

[0065] Next, for each session semantic unit, the server extracts the core semantics of the session semantic unit according to the semantic association knowledge base. For this session semantic unit of user B, the core semantics can be "unemployment leads to no ability to repay". Then, according to the semantic association knowledge base, the semantics is expanded. For the noun part "unemployment" in the core semantics, other nouns, adjectives and verbs related to it are found from the semantic association knowledge base and expanded into the session semantic unit, which can become "because of the industry recession, unemployment, no income source, really no ability to repay the loan"; for the verb part "repay" in the core semantics, its synonyms, near-synonyms and different tenses and voices are found and expanded into the session semantic unit, such as "unable to perform the obligation of repayment". After obtaining the set of session semantic units after semantic expansion, on this basis, semantic conversion is performed. According to the semantic association knowledge base, part of the semantics in the session semantic unit is converted, such as converting "no ability to repay the loan" to "difficult to bear the responsibility of repaying the loan". Based on the prior context information of the collection session, the set of session semantic units after semantic conversion is adjusted for context adaptability to obtain the set of context-adapted session semantic units. For example, if some preferential repayment policies of the bank are mentioned in the previous dialogue, then in this session semantic unit it can be adjusted to "although I know that the bank has preferential repayment policies, I am now unemployed, and still difficult to bear the responsibility of repaying the loan".

[0066] After that, the logical relationship between each session semantic unit in the set of context-adapted session semantic units is analyzed, and it is judged whether there is semantic jump and semantic contradiction between adjacent session semantic units. If so, the related session semantic units are adjusted to obtain a set of logically coherent session semantic units. Finally, the logically coherent session semantic units are recombined according to the order in the original session data stream to form an enhanced session data stream.

[0067] After the enhanced conversation data stream is generated, the server aggregates the enhanced conversation data stream and the original conversation data stream to generate the fabricated conversation data stream. First, the server obtains a plurality of key semantic node sets corresponding to the collection target user in the original conversation data stream. For example, for the conversation data stream of user B, a key semantic node set can be all expressions about repayment ability, including key semantic nodes such as "lost job and no income" and "no savings". Then for each key semantic node set, a semantic feature domain corresponding to the conversation data stream is generated according to the key semantic node set. Assuming that for the key semantic node set of repayment ability, the server obtains a first correlation degree between any one of the conversation semantic units in the conversation data stream and the penetration paragraph where the key semantic node set is located. For example, user B says "I recently spent a sum of money on medical treatment", and the first correlation degree between this conversation semantic unit and the penetration paragraph (the expression paragraph about the financial situation) where the key semantic node set of repayment ability is located can be high, because medical expenses also affect the financial situation and thus affect the repayment ability. Based on the first correlation degree, the semantic feature value of the conversation semantic unit is determined, and the semantic feature value of the conversation semantic unit is positively correlated with the first correlation degree. If the first correlation degree is 0.8, the semantic feature value can be 0.8. According to the semantic feature values of the conversation semantic units in the conversation data stream, the semantic feature domain is generated.

[0068] After obtaining the semantic feature domains corresponding to the plurality of key semantic node sets, for any one of the semantic feature domains, the enhanced conversation data stream is optimized according to the semantic feature domain to generate a first composite conversation data stream, and the original conversation data stream is optimized according to the inverse semantic feature domain corresponding to the semantic feature domain to generate a second composite conversation data stream. For example, for any one of the conversation semantic units in the enhanced conversation data stream corresponding to the semantic feature domain of repayment ability, the conversation semantic unit is multiplied by the semantic feature value corresponding to it in the semantic feature domain to generate the first composite conversation data stream; for any one of the conversation semantic units in the original conversation data stream, the conversation semantic unit is fused with the semantic feature value corresponding to it in the inverse semantic feature domain to generate the second composite conversation data stream. The sum of the semantic feature values of the conversation semantic units in the same penetration paragraph in the semantic feature domain and the inverse semantic feature domain is 1. Then, according to the first composite conversation data stream and the second composite conversation data stream, a composite conversation data stream corresponding to the semantic feature domain is generated. Finally, the composite conversation data streams corresponding to the plurality of semantic feature domains are aggregated to generate the fabricated conversation data stream. In the fabricated conversation data stream, the key semantic nodes of the collection target user exhibit inconsistency with the key semantic nodes in the original conversation data stream, and the non-key semantic nodes exhibit consistency with the non-key semantic nodes in the original conversation data stream.

[0069] The server covers the 100 session data streams in the above manner as corresponding fictitious session data streams, generating sample collection session training data. And in this sample collection session training data, the semantic representation vector associated with the key semantic node of the collection target user has a risk fraud event. For example, in the fictitious session data stream, user B may have some unreasonable expressions related to repayment ability, such as suddenly having a large amount of funding source but not repaying in a short period of time, which may indicate a risk fraud event.

[0070] Step S130, according to the sample collection session training data, the parameter learning of the collection session decision network is carried out, and the collection session decision network after completing the parameter learning is generated, which is used to decide whether the collection session data has a risk fraud event.

[0071] In the present embodiment, the sample collection session training data has been generated in the foregoing embodiment, and now these data are used to learn the parameters of the collection session decision network. The collection session decision network includes a time domain knowledge encoder and a space domain knowledge encoder, wherein the time domain knowledge encoder is generated according to the first initialization encoder and the first recurrent neural network, and the space domain knowledge encoder is generated according to the second initialization encoder and the second recurrent neural network.

[0072] Firstly, the server obtains the first mapping knowledge vector of the sample collection session training data by using the collection session decision network. The first mapping knowledge vector represents the semantic transition information of the corresponding semantic representation vector of the collection target user in the sample collection session training data in the time domain. Taking the fictitious session data stream of user B mentioned above as an example, in this session data stream, the expression of user B about repayment ability has semantic transition in the time domain as the conversation proceeds, from no repayment ability at the beginning to the unreasonable expression of funding source suddenly appearing later, and these semantic transition information is encoded in the first mapping knowledge vector.

[0073] Then, the server generates a second mapping knowledge vector based on the first mapping knowledge vector using the collection session decision network. In this process, the server performs attention weight distribution processing on the first mapping knowledge vector to generate a first mapping knowledge vector after attention weight distribution processing. Assuming that in the first mapping knowledge vector, the knowledge fragments about the early part and the sudden change part of the repayment ability expression of user B are arranged in the time domain. After attention weight distribution processing, these knowledge fragments are rearranged in the spatial domain according to their importance to form a first mapping knowledge vector after attention weight distribution processing. Then, the second mapping knowledge vector is generated based on the first mapping knowledge vector after attention weight distribution processing using the collection session decision network. The second mapping knowledge vector represents the semantic transition information of the semantic representation vector corresponding to the collection target user in the sample collection session training data in the spatial domain.

[0074] Next, the server generates a decision result of the sample collection session training data based on the second mapping knowledge vector using the collection session decision network. The decision result represents whether there is a risk fraud event in the semantic change of the semantic representation vector associated with the key semantic node of the collection target user in the sample collection session training data. For the fictitious conversation data stream of user B, if there is an obvious unreasonable place in the semantic change reflected by the second mapping knowledge vector, such as the expression of repayment ability does not conform to logic and matches the pattern related to risk fraud, the decision result will determine that there is a risk fraud event.

[0075] Finally, based on the decision result, the parameters of the collection session decision network are learned. The server determines the training error parameters of the collection session decision network based on the decision result. If it is determined that there is a risk fraud event but it is actually a false positive, or vice versa, a training error will be generated. By locking the weight parameter information of the first initialization encoder and the weight parameter information of the second initialization encoder, the weight parameter information of the first recurrent neural network and the weight parameter information of the second recurrent neural network are optimized based on the training error parameters, thereby generating a collection session decision network after parameter learning. This collection session decision network after parameter learning can be used to make decisions on new collection session data to determine whether there is a risk fraud event. For example, when new collection session data of user C enters, this collection session decision network can analyze it according to the same mechanism to determine whether there is a risk fraud event in the conversation data of user C.

[0076] Based on the above steps, the embodiments of the present application enhance the key semantic nodes by obtaining and organizing the collection of recovery session big data, generate sample recovery session training data, and then learn parameters of the recovery session decision network, finally generate a completed parameter learning recovery session decision network which can accurately determine whether the recovery session data exists risk fraud event, effectively improve the accuracy and efficiency of the recovery session data analysis, and help to discover and handle potential risk fraud events in time.

[0077] In a possible implementation, step S120 includes:

[0078] Step S121, enhancing each session semantic unit in the session data stream to generate an enhanced session data stream, and the session semantic units in the enhanced session data stream exhibit inconsistency with the session semantic units in the session data stream.

[0079] Step S122, aggregating the enhanced session data stream and the session data stream to generate a fictitious session data stream, and the key semantic nodes of the collection target user in the fictitious session data stream exhibit inconsistency with the key semantic nodes of the collection target user in the session data stream, and the non-key semantic nodes of the collection target user in the fictitious session data stream exhibit consistency with the non-key semantic nodes of the collection target user in the session data stream.

[0080] Step S123, covering the a session data streams into a variety of corresponding fictitious session data streams of the a session data streams to generate the sample recovery session training data.

[0081] In this embodiment, taking the session data stream of the credit card overdue user A mentioned earlier as an example, first, each session semantic unit in the session data stream is enhanced to generate an enhanced session data stream. In the original session data stream, when the customer service communicates with user A, user A says "I have recently spent too much and cannot pay the credit card debt temporarily". The server will first segment the session data stream into independent session semantic units according to semantic integrity. For this session semantic unit, the server identifies the lexical type, labels the key entity such as the debt information "credit card debt", and the semantic role, where user A is the agent and represents the situation of the patient that he cannot pay the debt due to overspending, and generates the labeled information. Then a semantic association knowledge base is constructed. For the noun "overspending", it is associated with "living expenses" and "consumption items", and for the verb "cannot pay", the semantic association with different subject-object combinations is established. Then the core semantic possibility is extracted, which is "overspending leads to inability to pay", and the semantic expansion is performed, expanding "overspending" to "financial strain due to unexpected large medical expenses and other living expenses", and similarly expanding the verb part. Then the semantic conversion operation is performed, such as changing to "the funds become short due to various expenses, making it difficult to repay the credit card debt", and then adjusting according to the prior context, such as the content of the bank repayment reminder message mentioned earlier, adjusting to "although the bank repayment reminder message is received, the funds are short due to various expenses, making it difficult to repay the credit card debt". Finally, the logical relationship is analyzed to ensure coherence and recombination to form an enhanced session data stream. The session semantic units in this enhanced session data stream and the session semantic units in the original session data stream exhibit inconsistency, for example, the expression is more detailed and specific.

[0082] After generating the enhanced conversation data stream, it is aggregated with the original conversation data stream to generate a fabricated conversation data stream. Still taking user A as an example, the server first acquires the key semantic node set corresponding to user A in the original conversation data stream, such as expressions about repayment ability and repayment willingness. For the key semantic node set of repayment ability, the server determines the relevance of each conversation semantic unit in the conversation data stream to the penetration paragraph of the key semantic node set, such as user A mentioning “I have other debts to repay”, which has a high relevance to repayment ability, determines its semantic feature value, and generates a semantic feature field according to the semantic feature values of all conversation semantic units. Then under the limitation of this semantic feature field, the enhanced conversation data stream is optimized to generate a first composite conversation data stream, and the original conversation data stream is optimized to generate a second composite conversation data stream. For example, a conversation semantic unit in the enhanced conversation data stream is multiplied by the semantic feature value to obtain the corresponding part in the first composite conversation data stream, and a conversation semantic unit in the original conversation data stream is fused with the inverse semantic feature value to obtain the part in the second composite conversation data stream. Then the composite conversation data stream corresponding to the semantic feature field is generated according to the first composite conversation data stream and the second composite conversation data stream, and multiple such composite conversation data streams are aggregated to obtain a fabricated conversation data stream. In this fabricated conversation data stream, the key semantic nodes of user A, such as repayment ability related expressions, are inconsistent with the key semantic nodes in the original conversation data stream, while some insignificant non-key semantic nodes, such as the irrelevant life trivia mentioned by user A, remain consistent.

[0083] Finally, the server converts the previously selected a conversation data streams (such as the previously mentioned 100) into their respective fabricated conversation data streams in the above manner, and these fabricated conversation data streams constitute the sample collection training data. This is like creating a special training material library for the collection conversation decision network, and the sample data in the library contains carefully constructed fabricated conversations with specific relationships with the original conversation, so that the network can learn the feature patterns of risk fraud events.

[0084] In a possible implementation, step S122 includes:

[0085] Step S1221, acquiring a plurality of key semantic node sets corresponding to the collection target user in the conversation data stream, each of the key semantic node sets including a plurality of key semantic nodes.

[0086] Step S1221, for any one of the plurality of key semantic node sets, generating a semantic feature field corresponding to the conversation data stream according to the key semantic node set, the semantic feature field representing the penetration paragraph of the key semantic node set in the conversation data stream.

[0087] Step S1222, under the restriction of the various corresponding semantic feature domains of the plurality of key semantic node sets, aggregate the enhanced conversation data stream and the conversation data stream to generate the fictitious conversation data stream.

[0088] In a possible implementation, step S1221 includes:

[0089] Step S1221-1, for any one conversation semantic unit in the conversation data stream, obtain a first correlation degree between the conversation semantic unit and the penetration paragraph where the key semantic node set is located.

[0090] Step S1221-2, based on the first correlation degree, determine a semantic feature value of the conversation semantic unit, the semantic feature value of the conversation semantic unit and the first correlation degree being in a positive correlation relationship.

[0091] Step S1221-3, generate the semantic feature domain according to the semantic feature values of the various conversation semantic units in the conversation data stream.

[0092] In this embodiment, taking the aforementioned credit card overdue user A as an example, the server first needs to obtain a plurality of key semantic node sets corresponding to the conversation data stream of user A. For the collection target user A, one key semantic node set can be about repayment ability, and this set contains multiple key semantic nodes such as "monthly income", "existing savings", "other debt situation", etc. Another key semantic node set can be about repayment willingness, and contains key semantic nodes such as "whether to promise to repay" and "attitude towards overdue".

[0093] Then, the server generates the semantic feature field corresponding to the session data stream according to the key semantic node set about the repayment ability. The server operates on any one of the session semantic units in the session data stream. For example, in a dialogue between a customer service and user A, user A says "I spent all my salary last month on rent, and there is not much left", which is related to the penetration paragraph (i.e. the paragraph about the impact of fund income and expenditure on repayment ability) of the key semantic node set of repayment ability. The server obtains the first correlation between the session semantic unit and the penetration paragraph, and the correlation here is relatively high, assuming 0.8. Based on the first correlation, since the semantic feature value of the session semantic unit is positively correlated with the first correlation, the semantic feature value of the session semantic unit is determined to be 0.8. For another example, user A says "I had bread for breakfast this morning", which has a very low correlation with the penetration paragraph of the key semantic node set of repayment ability, and the correlation is assumed to be 0.1, so the semantic feature value is 0.1. The server generates the semantic feature field according to the semantic feature values of the session semantic units in the session data stream. The semantic feature field is like a quantitative description of the session data stream about the relevant semantics of repayment ability, which represents the penetration paragraph of the key semantic node set of repayment ability in the entire session data stream, i.e. which parts are closely related to the session semantic units in terms of repayment ability, and which parts are loosely related.

[0094] The same operation process is also applied to the key semantic node set of repayment intention. For example, user A says "I know that it is not good to be in arrears, and I will repay as soon as possible", which has a high correlation with the penetration paragraph (paragraph about attitude towards overdue repayment) of the key semantic node set of repayment intention, and the correlation is assumed to be 0.9, so the semantic feature value is 0.9; if user A says "the weather has been bad recently", which has a very low correlation with the penetration paragraph of the key semantic node set of repayment intention, and the correlation is assumed to be 0.05, so the semantic feature value is 0.05, thereby generating the semantic feature field about repayment intention.

[0095] After obtaining the semantic feature domains corresponding to the multiple sets of key semantic nodes (such as the set of repayment ability and the set of repayment willingness), the server needs to aggregate the enhanced conversation data stream and the conversation data stream to generate the fabricated conversation data stream under the restriction of the semantic feature domains. Still taking user A as an example, for the semantic feature domain of repayment ability, the server will optimize the enhanced conversation data stream to generate the first composite conversation data stream and optimize the conversation data stream to generate the second composite conversation data stream according to the semantic feature domain. In the enhanced conversation data stream, for example, a certain conversation semantic unit has a semantic feature value of 0.7 in the previously determined semantic feature domain, and in the optimization process, the conversation semantic unit is adjusted according to the semantic feature value, for example, the semantic weight of the conversation semantic unit is adjusted according to 0.7, so as to generate the corresponding part in the first composite conversation data stream. For a certain conversation semantic unit in the conversation data stream, if the semantic feature value of the conversation semantic unit in the inverse semantic feature domain corresponding to the semantic feature domain of repayment ability is 0.3 (because the sum of the semantic feature values of the same penetration paragraph in the semantic feature domain and the inverse semantic feature domain is 1), the conversation semantic unit is fused according to the semantic feature value of 0.3 to generate the part in the second composite conversation data stream. Then, according to the first composite conversation data stream and the second composite conversation data stream, the composite conversation data stream corresponding to the semantic feature domain of repayment ability is generated.

[0096] The same operation is also applied to the semantic feature domains corresponding to other sets of key semantic nodes such as repayment willingness. Finally, the server aggregates the composite conversation data streams corresponding to all the sets of key semantic nodes to generate the fabricated conversation data stream. In the fabricated conversation data stream, the key semantic nodes (such as the expressions related to repayment ability and repayment willingness) of user A are inconsistent with the key semantic nodes in the original conversation data stream, while some non-key semantic nodes unrelated to repayment ability and repayment willingness, such as some daily trivia mentioned by user A, remain consistent. This completes the generation process from the original conversation data stream and the enhanced conversation data stream to the fabricated conversation data stream, providing a special data structure for subsequent operations, which is helpful for effective analysis and judgment of risk fraud events in the collection scenario.

[0097] In a possible implementation, step S1222 includes:

[0098] Step S1222-1, for any one of the semantic feature domains corresponding to the multiple sets of key semantic nodes, the enhanced conversation data stream is optimized according to the semantic feature domain to generate the first composite conversation data stream, and the conversation data stream is optimized according to the inverse semantic feature domain corresponding to the semantic feature domain to generate the second composite conversation data stream, and the sum of the semantic feature values of the same penetration paragraph in the semantic feature domain and the inverse semantic feature domain is 1.

[0099] Step S1222-2, generating a composite conversation data stream corresponding to the semantic feature domain according to the first composite conversation data stream and the second composite conversation data stream.

[0100] Step S1222-3, aggregating a plurality of composite conversation data streams corresponding to various semantic feature domains to generate the fictitious conversation data stream.

[0101] In a possible implementation, step S1222-1 includes:

[0102] Step S1222-11, multiplying, for any one conversation semantic unit in the enhanced conversation data stream, the conversation semantic unit with a semantic feature value corresponding to the semantic feature domain of the conversation semantic unit to generate the first composite conversation data stream.

[0103] Step S1222-12, fusing, for any one conversation semantic unit in the conversation data stream, the conversation semantic unit with a semantic feature value corresponding to the inverse semantic feature domain of the conversation semantic unit to generate the second composite conversation data stream.

[0104] In this embodiment, taking the case of the credit card overdue user A mentioned above as an example, after the server obtains the semantic feature domains corresponding to a plurality of key semantic node sets (such as repayment ability, repayment willingness, etc.) of the user A, the server starts the aggregation operation.

[0105] For the repayment ability semantic feature domain, the server optimizes the enhanced conversation data stream to generate the first composite conversation data stream according to it, and optimizes the conversation data stream to generate the second composite conversation data stream according to the inverse semantic feature domain corresponding thereto. For example, in the enhanced conversation data stream, there is a conversation semantic unit “I recently have a car loan to be repaid in addition to the credit card arrears, and the pressure is very large”. It is previously determined that the semantic feature value of this conversation semantic unit in the repayment ability semantic feature domain is 0.8. Then the server multiplies this conversation semantic unit with the corresponding 0.8 in this semantic feature domain to obtain an adjusted conversation semantic unit, and the adjusted conversation semantic unit becomes part of the first composite conversation data stream. In this way, the server operates any one conversation semantic unit in the enhanced conversation data stream, and finally generates the first composite conversation data stream.

[0106] For the conversation data stream, for the conversation semantic unit in it, such as one conversation semantic unit is "I originally intended to use the bonus to repay the credit card, but the bonus did not come down", assuming that the semantic feature value of this conversation semantic unit in the inverse semantic feature domain corresponding to the repayment ability semantic feature domain is 0.2 (because in the semantic feature domain and the inverse semantic feature domain, the sum of the semantic feature values of the same penetration paragraph conversation semantic unit is 1). The server will fuse this conversation semantic unit with its corresponding 0.2 in the inverse semantic feature domain. This fusion operation may be to adjust the semantics of this conversation semantic unit according to the weight of 0.2, so as to generate part of the second composite conversation data stream. In this way, the server operates on any one conversation semantic unit in the conversation data stream, and finally generates the second composite conversation data stream.

[0107] Then, according to the generated first composite conversation data stream and the second composite conversation data stream, a composite conversation data stream corresponding to the repayment ability semantic feature domain is generated. This composite conversation data stream integrates the information of the enhanced conversation data stream and the conversation data stream in the repayment ability semantic feature domain.

[0108] The same operation process is also applied to the semantic feature domain corresponding to the repayment willingness key semantic node set. For example, in the enhanced conversation data stream, the semantic feature value of the conversation semantic unit "I know that overdue will affect credit, I want to repay as soon as possible" in the repayment willingness semantic feature domain is 0.9. The adjusted conversation semantic unit obtained by multiplying it with this semantic feature value is used as part of the first composite conversation data stream. In the conversation data stream, the semantic feature value of the conversation semantic unit "I have always repaid on time before" in the inverse semantic feature domain corresponding to the repayment willingness semantic feature domain is 0.1. The part of the second composite conversation data stream obtained by fusing it with 0.1 is then generated. Finally, the composite conversation data stream corresponding to the repayment willingness semantic feature domain is generated.

[0109] Finally, the server aggregates all the composite conversation data streams corresponding to the semantic feature domains (such as repayment ability, repayment willingness, etc.). These composite conversation data streams are combined according to certain rules to generate a fictitious conversation data stream. In this fictitious conversation data stream, the key semantic nodes of user A (such as repayment ability and repayment willingness related expressions) exhibit inconsistency with the key semantic nodes in the original conversation data stream, while some non-key semantic nodes remain consistent. In this way, the generation process from the enhanced conversation data stream and the conversation data stream to the fictitious conversation data stream under the restriction of various semantic feature domains corresponding to multiple key semantic node sets is completed. This fictitious conversation data stream is of great significance for subsequent identification of risk fraud events in the collection scenario.

[0110] In one possible implementation, the step S121 comprises:

[0111] Step S1211, after the session data stream is preliminarily segmented according to semantic integrity to generate a plurality of independent session semantic units, the lexical types in each session semantic unit are identified, and the key entities and semantic roles in each session semantic unit are marked based on the lexical types to generate marking information of each session semantic unit, the key entities include debt information related to collection, financial products involved, identity information of the collection target user, and domain words related to risk fraud, and the semantic roles include agent and recipient.

[0112] Step S1212, based on the marking information of each session semantic unit, a corresponding semantic association knowledge base is constructed, the construction of the semantic association knowledge base is obtained by analyzing the collection-related corpus and general corpus, and specifically for the nouns in each session semantic unit, the association relationship with other related nouns is established in the semantic association knowledge base, and for the verbs in each session semantic unit, the semantic association under different combinations of subjects and objects is established, and for adjectives and adverbs, the semantic association between the modifiers and the modified objects and the semantic change association under different contexts are established.

[0113] Step S1213, for each session semantic unit, the core semantics of the session semantic unit are extracted according to the semantic association knowledge base to obtain a session semantic unit set containing core semantics.

[0114] Step S1214, for each session semantic unit containing core semantics, semantic expansion is performed according to the semantic association knowledge base, wherein for the noun part in the core semantics, other nouns, adjectives and verbs related to the noun are searched from the semantic association knowledge base and then expanded into the session semantic unit, and for the verb part in the core semantics, synonymous words, near-synonymous words and different tenses and voices are searched and then expanded into the session semantic unit to obtain a set of session semantic units after semantic expansion.

[0115] Step S1215, based on the set of session semantic units after semantic expansion, a semantic conversion operation is performed, part of the semantics in the session semantic units is converted according to the semantic association knowledge base, and a set of session semantic units after semantic conversion is obtained.

[0116] Step S1216, based on the prior context information of the collection session, the set of session semantic units after semantic conversion is adaptively adjusted in the context to obtain a set of session semantic units after context adaptation.

[0117] Step S1217, analyze the logical relationship between each of the conversation semantic units in the context-adapted conversation semantic unit set, determine whether there is a semantic jump and semantic contradiction between adjacent conversation semantic units, if so, adjust the relevant conversation semantic units, and obtain a logically coherent conversation semantic unit set.

[0118] Step S1218, recombine the logically coherent conversation semantic units in the original implementation conversation data stream to form an enhanced conversation data stream.

[0119] In this embodiment, taking the conversation data stream of the aforementioned credit card overdue user A as an example, the server first performs preliminary segmentation on this conversation data stream according to semantic integrity, thereby generating multiple independent conversation semantic units. For example, in a dialogue between the customer service and user A, “I am overdue because I spent a lot of money on hospitalization due to illness” will be segmented into a conversation semantic unit. Then, the server starts to identify the lexical types in this conversation semantic unit, such as “sick and hospitalized” is a noun phrase, “spent” is a verb, “a lot of money” is a noun phrase, “credit card” is a noun, and “overdue” is a verb. Based on these lexical types, the server labels the key entities and semantic roles. Among them, the key entity includes “credit card”, which is the debt information related to debt collection, and “I (user A)” is the agent, indicating that the action of “sick and hospitalized and spent money” is performed, thereby leading to the result of “credit card overdue”. In this way, the labeling information of this conversation semantic unit is generated.

[0120] For the labeling information of each conversation semantic unit, the server will construct a corresponding semantic association knowledge base. This knowledge base is obtained by analyzing the corpus related to debt collection and general corpus. For the noun “sick and hospitalized” in this conversation semantic unit, an association relationship with other related nouns is established in the semantic association knowledge base, such as “medical expenses” “medical insurance” and the like; for the verb “spent”, its semantic association with different subjects and objects is established, such as “he spent savings on travel” and the like; for the adjective “a lot”, its semantic association with the modified object “money” is established, as well as the semantic change association under different contexts, such as the amount range represented by “a lot” may be different when describing poor families and rich families.

[0121] Then, according to the semantic association knowledge base, the server extracts the core semantics of each session semantic unit. For the above session semantic unit, the core semantics can be "hospitalization expenses caused credit card overdue". Thus, the session semantic unit set containing the core semantics is obtained. Then, semantic expansion is performed for each session semantic unit containing the core semantics. For the noun part "hospitalization" in the core semantics, other nouns, adjectives and verbs related to it are searched from the semantic association knowledge base and expanded into the session semantic unit, for example, expanded to "because of the sudden serious illness, a large amount of medical expenses was spent"; for the verb part "spent" in the core semantics, synonyms, near-synonyms and different tenses and voices are searched and expanded into the session semantic unit, such as "consumed", thereby obtaining the session semantic unit set after semantic expansion.

[0122] On the basis of the session semantic unit set after semantic expansion, the server performs semantic conversion operation according to the semantic association knowledge base. For example, "spent a large amount of medical expenses" is converted to "generated high medical expenses", thereby obtaining the session semantic unit set after semantic conversion. Then, context adaptability adjustment is performed on the session semantic unit set after semantic conversion based on the prior context information of the collection session. If the previous conversation mentioned that the bank has a special policy for overdue due to illness, the session semantic unit can be adjusted to "although it is known that the bank has a special policy for overdue due to illness, because of the sudden serious illness, high medical expenses was generated, so the credit card is overdue", thereby obtaining the session semantic unit set after context adaptation.

[0123] The session semantic unit set after context adaptation is analyzed again to determine the logical relationship between each session semantic unit. For example, in this set, if one session semantic unit says "I have no money to pay the credit card", the next session semantic unit says "I just bought a very expensive computer", there is a semantic jump and semantic contradiction. The server will adjust the related session semantic units, such as adjusting the latter session semantic unit to "during my illness, my family borrowed money to buy a very expensive computer for my convenience, so I have no money to pay the credit card", thereby obtaining the session semantic unit set with logical coherence.

[0124] Finally, these logically coherent session semantic units are recombined in the order of the original session data stream. For example, the original session data stream first mentions hospitalization due to illness, and then mentions credit card overdue, so when recombining, the order is also followed, thus forming the enhanced session data stream. The session semantic units in the enhanced session data stream are processed and optimized in many aspects compared to the session semantic units in the original session data stream, and are more rich, accurate and logically coherent in semantic expression, thereby providing a better data basis for further operation.

[0125] In a possible implementation, step S130 comprises:

[0126] Step S131, obtaining, by the collection session decision network, a first mapping knowledge vector of the sample collection session training data, the first mapping knowledge vector representing semantic transition information of a corresponding semantic representation vector of the collection target user in the sample collection session training data in a time domain.

[0127] Step S132, generating, by the collection session decision network, a second mapping knowledge vector according to the first mapping knowledge vector, the second mapping knowledge vector representing semantic transition information of the corresponding semantic representation vector of the collection target user in the sample collection session training data in a space domain.

[0128] Step S133, generating, by the collection session decision network, a decision result of the sample collection session training data according to the second mapping knowledge vector, the decision result of the sample collection session training data representing whether a semantic change of a semantic representation vector associated with a key semantic node of the collection target user in the sample collection session training data exists a risk fraud event.

[0129] Step S134, performing parameter learning on the collection session decision network based on the decision result, to generate the collection session decision network after the parameter learning.

[0130] In this embodiment, taking the sample collection session training data related to the credit card overdue user A as an example, the server first obtains a first mapping knowledge vector of the sample collection session training data by using the collection session decision network. In the conversation data of user A, there is semantic transition information of the semantic representation vector in the time domain as the conversation proceeds. For example, at the beginning of the conversation, user A mentions that the credit card is overdue because of the huge cost of hospitalization due to illness, and the semantic representation vector at this time reflects a situation of difficulty in repayment due to unexpected expenses. As the conversation develops, user A also indicates that he is trying to raise funds to repay the loan, which is a change of the semantic representation vector in the time domain. The collection session decision network encodes the semantic transition information in the time domain to form the first mapping knowledge vector, and the first mapping knowledge vector can accurately represent the semantic transition of the corresponding semantic representation vector of user A in the entire sample collection session training data in the time domain.

[0131] Then, the server generates a second mapping knowledge vector according to the first mapping knowledge vector by using the collection session decision network. In this process, the network processes the first mapping knowledge vector to obtain the second mapping knowledge vector reflecting the semantic transition information of the semantic representation vector in the space. For example, for the conversation data of user A, the information in the first mapping knowledge vector can be arranged in the order of the occurrence of the dialogue, and after the processing of the network, the information is rearranged in the space according to the importance of the semantics and other factors. For example, part of the information about the repayment ability related expressions, such as income source, debt situation and other semantic information, is rearranged in the space according to the importance of judging the risk fraud event, thereby generating the second mapping knowledge vector, which can represent the semantic transition information of the semantic representation vector corresponding to user A in the sample collection session training data in the space.

[0132] Then, the server generates a second mapping knowledge vector according to the first mapping knowledge vector by using the collection session decision network. In this process, the network processes the first mapping knowledge vector to obtain the second mapping knowledge vector reflecting the semantic transition information of the semantic representation vector in the space. For example, for the conversation data of user A, the information in the first mapping knowledge vector can be arranged in the order of the occurrence of the dialogue, and after the processing of the network, the information is rearranged in the space according to the importance of the semantics and other factors. For example, part of the information about the repayment ability related expressions, such as income source, debt situation and other semantic information, is rearranged in the space according to the importance of judging the risk fraud event, thereby generating the second mapping knowledge vector, which can represent the semantic transition information of the semantic representation vector corresponding to user A in the sample collection session training data in the space.

[0133] Finally, based on the decision result, the parameter learning of the collection session decision network is performed. If the decision result determines that there is a risk fraud event, but in fact it is a false positive, or vice versa, a training error will be generated. The server determines the training error parameter of the collection session decision network according to the decision result. For example, if the judgment of user A is that there is a risk fraud event, but in fact user A is just not clear in expression, which is a false positive, and a certain training error is generated. Then, the server locks the weight parameter information of the first initialization encoder and the weight parameter information of the second initialization encoder in the collection session decision network, and optimizes the weight parameter information of the first recurrent neural network and the weight parameter information of the second recurrent neural network based on the training error parameter. Through such an optimization process, the network can better adapt to the data and continuously adjust its parameters to improve the accuracy of the judgment of risk fraud events, and finally generate a collection session decision network with completed parameter learning. The collection session decision network with completed parameter learning can more accurately analyze new collection session data and determine whether there is a risk fraud event.

[0134] In a possible implementation, step S132 includes:

[0135] Step S1321, performing attention weight allocation processing on the first mapping knowledge vector to generate a first mapping knowledge vector after attention weight allocation processing.

[0136] Step S1322, using the collection session decision network to generate the second mapping knowledge vector according to the first mapping knowledge vector after attention weight allocation processing.

[0137] In a possible implementation, the knowledge segments in the first mapping knowledge vector are arranged in the time domain, and the knowledge segments in the first mapping knowledge vector after attention weight allocation processing are arranged in the space domain.

[0138] In a possible implementation, the collection session decision network includes a time domain knowledge encoder and a space domain knowledge encoder, the time domain knowledge encoder is generated according to a first initialization encoder and a first recurrent neural network, and the space domain knowledge encoder is generated according to a second initialization encoder and a second recurrent neural network, the first recurrent neural network is used to obtain the first mapping knowledge vector according to the first initialization encoder, and the second recurrent neural network is used to obtain the second mapping knowledge vector according to the second initialization encoder.

[0139] Step S134 includes:

[0140] Step S1341, based on the decision result, determining a training error parameter of the collection session decision network.

[0141] Step S1342, lock the weight parameter information of the first initialization encoder and the weight parameter information of the second initialization encoder, optimize the weight parameter information of the first recurrent neural network and the weight parameter information of the second recurrent neural network based on the training error parameter, and generate the completed parameter learning collection session decision network.

[0142] In this embodiment, taking the collection session data of the credit card overdue user A mentioned earlier as an example, when the server generates the second mapping knowledge vector according to the first mapping knowledge vector by using the collection session decision network, the first mapping knowledge vector needs to be processed by attention weight distribution first. In the collection session of user A, the first mapping knowledge vector contains knowledge segments arranged in the time domain, which reflect the semantic change information of user A in the whole session process. For example, in the early stage of the session, user A mentioned that he spent a lot of money because of hospitalization due to illness, which led to the credit card overdue, and this information exists as a knowledge segment in the first mapping knowledge vector; as the session progresses, user A indicates that he is trying to repay the credit card debt by borrowing from relatives, which is another knowledge segment. These knowledge segments are arranged in the order of the occurrence of the dialogue, that is, in the time domain.

[0143] When the server performs attention weight distribution processing, it will distribute weights according to the importance of each knowledge segment in judging risk fraud events. For example, for knowledge segments directly related to repayment ability, such as income source, debt situation, etc., a higher attention weight is assigned. For relatively unimportant information such as user A's mention of the hospital environment during the hospitalization period, a lower weight is assigned. Through such weight distribution, the first mapping knowledge vector after attention weight distribution processing is generated, and the knowledge segments in it are arranged in the spatial domain at this time. This transformation from time domain arrangement to spatial domain arrangement is the result of reorganizing knowledge segments based on attention weights. For example, all knowledge segments related to repayment ability and with higher weights will be placed together, while knowledge segments not closely related to repayment ability will be placed in other positions.

[0144] Then, the first mapping knowledge vector processed by the attention weight distribution is used to generate a second mapping knowledge vector by the collection conversation decision network. The time domain knowledge encoder in the collection conversation decision network is generated based on the first initialization encoder and the first recurrent neural network, and the first recurrent neural network obtains the first mapping knowledge vector based on the first initialization encoder. The space domain knowledge encoder is generated based on the second initialization encoder and the second recurrent neural network. In this process, the second recurrent neural network takes the first mapping knowledge vector processed by the attention weight distribution as input, and generates the second mapping knowledge vector through a series of complex calculations and processing based on the second initialization encoder. This second mapping knowledge vector represents the semantic transformation information of the semantic representation vector corresponding to user A in the example collection conversation training data in the space domain. For example, in the space domain, the relationship between different semantic related knowledge fragments is reconstructed, which can more clearly reflect the change of semantics in space, which helps to more comprehensively analyze whether the conversation data of user A exists risk fraud event.

[0145] Next, when the collection conversation decision network is parameter learned based on the decision result to generate a collection conversation decision network with completed parameter learning, the server first determines the training error parameter of the collection conversation decision network based on the decision result. For the collection conversation decision result of user A, if it is determined that there is a risk fraud event, but further investigation finds that user A only expresses unclearly due to some misunderstanding in communication, and actually there is no fraud behavior, which is a misjudgment, thus generating a training error. Conversely, if it is determined that there is no risk fraud event, but in fact user A has fraud behavior, which is also a misjudgment.

[0146] After determining the training error parameter, the server will lock the weight parameter information of the first initialization encoder and the weight parameter information of the second initialization encoder. Because the weight parameters of the first initialization encoder and the second initialization encoder have certain stability in the network structure or have been optimized and determined in the early stage. Then, the weight parameter information of the first recurrent neural network and the weight parameter information of the second recurrent neural network are optimized based on the training error parameter. When optimizing the weight parameter information of the first recurrent neural network, the weights related to the processing of the first mapping knowledge vector will be adjusted according to the training error parameter. For example, if a certain weight causes deviation in the processing of the repayment ability related knowledge fragment of user A, thereby affecting the final decision result, then the weight will be adjusted according to the training error. The optimization of the weight parameter information of the second recurrent neural network is also a similar process, and the weights related to the generation of the second mapping knowledge vector are adjusted according to the training error parameter, so as to improve the accuracy of processing the semantic transformation information in the space domain.

[0147] Through this process of continuously adjusting the weight parameter information of the first recurrent neural network and the second recurrent neural network based on the decision result, the server gradually optimizes the collection session decision network, and finally generates a collection session decision network that has completed parameter learning. This collection session decision network that has completed parameter learning can more accurately determine whether there is a risk fraud event in the new collection session data when processing the new collection session data. For example, when processing the collection session data of the new credit card overdue user B, this network can more accurately analyze the session semantics of user B to determine whether there is a risk fraud behavior, reduce the possibility of misjudgment, and thus improve the efficiency and accuracy of the entire collection process. In actual application, as more and more sample collection session training data is used for training, the collection session decision network will be continuously optimized and can adapt to the risk fraud judgment needs in various complex collection scenarios.

[0148] Further illustrated with the case of user A, suppose that in a collection session, user A mentions that he is sick in hospital and spends a lot of money, has other debts in addition to credit card arrears, and is currently unemployed, and these information constitutes part of the knowledge fragments in the first mapping knowledge vector in the order of time domain. When performing attention weight allocation processing, the knowledge fragments such as "unemployed" and "has other debts" directly related to repayment ability are allocated higher weights, because these factors are crucial to determine whether user A has a risk fraud behavior. Some details in "sick in hospital and spend a lot of money", such as the name of the hospital, the level of the ward, and other relatively unimportant information are allocated lower weights. After the attention weight allocation processing, these knowledge fragments are rearranged in the spatial domain. Then, the collection session decision network generates a second mapping knowledge vector according to the processed first mapping knowledge vector, at this time the semantic transformation information in the spatial domain is more clear. For example, it can clearly see the relationship changes between the knowledge fragments related to repayment ability, and their association with other information.

[0149] In the parameter learning based on the decision result, if it is initially determined that user A has a risk fraud event, because the user A's attitude and ability to repay in the conversation is somewhat ambiguous, but it is later found that it is because the user A has just recovered from an illness and his thinking is not clear enough, so it is not actually a fraud. This misjudgment will produce a training error. The server optimizes the weight parameter information of the first recurrent neural network and the second recurrent neural network based on locking the weight parameter information of the first initialization encoder and the second initialization encoder. For the first recurrent neural network, the weight related to processing the knowledge fragment of the user A's repayment ability may be adjusted to more accurately reflect the actual situation in subsequent processing of similar situations. For the second recurrent neural network, the weight related to constructing the spatial semantic relationship may be adjusted to better process the semantic transformation information in the space and avoid similar misjudgments from occurring again. Through such continuous optimization, the collection conversation decision network will be more accurate and reliable in processing similar collection conversations.

[0150] In the entire collection business of the financial institution, there is a large amount of collection conversation data to be processed. Different users have different situations, for example, some users may be due to sudden economic difficulties, such as unemployment, investment failure, etc., leading to credit card overdue, while some users may have malicious fraud behavior, intentionally defaulting on the debt. The collection conversation decision network needs to accurately distinguish between these situations, and through the above parameter learning process, the network can continuously improve its accuracy. Taking another user C as an example, user C is due to investment failure and credit card overdue. In the collection conversation, user C mentions information such as his investment project, loss amount, current income source, etc. The server processes user C's collection conversation data according to the same process, from obtaining the first mapping knowledge vector, performing attention weight distribution processing to generate the second mapping knowledge vector, to parameter learning based on the decision result. In this process, each link is crucial to accurately determine whether user C has a risk fraud behavior. For example, if a misjudgment is made due to insufficient understanding of user C's investment project, then the problem needs to be accurately identified and adjusted when optimizing the network weight parameters to ensure that the network can make more accurate judgments when processing similar collection conversations of users.

[0151] As more different types of collection conversation data are used for training, the collection conversation decision network will continuously learn new patterns and features, and be able to better adapt to various complex collection scenarios. Whether it is facing a user who is overdue due to unexpected expenses, or a user who may have fraud behavior, the collection conversation decision network that has completed parameter learning can accurately analyze the conversation semantics to determine whether there is a risk fraud event, thereby providing strong support for the collection work of the financial institution, improving the collection efficiency, and reducing the risk loss.

[0152] Figure 2 The hardware structure of the information decision-making system 100 based on the collection session big data analysis for implementing the information decision-making method based on the collection session big data analysis is shown, and the information decision-making system 100 based on the collection session big data analysis is shown in Figure 1. Figure 2 As shown in Figure 1, the information decision-making system 100 based on the collection session big data analysis can include a processor 110, a machine readable storage medium 120, a bus 130, and a communication unit 140.

[0153] The machine readable storage medium 120 can store data and / or instructions. In some embodiments, the machine readable storage medium 120 can store data acquired from an external terminal. In some embodiments, the machine readable storage medium 120 can store data and / or instructions used by the information decision-making system 100 based on the collection session big data analysis to perform or use to complete the exemplary methods described in the present disclosure.

[0154] In the implementation process, the processor 110 executes the computer executable instructions stored in the machine readable storage medium 120, so that the processor 110 can perform the information decision-making method based on the collection session big data analysis of the method embodiments described above. The processor 110, the machine readable storage medium 120, and the communication unit 140 are connected through the bus 130, and the processor 110 can be used to control the transceiving action of the communication unit 140.

[0155] The implementation process of the processor 110 can refer to the various method embodiments executed by the information decision-making system 100 based on the collection session big data analysis described above, which has similar implementation principles and technical effects, and will not be described here.

[0156] In addition, the present disclosure also provides a readable storage medium, wherein computer executable instructions are preset in the readable storage medium, and when the processor executes the computer executable instructions, the information decision-making method based on the collection session big data analysis described above is implemented.

[0157] It should be noted that, in order to simplify the description of the present disclosure and help understand one or more embodiments of the present disclosure, in the foregoing description of the embodiments of the present disclosure, various features are sometimes combined into one embodiment, figure or description thereof.

Claims

1. An information decision-making method based on collection session big data analysis, characterized in that, The method comprises: obtaining collection session big data, wherein the collection session big data comprises a plurality of session data streams associated with a collection target user arranged according to a collection session node sequence; enhancing a key semantic node of the collection target user in a session data stream to generate sample collection session training data, wherein a semantic representation vector associated with the key semantic node of the collection target user in the sample collection session training data exists a risk fraud event, and a is a positive integer; performing parameter learning on a collection session decision network according to the sample collection session training data to generate a collection session decision network after parameter learning, wherein the collection session decision network after parameter learning is used to determine whether collection session data exists a risk fraud event; the enhancement of the key semantic node of the collection target user in the session data stream to generate the sample collection session training data comprises: enhancing each session semantic unit in the session data stream to generate an enhanced session data stream, wherein a session semantic unit in the enhanced session data stream exhibits inconsistency with a session semantic unit in the session data stream; aggregating the enhanced session data stream and the session data stream to generate a fictitious session data stream, wherein a key semantic node of the collection target user in the fictitious session data stream exhibits inconsistency with a key semantic node of the collection target user in the session data stream, and a non-key semantic node of the collection target user in the fictitious session data stream exhibits consistency with a non-key semantic node of the collection target user in the session data stream; covering the a session data streams into a plurality of corresponding fictitious session data streams of the a session data streams to generate the sample collection session training data; the enhancement of each session semantic unit in the session data stream to generate the enhanced session data stream comprises: after the session data stream is preliminarily segmented according to semantic integrity to generate a plurality of independent session semantic units, identifying a vocabulary type in each session semantic unit and marking a key entity and a semantic role in each session semantic unit based on the vocabulary type to generate marking information of each session semantic unit, wherein the key entity comprises debt information related to collection, a financial product involved, identity information of the collection target user, and domain words related to risk fraud, and the semantic role comprises an agent and a recipient; based on the marking information of each session semantic unit, a corresponding semantic association knowledge base is constructed, wherein the semantic association knowledge base is obtained by analyzing collection-related corpus and general corpus, and specifically, for a noun in each session semantic unit, an association relationship with other related nouns is established in the semantic association knowledge base, and for a verb in each session semantic unit, semantic associations under different combinations of subject and object are established, and for an adjective and an adverb, semantic associations with modified objects and semantic change associations under different contexts are established. For each session semantic unit, a core semantic of the session semantic unit is extracted according to the semantic association knowledge base, to obtain a set of session semantic units containing core semantics; For each session semantic unit containing core semantics, semantic expansion is performed according to the semantic association knowledge base, wherein for a noun part in the core semantics, other nouns, adjectives and verbs related to the noun part are searched from the semantic association knowledge base and expanded into the session semantic unit, and for a verb part in the core semantics, synonyms, near-synonyms and different tenses and voices of the verb part are searched and expanded into the session semantic unit, to obtain a set of session semantic units after semantic expansion; On the basis of the set of session semantic units after semantic expansion, a semantic conversion operation is performed, part of the semantics in the session semantic units is converted according to the semantic association knowledge base, to obtain a set of session semantic units after semantic conversion; Based on prior context information of the collection session, context adaptability adjustment is performed on the set of session semantic units after semantic conversion, to obtain a set of session semantic units after context adaptation; Logical relationships between each session semantic unit in the set of session semantic units after context adaptation are analyzed, whether there is semantic jump and semantic contradiction between adjacent session semantic units is judged, if there is, the related session semantic units are adjusted, to obtain a set of logically coherent session semantic units; The logically coherent session semantic units are recombined according to the order in the original implementation session data stream, to form an enhanced session data stream. 2.The information decision-making method based on collection session big data analysis according to claim 1, characterized in that, The enhanced session data stream and the session data stream are aggregated to generate a fictitious session data stream, including: A plurality of key semantic node sets corresponding to the collection target user in the session data stream are obtained, each key semantic node set including a plurality of key semantic nodes; For any one key semantic node set in the plurality of key semantic node sets, a semantic feature domain corresponding to the session data stream is generated according to the key semantic node set, the semantic feature domain representing a penetration paragraph of the key semantic node set in the session data stream; Under the restriction of the various corresponding semantic feature domains of the plurality of key semantic node sets, the enhanced session data stream and the session data stream are aggregated to generate the fictitious session data stream. 3.The information decision-making method based on collection session big data analysis according to claim 2, characterized in that, The generating of the semantic feature domain corresponding to the session data stream according to the key semantic node set includes: For any one session semantic unit in the session data stream, a first correlation degree between the session semantic unit and the penetration paragraph of the key semantic node set is obtained; Based on the first correlation degree, a semantic feature value of the session semantic unit is determined, the semantic feature value of the session semantic unit being in a positive correlation relationship with the first correlation degree; According to the semantic feature values of each session semantic unit in the session data stream, the semantic feature domain is generated. 4.The information decision-making method based on collection session big data analysis according to claim 3, characterized in that, The aggregation of the enhanced session data stream and the session data stream under the restriction of the various corresponding semantic feature domains of the plurality of key semantic node sets to generate the fictitious session data stream includes: According to any one semantic feature domain of various corresponding semantic feature domains of the plurality of key semantic node sets, the enhanced conversation data stream is optimized according to the semantic feature domain to generate a first composite conversation data stream, and the conversation data stream is optimized according to the inverse semantic feature domain corresponding to the semantic feature domain to generate a second composite conversation data stream, wherein the sum of the semantic feature values of the same penetrating paragraph conversation semantic units in the semantic feature domain and the inverse semantic feature domain is 1; According to the first composite conversation data stream and the second composite conversation data stream, a composite conversation data stream corresponding to the semantic feature domain is generated; Aggregating various corresponding composite conversation data streams of the semantic feature domains generates the fictitious conversation data stream.

5. The information decision making method based on collection session big data analysis according to claim 4, characterized in that, According to any one semantic feature domain of various corresponding semantic feature domains of the plurality of key semantic node sets, the enhanced conversation data stream is optimized according to the semantic feature domain to generate a first composite conversation data stream, and the conversation data stream is optimized according to the inverse semantic feature domain corresponding to the semantic feature domain to generate a second composite conversation data stream, wherein the sum of the semantic feature values of the same penetrating paragraph conversation semantic units in the semantic feature domain and the inverse semantic feature domain is 1; According to any one semantic feature domain of various corresponding semantic feature domains of the plurality of key semantic node sets, the enhanced conversation data stream is optimized according to the semantic feature domain to generate a first composite conversation data stream, and the conversation data stream is optimized according to the inverse semantic feature domain corresponding to the semantic feature domain to generate a second composite conversation data stream, wherein the sum of the semantic feature values of the same penetrating paragraph conversation semantic units in the semantic feature domain and the inverse semantic feature domain is 1; According to the sample collection conversation training data, the collection conversation decision network is parameter learned to generate a collection conversation decision network with completed parameter learning, including: 6.The information decision-making method based on collection session big data analysis of claim 1, wherein, Using the collection conversation decision network to obtain a first mapping knowledge vector of the sample collection conversation training data, the first mapping knowledge vector representing the semantic change information of the semantic representation vector corresponding to the collection target user in the sample collection conversation training data in the time domain; Using the collection conversation decision network to generate a second mapping knowledge vector according to the first mapping knowledge vector, the second mapping knowledge vector representing the semantic change information of the semantic representation vector corresponding to the collection target user in the sample collection conversation training data in the spatial domain; Using the collection conversation decision network to generate a decision result of the sample collection conversation training data according to the second mapping knowledge vector, the decision result of the sample collection conversation training data representing whether the semantic change of the semantic representation vector associated with the key semantic node of the collection target user in the sample collection conversation training data exists a risk fraud event; Based on the decision result, the collection conversation decision network is parameter learned to generate the collection conversation decision network with completed parameter learning. The collection conversation decision network generates a second mapping knowledge vector according to the first mapping knowledge vector, including: 7.The information decision-making method based on collection session big data analysis of claim 6, characterized in that, The first mapping knowledge vector is subjected to attention weight allocation processing to generate a first mapping knowledge vector after attention weight allocation processing; ​ The collection session decision network is used to generate the second mapping knowledge vector according to the first mapping knowledge vector after the attention weight distribution processing; The knowledge segments in the first mapping knowledge vector are arranged in a time domain, and the knowledge segments in the first mapping knowledge vector after the attention weight distribution processing are arranged in a space domain; The collection session decision network comprises a time domain knowledge encoder and a space domain knowledge encoder, the time domain knowledge encoder is generated according to a first initialization encoder and a first recurrent neural network, the space domain knowledge encoder is generated according to a second initialization encoder and a second recurrent neural network, the first recurrent neural network is used to obtain the first mapping knowledge vector according to the first initialization encoder, and the second recurrent neural network is used to obtain the second mapping knowledge vector according to the second initialization encoder; The parameter learning of the collection session decision network based on the decision result comprises: determining a training error parameter of the collection session decision network based on the decision result; locking weight parameter information of the first initialization encoder and weight parameter information of the second initialization encoder, optimizing weight parameter information of the first recurrent neural network and weight parameter information of the second recurrent neural network based on the training error parameter, and generating the collection session decision network after the parameter learning.

8. An information decision system based on collection session big data analysis, characterized in that, The information decision system based on collection session big data analysis comprises a processor and a memory, the memory is connected with the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to realize the information decision method based on collection session big data analysis in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Fraudulent behavior identification method and system applied to machine learning

    CN116542673A

  • Feature recognition model training method and ship and cargo information recognition method

    CN117422072A

  • Repayment willingness prediction method, device and equipment and readable storage medium

    CN117574232A

  • Artificial intelligence-based fraudulent behavior analysis method and digital financial big data system

    CN117575596A