Customer identification method and apparatus, electronic device, and readable storage medium
Patent Information
- Application Number
- CN202311301759.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-08
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2043-10-08
AI Technical Summary
[0004]本申请的主要目的在于提供一种客户识别方法、装置、电子设备及可读存储介质,旨在解决现有的客户识别方式的识别准确率低的技术问题
[0040] This application provides a customer identification method. First, the application obtains target text information associated with the target customer and target business scenario information corresponding to the target text information. Then, based on the target text information and the target business scenario information, a target customer identification model is constructed. Finally, by inputting the behavioral information of the customer to be identified into the target customer identification model, the identification of the customer to be identified can be completed, and the identification result can be obtained.
Smart Images

Figure CN117252208B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data identification technology, and in particular to a customer identification method, apparatus, electronic device, and readable storage medium. Background Technology
[0002] Customer identification refers to processing customer descriptions and characteristics through rules, models, and other methods to determine whether a customer meets the definition of a target customer group and output the identification result of whether the customer belongs to the target customer group.
[0003] Currently, customers are typically identified using textual information that describes them or represents their characteristics. However, since the same textual information may correspond to different semantics in different business scenarios, identifying customers solely through textual information may lead to misjudgments due to semantic differences. Therefore, the accuracy of existing customer identification methods is low. Summary of the Invention
[0004] The main objective of this application is to provide a customer identification method, apparatus, electronic device, and readable storage medium, which aims to solve the technical problem of low identification accuracy in existing customer identification methods.
[0005] To achieve the above objectives, this application provides a customer identification method, the customer identification method comprising:
[0006] Obtain target text information associated with the target customer, and obtain target business scenario information corresponding to the target text information;
[0007] Based on the target text information and the target business scenario information, a target customer identification model is constructed;
[0008] By inputting the behavioral information of the customer to be identified into the target customer identification model, the identification result of the customer to be identified is obtained.
[0009] Optionally, the step of obtaining target text information associated with the target customer includes:
[0010] The obtained text information set is processed by text segmentation to obtain the initial text information.
[0011] The relevance of each initial textual piece of information to the target customer is determined by chi-square verification and information value detection.
[0012] Initial text information with a relevance greater than a preset relevance threshold is used as the target text information associated with the target customer.
[0013] Optionally, the step of obtaining the target business scenario information corresponding to the target text information includes:
[0014] Extract keywords from the target text information;
[0015] Using the keywords as an index, the target business scenario information corresponding to the target text information is searched in the preset business scenario information configuration table.
[0016] Optionally, the step of constructing a target customer identification model based on the target text information and the target business scenario information includes:
[0017] Based on the target text information and the target business scenario information, a text enhancement feature set is constructed using a text enhancement function;
[0018] The text enhancement feature set is divided into a training sample set and a validation sample set;
[0019] Training samples are selected from the training sample set, and the samples to be identified are determined based on the training samples and their corresponding sample weights.
[0020] The sample to be identified is input into the customer identification model to be trained to obtain the sample identification result;
[0021] Based on the verification sample set and the sample recognition results, optimize the customer recognition model to be trained;
[0022] Return to the steps of selecting training samples from the training sample set and determining the samples to be identified based on the training samples and their corresponding sample weights, until the customer identification model to be trained meets the preset iterative training termination condition, and obtain the target customer identification model.
[0023] Optionally, the step of optimizing the customer identification model to be trained based on the verification sample set and the sample identification results includes:
[0024] Obtain the true results corresponding to the training samples from the validation sample set;
[0025] Based on the difference between the actual results and the sample identification results, a loss function is constructed, and it is verified whether the loss function is a convergent function.
[0026] If not, then optimize the customer identification model to be trained based on the calculated gradient of the loss function. Optionally, the step of constructing the target customer identification model based on the target text information and the target business scenario information includes:
[0027] Based on the target text information and the target business scenario information, a first customer identification model is constructed, and based on the target text information, a second customer identification model is constructed.
[0028] The target customer identification model is obtained by merging the first customer identification model and the second customer identification model.
[0029] Optionally, the target customer identification model is formed by fusing a first customer identification model and a second customer identification model. The step of obtaining the identification result of the customer to be identified by inputting the behavioral information of the customer to be identified into the target customer identification model includes:
[0030] After inputting the behavioral information of the customer to be identified into the target customer identification model, if the behavioral information carries business scenario information, then the first customer identification model is used as the model to be solved.
[0031] If the behavioral information does not carry business scenario information, then the second customer identification model will be used as the model to be solved.
[0032] Based on a preset solution algorithm, the optimal solution of the model to be solved is determined, and the optimal solution is used as the identification result of the customer to be identified.
[0033] This application also provides a customer identification device, the customer identification device comprising:
[0034] The acquisition module is used to acquire target text information associated with the target customer, and to acquire target business scenario information corresponding to the target text information;
[0035] The construction module is used to construct a target customer identification model based on the target text information and the target business scenario information;
[0036] The identification module is used to obtain the identification result of the customer to be identified by inputting the behavioral information of the customer to be identified into the target customer identification model.
[0037] This application also provides an electronic device, which is a physical device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the customer identification method as described above.
[0038] This application also provides a readable storage medium, which is a computer-readable storage medium, on which a program implementing a customer identification method is stored. The program implementing the customer identification method is executed by a processor to implement the steps of the customer identification method as described above.
[0039] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the customer identification method described above.
[0040] This application provides a customer identification method. First, the application obtains target text information associated with the target customer and target business scenario information corresponding to the target text information. Then, based on the target text information and the target business scenario information, a target customer identification model is constructed. Finally, by inputting the behavioral information of the customer to be identified into the target customer identification model, the identification of the customer to be identified can be completed, and the identification result can be obtained.
[0041] Therefore, this application determines the customer identification model by combining textual information and business scenario information. This allows the customer identification model to be used to assist in customer identification when it is used for customer identification. As a result, the customer identification model can distinguish the semantics of the same textual information in different business scenarios, which solves the technical problem of low identification accuracy in existing customer identification methods and improves the accuracy of customer identification. Attached Figure Description
[0042] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0043] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 This is a flowchart illustrating an embodiment of the customer identification method in this application.
[0045] Figure 2 This is a flowchart illustrating Embodiment 2 of the customer identification method in this application.
[0046] Figure 3 This is a simplified flowchart illustrating the construction of a target customer identification model provided in Embodiment 2 of this application;
[0047] Figure 4This is an interactive schematic diagram of the target customer identification model provided in Embodiment 2 of this application;
[0048] Figure 5 This is a schematic diagram of the module structure of the customer identification device according to an embodiment of this application;
[0049] Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the customer identification method in this application embodiment.
[0050] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0051] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0052] Example 1
[0053] Customer identification refers to processing customer descriptions and characteristics through rules, models, and other methods to determine whether a customer meets the definition of a target customer group and output the identification result of whether the customer belongs to the target customer group.
[0054] Currently, customers are typically identified using textual information that describes them or represents their characteristics. However, since the same textual information may correspond to different semantics in different business scenarios, identifying customers solely through textual information may lead to misjudgments due to semantic differences. Therefore, the accuracy of existing customer identification methods is low.
[0055] Currently, the mainstream methods for identifying customers using textual information include expert rule-based methods and natural language models. For expert rule-based methods, regular expressions, character segmentation, and recognition techniques can be used to formulate rules at the keyword, key-value, and term levels. In practical applications, individual textual information is typically weighted (if the textual information is matched, its weight in the recognition result is increased), or the textual information is organized into a classifier to serve as the basis for feature classification or segmentation models. However, expert rule-based methods cannot standardize the information contained in new texts. Furthermore, there is a possibility of misclassification due to overlap between individual texts and those involved in expert rules, even if their semantics differ. For natural language models, document topic recognition models (which extract abstract topics from a series of documents, usually represented by a set of words) can be used to extract customer-related topics, or language models, such as CTC (Connectionist Temporal) models, can be used. Text recognition is performed using models such as Classification (connected temporal classification), Sequence2Sequence, and Attention. However, using intelligent algorithms (i.e., the above-mentioned methods of finding text topics or summarizing semantics through document topic recognition models or language models) often results in overly scattered topic results or a bias towards the natural language connotation of the text, making it impossible to accurately identify customers based on business scenarios.
[0056] Based on this, this application proposes a customer identification method according to a first embodiment, please refer to... Figure 1 The customer identification method includes:
[0057] Step S10: Obtain target text information associated with the target customer, and obtain target business scenario information corresponding to the target text information;
[0058] It should be noted that a target customer refers to a customer within a specific target customer group. The number of target customers can be one or more, and this embodiment does not limit this. Textual information is used to describe customers and characterize their features. Describing a customer can refer to describing relevant attribute information, such as the customer's personal information. Customer features can refer to the customer's behavioral characteristics, such as the type of financial product the customer purchased, the customer's purchasing characteristics, etc. Target textual information refers to textual information that is highly relevant to the target customer. Business scenario information is used to characterize the scenario features within a corresponding business scenario. For example, this business scenario information may include the purchase quantity and purchase amount within a certain business scenario. Target business scenario information refers to business scenario information that is related to the target textual information.
[0059] When acquiring target text information associated with a target customer and target business scenario information corresponding to the target text information, the information can be acquired from the local device or from other devices connected to the local device. This embodiment does not limit the acquisition of such information.
[0060] Step S20: Construct a target customer identification model based on the target text information and the target business scenario information;
[0061] It should be noted that the target customer identification model is used to identify customers based on business scenario information and text information.
[0062] Step S30: By inputting the behavioral information of the customer to be identified into the target customer identification model, the identification result of the customer to be identified is obtained.
[0063] It should be noted that behavioral information is used to characterize a customer's financial behavior, and this behavioral information may include the number of financial products the customer purchases, the amount the customer spends on purchasing financial products, etc.
[0064] This application provides a customer identification method. First, it obtains target text information associated with a target customer and target business scenario information corresponding to that text information. Then, based on the target text information and the target business scenario information, it constructs a target customer identification model. Finally, by inputting the behavioral information of the customer to be identified into the target customer identification model, the identification of the customer to be identified is completed, and the identification result is obtained. Therefore, this application combines text information and business scenario information to determine the customer identification model. This allows the business scenario information to assist in customer identification when using the customer identification model, thus enabling the customer identification model to distinguish the semantics of the same text information in different business scenarios. This solves the technical problem of low identification accuracy in existing customer identification methods and improves the accuracy of customer identification. Furthermore, since the technical solution of this application uses a customer identification model constructed from text information and business scenario information for customer identification, the identification error is lower compared to the customer identification model constructed solely from text information in the prior art, improving the stability of the computer-output identification results.
[0065] Furthermore, by integrating different textual information with the customer's business scenario information in specific business scenarios, this application embodiment achieves the goal of closely linking textual information with business scenarios. Compared with the method of expert experience rules, this application embodiment can target and enhance the originally ambiguous textual information, thereby enhancing the business connotation of the textual information. Compared with the method of natural language models, this application embodiment has higher interpretability, which meets the requirements of interpretability of customer association in the field of financial risk control.
[0066] In one possible implementation, the step of obtaining the target text information associated with the target customer includes:
[0067] Step S101: Perform text segmentation on each text information in the obtained text information set to obtain each initial text information;
[0068] It should be noted that the initial text information refers to the text information obtained after text segmentation processing. Text segmentation processing refers to dividing the text information into several meaningful text blocks, such as generating a prefix tree or a trie, and scanning the word graph based on the structure of the prefix tree or trie to generate a directed acyclic graph (DAG) to represent all possible word combinations for each Chinese character in the sentence, or to filter possible words by calculating the hidden state sequence by labeling Chinese words.
[0069] Step S102: Determine the relevance between each initial text message and the target customer through chi-square verification and information value detection.
[0070] It should be noted that the chi-square test measures the deviation between observed and inferred values in a classification sample. The chi-square test includes at least two proportions or components, i.e., a comparison between the features contained in the initial text information and the target customer. Information Value (IV) measures the predictive power of a feature on the target by weighted summation of the proportion of positive and negative samples in different groups to the total positive and negative samples. This involves filtering features in the initial text information that are associated with the target customer. The correlation degree characterizes the degree of connection between the initial text information and the customer group characteristics corresponding to the target customer, i.e., the degree of relevance. For example, if the customer group characteristic corresponding to the target customer is "large-amount hardware purchases," and "large-amount hardware purchases" are related to "purchase amount" and "purchase quantity," then if the initial text information includes "purchase amount" and "purchase quantity," then the initial text information is considered associated with the target customer.
[0071] As an example, the step of determining the relevance between each initial text information and the target customer through chi-square verification and information value detection includes: for any initial text information, verifying and calculating the similarity value between each feature in the initial text information and the customer group feature corresponding to the target customer through chi-square verification, and calculating the information value between each feature in the initial text information and the customer group feature corresponding to the target customer through information value detection, calculating the sum of the similarity value and the information value, and obtaining the relevance between the initial text information and the target customer.
[0072] Step S103: Initial text information with a relevance greater than a preset relevance threshold is used as target text information associated with the target customer.
[0073] For example, suppose that the target customer belongs to a target customer group with the characteristic of "large-amount hardware purchases". Since "large-amount hardware purchases" are related to "purchase amount" and "purchase quantity", text information containing "purchase amount" and "purchase quantity" will be highly relevant to the target customer. In this case, the purpose of the preset relevance threshold is to determine whether the text information contains both "purchase amount" and "purchase quantity" at the same time. If it contains both, the result of the relevance is greater than the preset relevance threshold.
[0074] In this embodiment, the text information in the acquired text information set is first segmented to obtain initial text information. Then, the relevance of each initial text information to the target customer is determined by chi-square verification and information value detection. Next, initial text information with a relevance greater than a preset relevance threshold is selected as target text information associated with the target customer. Thus, this embodiment quantifies the text information through text segmentation, chi-square verification, and information value detection to filter out target text information associated with the target customer. This ensures that the deviation between the output of the constructed customer identification model and the actual result is small, guaranteeing the accuracy of the customer identification model. Furthermore, by segmenting the text information in the acquired text information set, the computer processes small units of data, avoiding the impact of excessively large data processing on computer operation and ensuring normal computer operation.
[0075] In one possible implementation, the step of obtaining the target business scenario information corresponding to the target text information includes:
[0076] Step S11: Extract keywords from the target text information;
[0077] It should be noted that keywords can be the most frequently occurring words in the target text information, or words in the target text information that match words recorded in the prediction keyword database. This embodiment does not limit this.
[0078] Step S12: Using the keyword as an index, search for the target business scenario information corresponding to the target text information in the preset business scenario information configuration table.
[0079] It should be noted that the business scenario information configuration table records the correspondence between different keywords and different business scenario information. For example, if the keyword is "hardware", the corresponding business scenario information for "hardware" includes "purchase amount" and "purchase quantity". The "purchase amount" and "purchase quantity" can be used to distinguish whether the customer is making a large-scale wholesale or a small-scale consumption when purchasing "hardware".
[0080] In this embodiment, by extracting keywords from the target text information, the target business scenario information corresponding to the target text information is searched in a preset business scenario information configuration table using these keywords. Thus, this embodiment improves the efficiency of determining the target business scenario information by directly using keywords in a business scenario information configuration table that records the correspondence between different keywords and different business scenario information.
[0081] In one possible implementation, the step of constructing a target customer identification model based on the target text information and the target business scenario information includes:
[0082] Step S21: Based on the target text information and the target business scenario information, construct a text enhancement feature set using a text enhancement function;
[0083] It should be noted that the text augmentation feature set includes multiple text augmentation features, which are formed by combining features from the target text information and features from the target business scenario information. This text augmentation function can be replaced by models that can handle categorical variables, such as LightGBM (LightGradient Boosting Machine).
[0084] As an example, the process of constructing a text enhancement feature set based on the target text information and the target business scenario information using a text enhancement function is as follows:
[0085] x strnew =f(x) i ,x string |x i ∈X,x string ∈X string )
[0086] Where, x strnew x represents the text enhancement features in the text enhancement feature set. i x represents the features in the target text information. string X represents a feature in the target business scenario information. i Represents the target text information, X string This represents the target business scenario information, and f(·) is the text enhancement function.
[0087] For example, for the business scenario of "purchasing hardware", the following text enhancement features can be constructed based on the business scenario information of "purchase amount" and "purchase quantity":
[0088]
[0089] For example, in the business scenario of "purchasing hardware", the following text enhancement features can be constructed based on the business scenario information of "purchase amount" and "time period":
[0090]
[0091] Step S22: Divide the text enhancement feature set to obtain a training sample set and a validation sample set;
[0092] It should be noted that the training sample set is used to train the target customer identification model, and the validation sample set is used to verify whether the trained target customer identification model has an overfitting problem. The overfitting problem means that the target customer identification model performs very well on the training sample set, but performs very poorly on the validation sample set.
[0093] As an example, the step of dividing the text enhancement feature set to obtain a training sample set and a validation sample set may include: dividing the text enhancement feature set proportionally to obtain a training sample set and a validation sample set; it may also include: dividing the text enhancement feature set according to a preset division rule to obtain a training sample set and a validation sample set. This example does not limit this step.
[0094] Step S23: Select training samples from the training sample set, and determine the sample to be identified based on the training samples and the sample weights corresponding to the training samples;
[0095] It should be noted that the samples to be identified are used for iterative training of the customer identification model to be trained. Training samples can be selected arbitrarily from the training sample set, or they can be selected sequentially from the training sample set according to the chronological order of their occurrence. This embodiment does not impose any limitations on this.
[0096] Step S24: Input the sample to be identified into the customer identification model to be trained to obtain the sample identification result;
[0097] It should be noted that the sample recognition result refers to the recognition result output by the customer recognition model to be trained based on the sample to be recognized. The customer recognition model to be trained can be an XGBoost (Xtreme Gradient Boosting) model.
[0098] As an example, the step of inputting the sample to be identified into the customer identification model to be trained and obtaining the sample identification result includes: mapping the sample to be identified to the sample identification result through the customer identification model to be trained.
[0099] Step S25: Optimize the customer identification model to be trained based on the verification sample set and the sample identification results;
[0100] As an example, the step of optimizing the customer recognition model to be trained based on the validation sample set and the sample recognition results includes: obtaining the real results corresponding to the training samples in the validation sample set, and optimizing the customer recognition model to be trained based on the difference between the real results and the sample recognition results. It can be understood that optimizing the customer recognition model to be trained mainly optimizes the internal parameters of the model so that the optimized customer recognition model to be trained can output recognition results with a smaller difference from the real results.
[0101] Step S26: Return to the step of selecting training samples from the training sample set and determining the samples to be identified according to the training samples and the sample weights corresponding to the training samples, until the customer identification model to be trained meets the preset iterative training termination condition, and obtain the target customer identification model.
[0102] It should be noted that the preset iterative training termination condition is used to characterize the termination condition for iterative training of the customer recognition model to be trained. The preset iterative training termination condition can be that the difference between the sample recognition result and the real result is less than the preset difference threshold or the loss function converges.
[0103] In this embodiment, by iteratively training the customer identification model using training and validation sample sets, the trained model can output identification results with minimal difference from the actual results, thus improving the model's accuracy. Furthermore, since the training and optimization process essentially optimizes the computational parameters in the computer, it enhances the computer's accuracy in customer identification.
[0104] In one possible implementation, the step of optimizing the customer identification model to be trained based on the verification sample set and the sample identification results includes:
[0105] Step S251: Obtain the true result corresponding to the training sample in the verification sample set;
[0106] Step S252: Based on the difference between the real result and the sample recognition result, construct a loss function and verify whether the loss function is a convergent function;
[0107] Step S253: If not, optimize the customer identification model to be trained based on the calculated gradient of the loss function.
[0108] In this embodiment, the real results corresponding to the training samples are first obtained from the validation sample set. Then, based on the difference between the real results and the recognition results output by the customer recognition model to be trained according to the samples to be identified, a loss function is constructed. When the loss function is found to be convergent, the customer recognition model to be trained is optimized based on the calculated gradient of the loss function. This ensures that the final target customer recognition model is jointly decided by the training sample set and the validation sample set, thereby improving the recognition accuracy of the target customer recognition model. Furthermore, since the target customer recognition model trained in this embodiment is jointly decided by the training sample set and the validation sample set, compared to the target customer recognition model obtained by using only the training sample set, its final output recognition result has a smaller error, thus improving the stability of the computer's output recognition result.
[0109] Example 2
[0110] Based on the first embodiment of this application, in another embodiment of this application, the same or similar content as in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 The step of constructing a target customer identification model based on the target text information and the target business scenario information includes:
[0111] Step E21: Based on the target text information and the target business scenario information, construct a first customer identification model, and based on the target text information, construct a second customer identification model;
[0112] As an example, the step of constructing a second customer identification model based on the target text information includes: dividing the target text information to obtain a first training sample set and a first verification sample set; selecting a first training sample from the first training sample set and determining a first sample to be identified based on the first training sample and the sample weights corresponding to the first training sample; inputting the first sample to be identified into the customer identification model to be trained to obtain a first sample identification result; optimizing the customer identification model to be trained based on the first verification sample set and the first sample identification result; returning to the step of selecting a first training sample from the first training sample set and determining a first sample to be identified based on the first training sample and the sample weights corresponding to the first training sample, until the customer identification model to be trained meets the preset iterative training termination condition, thereby obtaining the second customer identification model.
[0113] Step E22: Merge the first customer identification model and the second customer identification model to obtain the target customer identification model.
[0114] It should be noted that the fusion process of the first customer identification model and the second customer identification model can also be replaced by logistic regression.
[0115] In this embodiment, a first customer identification model incorporating business scenario information is first constructed based on the target text information and the target business scenario information. A second customer identification model, without incorporating business scenario information, is then constructed based on the target text information. The first and second customer identification models are then merged to obtain the target customer identification model. Subsequently, when using this target customer identification model for customer identification, the system can determine whether the information input into the target customer identification model carries business scenario information, thereby selecting the appropriate customer identification model and improving the flexibility of customer identification. Furthermore, the computer can flexibly select the process for outputting the identification result based on the input, which also improves the flexibility of computer operation.
[0116] The target customer identification model obtained in this embodiment is formed by fusing a first customer identification model constructed based on target text information and target business scenario information with a second customer identification model constructed based on target text information. In contrast, the target customer identification model obtained in Embodiment 1 is the first customer identification model in this embodiment. That is, the target customer identification model obtained in Embodiment 1 does not incorporate the second customer identification model constructed based on target text information. Therefore, the target customer identification model obtained in this embodiment has a certain degree of flexibility in customer identification compared to the target customer identification model obtained in Embodiment 1.
[0117] In one possible implementation, the target customer identification model is formed by fusing a first customer identification model and a second customer identification model. The step of obtaining the identification result of the customer to be identified by inputting the behavioral information of the customer to be identified into the target customer identification model includes:
[0118] Step S31: After inputting the behavioral information of the customer to be identified into the target customer identification model, if the behavioral information carries business scenario information, then the first customer identification model is used as the model to be solved.
[0119] Understandably, if the behavioral information of the customer to be identified carries business scenario information, the business scenario can be used as the basis for customer identification in order to ensure the accuracy of identification. Therefore, the customer to be identified can be identified by combining the first customer identification model generated with the business scenario information.
[0120] Step S32: If the behavioral information does not carry business scenario information, then the second customer identification model is used as the model to be solved.
[0121] Understandably, if the behavioral information of the customer to be identified does not carry business scenario information, a second customer identification model can be used to identify the customer to improve the efficiency of identification. Since the second customer identification model only needs to process the text information in the behavioral information and does not need to process the business scenario information, it requires less data to process than the first customer identification model and will run faster.
[0122] Step S33: Based on the preset solution algorithm, determine the optimal solution of the model to be solved, and use the optimal solution as the identification result of the customer to be identified.
[0123] It should be noted that the preset solution algorithm can be a single-objective solution algorithm or a dual-objective solution algorithm, such as particle swarm optimization, ant colony optimization, intelligent optimization algorithm, etc.
[0124] For example, assuming the input to the target customer identification model is "the amount of hardware purchased" and "the quantity of hardware purchased", the first customer identification model is selected to output the identification result. If the input to the target customer identification model is "the amount of hardware purchased" and "the quantity of hardware purchased", the second customer identification model is selected to output the identification result.
[0125] In this embodiment, the customer identification model is determined by judging whether the behavior information of the customer to be identified carries business scenario information. If the behavior information of the customer to be identified carries business scenario information, the customer to be identified is identified by combining the first customer identification model generated by the business scenario information, so as to use the business scenario as the basis for customer identification and ensure the accuracy of customer identification. If the behavior information of the customer to be identified does not carry business scenario information, the customer to be identified is identified by the second customer identification model generated without combining business scenario information, so as to improve the efficiency of customer identification.
[0126] For example, to aid in understanding the technical concept or principles of this application, please refer to Figure 3 and Figure 4 , Figure 3 A simplified flowchart for constructing a target customer identification model is provided. Figure 4 An interactive diagram of the target customer identification model is provided, as follows:
[0127] 1. Text Segmentation and Filtering: Based on the text information, perform text segmentation at the character and word level, and use chi-square test and information value detection to determine the relationship or information value between the text information and the target customer (corresponding to the above-mentioned relevance), and then filter out the target text information related to the target customer;
[0128] 2. Business scenario information filtering: For different business scenarios, other features of the same sample in the sample space can be observed, and features associated with the target text information can be selected, that is, the target business scenario information corresponding to the target text information can be determined.
[0129] 3. Constructing text-enhanced data: By combining features from the target text information and features from the target business scenario information through text enhancement functions, a text-enhanced feature set is obtained;
[0130] 4. Train and output customer recognition model: Train the XGBoost model using the text augmentation feature set to obtain the first customer recognition model, and train the XGBoost model using the target text information to obtain the second customer recognition model.
[0131] 5. Model Fusion: The first customer identification model and the second customer identification model are fused to obtain the target customer identification model. If the data input to the target customer identification model carries business scenario information, the identification result is output through the first customer identification model. If the data input to the target customer identification model does not carry business scenario information, the identification result is output through the second customer identification model.
[0132] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the customer identification method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0133] Example 3
[0134] This invention also provides a customer identification device, please refer to... Figure 5 The customer identification device includes:
[0135] The acquisition module 10 is used to acquire target text information associated with the target customer, and to acquire target business scenario information corresponding to the target text information;
[0136] Construction module 20 is used to construct a target customer identification model based on the target text information and the target business scenario information;
[0137] The identification module 30 is used to obtain the identification result of the customer to be identified by inputting the behavioral information of the customer to be identified into the target customer identification model.
[0138] Optionally, the acquisition module 10 is further configured to:
[0139] The obtained text information set is processed by text segmentation to obtain the initial text information.
[0140] The relevance of each initial textual piece of information to the target customer is determined by chi-square verification and information value detection.
[0141] Initial text information with a relevance greater than a preset relevance threshold is used as the target text information associated with the target customer.
[0142] Optionally, the acquisition module 10 is further configured to:
[0143] Extract keywords from the target text information;
[0144] Using the keywords as an index, the target business scenario information corresponding to the target text information is searched in the preset business scenario information configuration table.
[0145] Optionally, the building module 20 is further configured to:
[0146] Based on the target text information and the target business scenario information, a text enhancement feature set is constructed using a text enhancement function;
[0147] The text enhancement feature set is divided into a training sample set and a validation sample set;
[0148] Training samples are selected from the training sample set, and the samples to be identified are determined based on the training samples and their corresponding sample weights.
[0149] The sample to be identified is input into the customer identification model to be trained to obtain the sample identification result;
[0150] Based on the verification sample set and the sample recognition results, optimize the customer recognition model to be trained;
[0151] Return to the steps of selecting training samples from the training sample set and determining the samples to be identified based on the training samples and their corresponding sample weights, until the customer identification model to be trained meets the preset iterative training termination condition, and obtain the target customer identification model.
[0152] Optionally, the building module 20 is further configured to:
[0153] Obtain the true results corresponding to the training samples from the validation sample set;
[0154] Based on the difference between the actual results and the sample identification results, a loss function is constructed, and it is verified whether the loss function is a convergent function.
[0155] If not, then optimize the customer identification model to be trained based on the calculated gradient of the loss function.
[0156] Optionally, the building module 20 is further configured to:
[0157] Based on the target text information and the target business scenario information, a first customer identification model is constructed, and based on the target text information, a second customer identification model is constructed.
[0158] The target customer identification model is obtained by merging the first customer identification model and the second customer identification model.
[0159] Optionally, the target customer identification model is formed by fusing a first customer identification model and a second customer identification model, and the identification module 30 is further used for:
[0160] After inputting the behavioral information of the customer to be identified into the target customer identification model, if the behavioral information carries business scenario information, then the first customer identification model is used as the model to be solved.
[0161] If the behavioral information does not carry business scenario information, then the second customer identification model will be used as the model to be solved.
[0162] Based on a preset solution algorithm, the optimal solution of the model to be solved is determined, and the optimal solution is used as the identification result of the customer to be identified.
[0163] The customer identification device provided by this invention, employing the customer identification method in Embodiment 1 or Embodiment 2 described above, can solve the technical problem of low identification accuracy in existing customer identification methods. Compared with the prior art, the beneficial effects of the customer identification device provided by this invention are the same as those of the customer identification method provided in the above embodiments, and other technical features in the customer identification device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0164] Example 4
[0165] This invention provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the customer identification method in Embodiment 1 above.
[0166] The following is for reference. Figure 6 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of the present disclosure. The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0167] like Figure 6As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. While electronic devices with various systems are shown in the figures, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0168] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of embodiments of this disclosure.
[0169] The electronic device provided by this invention, employing the customer identification method described in the above embodiments, can solve the technical problem of low identification accuracy in existing customer identification methods. Compared with the prior art, the beneficial effects of the electronic device provided by this invention are the same as those of the customer identification method described in the above embodiments, and other technical features of this electronic device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0170] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.
[0171] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
[0172] Example 5
[0173] This invention provides a computer-readable storage medium having computer-readable program instructions stored thereon, which are used to execute the customer identification method in Embodiment 1 above.
[0174] The computer-readable storage medium provided in this embodiment of the invention may be, for example, a USB flash drive, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0175] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.
[0176] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by an electronic device, cause the electronic device to: acquire target text information associated with a target customer, and acquire target business scenario information corresponding to the target text information; construct a target customer identification model based on the target text information and the target business scenario information; and obtain the identification result of the target customer by inputting the behavioral information of the customer to be identified into the target customer identification model.
[0177] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0178] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0179] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0180] The readable storage medium provided by this invention is a computer-readable storage medium that stores computer-readable program instructions for executing the above-described customer identification method, thereby solving the technical problem of low identification accuracy in existing customer identification methods. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this invention are the same as those of the customer identification method provided in Embodiment 1 or Embodiment 2, and will not be repeated here.
[0181] Example 6
[0182] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the customer identification method described above.
[0183] The computer program product provided in this application can solve the technical problem of low recognition accuracy in existing customer identification methods. Compared with the prior art, the beneficial effects of the computer program product provided in the embodiments of this invention are the same as the beneficial effects of the customer identification methods provided in Embodiment 1 or Embodiment 2 above, and will not be repeated here.
[0184] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.
Claims
1. A method of customer identification, characterized by, The customer identification method includes: Obtain target text information associated with the target customer, and obtain target business scenario information corresponding to the target text information; Based on the target text information and the target business scenario information, a target customer identification model is constructed, wherein the target business scenario information enables the target customer identification model to distinguish the semantics of the target text information in different business scenarios; By inputting the behavioral information of the customer to be identified into the target customer identification model, the identification result of the customer to be identified is obtained; The step of constructing a target customer identification model based on the target text information and the target business scenario information includes: Based on the target text information and the target business scenario information, a first customer identification model is constructed, and based on the target text information, a second customer identification model is constructed. The target customer identification model is obtained by integrating the first customer identification model and the second customer identification model. The step of inputting the behavioral information of the customer to be identified into the target customer identification model to obtain the identification result of the customer to be identified includes: After inputting the behavioral information of the customer to be identified into the target customer identification model, if the behavioral information carries business scenario information, then the first customer identification model is used as the model to be solved. If the behavioral information does not carry business scenario information, then the second customer identification model will be used as the model to be solved. Based on a preset solution algorithm, the optimal solution of the model to be solved is determined, and the optimal solution is used as the identification result of the customer to be identified.
2. The client identification method of claim 1, wherein, The step of obtaining target text information associated with the target customer includes: The obtained text information set is processed by text segmentation to obtain the initial text information. The relevance of each initial textual piece of information to the target customer is determined by chi-square verification and information value detection. Initial text information with a relevance greater than a preset relevance threshold is used as the target text information associated with the target customer.
3. The customer identification method as described in claim 1, characterized in that, The step of obtaining the target business scenario information corresponding to the target text information includes: Extract keywords from the target text information; Using the keywords as an index, the target business scenario information corresponding to the target text information is searched in the preset business scenario information configuration table.
4. The customer identification method as described in claim 1, characterized in that, The step of constructing a first customer identification model based on the target text information and the target business scenario information includes: Based on the target text information and the target business scenario information, a text enhancement feature set is constructed using a text enhancement function; The text enhancement feature set is divided into a training sample set and a validation sample set; Training samples are selected from the training sample set, and the samples to be identified are determined based on the training samples and their corresponding sample weights. The sample to be identified is input into the customer identification model to be trained to obtain the sample identification result; Optimize the customer identification model to be trained based on the verification sample set and the sample identification results; Return to the steps of selecting training samples from the training sample set and determining the samples to be identified based on the training samples and their corresponding sample weights, until the customer identification model to be trained meets the preset iterative training termination condition, and obtain the first customer identification model.
5. The customer identification method as described in claim 4, characterized in that, The step of optimizing the customer identification model to be trained based on the verification sample set and the sample identification results includes: Obtain the true results corresponding to the training samples from the validation sample set; Based on the difference between the actual results and the sample identification results, a loss function is constructed, and it is verified whether the loss function is a convergent function. If not, then optimize the customer identification model to be trained based on the calculated gradient of the loss function.
6. A customer identification device, characterized in that, The customer identification device includes: The acquisition module is used to acquire target text information associated with the target customer, and to acquire target business scenario information corresponding to the target text information; The construction module is used to construct a target customer identification model based on the target text information and the target business scenario information, wherein the target business scenario information enables the target customer identification model to distinguish the semantics of the target text information in different business scenarios; The identification module is used to obtain the identification result of the customer to be identified by inputting the behavioral information of the customer to be identified into the target customer identification model; The building module is further used for: Based on the target text information and the target business scenario information, a first customer identification model is constructed, and based on the target text information, a second customer identification model is constructed. The target customer identification model is obtained by integrating the first customer identification model and the second customer identification model. The identification module is also used for: After inputting the behavioral information of the customer to be identified into the target customer identification model, if the behavioral information carries business scenario information, then the first customer identification model is used as the model to be solved. If the behavioral information does not carry business scenario information, then the second customer identification model will be used as the model to be solved. Based on a preset solution algorithm, the optimal solution of the model to be solved is determined, and the optimal solution is used as the identification result of the customer to be identified.
7. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the steps of the customer identification method as described in any one of claims 1 to 5.
8. A readable storage medium, characterized in that, The readable storage medium is a computer-readable storage medium, on which a program implementing the customer identification method is stored, the program implementing the customer identification method being executed by a processor to implement the steps of the customer identification method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Feature vector generation method and device and text classification method and device based on feature vector
CN110119445A
Semantic recognition method and device, equipment and storage medium
CN115422944A