Classified message pushing method and device, computer equipment and storage medium
By acquiring customer reports, online information, and claims information, generating behavioral summaries, and performing cluster analysis, the problem of insufficient accuracy in insurance industry message pushes has been solved, achieving more accurate information delivery.
Patent Information
- Application Number
- CN202510764385.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies cannot accurately identify the priority of customers' real needs in insurance industry push notifications, resulting in the generalization of push content and an inability to effectively distinguish between proactive and reactive insurance purchases, leading to a decrease in accuracy.
By acquiring customer reporting information, online information, and claims information, text processing is performed to generate behavioral summaries. These summaries are then input into a pre-trained language model for dimensionality transformation and cluster analysis to obtain customer classification results for information push.
It enables comprehensive analysis based on multiple dimensions of customer information, improving the accuracy and effectiveness of message push.
Smart Images

Figure CN120873273A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, computer device, and storage medium for classifying and pushing messages. Background Technology
[0002] In the digital marketing strategies of the insurance industry, targeted messaging has always been a core means of improving marketing conversion rates. However, the effectiveness of current insurance marketing messaging is constrained by multiple factors, exhibiting significant domain-specific differences. In the financial sector, the effectiveness of messaging for auto insurance products is affected by dynamic variables such as vehicle age, regional risk levels, and driver habits; while in the digital healthcare sector, medical insurance messaging needs to consider sensitive medical data such as the insured's health status, medical records, and drug consumption. This cross-domain differentiation makes it difficult to establish a unified, targeted marketing model using traditional messaging methods.
[0003] Current technologies primarily rely on customers' historical search behavior and insurance records for message push notifications, a method with significant technical limitations. First, when customers simultaneously search for or browse multiple insurance products (e.g., both car and medical insurance), the system struggles to accurately prioritize the customer's true needs, leading to indiscriminate push notifications. Second, for customers with complex insurance histories, current technologies cannot effectively distinguish between proactive and reactive insurance (e.g., corporate group insurance), resulting in misjudgments of needs and consequently, decreased accuracy in message push notifications, hindering the sustainable development of the insurance industry. Summary of the Invention
[0004] The purpose of this application is to provide a method, apparatus, computer device, and storage medium for classifying message push, so as to solve the problem of being unable to accurately and effectively push messages to customers.
[0005] Firstly, embodiments of this application provide a categorized message push method, which employs the following technical solution:
[0006] Obtain first customer information, extract reporting information from the first customer information, and obtain reporting behavior information;
[0007] Obtain second customer information, and extract online information from the second customer information to obtain online behavior information;
[0008] Obtain third-party customer information, extract claims information from the third-party customer information, and obtain claims behavior information;
[0009] Based on word engineering, text processing is performed on the reporting behavior information, the online behavior information, and the claims behavior information to generate behavior summaries;
[0010] The behavior summary is input into a pre-trained language model to obtain behavior word embedding vectors;
[0011] The behavior word embedding vector is subjected to dimensionality transformation to obtain a low-dimensional behavior vector;
[0012] Cluster analysis is performed on the low-dimensional vectors of the behaviors to obtain customer classification results, and information is pushed based on the customer classification results.
[0013] Secondly, this application also provides a categorized message push device, which adopts the following technical solution:
[0014] The first information extraction module is used to obtain first customer information, extract reporting information from the first customer information, and obtain reporting behavior information.
[0015] The second information extraction module is used to obtain second customer information, extract online information from the second customer information, and obtain online behavior information.
[0016] The third information extraction module is used to obtain third customer information, extract claims information from the third customer information, and obtain claims behavior information.
[0017] The information processing module is used to perform text processing on the reporting behavior information, the online behavior information, and the claims behavior information based on word engineering, and generate behavior summaries;
[0018] The model processing module is used to input the behavior summary into a pre-trained language model to obtain behavior word embedding vectors;
[0019] The dimension transformation module is used to perform dimension transformation processing on the behavior word embedding vector to obtain a low-dimensional behavior vector.
[0020] The results analysis module is used to perform cluster analysis on the low-dimensional vector of behavior to obtain customer classification results, and to push information based on the customer classification results.
[0021] Thirdly, embodiments of this application also provide a computer device that adopts the technical solution described below:
[0022] A computer device includes a memory and a processor, the memory storing computer-readable instructions, the processor executing the computer-readable instructions to implement the steps of the categorized message push method as described in any of the preceding claims.
[0023] Fourthly, embodiments of this application also provide a computer-readable storage medium, which adopts the technical solutions described below:
[0024] A computer-readable storage medium storing computer-readable instructions, which, when executed by a processor, implement the steps of the categorized message push method as described above.
[0025] Compared with existing technologies, the embodiments of this application have the following main advantages: This embodiment obtains first customer information, extracts reporting information from the first customer information to obtain reporting behavior information; obtains second customer information, extracts online information from the second customer information to obtain online behavior information; obtains third customer information, extracts claims information from the third customer information to obtain claims behavior information; performs text processing on the reporting behavior information, the online behavior information, and the claims behavior information based on word engineering to generate a behavior summary; inputs the behavior summary into a pre-trained language model to obtain behavior word embedding vectors; performs dimensional transformation processing on the behavior word embedding vectors to obtain low-dimensional behavior vectors; performs cluster analysis on the low-dimensional behavior vectors to obtain customer classification results, and pushes information based on the customer classification results. This effectively achieves comprehensive analysis and push based on multiple dimensions of customer information, thereby improving the accuracy and effectiveness of the push. Attached Figure Description
[0026] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is an exemplary system architecture diagram to which this application can be applied;
[0028] Figure 2 A flowchart of an embodiment of the categorized message push method according to this application;
[0029] Figure 3 This is a schematic diagram of the structure of one embodiment of the categorized message push device according to this application;
[0030] Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation
[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0032] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a non-related or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0033] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0034] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.
[0035] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0036] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptop computer 1011, tablet computer 1012 or mobile phone 1013, terminal device 101 can also be e-book reader, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer and desktop computer, etc.
[0037] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.
[0038] It should be noted that the categorized message push method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the categorized message push device is generally set in the server / terminal device.
[0039] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0040] Continue to refer to Figure 2 The diagram illustrates a flowchart of an embodiment of the categorized message push method according to this application. The categorized message push method includes the following steps:
[0041] Step S10: Obtain first customer information, extract reporting information from the first customer information, and obtain reporting behavior information;
[0042] In this embodiment, the first customer information includes customer non-motor vehicle quotation information and customer insurance information. The customer non-motor vehicle quotation information refers to the customer's quotation requests for non-motor vehicle insurance (such as health insurance, property insurance, accident insurance, etc.), including basic customer information, quotation request information, product preference information, etc. The customer insurance information refers to a detailed record of the customer's actual purchase of motor vehicle insurance products, including the type of insurance purchased, premium amount, insurance period, and insurance date, etc. Insurance information extraction refers to extracting the non-motor vehicle quotation information and insurance information from the first customer information. Reporting behavior information includes reporting frequency, average reporting amount, reporting type distribution, reporting status distribution, reporting time information, and reporting channel distribution, etc. The aforementioned first customer information can be historical customer non-motor vehicle quotation information and customer insurance information, or customer non-motor vehicle quotation information and customer insurance information within a historical time period, such as customer non-motor vehicle quotation information and customer insurance information within the past year. The first customer information can also be customer medical insurance quote information and customer medical insurance enrollment information in medical insurance. Among them, customer medical insurance quote information includes customer basic medical insurance information, medical insurance quote request information, medical insurance product preference information, etc., and customer medical insurance enrollment information includes medical insurance type, medical insurance premium amount, medical insurance period, medical insurance enrollment date, etc.
[0043] Step S20: Obtain second customer information, extract online information from the second customer information, and obtain online behavior information;
[0044] In this embodiment, the second customer information includes online vehicle owner information, which refers to information about the customer's online browsing and operations through channels such as insurance systems, websites, and platforms. This online vehicle owner information includes online behavior records, online needs information, and online preferences information. Online behavior information includes browsing behavior information, search behavior information, and interaction behavior information. The second customer information can also be online user information within medical insurance, where online user information includes medical insurance online behavior records, medical insurance online needs information, and medical insurance online preferences information.
[0045] Step S30: Obtain third customer information, extract claim information from the third customer information, and obtain claim behavior information;
[0046] In this embodiment, the third customer information includes customer non-motor vehicle accident information and customer claims information. Customer non-motor vehicle accident information refers to a series of relevant data and records generated when a customer experiences an insured event that meets the terms of the insurance policy within the coverage scope of non-motor vehicle insurance (such as health insurance, property insurance, accident insurance, etc.), including basic accident information, loss details, and report information. Customer claims information refers to a series of relevant data and records generated when a customer experiences an insured event that meets the terms of the insurance policy within the coverage scope of non-motor vehicle insurance (such as health insurance, property insurance, accident insurance, etc.), including claims application information, claims investigation information, claims review information, and claims payment information. Claims behavior information includes reporting behavior, claims application behavior, communication and cooperation behavior, and feedback and complaint behavior. Thirdly, customer information can also include medical insurance status information and medical insurance claim information. Among them, medical insurance status information includes a series of relevant data and records generated when a customer incurs an insurance claim within the scope of medical insurance coverage, which is in accordance with the terms of the insurance policy. Medical insurance claim information includes medical insurance claim application information, medical insurance claim investigation information, medical insurance claim review information, medical insurance claim payment information, etc.
[0047] Step S40: Based on word engineering, perform text processing on the reporting behavior information, the online behavior information, and the claims behavior information to generate a behavior summary;
[0048] In this embodiment, the behavioral summary includes a reporting behavior information summary, an online behavior information summary, and a claims behavior information summary. The reporting behavior information summary includes behavioral data related to the customer's insurance reporting process; the online behavior information summary includes behavioral records of the customer's activities related to insurance on internet systems and digital platforms; and the claims behavior information summary includes behavioral performance information of the customer and the insurance company during the insurance claims process. Text processing based on word engineering includes text conversion, word segmentation and filtering, and keyword / sentence extraction.
[0049] Step S50: Input the behavior summary into the pre-trained language model to obtain behavior word embedding vectors;
[0050] In this embodiment, the pre-trained language model uses a BERT model trained based on historical sample behavior summary data. The training steps of the language model include: cleaning and preprocessing the sample behavior data; designing the parameters of the Transformer encoder and constructing the Transformer encoder framework; mapping the preprocessed sample behavior summary data into vectors and inputting the mapped vectors into the Transformer encoder framework for training according to a predetermined task; and adjusting the model based on the training output to obtain the pre-trained language model.
[0051] Step S60: Perform dimensionality transformation on the behavior word embedding vector to obtain a low-dimensional behavior vector;
[0052] In this embodiment, PCA (Principal Component Analysis) is used to perform dimensionality transformation on the word embedding vectors. PCA is a linear dimensionality reduction technique used to project high-dimensional data into a low-dimensional space while preserving as much variance as possible. By performing dimensionality transformation on the behavioral word embedding vectors, more easily processed low-dimensional behavioral vectors are obtained.
[0053] Step S70: Perform cluster analysis on the low-dimensional vector of the behavior to obtain customer classification results, and push information based on the customer classification results.
[0054] In this embodiment, a preset clustering algorithm is used to perform clustering analysis on the low-dimensional behavioral vectors. The preset clustering algorithm can be the K-Means clustering algorithm. By determining the optimal number of clusters (K value) for the K-Means clustering algorithm and initializing K cluster centers, sample allocation and center updates are iteratively performed until convergence or the maximum number of iterations is reached, thereby completing the clustering of the low-dimensional behavioral vectors and obtaining vector clusters. Then, customer identification is performed based on the vector clusters, resulting in customer classification results representing different customer categories for message push operations. Customer classification results may include students, office workers, entrepreneurs, etc., and the information push may be related to car insurance or medical insurance.
[0055] This embodiment obtains a first customer information, extracts reporting information from it to obtain reporting behavior information; obtains a second customer information, extracts online information from it to obtain online behavior information; obtains a third customer information, extracts claims information from it to obtain claims behavior information; performs text processing on the reporting behavior information, online behavior information, and claims behavior information based on word engineering to generate a behavior summary; inputs the behavior summary into a pre-trained language model to obtain behavior word embedding vectors; performs dimensionality transformation on the behavior word embedding vectors to obtain low-dimensional behavior vectors; performs cluster analysis on the low-dimensional behavior vectors to obtain customer classification results, and pushes information based on the customer classification results. This effectively achieves comprehensive analysis and push based on multiple dimensions of customer information, improving the accuracy and effectiveness of the push notifications.
[0056] This method can be applied to categorized message pushes in online auto insurance systems. Users obtain and upload first, second, and third customer information related to auto insurance. The system processes this information separately to obtain different types of behavior information: reporting behavior information, online behavior information, and claims behavior information. Then, word engineering is used to process these information into text summaries, generating specific textual descriptions of the behaviors. These behavioral summaries are input into a pre-trained language model for analysis to obtain effective behavioral word embedding vectors representing the summary information. Dimensionality reduction and clustering analysis of these embedding vectors yield accurate customer classification results. Finally, different auto insurance-related information is pushed to customers based on their classification, achieving effective categorized information delivery.
[0057] In some optional implementations of this embodiment, obtaining the first customer information and extracting reporting information from the first customer information to obtain reporting behavior information includes the following steps:
[0058] Obtain a first information extraction identifier, and extract the customer's non-vehicle quotation information and the customer's insurance information from the database based on the first information extraction identifier;
[0059] In this embodiment, the first information extraction identifier is a unique identifier corresponding to the customer's non-vehicle price information and customer insurance information in the database. The corresponding customer's non-vehicle price information and customer insurance information are extracted by traversing the database using the first information extraction identifier as a query condition.
[0060] The customer's non-vehicle pricing information and the customer's insurance information are cleaned to obtain standard non-vehicle pricing information and standard customer insurance information.
[0061] In this embodiment, the data cleaning process includes removing duplicate records, filling in or deleting missing values, and standardizing date formats. By performing data cleaning processes including the above methods on customer non-vehicle quotation information and customer insurance information, standard and valid standard non-vehicle quotation information and standard customer insurance information are obtained.
[0062] Based on predefined insurance behavior feature rules, feature extraction is performed on the standard non-vehicle price information and the standard customer insurance information to obtain the insurance behavior information.
[0063] In this embodiment, the predefined insurance behavior characteristic rules include: insurance frequency: the number of times each customer files a claim within a specified time period; average insured amount: the average claim amount for each customer; insurance type distribution: the number of times each customer files a claim for different types of insurance; insurance channel distribution: the number of times each customer files a claim through different channels; earliest claim date: the date the customer first filed a claim; and most recent claim date: the date the customer most recently filed a claim. Based on these insurance behavior characteristic rules, features are extracted from standard non-motor vehicle pricing information and standard customer insurance information to effectively obtain insurance behavior information.
[0064] This embodiment obtains a first information extraction identifier, and extracts the customer's non-vehicle pricing information and insurance information from the database based on the first information extraction identifier; performs data cleaning processing on the customer's non-vehicle pricing information and insurance information to obtain standard non-vehicle pricing information and standard customer insurance information; and extracts features from the standard non-vehicle pricing information and standard customer insurance information according to predefined insurance behavior feature rules to obtain the insurance behavior information. This effectively achieves the acquisition of customer pricing and insurance-related information, providing reliable data for subsequent processing.
[0065] In some optional implementations of this embodiment, the second customer information includes online vehicle owner information. The steps of obtaining the second customer information and extracting online information from it to obtain online behavior information include the following:
[0066] Obtain the second information extraction identifier, and extract the vehicle owner's online information from the database based on the second information extraction identifier;
[0067] In this embodiment, the second information extraction identifier is a unique identifier corresponding to the online information of the vehicle owner in the database. The corresponding online information of the vehicle owner is extracted by traversing the database using the second information extraction identifier as a query condition.
[0068] Obtain time range definition information, and filter the online vehicle owner information based on the time range definition information to obtain valid online vehicle owner information;
[0069] In this embodiment, the time range definition information is a predefined time range, initially set to one year, which can be adjusted according to actual circumstances. By filtering vehicle owner online information using the time range definition information, valid vehicle owner online information is obtained.
[0070] The online behavior information is obtained by extracting features from the valid online information of the vehicle owners according to the predefined online behavior feature rules.
[0071] In this embodiment, the predefined online behavior feature rules include: access frequency: the number of times each customer accesses the platform within a specified time period; average dwell time: the average dwell time of each customer; module access distribution: the number of times each customer accesses different modules; operation path analysis: extracting the customer's operation path; and click behavior distribution: the click behavior distribution of each customer. Based on these online behavior feature rules, features are extracted from valid online vehicle owner information to effectively obtain online behavior information.
[0072] This embodiment obtains a second information extraction identifier and extracts the vehicle owner's online information from the database based on this identifier; it also obtains time range definition information and filters the online vehicle owner information based on this time range definition information to obtain valid online vehicle owner information; finally, it extracts features from the valid online vehicle owner information according to predefined online behavior feature rules to obtain the online behavior information. This effectively achieves the acquisition of customers' online vehicle owner-related information, providing reliable data for subsequent processing.
[0073] In some optional implementations of this embodiment, obtaining third-party customer information and extracting claims information from the third-party customer information to obtain claims behavior information includes the following steps:
[0074] Obtain a third information extraction identifier, and extract the customer's non-vehicle accident information and customer's claims information from the database based on the third information extraction identifier;
[0075] In this embodiment, the third information extraction identifier is a unique identifier corresponding to the customer's non-vehicle accident information and customer claim information in the database. By using the third information extraction identifier as a query condition to traverse and query the database, the corresponding customer's non-vehicle accident information and customer claim information can be extracted.
[0076] Key fields are extracted from the customer's non-vehicle accident information and the customer's claims information to obtain standard accident information and standard claims information;
[0077] In this embodiment, key fields can be extracted from customer non-vehicle accident information and customer claim information according to preset key field rules to obtain corresponding standard accident information and standard claim information.
[0078] The standard accident information and the standard claim information are feature extracted according to predefined claim behavior feature rules to obtain the claim behavior information.
[0079] In this embodiment, the predefined claims behavior characteristic rules include: claims frequency: the number of claims made by each customer within a specified time period; total claims amount: the total claims amount for each customer; average claims amount: the average claims amount for each customer; claims status distribution: the claims status distribution for each customer; and claims processing time: the average claims processing time for each customer. Based on these claims behavior characteristic rules, features are extracted from standard accident information and standard claims information to effectively obtain claims behavior information.
[0080] This embodiment obtains a third information extraction identifier and extracts the customer's non-vehicle accident information and claim information from the database based on the identifier. Key fields are extracted from the non-vehicle accident information and claim information to obtain standard accident information and standard claim information. Features are then extracted from the standard accident information and standard claim information according to predefined claim behavior feature rules to obtain claim behavior information. This effectively enables the acquisition of customer non-vehicle accident and claim-related information, providing reliable data for subsequent processing.
[0081] In some optional implementations of this embodiment, the text processing of the reporting behavior information, the online behavior information, and the claims behavior information based on word engineering to generate a behavior summary includes the following steps:
[0082] The reported behavior information, the online behavior information, and the claims behavior information are converted into text to obtain behavior information text.
[0083] In this embodiment, customer basic information, reporting frequency, average amount, insurance type, and channel characteristics from the reporting behavior information are integrated and described using natural language to generate reporting behavior description text. Online behavior description text is generated by combining access frequency, dwell time, module access status, operation path, and click behavior from online behavior information. Claims behavior description text is formed by integrating claims frequency, total amount, average amount, status distribution, and processing time from claims behavior information. The reporting behavior description text, online behavior description text, and claims behavior description text are then organized to obtain the behavior information text.
[0084] The behavioral information text is segmented and filtered to obtain valid information text;
[0085] In this embodiment, the word segmentation and filtering process includes word segmentation and stop word removal. By performing word segmentation on the behavioral information text to obtain the corresponding behavioral information words, and by removing stop words from the behavioral information words according to a preset stop word list, the effective information text is obtained.
[0086] Keyword extraction is performed on the effective information text to obtain effective information phrases;
[0087] In this embodiment, the TF-IDF algorithm can be used to extract keywords from the effective information text, and the TextRank algorithm can be used to extract key sentences from the effective information text. TF-IDF is a classic algorithm used in information retrieval and text mining to evaluate the importance of a word in a document collection or corpus. TextRank is a graph-based text processing algorithm that constructs a graph structure of the text and calculates the importance of nodes using the connections between nodes, thereby extracting key information (such as keywords and key sentences) from the text. By extracting keywords and key sentences from the effective information text and then integrating them, the effective information words and sentences are obtained.
[0088] Obtain a preset summary template, fill the effective information phrases into the preset summary template, and generate the behavior summary.
[0089] In this embodiment, the preset summary templates include: a reporting behavior summary template: within a {time range}, customer {customer ID} made {number of reports} reports, average amount {average amount}, and main insurance types {insurance type list}; an online behavior summary template: within a {time range}, customer {customer ID} accessed {number of times}, average stay {duration} seconds, and main modules {module list}; and a claims behavior summary template: within a {time range}, customer {customer ID} made {number of claims} claims, total amount {total amount}, average amount {average amount}, and status distribution {status distribution}. By filling in the corresponding valid information phrases into the preset summary templates and organizing the entered summaries, a behavior summary is effectively obtained.
[0090] This embodiment converts the reported behavior information, online behavior information, and claims behavior information into text to obtain behavior information text; it then performs word segmentation and filtering on the behavior information text to obtain valid information text; it extracts keywords and phrases from the valid information text to obtain valid information phrases; finally, it obtains a preset summary template, fills the preset summary template with the valid information phrases, and generates the behavior summary. This effectively achieves comprehensive analysis based on various aspects of customer-related information to generate a reliable behavior summary, facilitating subsequent model prediction processing.
[0091] In some optional implementations of this embodiment, the step of performing dimensionality transformation on the behavior word embedding vector to obtain a low-dimensional behavior vector includes the following steps:
[0092] The word embedding vectors are standardized to obtain standard word embedding vectors;
[0093] In this embodiment, for each dimension (feature) of each word embedding vector, its mean μ and standard deviation σ across the entire corpus are calculated using the following formula: x 标准化 = (x-μ) / σ; where x is a certain dimension value of the original word embedding vector, x 标准化 : The standardized value. For example: Suppose the word embedding vector has a dimension of 3, the original vector is [1.0, 2.0, 3.0], and the mean of this dimension in the entire corpus is μ = 2.0, and the standard deviation is σ = 1.0. The standardized vector is: [(1.0-2.0) / 1.0, (2.0...] / 3 ... - 2.0) / 1.0, (3.0) - 2.0) / 1.0]=[-1.0,0.0,1.0].
[0094] The standard word embedding vector is reduced in dimensionality using PCA to obtain the low-dimensional vector of the behavior.
[0095] In this embodiment, PCA (Principal Component Analysis) is a dimensionality reduction method that projects high-dimensional data into a low-dimensional space through linear transformation while preserving the main information of the data. The steps of PCA dimensionality reduction include calculating the covariance matrix: for the standardized set of word embedding vectors, calculate its covariance matrix Σ; the formula for calculating the covariance matrix is Σ = X. T X / n; where X is the standardized word embedding matrix (each row is a word vector), and n is the number of word vectors. Calculate eigenvalues and eigenvectors: Perform eigenvalue decomposition on the covariance matrix Σ to obtain the eigenvalues λ. i and the corresponding feature vector v i Principal component selection: Sort the eigenvalues from largest to smallest and select the top k eigenvectors (corresponding to the top k principal components). Dimensionality reduction projection: Project the original word embedding vectors onto the selected principal components to obtain a low-dimensional vector. The projection formula is: X 低维 =X*V k Among them, V k It is a matrix composed of the first k eigenvectors.
[0096] This embodiment standardizes the word embedding vectors to obtain standard word embedding vectors; then, it uses PCA to reduce the dimensionality of these standard word embedding vectors, resulting in low-dimensional behavioral vectors. This effectively reduces the dimensionality of the word embedding vectors, facilitating subsequent clustering analysis.
[0097] In some optional implementations of this embodiment, the step of performing cluster analysis on the low-dimensional behavioral vector to obtain customer classification results, and then pushing information based on the customer classification results, includes the following steps:
[0098] The low-dimensional vectors of the behavior are clustered according to a preset clustering algorithm to obtain the vector clustering results;
[0099] In this embodiment, the preset clustering algorithm adopts the K-Means clustering algorithm. By setting the number of clusters (K value of K-Means) or density threshold (ε and MinPts of DBSCAN), the low-dimensional vector is input into the K-Means clustering algorithm model, thereby effectively obtaining the vector clustering results.
[0100] Obtain the customer data table and extract the customer ID information corresponding to the vector clustering results. Map the customer ID information to the customer data table to obtain the clustered customer information.
[0101] In this embodiment, the customer data table includes fields such as customer ID, corresponding low-dimensional behavioral vectors, and specific customer information. The vector clustering result includes cluster labels for each vector. The customer ID list corresponding to each cluster is extracted using the vector clustering result. For example, assuming the cluster labels are 0, 1, and 2, the customer IDs might be distributed as follows: Cluster 0: Customer A, Customer B; Cluster 1: Customer C, Customer D; Cluster 2: Customer E. Then, based on the extracted customer ID information, the specific customer information (such as name, gender, age, etc.) is mapped and retrieved from the customer data table, thus obtaining detailed customer information for each cluster. The detailed customer information for all clusters is then organized to obtain clustered customer information.
[0102] Cluster feature analysis is performed on the clustered customer information to obtain feature analysis categories, and the feature analysis categories and the clustered customer information are integrated to obtain the customer classification result;
[0103] In this embodiment, feature statistical analysis is performed on the customer information of each cluster to extract common features (such as spending power, behavioral preferences, geographical location, etc.) and obtain feature analysis categories. For example: Cluster 0: high consumption, first-tier city, high proportion of males; Cluster 1: medium consumption, high proportion of females, preference for promotional activities; Cluster 2: low consumption, student group, frequent online shopping. The feature analysis categories (feature descriptions of each cluster) and clustered customer information (detailed customer information for each cluster) are integrated to obtain the customer classification results.
[0104] Obtain the category push information corresponding to the customer classification result, and push information to the customers corresponding to the customer classification result based on the category push information.
[0105] In this embodiment, the category-based push information is designed based on the characteristics of customer classification results (such as coupons, event notifications, product recommendations, etc.). The push information is tailored to the specific category of the customer classification result. For example, cluster 0 pushes car insurance product discount information, cluster 1 pushes limited-time promotional activities, and cluster 2 pushes exclusive offers, etc.
[0106] This embodiment clusters the low-dimensional vectors of behavior according to a preset clustering algorithm to obtain the vector clustering results; it obtains a customer data table and extracts the customer ID information corresponding to the vector clustering results, maps and searches the customer ID information in the customer data table to obtain clustered customer information; it performs clustering feature analysis on the clustered customer information to obtain feature analysis categories, and integrates the feature analysis categories and the clustered customer information to obtain the customer classification results; it obtains the category push information corresponding to the customer classification results, and pushes information to the customers corresponding to the customer classification results according to the category push information. This effectively achieves customer classification based on the vector clustering results to obtain accurate customer classification results and enables message push based on the customer classification results.
[0107] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When executed, the program can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0108] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0109] Further reference Figure 3 As a response to the above Figure 1 The implementation of the method shown in this application provides an embodiment of a categorized message push device, which is similar to... Figure 1Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0110] like Figure 3 As shown, the categorized message push device 800 described in this embodiment includes: a first information extraction module 801, a second information extraction module 802, a third information extraction module 803, an information processing module 804, a model processing module 805, a dimension conversion module 806, and a result analysis module 807. Wherein:
[0111] The first information extraction module 801 is used to obtain first customer information, extract reporting information from the first customer information, and obtain reporting behavior information.
[0112] The second information extraction module 802 is used to obtain second customer information, perform online information extraction on the second customer information, and obtain online behavior information.
[0113] The third information extraction module 803 is used to obtain third customer information, extract claims information from the third customer information, and obtain claims behavior information.
[0114] Information processing module 804 is used to perform text processing on the reporting behavior information, the online behavior information, and the claims behavior information based on word engineering, and generate behavior summaries;
[0115] The model processing module 805 is used to input the behavior summary into a pre-trained language model to obtain behavior word embedding vectors;
[0116] The dimension transformation module 806 is used to perform dimension transformation processing on the behavior word embedding vector to obtain a low-dimensional behavior vector.
[0117] The result analysis module 807 is used to perform cluster analysis on the low-dimensional vector of behavior to obtain customer classification results, and to push information based on the customer classification results.
[0118] This embodiment, by employing the aforementioned categorized message push device, can acquire first customer information, extract reporting information from the first customer information to obtain reporting behavior information; acquire second customer information, extract online information from the second customer information to obtain online behavior information; acquire third customer information, extract claim information from the third customer information to obtain claim behavior information; perform text processing on the reporting behavior information, the online behavior information, and the claim behavior information based on word engineering to generate a behavior summary; input the behavior summary into a pre-trained language model to obtain behavior word embedding vectors; perform dimensionality transformation processing on the behavior word embedding vectors to obtain behavior low-dimensional vectors; perform cluster analysis on the behavior low-dimensional vectors to obtain customer classification results, and push information based on the customer classification results. This effectively achieves comprehensive analysis and push based on multiple dimensions of customer information, thereby improving the accuracy and effectiveness of the push notifications.
[0119] In some optional implementations of this embodiment, the first information extraction module 801 includes: a first information extraction unit, a first information processing unit, and a first feature extraction unit.
[0120] in:
[0121] The first information extraction unit is used to obtain a first information extraction identifier and extract the customer's non-vehicle quotation information and the customer's insurance information from the database according to the first information extraction identifier;
[0122] The first information processing unit is used to perform data cleaning processing on the customer non-vehicle quotation information and the customer insurance information to obtain standard non-vehicle quotation information and standard customer insurance information;
[0123] The first feature extraction unit is used to extract features from the standard non-vehicle pricing information and the standard customer insurance information according to predefined insurance behavior feature rules, so as to obtain the insurance behavior information.
[0124] This embodiment effectively acquires customer quotes and insurance-related information by setting up a first information extraction module 801, which includes a first information extraction unit, a first information processing unit, and a first feature extraction unit, so as to provide reliable data basis for subsequent processing.
[0125] In some optional implementations of this embodiment, the second information extraction module 802 includes: a second information extraction unit, a second information processing unit, and a second feature extraction unit.
[0126] in:
[0127] The second information extraction unit is used to obtain a second information extraction identifier and extract the vehicle owner's online information from the database based on the second information extraction identifier;
[0128] The second information processing unit is used to acquire time range definition information, and to filter the online information of vehicle owners according to the time range definition information to obtain valid online information of vehicle owners;
[0129] The second feature extraction unit is used to extract features from the valid online information of the vehicle owner according to predefined online behavior feature rules, so as to obtain the online behavior information.
[0130] This embodiment effectively acquires online information related to the customer's vehicle owner by setting up a second information extraction module 802, which includes a second information extraction unit, a second information processing unit, and a second feature extraction unit, so as to provide reliable data basis for subsequent processing.
[0131] In some optional implementations of this embodiment, the third information extraction module 803 includes: a third information extraction unit, a third information processing unit, and a third feature extraction unit.
[0132] in:
[0133] The third information extraction unit is used to obtain a third information extraction identifier and extract the customer's non-vehicle accident information and the customer's claim information from the database based on the third information extraction identifier.
[0134] The third information processing unit is used to extract key fields from the customer's non-vehicle accident information and the customer's claim information to obtain standard accident information and standard claim information.
[0135] The third feature extraction unit is used to extract features from the standard accident information and the standard claim information according to predefined claim behavior feature rules to obtain the claim behavior information.
[0136] This embodiment effectively obtains information related to customers' non-vehicle accidents and claims by setting up a third information extraction module 803, which includes a third information extraction unit, a third information processing unit, and a third feature extraction unit, so as to provide reliable data basis for subsequent processing.
[0137] In some optional implementations of this embodiment, the information processing module 804 includes: a text conversion unit, a word segmentation and filtering unit, a word and sentence extraction unit, and a summary generation unit.
[0138] in:
[0139] The text conversion unit is used to convert the reporting behavior information, the online behavior information, and the claims behavior information into text to obtain behavior information text;
[0140] The word segmentation and filtering unit is used to perform word segmentation and filtering on the behavioral information text to obtain effective information text;
[0141] The word and phrase extraction unit is used to extract keywords and phrases from the effective information text to obtain effective information words and phrases;
[0142] The summary generation unit is used to obtain a preset summary template, fill the effective information words and phrases into the preset summary template, and generate the behavior summary.
[0143] This embodiment effectively achieves comprehensive analysis based on various aspects of customer-related information to generate reliable behavioral summaries, facilitating subsequent model prediction processing, by setting up an information processing module 804 that includes a text conversion unit, a word segmentation and filtering unit, a word and sentence extraction unit, and a summary generation unit.
[0144] In some optional implementations of this embodiment, the model processing module 806 includes: a vector normalization unit and a vector dimensionality reduction unit. Wherein:
[0145] The vector normalization unit is used to normalize the word embedding vector to obtain a standard word embedding vector;
[0146] The vector dimensionality reduction unit is used to perform dimensionality reduction processing on the standard word embedding vector based on PCA to obtain the behavior low-dimensional vector.
[0147] This embodiment effectively reduces the dimensionality of word embedding vectors by setting up a model processing module 806 that includes vector normalization units and vector dimensionality reduction units, so as to facilitate subsequent clustering analysis operations.
[0148] In some optional implementations of this embodiment, the dimension conversion module 807 includes: a vector clustering unit, a customer search unit, an information organization unit, and an information push unit.
[0149] in:
[0150] The vector clustering unit is used to perform clustering processing on the low-dimensional vectors of the behavior according to a preset clustering algorithm to obtain the vector clustering result.
[0151] The customer lookup unit is used to obtain a customer data table, extract customer ID information corresponding to the vector clustering results, map the customer ID information to the customer data table, and obtain clustered customer information.
[0152] The information processing unit is used to perform cluster feature analysis on the clustered customer information to obtain feature analysis categories, and integrate the feature analysis categories and the clustered customer information to obtain the customer classification result;
[0153] The information push unit is used to obtain category push information corresponding to the customer classification result, and push information to the customers corresponding to the customer classification result according to the category push information.
[0154] This embodiment effectively classifies customers based on the clustering results of vectors by setting up a dimension transformation module 807 that includes a vector clustering unit, a customer search unit, an information organization unit, and an information push unit, so as to obtain accurate customer classification results and push messages based on the customer classification results.
[0155] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0156] The computer device 9 includes a memory 91, a processor 92, and a network interface 93 that are interconnected via a system bus. It should be noted that only the computer device 9 with components 91-93 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0157] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0158] The memory 91 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 91 may be an internal storage unit of the computer device 9, such as the hard disk or memory of the computer device 9. In other embodiments, the memory 91 may also be an external storage device of the computer device 9, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 9. Of course, the memory 91 may also include both the internal storage unit and its external storage device of the computer device 9. In this embodiment, the memory 91 is typically used to store the operating system and various application software installed on the computer device 9, such as computer-readable instructions for categorized message push methods. In addition, the memory 91 can also be used to temporarily store various types of data that have been output or will be output.
[0159] In some embodiments, the processor 92 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 92 is typically used to control the overall operation of the computer device 9. In this embodiment, the processor 92 is used to execute computer-readable instructions stored in the memory 91 or to process data, for example, to execute computer-readable instructions for the categorized message push method.
[0160] The network interface 93 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 9 and other electronic devices.
[0161] This embodiment, by employing the aforementioned computer equipment, can acquire first customer information, extract reporting information from the first customer information to obtain reporting behavior information; acquire second customer information, extract online information from the second customer information to obtain online behavior information; acquire third customer information, extract claims information from the third customer information to obtain claims behavior information; perform text processing on the reporting behavior information, the online behavior information, and the claims behavior information based on word engineering to generate a behavior summary; input the behavior summary into a pre-trained language model to obtain behavior word embedding vectors; perform dimensionality transformation processing on the behavior word embedding vectors to obtain low-dimensional behavior vectors; perform cluster analysis on the low-dimensional behavior vectors to obtain customer classification results, and push information based on the customer classification results. This effectively achieves comprehensive analysis and push based on multiple dimensions of customer information, thereby improving the accuracy and effectiveness of the push notifications.
[0162] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the classified message push method described above.
[0163] This embodiment, by employing the aforementioned computer-readable storage medium, can acquire first customer information, extract reporting information from the first customer information to obtain reporting behavior information; acquire second customer information, extract online information from the second customer information to obtain online behavior information; acquire third customer information, extract claims information from the third customer information to obtain claims behavior information; perform text processing on the reporting behavior information, the online behavior information, and the claims behavior information based on word engineering to generate a behavior summary; input the behavior summary into a pre-trained language model to obtain behavior word embedding vectors; perform dimensionality transformation processing on the behavior word embedding vectors to obtain low-dimensional behavior vectors; perform cluster analysis on the low-dimensional behavior vectors to obtain customer classification results, and push information based on the customer classification results. This effectively achieves comprehensive analysis and push based on multiple dimensions of customer information, thereby improving the accuracy and effectiveness of the push notifications.
[0164] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0165] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.
[0166] The software tools or components not belonging to our company that appear in the embodiments of this application are merely examples and do not represent actual use.
Claims
1. A method for categorizing and pushing messages, characterized in that, Includes the following steps: Obtain first customer information, extract reporting information from the first customer information, and obtain reporting behavior information; Obtain second customer information, and extract online information from the second customer information to obtain online behavior information; Obtain third-party customer information, extract claims information from the third-party customer information, and obtain claims behavior information; Based on word engineering, text processing is performed on the reporting behavior information, the online behavior information, and the claims behavior information to generate behavior summaries; The behavior summary is input into a pre-trained language model to obtain behavior word embedding vectors; The behavior word embedding vector is subjected to dimensionality transformation to obtain a low-dimensional behavior vector; Cluster analysis is performed on the low-dimensional vectors of the behaviors to obtain customer classification results, and information is pushed based on the customer classification results.
2. The categorized message push method according to claim 1, characterized in that, The first customer information includes customer non-vehicle quotation information and customer insurance information. The step of obtaining the first customer information, extracting quotation information from the first customer information, and obtaining quotation behavior information specifically includes: Obtain a first information extraction identifier, and extract the customer's non-vehicle quotation information and the customer's insurance information from the database based on the first information extraction identifier; The customer's non-vehicle pricing information and the customer's insurance information are cleaned to obtain standard non-vehicle pricing information and standard customer insurance information. Based on predefined insurance behavior feature rules, feature extraction is performed on the standard non-vehicle price information and the standard customer insurance information to obtain the insurance behavior information.
3. The categorized message push method according to claim 1, characterized in that, The second customer information includes the vehicle owner's online information. The steps of obtaining the second customer information and extracting online information from it to obtain online behavior information specifically include: Obtain the second information extraction identifier, and extract the vehicle owner's online information from the database based on the second information extraction identifier; Obtain time range definition information, and filter the online vehicle owner information based on the time range definition information to obtain valid online vehicle owner information; The online behavior information is obtained by extracting features from the valid online information of the vehicle owners according to the predefined online behavior feature rules.
4. The categorized message push method according to claim 1, characterized in that, The third customer information includes customer non-motor vehicle accident information and customer claim information. The steps of obtaining the third customer information and extracting claim information from the third customer information to obtain claim behavior information specifically include: Obtain a third information extraction identifier, and extract the customer's non-vehicle accident information and customer's claims information from the database based on the third information extraction identifier; Key fields are extracted from the customer's non-vehicle accident information and the customer's claims information to obtain standard accident information and standard claims information; The standard accident information and the standard claim information are feature extracted according to predefined claim behavior feature rules to obtain the claim behavior information.
5. The categorized message push method according to claim 1, characterized in that, The step of performing text processing on the reporting behavior information, the online behavior information, and the claims behavior information based on word engineering to generate behavior summaries specifically includes: The reported behavior information, the online behavior information, and the claims behavior information are converted into text to obtain behavior information text. The behavioral information text is segmented and filtered to obtain valid information text; Keyword extraction is performed on the effective information text to obtain effective information phrases; Obtain a preset summary template, fill the effective information phrases into the preset summary template, and generate the behavior summary.
6. The categorized message push method according to claim 1, characterized in that, The step of performing dimensionality transformation on the behavior word embedding vector to obtain a low-dimensional behavior vector specifically includes: The word embedding vectors are standardized to obtain standard word embedding vectors; The standard word embedding vector is reduced in dimensionality using PCA to obtain the low-dimensional vector of the behavior.
7. The categorized message push method according to claim 1, characterized in that, The step of performing cluster analysis on the low-dimensional vectors of the behaviors to obtain customer classification results, and then pushing information based on the customer classification results, specifically includes: The low-dimensional vectors of the behavior are clustered according to a preset clustering algorithm to obtain the vector clustering results; Obtain the customer data table and extract the customer ID information corresponding to the vector clustering results. Map the customer ID information to the customer data table to obtain the clustered customer information. Cluster feature analysis is performed on the clustered customer information to obtain feature analysis categories, and the feature analysis categories and the clustered customer information are integrated to obtain the customer classification result; Obtain the category push information corresponding to the customer classification result, and push information to the customers corresponding to the customer classification result based on the category push information.
8. A categorized message push device, characterized in that, include: The first information extraction module is used to obtain first customer information, extract reporting information from the first customer information, and obtain reporting behavior information. The second information extraction module is used to obtain second customer information, extract online information from the second customer information, and obtain online behavior information. The third information extraction module is used to obtain third customer information, extract claims information from the third customer information, and obtain claims behavior information. The information processing module is used to perform text processing on the reporting behavior information, the online behavior information, and the claims behavior information based on word engineering, and generate behavior summaries; The model processing module is used to input the behavior summary into a pre-trained language model to obtain behavior word embedding vectors; The dimension transformation module is used to perform dimension transformation processing on the behavior word embedding vector to obtain a low-dimensional behavior vector. The results analysis module is used to perform cluster analysis on the low-dimensional vector of behavior to obtain customer classification results, and to push information based on the customer classification results.
9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the categorized message push method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the categorized message push method as described in any one of claims 1 to 7.