Data analysis method and device, computer device and storage medium
Patent Information
- Application Number
- CN202610790238.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本申请实施例的目的在于提出一种数据分析方法、装置、计算机设备及存储介质,以解决现有的代理人数据分析方法主要依赖于固定的规则模板,导致分析的准确性和智能性较低,进而无法为代理人提供精准的个性化业务指导建议的技术问题
[0010]上述数据分析方法、装置、计算机设备及存储介质所实现的方案中,首先从预设的多个业务系统获取用户的多维度数据;并基于预设的处理策略对所述多维度数据进行数据预处理,得到对应的目标数据;然后获取与所述用户对应的目标群体类型;并调用与所述目标群体类型对应的目标关联分析模型;之后基于所述目标关联分析模型对所述目标数据进行关联分析处理,得到对应的分析结果;后续从预设的建议知识库中获取与所述分析结果对应的目标优化建议;最后将所述目标优化建议发送给所述用户。基于以上的自动化处理流程,不同于现有的依赖于固定的规则模板的数据分析方法,本申请提供了一种智能化的新型数据分析方法,通过对从多个业务系统获取的用户的多维度数据进行数据预处理得到目标数据,然后根据用户的目标群体类型调用对应的目标关联分析模型,并基于目标关联分析模型的使用对目标数据进行关联分析以得到分析结果,进而基于建议知识库的使用获取与分析结果对应的目标优化建议并发送给用户。如此,本申请通过基于目标关联分析模型的使用对多维度的目标数据进行关联分析以得到分析结果,可以有效地提高数据分析的准确性和智能性。并基于建议知识库的使用生成与分析结果匹配的目标优化建议,从而能够为用户提供精准的个性化建议,提高了生成的目标优化建议的精准度。
Smart Images

Figure CN122840745A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology and can be applied to the financial technology field, particularly to data analysis methods, devices, computer equipment and storage media. Background Technology
[0002] In the financial and insurance sector, traditional methods for evaluating and analyzing agents' business capabilities and customer service quality rely heavily on fixed rule templates. This results in low accuracy and intelligence in the analysis, making it impossible to provide agents with precise, personalized business guidance. Specifically, traditional methods typically set fixed evaluation rules based on agents' basic performance indicators, lacking in-depth mining and comprehensive analysis of agents' multi-dimensional behavioral data. This extensive analysis approach struggles to accurately identify agents' true weaknesses and potential strengths in their business development, leading to generic improvement suggestions that fail to meet the specific business scenarios and development needs of different agents.
[0003] For example, in the context of agent performance analysis in the financial insurance sector, traditional methods might categorize agents as "excellent" or "needs improvement" solely based on their monthly sales volume, without considering other key factors. If an agent has a low sales volume but a high customer renewal rate and accurate customer profiling, traditional methods might misjudge them as inefficient and offer unreasonable suggestions like "increase visit frequency," ignoring their strengths in in-depth customer management. This analytical bias not only misleads agents' business improvement strategies but also leads to a misallocation of human resources for insurance companies, reducing overall business efficiency.
[0004] Therefore, there is an urgent need to provide an intelligent agent data analysis method to improve the accuracy and intelligence of the analysis, meet the personalized development needs of agents, and enhance the overall service efficiency of insurance business. Summary of the Invention
[0005] The purpose of this application is to propose a data analysis method, apparatus, computer equipment, and storage medium to solve the technical problem that existing agent data analysis methods mainly rely on fixed rule templates, resulting in low accuracy and intelligence of the analysis, and thus failing to provide agents with accurate and personalized business guidance suggestions.
[0006] Firstly, a data analysis method is provided, including: Obtain multi-dimensional user data from multiple pre-set business systems; Based on a preset processing strategy, the multi-dimensional data is preprocessed to obtain the corresponding target data; Obtain the target group type corresponding to the user; Invoke the target association analysis model corresponding to the target group type; Based on the target association analysis model, the target data is subjected to association analysis processing to obtain the corresponding analysis results; Obtain target optimization suggestions corresponding to the analysis results from a pre-set suggestion knowledge base; The proposed optimization suggestions are sent to the user.
[0007] Secondly, a data analysis device is provided, comprising: The first acquisition module is used to acquire multi-dimensional user data from multiple preset business systems; The preprocessing module is used to preprocess the multi-dimensional data based on a preset processing strategy to obtain the corresponding target data. The second acquisition module is used to acquire the target group type corresponding to the user; The calling module is used to call the target association analysis model corresponding to the target group type; The analysis module is used to perform association analysis on the target data based on the target association analysis model to obtain the corresponding analysis results; The third acquisition module is used to acquire target optimization suggestions corresponding to the analysis results from a preset suggestion knowledge base; The sending module is used to send the target optimization suggestions to the user.
[0008] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described data analysis method.
[0009] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described data analysis method.
[0010] The above-described data analysis method, apparatus, computer equipment, and storage medium implement a solution that first obtains multi-dimensional user data from multiple preset business systems; then, based on a preset processing strategy, preprocesses the multi-dimensional data to obtain corresponding target data; next, it obtains the target group type corresponding to the user; and calls the target association analysis model corresponding to the target group type; then, it performs association analysis processing on the target data based on the target association analysis model to obtain corresponding analysis results; subsequently, it obtains target optimization suggestions corresponding to the analysis results from a preset suggestion knowledge base; and finally, it sends the target optimization suggestions to the user. Based on the above automated processing flow, unlike existing data analysis methods that rely on fixed rule templates, this application provides an intelligent new data analysis method. It obtains target data by preprocessing multi-dimensional user data obtained from multiple business systems, then calls the corresponding target association analysis model according to the user's target group type, and performs association analysis on the target data based on the use of the target association analysis model to obtain analysis results. Finally, it obtains target optimization suggestions corresponding to the analysis results based on the use of the suggestion knowledge base and sends them to the user. Thus, this application, by using a target association analysis model to perform association analysis on multi-dimensional target data to obtain analysis results, can effectively improve the accuracy and intelligence of data analysis. Furthermore, by using a suggestion knowledge base to generate target optimization suggestions that match the analysis results, it can provide users with accurate and personalized suggestions, thereby improving the accuracy of the generated target optimization suggestions. Attached Figure Description
[0011] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 This is a flowchart of an embodiment of the data analysis method according to this application; Figure 3 This is a schematic diagram of the structure of one embodiment of the data analysis apparatus according to this application; Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation
[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0014] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0015] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0016] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0017] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0018] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.
[0019] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.
[0020] It should be noted that the data analysis method provided in the embodiments of this application is generally executed by a server / terminal device, and correspondingly, the data analysis device is generally located in the server / terminal device.
[0021] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0022] Continue to refer to Figure 2 A flowchart illustrating an embodiment of the data analysis method according to this application is shown. The order of steps in the flowchart can be changed, and some steps can be omitted, depending on different needs. The data analysis method provided by this application embodiment can be applied to any scenario requiring data analysis, and thus can be applied to products in these scenarios, such as data analysis products in the financial insurance field. The data analysis method includes the following steps: Step S201: Obtain multi-dimensional user data from multiple preset business systems.
[0023] In this embodiment, the data analysis method runs on an electronic device (e.g., Figure 1The server / terminal device shown can acquire multi-dimensional user data via wired or wireless connections. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wideband) connections, and other currently known or future-developed wireless connection methods. The executing entity of this application is specifically a data analysis system, which can be simply referred to as the system. This application can be applied to business data analysis scenarios related to insurance agents in the financial and insurance field. The user may refer to an insurance agent, which can be simply referred to as an agent. The aforementioned multi-dimensional data may refer to the agent's data for this month. The specific implementation process of acquiring multi-dimensional user data from multiple preset business systems will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated upon here.
[0024] Step S202: Based on a preset processing strategy, perform data preprocessing on the multi-dimensional data to obtain the corresponding target data.
[0025] In this embodiment, the specific implementation process of preprocessing the multi-dimensional data based on the preset processing strategy to obtain the corresponding target data will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0026] Step S203: Obtain the target group type corresponding to the user.
[0027] In this embodiment, the specific implementation process of obtaining the target group type corresponding to the user will be described in more detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0028] Step S204: Invoke the target association analysis model corresponding to the target group type.
[0029] In this embodiment, based on actual business needs, correlation analysis models corresponding to different group types are pre-constructed. Furthermore, a universal three-dimensional correlation analysis model corresponding to all agents is pre-constructed. The model construction process for the three-dimensional correlation analysis model will be described in further detail in subsequent specific embodiments, and will not be elaborated upon here.
[0030] The aforementioned target association analysis model refers to an association analysis model that matches the target group type of the aforementioned users. Specifically, this can be achieved by obtaining a pre-constructed universal three-dimensional association analysis model corresponding to all agents. Then, the target group type of the aforementioned users (newcomer group / high-performing group / inefficient group) is obtained, and the analysis weights of the aforementioned three-dimensional association analysis model are adjusted according to the target group type to obtain a target association analysis model that matches the target group type of the aforementioned users.
[0031] For example, if the users are newcomers, since the analysis of newcomers focuses on the behavioral process dimension, the system configures the first analysis weight for newcomers as follows: 70% for behavioral characteristics, 20% for performance characteristics, and 10% for compliance characteristics. Furthermore, the weights of the aforementioned three-dimensional correlation analysis model can be adjusted using this first analysis weight to obtain a target correlation analysis model that matches the user's target group type (newcomers).
[0032] If the user is a high-performing group, since the analysis of high-performing groups focuses on performance maintenance and compliance, the system assigns the following second analysis weights to this group: 50% for performance characteristics, 30% for behavioral characteristics, and 20% for compliance characteristics. Furthermore, these second analysis weights can be used to adjust the weights of the aforementioned three-dimensional correlation analysis model, thereby obtaining a target correlation analysis model that matches the user's target group type (high-performing group).
[0033] If the user group is an inefficient group, since the analysis of inefficient groups focuses on identifying behavioral shortcomings, the system configures the following third analysis weights for inefficient groups: 60% for behavioral characteristics, 25% for performance characteristics, and 15% for compliance characteristics. Furthermore, by using these third analysis weights, the weights of the above three-dimensional correlation analysis model can be adjusted to obtain a target correlation analysis model that matches the user's target group type (inefficient group).
[0034] Step S205: Perform association analysis on the target data based on the target association analysis model to obtain the corresponding analysis results.
[0035] In this embodiment, the specific implementation process of performing association analysis on the target data based on the target association analysis model to obtain the corresponding analysis results will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0036] Step S206: Obtain target optimization suggestions corresponding to the analysis results from a preset suggestion knowledge base.
[0037] In this embodiment, the specific implementation process of obtaining the target optimization suggestions corresponding to the analysis results from the preset suggestion knowledge base will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0038] Step S207: Send the target optimization suggestion to the user.
[0039] In this embodiment, the generated target optimization suggestions can be sent to the aforementioned users through methods such as email, message, or interface display.
[0040] This application first obtains multi-dimensional user data from multiple pre-defined business systems; then preprocesses the multi-dimensional data based on a pre-defined processing strategy to obtain corresponding target data; next, it obtains the target group type corresponding to the user; and calls the target association analysis model corresponding to the target group type; then, it performs association analysis processing on the target data based on the target association analysis model to obtain corresponding analysis results; subsequently, it obtains target optimization suggestions corresponding to the analysis results from a pre-defined suggestion knowledge base; finally, it sends the target optimization suggestions to the user. Based on the above automated processing flow, unlike existing data analysis methods that rely on fixed rule templates, this application provides an intelligent new data analysis method. It obtains target data by preprocessing multi-dimensional user data obtained from multiple business systems, then calls the corresponding target association analysis model according to the user's target group type, and performs association analysis on the target data based on the use of the target association analysis model to obtain analysis results. Finally, it obtains target optimization suggestions corresponding to the analysis results based on the use of the suggestion knowledge base and sends them to the user. Thus, this application, by performing association analysis on multi-dimensional target data based on the use of the target association analysis model to obtain analysis results, can effectively improve the accuracy and intelligence of data analysis. Furthermore, based on the use of the suggestion knowledge base, target optimization suggestions are generated and matched with the analysis results, thereby providing users with accurate and personalized suggestions and improving the accuracy of the generated target optimization suggestions.
[0041] In some optional implementations of this embodiment, step S201 includes the following steps: Step S2011: Obtain behavioral data, performance data, and compliance data corresponding to the user from the multiple business systems; In this embodiment, the aforementioned multiple business systems specifically include a pre-defined customer relationship management system, a performance management system, and a compliance management system. The process of obtaining behavioral data, performance data, and compliance data corresponding to the user from these multiple business systems includes: 1) By retrieving behavioral data corresponding to the user from the Customer Relationship Management System (CRM). Specifically, the system extracts user-related behavioral data from the CRM system (i.e., Customer Relationship Management System) through an interface or direct database connection. This behavioral data includes: daily customer visit records, with fields including visit date, number of visits, duration of each visit, and visit frequency distribution (i.e., the distribution of the agent's visit frequency on each day of the week); communication records with customers, with fields including call date, call duration, number of communication rounds (i.e., the number of back-and-forth exchanges in a single customer communication), communication quality score (scored post-event by the quality inspection system, out of ten), and script standardization (automatically determined by the voice recognition system, expressed as a percentage); and customer maintenance records, with fields including proactive follow-up date, number of proactive follow-ups, customer satisfaction score (out of ten), and customer churn marker (yes or no).
[0042] The aforementioned communication quality score refers to the scoring given by quality control personnel or intelligent quality control systems to the agent for the quality of each phone call or visit. Fields include the agent's identification code, scoring date, communication quality score (out of 10), and detailed scoring items (such as opening remarks score, product introduction score, and objection handling score). The aforementioned script compliance data refers to the semantic analysis performed by the voice recognition system on the agent's recorded calls to determine whether their scripts conform to company standards. Fields include the agent's identification code, call date, script compliance percentage, and type of inappropriate script (such as failure to provide risk warnings or exaggerating benefits).
[0043] 2) Obtain performance data corresponding to the user from the performance management system. Specifically, the system extracts user-related performance data from the performance management system through interfaces or direct database connections. This data includes: monthly premium income (in yuan), monthly number of policies issued, monthly renewal rate (expressed as a percentage), monthly surrender rate (expressed as a percentage), new customer conversion rate (expressed as a percentage), and average premium per customer (in yuan).
[0044] 3) Obtain compliance data corresponding to the user from the compliance management system. Specifically, the system extracts four types of data from the compliance management system as compliance data through interfaces or direct database connections. The first type is violation operation records, with fields including agent identification code, violation date, violation type (e.g., forged signature, commission rebate, misleading sales, etc.), and number of violations. The second type is the speech script standardization score (different from the speech script standardization score in the quality inspection system; this is a comprehensive score independently assessed by the compliance department, with a maximum score of 100 points). The third type is customer complaint records, with fields including agent identification code, complaint date, number of complaints, and complaint type. The fourth type is regulatory penalty records, with fields including agent identification code, penalty date, penalty type, and penalty level.
[0045] Step S2012: Based on the user's identification code, perform data alignment processing on the behavioral data, the performance data, and the compliance data to obtain the corresponding generated data.
[0046] In this embodiment, the system associates all data extracted from the three types of systems mentioned above using the agent's unique identifier as the primary key. Data from the same agent in the same month is merged into a single record. Specifically, the system first generates a master table framework with "agent identifier" and "month" as composite primary keys. Then, it aggregates and fills in the behavioral data from the CRM system by agent identifier and month, fills in the performance data from the performance management system by agent identifier and month, and fills in the violation and complaint data from the compliance system by agent identifier and month. The resulting original dataset (i.e., multi-dimensional) has each row representing all data for a specific agent in a given month, with columns containing all behavioral, performance, and compliance fields for that agent in that month.
[0047] Step S2013: Use the generated data as the multi-dimensional data.
[0048] This application obtains user-specific behavioral data, performance data, and compliance data from multiple business systems; then, based on the user's identifier, it performs data alignment processing on the behavioral data, performance data, and compliance data to obtain corresponding generated data; subsequently, the generated data is used as multi-dimensional data. Based on the above processing flow, this application extracts and aggregates three types of core data—behavioral data, performance data, and compliance data—from multiple business systems within the enterprise, according to a unified time period and a unified primary key identifier. This allows for the efficient and accurate integration of the three types of data into corresponding multi-dimensional data, improving the accuracy and richness of the generated multi-dimensional data.
[0049] In some alternative implementations, step S202 includes the following steps: Step S2021: Perform missing value processing on the multi-dimensional data to obtain the corresponding first processed data.
[0050] In this embodiment, the system checks the missing values of each field in the multi-dimensional data one by one, and adopts different imputation strategies according to the field type and degree of missing value. The missing value handling process includes: For missing behavioral data: If an agent's visit records are missing for a few days, and the consecutive missing days do not exceed three days, the system uses the mean of the previous and next days to fill the missing data. The specific calculation method is as follows: Suppose an agent's visit count is missing on day t, the number of visits on the previous day (day t-1) is V(t-1), and the number of visits on the next day (day t+1) is V(t+1). Then the fill value for day t is V_fill(t) = (V(t-1) + V(t+1)) / 2. If the missing data exceeds three consecutive days, no filling is performed; instead, the agent for that month is marked as having "incomplete behavioral data," and this sample is weighted less in subsequent analyses.
[0051] Regarding missing performance data: If an agent's surrender rate for a given month is not recorded, the system will use the median of that agent's surrender rates over the past six months to fill in the gaps. Let the agent's surrender rates for the past six months be R1, R2, R3, R4, R5, and R6. These six values will be sorted from smallest to largest, and the median will be used as the filler value. If all performance data for a given month is missing (i.e., premiums, number of policies, renewal rates, etc. are all empty), then that agent will be marked as an invalid sample for that month and directly removed from the dataset, not participating in any subsequent analysis.
[0052] Regarding missing compliance data: If an agent has no violations in a certain month, the system will default to zero violations, a perfect score in communication style, and zero complaints. However, the system will mark the agent with a "Compliance data source incomplete" label in the final report, reminding the administrator that the data is less reliable.
[0053] Step S2022: Perform outlier detection and correction on the first processed data to obtain the corresponding second processed data.
[0054] In this embodiment, the system performs outlier detection on the first processed data to identify and process obviously unreasonable data points, preventing them from distorting the model. The specific implementation process includes: For outliers in behavioral data: the system uses a rule based on three standard deviations of the individual's historical mean. Let the number of monthly visits by an agent over the past twelve months be V1, V2, ..., V12, with a historical mean of μ_V = (V1 + V2 + ... + V12) / 12 and a historical standard deviation of σ_V = sqrt(Σ(Vi - μ_V)² / 12). If the agent's monthly visit count V_current satisfies |V_current - μ_V| > 3 × σ_V, it is considered an outlier. There are two correction methods: the first is to replace the outlier with μ_V × 1.5 (i.e., 1.5 times the historical mean, as a reasonable upper limit); the second is not to replace the original value, but to assign a weight of 0.5 to the sample in subsequent modeling (i.e., deweighting), halving its impact on the model.
[0055] For outliers in performance data: If the premium income for a certain month is negative (usually caused by policy surrenders), the system does not delete the record directly, but splits it into two independent fields: positive premium (the total premium of new policies signed in that month) and surrender amount (the total premium of surrendered policies in that month), and includes them in the analysis separately.
[0056] Step S2023: Standardize the second processed data to obtain the corresponding third processed data.
[0057] In this embodiment, due to the significant differences in the dimensions of different features, directly including them in the model would cause features with larger dimensions to dominate the calculation results. For example, the numerical range of the number of visits is from zero to several hundred, the numerical range of call duration is from zero to several thousand minutes, and the numerical range of premium income is from zero to several hundred thousand yuan. Without processing, premium income would dominate the model due to its large value, and the influence of other features would be severely compressed.
[0058] Specifically, the system employs a min-max normalization method to map each continuous numerical feature to a range of zero to one. The specific formula is: X_normalized = (X_raw - X_min) / (X_max - X_min), where X_raw is the original value of the feature, X_min is the minimum value of the feature among all samples, and X_max is the maximum value of the feature among all samples. After normalization, all continuous features range from 0 to 1, allowing for weighted comparisons on the same scale.
[0059] Step S2024: Perform time-granularity unification processing on the third processed data to obtain the corresponding fourth processed data.
[0060] In this embodiment, the original behavioral data is recorded daily, while the performance data is summarized monthly. The two have inconsistent time granularities and must be unified to a monthly level for correlation analysis. Specifically, the system will aggregate all daily-recorded behavioral data monthly, using the following aggregation method: Total monthly visits = Σ(daily visits), which is the sum of the number of visits for all days in the month.
[0061] Total monthly call duration = Σ(daily call duration), which is the sum of call durations for all days in the month.
[0062] Monthly average communication quality score = Σ(daily communication quality score) / number of days with score records in the month, which is the arithmetic mean of the communication quality scores for all days with score records in the month.
[0063] Average monthly script standardization rate = Σ(daily script standardization percentage) / number of days with records in that month.
[0064] Monthly proactive follow-up visits = Σ(Daily proactive follow-up visits).
[0065] Monthly customer satisfaction score = Σ(daily customer satisfaction score) / number of customers with score records in that month.
[0066] In addition, to capture the dynamic trends of behavioral changes, the system also calculates the month-on-month change rate for each behavioral feature, adding it to the dataset as a dynamic feature. The formula for calculating the month-on-month change rate is: Month-on-month change rate = (Feature value this month - Feature value last month) / Feature value last month × 100%.
[0067] For example, if an agent's effective visit rate is 50% this month and 40% last month, then the month-on-month change rate is (0.50-0.40) / 0.40×100%=25%, indicating that the behavior is continuously improving.
[0068] Step S2025: Perform data encoding on the fourth processed data to obtain the corresponding fifth processed data.
[0069] In this embodiment, the raw data contains some categorical fields, such as the agent's team (e.g., Team A, Team B, Team C) and the main product type sold (e.g., life insurance, property insurance, health insurance). These fields are in text form and cannot be directly calculated by the association model; they must be converted into numerical form. The system uses one-hot encoding to handle this. For example, for the "team affiliation" field, if there are three teams, it is converted into three binary features: whether the agent belongs to Team A (1 if yes, 0 otherwise), whether the agent belongs to Team B (1 if yes, 0 otherwise), and whether the agent belongs to Team C (1 if yes, 0 otherwise). In this way, each categorical field is converted into a set of numerical features consisting of zeros and ones, which can be directly used by the model.
[0070] Step S2026: Use the fifth processed data as the target data.
[0071] This application obtains first-processed data by handling missing values in multi-dimensional data; then, outlier detection and correction are performed on the first-processed data to obtain second-processed data; next, standardization is applied to the second-processed data to obtain third-processed data; subsequently, time granularity unification is applied to the third-processed data to obtain fourth-processed data; further, data encoding is applied to the fourth-processed data to obtain fifth-processed data; finally, the fifth-processed data is used as the target data. Based on the above processing flow, this application first handles missing values in multi-dimensional data, then handles outliers, then performs dimensional standardization, then unifies the time granularity and aggregates features, and finally converts the category fields into numerical form. After these five sub-processes, a set of consistent, numerically comparable, and missing user-related feature vectors can be automatically and accurately generated, thereby effectively ensuring the accuracy and standardization of the generated target data.
[0072] In some alternative implementations, step S203 includes the following steps: Step S2031: Obtain the user's employment duration data and historical performance ranking data.
[0073] In this embodiment, user data can be queried based on the user's identification information (such as user name) according to the employment dimension to obtain the corresponding employment duration data. Simultaneously, user data can be queried based on the user's identification information (such as user name) according to historical performance ranking to obtain the corresponding historical performance ranking data.
[0074] Step S2032: Based on preset hierarchical rules, perform hierarchical analysis on the employment duration data and the historical performance ranking data to generate a target group type corresponding to the user.
[0075] In this embodiment, the aforementioned stratification rules refer to the rules for stratifying agents based on the dimensions of employment duration and historical performance ranking. The corresponding rule content may include: 1) Definition of the new agent group: Agents with employment duration between zero and six months, regardless of their performance, are all classified as new agents. This is because new agents have not yet established a stable business model, and their performance fluctuates greatly, making performance results unsuitable for evaluation. 2) Definition of the high-performing group: Agents whose performance ranking is consistently in the top 20% of all agents for the most recent three consecutive months. The system uses the three-month average for judgment to avoid misclassification due to occasional good performance in a single month. 3) Definition of the low-performing group: Agents whose performance ranking is consistently in the bottom 20% of all agents for the most recent three consecutive months. Again, the three-month average is used to avoid misjudgment due to occasional fluctuations. Furthermore, based on the above stratification rules, a group-matching stratification analysis can be performed on the aforementioned employment duration data and historical performance ranking data to generate target group types corresponding to users.
[0076] This application obtains users' work experience length data and historical performance ranking data; then, based on preset stratification rules, it performs stratified analysis on the work experience length data and historical performance ranking data to generate target group types corresponding to users. Based on the above processing flow, this application, by using stratification rules, can automatically and accurately complete the stratified analysis of users' work experience length data and historical performance ranking data, ensuring the accuracy of the generated target group types.
[0077] In some alternative implementations, step S205 includes the following steps: Step S2051: Perform feature engineering on the target data to obtain corresponding key behavioral features; wherein, the number of key behavioral features is multiple.
[0078] In this embodiment, the aforementioned feature engineering refers to extracting composite features from the aforementioned target data that better reflect the quality of behavior, and using them as corresponding key behavioral features (which can be simply referred to as behavioral features).
[0079] The first composite feature is the effective visit rate. Its calculation formula is: Effective Visit Rate = Number of visits lasting more than 15 minutes / Total number of visits per month × 100%. This feature reflects visit quality better than simply the "number of visits," because a high number of visits with only two minutes each contributes far less to performance than fewer visits with in-depth communication each time. The second composite feature is the overall communication score. Its calculation formula is: Overall Communication Score = Communication Quality Score × 0.6 + Percentage of Script Standardization × 0.4. This feature combines quality inspection scoring and voice recognition judgment, using a weighted average to form a comprehensive indicator. The weight allocation is 60% for communication quality and 40% for script standardization, because communication quality has a more direct impact on customer experience. The third composite feature is the customer maintenance index. Its calculation formula is: Customer Maintenance Index = Number of Proactive Follow-ups × 0.5 + Customer Satisfaction Score ÷ 10 × 0.5. This feature combines follow-up frequency and customer satisfaction, with each accounting for half the weight. This is because having many follow-up visits but no customer satisfaction, or having satisfied customers but no follow-up visits, does not necessarily indicate good customer maintenance.
[0080] Step S2052: Extract the influence weights corresponding to the key behavioral features from the target association analysis model.
[0081] In this embodiment, the influence weights corresponding to the key behavioral features can be extracted from the behavior-performance correlation model included in the target correlation analysis model.
[0082] Step S2053: Perform loss calculation processing on the key behavioral features and the influence weights to obtain the corresponding loss data.
[0083] In this embodiment, the specific implementation process of performing loss calculation on the key behavioral features and the influence weights to obtain the corresponding loss data will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0084] Step S2054: Sort the loss data in descending order to obtain the sorted target loss data.
[0085] In this embodiment, an ordered list is obtained by sorting the potential loss values (i.e., loss data) of all user behavior dimensions from largest to smallest. The sorting result is: Loss_1 ≥ Loss_2 ≥ Loss_3 ≥ … ≥ Loss_k. The behavior dimension corresponding to Loss_1, which ranks first, is the behavior weakness that has the greatest impact on the user's performance; the behavior dimension corresponding to Loss_2, which ranks second, is the secondary weakness, and so on.
[0086] Step S2055: Use the target loss data as the analysis result.
[0087] This application obtains key behavioral features by performing feature engineering on target data; there are multiple key behavioral features. Then, the influence weights corresponding to the key behavioral features are extracted from the target association analysis model. Next, loss calculation is performed on the key behavioral features and influence weights to obtain corresponding loss data. Subsequently, the loss data is sorted in descending order to obtain sorted target loss data. Finally, the target loss data is used as the analysis result. Based on the above processing flow, this application obtains key behavioral features by performing feature engineering on target data, extracts the influence weights corresponding to the key behavioral features from the target association analysis model, performs loss calculation on the key behavioral features and influence weights, and sorts the obtained loss data as the final target loss data. This enables automatic and accurate association analysis of target data based on the target association analysis model, improving the processing efficiency and intelligence of association analysis and ensuring the accuracy of the generated analysis results.
[0088] In some optional implementations, the process of constructing a general three-dimensional correlation analysis model includes: Step 1: Multi-dimensional Raw Data Collection. By extracting all three types of core data (behavioral data, performance data, and compliance data) related to all agents from multiple internal business systems, using a unified time period and primary key identifier, and aggregating them into a single data pool to obtain the raw dataset. The specific implementation method is similar to the aforementioned process of obtaining multi-dimensional user data from multiple pre-defined business systems, and will not be elaborated upon here.
[0089] The second step involves preprocessing the raw datasets from multiple business systems to obtain clean feature vectors, which will then serve as the original behavioral fields. The specific implementation method can refer to the aforementioned process of preprocessing the multi-dimensional data based on a preset processing strategy to obtain the corresponding target data; further details will not be elaborated upon here.
[0090] Step 3: Instead of directly using the cleaned original behavioral fields for modeling, the system first extracts composite features (effective visit rate, comprehensive communication score, and customer maintenance index) from the original behavioral fields to better reflect the quality of behavior.
[0091] Furthermore, 1) a behavior-performance correlation model is established. The system uses performance indicators as the dependent variable and all behavioral characteristics and composite characteristics as independent variables, employing multiple regression analysis to establish a quantitative model. Let the dependent variable Y be monthly premium income (normalized), and the independent variables be n behavioral characteristics X1, X2, ..., Xn (all normalized). Then, the expression of the multiple regression model is: Y = β0 + β1 × X1 + β2 × X2 + ... + βn × Xn + ε. Where β0 is the intercept term, β1, β2, ..., βn are the regression coefficients (i.e., influence weights) of each behavioral characteristic, and ε is the error term. The system solves for each regression coefficient using the least squares method to minimize the sum of squared residuals between the predicted and actual values. After obtaining each β value, the larger the absolute value of the β value, the greater the impact of the behavioral characteristic on performance; a positive β value indicates that the behavior has a positive promoting effect on performance, while a negative β value indicates a negative inhibiting effect.
[0092] For example, the model might yield the following results: β1 = 0.35 for effective visit rate, β2 = 0.28 for overall communication score, β3 = 0.22 for customer maintenance index, and β4 = 0.08 for total monthly visits. This indicates that for every unit increase in effective visit rate (on a normalized scale), monthly premium income is expected to increase by 0.35 units, while simply increasing the number of visits has an impact of only 0.08, far less than the impact of visit quality.
[0093] 2) Establish a behavior-compliance correlation model. The system uses compliance indicators as the dependent variable and behavioral characteristics as the independent variables to establish another regression model to analyze which behavioral characteristics are prone to compliance risks. Let the dependent variable Z be the monthly comprehensive violation score (already normalized; a higher value indicates a greater compliance risk), and the independent variables still be behavioral characteristics X1, X2, ..., Xn. Then the model expression is: Z = α0 + α1 × X1 + α2 × X2 + ... + αn × Xn + ε. The α coefficients are obtained by solving the model. For example, the model might find that agents with excessively long call durations but low communication quality scores have significantly positive α values, indicating that "long-duration, low-quality communication" is a high-risk violation behavior; agents with frequent visits but low customer maintenance indices also have significantly positive α values, indicating that "emphasizing development over maintenance" easily leads to customer complaints.
[0094] 3) Establish a cross-correlation between performance and compliance. The system further analyzes whether there is a contradictory relationship between performance and compliance. Specifically, it calculates the correlation coefficient between performance ranking and compliance risk ranking. Let all agents be ranked from highest to lowest monthly premium income, resulting in a performance ranking sequence R_y; and ranked from highest to lowest comprehensive violation score, resulting in a compliance risk ranking sequence R_z. Calculate the Spearman rank correlation coefficient between the two ranking sequences: ρ = 1 - (6 × Σ(R_yi - R_zi)²) / (n × (n² - 1)), where n is the total number of agents, R_yi is the performance ranking of the i-th agent, and R_zi is the compliance risk ranking of the i-th agent. If ρ is significantly positive, it indicates that agents with higher performance also have higher compliance risk, meaning that some high performance may stem from non-compliant behavior, requiring managerial vigilance. If ρ is close to zero or negative, it indicates that there is no significant contradiction between performance and compliance, and high performance is healthy and sustainable.
[0095] Once the 3D model is established, each agent possesses a complete set of relational profiles: beta coefficients indicating which behavioral characteristics have the greatest impact on their performance (from sub-step 1), alpha coefficients indicating which behavioral characteristics pose compliance risks (from sub-step 2), and ρ coefficients indicating whether there is a conflict between their performance and compliance (from sub-step 3). However, these analytical results are general conclusions applicable to all agents, while different agents are at different career stages (newcomers, high performers, low performers) and cannot be interpreted using the same set of standards. Therefore, the next step, stratified matching, is needed to adjust the analytical weights for different groups.
[0096] Specifically, the process of adjusting the analysis weights for different groups includes: (1) Assigning analysis weights to new recruits. The analysis of new recruits focuses on the behavioral process dimension. The system assigns the following weights to new recruits: 70% for behavioral characteristics, 20% for performance characteristics, and 10% for compliance characteristics. The logic behind this weighting is that new recruits have not yet developed stable performance, and it is unfair and inaccurate to evaluate them based on performance results. What should be focused on is whether their behavioral processes are developing in the right direction, such as whether the frequency of visits is increasing, whether the quality of communication is improving, and whether customer maintenance is accumulating. New recruits usually have fewer problems with compliance, so the weighting is the lowest.
[0097] (2) Assigning analysis weights to high-performing groups. The analysis of high-performing groups focuses on performance maintenance and compliance. The system assigns the following weights to high-performing groups: 50% for performance characteristics, 30% for behavioral characteristics, and 20% for compliance characteristics. The logic behind this weighting is that high-performing individuals have proven their abilities, and managers are most concerned about whether their high performance can be sustained, whether there are any compliance risks (because high performance is sometimes accompanied by high-risk behavior), and which behaviors are their core competencies that need to be maintained and replicated. Therefore, performance and compliance have higher weights, while the weight of behavior is relatively lower but still important.
[0098] (3) Assigning analysis weights to inefficient groups. The analysis of inefficient groups focuses on the dimension of identifying behavioral shortcomings. The system assigns the following weights to inefficient groups: 60% for behavioral characteristics, 25% for performance characteristics, and 15% for compliance characteristics. The logic behind this weight allocation is that inefficient individuals do not need to look at performance results (the results are already very poor, and looking at them will not change anything), but rather to accurately identify the biggest gap between their behavior and that of high-performing individuals, and then make targeted improvements. Therefore, behavioral characteristics have the highest weight, while performance and compliance have relatively lower weights.
[0099] After the stratification is completed, each agent is labeled with a group tag (newcomer, high performer, low performer), and the system has already loaded the corresponding weight configuration for that group. In the next step of correlation analysis, the system will use the agent's group tag to call the corresponding weight configuration and analysis dimensions to calculate the correlation between the agent's behavior and performance, thereby drawing targeted conclusions.
[0100] In some optional implementations of this embodiment, step S2053 includes the following steps: Step S20531: Determine the target group corresponding to the user based on the group type.
[0101] In this embodiment, the target group refers to a specific group that matches the group type. For example, if the group type is newcomers, the corresponding target group is the newcomer group. If the group type is high-performing, the corresponding target group is the high-performing group. If the group type is inefficient, the corresponding target group is the inefficient group.
[0102] Step S20532: Extract the mean value of target behavioral features corresponding to the specified key behavioral features from the target group.
[0103] In this embodiment, the specified key behavioral feature is any one of all the key behavioral features, such as effective visit rate / comprehensive communication score / customer maintenance index.
[0104] Specifically, based on the user's group type, the system can extract behavioral characteristic data of all agents within the target group and calculate the optimal mean for each behavioral characteristic. Taking the inefficient group as an example, the system extracts the effective visit rate data for all inefficient agents this month and calculates its mean as the optimal effective visit rate for the inefficient group. Assuming there are m agents in the inefficient group, and the effective visit rate of the j-th agent is E_j, then the optimal effective visit rate for the inefficient group is: E_best_low = (ΣE_j) / m, where j ranges from 1 to m. Similarly, the system calculates the optimal overall communication score mean G_best_low, the optimal customer maintenance index mean C_best_low, etc., for the inefficient group.
[0105] Step S20533: Calculate the gap data between the specified key behavioral feature and the mean of the target behavioral feature.
[0106] In this embodiment, the gap data is calculated by comparing the actual values of each behavioral characteristic of the user this month with the optimal mean of the target group to which the user belongs (i.e., the mean of the target behavioral characteristics).
[0107] Let E_actual be the effective visit rate of an inefficient agent this month, and E_best_low be the optimal effective visit rate for the inefficient group. Then, the gap in effective visit rates is: Gap_E = E_best_low - E_actual. Similarly, the gap in overall communication score is Gap_G = G_best_low - G_actual, and the gap in customer maintenance index is Gap_C = C_best_low - C_actual. If the gap is positive, it means that the agent is below the group's optimal level in that behavioral dimension; if the gap is negative (which is extremely rare), it means that the agent has exceeded the group's optimal level in that dimension.
[0108] Step S20534: Obtain the specified influence weight corresponding to the specified key behavioral feature.
[0109] In this embodiment, the specified influence weight corresponding to the specified key behavioral feature can be extracted from the above influence weight.
[0110] Step S20535: Calculate the gap data and the specified influence weight based on the preset loss calculation formula to obtain the corresponding calculation result.
[0111] In this embodiment, the loss calculation formula includes: Potential Loss = Influence Weight * Gap. The gap data and the specified influence weight can be substituted into the corresponding positions in the loss calculation formula for calculation, and the resulting calculation is used as the specified loss data corresponding to the specified key behavioral feature.
[0112] Taking the effective visit rate as an example, let β_E be the weight (regression coefficient) of the impact of the effective visit rate on premiums in the target correlation analysis model, and Gap_E be the gap in effective visit rates. Then, the potential performance loss due to this behavioral shortcoming of the effective visit rate is: Loss_E = β_E × Gap_E. For example, if β_E = 0.35 and Gap_E = 0.25 (meaning the agent's effective visit rate is 25 percentage points lower than the group's optimal value), then Loss_E = 0.35 × 0.25 = 0.0875, meaning this behavioral shortcoming may cause their performance to be about 8.75% lower than the group's optimal level. Similarly, the potential loss in the overall communication score is Loss_G = β_G × Gap_G, and the potential loss in the customer maintenance index is Loss_C = β_C × Gap_C.
[0113] Step S20536: Use the calculation result as the specified loss data corresponding to the specified key behavioral feature.
[0114] This application identifies a target group corresponding to a user based on group type; then extracts the mean of target behavioral features corresponding to a specified key behavioral feature from the target group; wherein the specified key behavioral feature is any one of all key behavioral features; next, it calculates the difference data between the specified key behavioral feature and the mean of the target behavioral feature; and obtains the specified influence weight corresponding to the specified key behavioral feature; subsequently, it calculates and processes the difference data and the specified influence weight based on a preset loss calculation formula to obtain the corresponding calculation result; finally, the calculation result is used as the specified loss data corresponding to the specified key behavioral feature. Based on the above processing flow, this application extracts the mean of target behavioral features corresponding to a specified key behavioral feature from the target group corresponding to the user, calculates the difference data between the specified key behavioral feature and the mean of the target behavioral feature, and then calculates and processes the difference data and the specified influence weight based on the use of a loss calculation formula, and uses the obtained calculation result as the specified loss data corresponding to the specified key behavioral feature. This enables efficient and accurate loss calculation processing between key behavioral features and influence weights, improves the processing efficiency of loss calculation, and ensures the accuracy of the obtained loss data.
[0115] In some optional implementations of this embodiment, step S206 includes the following steps: Step S2061: Invoke the pre-built suggestion knowledge base.
[0116] In this embodiment, the aforementioned suggestion knowledge base is a pre-built mapping relationship that stores a large number of historical success cases based on actual business needs, namely, "what kind of weakness + what kind of group, what kind of suggestion and what kind of expected effect." For example, it can be a database that stores a large number of "condition-suggestion-effect" triple mappings. The structure of each triple is: when the weakness type is A and the group type is B, the recommended suggestion is C, and the expected effect is D.
[0117] Step S2062: Based on the target group type and the analysis results, perform suggestion matching processing on the suggestion knowledge base to obtain the corresponding first optimization suggestion.
[0118] In this embodiment, the system can retrieve matching suggestions from the aforementioned suggestion knowledge base based on the user's target group type and analysis results (i.e., the type of weakness) as the corresponding first optimization suggestion.
[0119] For example, if the analysis result (weakness type) is "low effective visit rate" and the agent is a new recruit, the recommended suggestions from the knowledge base are: adjust the daily visit target from eight to five, but require each visit to last at least twenty minutes, and have the supervisor accompany the agent on two visits per week to improve communication quality; if the analysis result is "low overall communication score" and the agent is an inefficient group, the recommended suggestions from the knowledge base are: participate in specialized training on standardized communication techniques and conduct daily self-checks of call recordings for the next two weeks; if the analysis result is "low customer maintenance index" and the agent is a high-performing group, the recommended suggestions from the knowledge base are: classify customers by value, follow up with high-value customers at least twice a month, and use the company's customer care toolkit to improve satisfaction.
[0120] Furthermore, the system doesn't just generate one suggestion. Instead, it ranks the shortcomings (based on the ranked target loss data) and matches one suggestion to each of the top three to five shortcomings, forming a priority list. Specifically, the first priority suggestion generates a suggestion for the top-ranked behavioral shortcoming (the shortcoming with the greatest potential loss), along with an estimated expected effect. The formula for calculating the expected effect is: Expected performance improvement = Impact weight of the behavior × Gap value of the behavior × 100%. For example, if the impact weight of the effective visit rate β_E = 0.35 and the gap_E = 0.25, then the expected performance improvement = 0.35 × 0.25 × 100% = 8.75%. The second priority suggestion generates a suggestion for the second-ranked behavioral shortcoming, with the expected effect calculated using the same formula. The third and subsequent priority suggestions follow the same logic.
[0121] Step S2063: Add execution guidelines to the first optimization suggestion to obtain the processed second optimization suggestion.
[0122] In this embodiment, for each suggestion in the aforementioned suggestion knowledge base, corresponding execution guidelines are pre-configured based on actual guidance needs. For example, for the suggestion of "low effective visit rate," the execution guidelines are as follows: In the first week, reduce the daily visit target from eight to five times; the system automatically sets a new target in the CRM system; after each visit, the visit duration and customer needs record must be filled in the CRM system; visits lasting less than twenty minutes are not counted as effective visits; the supervisor accompanies the agent on two visits within the first week and provides immediate feedback after each visit. The target is: after two weeks, the effective visit rate should increase from the current 40% to over 50%.
[0123] Furthermore, by obtaining the specified execution guide corresponding to the first optimization suggestion mentioned above and supplementing the specified execution guide to the corresponding position in the first optimization suggestion, a corresponding second optimization suggestion can be obtained.
[0124] Step S2064: The second optimization suggestion is used as the target optimization suggestion.
[0125] This application utilizes a pre-built suggestion knowledge base; then, based on the target group type and analysis results, it performs suggestion matching processing on the suggestion knowledge base to obtain a corresponding first optimization suggestion; subsequently, it adds execution guidelines to the first optimization suggestion to obtain a processed second optimization suggestion; finally, it uses the second optimization suggestion as the target optimization suggestion. Based on the above processing flow, this application, by combining the use of target group type and analysis results, matches the most suitable target optimization suggestion from the suggestion knowledge base, thereby achieving suggestion recommendation based on cases that have been validated in historical data, effectively ensuring the relevance and operability of the generated target optimization suggestions.
[0126] In some alternative implementations, the user information obtained is subject to user consent and complies with relevant laws and policies.
[0127] Furthermore, any software tools or components not belonging to our company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.
[0128] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0129] It should be emphasized that, to further ensure the privacy and security of the above-mentioned optimization suggestions, these suggestions can also be stored in a node of a blockchain.
[0130] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0131] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0132] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0133] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0134] Further reference Figure 3 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a data analysis device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0135] like Figure 3 As shown, the data analysis device 300 described in this embodiment includes: a first acquisition module 301, a preprocessing module 302, a second acquisition module 303, a calling module 304, an analysis module 305, a third acquisition module 306, and a sending module 307. Wherein: The first acquisition module 301 is used to acquire multi-dimensional user data from multiple preset business systems; The preprocessing module 302 is used to preprocess the multi-dimensional data based on a preset processing strategy to obtain the corresponding target data; The second acquisition module 303 is used to acquire the target group type corresponding to the user; Module 304 is used to invoke the target association analysis model corresponding to the target group type. Analysis module 305 is used to perform association analysis processing on the target data based on the target association analysis model to obtain the corresponding analysis results; The third acquisition module 306 is used to acquire target optimization suggestions corresponding to the analysis results from a preset suggestion knowledge base; The sending module 307 is used to send the target optimization suggestions to the user.
[0136] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the data analysis method in the aforementioned embodiments, and will not be repeated here.
[0137] In some optional implementations of this embodiment, the first acquisition module 301 includes: The first acquisition submodule is used to acquire behavioral data, performance data, and compliance data corresponding to the user from the multiple business systems. The alignment submodule is used to perform data alignment processing on the behavioral data, the performance data, and the compliance data based on the user's identification code to obtain the corresponding generated data. The first determining submodule is used to use the generated data as the multi-dimensional data.
[0138] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the data analysis method in the aforementioned embodiments, and will not be repeated here.
[0139] In some optional implementations of this embodiment, the preprocessing module 302 includes: The first processing submodule is used to process the missing values of the multi-dimensional data to obtain the corresponding first processed data. The second processing submodule is used to perform outlier detection and correction processing on the first processed data to obtain the corresponding second processed data. The third processing submodule is used to standardize the second processed data to obtain the corresponding third processed data. The fourth processing submodule is used to perform time-granularity unification processing on the third processing data to obtain the corresponding fourth processing data; The fifth processing submodule is used to perform data encoding processing on the fourth processing data to obtain the corresponding fifth processing data; The second determining submodule is used to use the fifth processed data as the target data.
[0140] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the data analysis method in the aforementioned embodiments, and will not be repeated here.
[0141] In some optional implementations of this embodiment, the second acquisition module 303 includes: The second acquisition submodule is used to acquire the user's employment duration data and historical performance ranking data; The generation submodule is used to perform hierarchical analysis and processing on the employment duration data and the historical performance ranking data based on preset hierarchical rules, and generate the target group type corresponding to the user.
[0142] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the data analysis method in the aforementioned embodiments, and will not be repeated here.
[0143] In some optional implementations of this embodiment, the analysis module 305 includes: The sixth processing submodule is used to perform feature engineering on the target data to obtain corresponding key behavioral features; wherein, the number of key behavioral features is multiple; An extraction submodule is used to extract the influence weights corresponding to the key behavioral features from the target association analysis model. The calculation submodule is used to perform loss calculation processing on the key behavioral features and the influence weights to obtain the corresponding loss data; The sorting submodule is used to sort the loss data in descending order to obtain the sorted target loss data. The third determining submodule is used to take the target loss data as the analysis result.
[0144] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the data analysis method in the aforementioned embodiments, and will not be repeated here. In some optional implementations of this embodiment, the computation submodule includes: The first determining unit is used to determine the target group corresponding to the user based on the group type; An extraction unit is used to extract the mean value of target behavioral features corresponding to a specified key behavioral feature from the target group; wherein, the specified key behavioral feature is any one of all the key behavioral features; The calculation unit is used to calculate the difference data between the specified key behavioral feature and the mean of the target behavioral feature; The acquisition unit is used to acquire the specified influence weight corresponding to the specified key behavioral feature; The second calculation unit is used to calculate the gap data and the specified influence weight based on a preset loss calculation formula to obtain the corresponding calculation results. The second determining unit is used to use the calculation result as the specified loss data corresponding to the specified key behavioral feature.
[0145] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the data analysis method in the aforementioned embodiments, and will not be repeated here.
[0146] In some optional implementations of this embodiment, the third acquisition module 306 includes: The second submodule is used to invoke a pre-built suggestion knowledge base; The matching submodule is used to perform suggestion matching processing on the suggestion knowledge base based on the target group type and the analysis results to obtain the corresponding first optimization suggestion; Add a submodule to add execution guidelines to the first optimization suggestion to obtain the processed second optimization suggestion; The fourth determining submodule is used to take the second optimization suggestion as the target optimization suggestion.
[0147] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the data analysis method in the aforementioned embodiments, and will not be repeated here. To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0148] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0149] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0150] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may also include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for data analysis methods. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.
[0151] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, such as executing computer-readable instructions for the data analysis method.
[0152] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.
[0153] Compared with the prior art, the embodiments of this application have the following beneficial effects: In this embodiment, the application uses a target association analysis model to perform association analysis on multi-dimensional target data to obtain analysis results, which can effectively improve the accuracy and intelligence of data analysis. Furthermore, based on the use of a suggestion knowledge base, it generates target optimization suggestions that match the analysis results, thereby providing users with accurate and personalized suggestions and improving the accuracy of the generated target optimization suggestions.
[0154] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the data analysis method described above.
[0155] Compared with the prior art, the embodiments of this application have the following main advantages: In this embodiment, the application uses a target association analysis model to perform association analysis on multi-dimensional target data to obtain analysis results, which can effectively improve the accuracy and intelligence of data analysis. Furthermore, based on the use of a suggestion knowledge base, it generates target optimization suggestions that match the analysis results, thereby providing users with accurate and personalized suggestions and improving the accuracy of the generated target optimization suggestions.
[0156] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0157] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.
[0158] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.
Claims
1. A data analysis method, characterized in that, Includes the following steps: Obtain multi-dimensional user data from multiple pre-set business systems; Based on a preset processing strategy, the multi-dimensional data is preprocessed to obtain the corresponding target data; Obtain the target group type corresponding to the user; Invoke the target association analysis model corresponding to the target group type; Based on the target association analysis model, the target data is subjected to association analysis processing to obtain the corresponding analysis results; Obtain target optimization suggestions corresponding to the analysis results from a pre-set suggestion knowledge base; The proposed optimization suggestions are sent to the user.
2. The data analysis method according to claim 1, characterized in that, The steps of obtaining multi-dimensional user data from multiple preset business systems specifically include: Obtain behavioral data, performance data, and compliance data corresponding to the user from the multiple business systems; Based on the user's identification code, the behavioral data, the performance data, and the compliance data are aligned to obtain the corresponding generated data. The generated data is used as the multi-dimensional data.
3. The data analysis method according to claim 1, characterized in that, The step of preprocessing the multi-dimensional data based on a preset processing strategy to obtain the corresponding target data specifically includes: Missing values are processed on the multi-dimensional data to obtain the corresponding first processed data; The first processed data is subjected to outlier detection and correction to obtain the corresponding second processed data; The second processed data is standardized to obtain the corresponding third processed data; The third processed data is subjected to time-granularity unification processing to obtain the corresponding fourth processed data; The fourth processed data is encoded to obtain the corresponding fifth processed data; The fifth processed data is used as the target data.
4. The data analysis method according to claim 1, characterized in that, The step of obtaining the target group type corresponding to the user specifically includes: Obtain the user's employment duration data and historical performance ranking data; Based on preset stratification rules, the employment duration data and the historical performance ranking data are analyzed and processed in a stratified manner to generate target group types corresponding to the users.
5. The data analysis method according to claim 1, characterized in that, The step of performing association analysis on the target data based on the target association analysis model to obtain the corresponding analysis results specifically includes: Feature engineering is performed on the target data to obtain corresponding key behavioral features; wherein, the number of key behavioral features is multiple. Extract the influence weights corresponding to the key behavioral features from the target association analysis model; The key behavioral features and the influence weights are subjected to loss calculation processing to obtain the corresponding loss data; The loss data is sorted from largest to smallest to obtain the sorted target loss data. The target loss data is used as the analysis result.
6. The data analysis method according to claim 5, characterized in that, The step of performing loss calculation processing on the key behavioral features and the influence weights to obtain the corresponding loss data specifically includes: Based on the group type, the target group corresponding to the user is determined; Extract the mean of target behavioral features corresponding to a specified key behavioral feature from the target group; wherein, the specified key behavioral feature is any one of all the key behavioral features; Calculate the difference data between the specified key behavioral feature and the mean of the target behavioral feature; Obtain the specified influence weight corresponding to the specified key behavioral feature; The gap data and the specified influence weight are calculated based on the preset loss calculation formula to obtain the corresponding calculation results. The calculation results are used as the specified loss data corresponding to the specified key behavioral features.
7. The data analysis method according to claim 1, characterized in that, The step of obtaining target optimization suggestions corresponding to the analysis results from a preset suggestion knowledge base specifically includes: Call upon a pre-built suggestion knowledge base; Based on the target group type and the analysis results, the suggestion knowledge base is subjected to suggestion matching processing to obtain the corresponding first optimization suggestion; The first optimization suggestion is processed by adding execution guidelines to obtain the processed second optimization suggestion; The second optimization suggestion is taken as the target optimization suggestion.
8. A data analysis device, characterized in that, include: The first acquisition module is used to acquire multi-dimensional user data from multiple preset business systems; The preprocessing module is used to preprocess the multi-dimensional data based on a preset processing strategy to obtain the corresponding target data. The second acquisition module is used to acquire the target group type corresponding to the user; The calling module is used to call the target association analysis model corresponding to the target group type; The analysis module is used to perform association analysis on the target data based on the target association analysis model to obtain the corresponding analysis results; The third acquisition module is used to acquire target optimization suggestions corresponding to the analysis results from a preset suggestion knowledge base; The sending module is used to send the target optimization suggestions to the user.
9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the data analysis method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the data analysis method as described in any one of claims 1 to 7.