Data analysis decision method and device, equipment and medium
By building and updating target portraits through distributed data collection and cleaning, text mining, and association rule mining, we solve the problem of low efficiency of traditional analysis methods, achieve efficient and accurate data-driven decision support, and optimize corporate strategy and customer service.
Patent Information
- Application Number
- CN202510716289.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-16
AI Technical Summary
Traditional manual analysis methods are inefficient and have limited accuracy. They are unable to build comprehensive and accurate portraits of products, customers, and agents from complex customer service data, making it difficult for decision makers to obtain in-depth and effective information to guide corporate strategy formulation, product optimization, and customer service improvement.
Data cleaning is performed through distributed data collection, preset segmentation rules and pre-trained detection models. Combined with text mining and association rule mining, target portraits are constructed and updated in real time, and decision reports are generated using preset decision algorithms.
It significantly improves the efficiency and accuracy of data processing and decision support, helping companies extract useful business information from massive customer service data and optimize corporate strategic decision-making.
Smart Images

Figure CN120655121A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing and analysis technology in financial and medical application scenarios, and in particular to a data analysis decision-making method, device, equipment and medium. Background Art
[0002] In today's financial and healthcare sectors, major companies have accumulated vast amounts of customer service data. This data contains valuable information on the usage of financial or medical products, user preferences, and agent performance. However, traditional manual analysis methods are inefficient and have limited accuracy, making it difficult to fully tap the potential value of this data. Furthermore, existing simple statistical methods are unable to construct comprehensive and accurate profiles of products, customers, and agents from complex customer service data, making it difficult for decision makers to obtain in-depth and effective information to guide corporate strategy formulation, product optimization, and customer service enhancement. Therefore, developing an efficient and intelligent customer service data analysis solution is of great practical significance. Summary of the Invention
[0003] The embodiments of the present invention provide a data analysis and decision-making method, apparatus, device and medium, which aim to solve the problem that data analysis and decision-making under the existing technology are inefficient and cannot meet the practical requirements of the application field.
[0004] In a first aspect, an embodiment of the present invention provides a data analysis and decision-making method, which includes: performing distributed data collection from a target data source according to preset collection rules to collect raw data; performing data segmentation on the raw data according to preset segmentation rules, and then cleaning the data through a pre-trained detection model to obtain data to be analyzed; performing text mining analysis and association rule mining on the data to be analyzed, and writing the results into an analysis result library; constructing a target portrait based on the data in the analysis result library, and updating the target portrait as the raw data is updated in real time; performing decision analysis based on the target portrait through a preset decision algorithm to generate a decision report.
[0005] In a second aspect, an embodiment of the present invention further provides a data analysis and decision-making device for executing the data analysis and decision-making method described above.
[0006] In a third aspect, an embodiment of the present invention further provides a computer device, comprising a memory and a processor connected to the memory; the memory is used to store computer programs; and the processor is used to run the computer programs stored in the memory to execute the steps of the above-mentioned data analysis and decision-making method.
[0007] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the steps of the above-mentioned data analysis and decision-making method can be implemented.
[0008] Compared with the prior art, the present invention has the following beneficial effects:
[0009] In the technical solution of the present invention, the data analysis and decision-making method first performs distributed data collection from the target data source by combining passive collection with active collection to obtain raw data. Then, the raw data is efficiently cleaned and processed through preset segmentation rules and pre-trained detection models. Then, through text mining analysis and association rule mining technology, the key patterns and potential associations in a large amount of data are extracted, and then an accurate target portrait is constructed, and dynamically adjusted according to the real-time updated raw data. Finally, through the preset decision algorithm, an in-depth analysis is performed based on the target portrait and a decision report is generated. This method significantly improves the efficiency and accuracy of data processing and decision support. It helps companies extract useful business information from massive customer service data, thereby optimizing the company's strategic decisions and meeting the high demand of companies in the financial and medical fields for data-driven decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0011] Figure 1 A flow chart of the data analysis and decision-making method provided by the present invention;
[0012] Figure 2 This is a first sub-flowchart of the data analysis and decision-making method provided by the present invention;
[0013] Figure 3 A sub-flowchart of the second sub-flowchart of the data analysis and decision-making method provided by the present invention;
[0014] Figure 4 This is a third sub-flowchart of the data analysis and decision-making method provided by the present invention;
[0015] Figure 5 This is a fourth sub-flowchart of the data analysis and decision-making method provided by the present invention;
[0016] Figure 6 This is a fifth sub-flowchart of the data analysis and decision-making method provided by the present invention;
[0017] Figure 7 This is a sixth sub-flowchart of the data analysis and decision-making method provided by the present invention;
[0018] Figure 8 A schematic block diagram of the units of the data analysis and decision-making device provided by the present invention;
[0019] Figure 9 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0021] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0022] It should also be understood that the terminology used in this specification is for the purpose of describing medical embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0023] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0024] The present invention aims to solve the problem that the data analysis and decision-making efficiency under the existing technology is low and cannot meet the practical requirements of the application field, and provides a data analysis and decision-making method, device, equipment and medium. Figures 1 to 7 ,wherein the data analysis and decision-making method includes the following steps.
[0025] S110, performing distributed data collection from the target data source according to preset collection rules to collect raw data;
[0026] S120, segmenting the raw data using a preset segmentation rule, and then cleaning the data using a pre-trained detection model to obtain data to be analyzed;
[0027] S130, performing text mining analysis and association rule mining on the data to be analyzed, and writing the results into an analysis result database;
[0028] S140: Construct a target profile based on the data in the analysis result library, and update the target profile as the original data is updated in real time;
[0029] S150: Perform decision analysis based on the target portrait using a preset decision algorithm to generate a decision report.
[0030] During the data collection phase, this method utilizes a combination of passive and active data collection methods to perform distributed data collection from target data sources. Through these methods, the system can collect a wealth of raw data from multiple channels. This collected raw data requires cleaning to ensure the accuracy of subsequent analysis. During the data cleaning phase, the present invention segments the data using preset segmentation rules and cleans the data using a pre-trained detection model. This process removes irrelevant information, formats the data, and eliminates noise, providing clean, high-quality data for the subsequent analysis phase. This significantly improves data cleaning efficiency and ensures the accuracy of the data to be analyzed. After data cleaning is complete, the data to be analyzed enters the data analysis phase. In this phase, the system utilizes text mining techniques and association rule mining algorithms to conduct in-depth data analysis. Text mining analyzes customer service conversation text to extract key information, such as customer needs and product feedback, thereby helping to identify potential issues or opportunities. Furthermore, the system utilizes an improved Apriori algorithm and association rule mining techniques to discover potential relationships between data. This process can provide valuable business insights and data support for decision-making. Based on the results of data analysis, a target profile is further constructed. By analyzing data in the result database, the system constructs and updates customer, product, and agent profiles in real time, ensuring the accuracy and timeliness of target profiles. This profile construction utilizes big data cluster analysis and distributed storage technologies to analyze customer behavior, product feedback, and other information in real time, presenting this information visually to decision makers. The target profile encompasses not only individual customer characteristics but also their behavioral traits and preferences, helping companies better understand customer needs and optimize their products and services. Finally, based on the target profile and the results of the aforementioned data analysis, the system uses pre-defined decision-making algorithms to conduct decision analysis and generate decision reports. These reports provide clear support for decision makers, helping them make more informed and accurate decisions. To enhance decision effectiveness, the system also uses intelligent recommendation algorithms to proactively recommend potential decision options or perspectives based on their historical decision-making behavior and current data. This mechanism significantly shortens the decision cycle and improves the accuracy and reliability of decisions.
[0031] In one embodiment, referring to Figure 2 , the steps of S110 include:
[0032] S111. Periodically and automatically passively collecting and acquiring the raw data from the recorded target data at preset time intervals, and / or passively collecting and acquiring the raw data when a preset business event is detected based on an event triggering mechanism;
[0033] S112, actively collecting and acquiring the raw data by initiating active inquiries to the target object under a preset specific business scenario;
[0034] S113, centralizing and storing the raw data collected from different data sources through distributed collection;
[0035] S114: Perform desensitization processing on the original data, where the desensitization processing includes at least one of masking desensitization, encryption desensitization, and generalization desensitization.
[0036] During the data collection phase, passive data collection is achieved through pre-set scheduled tasks and event-triggered mechanisms. Examples include periodic time intervals, such as automatically capturing a full snapshot of the customer service ticket database at 2:00 AM each day to capture data. Event-triggered mechanisms, such as capturing the associated call recordings in real time when a customer submits a claim, are also implemented. The event-triggered mechanism incorporates a built-in business rules engine that instantly initiates the data collection process upon detecting a pre-set business event. Furthermore, for high-concurrency scenarios, the system automatically enables a backup collection channel, buffering the data stream via a Kafka message queue to ensure that transaction logs are not lost during the collection process.
[0037] Active collection is the process of proactively initiating questionnaires or data collection requests to customers or agents in specific business scenarios. For example, in the first 30 days after a new product launch, a user experience questionnaire will be automatically pushed to all customers who purchased the product. Or, when applying for a claim for a highly controversial insurance product, a real-time pop-up window will be displayed asking the customer to upload a list of supporting documents, and then the authenticity of the documents will be automatically verified through OCR technology. Among them, the so-called specific business scenarios can specifically be the new product launch cycle, high-risk product claim requests, customer activity reaching the standard, or agent risk indicators exceeding the threshold. Specifically, after review, the actively collected data will be written into the analysis database in JSON format and an associated index will be established with the passively collected baseline data.
[0038] To efficiently process large amounts of data, this method also employs big data distributed collection technology. This involves building a distributed data collection system with the help of tools such as Apache Flume, which centrally stores raw data collected from various data sources. As a distributed data collection tool, Apache Flume can efficiently collect data from multiple data sources and, through the rational configuration of multiple proxy nodes, ensure efficient data transmission and system scalability. Specifically, a three-tier distributed collection network is constructed using Apache Flume: edge-layer Flume Agents are deployed in local customer service centers to handle preliminary local data processing; aggregation-layer Agent clusters are set up in regional centers to receive edge-layer data through load balancing strategies; and the central storage layer utilizes the HDFS distributed file system. During transmission, SSL encryption is enabled across inter-center links, and local transmission uses the Avro RPC protocol. Collection logs are written to the Elasticsearch cluster in real time, providing collection progress monitoring and exception alerting capabilities.
[0039] During the data collection process, especially when sensitive information such as customer ID numbers and bank card numbers is involved, the system will perform desensitization after data collection to ensure data security and privacy protection. In practical applications in the financial and medical fields, the following desensitization methods are specifically used:
[0040] Masking and desensitization. For example, for a bank card number, the system retains the first 6 and last 4 digits, replacing the rest with asterisks (e.g., 622888*****1234). This desensitization method is suitable for real-time data display to ensure that sensitive information is not leaked.
[0041] Encryption and desensitization: Sensitive data is encrypted using the national secret SM4 encryption algorithm, making it impossible to directly obtain the true information even if the data is leaked. This method is used for internal risk control system transmission, ensuring the security of data transmission and storage.
[0042] Generalized desensitization, for example, generalizes the customer's age and converts the specific age value into an age range, such as converting 30 years old into the 20-35 age range, to ensure the anonymity of the data while not affecting the accuracy of data analysis. This method is used in external data analysis reports.
[0043] Hash desensitization, such as using the SHA-256 algorithm to hash a mobile phone number with salt, renders the data irreversible, preventing the user's real information from being exposed when linking data across systems. This method is suitable for cross-system data association analysis.
[0044] Through these desensitizing methods, the system can ensure data privacy and security while still providing usable data for subsequent analysis and decision-making.
[0045] In one embodiment, referring to Figure 3 , the steps of S120 include
[0046] S121, dividing the original data into data blocks divided into natural time units according to the data timestamp;
[0047] S122, dividing the data block into two parts according to the service type, so that the same data block only contains data of a single service type;
[0048] S123. Allocate storage nodes to the data blocks after the secondary division based on customer attributes, where the customer attributes include geographical distribution or customer level.
[0049] In order to ensure that the data can be processed and stored effectively and accurately, thereby improving the efficiency and accuracy of subsequent analysis, it is necessary to set reasonable data segmentation rules. In this embodiment, the system performs structured sharding processing on the original data based on multidimensional features. First, primary division is performed based on the timestamp when the data is generated. For time-sensitive data such as customer service hotline recordings and online chat records, they are cut into independent data blocks according to natural time units based on the different timestamps of the data, that is, the original data is divided into data blocks divided by natural time units. Usually, the time unit can be a day, a week, a month, etc. For example, customer service recording data can be divided by week, and all customer service recordings in a quarter of the historical data can be generated into multiple data blocks per week. This division method by natural time units can ensure that in the subsequent analysis process, analysts can view the data changes in different time periods according to the timeline and make periodic comparisons.
[0050] Next, a secondary segmentation based on business type is performed. To ensure data consistency and relevance within each data block, the data blocks are further segmented based on business type. For example, data can be segmented based on different insurance codes, such as auto insurance and health insurance, or based on different service types, such as reporting calls and consultation calls. This ensures that each data block contains only data from a single business type, preventing data from different business types from being mixed together and affecting subsequent analysis.
[0051] After completing the secondary division of the business, the system performs storage node allocation based on customer attributes. Storage nodes are allocated to the divided data blocks according to different customer attributes to ensure that each storage node is only responsible for data in a specific region or a specific customer type. For example, based on regional distribution characteristics, a regional label storage strategy is configured in the HDFS cluster to independently store customer data in large regions such as North China and East China. The so-called HDFS cluster is a distributed file system used to store and manage large-scale data sets. For example, based on customer level characteristics, a dedicated storage pool for VIP customers is set up for independent storage. This method of segmentation based on customer attributes makes the business data of different customers more targeted and convenient when storing and processing.
[0052] Specifically, in actual operations, the MapReduce framework is used to parallelize data cleaning tasks. The MapReduce framework consists of two core steps: the Map phase and the Reduce phase. For massive amounts of customer service data, the Map phase processes each data block in parallel, splitting the data into multiple blocks. Data cleaning operations are then performed to remove duplicate records and correct format errors. The Reduce phase merges and aggregates the cleaned data, significantly reducing data cleaning time and improving cleaning efficiency.
[0053] In this embodiment, through intelligent segmentation rule design, not only the efficiency of data storage and analysis is improved, but also the security and accuracy of data processing are guaranteed, which helps enterprises better utilize customer service data for analysis and decision-making, and improve business operation efficiency and customer service level.
[0054] In one embodiment, referring to Figure 4 The step of S120 further includes:
[0055] S124, performing anomaly detection on the numerical data in the original data using an isolation forest algorithm, and marking abnormal data that exceeds a preset confidence interval;
[0056] S125, using regular expression matching to remove garbled characters from the text data in the original data, performing semantic legitimacy verification based on a pre-trained text classification model, and marking abnormal data that fails the verification;
[0057] S126. Manually review or automatically correct the abnormal data;
[0058] S127: Create a data cleaning log to record the manual review or automatic correction processing operation performed on the abnormal data.
[0059] For numerical data such as customer service ticket processing time and customer spending amount, the system loads a pre-trained isolation forest model for anomaly detection. The isolation forest is an unsupervised learning algorithm. In specific applications, the system first trains the isolation forest model based on historical normal data to construct the distribution pattern of normal data. When cleaning newly imported data, the model automatically marks data that exceeds the threshold interval determined by the model as abnormal data. For example, in the analysis of auto insurance claims data, when it is detected that the processing time of a case exceeds N times the preset standard deviation of the industry average, the system triggers an alarm and generates an abnormal work order queue.
[0060] For text fields such as customer messages, customer service records, and case reports, the cleaning process includes two steps: removing garbled characters and verifying semantic legitimacy. Regular expression filtering removes garbled special characters such as "[#&*?]" and invalid encodings, followed by word segmentation and part-of-speech tagging. The system integrates pre-trained text classification models such as BERT or its lightweight version, ERNIE-Tiny. Based on this cleaned contextual understanding, the system determines whether the input text possesses business semantic integrity and logic, preventing invalid or haphazardly entered data from entering the analysis process.
[0061] Abnormal data is reviewed and corrected through a collaborative human-machine mechanism. The system performs rule-based remediation for abnormalities that can be automatically corrected. For abnormalities that cannot be automatically handled, pending review tasks are generated and pushed to the manual review platform. The manual review platform's review interface displays the full context of abnormal data and provides annotation tools to assist in decision-making.
[0062] To ensure traceability and auditability of data processing, the system introduces a cleaning log mechanism that records every cleaning operation in detail. This data cleaning log includes the anomaly detection type, handling method, values before and after correction, operator information, timestamp, and whether the data was finally confirmed for subsequent analysis. This log file is stored in JSON or Parquet format and archived to a data lake such as HDFS / S3 to support compliance checks and system backtracking.
[0063] The entire data cleaning process uses parallel processing based on the MapReduce framework, supporting distributed cleaning of massive customer service data. MapReduce is a programming model and processing framework for distributed computing on large datasets. In its implementation, each field type serves as input for a Map task. After being processed by different cleaning functions, the cleansing results are merged in the Reduce stage, significantly improving processing efficiency. By integrating isolation forests, regular expressions, semantic models, and MapReduce technologies, the method of this embodiment forms a complete process for cleaning structured and unstructured data.
[0064] In one embodiment, referring to Figure 5 , S130 includes:
[0065] S131. Analyze the data to be analyzed using a natural language processing model based on deep learning, and mine and extract sentiment, customer intent, and key entities from the data to be analyzed;
[0066] S132. Generate a question type distribution map and customer intention classification results based on the sentiment tendency, customer intention, and key entities;
[0067] S133. Performing association rule mining on the question type distribution map and the customer intention classification results using a preset distributed Apriori algorithm, and then compressing the transaction item set using a bitmap and generating a candidate item set across layers;
[0068] S134: Writing the candidate item set into the analysis result library.
[0069] A deep learning-based natural language processing model is used to analyze the data to be analyzed. Pre-trained language models such as BERT are used to perform sentiment analysis on customer service conversation text. By fine-tuning the BERT model, it can identify sentiment in text, such as positive, negative, or neutral. For example, when customer service records contain terms such as "very satisfied" or "poor service," the model can accurately identify the customer's sentiment. A deep learning-based intent recognition model extracts customer intent from customer service conversations, including inquiries, complaints, and suggestions. The model analyzes contextual information to identify the customer's actual needs. For example, if a customer enters "I want to learn more about a new insurance product," the model accurately identifies the customer's intent as an inquiry. Named entity recognition technology is used to extract key entities from text, such as product name, customer name, and question type. For example, in a conversation containing "My car insurance policy number is 123456," the model can extract the key entity "car insurance policy number."
[0070] After completing text mining, a question type distribution map and customer intent classification results are generated based on the extracted sentiment, customer intent, and key entities. A chi-square test is used to calculate the strength of association between question type and customer attributes, and a force-directed graph algorithm is used to construct a visual question type distribution map. Graph nodes represent specific question categories, and edge weights reflect the co-occurrence frequency between questions. Identified customer intents are classified to generate intent classification results. For example, the accuracy of customer intent classification results can be assessed using a confusion matrix. Classification statistics are then generated by business lines such as auto insurance and health insurance, showing the distribution and trend of each intent category.
[0071] In order to discover potential associations between data, this embodiment introduces an improved distributed Apriori algorithm to mine association rules. The improved distributed Apriori algorithm is as follows:
[0072] First, transaction itemset compression is performed. Each problem type combination is encoded as a bitmap, and bitwise operations replace traditional item set concatenation operations to reduce memory usage. Then, candidate itemsets are generated. Compared to traditional methods that only allow connections between adjacent levels, this improved distributed Apriori algorithm adopts a cross-level concatenation strategy, allowing k-itemsets and 1-itemsets to be directly combined to generate (k+1)-itemsets. Generating candidate itemsets through cross-level concatenation avoids the inefficiency of layer-by-layer concatenation in traditional Apriori algorithms.
[0073] The generated candidate item sets are written to the analysis results database for subsequent analysis and decision support. The data storage in the analysis results database uses a columnar partitioning strategy, with physical sharding based on business types such as auto insurance and health insurance, and date ranges. A Bloom filter is used to establish a fast search index.
[0074] The solution of this embodiment not only improves the efficiency and accuracy of data processing, but also provides enterprises with more comprehensive and in-depth business insights.
[0075] In one embodiment, referring to Figure 6 The steps of S140 include:
[0076] S141. Extract multidimensional attributes from the data in the existing analysis result library, and perform cluster analysis on the multidimensional attributes based on a dynamic weight allocation rule to generate the target profile;
[0077] S142: Setting sensitivity levels for changes in the original data, and configuring differentiated update conditions for different sensitivity levels;
[0078] S143: When it is detected that the raw data collected in real time meets the update condition, triggering an incremental update of the target portrait;
[0079] S144. Synchronize the updated target portrait to all business terminals in real time through distributed storage.
[0080] Perform multidimensional attribute extraction on the data in the existing analysis result library. Multidimensional attributes include, but are not limited to, basic information about the customer's age or gender, such as purchase records or browsing history behavior data, such as customer service conversation record interaction data. Based on these multidimensional attributes, dynamic weight assignment rules are used to assign different weights to each attribute. The weight assignment can be adjusted according to business needs and data importance. For example, purchase records may have a higher weight than browsing history. Then, cluster analysis is performed based on these multidimensional attributes to generate target profiles. Specifically, clustering algorithms such as K-means and DBSCAN are used to cluster customers, products, or agents with similar attributes to form multiple profile categories. Each profile category represents a class of objects with similar characteristics. For example, through cluster analysis, customers can be divided into different profile categories such as high-value customers, potential churn customers, and loyal customers.
[0081] Set sensitivity levels for changes in raw data, categorizing them into high, medium, and low based on the degree of change and business impact. Configure differentiated update conditions for each sensitivity level. For example, with high sensitivity, significant behavioral changes such as large purchases or complaints trigger immediate profile updates. With medium sensitivity, changes in customer behavior, such as frequent browsing of a certain product category, trigger regular profile updates. With low sensitivity, minimal changes in customer behavior, such as occasional account logins, trigger regular batch profile updates.
[0082] When the system detects that the raw data collected in real time meets the preset update conditions, it triggers an incremental update of the target profile. First, it filters out data that meets the update conditions and updates the relevant attribute values in the customer profile based on the new data. Then, based on business needs and data changes, it dynamically adjusts the weights of each attribute to ensure the accuracy and timeliness of the profile.
[0083] To ensure that the updated target profile can be synchronized to all business terminals in real time, the system uses distributed storage technologies such as HDFS. The updated profile data is distributed and stored on multiple nodes. When new data is needed to update the customer profile, the system quickly locates and updates the corresponding data blocks through the distributed computing framework. For example, when a large number of customers generate new consumer behavior data at the same time, HDFS can efficiently handle data writing and updating operations, ensuring that the real-time update of the customer profile is not affected by the growth of data volume. Through the distributed storage system, the updated profile data is synchronized to all business terminals in real time, ensuring that the profile data obtained by each terminal is always the latest.
[0084] Through the method of this embodiment, it is possible to construct a target portrait based on the data in the analysis result library and dynamically update the portrait with real-time original data. This not only improves the efficiency of portrait generation and updating, but also ensures the accuracy and timeliness of the portrait.
[0085] Further, refer to Figure 7 , the steps of S150 include:
[0086] S151. Analyze the decision maker's historical operation records through a multimodal deep network to generate multiple candidate decision plans;
[0087] S152. Predicting the response after the candidate decision plan is implemented using a time series model to generate a prediction result;
[0088] S153: Generate an interactive decision report based on the candidate decision solutions and the prediction results, and push it to the decision terminal;
[0089] S154: synchronously push the candidate decision adopted by the decision terminal to the enterprise resource management system to perform a closed-loop operation.
[0090] The decision maker's historical operation records are analyzed through a multimodal deep learning network (MDRecNet) to generate multiple candidate decision solutions. The decision maker's historical operation records in the system, including decision-making behavior, operation frequency, time points, etc., are collected and data is cleaned and normalized. A bidirectional long short-term memory network (Bi-LSTM) is used to capture the decision maker's historical behavior sequence pattern and embed the decision actions. The formula is as follows:
[0091]
[0092] in, is the embedding representation of the decision action, h t In hidden state.
[0093] The real-time business data is integrated through the multi-head attention mechanism. The formula is:
[0094]
[0095] Where Q is the query vector (Query); K is the key vector (Key); d is the dimension of the vector; exp is the exponential function; α ij Represents the attention weight between the i-th query vector and the j-th key vector. Based on the model prediction results, the response prediction after the implementation of the candidate decision plan is generated, including the change trend and possible impact of key indicators.
[0096] Comprehensively consider historical behavior and current business data to generate multiple candidate decision plans.
[0097] Next, we use time series models to predict the response after the candidate decision options are implemented and generate forecast results. We first perform time series modeling on historical data and select key indicators related to the decision options, such as sales and customer satisfaction. We then use time series forecasting models such as the Autoregressive Integrated Moving Average (ARIMA) model to predict each candidate decision option. The formula is as follows:
[0098] Y t =c+φ1Y t-1 +φ2Y t-2 +…+φ p Y t-p +θ1∈ t-1 +θ2∈ t-2 +…+θ q ∈ t-q +∈ t
[0099] Among them, Y t is the predicted index value; φ and θ are model parameters; ∈ t is the random error term. Based on the model prediction results, the response prediction after the implementation of the candidate decision plan is generated, including the change trend and possible impact of key indicators.
[0100] An interactive decision report is generated based on candidate decision options and forecast results and pushed to the decision-making terminal. First, the report's content structure is designed, including candidate decision options, forecast results, and key indicator analysis. Then, using big data visualization technologies such as Echarts and D3.js, a highly interactive decision-support visualization interface is constructed. Dynamic interactive charts display the expected effects of different candidate options, supporting decision-makers in data filtering, sorting, and drill-down analysis. The resulting interactive decision report is pushed to the decision-making terminal, ensuring that decision-makers can access and view the report content in real time.
[0101] Finally, the candidate decisions adopted by the decision-making terminal are synchronously pushed to the enterprise resource planning system (ERP system) for closed-loop execution. The specific candidate decision solutions adopted by the decision-making terminal are recorded, and corresponding decision execution instructions are generated. These execution instructions are synchronously pushed to the ERP system via an API or data bus, ensuring timely delivery to the relevant business modules for specific execution. After the ERP system executes the decision solution, the execution results are monitored in real time and fed back to the decision-making system, improving the closed-loop decision-making process and forming a continuous optimization mechanism.
[0102] The method and technical solution of this embodiment can not only provide multiple candidate decision plans and their predicted results, but also help decision makers more intuitively understand and evaluate the potential effects of each candidate plan through interactive decision reports, and realize closed-loop decision operations.
[0103] Figure 8 FIG is a schematic block diagram of a data analysis and decision-making device 600 provided by an embodiment of the present invention. Figure 8 As shown, corresponding to the above data analysis and decision-making method, the present invention also provides a data analysis and decision-making device 600. The data analysis and decision-making device 600 includes a unit for executing the above data analysis and decision-making method, and the device can be configured in a desktop computer, tablet computer, smart phone, etc.
[0104] Specifically, see Figure 8 , the data analysis and decision-making device 600 includes:
[0105] The raw data unit 610 is used to perform distributed data collection from the target data source according to preset collection rules to collect raw data;
[0106] The data cleaning unit 620 is used to segment the raw data using a preset segmentation rule, and then clean the data using a pre-trained detection model to obtain data to be analyzed;
[0107] The data analysis unit 630 is used to perform text mining analysis and association rule mining on the data to be analyzed, and write the results into the analysis result library;
[0108] A portrait construction unit 640 is configured to construct a target portrait based on the data in the analysis result library, and to update the target portrait as the original data is updated in real time;
[0109] The decision support unit 650 is used to perform decision analysis based on the target portrait through a preset decision algorithm and generate a decision report.
[0110] In one embodiment, the original data unit 610 includes:
[0111] A passive collection unit, configured to periodically and automatically passively collect and obtain the raw data from the recorded target data at preset time intervals, and / or passively collect and obtain the raw data when a preset business event is detected based on an event trigger mechanism;
[0112] An active collection unit, configured to actively collect and acquire the raw data by initiating an active inquiry to a target object under a preset specific business scenario;
[0113] A distributed collection unit, configured to centralize and store the raw data collected from different data sources through distributed collection;
[0114] The desensitization processing unit is used to perform desensitization processing on the original data, and the desensitization processing includes at least one of masking desensitization, encryption desensitization, and generalization desensitization.
[0115] In one embodiment, the data cleaning unit 620 includes:
[0116] A preliminary segmentation unit, configured to segment the original data into data blocks segmented by natural time units according to data timestamps;
[0117] A secondary segmentation unit, configured to perform secondary segmentation on the data block according to the service type, so that the same data block only contains data of a single service type;
[0118] The storage allocation unit is used to allocate storage nodes to the data blocks after the secondary division based on customer attributes, wherein the customer attributes include regional distribution or customer level.
[0119] In one embodiment, the data cleaning unit 620 further includes:
[0120] a first anomaly detection unit, configured to perform anomaly detection on numerical data in the original data using an isolation forest algorithm, and mark abnormal data that exceeds a preset confidence interval;
[0121] A second anomaly detection unit is configured to use regular expression matching to remove garbled characters from the text data in the original data, perform semantic legitimacy verification based on a pre-trained text classification model, and mark abnormal data that fails the verification;
[0122] A review and correction unit, used for manually reviewing or automatically correcting the abnormal data;
[0123] The data cleaning log unit is used to establish a data cleaning log to record the manual review or automatic correction processing operation performed on the abnormal data.
[0124] In one embodiment, the data analysis unit 630 includes:
[0125] A text mining unit is used to parse the data to be analyzed using a natural language processing model based on deep learning, and to mine and extract sentiment tendencies, customer intentions, and key entities from the data to be analyzed;
[0126] An icon generation unit, configured to generate a question type distribution map and a customer intention classification result based on the sentiment tendency, customer intention, and key entities;
[0127] A candidate item set generation unit is used to perform association rule mining on the problem type distribution map and customer intention classification results using a preset distributed Apriori algorithm, and then compress the transaction item set through bitmap compression and generate candidate item sets across layers;
[0128] The candidate item set storage unit is used to write the candidate item set into the analysis result library.
[0129] Furthermore, the portrait construction unit 640 includes:
[0130] A target portrait generation unit is used to extract multidimensional attributes from the data in the existing analysis result library, and perform cluster analysis on the multidimensional attributes based on a dynamic weight allocation rule to generate the target portrait;
[0131] A sensitivity level unit, configured to set a sensitivity level for changes in the original data and configure differentiated update conditions for different sensitivity levels;
[0132] A target portrait updating unit, configured to trigger an incremental update of the target portrait when detecting that the raw data collected in real time meets an update condition;
[0133] The target portrait synchronization unit is used to synchronize the updated target portrait to all business terminals in real time through distributed storage.
[0134] In one embodiment, the decision support unit 650 includes:
[0135] A candidate decision solution generation unit is used to analyze the decision maker's historical operation records through a multimodal deep network to generate multiple candidate decision solutions;
[0136] A prediction result generating unit, configured to predict the response after the candidate decision plan is implemented by using a time series model, and generate a prediction result;
[0137] A report pushing unit, configured to generate an interactive decision report based on the candidate decision solutions and the prediction results, and push the report to the decision terminal;
[0138] The closed-loop operation unit is used to synchronously push the candidate decision adopted by the decision terminal to the enterprise resource management system to perform a closed-loop operation.
[0139] The data analysis and decision-making device 600 can be implemented in the form of a computer program. Figure 9 Runs on the computer equipment shown.
[0140] See also Figure 9 , Figure 9 This is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 500 can be a terminal or a server. The terminal can be a desktop computer, tablet computer, smartphone, or other electronic device with communication capabilities. The server can be a standalone server or a server cluster consisting of multiple servers.
[0141] See Figure 9The computer device 500 includes a processor 502 , a memory, and a network interface 505 connected via a system bus 501 , wherein the memory may include a non-volatile storage medium 503 and an internal memory 504 .
[0142] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, which, when executed, can enable the processor 502 to execute a data analysis and decision-making method.
[0143] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.
[0144] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a data analysis and decision-making method.
[0145] The network interface 505 is used to communicate with other devices over the network. Figure 9 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0146] The processor 502 is configured to run a computer program 5032 stored in the memory to implement the steps of the above method.
[0147] It should be understood that in the embodiment of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0148] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.
[0149] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor performs the steps of the above method.
[0150] The storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.
[0151] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0152] In the several embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the various units is merely a logical functional division, and actual implementation may employ other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented.
[0153] The steps in the methods of the embodiments of the present invention may be adjusted in order, combined, or deleted as needed. The units in the devices of the embodiments of the present invention may be combined, divided, or deleted as needed. Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.
[0154] If this integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, terminal, or network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present invention.
[0155] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A data analysis and decision-making method, characterized in that: The method comprises: Perform distributed data collection from target data sources according to preset collection rules to collect raw data; The raw data is segmented using a preset segmentation rule, and then the data is cleaned using a pre-trained detection model to obtain data to be analyzed; Perform text mining analysis and association rule mining on the data to be analyzed, and write the results into the analysis result library; Constructing a target portrait based on the data in the analysis result library, and updating the target portrait as the original data is updated in real time; A decision analysis is performed based on the target profile using a preset decision algorithm to generate a decision report.
2. The data analysis and decision-making method according to claim 1, characterized in that: The steps of collecting raw data from the target data source according to the preset collection rules include: Passively collecting and acquiring the raw data from the recorded target data automatically and periodically at preset time intervals, and / or passively collecting and acquiring the raw data when a preset business event is detected based on an event trigger mechanism; Actively collect and obtain the raw data by initiating active inquiries to the target object under a preset specific business scenario; Centralize and store the raw data collected from different data sources through distributed collection; Desensitization processing is performed on the original data, where the desensitization processing includes at least one of masking desensitization, encryption desensitization, and generalization desensitization.
3. The data analysis and decision-making method according to claim 1, characterized in that: The step of segmenting the original data using a preset segmentation rule includes: Dividing the original data into data blocks divided into natural time units according to the data timestamp; Dividing the data block into two parts according to the service type, so that the same data block only contains data of a single service type; The data blocks after the secondary division are allocated storage nodes based on customer attributes, where the customer attributes include geographical distribution or customer level.
4. The data analysis and decision-making method according to claim 1, characterized in that: The step of cleaning data using the pre-trained detection model includes: Performing anomaly detection on the numerical data in the original data using an isolation forest algorithm, marking abnormal data that exceeds a preset confidence interval; Regular expression matching is used to remove garbled characters from the text data in the original data, and semantic legitimacy verification is performed based on a pre-trained text classification model, marking abnormal data that fails the verification; Manually review or automatically correct the abnormal data; A data cleaning log is established to record the manual review or automatic correction processing operation performed on the abnormal data.
5. The data analysis and decision-making method according to claim 1, characterized in that: The step of performing text mining analysis and association rule mining on the data to be analyzed and writing the results into an analysis result database comprises: Analyze the data to be analyzed using a natural language processing model based on deep learning, and mine and extract emotional tendencies, customer intentions, and key entities from the data to be analyzed; Generate a question type distribution map and customer intention classification results based on the sentiment tendency, customer intention, and key entities; Association rule mining is performed on the problem type distribution map and customer intention classification results using a preset distributed Apriori algorithm, and then the transaction item set is compressed using a bitmap and candidate item sets are generated across layers; The candidate item set is written into the analysis result library.
6. The data analysis and decision-making method according to claim 1, characterized in that: The step of constructing a target portrait based on the data in the analysis result library and updating the target portrait as the original data is updated in real time includes: Extracting multidimensional attributes from the data in the existing analysis result library, and performing cluster analysis on the multidimensional attributes based on dynamic weight allocation rules to generate the target portrait; Setting sensitivity levels for changes in the original data, and configuring differentiated update conditions for different sensitivity levels; When it is detected that the raw data collected in real time meets the update condition, an incremental update of the target portrait is triggered; The updated target profile is synchronized to all business terminals in real time through distributed storage.
7. The data analysis and decision-making method according to claim 1, characterized in that: The step of performing decision analysis based on the target profile using a preset decision algorithm to generate a decision report includes: Analyze the decision maker's historical operation records through a multimodal deep network to generate multiple candidate decision plans; Predicting the response after the candidate decision plan is implemented through a time series model to generate a prediction result; Generate an interactive decision report based on the candidate decision solutions and the prediction results, and push it to the decision terminal; The candidate decision adopted by the decision terminal is synchronously pushed to the enterprise resource management system to perform a closed-loop operation.
8. A data analysis and decision-making device, characterized in that: Used to execute the data analysis and decision-making method according to any one of claims 1 to 7.
9. A computer device, characterized in that: The computer device includes a memory and a processor connected to the memory; the memory is used to store a computer program; the processor is used to run the computer program stored in the memory to perform the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 can be implemented.
Citation Information
Cited By
Data cleaning and repairing method and device
CN121579461A
A data cleaning and repairing method and device
CN121579461B