A member big data analysis system and method based on data mining
By constructing a data fusion framework for members across all scenarios and mining causal relationships in behavioral time series, the problems of homogenization in multi-source data fusion and the disconnect between time series features and behavioral intentions have been solved. This has enabled the accurate extraction and standardized output of member features, thereby enhancing the enterprise's data support capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN GEEKSEN OPERATION MANAGEMENT CO LTD
- Filing Date
- 2026-03-06
- Publication Date
- 2026-06-09
AI Technical Summary
Existing big data analytics technologies for members have failed to effectively identify the correlation value of multi-source data, resulting in homogenized data fusion, a disconnect between temporal characteristics and behavioral intentions, and an inability to accurately extract members' real needs, thus affecting the effectiveness of personalized services and marketing strategies for enterprises.
We construct a data fusion framework for members across all scenarios, dynamically allocate the weights of multi-source data through scenario correlation, combine behavioral temporal causal correlation mining, accurately extract the real needs characteristics of members, and achieve standardized output.
It improves the accuracy and timeliness of member feature extraction, provides direct data support for enterprises, helps with personalized service delivery, marketing strategy formulation and product optimization, and enhances market competitiveness and operating efficiency.
Smart Images

Figure CN122175634A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of membership data analysis technology, specifically to a membership big data analysis system and method based on data mining. Background Technology
[0002] Membership serves as a long-term, stable identifier for businesses to connect with users. By providing exclusive benefits and personalized services, businesses achieve deep engagement with their members, who constitute a crucial part of their core user base. With the acceleration of digital transformation, membership data sources are becoming increasingly diverse, encompassing multiple dimensions such as consumption, interaction, smart device usage, and social sharing. The scale of this data continues to expand. Membership big data analytics refers to the process of collecting, organizing, analyzing, and mining this multi-dimensional membership data to extract key information such as member behavior patterns, demand tendencies, and consumption preferences. Its core significance lies in helping businesses break down information barriers, accurately understand core member needs, optimize product design, improve service quality, and develop targeted marketing strategies, ultimately enhancing member stickiness and loyalty, and improving the company's market competitiveness and operational efficiency.
[0003] However, existing membership big data analysis technologies still have certain shortcomings. The processing of multi-source membership data is limited to simple integration, without considering the correlation value of various types of data in different application scenarios. The use of a fixed weight allocation model leads to serious homogenization of data fusion. The analysis of the temporal characteristics of membership behavior is relatively isolated, relying solely on single time slice data for judgment. It fails to identify the causal relationships between continuous behavioral nodes, resulting in a disconnect between temporal characteristics and members' true behavioral intentions. Ultimately, this leads to insufficient accuracy in feature extraction and an inability to provide effective decision support for enterprises. Therefore, developing a membership big data analysis system and method based on data mining is of great significance. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a membership big data analysis system and method based on data mining. It can construct a membership full-scenario data association and fusion framework, dynamically allocate the weight of multi-source data based on scenario correlation, and combine behavioral temporal causal correlation mining to correct analysis biases. It can accurately extract the real needs characteristics of members and achieve standardized output, solving the problems of homogenization of multi-source data fusion and the disconnect between temporal characteristics and behavioral intentions in existing technologies. It not only improves the accuracy and timeliness of membership feature extraction, but also provides direct data support for enterprises to push personalized services, formulate marketing strategies, and optimize and upgrade products.
[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a membership big data analysis system based on data mining, the system comprising: a data acquisition module, a scene recognition module, a dynamic weight allocation module, a time-series correlation mining module, a feature extraction module, and a feature output module, wherein each module is connected by a preset process signal to form a data transmission link; The data acquisition module collects structured consumption data, unstructured comment text data, smart device behavior trajectory data, and social interaction data from members, and then transmits the data to the scene recognition module and the dynamic weight allocation module, respectively. The scene recognition module identifies the application scene based on the channel, environment, and type of member behavior, and sends the recognition result to the dynamic weight allocation module. The dynamic weight allocation module assigns weights to various types of data based on scene relevance and outputs fused data to the feature extraction module. The temporal correlation mining module mines the causal relationships between continuous behavior nodes of members and corrects analysis biases, and synchronizes the results to the feature extraction module; The feature extraction module extracts the real needs features of members based on the fused data and time series correlation results, and transmits them to the feature output module. The feature output module outputs the requirement features in a standardized format to the relevant business systems of the enterprise.
[0006] Furthermore, the data acquisition module performs the following operations when collecting multi-source member data: Connect to enterprise transaction systems, member evaluation platforms, smart device terminals, and social interaction interfaces; configure data transmission protocols and encryption rules; and establish a stable data transmission link. According to the preset data field specifications, structured consumption data, unstructured comment text data, smart device behavior trajectory data, and social interaction data are extracted by category. The missing value detection and validity verification algorithm is used to verify the integrity of the collected data, remove invalid missing data that does not meet the specifications, and establish a categorized data storage index.
[0007] Furthermore, the scene recognition module performs the following operations when recognizing application scenarios: Analyze member behavior data to extract channel identifiers, environmental sensor data, and behavioral operation records; Call the preset scene feature library, which contains feature thresholds and matching rules for offline consumption scenarios, online interaction scenarios, and consultation service scenarios; The extracted information is compared and matched with the feature library. Based on the matching results, the current application scenario of the member is determined, and a recognition result containing the scenario type and scenario attributes is generated and synchronized to the dynamic weight allocation module. At the same time, the scenario recognition log is updated.
[0008] Furthermore, the dynamic weight allocation module performs the following operations when allocating data weights: Receive the scene recognition results output by the scene recognition module, and analyze the scene type and core requirement dimensions; Retrieve the preset scenario-weight association rule. This rule divides the data into multiple weight levels based on the degree of correlation between the data and the scenario requirements, using a formula. Calculate the weights of each data type, where For the first Class data in the first Scene weights To determine the correlation between data and scenarios, For data credibility, The balancing coefficient is determined based on statistical analysis of historical scenario-data matching effects, reflecting the relative importance of relevance and credibility. Based on the calculation results, corresponding weight values are assigned to each type of collected data. The data is then integrated through a data fusion algorithm to form differentiated fused data, which is then output to the feature extraction module.
[0009] Furthermore, the time-series correlation mining module traces members' historical behavior records, sorts them by timestamp to form a continuous behavior sequence, and then uses a formula... Calculate the association strength between behavioral nodes, where For the first The and the first The association strength of each behavioral node , For the time when the behavior occurs, For behavioral logical relevance, The time-series adjustment coefficient is determined by verifying the correlation validity of the member's historical behavior links. It adapts to the time-series characteristics of different behavior types, analyzes the triggering conditions and subsequent responses of each behavior node in the sequence, identifies logically related causal relationships, and constructs a behavior link map that includes the order of behavior and the strength of correlation. Based on this map, the analysis results that rely solely on a single time slice data are corrected.
[0010] Furthermore, the feature extraction module employs a data filtering algorithm to remove redundant information and noisy data from the fused data after dynamic weight allocation, using the formula... Integrating feature parameters, where For the first Confidence level of each demand characteristic For the first In the class of data and the first Parameter values related to each feature The feature integration coefficient is determined by consistency analysis of feature extraction results from multiple scenarios based on feature classification standards. It balances the contribution of multi-source data and time-series correlation, performs correlation analysis between the filtered data and the behavioral link map output by the time-series correlation mining module, extracts key feature parameters directly related to members' behavioral intentions, and integrates them into consumption preference features and service demand features according to preset feature classification standards.
[0011] Furthermore, the feature output module incorporates interface adaptation protocols for various business systems within the enterprise. These protocols include data format specifications and transmission protocol standards for each business system. The extracted real needs features of members are mapped and format converted according to the requirements of the target business system to generate standardized feature data packets. These data packets are then pushed to the enterprise decision-making system, marketing system, or member service system through preset interfaces, and push status receipts are output simultaneously.
[0012] Furthermore, the data acquisition module adopts a real-time acquisition mode, sets a fixed acquisition cycle and a triggered acquisition mechanism, continuously connects to various data source ports to obtain the latest member behavior data, periodically scans the data types of member behavior, updates the data collection directory, and automatically configures the corresponding acquisition parameters, storage path and data verification rules when new data types appear, and generates a data type update report.
[0013] A data mining-based method for analyzing member big data, applicable to the aforementioned data mining-based member big data analysis system, includes the following steps: S1. Comprehensively collect members' structured consumption data, unstructured comment text data, smart device behavior trajectory data, and social interaction data through the data collection module; S2. Analyze the channel, environment, and type attributes of member behavior through the scene recognition module to determine the corresponding application scenarios; S3. The dynamic weight allocation module assigns dynamic weights to various types of collected data based on scene relevance to form fused data. S4. By using the time-series correlation mining module, we can sort out the continuous behavior nodes of members, identify the causal relationship between behaviors, and correct the analysis bias. S5. Extract the true needs features of members based on the fused data and time series correlation results through the feature extraction module; S6. After standardizing the requirement features through the feature output module, the feature output module outputs them to the relevant business systems of the enterprise.
[0014] Furthermore, in S4, the time-series correlation mining module first retrieves the member's historical behavior data, arranges it in the order of timestamps to form a complete behavior sequence, then analyzes the logical correlation and dependency of each node in the behavior sequence through a preset correlation strength calculation formula, identifies the behavior links with causal orientation, and finally systematically corrects the analysis results of a single time slice data based on the behavior links to form a time-series correlation analysis result, which is synchronized to the feature extraction module.
[0015] Compared with existing technologies, this data mining-based membership big data analysis system and method have the following beneficial effects: This invention constructs a data fusion framework for members across all scenarios, dynamically allocates multi-source data weights based on scenario relevance, and combines behavioral temporal causal correlation mining to correct analytical biases. It accurately extracts the true needs of members and achieves standardized output, solving the problems of homogenization in multi-source data fusion and the disconnect between temporal features and behavioral intentions in existing technologies. This not only improves the accuracy and timeliness of member feature extraction but also provides direct data support for personalized service delivery, marketing strategy formulation, and product optimization and upgrading. It helps enterprises optimize resource allocation, improve member service experience and stickiness, and further enhance their market competitiveness and operating efficiency.
[0016] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0018] Figure 1 This is a schematic diagram of the structure of a membership big data analysis system based on data mining; Figure 2 A flowchart of a membership big data analysis method based on data mining; Figure 3 This is a flowchart of a data mining-based big data analysis method for members. Detailed Implementation
[0019] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0020] This invention provides a membership big data analysis system and method based on data mining, aiming to solve the problems of homogenization in multi-source data fusion and the disconnect between time-series features and behavioral intentions in existing technologies. It improves the accuracy and timeliness of membership feature extraction, providing effective data support for enterprises. (See also...) Figure 1 , Figure 2 and Figure 3 The specific technical solution is as follows: The system consists of a data acquisition module, a scene recognition module, a dynamic weight allocation module, a time-series correlation mining module, a feature extraction module, and a feature output module. Each module is connected in sequence to form a complete data transmission link.
[0021] The data acquisition module adopts a real-time acquisition mode, connecting to multiple data sources such as enterprise transaction systems and member evaluation platforms. It collects structured consumption data, unstructured comment text data, smart device behavior trajectory data, and social interaction data from members. The module performs data verification through missing value detection and validity validation algorithms, removes invalid and missing data, and establishes a categorized storage index.
[0022] The scene recognition module analyzes member behavior data, extracts information such as channel identifiers and environmental sensor data, calls a preset feature library containing scenarios such as offline consumption, online interaction, and consultation services, determines the member's current application scenario through comparison and matching, and generates recognition results containing scenario type and attributes.
[0023] The dynamic weight allocation module assigns dynamic weights to various types of data based on the scene recognition results and the scene-weight association rules, integrating them to form differentiated fused data.
[0024] The time-series correlation mining module traces members' historical behavior records, forms continuous behavior sequences by timestamp, identifies causal relationships between behavior nodes, constructs behavior link graphs, and corrects analytical biases in single time slice data.
[0025] The feature extraction module filters out redundant and noisy information in the fused data and extracts key features such as member consumption preferences and service needs by combining the time-series correlation results.
[0026] The feature output module adapts to the interface protocols of various business systems within the enterprise, standardizes the required features, pushes them to the decision-making, marketing, or membership service systems, and provides synchronous feedback on the push status.
[0027] This method follows the steps of "collecting multi-source data, identifying behavioral scenarios, dynamically allocating data weights, mining temporal causal relationships, extracting real demand characteristics, and standardizing output," effectively helping enterprises optimize resource allocation, improve member service experience and stickiness, and enhance market competitiveness and operating efficiency. Example
[0028] This embodiment applies to the membership operation scenario of a chain retail enterprise. This enterprise has multiple physical stores, an online shopping platform, member communities, and smart shopping guide devices. Members generate multi-dimensional behavioral data daily, including offline shopping, online product browsing and ordering, community interaction, smart device operation, and service consultation feedback. As the enterprise's membership scale continues to expand, the multi-source data is experiencing explosive growth. Existing data analysis methods struggle to fully explore the correlation value of data across different scenarios. Fixed weight allocation leads to a lack of targeted data fusion and fails to effectively link the causal relationships of continuous member behavior, making it difficult for the enterprise to accurately grasp the true needs of members and affecting the effectiveness of personalized marketing and services. Therefore, this embodiment, based on the actual business scenario of this chain retail enterprise, deploys a data mining-based membership big data analysis system. Through end-to-end data analysis and processing, it achieves accurate extraction and standardized output of member demand characteristics.
[0029] See Figure 1 , Figure 2 and Figure 3 The specific implementation process of this embodiment is as follows: As the core component of the system's data input, the data acquisition module first connects to the enterprise's existing transaction systems, including the POS systems of offline stores and the order management system of the online mall. It also integrates with the member evaluation platform, encompassing the in-app evaluation section, the mini-program feedback entry, and smart devices such as smart shopping guide tablets and smart membership bracelets worn by members in stores. Furthermore, it connects to social interaction channels such as the enterprise's WeChat member community and official Weibo interactive interface. During the integration process, the system is configured with a unified data transmission protocol and encryption rules to ensure the security and stability of data during transmission, establishing a stable data transmission link covering multiple channels.
[0030] According to the preset data field specifications, the data collection module extracts four types of core data: structured consumption data, including members' spending amount, product categories purchased, purchase frequency, payment methods, etc.; unstructured review text data, covering members' written evaluations and feedback on product quality, service attitude, logistics speed, etc.; smart device behavior trajectory data, including members' records of searching for products through smart shopping guide devices in stores, device operation paths, dwell time, etc.; and social interaction data, involving members' posts in social groups, likes and reposts, questions asked to @official accounts, etc.
[0031] After data collection is complete, the data acquisition module initiates missing value detection and validity verification algorithms to perform integrity checks on various data types, eliminating invalid data with missing key fields, format errors, or logical contradictions. Simultaneously, it establishes a categorized storage index for valid data that passes verification, classifying and archiving it by data type and collection time for easy retrieval by subsequent modules. Furthermore, this module employs a real-time acquisition mode with a fixed acquisition cycle and a triggered acquisition mechanism. When a member engages in activities such as consumption, reviews, or interactions, data acquisition is immediately triggered, continuously connecting to various data source ports to obtain the latest member behavior data. The system also periodically scans for new member behavior data types. When new data types emerge, such as member interaction data on short video platforms, it automatically configures the corresponding acquisition parameters, storage paths, and data verification rules, generating a data type update report and providing it to the enterprise's technical management department.
[0032] After receiving the member behavior data transmitted by the data acquisition module, the scene recognition module first parses and processes the data to extract key information, including channel identifiers such as offline store codes, online mall identifiers, and social platform names; environmental sensor data such as temperature and humidity data in the store and device environment information of members logging into the online platform; and behavior operation records such as product purchase operations, consultation dialogue content, and device query instructions.
[0033] Subsequently, the scene recognition module calls the preset scene feature library, which has been pre-loaded with feature thresholds and matching rules corresponding to offline consumption scenarios, online interaction scenarios, and consultation service scenarios. The features of offline consumption scenarios include store channel identification, consumption payment operations, and on-site environmental data. The features of online interaction scenarios cover online platform identification, browsing and ordering behavior, and text evaluation records. The features of consultation service scenarios include consultation dialogue keywords, @official account operations, and problem feedback types.
[0034] The extracted key information about member behavior is compared one by one with the feature thresholds and matching rules in the scene feature library. Based on the matching results, the current application scene of the member is determined. For example, if a member completes a payment for goods in a store through a POS machine, and the data includes the store's channel identifier and on-site environmental sensor data, it is determined to be an offline consumption scene; if a member submits a product review in the APP, it is determined to be an online interaction scene. After determining the scene, a recognition result containing scene type and scene attributes is generated and synchronously transmitted to the dynamic weight allocation module. At the same time, the scene recognition log is updated to record key information such as the time of scene recognition, member ID, and recognition result.
[0035] After receiving the recognition results from the scene recognition module, the dynamic weight allocation module first analyzes the scene type and its corresponding core requirement dimensions. For example, the core requirement dimensions for offline consumption scenes are product preferences and purchasing power; for online interaction scenes, the core requirement dimensions are willingness to interact and feedback requests; and for consultation service scenes, the core requirement dimensions are problem-solving needs and service experience expectations. Next, the module retrieves preset scene-weight association rules. These rules divide data into multiple weight levels based on the degree of correlation between the data and scene requirements, with different correlation levels for different types of data in different scenes. In the specific implementation of this embodiment, the formula... Calculate the weights of each data type, where For the first Class data in the first Scene weights The degree of relevance between data and scenarios is determined based on the degree of fit between data types and the core requirements of the scenario. Data credibility is assessed by comprehensively evaluating the reliability of the data collection channels and the results of data integrity verification. The balancing coefficient is determined based on the statistical analysis results of the enterprise's historical scenario-data matching effect, and is used to reflect the relative importance of relevance and credibility in the weight calculation.
[0036] After calculating the weight values of various types of data according to the above formula, the dynamic weight allocation module uses a data fusion algorithm to integrate and process the collected data with different weights to form differentiated fusion data with specific scenarios. For example, in offline consumption scenarios, the weight of structured consumption data is higher than that of other types of data. The fused data will highlight consumption-related information, and then the differentiated fusion data will be output to the feature extraction module.
[0037] Once the time-series correlation mining module is activated, it first traces the historical behavior records of the chain retail enterprise's members, retrieves all the members' behavior data over a period of time from the data storage index, sorts them according to the order of timestamps, and forms a continuous behavior sequence. For example, a member first browses a clothing product in an online mall, then inquires about the size of the product in the member community, and three days later goes to a physical store to try it on and complete the purchase. This series of behaviors will form a complete behavior sequence in chronological order.
[0038] In the specific implementation process of this embodiment, through formula Calculate the association strength between behavior nodes, where For the first The and the first The association strength of each behavioral node and These represent the occurrence times of the two behavioral nodes. The degree of logical correlation between behaviors is determined based on the causal relationship and the degree of functional correlation between behaviors. The time-series adjustment coefficient is determined by the verification results of the correlation effectiveness of the member's historical behavior links. It can adapt to the time-series characteristics of different behavior types. For example, the time-series correlation characteristics of consumption behavior and consultation behavior are different from those of browsing behavior and interaction behavior.
[0039] Based on the calculated association strength, the triggering conditions and subsequent responses of each behavioral node in the behavioral sequence are analyzed to identify logically related causal relationships. For example, a member's inquiry about product sizes is a key reason for subsequent offline try-on and purchase behavior. A behavioral link graph, including the sequence of behaviors and the strength of association, is then constructed. Based on this graph, the analysis results relying solely on single time slice data are corrected. For instance, a single time slice might show a member browsing high-end home appliances, but the behavioral link graph reveals subsequent inquiries about entry-level home appliances. After correction, it is determined that the member's actual need is for entry-level appliances, not high-end ones. Finally, the temporal association analysis results are synchronized to the feature extraction module.
[0040] After receiving the differentiated fusion data output by the dynamic weight allocation module, the feature extraction module first uses a data filtering algorithm to remove redundant information and noise data, such as irrelevant chat content of members in the community, duplicate submissions of the same evaluation, and device records generated by accidental operation, and retains the effective data related to the member's behavioral intention.
[0041] In the specific implementation process of this embodiment, through formula Integrating feature parameters, where For the first Confidence level of each demand characteristic For the first In the class of data and the first Parameter values related to each feature The feature integration coefficient is determined based on a preset feature classification standard and through consistency analysis of feature extraction results from multiple scenarios. It is used to balance the contribution of multi-source data and time-series correlation results in feature extraction.
[0042] The filtered valid data is correlated with the behavioral link graph output by the time-series correlation mining module to extract key feature parameters directly related to members' behavioral intentions. For example, by combining consumption data and behavioral links, members' preferred product categories, price ranges, and purchase cycles are extracted. Based on consultation records and interaction data, service demand features such as members' requirements for service response speed and their concerns about after-sales guarantees are extracted. Finally, according to preset feature classification standards, the key feature parameters are integrated to form a complete set of members' real needs features, which is then transmitted to the feature output module.
[0043] The feature output module incorporates interface adaptation protocols for various enterprise business systems. These protocols include data format specifications and transmission protocol standards for the enterprise decision-making system, marketing system, and membership service system. After receiving the set of real member demand features transmitted by the feature extraction module, the module performs field mapping and format conversion on the demand features according to the specific requirements of the target business system. For example, for the marketing system, it converts consumer preference features into a product recommendation tag format that the system can recognize; for the membership service system, it converts service demand features into a service priority identifier format.
[0044] After format conversion, standardized feature data packages are generated. These packages are then pushed to the corresponding enterprise business systems via preset interfaces. For example, a data package containing member consumption preferences and service needs is pushed to the marketing system to provide data support for targeted promotional activities; the feature data package is pushed to the member service system to assist customer service personnel in providing personalized services. Simultaneously, the feature output module outputs push status receipts, providing feedback on the data package push progress and whether it was successfully delivered, facilitating monitoring of data transmission by enterprise technical personnel.
[0045] In summary, this embodiment, through actual deployment and application in a chain retail enterprise, fully demonstrates the entire process of a data mining-based membership big data analysis system. The system achieves comprehensive collection, verification, and archiving of multi-channel membership data through a data acquisition module; accurately locates membership behavior scenarios through a scene recognition module; achieves differentiated data fusion based on scene characteristics through a dynamic weight allocation module; effectively identifies the causal relationships of continuous membership behaviors and corrects analytical biases through a time-series correlation mining module; accurately extracts membership consumption preferences and service demand characteristics through a feature extraction module; and completes standardized data output and integration with business systems through a feature output module. The entire implementation process solves the core problems of homogenization in multi-source data fusion and the disconnect between time-series features and behavioral intentions in existing technologies, significantly improving the accuracy and timeliness of membership feature extraction. Example
[0046] This embodiment applies to the membership service scenario of an online education platform. This platform encompasses multiple service modules, including live courses, recorded learning, assignment submission, Q&A communities, and learning groups. Members generate various behavioral data as learners, including course learning records, assignment completion status, Q&A interaction content, community discussions, and smart learning device usage patterns. As the platform's membership grows, the needs of members of different ages and with different learning goals vary significantly. Existing data analysis methods have obvious limitations: simply piling up multi-source data obscures core learning needs; fixed weight allocation cannot adapt to the data analysis focus of different learning scenarios; and isolated analysis of learning behavior at a single point in time makes it difficult to capture the causal relationship of the complete learning chain from previewing, learning, practicing to consolidation, resulting in insufficient targeted course recommendations, untimely Q&A service responses, and a lack of personalized learning plans. Therefore, this embodiment, based on the aforementioned embodiments, deploys this membership big data analysis system in conjunction with the business characteristics of the online education platform to accurately uncover students' real learning and service needs.
[0047] See Figure 1 , Figure 2 and Figure 3 The specific implementation process of this embodiment is as follows: Building upon the data acquisition process described in the aforementioned embodiments, the data acquisition module specifically interfaces with the core business systems and data sources of the online education platform. First, it integrates with the learning management system, encompassing live-streaming tools, recorded course players, and assignment submission platforms. Simultaneously, it connects to the member evaluation area, Q&A community interface, learning community management system, and commonly used smart learning devices such as tablets and smart pens. At the data transmission level, it utilizes pre-defined data transmission protocols and encryption rules to ensure the security and stability of student learning and interaction data transmission, constructing a data transmission link covering the entire learning process.
[0048] Based on the pre-defined data field specifications for online education scenarios, four types of core data were extracted: structured learning data, including course completion rate, learning duration, homework accuracy, exam scores, and course collection records; unstructured comment text data, covering students' written evaluations of course content, teacher teaching style, and platform services, as well as question descriptions and feedback messages posted in the Q&A community; smart device behavior trajectory data, including course playback progress on learning tablets, knowledge point marking records, homework answering trajectories, and key point annotations in smart pen writing; and social interaction data, involving topic discussions within learning communities, mutual Q&A messages among students, course sharing behavior, and bullet screen content in live streaming interactions.
[0049] After data collection, missing value detection and validity verification algorithms are activated to remove invalid information such as invalid answer records, blank evaluations, and device data generated by erroneous operations. A categorized storage index is established for valid data, archived according to the categories of "learning data - interaction data - device data - evaluation data" and collection time. Simultaneously, a real-time collection mode is maintained, with a fixed collection cycle and a triggered collection mechanism. Data collection is immediately triggered when students complete course learning, submit assignments, or post Q&A. The system periodically scans data types. When new types of data are added, such as live-stream interaction data or virtual classroom operation data, the system automatically completes the configuration of collection parameters, storage path settings, and verification rule adaptation, and generates a data type update report for submission to the platform's technical department.
[0050] After receiving student behavior data transmitted by the data acquisition module, the scene recognition module prioritizes parsing key information, including behavior channel identifiers such as live classroom ID, recorded course player identifier, and Q&A community entry identifier; environmental sensor data such as the network environment of the student's login device and the learning time period; and specific behavior operation records such as course playback pause operation, homework submission action, and Q&A question posting behavior.
[0051] Then, a pre-defined feature library adapted for online education scenarios is invoked. This feature library contains feature thresholds and matching rules corresponding to course learning scenarios, assignment submission scenarios, Q&A scenarios, and community mutual assistance scenarios. Features for course learning scenarios include learning duration thresholds, course playback operation records, and knowledge point marking behaviors; features for assignment submission scenarios cover assignment upload operations, answering time, and incorrect question submission records; features for Q&A scenarios include question description keywords, @teacher actions, and multiple follow-up questions; and features for community mutual assistance scenarios involve topic discussion keywords, mutual assistance Q&A messages, and course resource sharing records.
[0052] The extracted student behavior information is compared and matched with a scene feature database to determine the student's current application scenario based on the matching results. For example, if a student continuously plays a recorded course and marks knowledge points, it is determined to be a course learning scenario; if a student posts a message containing specific knowledge point questions and @ the teacher answering the question, it is determined to be a Q&A consultation scenario. Recognition results containing scene type and scene attributes are generated and synchronously transmitted to the dynamic weight allocation module, while the scene recognition log is updated to record key information of scene recognition.
[0053] After receiving the scene recognition results, the dynamic weight allocation module analyzes the scene type and its corresponding core requirement dimensions. The core requirement dimensions for a course learning scene are the level of knowledge mastery and the suitability of the learning pace; for an assignment submission scene, the core requirement dimensions are knowledge gaps and the need for answering methods; for a Q&A consultation scene, the core requirement dimensions are problem-solving efficiency and the need for knowledge extension; and for a community mutual assistance scene, the core requirement dimensions are the willingness to communicate and the need for resource acquisition.
[0054] The preset scene-weight association rules are retrieved, and weight levels are assigned based on the correlation between the data and scene requirements. In the specific implementation of this embodiment, the formula is used... Calculate the weights of each data type. Based on the calculation results, assign corresponding weight values to each type of collected data. For example, in a course learning scenario, structured learning data has a higher weight than other types of data, highlighting core information such as learning duration and course completion rate; in a Q&A scenario, unstructured comment text data has the highest weight, prioritizing the capture of students' question descriptions and expressed needs. Integrate the data using a data fusion algorithm to form differentiated fused data, and output it to the feature extraction module.
[0055] The temporal correlation mining module first retrieves the learner's historical behavior records, extracts complete behavioral data of the learner over a period of time from the categorized storage index, and sorts it according to the chronological order of timestamps to form a continuous learning behavior sequence. For example, a learner first previews a recorded course on a certain chapter, then participates in a live course on the same chapter, submits related assignments after the live course, makes mistakes on geometry questions in the assignments, then asks questions about geometry problem-solving methods in the Q&A community, and finally watches a specialized course explaining geometry knowledge points, thus forming a complete learning behavior chain.
[0056] In the specific implementation process of this embodiment, through formula Calculate the correlation strength between behavioral nodes. Based on the correlation strength analysis, identify the triggering conditions and subsequent responses of each behavioral node, and pinpoint the causal relationships. For example, a geometry error in a homework assignment is the direct cause of the Q&A session, and watching a specialized lecture after the Q&A session is a reinforcement behavior for that problem. Construct a learning behavior chain graph that includes the sequence of behaviors and the correlation strength, and use this graph to correct the analytical bias of single time slice data. For example, a single time slice shows that a student is watching an introductory English course, but the behavior chain graph shows that they have already completed an intermediate English course, and this viewing is for reviewing basic knowledge points. After correction, it is determined that their actual need is to reinforce basic knowledge points rather than introductory learning, and the corrected temporal correlation results are synchronized to the feature extraction module.
[0057] After receiving the differentiated fused data, the feature extraction module uses a data filtering algorithm to remove redundant information and noisy data, such as irrelevant chat content in the learning community, course playback records generated by accidental clicks, and duplicate submissions of the same evaluations, retaining effective data that is directly related to students' learning intentions and service needs.
[0058] In the specific implementation process of this embodiment, through formula Integrate feature parameters. Perform correlation analysis between the filtered valid data and the learning behavior link graph output by the time-series correlation mining module to extract key feature parameters. Combined with the preset feature classification standards for online education scenarios, integrate them into two core feature categories: learning preference features, including preferred course types (e.g., live or recorded courses), areas of strength and weakness, distribution of learning time periods, and preferred teaching pace; and service demand features, covering requirements for Q&A response speed, level of detail in homework feedback, personalized learning planning, and access to supplementary materials. The integrated set of real student demand features is then transmitted to the feature output module.
[0059] The feature output module incorporates interface adaptation protocols for various business systems within the online education platform, including data format specifications and transmission protocol standards for the course recommendation system, teaching service system, teacher management system, and membership operation system. Upon receiving the requirement feature set, it performs field mapping and format conversion according to the requirements of the target business system. For example, it converts learning preference features into a course tag matching format for the course recommendation system, and converts service requirement features into a service priority and type identifier format for the teaching service system.
[0060] After generating standardized feature data packages, they are pushed to the corresponding business systems through preset interfaces: learning preference features are pushed to the course recommendation system to support personalized course recommendations; service demand features are pushed to the teaching service system to assist teachers in accurately responding to student questions and optimizing homework grading feedback; student knowledge point mastery features are pushed to the teacher management system to help teachers adjust their teaching strategies; and comprehensive demand features are pushed to the membership operation system to support the optimized configuration of membership benefits and services. Simultaneously, push status receipts are output to provide feedback on the data package transmission progress and delivery status, facilitating monitoring and maintenance by technical personnel.
[0061] In summary, this embodiment, combined with the business characteristics of the online education platform, completed the targeted deployment and application of the system based on the aforementioned embodiments, demonstrating the adaptability and practicality of the analysis system in the vertical field. Through a data collection scheme adapted to the education scenario, it achieved comprehensive capture of students' learning data throughout the entire process; based on the identification mechanism of the education-specific scenario feature library, it accurately located student behavior scenarios; dynamic weight allocation and time-series correlation mining effectively focused on core needs, correcting the analytical bias of isolated data; feature extraction and standardized output achieved precise alignment between requirements and business systems.
[0062] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A data mining based member big data analysis system, characterized by, The system includes: a data acquisition module, a scene recognition module, a dynamic weight allocation module, a temporal correlation mining module, a feature extraction module, and a feature output module; The data acquisition module collects structured consumption data, unstructured comment text data, smart device behavior trajectory data, and social interaction data from members, and then transmits the data to the scene recognition module and the dynamic weight allocation module, respectively. The scene recognition module identifies the application scene based on the channel, environment, and type of member behavior, and sends the recognition result to the dynamic weight allocation module. The dynamic weight allocation module assigns weights to various types of data based on scene relevance and outputs fused data to the feature extraction module. The temporal correlation mining module mines the causal relationships between continuous behavior nodes of members and corrects analysis biases, and synchronizes the results to the feature extraction module; The feature extraction module extracts the real needs features of members based on the fused data and time series correlation results, and transmits them to the feature output module. The feature output module outputs the requirement features in a standardized format to the relevant business systems of the enterprise.
2. The member big data analysis system based on data mining according to claim 1, characterized in that, The data acquisition module performs the following operations when collecting multi-source member data: Connect to enterprise transaction systems, member evaluation platforms, smart device terminals, and social interaction interfaces; configure data transmission protocols and encryption rules; and establish a stable data transmission link. According to the preset data field specifications, structured consumption data, unstructured comment text data, smart device behavior trajectory data, and social interaction data are extracted by category. The missing value detection and validity verification algorithm is used to verify the integrity of the collected data, remove invalid missing data that does not meet the specifications, and establish a categorized data storage index.
3. The member big data analysis system based on data mining according to claim 1, characterized in that, The scene recognition module performs the following operations when recognizing application scenarios: Analyze member behavior data to extract channel identifiers, environmental sensor data, and behavioral operation records; Call the preset scene feature library, which contains feature thresholds and matching rules for offline consumption scenarios, online interaction scenarios, and consultation service scenarios; The extracted information is compared and matched with the feature library. Based on the matching results, the current application scenario of the member is determined, and a recognition result containing the scenario type and scenario attributes is generated and synchronized to the dynamic weight allocation module. At the same time, the scenario recognition log is updated.
4. The member big data analysis system based on data mining according to claim 1, characterized in that, The dynamic weight allocation module performs the following operations when allocating data weights: Receive the scene recognition results output by the scene recognition module, and analyze the scene type and core requirement dimensions; Retrieve the preset scenario-weight association rule. This rule divides the data into multiple weight levels based on the degree of correlation between the data and the scenario requirements, using a formula. Calculate the weights of each data type, where For the first Class data in the first Scene weights To determine the correlation between data and scenarios, For data credibility, This is the balance coefficient; Based on the calculation results, corresponding weight values are assigned to each type of collected data. The data is then integrated through a data fusion algorithm to form differentiated fused data, which is then output to the feature extraction module.
5. A membership big data analysis system based on data mining according to claim 1, characterized in that, The time-series correlation mining module traces members' historical behavior records, sorts them by timestamp to form a continuous behavior sequence, and then uses a formula... Calculate the association strength between behavioral nodes, where For the first The and the first The association strength of each behavioral node , For the time when the behavior occurs, For behavioral logical relevance, As a time-series adjustment coefficient, the triggering conditions and subsequent responses of each behavioral node in the sequence are analyzed to identify logically related causal relationships. A behavioral link graph containing the order of behavior and the strength of association is constructed, and the analysis results that rely solely on a single time slice data are corrected based on this graph.
6. The membership big data analysis system based on data mining according to claim 1, characterized in that, The feature extraction module uses a data filtering algorithm to remove redundant information and noisy data from the fused data after dynamic weight allocation, through a formula... Integrating feature parameters, where For the first Confidence level of each demand characteristic For the first In the class of data and the first Parameter values related to each feature The feature integration coefficient is used to perform correlation analysis between the filtered data and the behavioral link map output by the time series correlation mining module, extract key feature parameters directly related to members' behavioral intentions, and integrate them into consumption preference features and service demand features according to preset feature classification standards.
7. A membership big data analysis system based on data mining according to claim 1, characterized in that, The feature output module has built-in interface adaptation protocols for various business systems of the enterprise. These protocols include data format specifications and transmission protocol standards for each business system. According to the requirements of the target business system, the extracted real needs features of members are mapped and format converted to generate standardized feature data packets. The data packets are pushed to the enterprise decision-making system, marketing system or member service system through preset interfaces, and push status receipts are output simultaneously.
8. A membership big data analysis system based on data mining according to claim 1, characterized in that, The data acquisition module adopts a real-time acquisition mode, sets a fixed acquisition cycle and a triggered acquisition mechanism, continuously connects to various data source ports to obtain the latest member behavior data, periodically scans the data types of member behavior, updates the data collection directory, and automatically configures the corresponding acquisition parameters, storage path and data verification rules when new data types appear, and generates a data type update report.
9. A method for analyzing member big data based on data mining, applicable to the member big data analysis system based on data mining as described in any one of claims 1-8, characterized in that, The method includes the following steps: S1. Comprehensively collect members' structured consumption data, unstructured comment text data, smart device behavior trajectory data, and social interaction data through the data collection module; S2. Analyze the channel, environment, and type attributes of member behavior through the scene recognition module to determine the corresponding application scenarios; S3. The dynamic weight allocation module assigns dynamic weights to various types of collected data based on scene relevance to form fused data. S4. By using the time-series correlation mining module, we can sort out the continuous behavior nodes of members, identify the causal relationship between behaviors, and correct the analysis bias. S5. Extract the true needs features of members based on the fused data and time series correlation results through the feature extraction module; S6. After standardizing the requirement features through the feature output module, the feature output module outputs them to the relevant business systems of the enterprise.
10. A method for analyzing member big data based on data mining according to claim 9, characterized in that, In step S4, the time-series correlation mining module first retrieves the member's historical behavior data, arranges it in the order of timestamps to form a complete behavior sequence, then analyzes the logical correlation and dependency of each node in the behavior sequence through a preset correlation strength calculation formula, identifies the behavior links with causal orientation, and finally systematically corrects the analysis results of a single time slice data based on the behavior links to form a time-series correlation analysis result, which is synchronized to the feature extraction module.