A Big Data Processing Method and System Based on User Behavior Information
By implanting the buried point SDK in the news client to collect user behavior data and perform data cleaning and analysis, the problem of incomplete data collection is solved, personalized recommendation and real-time operation are achieved, and user experience and operation efficiency are improved.
Patent Information
- Application Number
- CN202510351953.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The existing data analysis methods for user behavior of news clients have problems such as incomplete data collection, single analysis methods, lack of personalized services and low operational efficiency, and cannot meet the personalized recommendation and real-time operation requirements of user needs.
By implanting buried point SDK on the client, uploading it to the backend server using the HTTP protocol to perform data cleaning, format checksum conversion, and data separation and prediction are combined with data filtering nodes to display user behavior data trends in real time, and data expansion and personalized recommendations are used to use generative adversarial networks.
It realizes comprehensive collection and in-depth analysis of user behavior data, provides personalized services, improves operational efficiency, can monitor the number of online users in real time and adjusts operation strategies to meet diversified operational needs.
Smart Images

Figure CN119883824B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet information technology, particularly to the data analysis and operation of news clients in the media industry, and relates to a method for identifying data violation operation behaviors of a business system based on artificial intelligence. Background Art
[0002] In today's digital information age, the way of news dissemination and acquisition has undergone a huge transformation. As a convenient and efficient news dissemination platform, news clients have become an important channel for people to obtain news information. With the rapid development of mobile Internet and the wide popularity of smart phones, the number of users of news clients is constantly increasing, and the market competition is becoming increasingly fierce.
[0003] In order to stand out in the competition, news client providers need to deeply understand user needs, provide personalized news services, and improve user experience and satisfaction. The key to achieving this goal lies in the in-depth analysis of news client user behaviors.
[0004] Traditional news dissemination methods mainly rely on the subjective judgment and experience of editors, and cannot accurately understand the real needs and interest preferences of users. With the rise of big data technology, news client providers have begun to use data analysis technology to understand user behaviors, optimize news recommendation algorithms, and improve the accurate push ability of news. However, the existing user behavior data analysis methods have the following problems:
[0005] Incomplete data collection: The current news client user behavior data collection mainly focuses on aspects such as user browsing history and click behaviors, while the collection of important data such as user reading duration, sharing behaviors, and comment content is not comprehensive enough, resulting in inaccurate data analysis results.
[0006] Single data analysis method: The existing user behavior data analysis mainly adopts statistical analysis methods, lacking in-depth mining and analysis of data. For example, it is impossible to accurately classify and predict user interest preferences, nor can it deeply understand and analyze user behavior patterns.
[0007] Lack of personalized services: Although the existing news clients provide a certain degree of personalized recommendation services, they often only make simple recommendations based on user browsing history and interest tags, and cannot truly meet the personalized needs of users. For example, it is impossible to make accurate personalized recommendations according to factors such as user reading habits, time preferences, and geographical locations.
[0008] Low operation efficiency: Due to the lack of effective user behavior data analysis and operation management tools, news client providers are inefficient in content planning, recommendation algorithm optimization, user operation, etc., and cannot respond to market changes and user needs in a timely manner.
[0009] In summary, to solve the problems existing in the existing news client user behavior data analysis method and improve the user experience and satisfaction of the news client, there is an urgent need for a new news client user behavior data analysis and operation system that can comprehensively and accurately collect user behavior data, use advanced data analysis methods for in-depth mining and analysis, provide personalized news services, and improve operation efficiency. Summary of the Invention
[0010] The purpose of this part is to outline some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Simplifications or omissions may be made in this part, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this part, the abstract, and the title, and such simplifications or omissions shall not be used to limit the scope of the present invention.
[0011] In view of the above and / or the problems existing in an existing big data processing method based on user behavior information, the present invention is proposed.
[0012] Therefore, the problems to be solved by the present invention are
[0013] Currently, there are also SaaS platforms for collecting and analyzing user behavior data on the market, but the data sources for analysis can only be the user behavior data of the client, the data sources are relatively single, the industry characteristics are not obvious, and customized analysis cannot be performed for specific scenarios. At the same time, these SaaS platforms are basically based on offline computing and analysis, and cannot meet the real-time requirements of operation.
[0014] To solve the above technical problems, the present invention provides the following technical solution: a big data processing method based on user behavior information, which includes,
[0015] By implanting a buried point SDK in the client to collect user behavior data and uploading it to the backend server through the HTTP protocol, the user behavior data includes the click preview behavior of the user on the client;
[0016] After the load balancer of the backend server receives the user behavior data in the HTTP protocol, it balances the pressure on each data cleaning node. The data cleaning node parses the original data of the client and puts the parsed behavior data into a log file;
[0017] The data processing node checks and converts the format of the behavior data to form converted data and writes it into the stored message queue;
[0018] The data screening node listens to and calculates the converted data in the message queue through a self-calculated factor, and divides the converted data into user static data information and user dynamic data information;
[0019] Put the user static data information into the data warehouse, and regularly read the user dynamic data information in real time via the data recording node. Put the recorded user dynamic data information content into the data warehouse in units of recording time, and arrange the user dynamic data information and user static data information recorded in units of recording time in the database in units of time;
[0020] Via the data screening node, calculate and predict the user dynamic data information by self-calculating factors to obtain the predicted data information. Combine the data warehouse information and the predicted data information to draw the trend of user behavior data and present it in the form of a chart on the client side.
[0021] As a preferred solution of the big data processing method based on user behavior information of the present invention, wherein: collecting the user's behavior data by implanting a buried point SDK on the client side and uploading it to the backend server through the HTTP protocol includes;
[0022] The user behavior data includes the start and exit of the client, and the click and browse of advertisements, news, live broadcasts and operation activities;
[0023] After the client collects the user's behavior data, it is temporarily stored locally on the client side. Through a pre-designed timing mechanism, the collected behavior data is defined and summarized and reported to the backend server through the HTTP protocol, and is received by the load balancer of the backend server.
[0024] As a preferred solution of the big data processing method based on user behavior information of the present invention, wherein: the data cleaning node parses the original data of the client including;
[0025] IP address, User-Agent information;
[0026] The User-Agent information includes browser type information, version information, and operating system information.
[0027] As a preferred solution of the big data processing method based on user behavior information of the present invention, wherein: the data processing node forms conversion data after validating and converting the behavior data format including:
[0028] The data format verification includes verifying and disassembling the user behavior data into three-level verification layers of syntax, semantics, and behavior, and each layer is independent but the information is interconnected;
[0029] The verification result reversely trains the rule generation model to form a closed-loop optimization;
[0030] Allocate different types of data to different processing pipelines, and convert the user behavior data qualified after verification through the processing pipelines;
[0031] The reverse training rule generation model includes automatically adjusting the verification strictness according to data timeliness, performing spatio-temporal alignment verification on client text logs and real-time operation videos, encoding verification rules into qubit states, generating a unique spatio-temporal encoding for each piece of data, and realizing a superlinear rule combination.
[0032] As a preferred solution of the big data processing method based on user behavior information according to the present invention, wherein: the superlinear rule combination includes:
[0033] Calculate the comprehensive verification score S of the unique spatio-temporal encoding, and select a processing strategy according to the score:
[0034] When S≥0.9, the data processing node directly performs data conversion on the user behavior data and marks it as trusted data;
[0035] When 0.7≤S<0.9, the data processing node writes the user behavior data into a temporary queue for manual review;
[0036] When S<0.7, the data processing node transfers the user behavior data to an error correction pipeline for interception;
[0037] The calculation method of the comprehensive verification score S includes:
[0038]
[0039] Where a is the semantic score, b is the language score, λ(t) identifies the timeliness sensitive function, the semantic score is verified by the semantic verification layer, and the language score is verified by the language verification layer;
[0040] The timeliness sensitive calculation method is:
[0041]
[0042] Where t is the data generation time.
[0043] In view of the above and / or existing problems in a big data processing method based on user behavior information, the present invention is proposed.
[0044] Therefore, to solve the above technical problems, the present invention provides the following technical solution: a big data processing system based on user behavior information, which includes,
[0045] A front-end server module 100, a back-end server module 200 and a storage module 300;
[0046] The front-end server module 100 includes a client input unit 101 and a client display unit 102. The client input unit contains an implanted buried point SDK for collecting data information input by users in the client input unit 101. The client display unit 102 displays the user static data, predicted user data, and the trend of user behavior data of the savings module 300;
[0047] The back-end server module 200 includes a load balancer unit 201, a data cleaning unit 202, a data processing unit 203, and a data screening unit 204;
[0048] After receiving the data input by the client input unit 101, the load balancer unit 201 analyzes the load capacity of the data cleaning unit 202 and allocates the user behavior data to the data cleaning unit 202 for cleaning. The data processing unit 203 checks and converts the user behavior data after data cleaning to form converted data, which is written into the message queue of the storage module 300;
[0049] The data screening unit 204 listens to and calculates the converted data in the message queue of the storage module 300 through self-calculating factors to form user static data information and user dynamic data information. The data screening unit 204 processes the user dynamic data information to form user data information and predicted data information within a stage, inserts the user data information within the stage into the user static data information in chronological order to form complete user static data information, puts the complete user static data information into the first block of the storage module 300, puts the predicted data information into the second block of the storage module 300, combines the data warehouse information and the predicted data information to draw the trend of user behavior data, and presents it in the form of a chart in the third block of the storage module 300;
[0050] The stored content of the storage module 300 can be displayed on the client display unit 102.
[0051] As a preferred solution of the big data processing system based on user behavior information of the present invention, wherein: the data cleaning method of the data cleaning unit 202 for user data information includes:
[0052] Construct a user behavior knowledge graph and present the user abnormal pattern data in the client input unit 101 for the customer to confirm. If the abnormal data appears in the client input unit 101 and the customer selects to confirm and fills in the input password, the behavior data information confirmed by the user is submitted to the data processing unit 203. If the abnormal data appears in the client input unit 101 and the customer does not select to confirm or does not fill in the correct password, the user abnormal data is deleted;
[0053] Building a user behavior knowledge graph includes recording user nodes to establish a knowledge graph, where the user nodes include user IDs, user device types, and user's commonly used functions.
[0054] As a preferred solution of the big data processing system based on user behavior information according to the present invention, wherein: the stored content of the storage module 300 can be displayed on the client display unit 102, including a first block, a second block, and a third block, and the first block, the second block, and the third block are arranged in sequence in the interface display of the client display unit 102.
[0055] The present invention provides the following technical solution: An electronic device includes:
[0056] One or more processors;
[0057] A storage device having one or more programs stored thereon;
[0058] When the one or more programs are executed by the one or more processors, the one or more processors implement a big data processing method based on user behavior information.
[0059] The present invention provides the following technical solution: An electronic device includes:
[0060] A computer-readable storage medium having executable instructions stored thereon, and when the instructions are executed by a processor, the processor implements a big data processing method based on user behavior information.
[0061] As a preferred solution of the big data processing system based on user behavior information according to the present invention, wherein: after the display module displays the data input by the display classifier module, it performs a display of illegal operations and processes warning signals for the illegal behavior and the location where the illegal behavior data is generated.
[0062] The present invention proposes a data augmentation method and system based on a generative adversarial network. Compared with the prior art, the present invention can monitor the number of online users on the client in real time on the data operation platform, facilitating the operation personnel to adjust the operation and promotion strategies in a timely manner and achieve precise operation. By calculating and analyzing, users are grouped and feature tags are added, enabling the operation personnel to perform personalized recommendations based on aspects such as user location and interests, meeting diverse operation requirements. At the same time, on the operation platform, statistical data and trend charts of client news, live broadcasts, activities, etc. can be presented, enabling decision-making leaders to more conveniently grasp the operation status of the client application. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings. Among them:
[0064] Figure 1 It is a flowchart of the implementation method of a big data processing method based on user behavior information in Embodiment 1.
[0065] Figure 2 It is a system structure flowchart of a big data processing system based on user behavior information in Embodiment 2. Detailed implementation manners
[0066] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific implementation manners of the present invention with reference to the accompanying drawings of the specification.
[0067] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0068] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The "in one embodiment" appearing in different places in this specification does not all refer to the same embodiment, nor is it an independent or selectively exclusive embodiment from other embodiments.
[0069] Embodiment 1
[0070] Refer to Figure 1 This is the first embodiment of the present invention. This embodiment provides a big data processing method based on user behavior information, which includes:
[0071] Step 1, collection of user behavior data
[0072] As Figure 1 shown, by implanting a buried point SDK in the client, collect the user's behavior data, including behaviors such as client startup, exit, click and browse of advertisements, news, live broadcasts, and operation activities. After the data is collected, it is temporarily stored locally on the client. Through a pre-designed timing mechanism, the collected behavior data is reported to the backend service in batches through the HTTP protocol, avoiding the performance impact on the client and the backend service caused by frequent calls to the data interface.
[0073] Step 2, transfer and transmission of user behavior data
[0074] After the user behavior data is reported via the HTTP protocol, it reaches the backend load balancer. The load balancer is responsible for balancing the pressure on each data processing node. The processing mainly includes parsing the original information of the client request, such as the IP address, User-Agent (including information such as browser type, version, and operating system). The processed data is recorded in the log file on the local disk, waiting for further processing by Flume. Flume is responsible for reading the behavior data in the log file, and after simple format verification and conversion, writes it into the message queue Kafka for storage.
[0075] Step 3, Real-time Analysis and Processing of User Behavior Data
[0076] The real-time data processing is implemented based on the Flink framework. Through self-developed operators, it listens to the data in the message queue Kafka in real time, and calculates real-time data such as the number of online users and the number of visits according to business requirements. The calculated result data is stored in the Redis cache, and combined with a scheduled task, the cached data is updated to MySQL regularly. In addition, the operator is also responsible for batch inserting the processed behavior data into ElasticSearch for detailed query and backtracking by the subsequent business system.
[0077] In the inductive example, the core creativity lies in the dynamic coupling of hierarchical processing and feedback loop:
[0078] Hierarchical abstraction: Decompose data verification into three-level verification layers of syntax, semantics, and behavior. Each layer is independent but information-interconnected;
[0079] Feedback-driven: The verification results reverse-train the rule generation model to form a closed-loop optimization;
[0080] Heterogeneous parallelism: Different types of data are assigned to different processing pipelines to break through the bottleneck of a single pipeline.
[0081] Deepen into a multi-dimensional verification architecture with spatio-temporal overlap. The specific innovation points include:
[0082] Dynamic weight allocation: Automatically adjust the verification strictness according to data timeliness;
[0083] Cross-modal association verification: Spatio-temporal alignment verification of text logs and real-time operation videos;
[0084] Quantum rule engine: Encode verification rules into quantum bit states to achieve super-linear rule combination.
[0085] Specific Implementation Architecture
[0086] I. Construction of Heterogeneous Data Pipelines
[0087] Real-time stream pipeline (<50ms latency):
[0088] Allocate a syntax check unit accelerated by FPGA, and use a convolutional neural network (CNN) to detect log format anomalies in real time, processing 100,000 logs per second.
[0089] Batch processing pipeline (>1TB / batch):
[0090] Deploy a distributed semantic analysis cluster, and identify logical contradictions across log files through knowledge graph association verification.
[0091] Hybrid pipeline (dynamic load):
[0092] When the data flow rate exceeds the threshold, automatically route some data to the edge node for preprocessing, and the core node focuses on complex verification.
[0093] II. Quantum quantization rule engine
[0094] Encode traditional rules (such as regular expressions) into quantum states:
[0095]
[0096] Among them, the weights of α and β are dynamically adjusted by the historical verification accuracy;
[0097] Construct quantum entanglement verification pairs:
[0098] Establish an entangled association between the log timestamp field and the NTP time signal of the user operation video. If the verification fails on either side, a cross-modal alarm will be triggered immediately.
[0099] III. Space-time overlapping verification
[0100] Longitudinal verification (time dimension):
[0101] Establish a Markov chain model to predict the probability distribution of the log sequence, and intercept abnormal sequences that deviate from the 3σ interval.
[0102] Transverse verification (space dimension):
[0103] Use blockchain to record the association relationship of cross-server logs, and use zero-knowledge proof to verify distributed consistency.
[0104] IV. Dynamic weight allocation mechanism
[0105] Define the timeliness sensitivity coefficient:
[0106] (t is the data generation time, unit: minute)
[0107] Verification strictness formula:
[0108]
[0109] When S < 0.7, trigger data reprocessing.
[0110] V. Closed-loop Optimization System
[0111] Input the verification result into the reinforcement learning model to generate new quantization rules;
[0112] Perform dimensionality compression analysis every week to eliminate redundant rules with a contribution of < 5%;
[0113] Simulate attack scenarios through a generative adversarial network and continuously train the anomaly detection model.
[0114] Implementation Process
[0115] 1. Data Ingestion and Pipeline Allocation
[0116] When the log file enters the system, it is quickly classified through the leading 128-byte feature vector;
[0117] Automatically allocate processing pipelines according to the data type (such as audit log / performance log) and urgency;
[0118] Generate a unique spatio-temporal encoding (Timestamp⊕ServerIP⊕LogType) for each piece of data.
[0119] 2. Parallel Quantum Verification
[0120] Real-time Streaming Pipeline Execution:
[0121] CNN Syntax Detection (detecting character set anomalies, field missing, etc.), Quantum State Rule Matching (verifying 1000 regular expressions simultaneously);
[0122] Batch Pipeline Execution:
[0123] Knowledge Graph Association Analysis (such as temporal verification of user login logs and permission change logs);
[0124] Byzantine Fault Tolerance Verification for Cross-server Logs.
[0125] 3. Cross-modal Association Verification
[0126] Extract operation videos of the same time period from the video storage system;
[0127] Verify using a multi-modal alignment algorithm:
[0128] The spatio-temporal consistency between the operation commands in the log and the actual operations in the video;
[0129] Association analysis between server performance logs and cabinet infrared thermal imaging.
[0130] 4. Dynamic Decision-making and Routing
[0131] Calculate the comprehensive verification score S;
[0132] Select a processing strategy based on the score:
[0133] S ≥ 0.9: Directly write to Kafka and mark it as trusted data;
[0134] 0.7 ≤ S < 0.9: Write to a temporary queue and wait for manual review;
[0135] S < 0.7: Transfer to the error correction pipeline for re - cleaning.
[0136] 5. Loop optimization iteration
[0137] Collect misjudgment samples (False Positive / Negative) from each pipeline;
[0138] Optimize the rule weight distribution through the quantum annealing algorithm;
[0139] Generate a new set of quantum verification rules every week and verify the effect through A / B testing.
[0140] Step 4, Import user behavior data into the data warehouse
[0141] The user behavior data processed in the previous step is stored in ElasticSearch. To meet the requirements of offline calculation, the data needs to be written to the Hive - based data warehouse through Spark.
[0142] Step 5, Import business system data into the data warehouse
[0143] To solve the problem of incomplete data collection, in addition to user behavior data, the offline calculation data also requires business data precipitated by application systems. Business data is mainly used to provide multi - dimensional statistical solutions and information expansion during offline calculation, such as manuscript information, channel information, activity information, and the association relationship between channel manuscripts, etc. Since business systems have their own independent storage media and data formats and cannot adopt a unified data import solution, it is necessary to customize and develop relevant import programs to extract business data and import it according to the preset format of the data warehouse.
[0144] Step 6, Data analysis and processing
[0145] This step mainly performs offline analysis and calculation on the data in the data warehouse through the Spark computing platform, and finally writes the statistical result data into MySQL. For example, at 1:00 am every day, the daily active user statistics task of the scheduling client is calculated to obtain indicators such as the number of client users, the number of visits, and the stay duration of the previous day. Combining with the session ID identifier of user behavior, the behaviors of users can be concatenated to form the user behavior trajectory, which is different from tagging the interests and attributes of users through rules, so as to operate specific user groups more refinedly and provide more personalized services and information.
[0146] Step 7, Presentation of data analysis results
[0147] The result data statistically calculated in the previous step is stored in MySQL and presented to the operation personnel in the "Data Operation Platform" by means of drawing data trend and distribution charts.
[0148] Embodiment 2
[0149] Refer to Figure 2 , which is the second embodiment of the present invention. A big data processing system based on user behavior information includes:
[0150] A front-end server module 100, a back-end server module 200 and a storage module 300;
[0151] The front-end server module 100 includes a client input unit 101 and a client display unit 102. The client input unit contains an implanted buried point SDK for collecting data information input by users in the client input unit 101. The client display unit 102 displays the user static data, predicted user data and user behavior data trend of the storage module 300;
[0152] The back-end server module 200 includes a load balancer unit 201, a data cleaning unit 202, a data processing unit 203 and a data screening unit 204;
[0153] After the load balancer unit 201 receives the data input by the client input unit 101, after analyzing the load capacity of the data cleaning unit 202, it distributes the user behavior data to the data cleaning unit 202 for cleaning. The data processing unit 203 checks and converts the user behavior data after data cleaning to form conversion data and writes it into the message queue of the storage module 300;
[0154] The data screening unit 204 listens to and calculates the conversion data in the message queue of the storage module 300 through self-calculating factors to form user static data information and user dynamic data information. The data screening unit 204 processes the user dynamic data information to form user data information and prediction data information within a stage, inserts the user data information within the stage into the user static data information in chronological order to form complete user static data information, places the complete user static data information into the first block of the storage module 300, places the prediction data information into the second block of the storage module 300, combines the data warehouse information and the prediction data information to draw the trend of user behavior data, and presents it in the form of a chart in the third block of the storage module 300;
[0155] The stored content of the storage module 300 can be displayed on the client display unit 102.
[0156] The data cleaning method for user data information by the data cleaning unit 202 includes:
[0157] Construct a user behavior knowledge graph and present the user abnormal pattern data in the client input unit 101 for the client to confirm. If abnormal data appears in the client input unit 101 and the client selects to confirm and fills in the input password, submit the user-confirmed behavior data information to the data processing unit 203. If abnormal data appears in the client input unit 101 and the client does not select to confirm or does not fill in the correct password, delete the user abnormal data;
[0158] Constructing a user behavior knowledge graph includes recording user nodes to establish a knowledge graph, and the user nodes include user ID, user device type, and user common functions.
[0159] The stored content of the storage module 300 can be displayed on the client display unit 102, including the first block, the second block, and the third block. The first block, the second block, and the third block are arranged in sequence in the interface display of the client display unit 102.
[0160] Embodiment 3
[0161] The third embodiment of the present invention, a big data processing system based on user behavior information includes:
[0162] I. Test preparation and implementation process
[0163] To verify the innovation of the big data processing method of the present invention, a test environment simulating a real scenario was built. The test platform includes 3 groups of client clusters (10,000 simulated devices each for Android / iOS / Web), 6 data cleaning node servers (Intel Xeon Gold 6338, 128GB RAM), 2 data processing nodes (AMD EPYC 7763, 1TB NVMe SSD), and a distributed storage system (Ceph cluster, with a total capacity of 500TB). The experimental data comes from 50 million behavior logs generated by simulating 100,000 users for 72 hours, including 12 types of events such as client startup, news click, and advertisement browsing.
[0164] The implementation process strictly follows the technical route of the present invention:
[0165] Data collection stage: Customized buried point SDKs were implanted in the simulated clients, and a 1-minute timed reporting mechanism was set. The collected fields include basic metadata such as device ID, timestamp, and event type, and at the same time, the screen coordinate sequence of the operation trajectory was recorded (with an accuracy of 0.1% of the resolution).
[0166] Load balancing optimization: A dynamic weight distribution algorithm was deployed on the backend servers, and the traffic distribution ratio was automatically adjusted according to the real-time CPU occupancy rate of each cleaning node (with a collection frequency of 10Hz). When the node load exceeds 70%, horizontal expansion is triggered, and new nodes can complete containerized deployment within 30 seconds.
[0167] Multi-level data verification: Build a three-level verification pipeline of syntax-semantics-behavior:
[0168] Verify the compliance of the JSON format and the integrity of fields at the syntax layer (18 mandatory verification rules)
[0169] Detect abnormal spatio-temporal sequences at the semantic layer through a pre-trained model (such as a user operating across regions within 0.1 seconds). At the behavior layer, compare the consistency between the client video hash value and the log description (using the SHA-3 algorithm).
[0170] Dynamic data processing: Build a user portrait cube in the data warehouse. Static data (such as device model, registration information, etc.) is stored in columnar format, and dynamic data (such as click stream, dwell time, etc.) is sliced by a 5-minute time window. The prediction model uses an improved LSTM network, and an attention mechanism is introduced to capture mutation points of behavior patterns.
[0171] Visualization feedback: Implement a multi-dimensional data drilling function on the front-end management interface, supporting the generation of heat maps and trend curves according to combined conditions such as user groups, time granularity (1 minute to 24 hours), and event types.
[0172] II. Experimental data table
[0173] Table 1: Comparison of Load Balancing Performance (Response Time Unit: ms)
[0174]
[0175] Table 2: Comparison of Data Verification Effects (Unit: %)
[0176]
[0177] Table 3: Timeliness of Data Processing (Unit: ms)
[0178]
[0179] Table 4: Accuracy of Behavior Prediction (Unit: %)
[0180]
[0181] III. Data Analysis and Conclusions
[0182] Experimental data show that the present invention demonstrates significant advantages in the entire data processing link:
[0183] Improved load balancing efficiency: As shown in Table 1, in the scenario of 100,000 concurrent requests, the dynamic weight algorithm reduces the response time from the timeout state of the traditional scheme to 397 ms (P < 0.01, t-test), while reducing the hardware cost by 15.8%. This benefits from the collaborative effect of the timing mechanism and the load sensor. By real-time collecting the node resource utilization rate (sampling error < 0.5%), sub-second dynamic adjustment of traffic distribution is achieved.
[0184] Breakthrough in data verification accuracy: Table 2 shows that the three-level verification system increases the comprehensive interception rate to 99.9%, an increase of 7.4 percentage points compared to the single-layer verification. Especially in the detection of spatio-temporal contradictions, the combination of semantic layer and behavior layer verification achieves an accuracy rate of 99.1%, effectively identifying 0.03% of maliciously forged data. This verifies the effectiveness of the spatio-temporal coding technology - by quantum coding and comparing the hash value (SHA3-512) of the client operation video with the spatio-temporal description of the log, anti-tampering verification is achieved.
[0185] Optimization of processing timeliness: As shown in the data of Table 3, in the processing of millions of data, the present invention reduces the end-to-end delay from 1832 ms of the existing technology to 1205 ms (a decrease of 34.2%). This is attributed to the application of the superlinear rule combination: when the comprehensive verification score S ≥ 0.9 (accounting for 92.7% of the data), it directly enters the high-speed processing channel to avoid redundant calculations. The parallel parsing ability of spatio-temporal coding increases the cleaning speed by 1.8 times.
[0186] Innovation of the prediction model: Table 4 proves that the improved LSTM+Attention model has an average improvement of 25.4% in key metrics. By analyzing the temporal correlation of user dynamic data, the model can automatically identify behavioral chain patterns such as "evening news click → next-day advertisement conversion". The prediction latency is controlled within 100ms, meeting the real-time recommendation requirements, which benefits from the 5-minute granularity optimization of the dynamic data window (comparative experiments show that the prediction accuracy only increases by 1.2% with a 1-minute granularity but the latency increases by 3 times).
[0187] The above data confirm that the present invention exceeds the prior art solutions in core metrics such as data processing efficiency, system stability, and prediction accuracy, providing a new technical paradigm for large-scale user behavior analysis.
[0188] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the present invention.
Claims
1. A big data processing method based on user behavior information, characterized in that, The method includes: Collecting user behavior data by implanting a data tracking SDK on the client side and uploading it to the backend server via the HTTP protocol. The user behavior data includes the user's click and preview behavior on the client side. After receiving the user behavior data in the HTTP protocol, the load balancer of the backend server balances the pressure on each data cleaning node. The data cleaning node parses the original data of the client and puts the parsed behavior data into a log file. The data processing node forms transformed data after validating and converting the format of the behavior data and writes it into the stored message queue. The data screening node listens to and calculates the transformed data in the message queue and classifies the transformed data into user static data information and user dynamic data information. Put the user static data information into the data warehouse. The data recording node periodically reads the user dynamic data information in real time and puts the recorded user dynamic data information content into the data warehouse in units of recording time. Arrange the user dynamic data information and user static data information recorded in units of recording time in the database in units of time. The data screening node calculates and predicts the user dynamic data information to obtain predicted data information, combines the data warehouse information and the predicted data information to draw the trend of user behavior data, and presents it in the form of a chart on the client side. The formation of the transformed data by the data processing node after validating and converting the format of the behavior data includes: The data format validation includes decomposing the validation of user behavior data into three-level validation layers of syntax, semantics, and behavior. Each layer is independent but the information is interconnected. The verification result is used to reverse-train the rule generation model to form a closed-loop optimization. Different types of data are assigned different processing pipelines, and the qualified user behavior data after verification is converted through the processing pipeline. The reverse-training rule generation model includes automatically adjusting the verification strictness of the syntax verification layer and the semantics verification layer according to data timeliness, verifying the spatio-temporal alignment of the client text log and the real-time operation video, encoding the verification rules into the state of quantum bits, and generating a unique spatio-temporal code for each piece of data to achieve a superlinear rule combination. The superlinear rule combination includes: Calculating the comprehensive verification score S of the unique spatio-temporal code and selecting a processing strategy according to the score: When S≥0.9, the data processing node directly converts the user behavior data and marks it as trusted data. When 0.7≤S<0.9, the data processing node writes the user behavior data into a temporary queue for manual review. When S<0.7, the data processing node transfers the user behavior data to the error correction pipeline for interception. The calculation method of the comprehensive verification score S includes: Where a is the syntax score, b is the semantics score, λ(t) represents the time-sensitive function, the syntax score is obtained by the syntax verification layer, and the semantics score is obtained by the semantics verification layer. The time-sensitive calculation method is: Where t is the data generation time.
2. The big data processing method based on user behavior information according to claim 1, characterized in that: The collection of user behavior data by implanting a data tracking SDK on the client side and uploading it to the backend server via the HTTP protocol includes: The user behavior data includes client startup and exit, and clicks and views of advertisements, news, live broadcasts, and operation activities; After the client collects the user's behavior data, it is temporarily stored locally on the client. Through a pre-designed timing mechanism, the collected behavior data is defined and summarized and reported to the backend server via the HTTP protocol, and is received by the load balancer of the backend server.
3. A big data processing method based on user behavior information according to claim 2, characterized in that: The data cleaning node parses the original data of the client, including: IP address and User-Agent information; The User-Agent information includes browser type information, version information, and operating system information.
4. A big data processing system based on user behavior information, characterized in that, The system includes a front-end server module (100), a back-end server module (200), and a storage module (300); The front-end server module (100) includes a client input unit (101) and a client display unit (102). The client input unit contains an implanted buried point SDK for collecting data information input by the user in the client input unit (101). The client display unit (102) displays the user static data information, prediction data information, and user behavior data trend in the storage module (300); The back-end server module (200) includes a load balancer unit (201), a data cleaning unit (202), a data processing unit (203), and a data screening unit (204); After the load balancer unit (201) receives the data input by the client input unit (101), after analyzing the load capacity of the data cleaning unit (202), it allocates the user behavior data to the data cleaning unit (202) for cleaning. The data processing unit (203) verifies and converts the user behavior data after cleaning to form conversion data, and writes it into the message queue of the storage module (300); The data screening unit (204) listens to and calculates the conversion data in the message queue of the storage module (300) to form user static data information and user dynamic data information. The data screening unit (204) processes the user dynamic data information to form user data information and prediction data information within a stage, inserts the user data information within the stage into the user static data information in chronological order to form complete user static data information, puts the complete user static data information into the first block of the storage module (300), puts the prediction data information into the second block of the storage module (300), combines the data warehouse information and the prediction data information to draw the user behavior data trend, and presents it in the form of a chart in the third block of the storage module (300); The storage content of the storage module (300) can be displayed on the client display unit (102); The data processing unit (203) forms conversion data after verifying and converting the behavior data format, including: The data format verification includes verifying and disassembling the user behavior data into three-level verification layers of syntax, semantics, and behavior, with each layer being independent but information-interconnected; The verification result is used to reverse-train the rule generation model to form a closed-loop optimization; Differentiated processing pipelines for different types of data distribution, and the qualified user behavior data is subjected to data conversion through the processing pipeline; The reverse training rule generation model includes automatically adjusting the verification strictness of the syntax verification layer and the semantic verification layer according to data timeliness, verifying the spatio-temporal alignment of the client text log and the real-time operation video, encoding the verification rules into qubit states, generating a unique spatio-temporal encoding for each piece of data, and realizing a superlinear rule combination; The superlinear rule combination includes: Calculate the comprehensive verification score S of the unique spatio-temporal encoding, and select a processing strategy according to the score: When S≥0.9, the data processing unit (203) directly performs data conversion on the user behavior data and marks it as trusted data; When 0.7≤S<0.9, the data processing unit (203) writes the user behavior data into a temporary queue for manual review; When S<0.7, the data processing unit (203) transfers the user behavior data to an error correction pipeline for interception; The calculation method of the comprehensive verification score S includes: Where a is the syntax score, b is the semantic score, λ(t) identifies the time-sensitive function, the syntax score is obtained by the syntax verification layer, and the semantic score is obtained by the semantic verification layer; The time-sensitive calculation method is: Where t is the data generation time.
5. A big data processing system based on user behavior information according to claim 4, characterized in that: The data cleaning unit (202) cleans the user data information in the following ways: Construct a user behavior knowledge graph and present the user abnormal pattern data in the client input unit (101) for the client to confirm. If the abnormal data appears in the client input unit (101) and the client selects to confirm and fills in the input password, the user-confirmed behavior data information is submitted to the data processing unit (203). If the abnormal data appears in the client input unit (101) and the client does not select to confirm or does not fill in the correct password, the user abnormal data is deleted; The construction of the user behavior knowledge graph includes recording user nodes to establish a knowledge graph, and the user nodes include user ID, user device type, and user common functions.
6. A big data processing system based on user behavior information according to claim 4, characterized in that: The storage content of the storage module (300) can be displayed on the client display unit (102), including a first block, a second block, and a third block, and the first block, the second block, and the third block are arranged in sequence in the interface display of the client display unit (102).
7. An electronic device, including: One or more processors; A storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-3.
8. A computer-readable storage medium having executable instructions stored thereon, which when executed by a processor cause the processor to implement the method according to any one of claims 1-3.
Citation Information
Patent Citations
Fuzzy clustering system based on user behavior data
CN111666351A
Power customer data sampling method based on Kafka
CN115630111A