An intelligent management system for C-end user data based on large models

Through the C-end user data intelligent management system based on the big model, the problem of unstructured data in traditional methods is solved, and the automation and intelligent management of user data is realized, data collection and analysis efficiency is improved, user behavior is deeply understood, and user experience and service quality are enhanced.

CN119862334BActive Publication Date: 2025-07-08SHANDONG TAIYING TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510352420.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-08
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

Traditional data processing methods are difficult to effectively mine unstructured data, resulting in a large amount of valuable user information being ignored, and user behavior analysis is difficult to deeply explore potential patterns and relationships.

Method used

The C-end user data intelligent management system based on large models is adopted, including data acquisition, transmission, generation, analysis, extraction and intelligent management modules. Through technical means such as crawling, block processing, dynamic parameter calculation, user behavior tracking, data correction and feature selection, data automation and intelligent management are realized.

Benefits of technology

It improves data collection efficiency and accuracy, optimizes data processing processes, deeply understands user behavior, improves data analysis and prediction accuracy, and enhances user experience and service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119862334B_ABST
    Figure CN119862334B_ABST
Patent Text Reader

Abstract

The present invention relates to an intelligent management system for C-end user data based on a large model, belonging to the field of big data technology, including: a data transmission module that divides the C-end user data into blocks to calculate the dynamic parameters of each data block; determines the urgency level according to the dynamic parameters of each data block, sorts and transmits the data blocks according to the urgency level; a generation module that is used to receive the data blocks, generates heat data and mines historical behavior data by tracking the user browsing path, analyzing the jump relationship, and identifying the loss nodes, so as to record the operation behavior data of the user on the client side; a data analysis module that identifies and corrects the errors in the user data and operation behavior data, performs feature selection and data reduction, and analyzes the reduced data to obtain the analyzed data. The present invention can efficiently mine the value of user data and realize the automated and intelligent management of data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data technology, and specifically, to an intelligent management system for C-end user data based on a large model. Background Art

[0002] With the rapid development and popularization of the Internet, C-end (consumer-side) user data has shown an explosive growth trend. These data contain valuable information such as various behavioral trajectories, preferences, and demands of users on the Internet, which is of great significance for enterprises to understand market dynamics, optimize product services, and formulate marketing strategies.

[0003] Some traditional data processing methods can only process structured data, and have limited processing capabilities for unstructured data (such as text, images, videos, etc.), which leads to a large amount of valuable information being ignored or wasted. In addition, some traditional user behavior analyses are based on simple statistical methods and rules, making it difficult to deeply explore the potential patterns and relationships of user behavior. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide an intelligent management system for C-end user data based on a large model, which can efficiently mine the value of user data and realize the automated and intelligent management of data.

[0005] The basic concept of the technical solution adopted by the present invention to solve the above technical problem is:

[0006] An intelligent management system for C-end user data based on a large model, comprising:

[0007] A data acquisition module, starting from an initial URL, obtaining new URLs by requesting and parsing target site data, de-duplicating and filtering, and putting them into a queue for cyclic crawling to obtain C-end user data;

[0008] A data transfer module, which divides the C-end user data into blocks to calculate the dynamic parameters of each data block; determines the urgency level according to the dynamic parameters of each data block, and sorts and transmits the data blocks according to the urgency level;

[0009] A generation module, which is used to receive data blocks, generate heat data and mine historical behavior data by tracking the user browsing path, analyzing jump relationships, and identifying lost nodes, so as to record the operation behavior data of the user on the client side;

[0010] A data analysis module, which identifies and corrects errors in user data and operation behavior data, performs feature selection and data reduction, and analyzes the reduced data to obtain the analyzed data;

[0011] The extraction module determines the data classification task, extracts the user data information of each client, including user basic information, behavior and transaction data, and classifies it to obtain the classified user data;

[0012] The training module integrates the analyzed data and the classified user data, trains and infers the integrated data, and obtains the patterns and relationships between the data;

[0013] The intelligent management module predicts user behavior based on the patterns and relationships between the data to achieve intelligent management of C-end user data.

[0014] Furthermore, starting from the initial URL, by requesting and parsing the data of the target site, after deduplication and filtering, new URLs are obtained and put into a queue for cyclic crawling to obtain C-end user data, including:

[0015] Determine the crawling target of the C-end user data to be crawled, and the crawling target includes the corresponding website, forum or social media platform;

[0016] According to the crawling target, obtain the web page link as the starting point of the crawler, that is, the initial URL;

[0017] Use the HTTP library to initiate a network request to the target site, obtain the web page content, and parse the web page through the parsing library to extract the required data;

[0018] Perform deduplication and filtering operations on the required data to extract new URL links;

[0019] Put the new URL links into a queue, repeat the extraction process, continuously take out URLs from the queue for crawling until the preset crawling depth is reached to obtain C-end user data.

[0020] Furthermore, the C-end user data is chunked to calculate the dynamic parameters of each data chunk; according to the dynamic parameters of each data chunk, determine the urgency level, and sort and transmit the data chunks according to the urgency level, including:

[0021] Calculate the average value of the C-end user data to obtain a center point; calculate the standard deviation according to the C-end user data;

[0022] Determine a corresponding radius based on the standard deviation to define a circle;

[0023] Traverse each data point, calculate its Euclidean distance from the center point. If the distance ≤ radius, the data point is regarded as an "inner point of the circle"; if the distance > radius, the data point is "projected" onto the circle by a mapping method to become an "on-the-circle point"; select the "inner points of the circle" and "on-the-circle points" as the dataset to be processed;

[0024] Divide the processed data set into multiple data blocks using an equidistant segmentation algorithm according to the distribution and quantity of the data;

[0025] For each data block, calculate the dynamic parameter;

[0026] Analyze the dynamic parameters to obtain the urgency index corresponding to each data block;

[0027] According to the urgency index, sort all data blocks from high to low, and transmit the data blocks in the sorted order.

[0028] Furthermore, the calculation formula for the dynamic parameter is:

[0029] ;

[0030] Wherein, is the dynamic parameter, is the weight coefficient of the real-time factor, is the base of the natural logarithm, is the decay rate, is the current time, is the creation time of the data block, is the weight coefficient, with a value ranging from 0.1 to 10; is the number of data blocks, is the t th data block coefficient, is the t th frequency of the data block appearing in user behavior, is the data size weight coefficient, is the maximum number of data block bytes, is the number of data block bytes, is the index.

[0031] Furthermore, analyze the dynamic parameters to obtain the urgency index corresponding to each data block, including:

[0032] Analyze the dynamic parameters to determine the range of the parameters;

[0033] According to the range of the parameters, map different value ranges of the dynamic parameters to different urgency levels, and represent the mapping rules through the structure of a decision tree, wherein each node represents a judgment condition of the dynamic parameter, the branches represent different value ranges, and the leaf nodes represent the results of the urgency;

[0034] For each data block, extract its dynamic parameter, start from the root node, divide the data set into subsets according to the dynamic parameter, and create child nodes. Evaluate each node. If the performance improvement after division is not obvious, stop dividing and mark the current node as a leaf node;

[0035] Recursively repeat the above process on each child node until the stopping condition is met, and traverse each non-leaf node from bottom to top to evaluate its performance after being replaced by a leaf node;

[0036] According to the performance evaluation results after the leaf nodes, mark each leaf node as the corresponding urgency category to obtain the urgency index corresponding to each data block.

[0037] Furthermore, by tracking the user's browsing path, analyzing the jump relationship, and identifying the lost nodes, generate heat data and mine historical behavior data to record the user's operation behavior data on the client side, including:

[0038] By tracking the user's browsing path within the client, construct the user's navigation map, analyze the jump relationship between different pages, and identify the key nodes where the user is lost to obtain the path analysis result;

[0039] Generate heat data for the page based on the path analysis result and operation behavior to obtain the user's historical behavior data;

[0040] Mine and analyze the user's historical behavior data, track events, and record the user's operation behavior data on the client side.

[0041] Furthermore, identify and correct errors in the user data and operation behavior data, perform feature selection and data reduction, and analyze the reduced data to obtain the analyzed data, including:

[0042] Identify the incorrect data in the user data and operation behavior data, and correct the incorrect data to obtain the corrected data;

[0043] Perform data selection on the corrected data, standardize the selected feature data to obtain the standardized data matrix;

[0044] For the standardized data matrix, calculate the covariance matrix, and perform eigenvalue decomposition on the covariance matrix to obtain the eigenvalues;

[0045] Sort the eigenvalues from largest to smallest, and k take the eigenvectors corresponding to the largest

[0046] Project the standardized feature matrix of the original data onto the principal component matrix to obtain the reduced low-dimensional data matrix;

[0047] Perform transformation analysis on the reduced data to obtain the analyzed data.

[0048] Further, determine the data classification task, extract the user data information of each client, including basic user information, behavior and transaction data, and classify them to obtain classified user data, including:

[0049] Determine the task plan for data classification, divide the tasks into different clients, and extract the user data information of the client, including basic user information, behavior data and transaction data. The basic user information includes user ID, name, gender, age, address and contact information;

[0050] Set classification standards and rules, classify the user data information data of the client, and obtain the classified user data.

[0051] Furthermore, the analyzed data and classified user data are integrated, and training and reasoning are performed on the integrated data to obtain patterns and relationships between the data, including:

[0052] Integrate the analyzed data and classified user data to obtain a unified data set;

[0053] Using the data set, training the deep learning network to obtain a trained deep learning network;

[0054] New user data is predicted and analyzed through the trained deep learning network to discover potential patterns and relationships in the data.

[0055] Furthermore, user behavior prediction is performed based on the patterns and relationships between data to achieve intelligent management of C-end user data, including:

[0056] Build user profiles based on patterns and relationships between data;

[0057] Calculate the similarity between users based on user portraits and historical behavior data;

[0058] Based on the similarity, for the user behavior data with time sequence, time series analysis is performed to predict the user's future behavior trend;

[0059] Based on the similarity, for discrete user behavior data, analyze the user's historical behavior data and predict the user's future behavior trend through the recommendation of user portraits;

[0060] Based on the predicted future behavior trends of users, intelligent management of C-end user data can be achieved.

[0061] After adopting the above technical scheme, the present invention has the following beneficial effects compared with the prior art.

[0062] Automation scripts reduce the time and labor costs of manual data collection, improve data collection efficiency, ensure data uniqueness and accuracy through deduplication and filtering, avoid duplicate processing, use queues for cyclic crawling to achieve continuous and efficient data collection, and ensure the timeliness and comprehensiveness of data. Process C-end user data in chunks, optimize the data processing flow, improve processing efficiency, determine the urgency level by calculating dynamic parameters to achieve data prioritization, ensure that important data is processed in a timely manner, improve the efficiency and response speed of data transmission, and meet real-time requirements. Track the user browsing path and analyze the jump relationship, deeply understand user behavior, accurately identify the key nodes of user churn, generate heat map data and mine historical behavior data, and record the operation behavior data of users on the client side. Identify and correct data errors, improve data quality, ensure the accuracy of analysis results, perform feature selection and data reduction to improve analysis efficiency, deeply analyze the reduced data, and reveal potential trends and patterns.

[0063] Define clear data classification tasks to make data processing more focused and improve processing efficiency. Extract comprehensive user data information, including user basic information, behavior, and transaction data. Classify user data to make the data more orderly and clear, and improve the usability and value of the data. Integrate the analyzed data and the classified user data to form a more complete data set, improve data utilization, use high-quality data to train a deep learning network, enhance the performance and accuracy of the model, and through training and inference, discover the deep patterns and correlation relationships between data, providing a strong data foundation for prediction and decision support. Build a refined user profile, predict the future behavior trends of users through time series analysis and user profile recommendations, and improve prediction accuracy. Based on the predicted future behavior trends of users, formulate intelligent management strategies for C-end users to enhance the user experience and service quality. Brief Description of the Drawings

[0064] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0065] Figure 1 It is a schematic diagram of the C-end user data intelligent management system based on the large model of the present invention. Detailed Embodiments

[0066] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0067] In the following embodiments of this application, a case of an intelligent management system for C-end user data based on a large model is used to illustrate the solution of this application in detail, but this embodiment does not limit the protection scope of this application.

[0068] As Figure 1 shown, the present invention provides an intelligent management system for C-end user data based on a large model, including:

[0069] A data acquisition module 11, starting from the initial URL, obtains new URLs by requesting and parsing the target site data, de-duplicates and filters them, and puts them into a queue for cyclic crawling to obtain C-end user data;

[0070] A data transfer module 12 divides the C-end user data into blocks to calculate the dynamic parameters of each data block; determines the urgency according to the dynamic parameters of each data block, and sorts and transmits the data blocks according to the urgency;

[0071] A generation module 13 is used to receive data blocks, generate heat data and mine historical behavior data by tracking the user browsing path, analyzing jump relationships and identifying churn nodes, so as to record the operation behavior data of users on the client side;

[0072] A data analysis module 14 identifies and corrects errors in user data and operation behavior data, performs feature selection and data reduction, and analyzes the reduced data to obtain the analyzed data;

[0073] An extraction module 15 determines data classification tasks, extracts user data information of each client, including user basic information, behavior and transaction data, and classifies them to obtain classified user data;

[0074] A training module 16 integrates the analyzed data and the classified user data, trains and infers the integrated data, and obtains the patterns and relationships between the data;

[0075] An intelligent management module 17 predicts user behavior according to the patterns and relationships between the data to achieve intelligent management of C-end user data.

[0076] In a specific embodiment of the present invention, data is automatically crawled from a target site to improve data collection efficiency. Through deduplication and filtering, the uniqueness and accuracy of the data are ensured, and duplicate processing is avoided. A queue is used for cyclic crawling to achieve continuous and efficient data collection. The C-end user data is processed in chunks to optimize the data processing flow. By calculating dynamic parameters to determine the urgency, the priority sorting of the data is realized, ensuring that important data is processed in a timely manner, and the efficiency and response speed of data transmission are improved. The user browsing path is traced and the jump relationship is analyzed to deeply understand user behavior. By generating heat map data and mining historical behavior data, user preferences and usage habits are revealed, and the operation behavior data of users on the client side is recorded. Data errors are identified and corrected to improve data quality. Feature selection and data reduction are performed to simplify the data model and improve analysis efficiency. The reduced data is deeply analyzed to reveal potential trends and patterns. Clear data classification tasks are determined to make data processing more focused. Comprehensive user data information, including user basic information, behavior, and transaction data, is extracted to lay a foundation for in-depth analysis, and the user data is classified. The analyzed data and the classified user data are integrated to form a more complete data set. Through training and inference, deep patterns and correlation relationships between the data are discovered, providing a strong data foundation for prediction and decision support. The patterns and relationships between the data are used to predict user behavior, enhance market insight ability, realize intelligent management of C-end user data, and improve user experience and service quality.

[0077] In a preferred embodiment of the present invention, starting from the initial URL, the data of the target site is requested and parsed. After deduplication and filtering, new URLs are obtained and put into the queue for cyclic crawling to obtain C-end user data, including:

[0078] Determine the crawling target of the C-end user data to be crawled. The crawling target includes the corresponding website, forum, or social media platform, specifically including: determining the type of C-end user data to be crawled, such as user comments, product evaluations, user behavior data, etc. Analyze and select one or more suitable websites, forums, or social media platforms as data sources, evaluate the data quality, update frequency, and ease of crawling of these platforms, and conduct a preliminary analysis of the selected website to understand its web page structure, data layout, and existing anti-crawling mechanisms.

[0079] According to the crawling target, obtain the web page link that serves as the starting point of the crawler, that is, the initial URL, specifically including: find the URL of the website area or page containing the required data, usually the home page, category page, or search result page of the website, ensure that the obtained URL is accessible and actually contains the required data, and prepare multiple initial URLs to expand the scope of data crawling.

[0080] Use an HTTP library to initiate a network request to the target site, obtain the web page content, and parse the web page through a parsing library to extract the required data. Specifically, it includes: using the HTTP library to set up and send a network request, configuring the request headers to simulate browser behavior to avoid being recognized as a crawler by the target website. After sending the request, receive the server's response and obtain the HTML content of the web page. Use the parsing library to parse the HTML, locate and extract the required data, and write selectors according to the website structure to accurately capture the target data.

[0081] Perform deduplication and filtering operations on the required data to extract new URL links. Specifically, it includes: using a data structure to store the identifiers of the already crawled data. Before each new data crawl, check whether it already exists in the crawled dataset, and filter out invalid, duplicate, or irrelevant data according to preset rules. Regular expressions, conditional judgments, etc. can be used for data cleaning.

[0082] Put the new URL links into a queue, repeat the extraction process, continuously take out URLs from the queue for crawling until the preset crawling depth is reached to obtain C-end user data. Specifically, it includes: creating an empty queue to store the URLs to be crawled, adding the initial URL to the queue, and taking out a URL from the queue for crawling: taking out a URL from the queue and performing data crawling; during the crawling process, add the newly discovered valid URLs to the queue, ensuring that the newly added URLs meet the crawling conditions and have not been crawled before. Repeat the above process until the queue is empty or the preset crawling depth is reached. Set appropriate stop conditions, such as crawling depth, crawling time, and the amount of crawled data; store the crawled valid data in a database, file, or other storage media.

[0083] In a specific embodiment of the present invention, through an automated script, the time and labor costs of manual data collection can be reduced, the scope and source of data collection can be clearly defined, making data collection more focused and efficient. Select data sources targeted, avoiding wasting time and resources on unnecessary or low-quality data sources, providing a clear starting point for data crawling, and ensuring the controllability of the crawling process; using an HTTP library and a parsing library can efficiently obtain and parse web page content, improving the speed of data collection. Through a professional parsing library, the key information in the web page can be more accurately extracted, removing duplicate and irrelevant data, ensuring the quality and accuracy of the collected C-end user data; reducing waste of storage space and only saving valuable data. By continuously taking out URLs from the queue for crawling, the data of the target site can be deeply mined to obtain more comprehensive C-end user data. The automated loop crawling process reduces the need for manual intervention, improving the efficiency and sustainability of data collection.

[0084] In a preferred embodiment of the present invention, the C-end user data is chunked to calculate the dynamic parameters of each data chunk; based on the dynamic parameters of each data chunk, the urgency level is determined, and the data chunks are sorted according to the urgency level and transmitted, including:

[0085] Calculate the average value of the C-end user data to obtain a central point; calculate the standard deviation according to the C-end user data, specifically including: obtaining the C-end user data, calculating the average value for each dimension in the data set respectively, and these average values will form a central point, representing the "average user" of the data set, calculating the standard deviation for each dimension in the data set, and the standard deviation is an important indicator to measure the degree of dispersion of the data distribution.

[0086] Based on the standard deviation, determine a corresponding radius to define a circle, specifically including: based on the standard deviation calculated above, select a suitable multiple (such as 1 times, 2 times the standard deviation, etc.) as the radius, and this radius will be used to define a circle, and the data points within the circle will be regarded as data points similar or close to the central point, using the calculated central point as the center of the circle and drawing a circle with the determined radius.

[0087] Traverse each data point, calculate its Euclidean distance from the central point, if the distance ≤ radius, the data point is regarded as an "inner-circle point"; if the distance > radius, project the data point "onto" the circle through a mapping method to become an "on-circle point"; select the "inner-circle points" and "on-circle points" as the data set to be processed, specifically including: traverse each data point in the C-end user data, for each data point, calculate its Euclidean distance from the central point, if the Euclidean distance between the data point and the central point is less than or equal to the radius, then regard the data point as an "inner-circle point", if the Euclidean distance between the data point and the central point is greater than the radius, then project the data point "onto" the circle through a mapping method to become an "on-circle point", and combine the "inner-circle points" and "on-circle points" to form a new data set to be processed.

[0088] According to the distribution and quantity of the data, use the equidistant segmentation algorithm to divide the processed data set into multiple data chunks, specifically including: perform a distribution analysis on the data in the processed data set, understand the dense areas and sparse areas of the data, and according to the distribution and quantity of the data, use the equidistant segmentation algorithm to divide the processed data set into multiple data chunks, and the equidistant segmentation algorithm ensures that each data chunk contains a similar number of data points and the intervals between the data chunks are relatively uniform.

[0089] For each data chunk, calculate the dynamic parameter;

[0090] Analyze the dynamic parameters to obtain the urgency level index corresponding to each data chunk;

[0091] According to the urgency level index, sort all the data chunks from high to low, and transmit the data chunks in the sorted order in turn.

[0092] In a specific embodiment of the present invention, by calculating the standard deviation and determining a corresponding radius, the process actually performs a kind of normalization processing on the data, which can make the data points more concentrated. By mapping the data points to "points inside the circle" and "points on the circle", the influence of extreme values can be removed, making the data set more robust. By splitting the processed data set into multiple data blocks, these data blocks can be processed in parallel, thereby improving the efficiency of data processing. Calculating dynamic parameters for each data block can more flexibly reflect the characteristics and changing trends of the data to adapt to different data distributions and quantities, thus providing more accurate data analysis results. By analyzing the dynamic parameters, an emergency degree index corresponding to each data block is obtained, and this index can be used as the basis for prioritizing data processing, improving the pertinence and efficiency of data processing. Sorting all data blocks according to the emergency degree index and transmitting the data blocks in the sorted order can ensure that important data blocks can be processed and transmitted first.

[0093] In a preferred embodiment of the present invention, the calculation formula of the dynamic parameter is:

[0094] ;

[0095] Wherein, is the dynamic parameter, is the weight coefficient of the real-time factor, is the base of the natural logarithm, is the decay rate, is the current time, is the creation time of the data block, is the weight coefficient, with a value ranging from 0.1 to 10; is the number of data blocks, is the t th data block coefficient, is the t th frequency of the user behavior in the data block, is the data size weight coefficient, is the maximum number of bytes of the data block, is the number of bytes of the data block, is the index.

[0096] In a specific embodiment of the present invention, by introducing a weight coefficient for real-time factors, new generated data can be better processed. By using the natural logarithm and decay rate to adjust the importance of data, the priority of old data can be gradually reduced. The adjustability of the weight coefficient and decay rate enables customization according to the specific requirements of different application scenarios. By considering the number, size, and occurrence frequency of data blocks in user behavior, processing resources can be more effectively allocated, which helps to avoid processing bottlenecks and ensures that frequently occurring or larger data blocks receive appropriate attention. By comprehensively considering multiple dynamic parameters, the priority of data processing can be more intelligently determined, thereby improving the overall processing efficiency. This can not only reduce processing latency but also contribute to an increase in throughput and response speed. By considering the data size weight coefficient and the number of bytes of data blocks, storage and transmission resources can be more reasonably allocated, which helps to avoid resource waste and ensures that critical data receives sufficient resource support.

[0097] In a preferred embodiment of the present invention, the dynamic parameters are analyzed to obtain an urgency index corresponding to each data block, including:

[0098] The dynamic parameters are analyzed to determine the range of the parameters, specifically including: collecting a historical data set containing dynamic parameters, calculating the average value of the dynamic parameter values in the entire historical data set to obtain the mean value of the dynamic parameter; by observing the distribution of the dynamic parameter values in the data set, the change range and central tendency of the parameter can be initially understood; by calculating statistical quantities such as the maximum value, minimum value, and standard deviation of the dynamic parameter, its change range can be analyzed; combining the statistical analysis results, a reasonable range of the dynamic parameter can be determined. For example, the upper and lower limits of the range can be set as the mean value plus or minus the standard deviation. Through the above steps, based on the historical data set and the mean value method, the distribution and change range of the dynamic parameter can be effectively analyzed, and its reasonable range can be determined accordingly.

[0099] According to the range of parameters, map different value ranges of dynamic parameters to different levels of urgency, and represent the mapping rules through the structure of a decision tree. Among them, each node represents a judgment condition of a dynamic parameter, the branches represent different value ranges, and the leaf nodes represent the results of the urgency level. Specifically, it includes: determining the number of urgency levels to be divided, for example, three levels: low, medium, and high; setting the level names and identifiers: set a definite name and identifier for each level, such as "low" corresponding to identifier 1, "medium" corresponding to identifier 2, and "high" corresponding to identifier 3; conduct statistical analysis on the historical data of dynamic parameters collected, including minimum value, maximum value, average value, standard deviation, etc.; based on the results of statistical analysis, divide the value range of dynamic parameters into several intervals, and each interval corresponds to an urgency level; determine the urgency level corresponding to each value range of dynamic parameters, for example, the parameter value in the range of 0 - 30 corresponds to "low", 31 - 70 corresponds to "medium", and 71 - 100 corresponds to "high"; record the mapping rules in detail in the document, including the name of the dynamic parameter, the value range, and the corresponding urgency level; select a key dynamic parameter as the root node of the decision tree, and this parameter should have a significant impact on the urgency level; according to the mapping rules, add branches to the value range of each dynamic parameter, and create leaf nodes at the end of the branches; mark each leaf node with the result of the corresponding urgency level.

[0100] For each data block, extract its dynamic parameters. Starting from the root node, divide the data set into subsets according to the dynamic parameters, and create child nodes. Evaluate each node. If the performance improvement after division is not obvious, stop the division and mark the current node as a leaf node. Specifically, it includes: for each data block, extract its dynamic parameter value, divide the data set into different subsets according to the judgment condition of the root node, create a child node for each divided subset, evaluate the performance after division. If the performance improvement is not obvious, stop the division and mark the current node as a leaf node. Among them, the specific calculation formula for evaluating the performance after division is:

[0101] ;

[0102] Among them, is the probability of the i th class sample in subset D; m is the total number of classes; is the weight of subset D i , representing the proportion of its sample number to the sample number of the entire data set; represents the index of the subset obtained after dividing the data set D by a certain feature A; represents the class index in subset D i ; The total number of subsets into which the representative feature A divides the data set D; Indicates the weighted Gini impurity.

[0103] Recursively repeat the above process on each child node until the stopping condition is met, and traverse each non-leaf node from bottom to top to evaluate its performance after being replaced by a leaf node, specifically including: for each newly created child node, repeat the above process, that is, continue to divide the data set according to the dynamic parameter and create deeper child nodes; set the stopping condition for recursive division, such as reaching a predetermined tree depth; when the stopping condition is met, terminate the recursive process.

[0104] According to the performance evaluation results after the leaf nodes, mark each leaf node with the corresponding urgency level category to obtain the urgency level index corresponding to each data block, specifically including: starting from the bottom of the decision tree, traverse each non-leaf node upward to evaluate the performance after replacing the current non-leaf node with a leaf node (such as using methods like cross-validation); if the performance improves after replacement or meets other optimization criteria, replace the non-leaf node with a leaf node and mark the leaf node with the most common urgency level category in the current subset, ensuring that all leaf nodes are marked with the corresponding urgency level category.

[0105] In a specific embodiment of the present invention, through the structure of the decision tree, the relationship between different dynamic parameter values and the urgency level can be intuitively seen, automating the process of judging the urgency level, reducing the need for manual intervention, and thus improving the processing efficiency. Since the construction of the decision tree is based on the range and values of the dynamic parameters, when the distribution or range of the dynamic parameters changes, the decision tree can be adjusted accordingly to adapt to the new data situation. By recursively constructing and evaluating the decision tree, it can be ensured that better performance is achieved when dividing data blocks, which helps to more accurately judge the urgency level of data blocks and avoid misjudgment or missed judgment. According to the urgency level index of the data blocks, processing resources can be allocated more reasonably. The structure of the decision tree is convenient for expansion and maintenance. When new dynamic parameters need to be introduced or existing parameters need to be adjusted, the decision tree can be updated conveniently without large-scale changes to the whole. The urgency level index can provide valuable information for decision-makers and help them respond faster.

[0106] In a preferred embodiment of the present invention, receive a data block, generate heat data and mine historical behavior data by tracking the user browsing path, analyzing the jump relationship, and identifying the loss nodes to record the operation behavior data of the user on the client side, including:

[0107] By tracking the browsing paths of users within the client, constructing the user's navigation graph, analyzing the jump relationships between different pages, and identifying the key nodes of user churn, the path analysis results can be obtained. Specifically, it includes: Using front-end buried-point technology, such as JavaScript tags or SDKs, to capture user behavior data such as clicks, browsing, and scrolling within the client, ensuring that data collection complies with privacy policies and obtains user consent; Through the collected data, draw the jump paths of users between different pages, and use graph databases or visualization tools to present the user's navigation paths to form a graph; Analyze the flow of users between each page, identify which pages are frequently visited by users and which jump paths are often taken by users; Use data analysis tools, such as pivot tables or data analysis software, to count and analyze the jump data; By analyzing indicators such as the stay time and bounce rate of users on different pages, identify the key nodes of user churn, and use tools such as the funnel model to analyze in which links the most users are lost, so as to locate the problem.

[0108] According to the path analysis results and operation behaviors, generate the heat data of the page and obtain the historical behavior data of the user. Specifically, it includes: Through front-end buried-point technology, capture and record the complete access path of the user on the page, including the order of visited pages, stay time, etc.; At the same time, collect the specific operation behaviors of the user on each page, such as clicks, scrolls, hovers, etc.; Integrate the collected path data and operation behavior data to ensure the relevance and timeliness between the data, and map the user's operation behaviors to specific page elements or regions; According to operation behavior indicators such as the click frequency, scroll depth, and stay time of the user in each region of the page, calculate the heat value of each region. The heat value can be a comprehensive indicator reflecting the user's attention degree and interaction activity in this region; The specific calculation formula for the heat value of each region includes:

[0109] Count the number of clicks of users in each region, which can be directly used as part of the heat value of this region; Accumulate the stay time of each user in each region and then calculate the average value to reflect the average attention duration of the user in this region. This average duration can be used as a weighted factor for the heat value; According to the scroll distance or scroll time of the user within the region to judge their attention degree to this region, the scroll depth can be converted into a specific value and added to the calculation of the heat value; Perform weighted summation on indicators such as the number of clicks, stay time, and scroll depth to obtain the comprehensive heat value of each region.

[0110] Mining and analyzing the historical behavior data of users, tracking events, and recording the operation behavior data of users on the client side, specifically including: continuously collecting and saving all historical behavior data of users through methods such as client log records and database storage. The data includes page browsing records, button click events, form filling behaviors, etc.; cleaning the collected historical behavior data to remove invalid, duplicate, or incorrect data records, and performing data preprocessing such as formatting and standardization to ensure data quality and consistency; tracking events for the key operation behaviors of users, such as purchases, registrations, adding items to the shopping cart, etc., assigning a unique identifier to each event, and recording relevant information such as the specific time of the event and the user ID; using data mining techniques such as association rule mining and sequence pattern mining to analyze the correlation and sequential patterns between user behaviors, discover user preferences, habits, and potential needs, provide data support for personalized recommendations and marketing strategies, and generate a user operation behavior data report based on the mining and analysis results.

[0111] In a specific embodiment of the present invention, by constructing a navigation map of users, the browsing path and jump relationship of users within the client can be intuitively understood. By analyzing the jump relationship between different pages of users, the key nodes leading to user loss can be accurately identified, providing a strong basis for optimizing the user experience and reducing user loss. The path analysis results can provide data support for product optimization, page design, content recommendation, etc., helping to improve user satisfaction and retention rate. By generating heat data of pages, the click, browsing and other operation behaviors of users on the page can be intuitively displayed, which helps to discover the hot areas of the page and the focus of user attention. Combining the path analysis results and operation behaviors, more comprehensive historical behavior data of users can be obtained. The heat data can provide feedback for page design. By mining and analyzing the historical behavior data of users, the laws and trends of user behaviors can be revealed. Tracking and recording the operation behavior data of users on the client side can ensure the integrity and accuracy of the data.

[0112] In a preferred embodiment of the present invention, identifying and correcting errors in user data and operation behavior data, performing feature selection and data reduction, and analyzing the reduced data to obtain the analyzed data, including:

[0113] Identify the error data in the user data and operation behavior data, and correct the error data to obtain the corrected data, which specifically includes: reviewing the user data and operation behavior data item by item to check for problems such as missing values, outliers, duplicate values, or format errors, and using statistical methods (such as box plots, IQR rules, etc.) to identify outliers; for missing values, select a filling method according to the data distribution, such as mean filling, median filling, or using interpolation methods; for outliers, decide whether to delete or replace them with reasonable values according to the data characteristics, correct the format errors, and ensure that all data conforms to the expected format standard.

[0114] Perform data selection on the corrected data, and perform standardization processing on the selected feature data to obtain a standardized data matrix, which specifically includes:

[0115] Analyze the user data and operation behavior data, select features highly relevant to the business goal, and use feature selection techniques (such as analysis of variance, correlation analysis, etc.) to remove redundant or irrelevant features; perform standardization processing on the selected features, that is, scale each feature value so that its mean is 0 and its standard deviation is 1.

[0116] For the standardized data matrix, calculate the covariance matrix and perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues; sort the eigenvalues from largest to smallest, and k Take the eigenvectors corresponding to the k largest eigenvalues as the principal components to obtain the principal component matrix, which specifically includes: determining the standardized data matrix, where each column represents a feature and each row represents a sample, and the standardized data ensures that the mean of each feature is 0 and the standard deviation is; for each pair of features (columns) in the standardized data matrix, calculate their covariance; fill the covariance values between all pairs of features into a matrix to form a square matrix, and the diagonal elements of the matrix are the covariance of each feature with itself, that is, the variance, which is 1 in the standardized data (because the standard deviation is 1); use eigenvalue decomposition methods in linear algebra (such as QR algorithm, Jacobi iteration, etc.) to decompose the covariance matrix, and after decomposition, two sets of outputs are obtained: one set is the eigenvalues (scalars), which represent the "importance" of the eigenvectors corresponding to the covariance matrix; the other set is the corresponding eigenvectors (vectors), which represent the main change directions of the data, sort the eigenvalues from largest to smallest, and the size of the eigenvalues reflects the amount of data variability represented by the corresponding eigenvectors; larger eigenvalues mean that the corresponding eigenvectors explain more data changes; select the first k largest eigenvalues and their corresponding eigenvectors as needed, and these eigenvectors constitute the "principal components" of the principal component analysis, and the selection of k can be based on the cumulative contribution rate of the eigenvalues (that is, the sum of the first k eigenvalues accounts for the proportion of the sum of all eigenvalues) or a preset threshold.

[0117] Project the standardized feature matrix of the original data onto the principal component matrix to obtain the low-dimensional data matrix after dimensionality reduction, which specifically includes: Select the eigenvectors corresponding to the top k largest eigenvalues through eigenvalue decomposition, and form the principal component matrix by arranging these eigenvectors column by column, denoted as W; Use matrix multiplication to project the standardized feature matrix X onto the principal component matrix W, and the projection formula is: Y = XW, where Y is the low-dimensional data matrix after dimensionality reduction. Each row in Y still corresponds to a sample, but the number of columns is now reduced to k, that is, the number of selected principal components. In this way, the original high-dimensional data is projected into a low-dimensional space while retaining the most important change directions and information.

[0118] Perform transformation analysis on the reduced data to obtain the analyzed data, which specifically includes: Apply the selected transformation method to each element in the reduced low-dimensional data matrix. For example, if logarithmic transformation is selected, take the natural logarithm of each non-zero element in the data matrix; Perform descriptive statistical analysis on the transformed data, including calculating the mean, median, standard deviation, quartiles, etc.; Understand the central tendency of the data through the mean and median, understand the dispersion degree of the data through the standard deviation, and understand the distribution form and possible outliers of the data through the quartiles.

[0119] In a specific embodiment of the present invention, by identifying and correcting errors in user data and operation behavior data, the data quality is significantly improved. The data is selected, and the selected features are standardized, which helps to eliminate the dimensionality differences between different features. The dimensionality of the data is reduced through principal component analysis, which can not only reduce the computational complexity, improve the processing speed, but also help to eliminate redundant information in the data, making the analysis results clearer and more intuitive. The principal components extracted by PCA often represent the main change trends and key information in the data. After the data processed through the above process, the analysis efficiency and accuracy will be significantly improved; Performing transformation analysis on the reduced data provides more flexibility and possibilities.

[0120] In a preferred embodiment of the present invention, determine the data classification task, extract the user data information of each client, including user basic information, behavior and transaction data, and perform classification to obtain the classified user data, including:

[0121] Determine the purpose of data classification, such as customer segmentation, behavioral analysis, risk assessment, etc.; Decompose the overall task into several subtasks, such as data collection, data preprocessing, feature extraction, classification model selection and evaluation, etc., set time nodes for each subtask, and formulate a project schedule; Determine the different clients to be processed, such as different business departments, product lines or geographical regions, and ensure that the data of different clients can be accessed and processed independently; Identify the data sources of each client, including databases, APIs, CSV files, etc.; Use ETL (Extract, Transform, Load) tools or SQL queries to extract data from the data sources; Extract the basic user information of users, including user ID, name, gender, age, address and contact information; Process missing values, duplicate values and outliers to ensure data quality. Extract the behavioral data of users, such as the number of logins, page views, click-through rates, search records, etc., analyze the time series characteristics of the behavioral data, and understand the patterns and trends of user activities. Extract the transaction data of users, such as purchase records, transaction amounts, payment methods, transaction times, etc., and calculate indicators such as the total transaction amount, average transaction amount, and transaction frequency of users.

[0122] Select features related to the classification objective, such as user value, activity level, risk level, etc.; Create new features, such as customer lifetime value (CLV), recency, etc.; Define classification rules according to business requirements, such as classifying users into high-value users, ordinary users and low-value users, and set the classification threshold. For example, users with an annual transaction amount exceeding 10,000 yuan are defined as high-value users.

[0123] Analyze each feature in the dataset to determine which features are closely related to the classification objective. For example, for user classification, the possibly relevant features may include age, gender, purchase history, transaction amount, etc.; Define the classification objective, such as classifying users into high-value users, ordinary users and low-value users, and convert the classification objective into a numerical form so that the model can process it. For example, represent the user value level with numbers.

[0124] Based on the data features and classification objective, select a decision tree model, divide the dataset into a training set and a test set, use 80% of the data for training and 20% for testing to ensure the independence of model evaluation; Set a random seed to ensure the repeatability of the data division and model training results, instantiate the decision tree classification model in the selected tool, use the training dataset to train the model, and adjust the model parameters to fit the data.

[0125] Adjust the hyperparameters of the decision tree (such as maximum depth, minimum number of samples for splitting) through methods like grid search or random search to optimize the model performance. Use cross - validation to evaluate the model performance under different parameter combinations during the parameter adjustment process, and select the parameter settings with the best performance.

[0126] Apply the trained decision tree model to the new user dataset, conduct classification prediction for each user, interpret the classification results output by the model, use the cross - validation method to evaluate the model performance on the training dataset, calculate evaluation metrics such as accuracy, recall, and F1 - score; analyze the results of different evaluation metrics to identify potential weaknesses and improvement spaces of the model; conduct A / B testing in the actual business environment, apply the classification results of the model to actual users, compare the impacts of different classification strategies on business metrics, and evaluate the actual effectiveness of the classification model according to business metrics (such as conversion rate, user satisfaction); select a suitable database or data warehouse (such as MySQL, PostgreSQL, Amazon Redshift) to store the classification results according to the data volume and access requirements, design the corresponding data table structure, including fields such as user ID, classification label, and classification time, write the classification results into the selected database or data warehouse, and create indexes for commonly queried fields to optimize data retrieval performance.

[0127] In a specific embodiment of the present invention, by determining the task plan for data classification, the goals and scope of classification can be clarified to ensure the orderly progress of the classification work. Dividing the tasks among different clients can make full use of resources, achieve division of labor and cooperation, and improve the efficiency of data processing. Extract the user data information of the clients, including user basic information, behavior data, and transaction data, ensuring the comprehensiveness and accuracy of the data. By setting clear classification criteria and rules, the consistency and accuracy of data classification can be ensured, avoiding subjectivity and randomness in the classification process. The classified user data is more orderly and clear, improving the usability and value of the data. The classified user data can provide strong support for precision marketing.

[0128] In a preferred embodiment of the present invention, integrate the analyzed data and the classified user data, and conduct training and inference on the integrated data to obtain the patterns and relationships between the data, including:

[0129] Integrate the analyzed data and the classified user data to obtain a unified dataset, specifically including: merge the datasets from different sources according to a specific key (such as user ID) to ensure the integrity of each user's data entries; splice the data with different features together to form a complete data entry containing all relevant information; for missing data, use methods such as mean filling, median filling, most frequent value filling, or interpolation to process it; dataset division:

[0130] Training set: A subset of data used for model training, usually accounting for 60 - 80% of the entire dataset.

[0131] Validation set: A subset of data used for model tuning and selection, usually accounting for 10 - 20%.

[0132] Test set: A subset of data used for model evaluation, usually accounting for 10 - 20%.

[0133] Using the dataset to train a deep learning network to obtain a trained deep learning network, specifically including: selecting a suitable deep learning architecture (such as Convolutional Neural Network CNN) according to data characteristics and task requirements; determining hyperparameters such as the number of network layers, the number of neurons in each layer, activation functions, optimizers, etc.; performing standardization, normalization, or other necessary preprocessing operations on the training set data; inputting the preprocessed training set data into the selected deep learning network, calculating the output through forward propagation; using the backpropagation algorithm to calculate the gradient of the loss function, updating the weights and biases of the network through the gradient descent algorithm, repeating the above process, iteratively training the network multiple times until reaching a preset stopping condition (such as reaching the maximum number of iterations); evaluating the performance of the model on the validation set, adjusting hyperparameters to optimize the model, and selecting the model with the best performance according to the performance on the validation set. After training and validation, a trained deep learning network is obtained, which can capture complex patterns and relationships between data.

[0134] A deep learning network can identify the correlations between different features. For example, in e-commerce data, there may be some associations between purchase history and user age, gender, and the network can learn how these features jointly affect user purchase behavior. For time series data or data with sequential properties, deep learning networks (such as Recurrent Neural Network RNN or its variants LSTM, GRU) can capture the patterns of data changing over time and how data at different time points interact with each other. In image or video data, deep learning networks (such as Convolutional Neural Network CNN) can learn the spatial structural relationships between pixels or frames to identify objects, scenes, or actions. Through a multi-layer structure, a deep learning network can extract and represent hierarchical features of data layer by layer. For example, in image recognition, the underlying network may learn low-level features such as edges and textures, while the high-level network may learn high-level features of object parts or the whole.

[0135] Predict and analyze new user data through the trained deep learning network to discover potential patterns and relationships in the data, specifically including: validating the trained deep learning network using a validation dataset. The validation process includes calculating metrics such as accuracy, recall, and F1-score on the validation set to evaluate the performance. Use the validated deep learning network to predict and analyze new user data, which involves predicting data such as user behavior and transaction records, or analyzing data such as user portraits and user preferences. Through prediction and analysis, discover potential patterns and relationships in the data, and integrate the prediction and analysis results to form a large processing result.

[0136] In a specific embodiment of the present invention, by integrating the analyzed data and the classified user data, a unified dataset containing richer and more comprehensive information is obtained. Preprocessing the data, such as cleaning, denoising, and normalizing, can improve the quality of the data, reduce outliers and noise in the data, and make the data more accurate and reliable. The integration and preprocessing processes ensure the consistency of the data, enabling subsequent data analysis and training to be carried out on a unified data basis, and avoiding problems caused by data inconsistency. Training a deep learning network using the preprocessed high-quality data can enable it to better learn the features and patterns in the data, improving performance and accuracy. The deep learning network has a powerful automatic feature extraction ability, which can automatically extract useful features from complex data, reducing the workload of manual feature engineering. By training the deep learning network, it is possible to learn the general patterns in the data, thereby having better generalization ability for new data and being able to adapt to the data analysis and prediction needs in different scenarios.

[0137] In a preferred embodiment of the present invention, according to the patterns and relationships between the data, perform user behavior prediction to achieve intelligent management of C-end user data, including:

[0138] Construct user portraits according to the patterns and relationships between the data, specifically including: extracting data related to users and integrating these data into a comprehensive user portrait. Each user portrait contains multiple features, such as age, gender, geographical location, browsing records, etc. According to the characteristics of the data, select features that have an important impact on user behavior prediction, and preprocess the features, such as standardization and normalization, to ensure the comparability and consistency between the features. Combine the processed feature data into user portraits, with each user corresponding to a unique portrait.

[0139] According to the user portrait and the historical behavior data of the user, through Calculate the similarity between users, where is the total number of features in the user portrait, is the user at the The value on the -th feature, is the value of the user on the -th feature, is the weight of the -th feature, is the mean value of the -th feature among all users, is the standard deviation of the -th feature among all users, is the similarity, and is the user profile,

[0140] Based on the similarity, for the user behavior data with time order, through time series analysis, to predict the future behavior trend of users; based on the similarity, for the discrete user behavior data, analyze the historical behavior data of users, through the recommendation of the user profile, to predict the future behavior trend of users; based on the predicted future behavior trend of users, to realize the intelligent management of C-end user data, specifically including:

[0141] Extract time-related features, such as time interval, periodicity, trend, etc.; for discrete behavior features, extract features such as the frequency, type, order of user behavior, etc.; based on the historical behavior data of users, generate a user profile, including interest preferences, consumption ability, activity, etc.

[0142] Collect the discrete behavior data of users, such as clicks, comments, shares, etc., remove invalid data, and encode the discrete behavior into a numerical form; based on the discrete behavior data of users, construct a user profile, including interest preferences, consumption habits, activity, etc.; according to the characteristics of time series data (such as trend, seasonality), select appropriate ARIMA model parameters (p, d, q), where p represents the number of autoregressive terms; d represents the number of differences used to make the time series stationary; q represents the number of moving average terms; use the preprocessed time series data to train the ARIMA model, use the historical data as the training set, estimate the model parameters through methods such as maximum likelihood estimation, and use the cross-validation method to verify the prediction performance of the model; use the trained ARIMA model to predict the future behavior trend of users, such as future purchase volume, page view volume, etc.

[0143] Based on the user profile and discrete behavior data, calculate the cosine similarity between users; the user profile is a description of the image and preferences of real users in the system, usually including the demographic characteristics of users (such as gender, age, region, etc.), behavior characteristics (such as browsing, clicking, purchase records, etc.) and content characteristics (such as categories, topics, keywords liked by users, etc.). The process of constructing a user profile usually involves data collection, cleaning, feature extraction and labeling.

[0144] Collect discrete behavior data:

[0145] Discrete behavior data refers to specific behavior records generated during the interaction between users and the system, such as user clicks, purchases, comments on goods, etc. These data usually exist in the form of logs and are the basis for building a recommendation model.

[0146] Construct a user-item behavior matrix: Based on the discrete behavior data of users, construct a user-item matrix, where the elements in the matrix represent the degree of user preference for items (such as ratings, click counts, etc.); According to the calculated user similarity, select the top K users with the highest similarity to the target user as the set of similar users; For an item i that the target user u has not heard of, calculate the recommendation score of the target user u for the item i based on the preference degree of the users in the similar user set S(u, K) for the item i. Among them, the calculation formula for the recommendation score of the target user u for the item i is:

[0147] ;

[0148] Where, is the predicted rating of user for item , that is, estimate the interest in based on the information of similar users; is the set of the K users most similar to user , that is, select the K users with the greatest similarity to and use their ratings to infer the rating; is the cosine similarity between user and user , which is used to measure the similarity of interests between user and . The larger the value, the more similar the two users are; is the average rating of user for all items. By subtracting the mean value, the differences in user rating habits can be eliminated. For example, some users are used to giving higher ratings while others are lower; is the set of items that user has rated. In order to predict the rating of user for item , the rated items are needed to help with the prediction; is the actual rating of user for item If the user If a score has been given to the item then directly use this score; is the absolute value of the similarity between users, and this value is used for normalization to ensure that the weight contributed by each similarity is within a reasonable range.

[0149] Using the trained collaborative filtering model and combining with user profile information, predict the content or products that the user may be interested in; integrate the results of time series prediction and collaborative filtering recommendation. For example, according to the predicted future behavior trend of the user by time series, adjust the weight or priority of collaborative filtering recommendation. According to the integrated results, provide personalized content or product recommendations for the user; evaluate the recommendation effect through indicators such as user feedback, click-through rate, and conversion rate, and optimize and adjust the model according to the evaluation results; as new data is continuously generated, regularly update the dataset and retrain the ARIMA model and collaborative filtering model; according to the changes in business requirements and user behavior, adjust the model parameters or select a new model to improve the accuracy of prediction and recommendation. Through the above steps, it is possible to predict the future behavior trend of the user and recommend the content or products that the user may be interested in, thereby improving the user experience and business effect.

[0150] In the specific embodiments of the present invention, by constructing a refined user profile, it is possible to more accurately understand the preferences, needs, and behavior patterns of each user. The user profile helps the enterprise accurately target the target user group, optimize the marketing strategy, and improve the response rate and conversion rate of marketing activities. By calculating the similarity between users, it is possible to discover user groups with similar interests or behaviors. Similarity calculation is the basis for constructing recommendations and helps to improve the accuracy of recommendations and user satisfaction. Using time series analysis, it is possible to capture the trend of user behavior changing over time, more accurately predict the future behavior of the user, and thus make more accurate decisions. By predicting user behavior, the enterprise can provide appropriate services or products when the user may need them, enhancing user stickiness and loyalty. Intelligent management of user data can help the enterprise automate and optimize many operation processes, reduce manual intervention, and improve work efficiency.

[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. An intelligent management system for C-end user data based on large models, characterized in that, Including: A data acquisition module that starts from an initial URL, obtains new URLs by requesting and parsing target site data, performs deduplication and filtering, and puts them into a queue for cyclic crawling to obtain C-end user data; A data transfer module that divides the C-end user data into chunks to calculate the dynamic parameters of each data chunk; Based on the dynamic parameters of each data chunk, determine the urgency level, sort and transmit the data chunks according to the urgency level, including: calculating the average value of the C-end user data to obtain a center point; calculating the standard deviation based on the C-end user data; determining a corresponding radius based on the standard deviation to define a circle; traversing each data point and calculating its Euclidean distance from the center point. If the distance ≤ radius, the data point is regarded as an "inner-circle point"; if the distance > radius, the data point is "projected" onto the circle through a mapping method to become an "on-circle point"; select the "inner-circle points" and "on-circle points" as the dataset to be processed; according to the distribution and quantity of the data, use the equidistant segmentation algorithm to divide the processed dataset into multiple data chunks; for each data chunk, calculate the dynamic parameters; analyze the dynamic parameters to obtain the urgency level index corresponding to each data chunk; sort all data chunks from high to low according to the urgency level index, and transmit the data chunks in the sorted order; A generation module that is used to receive data chunks, generate heat data and mine historical behavior data by tracking the user browsing path, analyzing jump relationships, and identifying churn nodes to record the operation behavior data of users on the client side; A data analysis module that identifies and corrects errors in user data and operation behavior data, performs feature selection and data reduction, and analyzes the reduced data to obtain the analyzed data; An extraction module that determines the data classification task, extracts user data information of each client, including user basic information, behavior and transaction data, and performs classification to obtain classified user data; A training module that integrates the analyzed data and the classified user data, trains and infers the integrated data to obtain the patterns and relationships between the data; An intelligent management module that predicts user behavior according to the patterns and relationships between the data to achieve intelligent management of C-end user data. The calculation formula for the dynamic parameter is: ; Among them, is a dynamic parameter, is the weight coefficient of the real-time factor, is the base of the natural logarithm, is the decay rate, is the current time, is the creation time of the data block, is the weight coefficient; is the number of data blocks, is the t th data block coefficient, is the t th frequency of the data block appearing in user behavior, is the data size weight coefficient, is the maximum number of bytes of the data block, is the number of bytes of the data block, is the index.

2. The intelligent management system for C-end user data based on large models according to claim 1, characterized in that, Starting from the initial URL, obtain new URLs by requesting and parsing target site data, perform deduplication and filtering, and put them into a queue for cyclic crawling to obtain C-end user data, including: Determine the crawling target of the C-end user data to be crawled, and the crawling target includes the corresponding website, forum or social media platform; According to the crawling target, obtain the web page link as the starting point of the crawler, that is, the initial URL; Use the HTTP library to initiate a network request to the target site, obtain the web page content, and parse the web page through a parsing library to extract the required data; Perform deduplication and filtering operations on the required data to extract new URL links; Put the new URL links into a queue, repeat the extraction process, continuously take out URLs from the queue for crawling until the preset crawling depth is reached to obtain C-end user data.

3. The intelligent management system for C-end user data based on large models according to claim 2, wherein Analyze dynamic parameters to obtain the urgency index corresponding to each data block, including: Analyze dynamic parameters to determine the parameter range; According to the parameter range, map different value ranges of the dynamic parameters to different urgency levels, and represent the mapping rules through the structure of a decision tree, where each node represents a judgment condition of a dynamic parameter, the branches represent different value ranges, and the leaf nodes represent the results of the urgency level; For each data block, extract its dynamic parameters, start from the root node, divide the data set into subsets according to the dynamic parameters, and create child nodes. Evaluate each node. If the performance improvement after division is not obvious, stop the division and mark the current node as a leaf node; Recursively repeat the above process on each child node until the stop condition is met, and traverse each non-leaf node from bottom to top to evaluate its performance after being replaced by a leaf node; According to the performance evaluation results after the leaf nodes, mark each leaf node with the corresponding urgency category to obtain the urgency index corresponding to each data block.

4. The intelligent management system for C-end user data based on large models according to claim 3, wherein, Generate heat data and mine historical behavior data by tracking the user browsing path, analyzing jump relationships, and identifying lost nodes to record the operation behavior data of the user on the client, including: Construct a navigation map of the user by tracking the browsing path of the user within the client, analyze the jump relationships between different pages of the user, and identify the key nodes where the user is lost to obtain the path analysis result; Generate heat data of the page according to the path analysis result and operation behavior to obtain the historical behavior data of the user; Mine and analyze the historical behavior data of the user, track events, and record the operation behavior data of the user on the client.

5. The intelligent management system for C-end user data based on large models according to claim 4, wherein Identify and correct errors in the user data and operation behavior data, perform feature selection and data reduction, and analyze the reduced data to obtain the analyzed data, including: Identify the error data in the user data and operation behavior data, and correct the error data to obtain the corrected data; Perform data selection on the corrected data, perform standardization processing on the selected feature data to obtain the standardized data matrix; For the standardized data matrix, calculate the covariance matrix, and perform eigenvalue decomposition on the covariance matrix to obtain the eigenvalues; Sort the eigenvalues from largest to smallest, and use the eigenvectors corresponding to the k largest eigenvalues as the principal components to obtain the principal component matrix; Project the standardized feature matrix of the original data onto the principal component matrix to obtain the reduced low-dimensional data matrix; Perform transformation analysis on the reduced data to obtain the analyzed data.

6. The intelligent management system for C-end user data based on large models according to claim 5, characterized in that, Determine the data classification task, extract the user data information of each client, including user basic information, behavior and transaction data, and perform classification to obtain the classified user data, including: Determine the task plan for data classification, divide the tasks into different clients, extract the user data information data of the clients, including user basic information, behavior data and transaction data, and the user basic information includes user ID, name, gender, age, address and contact information; Set classification criteria and rules, and perform classification operations on the user data information data of the clients to obtain the classified user data.

7. The intelligent management system for C-end user data based on large models according to claim 6, characterized in that, Integrate the analyzed data and classified user data, perform training and reasoning on the integrated data, and obtain patterns and relationships between the data, including: Integrate the analyzed data and classified user data to obtain a unified data set; Using the data set, training the deep learning network to obtain a trained deep learning network; New user data is predicted and analyzed through the trained deep learning network to discover potential patterns and relationships in the data.

8. The intelligent management system for C-end user data based on large models according to claim 7, characterized in that, Based on the patterns and relationships between data, user behavior prediction is performed to achieve intelligent management of C-end user data, including: Build user profiles based on patterns and relationships between data; Calculate the similarity between users based on user portraits and historical behavior data; Based on the similarity, for user behavior data with time sequence, time series analysis is performed to predict the user's future behavior trend; Based on the similarity, for discrete user behavior data, analyze the user's historical behavior data and predict the user's future behavior trend through the recommendation of user portraits; Based on the predicted future behavior trends of users, intelligent management of C-end user data can be achieved.

Citation Information

Patent Citations

  • Client information management system based on data analysis

    CN118628147A

  • Environment data analysis system based on cloud computing

    CN118982255A

  • Creation method and system based on multi-modal data processing and dynamic task planning

    CN119515303A