Big data skew regulation and control method based on multi-dimensional feature analysis and terminal

By combining multi-dimensional feature analysis and graph computing, we construct data feature graphs, identify and adjust data skew, optimize resource allocation, solve the data skew and uneven resource allocation problems in traditional methods, and improve the stability and efficiency of big data systems.

CN120632408APending Publication Date: 2025-09-12FUJIAN TIANQUAN EDUCATION TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510506619.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Traditional big data processing methods ignore the multidimensionality and complexity of data, resulting in data skew problems, uneven distribution of computing resources, inefficient data processing, high risk of system failure, and lack of comprehensiveness in operational strategies.

Method used

A multidimensional feature analysis method is used to construct a data feature graph through a graph construction algorithm. Key features are extracted for multidimensional analysis and graph calculation. The results are integrated to determine the tilt situation, formulate data control strategies, and monitor and adjust in real time.

Benefits of technology

Accurately identify data skew, optimize resource allocation, improve data processing efficiency, ensure system stability and performance, discover deep data connections, and support precise operational strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632408A_ABST
    Figure CN120632408A_ABST
Patent Text Reader

Abstract

The invention discloses a big data skew regulation and control method based on multi-dimensional feature analysis and a terminal, and the method comprises the steps: collecting data objects of all data sources from a big data system, preprocessing the data objects, storing the preprocessed data objects in a preset graph storage structure, and storing the preprocessed data objects in the preset graph storage structure; constructing a data feature graph by using a corresponding graph construction algorithm according to a preset graph storage structure; key features of the data object are extracted, multi-dimensional feature analysis is carried out on the data object according to the key features, meanwhile, graph calculation is carried out on a data feature graph, a multi-dimensional feature analysis result and a graph calculation result are integrated to determine the inclination condition of the data object, and a data regulation and control strategy of the big data system is formulated and executed according to the inclination condition; the inclination condition of the data object is monitored in real time, if the inclination condition changes, the data regulation and control strategy is adjusted according to the change condition, comprehensive multi-dimensional analysis and complex relation mining of the data can be achieved, the inclination condition can be accurately recognized, and therefore resource allocation is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a big data tilt control method and terminal based on multidimensional feature analysis. Background Art

[0002] In the field of big data processing, data skew is a common challenge. It refers to the uneven distribution of data across different processing nodes or partitions. For example, in large-scale e-commerce transaction data, most orders may be concentrated on a few popular products or merchants. As a result, when processing and analyzing data based on products or merchants, the amount of data in some partitions or nodes far exceeds that in other parts, which leads to a series of problems, including uneven distribution of computing resources, inefficient data processing, and potential risks of system failure.

[0003] Traditional big data processing methods often use relatively simple data models and analytical methods. In terms of feature processing, traditional methods tend to focus on a single or a few key features, while ignoring the multidimensionality and complexity of the data. For example, in early customer relationship management systems, customer levels might be divided based solely on the amount of their purchases, while ignoring other important features such as the frequency of their purchases and their preferences, resulting in an inaccurate and incomplete assessment of customer value. In addition, traditional methods mainly perform simple subject classification, parsing, and association based on data fields, lacking the ability to mine and analyze the deep connections between data. This makes it difficult to discover important information such as the correlation between user behavior patterns in different business links and the impact of social relationships between user groups on the business. This means that the formulation of operational strategies may be based only on surface data, which is prone to data skew. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a big data tilt control method and terminal based on multidimensional feature analysis, which can realize comprehensive multidimensional analysis of data and mining of complex relationships, accurately identify tilt conditions, and thus optimize resource allocation.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is: A big data tilt control method based on multidimensional feature analysis, comprising the steps of: S1. Collect data objects from various data sources in a big data system, preprocess the data objects, store the preprocessed data objects in a preset graph storage structure, and construct a data feature graph using a corresponding graph construction algorithm based on the preset graph storage structure; S2. Extract key features of the data object, perform multidimensional feature analysis on the data object based on the key features, perform graph calculation on the data feature graph, integrate the multidimensional feature analysis results and the graph calculation results to determine the tilt of the data object, and formulate and implement a data control strategy for the big data system based on the tilt; S3. Monitor the tilt of the data object in real time. If the tilt changes, adjust the data control strategy according to the change.

[0006] In order to solve the above technical problems, another technical solution adopted by the present invention is: A big data tilt control terminal based on multidimensional feature analysis includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, each step of the above-mentioned big data tilt control method based on multidimensional feature analysis is implemented.

[0007] The beneficial effects of the present invention are: a big data tilt control method and terminal based on multidimensional feature analysis provided by the present invention collects data objects from various data sources from a big data system, preprocesses the data objects, ensures the quality and availability of the data, stores the preprocessed data objects in a preset graph storage structure, and constructs a data feature graph based on the preset graph storage structure using a corresponding graph construction algorithm to intuitively and effectively display the relationship between data objects; extracts key features of the data objects, performs multidimensional feature analysis on the data objects based on the key features, and performs graph calculation on the data feature graph at the same time, integrates the multidimensional feature analysis results and the graph calculation results to determine the tilt of the data objects, can explore the deep connections between the data, and accurately identify the tilt of the data objects, so as to subsequently formulate and execute the data control strategy of the big data system according to the tilt situation to solve the data tilt problem; monitors the tilt of the data objects in real time, and if the tilt situation changes, adjusts the data control strategy according to the change to ensure the stability and performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 This is a flow chart of a big data tilt control method based on multidimensional feature analysis according to an embodiment of the present invention; Figure 2 Another flow chart of a big data tilt control method based on multidimensional feature analysis according to an embodiment of the present invention; Figure 3 This is a flowchart of the steps of constructing a data feature graph based on an adjacency matrix in an embodiment of the present invention; Figure 4 A flowchart of the steps for constructing a data feature graph based on an adjacency list in an embodiment of the present invention; Figure 5A schematic diagram of integrating multi-dimensional feature analysis results and graph calculation results in an embodiment of the present invention; Figure 6 Schematic diagram of a big data tilt control terminal based on multi-dimensional feature analysis according to an embodiment of the present invention; Description of labels: 1. A big data tilt control terminal based on multi-dimensional feature analysis; 2. Memory; 3. Processor. DETAILED DESCRIPTION

[0009] To illustrate the technical content, achieved objectives and effects of the present invention in detail, the following description is given in conjunction with the embodiments and accompanying drawings.

[0010] Please refer to Figure 1 The embodiment of the present invention provides a big data tilt control method based on multi-dimensional feature analysis, comprising the steps of: S1. Collect data objects from various data sources in a big data system, preprocess the data objects, store the preprocessed data objects in a preset graph storage structure, and construct a data feature graph using a corresponding graph construction algorithm based on the preset graph storage structure; S2. Extract key features of the data object, perform multidimensional feature analysis on the data object based on the key features, perform graph calculation on the data feature graph, integrate the multidimensional feature analysis results and the graph calculation results to determine the tilt of the data object, and formulate and implement a data control strategy for the big data system based on the tilt; S3. Monitor the tilt of the data object in real time. If the tilt changes, adjust the data control strategy according to the change.

[0011] From the above description, it can be seen that the beneficial effects of the present invention are: by collecting data objects from various data sources from the big data system, preprocessing the data objects, ensuring the quality and availability of the data, storing the preprocessed data objects in a preset graph storage structure, and using the corresponding graph construction algorithm to construct a data feature graph according to the preset graph storage structure, the relationship between data objects is intuitively and effectively displayed; the key features of the data objects are extracted, and multidimensional feature analysis is performed on the data objects according to the key features. At the same time, graph calculation is performed on the data feature graph, and the multidimensional feature analysis results and the graph calculation results are integrated to determine the inclination of the data objects. It is possible to explore the deep connections between the data and accurately identify the inclination of the data objects, so as to subsequently formulate and execute the data control strategy of the big data system according to the inclination situation and solve the data inclination problem; the inclination of the data objects is monitored in real time. If the inclination situation changes, the data control strategy is adjusted according to the change to ensure the stability and performance of the system.

[0012] Furthermore, extracting key features of the data object and performing multi-dimensional feature analysis on the data object based on the key features includes: The key features of the data object are extracted using a feature selection algorithm, and the principal component analysis technology is used to reduce the dimensionality of the extracted key features. According to the key features after dimensionality reduction, the data object is statistically analyzed and correlated to obtain multidimensional feature analysis results.

[0013] From the above description, it can be seen that by extracting key features through feature selection algorithms, data redundancy can be effectively reduced and analysis efficiency can be improved. The principal component analysis technology is used to reduce the dimensions of key features, simplify the data content, facilitate subsequent statistical analysis and correlation analysis, and explore the inherent continuity and regularity between data objects to obtain multi-dimensional feature analysis results, which helps to improve the accuracy and efficiency of data analysis, thereby revealing the deep-seated characteristics of data objects.

[0014] Furthermore, graph calculation is performed on the data feature graph, including: The degree and centrality index of each node in the data feature graph are calculated by a graph calculation algorithm, and the degree and centrality index of each node are used to mine the connected subgraph in the data feature graph to obtain a graph calculation result.

[0015] From the above description, we can see that the graph computing algorithm can quantify the degree of connectivity and importance of nodes in the data feature graph, effectively identify key nodes and connected subgraphs, and provide strong support for data analysis and mining, thereby improving the accuracy and efficiency of data analysis and mining.

[0016] Furthermore, the step of integrating the multi-dimensional feature analysis results and the graph calculation results to determine the tilt of the data object includes: Unifying the data format of the data object, and if the data object has time or space dimension information, aligning the key features of the data object and the data feature graph according to the time or space dimension information; The multidimensional feature analysis results and the graph calculation results are associated according to the identification of the data object to form a data set, and the key features of the data object and the graph calculation features in the graph calculation results are combined to generate a composite feature set. The composite feature set is filtered using a feature selection algorithm, and the inclination of the data object is determined based on the filtered composite feature set.

[0017] From the above description, it can be seen that by unifying the data format and aligning key features and data feature graphs according to time or spatial dimension information, the consistency and comparability of the data can be ensured; by associating the multidimensional feature analysis results and the graph calculation results, the comprehensiveness of data consideration is ensured, the accuracy of tilt judgment is improved, and the generation of composite feature sets and the use of feature selection algorithms for screening are helpful to remove redundant features, improve computing efficiency and the effectiveness of results, and determine the tilt of data objects based on the screened composite feature sets, providing strong support for the subsequent formulation of data tilt control strategies for big data systems.

[0018] Furthermore, a data control strategy for the big data system is formulated and executed according to the tilt situation, including: The key nodes and their associated data with a degree higher than a preset degree threshold are obtained from the tilt situation, the associated data are redistributed by data splitting or replication, the processing tasks of the associated data are assigned to different computing nodes for parallel execution, and the computing resource allocation of the computing nodes is optimized and adjusted.

[0019] From the above description, we can see that by redistributing the associated data of key nodes, we can avoid overloading a single node, balance the load of each computing node, improve resource utilization, and assign the processing tasks of associated data to different computing nodes for parallel execution, shortening data processing time, improving overall system performance, and optimizing and adjusting the allocation of computing node resources to ensure efficient utilization of system resources and further improve data processing efficiency.

[0020] Please refer to Figure 6 Another embodiment of the present invention provides a big data tilt control terminal based on multidimensional feature analysis, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the various steps of the above-mentioned big data tilt control method based on multidimensional feature analysis.

[0021] The present invention relates to a method and terminal for controlling big data tilt based on multidimensional feature analysis, which is applicable to big data tilt control scenarios. It can achieve comprehensive multidimensional analysis of data and mining of complex relationships, accurately identify tilt conditions, and thus optimize resource allocation. The following is an explanation of the specific implementation methods: Please refer to Figures 1 to 5 , embodiment 1 of the present invention is: A big data tilt control method based on multidimensional feature analysis, comprising the steps of: S1. Collect data objects from various data sources in the big data system, preprocess the data objects, store the preprocessed data objects in a preset graph storage structure, and construct a data feature graph using a corresponding graph construction algorithm based on the preset graph storage structure.

[0022] In this embodiment, data sources include databases, log files, sensors, etc., and data objects include structured, semi-structured, and unstructured data. Specifically, in e-commerce scenarios, user behavior data, product information data, order data, etc. can be collected, and noise, duplicate data, and erroneous data in the collected data can be removed. Missing values ​​can be processed, such as filling or deleting. For example, outliers in the user age field can be corrected, and missing values ​​in product prices can be filled with the average price of similar products.

[0023] Furthermore, in this embodiment, constructing a data feature graph includes defining nodes and edges. Data objects are defined as nodes, and edges are defined based on the relationships between data. For example, in e-commerce data, users and products are nodes, and user purchases can be edges. The weight of an edge can represent the number of purchases or the amount. The data feature graph is constructed using a corresponding graph construction algorithm based on the preset graph storage structure. The preset graph storage structure can be an adjacency matrix or an adjacency list, ensuring that the data feature graph accurately reflects the relationships between data.

[0024] Please refer to Figure 3 , the steps to construct the data feature graph based on the adjacency matrix are: 1. Clarify the basic information of the graph: Before building the graph, determine the total number of vertices in the graph, denoted as V. For example, the number of users in a social network can be considered the number of vertices in the graph. Determine whether the graph is directed or undirected. 2. Initialize the adjacency matrix: Create a V×V two-dimensional array (matrix) and initialize all elements of the matrix to 0. In an undirected graph, the matrix is ​​symmetric, but in a directed graph, it is not necessarily symmetric. 3. Add edges: For each edge, update the value of the corresponding position in the adjacency matrix according to the numbers of the two endpoints (vertices) of the edge and the weight of the edge (if it is a weighted graph). Specifically: in an undirected graph, if there is an edge between vertex i and vertex j, and the weight of the edge is w, then set the values ​​of A[i][j] and A[j][i] in the adjacency matrix to w; in a directed graph, if there is an edge from vertex i to vertex j, and the weight of the edge is w, then set the value of A[i][j] in the adjacency matrix to w; 4. Complete the graph construction: Repeat the steps in point 3 until all edges are added to the adjacency matrix, at which point the graph construction is complete.

[0025] Please refer to Figure 4 ,The steps to construct the data feature graph based on the adjacency list are: 1. Define the data structure Vertex structure: Create a vertex structure or class to represent each vertex in the graph. The vertex structure or class contains the vertex identification information, for example, an integer or string can be used to uniquely identify each vertex, and also contains a pointer to a linked list that can be used to store the vertex information adjacent to the vertex. Edge structure: Create an edge structure or class to represent the edges in the graph. The edge structure or class contains information such as the identifier of the target vertex connected by the edge and the weight of the edge (if it is a weighted graph). It also contains a pointer to the next edge, which can be used to construct the linked list structure in the adjacency list. 2. Initialize the graph Determine the number of vertices: determine the total number V of vertices in the graph. For example, in a graph representing a city transportation network, the number of cities is the number of vertices. Create a vertex array: Create an array containing V elements to store all vertices. Each element in the array is a vertex object, and initialize the adjacency list pointer of each vertex to NULL, indicating that there are no adjacent vertices at the beginning. 3. Add edges Parse edge information: For each edge to be added, obtain the identifiers of the two vertices it connects and the weight of the edge (if any). For example, to add an edge from vertex u to vertex v with a weight of w, you need to add edge information to the adjacency list of vertex u. Specifically, find the position of vertex u in the vertex array, find the corresponding linked list according to its adjacency list pointer, create a new edge node, store the identifier and weight w of the target vertex v in the node, and insert the node into the head or tail of the adjacency list of vertex u. If it is an undirected graph, you also need to perform the same operation in the adjacency list of vertex v to indicate the bidirectionality of the edge. 4. Complete the construction of the graph Repeat the steps of adding edges in point 3 until all edges are added to the graph. At this point, the construction of the entire graph is completed, and the graph structure is completely represented by the vertex array and the adjacency list of each vertex.

[0026] In this embodiment, the graph construction method to be adopted depends on the graph computing algorithm. If the graph computing algorithm needs to frequently query the connection relationship between any two vertices, or needs to perform matrix operations and other operations, such as calculating the transitive closure of the graph, the Floyd algorithm for finding the shortest path, etc., then the graph construction is implemented based on the adjacency matrix; if the graph computing algorithm mainly involves traversing the graph, such as depth-first search (DFS) and breadth-first search (BFS), or needs to frequently visit the adjacent vertices of a vertex, then the graph construction is implemented based on the adjacency list. By adopting a suitable graph construction method to match the graph computing algorithm, the time complexity of the algorithm is reduced.

[0027] S2. Extract the key features of the data object, perform multidimensional feature analysis on the data object based on the key features, and perform graph calculation on the data feature graph. Integrate the multidimensional feature analysis results and graph calculation results to determine the inclination of the data object. Formulate and execute the data regulation strategy of the big data system based on the inclination, specifically including S2.1-S2.4.

[0028] S2.1. Extracting key features of the data object and performing multi-dimensional feature analysis on the data object based on the key features, including: The key features of the data object are extracted using a feature selection algorithm, and the principal component analysis technology is used to reduce the dimensionality of the extracted key features. According to the key features after dimensionality reduction, the data object is statistically analyzed and correlated to obtain multidimensional feature analysis results.

[0029] In this embodiment, feature selection algorithms (such as chi-square test, mutual information, etc.) are used to extract key features of data objects, and principal component analysis (PCA) technology is used to reduce the dimensionality of the extracted key features, thereby reducing the complexity of data processing while retaining the main features.

[0030] Furthermore, in this embodiment, statistical analysis is performed on the data objects based on the key features after dimensionality reduction processing, including: calculating the statistics of each key feature, such as the mean, median, standard deviation, etc., to understand the basic distribution of the data, for example, calculating the average sales volume and standard deviation of prices of different product categories; performing correlation analysis on the data objects based on the key features after dimensionality reduction processing, analyzing the correlation between different key features, and finding strongly correlated features to provide a reference for subsequent analysis, wherein the correlation analysis includes but is not limited to correlation analysis between numerical features, correlation analysis between categorical features and numerical features, and correlation analysis between categorical features, specifically: 1. Correlation Analysis of Numerical Features: Use the Pearson correlation coefficient to measure the linear correlation between two numerical features. The Pearson correlation coefficient ranges from -1 to 1, with a value of 1 indicating a perfect positive correlation, a value of -1 indicating a perfect negative correlation, and a value of 0 indicating no linear correlation. For example, in e-commerce data, if the Pearson correlation coefficient between a user's purchase frequency and purchase amount is non-zero, the relationship between these two features is linearly correlated.

[0031] 2. Correlation Analysis between Categorical and Numerical Features: For categorical features, group the numerical features according to their different values ​​and then calculate statistics for each group, such as the mean, median, and standard deviation. For example, in e-commerce data, group user purchase amounts by product category and calculate the average purchase amount for each category to understand the impact of different product categories on purchase amounts.

[0032] 3. Correlation Analysis between Categorical Features: Use a contingency table to analyze the correlation between two categorical features. Rows represent different values ​​of one categorical feature, and columns represent different values ​​of the other categorical feature. Calculate the frequency and count in the contingency table to analyze the relationship between the two categorical features.

[0033] S2.2. Performing graph calculation on the data feature graph, including: The degree and centrality index of each node in the data feature graph are calculated by a graph calculation algorithm, and the degree and centrality index of each node are used to mine the connected subgraph in the data feature graph to obtain a graph calculation result.

[0034] In this embodiment, a graph computing algorithm is used to calculate the degree (including out-degree and in-degree) of each node. Nodes with high degrees may be key nodes that cause data skew. For example, in a social network, internet celebrities with many followers have high degrees. Node centrality metrics, such as betweenness centrality and closeness centrality, are calculated to identify nodes in key positions in the network, which may cause data skew. Densely connected subgraphs in the graph are mined. These subgraphs contain highly correlated data, and excessive amounts of data can lead to skew. For example, in an e-commerce graph, a dense subgraph consisting of users who frequently purchase from each other is discovered. By quantifying the degree of connectivity and importance of nodes in the data feature graph through a graph computing algorithm, key nodes and connected subgraphs can be effectively identified, providing strong support for data analysis and mining, thereby improving the accuracy and efficiency of data analysis and mining.

[0035] S2.3. Integrating the multi-dimensional feature analysis results and the graph calculation results to determine the tilt of the data object includes: Unifying the data format of the data object, and if the data object has time or space dimension information, aligning the key features of the data object and the data feature graph according to the time or space dimension information; The multidimensional feature analysis results and the graph calculation results are associated according to the identification of the data object to form a data set, and the key features of the data object and the graph calculation features in the graph calculation results are combined to generate a composite feature set. The composite feature set is filtered using a feature selection algorithm, and the inclination of the data object is determined based on the filtered composite feature set.

[0036] In this embodiment, by unifying the data format of data objects, the data used for graph computation and multidimensional feature analysis is ensured to have consistent formats and identifiers. For example, if the nodes in the graph represent users, the multidimensional feature data is also recorded per user and linked using user identifiers (such as user IDs), which are identical and unique. If the data has temporal or spatial dimensional information, the key features of the data object and the data feature graph are aligned based on this temporal or spatial dimensional information. For example, when analyzing urban transportation networks (graph data) and multidimensional features such as population and economy of various urban regions, it is ensured that the data for both are from the same time period and geographic region.

[0037] Further, in this embodiment, please refer to Figure 5 , integrating the multi-dimensional feature analysis results and the graph calculation results to determine the tilt of the data object, including data level integration and feature level integration, specifically: 1. Data-level integration: Based on the data's common identifiers (e.g., node IDs), graph computation results (e.g., node degree, centrality, etc.) are linked with multidimensional feature analysis results (e.g., user age, spending amount, etc.) to form a dataset containing more information. For example, in social network analysis, the graph centrality metric of each user node is integrated with multidimensional feature information such as the user's age, gender, and interests. The linked data is then reorganized into a new dataset, which can be stored in a table format, with each row representing a node or data object and each column representing a feature (including both graph computation features and multidimensional features). 2. Feature-level integration: Combine graph computing features and multidimensional features to create new composite features. For example, multiplying a node's degree by a user's consumption frequency yields a new feature that serves as a comprehensive indicator of a user's social network activity and spending power. Feature selection algorithms (such as correlation analysis, chi-square tests, and random forest feature importance) are used to filter out the most valuable features for subsequent analysis or modeling from the integrated feature set, removing redundant or irrelevant features to reduce data dimensionality and computational complexity.

[0038] S2.4. Develop and implement a data control strategy for the big data system based on the tilt, including: The key nodes and their associated data with a degree higher than a preset degree threshold are obtained from the tilt situation, the associated data are redistributed by data splitting or replication, the processing tasks of the associated data are assigned to different computing nodes for parallel execution, and the computing resource allocation of the computing nodes is optimized and adjusted.

[0039] In this embodiment, for the determined key nodes and their associated data, data splitting or replication is used for redistribution. For example, the fan data of the Internet celebrity node is split into different computing nodes by region, or the popular product data is replicated to multiple storage nodes. In addition, the computing resource allocation is dynamically adjusted according to the data skew and the importance of the node, and more CPU, memory and other resources are allocated to the tasks of processing key node data. At the same time, tasks are reasonably scheduled to avoid tasks being concentrated in areas with severe data skew. Tasks related to key nodes are assigned to different computing nodes for parallel execution to improve processing efficiency.

[0040] S3. Monitor the tilt of the data object in real time. If the tilt changes, adjust the data control strategy according to the change.

[0041] In this embodiment, the tilt of the data object is monitored in real time. If the tilt changes, the data control strategy is adjusted according to the change to ensure the stability and performance of the system.

[0042] This embodiment also provides the following specific application scenarios: Scenario 1: Applying the solution of the present invention to social network analysis, through graph computing to analyze the closeness of social relationships between users and the scope of influence, such as calculating the degree and centrality of nodes to measure users' social activity and influence. At the same time, combined with the multi-dimensional characteristics of users, for example, correlation analysis between users' age and social activity can be performed to understand the behavioral characteristics of users of different age groups in social networks; or combining interests and hobbies with social circles to analyze the impact of interests on the formation of social relationships, etc., thereby providing social platforms with accurate recommendation services, user profile construction and other support, improving user experience and platform operational efficiency.

[0043] Scenario 2: Applying the solution of this invention to urban traffic planning, using graph computing to study traffic network flow distribution and congestion conditions, such as calculating the traffic load of each road and the congestion index of each node. This is combined with the multidimensional characteristics of the city, such as combining population density with traffic flow to analyze traffic pressure in densely populated areas. Correlating economic development level with road utilization to understand the relationship between traffic demand and economic activity in different regions provides a scientific basis for urban traffic planning, road construction, and public transportation route optimization.

[0044] Scenario 3: Applying this solution to financial risk assessment, graph computing is used to analyze capital flow networks and related transactions, calculating indicators such as the risk propagation probability of nodes. By combining multi-dimensional customer characteristics, such as income, liabilities, and credit history, the customer's repayment ability and default risk can be assessed. Consumer behavior can be linked to investment preferences to analyze the customer's investment risk tolerance, providing comprehensive information support for financial institutions' credit approval, risk warning, and investment decision-making, thereby reducing financial risks.

[0045] Please refer to Figure 6 , embodiment 2 of the present invention is: a big data tilt control terminal 1 based on multidimensional feature analysis, including a memory 2, a processor 3 and a computer program stored in the memory 2 and executable on the processor 3. When the processor 3 executes the computer program, each step of a big data tilt control method based on multidimensional feature analysis of embodiment 1 is implemented.

[0046] In summary, the present invention provides a big data tilt control method and terminal based on multidimensional feature analysis, which collects data objects from various data sources from the big data system, preprocesses the data objects, ensures data quality and availability, stores the preprocessed data objects in a preset graph storage structure, and constructs a data feature graph based on the preset graph storage structure using a corresponding graph construction algorithm to effectively present the complex relationship between data objects; extracts key features of the data objects, performs multidimensional feature analysis on the data objects based on the key features, and calculates the data feature graph at the same time, integrates the multidimensional feature analysis results and the graph calculation results to determine the tilt of the data objects, accurately identifies the data tilt problem, so as to subsequently formulate and execute the data control strategy of the big data system according to the tilt situation, including data redistribution, resource allocation adjustment and task scheduling optimization; monitors the tilt of the data objects in real time, and if the tilt situation changes, adjusts the data control strategy according to the change to ensure the stability and performance of the system.

[0047] In summary, the embodiments of the present invention can achieve the following beneficial effects: 1. When constructing a graph structure, you can freely choose to store it in an adjacency matrix or adjacency table format based on the characteristics of the data. The adjacency matrix is ​​simple and intuitive to construct, making it easy to quickly determine the connection relationships between nodes. When processing large-scale sparse graphs, the adjacency table can effectively save space and improve storage and processing efficiency. By clearly defining nodes and edges (such as users and products as nodes and purchasing behavior as edges in e-commerce data), the complex relationships between data are accurately reflected, providing strong support for subsequent analysis. Graph computing algorithms are used to calculate indicators such as node degree and centrality. Combined with the results of multi-dimensional feature analysis, data skew can be accurately identified. Compared with traditional methods that rely solely on graph computing or feature analysis, this solution's analysis is more comprehensive and accurate, helping to discover hidden patterns and problems in the data.

[0048] 2. Combine graph computation results with multidimensional feature analysis results, integrating them at both the data and feature levels. At the data level, the two types of results are linked based on common identifiers to form a dataset containing more information. At the feature level, the two types of features are combined to create composite features. Feature selection algorithms are then used to select the most valuable features, reducing data dimensionality and computational complexity.

[0049] 3. By integrating the above-mentioned graph calculation results with the multidimensional feature analysis results, deeper insights into the data can be gained. In the e-commerce field, comprehensive analysis of the relationship between users and products, multidimensional user characteristics, and product attributes can be used to uncover the connection between users' potential purchasing needs and products. This provides stronger data support for e-commerce platforms' precision marketing and product recommendations. Compared to traditional methods, this method can uncover more valuable information and improve the scientific nature and effectiveness of business decisions.

[0050] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent transformations made using the contents of the present invention's description and drawings, or directly or indirectly applied in related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A big data tilt control method based on multidimensional feature analysis, characterized in that: Including steps: S1. Collect data objects from various data sources in a big data system, preprocess the data objects, store the preprocessed data objects in a preset graph storage structure, and construct a data feature graph using a corresponding graph construction algorithm based on the preset graph storage structure; S2. Extract key features of the data object, perform multidimensional feature analysis on the data object based on the key features, perform graph calculation on the data feature graph, integrate the multidimensional feature analysis results and the graph calculation results to determine the tilt of the data object, and formulate and implement a data control strategy for the big data system based on the tilt; S3. Monitor the tilt of the data object in real time. If the tilt changes, adjust the data control strategy according to the change.

2. The method for controlling the big data tilt based on multidimensional feature analysis according to claim 1, characterized in that: Extracting key features of the data object and performing multi-dimensional feature analysis on the data object based on the key features, including: The key features of the data object are extracted using a feature selection algorithm, and the principal component analysis technology is used to reduce the dimensionality of the extracted key features. According to the key features after dimensionality reduction, the data object is statistically analyzed and correlated to obtain multidimensional feature analysis results.

3. The method for controlling the big data tilt based on multidimensional feature analysis according to claim 2, characterized in that: Performing graph calculation on the data feature graph includes: The degree and centrality index of each node in the data feature graph are calculated by a graph calculation algorithm, and the degree and centrality index of each node are used to mine the connected subgraph in the data feature graph to obtain a graph calculation result.

4. The method for controlling the big data tilt based on multidimensional feature analysis according to claim 3, characterized in that: The step of integrating the multi-dimensional feature analysis result and the graph calculation result to determine the tilt of the data object includes: Unifying the data format of the data object, and if the data object has time or space dimension information, aligning the key features of the data object and the data feature graph according to the time or space dimension information; The multidimensional feature analysis results and the graph calculation results are associated according to the identification of the data object to form a data set, and the key features of the data object and the graph calculation features in the graph calculation results are combined to generate a composite feature set. The composite feature set is filtered using a feature selection algorithm, and the inclination of the data object is determined based on the filtered composite feature set.

5. The method for controlling the big data tilt based on multidimensional feature analysis according to claim 3 is characterized in that: Formulate and implement data control strategies for the big data system based on the tilt situation, including: The key nodes and their associated data with a degree higher than a preset degree threshold are obtained from the tilt situation, the associated data are redistributed by data splitting or replication, the processing tasks of the associated data are assigned to different computing nodes for parallel execution, and the computing resource allocation of the computing nodes is optimized and adjusted.

6. A big data tilt control terminal based on multi-dimensional feature analysis, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the following steps are implemented: S1. Collect data objects from various data sources in a big data system, preprocess the data objects, store the preprocessed data objects in a preset graph storage structure, and construct a data feature graph using a corresponding graph construction algorithm based on the preset graph storage structure; S2. Extract key features of the data object, perform multidimensional feature analysis on the data object based on the key features, perform graph calculation on the data feature graph, integrate the multidimensional feature analysis results and the graph calculation results to determine the tilt of the data object, and formulate and implement a data control strategy for the big data system based on the tilt; S3. Monitor the tilt of the data object in real time. If the tilt changes, adjust the data control strategy according to the change.

7. The big data tilt control terminal based on multi-dimensional feature analysis according to claim 6, characterized in that: Extracting key features of the data object and performing multi-dimensional feature analysis on the data object based on the key features, including: The key features of the data object are extracted using a feature selection algorithm, and the principal component analysis technology is used to reduce the dimensionality of the extracted key features. According to the key features after dimensionality reduction, the data object is statistically analyzed and correlated to obtain multidimensional feature analysis results.

8. The big data tilt control terminal based on multi-dimensional feature analysis according to claim 7, characterized in that: Performing graph calculation on the data feature graph includes: The degree and centrality index of each node in the data feature graph are calculated by a graph calculation algorithm, and the degree and centrality index of each node are used to mine the connected subgraph in the data feature graph to obtain a graph calculation result.

9. The big data tilt control terminal based on multi-dimensional feature analysis according to claim 8, characterized in that: The step of integrating the multi-dimensional feature analysis result and the graph calculation result to determine the tilt of the data object includes: Unifying the data format of the data object, and if the data object has time or space dimension information, aligning the key features of the data object and the data feature graph according to the time or space dimension information; The multidimensional feature analysis results and the graph calculation results are associated according to the identification of the data object to form a data set, and the key features of the data object and the graph calculation features in the graph calculation results are combined to generate a composite feature set. The composite feature set is filtered using a feature selection algorithm, and the inclination of the data object is determined based on the filtered composite feature set.

10. The big data tilt control terminal based on multi-dimensional feature analysis according to claim 8, characterized in that: Formulate and implement data control strategies for the big data system based on the tilt situation, including: The key nodes and their associated data with a degree higher than a preset degree threshold are obtained from the tilt situation, the associated data are redistributed by data splitting or replication, the processing tasks of the associated data are assigned to different computing nodes for parallel execution, and the computing resource allocation of the computing nodes is optimized and adjusted.