Multi-dimensional big data real-time analysis and management system

Through the multi-dimensional big data real-time analysis and management system, the integration and classification problems of traditional systems in multi-source heterogeneous data processing have been solved, and accurate data division, real-time monitoring and exception handling have been achieved, which has improved the accuracy of data analysis and the timeliness of corporate decision-making.

CN120705196APending Publication Date: 2025-09-26ZHUHAI WANDU TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510825210.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Traditional big data analysis and management systems have difficulty in efficiently integrating and classifying multi-source heterogeneous data, cannot meet real-time requirements, and lack effective data dimension processing and anomaly correction capabilities, resulting in inaccurate data analysis and delayed decision-making.

Method used

A multi-dimensional big data real-time analysis and management system is adopted to achieve accurate division, real-time monitoring and exception processing of multi-source heterogeneous data through data dimension collection unit, dimension classification processing unit, data flow tracking unit and data trend prediction unit, generate cross-dimensional data association map, optimize data flow and predict data trends.

Benefits of technology

It achieves precise analysis and real-time processing of multi-source heterogeneous data, improves the accuracy and efficiency of data analysis, ensures the stability of data flow, supports the security and reliability of enterprises' precision marketing, inventory management and financial transactions, and improves market forecasting and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705196A_ABST
    Figure CN120705196A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of big data analysis and management, and discloses a multi-dimensional big data real-time analysis and management system. The system comprises a data integration platform which is in communication connection with a plurality of units such as a data dimension acquisition unit and a dimension classification processing unit. The data dimension acquisition unit divides data dimensions and acquires data features to judge data flow stability; the dimension classification processing unit determines the data flow direction and dimension, and calculates coefficient judgment classification logic; the data flow direction tracking unit marks the data to be corrected and verifies and corrects the data; the data trend pre-judging unit evaluates the data trend; and the multi-dimensional data fusion unit fuses the data to generate a graph and adjusts a flow direction weight. According to the system, multi-source heterogeneous big data can be efficiently processed, real-time analysis and management are achieved, the defects of a traditional system in the aspect of data processing are effectively overcome, the accuracy, real-time performance and efficiency of big data processing are improved, and the system is widely applied to various big data application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data analysis and management, and in particular to a multi-dimensional big data real-time analysis and management system. Background Art

[0002] In today's digital age, the scale and complexity of big data are growing exponentially, and a large amount of multi-source heterogeneous data is emerging, which poses unprecedented challenges to the analysis and management technology of big data.

[0003] Traditional big data analysis and management systems have numerous shortcomings when processing multi-source, heterogeneous data. First, they struggle to efficiently integrate and categorize data from diverse sources with varying structures and formats. For example, in the internet industry, data such as user browsing history, transaction data, and device information comes from a wide range of sources, including both structured transaction data and unstructured user review data. Traditional systems are unable to accurately segment this data according to pre-set dimensions, significantly compromising the accuracy and completeness of data collection and hindering in-depth analysis.

[0004] The real-time requirements for data cannot be met. With the rapid development of business, many scenarios require real-time data analysis and processing to make timely decisions. For example, in the financial sector, stock trading, market conditions are constantly changing. Without real-time analysis of fluctuations in data such as stock prices and trading volume, investors and financial institutions will struggle to grasp trading opportunities, potentially missing out on profitable opportunities or even suffering significant losses. However, traditional systems are limited in data collection frequency and analysis speed. They cannot accurately assess data fluctuations based on real-time update rates, making it difficult to accurately assess the stability of data streams, leading to delayed decision-making.

[0005] When it comes to processing data dimensions, traditional systems lack effective classification and correlation capabilities. They can't clearly identify core and secondary data flows, making it difficult to deeply explore and leverage the complex relationships between data of different dimensions. For example, in the e-commerce industry, product sales data, user behavior data, logistics data, and other interrelated dimensions are interrelated. Traditional systems are unable to accurately label primary and associated processing dimensions, resulting in superficial data analysis and a failure to provide effective support for precision marketing, inventory management, and other aspects of a company's operations.

[0006] When data anomalies occur, traditional systems lack the ability to correct and predict them. They can't promptly and accurately mark and process mismatched, abnormal data nodes or data nodes with incomplete processing flows, nor can they effectively predict based on data trends. In industrial production, if equipment operating data shows anomalies, traditional systems can't quickly locate and correct the problem, nor can they predict the risk of equipment failure in advance, impacting production continuity and stability and causing economic losses. Summary of the Invention

[0007] The purpose of the present invention is to provide a multi-dimensional big data real-time analysis and management system to solve the problems raised in the above background technology.

[0008] To achieve the above objectives, the present invention provides the following technical solution: a multi-dimensional big data real-time analysis and management system, the system comprising:

[0009] The data dimension collection unit is used to divide multi-source heterogeneous data streams into structured data dimensions and unstructured data dimensions according to preset dimensions, and continuously collect data in each dimension. It obtains the data fluctuation amplitude based on the real-time update frequency corresponding to the data dimension, sets the change process of the current data stream as a single data cycle based on the data fluctuation amplitude, collects the static data features of the dimension and the dynamic data features of the dimension, and determines whether there are any abnormalities in the data stream stability based on feature comparison;

[0010] The dimension classification processing unit is used to set the main analysis direction in the current data dimension as the data core flow direction, and mark other data interaction paths that are cross-related with the data core flow direction as data auxiliary flow directions. At the same time, the dimension where the data core flow direction is located is marked as the main processing dimension, and the dimension where the data auxiliary flow direction is located is marked as the associated processing dimension; the dimension classification processing parameters are collected, substituted into the rule calculation to obtain the classification processing coefficient of the current data dimension, and the classification processing logic of the data dimension is determined to be invalid based on the coefficient comparison.

[0011] Preferably, the data integration platform is also communicatively connected to:

[0012] The data flow tracking unit is used to uniformly mark abnormal data nodes that do not match the data core flow in the main processing dimension and data nodes that are consistent with the data auxiliary flow but have not completed the complete processing flow in the associated processing dimension as data to be corrected. It obtains the input and output trajectories of the data to be corrected, performs data integrity verification based on trajectory data analysis, and generates real-time correction instructions;

[0013] The data trend prediction unit is used to collect data cycle parameters and data scale parameters, and evaluate whether there is any deviation in the real-time data trend prediction based on parameter analysis.

[0014] Preferably, the dimension static data characteristics and dimension dynamic data characteristics are respectively the baseline data mean during the period of low fluctuation of the main processing dimension data volume corresponding to adjacent single data cycles and the peak span change rate during the period of growth of the main processing dimension data volume corresponding to a single data cycle.

[0015] Preferably, if the static data feature of the dimension exceeds the preset baseline mean threshold, or the dynamic data feature of the dimension exceeds the peak change rate threshold, it is determined that there is an abnormality in the data flow stability; if the static data feature of the dimension does not exceed the preset baseline mean threshold, and the dynamic data feature of the dimension does not exceed the peak change rate threshold, it is determined that the data flow stability is normal.

[0016] Preferably, the dimension classification processing parameters include the incremental span of the path switching frequency of the associated processing dimension in the main processing dimension, the incremental span of the amount of data continuously accessed by non-identical associated processing dimensions in the main processing dimension, and the peak incremental span of the main processing dimension data density when the main processing dimension is continuously interacting with the associated processing dimension data;

[0017] If the classification processing coefficient of the current data dimension exceeds the classification processing coefficient threshold, the data dimension classification processing logic is determined to be invalid; if the classification processing coefficient of the current data dimension does not exceed the classification processing coefficient threshold, the data dimension classification processing logic is determined to be valid.

[0018] Preferably, the input trajectory and output trajectory of the data to be corrected are respectively the numerical ratio of the retention time of the data to be corrected at the entrance identification position of the main processing dimension to the probability of not triggering the core processing process, and the ratio of the retention time of the data to be corrected in the main processing dimension to the processing time of the associated processing dimension.

[0019] Preferably, if the input trajectory of the data to be corrected exceeds the input trajectory threshold, and the output trajectory of the data to be corrected exceeds the output trajectory threshold, a data redundancy alarm signal is generated; if the input trajectory of the data to be corrected does not exceed the input trajectory threshold, and the output trajectory of the data to be corrected does not exceed the output trajectory threshold, a data flow normal signal is generated.

[0020] Preferably, the data cycle parameter and the data scale parameter are respectively the ratio of the processing time interval of the main processing dimension in adjacent data cycles to the duration of the data volume growth cycle, and the quantitative deviation value between the current total amount of data in any main processing dimension and the preset processing capacity upper limit.

[0021] Preferably, if the data cycle parameter does not exceed the duration ratio threshold, and the data scale parameter does not exceed the quantity deviation value threshold, then it is determined that the data trend prediction is without deviation; if the data cycle parameter exceeds the duration ratio threshold, or the data scale parameter exceeds the quantity deviation value threshold, then it is determined that the data trend prediction is biased.

[0022] Preferably, the data integration platform is also communicatively connected to:

[0023] The multi-dimensional data fusion unit is used to dynamically match the real-time data of the main processing dimension and the associated processing dimension, generate a cross-dimensional data association map based on the matching results, and adjust the priority weights of the data core flow and the data auxiliary flow based on the map analysis results.

[0024] Compared with the prior art, the present invention has the following beneficial effects:

[0025] In terms of data collection and stability monitoring, the system's data dimension acquisition unit can accurately divide multi-source heterogeneous data streams into structured and unstructured data dimensions, and continuously collect data from each dimension. By obtaining the data fluctuation amplitude to set a single data cycle, and collecting the static and dynamic data characteristics of the dimension, it can accurately determine whether the data stream stability is abnormal. This allows the system to promptly detect abnormal data fluctuations. For example, in the Internet of Things environment, the large amount of data generated by sensors may be abnormal due to equipment failure or external interference. This system can quickly detect and issue early warnings, providing a reliable data foundation for subsequent data processing and decision-making, and avoiding analytical errors and decision-making errors caused by data anomalies.

[0026] The design of the dimension classification processing unit is highly advantageous. It clearly identifies the primary analysis direction and core data flow, identifies associated auxiliary flows and corresponding dimensions, and determines whether the classification processing logic has failed by collecting classification processing parameters and calculating coefficients. For example, in supply chain management, cargo flow data can be considered the core flow, with associated supplier information, inventory data, and other data as auxiliary flows, accurately determining the primary and associated processing dimensions. This allows the system to efficiently process complex supply chain data based on the inherent connections between data, optimizing business processes, improving overall supply chain efficiency, and reducing operating costs.

[0027] The data flow tracking unit is extremely powerful. It marks abnormal data nodes and those with incomplete processing flows as data to be corrected, captures input and output traces for integrity verification, and generates real-time correction instructions. In financial transaction systems, when transaction data anomalies or incomplete transactions occur, the system can quickly locate the problematic data and make timely corrections based on trace analysis, ensuring the accuracy and integrity of transaction data, safeguarding the security and reliability of financial transactions, and effectively mitigating financial risks caused by data errors.

[0028] The data trend prediction unit collects data cycle and scale parameters to assess whether there are any deviations in real-time data trend predictions. This is of great significance for a company's market forecasting and resource planning. For example, in the retail industry, by analyzing the cyclical changes and scale growth trends of historical sales data, combined with current data parameters, companies can predict market demand in advance and rationally arrange inventory, production, and distribution plans, avoiding inventory backlogs and stockouts, thereby improving the company's economic efficiency and market competitiveness.

[0029] The multidimensional data fusion unit dynamically matches real-time data from the primary and associated processing dimensions, generating a cross-dimensional data correlation map and adjusting flow priority weights. In social media data analysis, the fusion of multidimensional data—including users' personal information, social relationships, and published content—builds detailed user profiles and social relationship maps, helping companies better understand user behavior and needs, enabling precision marketing and personalized service recommendations, and improving user experience and market share. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 This is a working principle diagram of the multi-dimensional big data real-time analysis and management system of the present invention;

[0031] Figure 2 This is a working principle diagram of dimension classification processing coefficient calculation and logical judgment;

[0032] Figure 3 A working principle diagram for obtaining the data trajectory to be corrected;

[0033] Figure 4 This is a diagram showing the working principle of determining data flow status. DETAILED DESCRIPTION

[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0035] See also Figure 1-Figure 4 The present invention relates to a multi-dimensional big data real-time analysis and management system, and the specific implementation steps are as follows:

[0036] As the core hub of the entire system, the data integration platform is responsible for coordinating data interaction and collaboration between various functional units. Its communication connection, the data dimension collection unit, plays a crucial and fundamental role in the entire system's data processing flow. This unit undertakes the critical task of classifying and processing multi-source heterogeneous data streams. It clearly divides multi-source heterogeneous data streams into structured and unstructured data dimensions according to preset dimensions. In actual operation, the setting of preset dimensions is determined by a combination of factors such as the type and source of the data processed by the system, as well as the intended analysis objectives. Once the division is completed, the data dimension collection unit continuously and uninterruptedly collects data from each dimension.

[0037] During the collection process, the data dimension collection unit determines the data fluctuation range based on the corresponding real-time update frequency of the data dimension. Different data dimensions often have different update frequencies. For example, some monitoring data dimensions may have a high update frequency, while some basic information data dimensions may have a relatively low update frequency. By accurately understanding the update frequency, the data fluctuation range can be accurately calculated. Based on the data fluctuation range, the unit then defines the current data stream change process as a single data cycle. This setting is crucial for accurately analyzing data characteristics and determining the stability of the data stream.

[0038] After completing the setting of a single data cycle, the data dimension collection unit will collect dimensional static data features and dimensional dynamic data features. Among them, the dimensional static data features and dimensional dynamic data features are respectively the baseline data mean of the main processing dimension data volume fluctuation trough period corresponding to the adjacent single data cycle, and the peak span change rate of the main processing dimension data volume growth period corresponding to the single data cycle. By collecting and analyzing these features, the unit can judge whether there is an abnormality in the data flow stability based on feature comparison. Specifically, if the dimensional static data feature exceeds the preset baseline mean threshold, or the dimensional dynamic data feature exceeds the peak change rate threshold, it is determined that there is an abnormality in the data flow stability; if the dimensional static data feature does not exceed the preset baseline mean threshold, and the dimensional dynamic data feature does not exceed the peak change rate threshold, it is determined that the data flow stability is normal.

[0039] The dimension classification processing unit, which communicates with the data integration platform, also plays an indispensable role. This unit first sets the primary analysis direction within the current data dimension as the core data flow. This designation is based on a deep understanding of the data's business logic and the key points of system analysis. It also marks other data interaction paths that cross-correlate with the core data flow as auxiliary data flows. Based on this, the dimension classification processing unit marks the dimension containing the core data flow as the primary processing dimension and the dimension containing the auxiliary data flow as the associated processing dimension.

[0040] The dimension classification processing unit also collects dimension classification processing parameters. These parameters include the incremental span of the path switching frequency of the associated processing dimension in the main processing dimension, the incremental span of the amount of data continuously accessed by non-identical associated processing dimensions in the main processing dimension, and the peak incremental span of the main processing dimension data density when the main processing dimension is accompanied by the associated processing dimension data interaction. After collecting these parameters, the unit will substitute the rules into the calculation to obtain the classification processing coefficient of the current data dimension. Finally, based on the coefficient comparison, it is determined whether the classification processing logic of the data dimension is invalid. If the classification processing coefficient of the current data dimension exceeds the classification processing coefficient threshold, the data dimension classification processing logic is determined to be invalid; if the classification processing coefficient of the current data dimension does not exceed the classification processing coefficient threshold, the data dimension classification processing logic is determined to be valid.

[0041] The present invention will be further described below in conjunction with Examples 1 to 5:

[0042] Example 1:

[0043] The data flow tracking unit connected to the data integration platform is mainly responsible for monitoring and correcting data nodes. During actual operation, the unit will uniformly mark abnormal data nodes in the main processing dimension that do not match the core flow of data, and data nodes in the associated processing dimension that are consistent with the auxiliary flow of data but have not completed the complete processing process as data to be corrected. This marking process is based on strict monitoring of the data flow and judgment of the integrity of the data processing process. After the marking is completed, the data flow tracking unit will obtain the input trajectory and output trajectory of the data to be corrected. Among them, the input trajectory and output trajectory of the data to be corrected are the numerical ratio of the residence time of the data to be corrected at the entry identification position of the main processing dimension to the probability of not triggering the core processing process, and the ratio of the residence time of the data to be corrected in the main processing dimension to the processing time of the associated processing dimension.

[0044] After acquiring these trajectory data, the data flow tracking unit will perform data integrity verification based on the trajectory data analysis. During the data integrity verification process, the system will comprehensively consider various data indicators of the input trajectory and the output trajectory. For example, by analyzing the retention time of the data to be corrected at the entry identification position of the main processing dimension and the corresponding numerical ratio of the probability of not triggering the core processing process, it can be determined whether there is an abnormal delay when the data enters the main processing dimension or whether the processing process is not started normally. The ratio of the retention time of the data to be corrected in the main processing dimension to the processing time of the associated processing dimension helps to determine whether the data flows smoothly between different dimensions.

[0045] After completing the data integrity check, the data flow tracking unit generates real-time correction instructions based on the verification results. If the input trajectory of the data to be corrected exceeds the input trajectory threshold, and the output trajectory of the data to be corrected exceeds the output trajectory threshold, this indicates that there may be redundancy or abnormality in the data flow process. At this time, the data flow tracking unit will generate a data redundancy alarm signal so that system administrators can take timely measures to address it. If the input trajectory of the data to be corrected does not exceed the input trajectory threshold, and the output trajectory of the data to be corrected does not exceed the output trajectory threshold, it indicates that the data flow is normal, and the data flow tracking unit will generate a data flow normal signal.

[0046] In the application scenario of a multi-dimensional big data real-time analysis management system of an e-commerce platform, the e-commerce platform generates massive order data, user browsing data, product information data, etc. every day. These data constitute multi-source heterogeneous data streams.

[0047] During the order processing process, the data flow tracking unit begins operating. Assume that the system defines order generation, payment confirmation, product shipment, and logistics distribution as the primary processing dimensions for order processing, with the core data flow being the forward process from order generation to final delivery. During a certain time period, the data flow tracking unit detects that a batch of order data is significantly longer than normal at the order generation stage (the entry point for the primary processing dimension), and the probability that these orders fail to trigger subsequent core processing steps (such as payment confirmation) has also increased significantly. For example, under normal circumstances, the average retention time of orders in the order generation stage is 1 minute, with a 5% probability of failing to trigger the core processing step. However, at this point, some orders are detected to be held for up to 5 minutes, with a probability of failing to trigger the core processing step rising to 30%. This results in a significant difference between the retention time of the data to be corrected at the entry point for the primary processing dimension and the probability of failing to trigger the core processing step.

[0048] At the same time, in the associated processing dimension (such as the user information associated dimension, which is used to verify the accuracy and completeness of user information and assist in order processing), the data flow tracking unit found that although some data were consistent with the data auxiliary flow (that is, they participated in the user information verification process related to order processing), they did not complete the entire processing process. For example, when verifying user address information, the verification of user address information corresponding to some orders was interrupted, but no effective processing was carried out subsequently. The retention time of this data in the main processing dimension (order processing dimension) far exceeded the processing time of the associated processing dimension (user information verification dimension). Assuming that under normal circumstances, the ratio of the order retention time in the main processing dimension to the processing time in the associated processing dimension is 3:1, but at this time, this ratio for some orders has reached 10:1.

[0049] Based on the above situation, these data are marked as pending correction data. After the data flow tracking unit obtains the input trajectory (the aforementioned retention time and probability corresponding value ratio) and output trajectory (the ratio of the main processing dimension to the associated processing dimension duration) of these pending correction data, it begins to perform data integrity verification.

[0050] During the data integrity verification process, the system comprehensively analyzes the data from the input and output trajectories. From the input trajectory, there's a high probability that an order remains in the order generation process for too long and fails to trigger the core processing flow. This could indicate a malfunction in the order generation system, such as a slow server response, preventing new orders from entering the subsequent processing flow in a timely manner. Alternatively, there could be an anomaly in the front-end user's order placement, resulting in a large number of invalid orders being generated. From the output trajectory analysis, the ratio of the duration of the primary processing dimension to the duration of the associated processing dimension is too large, indicating that after a problem with the associated processing dimension occurred, the primary processing dimension failed to adjust in time or wait for the associated processing to complete. This could be due to a flaw in the data interaction mechanism and the lack of effective exception handling logic.

[0051] After completing the data integrity check, the system generates corresponding instructions based on the check results. Since the input trajectory of the data to be corrected exceeds the input trajectory threshold (assuming the input trajectory threshold is set to 1.5 times the normal ratio), and the output trajectory of the data to be corrected exceeds the output trajectory threshold (assuming the output trajectory threshold is set to 2 times the normal ratio), the data flow tracking unit generates a data redundancy alarm signal. This signal will be sent to the system's monitoring center and operation and maintenance team. After receiving the signal, the operation and maintenance personnel will immediately check and repair the order generation system, data interaction mechanism, etc. to ensure the smooth progress of the order processing process, avoid data redundancy and processing anomalies, and ensure the normal operation of the e-commerce platform. If in other cases, the input trajectory of the data to be corrected does not exceed the input trajectory threshold, and the output trajectory of the data to be corrected does not exceed the output trajectory threshold, for example, the order retention time in the order generation link and the probability of not triggering the core processing process are within the normal range, and the ratio of the main processing dimension to the associated processing dimension time is also normal, then the data flow tracking unit will generate a data flow normal signal, indicating that the current order processing process and related data interaction are in normal state.

[0052] Example 2:

[0053] The data trend prediction unit is also a crucial component of the communication connection with the data integration platform. This unit is primarily responsible for collecting data cycle and data scale parameters, and using these parameters to analyze and evaluate whether there are any deviations in the real-time data trend prediction. The data cycle and data scale parameters are the ratio of the processing interval of the primary processing dimension to the data volume growth period within adjacent data cycles, and the deviation between the current total amount of data in any primary processing dimension and the preset processing capacity limit, respectively.

[0054] When collecting data cycle parameters, the system continuously monitors the processing intervals of the primary processing dimension and the duration of the data growth cycle within adjacent data cycles. By calculating the ratio of these two, we can reflect the data processing rhythm and the relative speed of data growth. For example, if the processing intervals of the primary processing dimension are long while the data growth cycle is short, the ratio will be large, which may indicate that the system is facing certain pressure when processing data, and the data growth rate is exceeding the system's processing capacity.

[0055] To collect data scale parameters, the system collects the current total data volume for any primary processing dimension in real time, compares it to the preset processing capacity limit, and calculates a quantity deviation value. This deviation value provides a direct reflection of the current data load in the primary processing dimension. If the quantity deviation value is large, it indicates that the current total data volume is approaching or exceeding the preset processing capacity limit, and the system may need to adjust resources or divert data.

[0056] After collecting the data cycle parameters and data scale parameters, the data trend prediction unit will analyze and evaluate whether there is any deviation in the real-time data trend prediction based on these parameters. If the data cycle parameter does not exceed the duration ratio threshold, and the data scale parameter does not exceed the quantity deviation threshold, the data trend prediction is determined to be unbiased, indicating that the system's current data processing is relatively stable, and data growth and processing rhythm are within a reasonable range. If the data cycle parameter exceeds the duration ratio threshold, or the data scale parameter exceeds the quantity deviation threshold, the data trend prediction is determined to be biased. At this time, the system needs to adjust the data processing strategy to cope with possible data processing pressure or insufficient resources.

[0057] In a real-time monitoring and analysis system for urban traffic flow, the data trend prediction unit processes traffic data to evaluate whether there is any deviation in the real-time data trend prediction.

[0058] In this system, the primary processing dimension is the traffic flow monitoring dimension of the city's main roads, responsible for collecting key data such as traffic volume and speed on each road section. The data cycle parameter is the ratio of the processing time interval of the primary processing dimension to the duration of the data volume growth period within adjacent data cycles, expressed as: Among them, R represents the data cycle parameter, T interval Indicates the processing time interval of the main processing dimension in adjacent data cycles, that is, the time interval between the completion of the previous data cycle processing and the start of the next data cycle processing; T growth Indicates the duration of the data volume growth cycle, that is, the time it takes for traffic volume to grow from a relatively stable value to a peak value.

[0059] For example, during the morning rush hour, the system sets a data cycle of every 15 minutes to process traffic flow data. In a certain day's monitoring, the interval time T between the completion of the previous data cycle and the start of the next data cycle is interval is 10 minutes, and the time T from the beginning of traffic growth to the peak of the morning rush hour is growth is 60 minutes, then the formula can be calculated to get

[0060] The data scale parameter is the deviation between the total amount of current data in any main processing dimension and the preset processing capacity limit. Assume that the traffic volume (i.e. the total amount of current data) of a main road in a city at a certain moment is N, and the maximum traffic volume (preset processing capacity limit) that the road can carry is N. max , then the data scale parameter D = NN max For example, a road has a preset processing capacity limit N max The number of vehicles per hour is 3,000, and the actual monitored traffic flow N at a certain moment is 3,200 vehicles. Then the data scale parameter D = 3,200 - 3,000 = 200 vehicles.

[0061] After collecting the data cycle parameters and data scale parameters, the data trend prediction unit begins to evaluate whether there is a deviation in the real-time data trend prediction. The system presets the duration ratio threshold as The quantity deviation threshold is 150 vehicles. The data scale parameter D=200>150, wherein the data scale parameter exceeds the quantity deviation value threshold, so it is determined that there is a deviation in the data trend prediction.

[0062] This indicates that the current urban traffic situation is abnormal. It may be that there are emergencies such as traffic accidents or road construction on the road, which have caused the traffic volume to exceed the road carrying capacity and is inconsistent with the originally expected traffic flow trend. At this time, the system will promptly issue an early warning message to notify the traffic management department to take corresponding measures, such as increasing the number of traffic police to divert traffic, issuing road condition information to guide vehicles to detour, etc., to alleviate traffic pressure and ensure the normal operation of urban traffic. If in other cases, the data cycle parameter does not exceed the duration ratio threshold, and the data scale parameter does not exceed the quantity deviation threshold, for example If D=100, it is determined that the data trend prediction has no deviation, indicating that the current urban traffic flow is within the normal fluctuation range and the traffic operation is in good condition.

[0063] Example 3:

[0064] When collecting static data features for a dimension, the system monitors the peak and valley periods of data volume for the corresponding primary processing dimension within each data cycle. Within each data cycle, data volume fluctuates over time, often with low peaks. During these low peaks, the system continuously collects data for the primary processing dimension and calculates the average baseline data value. This average value reflects the data volume level of the primary processing dimension under relatively stable conditions.

[0065] To collect dynamic data features for a dimension, the system focuses on a single data cycle corresponding to the period of data growth in the primary processing dimension. As data volume grows, the system records the change from the starting point to the peak value and calculates the rate of change of the peak span. This rate reflects the speed of data volume change during the growth phase and is an important indicator of data dynamics.

[0066] After collecting the static and dynamic data features of a dimension, the system will make an assessment based on preset thresholds. The preset baseline mean threshold and peak rate of change threshold are pre-set based on various factors, including the system's historical data, business requirements, and performance metrics. If the static data feature of a dimension exceeds the preset baseline mean threshold, this may indicate an abnormal increase in data volume during a period of relatively stable data volume, potentially indicating issues such as abnormal data input or system failure. If the dynamic data feature of a dimension exceeds the peak rate of change threshold, this indicates that data growth during a period of rapid growth exceeded the system's expected range, potentially impacting system stability and processing capacity. Therefore, if either of these two conditions occurs, the system determines that there is an abnormality in data flow stability. Conversely, if the static data feature of a dimension does not exceed the preset baseline mean threshold and the dynamic data feature does not exceed the peak rate of change threshold, then both static and dynamic data changes are within normal ranges, and the system determines that the data flow stability is normal.

[0067] For example, a multi-dimensional big data real-time analysis and management system for an internet video platform generates massive amounts of heterogeneous data from multiple sources, including video playback data and user interaction data, every day. Processing video playback data is a crucial component of the primary processing dimension. The system processes this data in a single data cycle, set at hourly intervals.

[0068] When collecting static data features for a dimension, the system focuses on periods of low data volume fluctuation for the primary processing dimension corresponding to adjacent single data cycles. For example, between 2:00 AM and 3:00 AM, most users are resting, and video play volume is relatively low, representing a period of low data volume fluctuation. The system continuously collects video play data during this period. Suppose that the video play data for this period for two consecutive days is as follows: on the first day, play volume is recorded every 5 minutes, with the data being 200, 210, 205, 208, and 203, respectively; on the second day, the corresponding data is 220, 215, 218, 213, and 216. Calculation shows that the average baseline data for this period on the first day is (200 + 210 + 205 + 208 + 203) ÷ 5 = 205.2; and on the second day, the average baseline data for this period is (220 + 215 + 218 + 213 + 216) ÷ 5 = 216.4. This benchmark data mean reflects the average level of video playback volume during this relatively stable low period and is a key indicator of the static data characteristics of the dimension.

[0069] To collect dynamic data features for a dimension, the system focuses on periods of data growth for the primary processing dimension within a single data cycle. For example, 7:00 PM to 9:00 PM is peak user activity for a video platform, and video play counts surge. This represents a period of data growth. Suppose that video play counts are recorded every 10 minutes between 7:00 PM and 9:00 PM on a particular day. The initial play count is 5,000. Over time, the play count gradually increases, reaching a peak of 12,000 at 8:30 PM. During this growth process, the peak span change rate is calculated. From 7:00 PM to 8:30 PM, there were nine 10-minute intervals (i.e., nine statistical nodes), with a starting value of 5,000 and a peak value of 12,000. Therefore, the peak span change rate = (12,000 - 5,000) ÷ 9 ≈ 777.78 (here, the average increase in play count per interval). This rate reflects the speed of change in data volume during the growth phase and is a key indicator of the dynamic data features of the dimension.

[0070] After completing the collection of dimensional static data features and dimensional dynamic data features, the system will make judgments based on preset thresholds. Assume that the system presets a baseline mean threshold of 250 and a peak change rate threshold of 1000 based on historical data and platform operation experience. In the above example, the baseline data mean during the early morning trough period on the first and second days did not exceed the preset baseline mean threshold of 250, and the peak change rate of 777.78 from 7 pm to 9 pm did not exceed the peak change rate threshold of 1000. Therefore, the system determined that the data flow stability of the video platform during these two data periods was normal. This means that during this period, the user behavior and data generation of the video platform were within the normal fluctuation range, and the system was operating stably.

[0071] However, if on a certain day, the baseline data average during the early morning off-peak period suddenly reaches 300, exceeding the preset baseline average threshold of 250, this may indicate abnormal video playback behavior, such as malicious volume manipulation, which leads to an abnormal increase in playback volume during the normal off-peak period. Or during the peak period of a certain evening, the peak change rate reaches 1200, exceeding the peak change rate threshold of 1000. This indicates that the playback volume is growing too fast during the period of data growth. It may be that the platform launched a popular event to attract a large number of new users, but there may also be the risk of abnormal data growth, such as data collection errors or network attacks that lead to an increase in false traffic. When these situations occur, the system will determine that there is an abnormality in the stability of the data flow, and then trigger the corresponding early warning mechanism, notifying the platform operators to investigate and handle it to ensure the normal operation of the platform and the accuracy of the data.

[0072] Example 4:

[0073] When collecting the incremental span of path switching frequency for associated processing dimensions within the primary processing dimension, the system monitors the switching of data interaction paths within the associated processing dimensions in real time. As data processing progresses, the data interaction paths in the associated processing dimensions may change. The system records the frequency of path switching and calculates the incremental span of path switching frequency within adjacent time periods. This parameter reflects the frequency and changing trend of path switching in the associated processing dimension.

[0074] When collecting the growth span of data volume continuously accessed by non-identical associated processing dimensions within the primary processing dimension, the system focuses on the volume of data continuously accessed by each associated processing dimension. Over a period of time, the system counts the volume of data continuously accessed by each non-identical associated processing dimension and calculates its growth span. This parameter reflects the contribution of different associated processing dimensions to the growth of the primary processing dimension's data volume and the degree of fluctuation in this growth.

[0075] When collecting the peak incremental span of the primary processing dimension's data density during ongoing data interaction with the associated processing dimension, the system monitors changes in the primary processing dimension's data density in real time. Data density refers to the amount of data per unit space, and the data density of the primary processing dimension changes as data interaction progresses. The system records the peak data density and calculates the incremental span of this peak during this ongoing data interaction. This parameter reflects the changes in the primary processing dimension's data density during this interaction and is crucial for evaluating the system's data processing pressure and performance.

[0076] After collecting these dimension classification processing parameters, the system calculates the classification processing coefficient for the current data dimension according to established rules. The classification processing coefficient threshold is also pre-set based on multiple factors, including the system's historical data, business logic, and performance requirements. If the classification processing coefficient for the current data dimension exceeds the classification processing coefficient threshold, this indicates an abnormal change in the dimension classification processing parameters, potentially preventing the data dimension classification processing logic from executing properly. The system will then determine that the data dimension classification processing logic has failed. Conversely, if the classification processing coefficient for the current data dimension does not exceed the classification processing coefficient threshold, it indicates that the dimension classification processing parameters are within the normal range, and the data dimension classification processing logic is functioning effectively.

[0077] Take an online education platform as an example. This platform holds a vast amount of course data, student learning data, and teacher teaching data, forming a complex multi-source, heterogeneous data stream. Within this platform, the course learning process is set as the primary analysis direction for the current data dimension, which is also known as the core data flow. For example, the entire process from student course selection, entering the course learning interface, watching instructional videos, completing homework, and taking course exams is the primary processing dimension. The data interaction paths involved in related teacher teaching activities (such as teacher preparation, publishing course materials, and grading homework) and student social interactions (such as communication in course discussion forums and collaborative learning groups) are other data interaction paths that cross-correlate with the core data flow and are labeled as auxiliary data flows. Accordingly, the course learning dimension is the primary processing dimension, while the teacher teaching dimension and the student social interaction dimension are associated processing dimensions.

[0078] The dimension classification processing unit begins collecting dimension classification processing parameters. Let's first look at the incremental span of the path switching frequency of the associated processing dimension in the main processing dimension. Assume that within a week, the platform monitors the data interaction path between the teacher teaching dimension and the course learning dimension. In the first week, the average number of times teachers push teaching materials to the course learning dimension (this is a data interaction path) is 5 times a day. In the second week, due to the course update schedule, the average number of times teachers push teaching materials per day becomes 8 times. The incremental span of the path switching frequency is (8-5) ÷ 5 = 0.6. This value reflects the magnitude of the change in the switching frequency of the data interaction path between the associated processing dimension (teacher teaching dimension) and the main processing dimension (course learning dimension).

[0079] Next, let's examine the growth span of continuously accessed data in non-identical associated processing dimensions within the primary processing dimension. Within the course learning dimension (primary processing dimension), the focus is on the continuously accessed data in the student social interaction dimension (non-identical associated processing dimension). For example, within a given course's learning cycle, the number of posts published by students in the course discussion forum (belonging to the student social interaction dimension) in the first three days was 100, 120, and 130, respectively. The number of posts in the last three days was 150, 180, and 200, respectively. Calculate the increase span of the continuous access data volume for the first three days: First, calculate the increase in data volume for two consecutive days. The increase on the second day relative to the first is (120 - 100) ÷ 100 = 0.2, and the increase on the third day relative to the second is (130 - 120) ÷ 120 ≈ 0.083. The average increase span = (0.2 + 0.083) ÷ 2 ≈ 0.142. The same calculation is performed for the next three days: the average increase span = [((180 - 150) ÷ 150) + ((200 - 180) ÷ 180)] ÷ 2 ≈ 0.128. This parameter reflects the contribution of different related processing dimensions (student social interaction dimension) to the data volume growth of the main processing dimension (course learning dimension).

[0080] Finally, the peak incremental span of the data density in the primary processing dimension as data interaction with the associated processing dimension continues. Taking the course learning dimension as an example, during the teacher's grading of assignments (the data interaction between the associated processing dimension and the primary processing dimension), the data density of the course learning dimension (assuming it is measured by the amount of course learning-related data per unit time, such as student assignment submission data and teacher feedback data) will change. Assume that before grading, the peak data density of the course learning dimension in a certain period of time is 100 (the amount of data per unit time). During the grading process, as teacher feedback data continues to increase, the peak data density rises to 150. Then, the peak incremental span of the data density is (150 - 100) ÷ 100 = 0.5. This parameter reflects the degree of change in the data density of the primary processing dimension during this data interaction.

[0081] After collecting these dimension classification processing parameters, the system calculates the classification processing coefficient of the current data dimension according to established rules (the specific calculation rules are set by the system based on business logic, and the detailed formula calculation is not expanded here). Assume that the classification processing coefficient threshold is 0.8. If the calculated classification processing coefficient of the current data dimension is 1.2, which exceeds the classification processing coefficient threshold of 0.8, this indicates that there has been an abnormal change in the dimension classification processing parameters. This may be because the platform recently adjusted the teaching model, resulting in too frequent data interaction between teachers and students, exceeding the processing logic range originally set by the system, and then determining that the data dimension classification processing logic is invalid. At this time, the platform needs to re-evaluate and adjust the data processing strategy to ensure the accuracy and efficiency of data processing. Conversely, if the calculated classification processing coefficient is 0.7, which does not exceed the classification processing coefficient threshold of 0.8, it means that the dimension classification processing parameters are within the normal range, the data dimension classification processing logic can operate effectively, and the platform's teaching data processing process can proceed normally.

[0082] Example 5:

[0083] The multidimensional data fusion unit communicates with the data integration platform and is responsible for dynamically matching real-time data from the primary and associated processing dimensions within the system. During this dynamic matching process, the system compares and correlates real-time data from the primary and associated processing dimensions based on various factors, including data characteristics, timestamps, and business relationships. For example, timestamps ensure matching between data within the same timeframe, while business relationships enable accurate correlation of data with real business connections.

[0084] After completing dynamic matching, the multidimensional data fusion unit generates a cross-dimensional data association map based on the matching results. This map graphically illustrates the data associations between the primary processing dimension and the associated processing dimensions, making the connections between the data more intuitive and clear. By analyzing the cross-dimensional data association map, the system can gain a deeper understanding of the interactions and influences between data from different dimensions.

[0085] Based on the results of the graph analysis, the multidimensional data fusion unit adjusts the priority weights of the core data flow and the auxiliary data flow. If the graph analysis results show that the data in certain associated processing dimensions is closely related to the data in the main processing dimension and has a significant impact on the system's core business, the system may increase the priority weights of the auxiliary data flows corresponding to these associated processing dimensions to prioritize this data during data processing. Conversely, if the data in certain associated processing dimensions is relatively weakly correlated and has less impact on the core business, the system may lower the priority weights of their auxiliary data flows, thereby optimizing the system's data processing resource allocation and improving overall data processing efficiency.

[0086] Take a large-scale logistics and distribution platform as an example. During its operation, it generates massive amounts of heterogeneous data from multiple sources, including order data, vehicle transportation data, and warehouse inventory data. This data is divided into different dimensions for management and analysis. The dimension of the order delivery process serves as the primary processing dimension, while related dimensions such as vehicle scheduling and warehouse inventory management serve as associated processing dimensions. The multidimensional data fusion unit plays a crucial role in this platform. The following details its working process.

[0087] During a peak delivery period, the logistics platform processes a large number of orders simultaneously. The multidimensional data fusion unit dynamically matches real-time data from the primary processing dimension (order delivery) with associated processing dimensions (such as vehicle scheduling and warehouse inventory).

[0088] For data matching in the order delivery dimension and vehicle scheduling dimension, the system will match based on data features such as the order's shipping address, delivery address, estimated delivery time, and the vehicle's current location, cargo capacity, and driving speed. For example, there is a batch of orders sent from warehouse A to area B, and the orders are required to be delivered before 5 pm on the same day. At the same time, there are multiple vehicles near warehouse A with corresponding carrying capacity. The system will associate these orders with qualified vehicles and determine the orders that each vehicle needs to load by comparing the order quantity, cargo weight and volume, and the vehicle's remaining carrying capacity. If vehicle 1 has a remaining carrying capacity of 5 tons, and there are 10 orders from warehouse A to area B weighing no more than 5 tons, the system will match these 10 orders with vehicle 1.

[0089] To match data between order delivery and warehouse inventory, the system associates the product information in the order with the warehouse's inventory data. For example, if an order contains 50 items of item X, the system will search the inventory data of each warehouse for the inventory quantity of item X. If Warehouse A has 80 items of item X in stock, the order is associated with Warehouse A's inventory data.

[0090] After dynamic matching is complete, the multidimensional data fusion unit generates a cross-dimensional data association graph based on the matching results. In this graph, data entities such as orders, vehicles, and warehouses are presented as nodes, and the matching relationships between them are represented by lines. The graph clearly shows which vehicle is transporting a particular order, which warehouse the vehicle is loading the goods from, and the relationship between the inventory of each warehouse and the order demand. For example, the graph shows that order 1 is transported to the destination by vehicle A after loading the goods from warehouse 1, and order 2 is delivered by vehicle B after loading the goods from warehouse 2. The degree of matching between the inventory status of various products in warehouses 1 and 2 and the order demand can be intuitively seen.

[0091] Based on the results of graph analysis, the multidimensional data fusion unit adjusts the priority weights of core and auxiliary data flows. Suppose the graph analysis reveals that due to an emergency in a certain region, demand for certain supplies has surged, leading to a sharp increase in orders destined for that region, while the number of vehicles available for delivery is limited. To prioritize the delivery of these urgent orders, the system increases the priority weights of auxiliary data flows related to these orders (such as the flows that allocate vehicles to these orders in the vehicle scheduling dimension and the flows that prioritize the allocation of supplies to these orders in the warehouse inventory dimension). Specifically, while vehicle scheduling may have previously been allocated on a first-come, first-served basis, priority will now be given to orders destined for the emergency region. Warehouse inventory allocation will also prioritize the supply needs of these urgent orders, reducing the inventory allocation priority for other, less urgent orders. These adjustments ensure the efficient execution of critical tasks (delivery of urgent orders), optimize the data processing and business execution processes of the entire logistics and distribution system, improve logistics and distribution efficiency and service quality, and ensure customer satisfaction.

[0092] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0093] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A multi-dimensional big data real-time analysis and management system, characterized in that: Including data integration platform, the data integration platform communication connections are: The data dimension collection unit is used to divide multi-source heterogeneous data streams into structured data dimensions and unstructured data dimensions according to preset dimensions, and continuously collect data in each dimension. It obtains the data fluctuation amplitude based on the real-time update frequency corresponding to the data dimension, sets the change process of the current data stream as a single data cycle based on the data fluctuation amplitude, collects the static data features of the dimension and the dynamic data features of the dimension, and determines whether there are any abnormalities in the data stream stability based on feature comparison; The dimension classification processing unit is used to set the main analysis direction in the current data dimension as the data core flow direction, and mark other data interaction paths that are cross-related with the data core flow direction as data auxiliary flow directions. At the same time, the dimension where the data core flow direction is located is marked as the main processing dimension, and the dimension where the data auxiliary flow direction is located is marked as the associated processing dimension; the dimension classification processing parameters are collected, substituted into the rule calculation to obtain the classification processing coefficient of the current data dimension, and the classification processing logic of the data dimension is determined to be invalid based on the coefficient comparison.

2. A multi-dimensional big data real-time analysis and management system according to claim 1, characterized in that: The data integration platform also communicates with: The data flow tracking unit is used to uniformly mark abnormal data nodes that do not match the data core flow in the main processing dimension and data nodes that are consistent with the data auxiliary flow but have not completed the complete processing flow in the associated processing dimension as data to be corrected. It obtains the input and output trajectories of the data to be corrected, performs data integrity verification based on trajectory data analysis, and generates real-time correction instructions; The data trend prediction unit is used to collect data cycle parameters and data scale parameters, and evaluate whether there is any deviation in the real-time data trend prediction based on parameter analysis.

3. A multi-dimensional big data real-time analysis and management system according to claim 1, characterized in that: The static data characteristics of the dimension and the dynamic data characteristics of the dimension are respectively the mean of the baseline data during the period of low fluctuation of the main processing dimension data volume in adjacent single data cycles and the peak span change rate during the period of growth of the main processing dimension data volume in a single data cycle.

4. A multi-dimensional big data real-time analysis and management system according to claim 3, characterized in that: If the static data feature of a dimension exceeds the preset baseline mean threshold, or the dynamic data feature of a dimension exceeds the peak change rate threshold, then it is determined that there is an abnormality in the data flow stability; If the static data characteristics of the dimension do not exceed the preset baseline mean threshold, and the dynamic data characteristics of the dimension do not exceed the peak change rate threshold, the data flow stability is determined to be normal.

5. The multi-dimensional big data real-time analysis and management system according to claim 1, characterized in that: Dimension classification processing parameters include the incremental span of the path switching frequency of the associated processing dimension in the main processing dimension, the incremental span of the continuous access data volume of different associated processing dimensions in the main processing dimension, and the peak incremental span of the main processing dimension data density when the main processing dimension is continuously interacting with the associated processing dimension data. If the classification processing coefficient of the current data dimension exceeds the classification processing coefficient threshold, the data dimension classification processing logic is determined to be invalid; if the classification processing coefficient of the current data dimension does not exceed the classification processing coefficient threshold, the data dimension classification processing logic is determined to be valid.

6. A multi-dimensional big data real-time analysis and management system according to claim 2, characterized in that: The input trajectory and output trajectory of the data to be corrected are respectively the ratio of the retention time of the data to be corrected at the entry identification position of the main processing dimension to the corresponding numerical value of the probability of not triggering the core processing process, and the ratio of the retention time of the data to be corrected in the main processing dimension to the processing time of the associated processing dimension.

7. A multi-dimensional big data real-time analysis and management system according to claim 6, characterized in that: If the input trajectory of the data to be corrected exceeds the input trajectory threshold, and the output trajectory of the data to be corrected exceeds the output trajectory threshold, a data redundancy alarm signal is generated; If the input trajectory of the data to be corrected does not exceed the input trajectory threshold, and the output trajectory of the data to be corrected does not exceed the output trajectory threshold, a data flow normal signal is generated.

8. The multi-dimensional big data real-time analysis and management system according to claim 2, characterized in that: The data cycle parameter and data scale parameter are respectively the ratio of the processing time interval of the main processing dimension to the duration of the data volume growth period in adjacent data cycles, and the quantity deviation value between the current total amount of data in any main processing dimension and the preset processing capacity upper limit.

9. A multi-dimensional big data real-time analysis and management system according to claim 8, characterized in that: If the data cycle parameter does not exceed the duration ratio threshold, and the data scale parameter does not exceed the quantity deviation value threshold, then the data trend prediction is determined to be without deviation; if the data cycle parameter exceeds the duration ratio threshold, or the data scale parameter exceeds the quantity deviation value threshold, then the data trend prediction is determined to be biased.

10. The multi-dimensional big data real-time analysis and management system according to claim 1, characterized in that: The data integration platform also communicates with: The multi-dimensional data fusion unit is used to dynamically match the real-time data of the main processing dimension and the associated processing dimension, generate a cross-dimensional data association map based on the matching results, and adjust the priority weights of the data core flow and the data auxiliary flow based on the map analysis results.