Fusion platform construction method and system based on intelligent information processing and data analysis
By deeply analyzing the needs of business scenarios, building an efficient server cluster and streaming data processing framework, and optimizing data processing indicators, the problem of insufficient real-time analysis capabilities and scalability of traditional platforms in complex business scenarios is solved, and the timeliness and stability of information processing and data analysis is achieved.
Patent Information
- Application Number
- CN202510095204.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-03
AI Technical Summary
The traditional intelligent information processing and data analysis platform lacks real-time analysis capabilities, poor system scalability, and serious data island phenomena, which limits its effective application in complex and changing business scenarios.
By determining the business scenarios, data types, data sources and data volumes of the data to be processed, analyzing the scenario business needs, establishing a server cluster, calculating the cluster load and security coefficients, building a cluster network environment, configuring a streaming data processing framework, optimizing data processing indicators, and building multimodal data preprocessing components and data buffers.
It improves the timeliness and stability of information processing and data analysis, ensures the efficient operation of the data processing framework, adapts to different data processing needs, processes data of different types and sources, and ensures the fluency and stability of data during processing.
Smart Images

Figure CN120086013A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a construction method and system for an intelligent information processing and data analysis integration platform, and belongs to the field of data analysis. Background Art
[0002] An intelligent information processing and data analysis integration platform refers to a comprehensive system integrating advanced information processing technologies and data analysis tools. The design of this platform aims to extract valuable information from a large amount of complex data and process and analyze this information through intelligent means to support key business activities such as decision-making, business optimization, and risk management.
[0003] Traditional intelligent information processing and data analysis platforms usually adopt static data processing methods and isolated system architectures. This method lacks real-time analysis capabilities, has poor system scalability, and serious data island phenomena, restricting its effective application in complex and changeable business scenarios, thus affecting the timeliness and accuracy of decision-making. Summary of the Invention
[0004] The present invention provides a construction method and system for an intelligent information processing and data analysis integration platform, and its main purpose is to improve the timeliness and stability of information processing and data analysis.
[0005] To achieve the above object, a construction method for an intelligent information processing and data analysis integration platform provided by the present invention includes:
[0006] Determine the business scenario of the data to be processed, define the data type, data source, and data volume of the data to be processed, analyze the scenario business requirements of the business scenario through the data type, data source, and data volume, and establish a server cluster for the data to be processed based on the scenario business requirements;
[0007] Calculate the cluster load of the server cluster, identify the security factor of the server cluster, calculate the security coefficient of the server cluster according to the security factor, and construct the cluster network environment of the server cluster in combination with the cluster load and the security coefficient;
[0008] Configure the streaming data processing framework of the server cluster in the cluster network environment, and calculate the data processing metrics of the streaming data processing framework, where the data processing metrics include real-time performance, scalability, fault tolerance, and integration;
[0009] Optimize the streaming data processing framework based on the data processing metrics to obtain an optimized streaming data processing framework. Construct a multi-modal data preprocessing component for the data to be processed based on the data type, data source, and data volume, and establish a data buffer for the data to be processed and the data processing framework.
[0010] Use the multi-modal data preprocessing component to transfer the data to be processed into the data buffer to obtain buffer data. Analyze the streaming processing logic of the buffer data, and based on the streaming processing logic, construct a continuous processing algorithm set for the optimized streaming data processing framework. Continuously process the buffer data based on the continuous processing algorithm set to obtain the final data tuple.
[0011] Optionally, analyzing the scenario business requirements of the business scenario through the data type, data source, and data volume includes:
[0012] Analyze the data complexity of the business scenario based on the data type;
[0013] Define the data source quality of the data source;
[0014] Analyze the data growth status of the business scenario through the data volume;
[0015] Combine the data complexity, the data source quality, and the data growth status to analyze the scenario business requirements of the business scenario.
[0016] Optionally, defining the data source quality of the data source includes:
[0017] Define the data source quality metrics of the data source, where the data source quality metrics include data reliability, data accuracy, data integrity, and data timeliness;
[0018] Determine the metric weights of the data source quality metrics;
[0019] Based on the data reliability, data accuracy, data integrity, data timeliness, and the metric weights, calculate the data source quality of the data source using the following formula:
[0020] Q = 1 / (e -(μR+σA+γI+θT) )
[0021] where Q represents the data source quality of the data source, e represents the exponential function, R represents the data reliability of the data source, μ represents the metric weight of data reliability, A represents the data accuracy of the data source, σ represents the metric weight of data accuracy, I represents the data integrity of the data source, γ represents the metric weight of data integrity, T represents the data timeliness of the data source, and θ represents the metric weight of data timeliness
[0022] Optionally, establishing the server cluster for the data to be processed based on the scenario service requirements includes:
[0023] Determining the data processing performance metrics of the data to be processed according to the scenario service requirements;
[0024] Defining the server cluster type for the data to be processed based on the data processing performance metrics;
[0025] Determining the number of server nodes and the server node configuration of the server cluster type through the scenario service requirements;
[0026] Establishing the cluster topology structure of the server cluster type according to the number of server nodes and the server node configuration;
[0027] Constructing the server cluster for the data to be processed based on the cluster topology structure.
[0028] Optionally, calculating the cluster load of the server cluster includes:
[0029] Analyzing the cluster load metrics of the server cluster;
[0030] Configuring the monitoring devices of the server cluster according to the cluster load metrics;
[0031] Collecting the load test data of the server cluster based on the monitoring devices;
[0032] Calculating the CPU load, memory load, disk I / O load, and network bandwidth load of the server cluster through the load test data;
[0033] Determining the cluster load of the server cluster based on the CPU load, memory load, disk I / O load, and network bandwidth load of the server cluster.
[0034] Optionally, calculating the data processing metrics of the streaming data processing framework includes:
[0035] Obtaining the benchmark test data of the streaming data processing framework;
[0036] Extracting the latency data from the benchmark test data and calculating the average processing latency of the streaming data processing framework based on the latency data;
[0037] Calculating the real-time performance of the streaming data processing framework based on the average processing latency;
[0038] Extracting the node throughput data from the benchmark test data;
[0039] Analyze the node-throughput relationship of the streaming data processing framework according to the node throughput data;
[0040] Calculate the scalability of the streaming data processing framework through the node-throughput relationship;
[0041] Extract the node failure data of the benchmark test data, and analyze the failure integrity coefficient of the streaming data processing framework based on the node failure data;
[0042] Determine the fault tolerance of the streaming data processing framework through the failure integrity coefficient;
[0043] Extract the integrated data of the benchmark test data, and analyze the integration complexity of the streaming data processing framework based on the integrated data;
[0044] Calculate the integration of the streaming data processing framework using the following formula through the integration complexity:
[0045]
[0046] where I represents the integration of the streaming data processing framework, S c represents the c-th dimension of the integration complexity, ω c represents the dimension weight of the c-th dimension of the integration complexity, m represents the number of dimensions of the integration complexity, and P represents the testability score of the streaming data processing framework;
[0047] Combine the real-time performance, scalability, fault tolerance, and integration to determine the data processing metrics of the streaming data processing framework.
[0048] Optionally, the analyzing the failure integrity coefficient of the streaming data processing framework based on the node failure data includes:
[0049] Analyze the failure messages of the streaming data processing framework after a failure based on the node failure data;
[0050] Mark the total number of messages, the number of successfully recovered messages, the number of failed recovered messages, and the number of unique messages of the failure messages;
[0051] Calculate the message delivery rate, message recovery rate, and message uniqueness rate of the streaming data processing framework based on the total number of messages, the number of successfully delivered messages, the number of successfully recovered messages, the number of failed recovered messages, and the number of unique messages;
[0052] Combine the message delivery rate, message recovery rate, and message uniqueness rate, and calculate the failure integrity coefficient of the streaming data processing framework using the following formula:
[0053] F = (MDR * MRR * MUR) / (1 - (1 - MDR) * (1 - MRR) * (1 - MUR))
[0054] Among them, F represents the fault integrity coefficient of the streaming data processing framework, MDR represents the message delivery rate of the streaming data processing framework, MRR represents the message recovery rate of the streaming data processing framework, and MUR represents the message uniqueness rate of the streaming data processing framework.
[0055] Optionally, establishing the data buffer for the data to be processed and the data processing framework includes:[[]]
[0056] Determine the data buffer requirements of the data processing framework;
[0057] Define the data buffer type of the data processing framework according to the data buffer requirements;
[0058] Construct a distributed buffer architecture for the data buffer type;
[0059] Establish an initial data buffer for the data to be processed and the data processing framework through the distributed buffer architecture;
[0060] Simulate a high-load scenario of the initial data buffer;
[0061] Calculate the buffer performance of the initial data buffer based on the high-load scenario;
[0062] When the buffer performance meets the preset buffer performance threshold, use the initial data buffer as the data buffer for the data to be processed and the data processing framework.
[0063] Optionally, constructing the continuous processing algorithm set of the optimized streaming data processing framework based on the streaming processing logic includes:[[]]
[0064] Define the data access algorithm, preprocessing algorithm, core processing algorithm, and data output algorithm of the optimized streaming data processing framework according to the streaming processing logic;
[0065] Analyze the algorithm continuity coefficients of the data access algorithm, preprocessing algorithm, core processing algorithm, and data output algorithm;
[0066] When the algorithm continuity coefficients meet the preset algorithm continuity threshold, integrate the data access algorithm, preprocessing algorithm, core processing algorithm, and data output algorithm into the optimized streaming data processing framework to obtain the continuous processing algorithm set of the optimized streaming data processing framework.
[0067] To solve the above problems, the present invention also provides a system for constructing an intelligent information processing and data analysis fusion platform, the system comprising:
[0068] A server cluster establishment module, configured to determine the business scenario of the data to be processed, define the data type, data source, and data volume of the data to be processed, analyze the scenario business requirements of the business scenario through the data type, data source, and data volume, and establish a server cluster for the data to be processed based on the scenario business requirements;
[0069] A cluster network environment construction module, configured to calculate the cluster load of the server cluster, identify the security factors of the server cluster, calculate the security coefficient of the server cluster according to the security factors, and construct the cluster network environment of the server cluster by combining the cluster load and the security coefficient;
[0070] A data processing index determination module, configured to configure a streaming data processing framework for the server cluster in the cluster network environment, and calculate the data processing indexes of the streaming data processing framework, wherein the data processing indexes include real-time performance, scalability, fault tolerance, and integration;
[0071] A data buffer construction module, configured to optimize the streaming data processing framework based on the data processing indexes to obtain an optimized streaming data processing framework, construct a multi-modal data preprocessing component for the data to be processed based on the data type, data source, and data volume, and establish a data buffer between the data to be processed and the data processing framework;
[0072] A final data tuple generation module, configured to use the multi-modal data preprocessing component to transfer the data to be processed into the data buffer to obtain buffer data, analyze the streaming processing logic of the buffer data, construct a continuous processing algorithm set for the optimized streaming data processing framework based on the streaming processing logic, and perform continuous processing on the buffer data based on the continuous processing algorithm set to obtain final data tuples.
[0073] Compared with the problems described in the background art, in the present invention, by precisely defining the business scenarios, data types, data sources, and data volumes of the data to be processed, we can deeply analyze the business requirements of the scenarios, and accordingly establish an efficient server cluster, calculate the cluster load, and identify security factors, enabling us to calculate the security coefficient of the server cluster. Furthermore, a stable and secure cluster network environment is constructed. In this environment, a streaming data processing framework is configured, and its data processing metrics, including real-time performance, scalability, fault tolerance, and integration, are calculated to ensure the efficient operation of the framework. Through the optimization of the streaming data processing framework, an optimized streaming data processing framework is obtained, which can better adapt to different data processing requirements. The constructed multimodal data preprocessing component can effectively process different types and sources of data, and the establishment of the data buffer ensures the smoothness and stability of the data during the processing. Therefore, the present invention can improve the timeliness and stability of information processing and data analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 FIG. is a schematic flowchart of a method for constructing a platform based on the integration of intelligent information processing and data analysis provided by an embodiment of the present invention;
[0075] Figure 2 FIG. is a schematic diagram of a module for implementing the method for constructing a platform based on the integration of intelligent information processing and data analysis provided by an embodiment of the present invention.
[0076] The implementation, functional features, and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0077] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0078] An embodiment of the present application provides a method for constructing a platform based on the integration of intelligent information processing and data analysis. The execution subject of the method for constructing a platform based on the integration of intelligent information processing and data analysis includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the method for constructing a platform based on the integration of intelligent information processing and data analysis can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc.
[0079] Embodiment 1:
[0080] Refer to Figure 1As shown in the figure, it is a schematic flowchart of a method for constructing a platform based on the integration of intelligent information processing and data analysis provided by an embodiment of the present invention. In this embodiment, the method for constructing a platform based on the integration of intelligent information processing and data analysis includes:
[0081] S1. Determine the business scenario of the data to be processed, define the data type, data source, and data volume of the data to be processed. Analyze the scenario business requirements of the business scenario through the data type, data source, and data volume. Based on the scenario business requirements, establish a server cluster for the data to be processed.
[0082] It should be explained that the business scenario refers to the specific application environment or background where the data will be processed and analyzed. The data type refers to the manifestation form and structure of the data. The data source refers to the source of the data. The data volume refers to the scale of the data to be processed, usually measured in bytes (B), kilobytes (KB), megabytes (MB), gigabytes (GB), terabytes (TB), or larger units.
[0083] Through the data type, data source, and data volume, the present invention can systematically analyze the business requirements of the business scenario and formulate a detailed data processing and analysis plan.
[0084] Specifically, the analysis of the scenario business requirements of the business scenario through the data type, data source, and data volume includes:
[0085] Based on the data type, analyze the data complexity of the business scenario;
[0086] Define the data source quality of the data source;
[0087] Through the data volume, analyze the data growth status of the business scenario;
[0088] Combining the data complexity, the data source quality, and the data growth status, analyze the scenario business requirements of the business scenario.
[0089] Among them, the data complexity refers to the diversity and processing difficulty of data. It includes the degree of data structuring, the complexity of data relationships, the proportion of noise and outliers in the data, and the complexity of the algorithms and technologies required to process the data. The data source quality refers to the reliability, accuracy, integrity, and timeliness of the data provided by the data source. The data growth state refers to the changing trend of the data volume over time, including the growth rate of the data, the growth pattern (linear, exponential, or seasonal), and the expected data scale. The scenario business requirements refer to the specific requirements for data processing and analysis required to achieve business goals in a specific business scenario, including requirements for data collection, storage, processing, analysis, and display, etc.
[0090] Furthermore, defining the data source quality of the data source includes:
[0091] Defining the data source quality indicators of the data source, where the data source quality indicators include data reliability, data accuracy, data integrity, and data timeliness;
[0092] Determining the index weights of the data source quality indicators;
[0093] Based on the data reliability, data accuracy, data integrity, data timeliness, and the index weights, use the following formula to calculate the data source quality of the data source:
[0094] Q = 1 / (e -(μR+σA+γI+θT) )
[0095] Among them, Q represents the data source quality of the data source, e represents the exponential function, R represents the data reliability of the data source, μ represents the index weight of data reliability, A represents the data accuracy of the data source, σ represents the index weight of data accuracy, I represents the data integrity of the data source, γ represents the index weight of data integrity, T represents the data timeliness of the data source, and θ represents the index weight of data timeliness.
[0096] Based on the scenario business requirements, the present invention can establish a server cluster for the data to be processed to establish a server cluster that meets the specific scenario business requirements.
[0097] Specifically, establishing the server cluster for the data to be processed based on the scenario business requirements includes:
[0098] According to the scenario business requirements, determine the data processing performance indicators of the data to be processed;
[0099] Based on the data processing performance indicators, define the server cluster type of the data to be processed;
[0100] Determine the number of server nodes and the server node configuration of the server cluster type based on the described scenario service requirements;
[0101] Establish the cluster topology structure of the server cluster type according to the number of server nodes and the server node configuration;
[0102] Construct the server cluster for the data to be processed based on the cluster topology structure.
[0103] Among them, the data processing performance index refers to the standard for measuring the data processing ability of the server cluster, such as standards like throughput, response time, number of concurrent users, etc. The server cluster type refers to different cluster architectures selected according to service requirements. The number of server nodes refers to the number of servers required based on the expected load and processing capacity. The server node configuration includes configurations such as the number of CPU cores, memory size, storage type and capacity, network interface speed, etc. The cluster topology structure refers to the physical and logical layout of the server cluster. The server cluster refers to a system composed of multiple server nodes.
[0104] S2. Calculate the cluster load of the server cluster, identify the security factors of the server cluster, calculate the security coefficient of the server cluster according to the security factors, and construct the cluster network environment of the server cluster by combining the cluster load and the security coefficient.
[0105] The present invention calculates the cluster load of the server cluster, which can effectively calculate and manage the load of the working server cluster and ensure the high performance and reliability of the cluster.
[0106] Specifically, the calculation of the cluster load of the server cluster includes:
[0107] Analyze the cluster load indicators of the server cluster;
[0108] Configure the monitoring devices of the server cluster according to the cluster load indicators;
[0109] Collect the load test data of the server cluster based on the monitoring devices;
[0110] Calculate the CPU load, memory load, disk I / O load, and network bandwidth load of the server cluster through the load test data;
[0111] Determine the cluster load of the server cluster based on the CPU load, memory load, disk I / O load, and network bandwidth load of the server cluster.
[0112] Among them, the cluster load metrics refer to a series of parameters used to measure the performance and resource usage of a server cluster. The monitoring device refers to the hardware and software tools used to collect and analyze the performance data of the server cluster. The load test data refers to the data on the resource usage of the server cluster collected by the monitoring device within a specific time period. The CPU load refers to the processor usage rate. The memory load refers to the memory usage rate. The disk I / O load refers to the frequency and speed of disk read and write operations. The network bandwidth load refers to the data transmission rate of the network interface. The cluster load refers to the overall resource usage and workload of the server cluster when processing data and executing tasks.
[0113] It should be explained that the security factors refer to various factors that affect the security of the cluster, which can be characteristics or conditions in aspects such as hardware, software, network, configuration, and personnel operations.
[0114] The present invention calculates the security coefficient of the server cluster, which can evaluate the risks of the server cluster and thus prevent them to improve the security of data processing. Among them, the security coefficient refers to the degree of security of the server cluster in processing data.
[0115] Optionally, the present invention combines the cluster load and the security coefficient to construct the cluster network environment of the server cluster, which can construct a cluster network environment that can not only meet the cluster load requirements but also ensure the security coefficient. Among them, the cluster network environment refers to a computing environment composed of multiple servers, and these servers are interconnected through a network and work together to provide characteristics such as high availability, load balancing, failover, and resource sharing.
[0116] S3. Under the cluster network environment, configure the streaming data processing framework of the server cluster and calculate the data processing metrics of the streaming data processing framework, where the data processing metrics include real-time performance, scalability, fault tolerance, and integration.
[0117] It should be explained that the streaming data processing framework refers to a Kafka software architecture used to process and analyze continuous data streams, which can ingest, process, and forward data in real time or near real time.
[0118] The present invention calculates the data processing metrics of the streaming data processing framework, which can be used as the basis for later framework optimization.
[0119] Specifically, calculating the data processing metrics of the streaming data processing framework includes:
[0120] Obtain the benchmark test data of the streaming data processing framework;
[0121] Extract the latency data of the benchmark test data, and calculate the average processing latency of the streaming data processing framework based on the latency data;
[0122] Calculate the real-time performance of the streaming data processing framework based on the average processing latency;
[0123] Extract the node throughput data of the benchmark test data;
[0124] Analyze the node-throughput relationship of the streaming data processing framework according to the node throughput data;
[0125] Calculate the scalability of the streaming data processing framework through the node-throughput relationship;
[0126] Extract the node failure data of the benchmark test data, and analyze the failure integrity coefficient of the streaming data processing framework based on the node failure data;
[0127] Determine the fault tolerance of the streaming data processing framework through the failure integrity coefficient;
[0128] Extract the integration data of the benchmark test data, and analyze the integration complexity of the streaming data processing framework based on the integration data;
[0129] Calculate the integration of the streaming data processing framework using the following formula through the integration complexity:
[0130]
[0131] where I represents the integration of the streaming data processing framework, S c represents the c-th dimension of the integration complexity, ω c represents the dimension weight of the c-th dimension of the integration complexity, m represents the number of dimensions of the integration complexity, and P represents the testability score of the streaming data processing framework;
[0132] Combine the real-time performance, scalability, fault tolerance, and integration to determine the data processing metrics of the streaming data processing framework.
[0133] Among them, the benchmark test data refers to the performance data collected from a series of tests on the streaming data processing framework under specific conditions, including latency, throughput, fault recovery time, and integration-related information. The latency data refers to the record of the time required for data to go from input to output. The average processing latency refers to the average value of all latency data recorded in one or more benchmark tests. The real-time performance refers to how quickly the system can process and respond to the data stream. The node throughput data refers to the ability to process data on a single node or multiple nodes. The node-throughput relationship refers to the trend of the overall throughput of the system as the number of nodes increases. The scalability refers to the ability of the system to linearly or super-linearly increase the throughput as resources (such as nodes) increase without affecting performance. The node failure data refers to the data on how the system responds and recovers in the case of simulated node failures. The fault integrity coefficient refers to the ability of the system to maintain data integrity and system availability after a node failure. The fault tolerance refers to the ability of the system to remain operational and recover data in the face of node failures. The integration data refers to the data and configuration information required for the system to integrate with other systems or services. The integration complexity refers to the difficulty of integrating the streaming data processing framework into an existing system. The integratability refers to the ability of the streaming data processing framework to integrate with other systems. The data processing metrics refer to a set of quantitative metrics used to measure the performance of the streaming data processing framework, including real-time performance, scalability, fault tolerance, and integratability. The dimensions include API compatibility, data format conversion, configuration difficulty, documentation completeness, etc. The dimension weights refer to the importance of different dimensions of the integration complexity in the integratability analysis.
[0134] Further, analyzing the fault integrity coefficient of the streaming data processing framework based on the node failure data includes:
[0135] Analyzing the fault messages of the streaming data processing framework after a fault occurs based on the node failure data;
[0136] Marking the total number of messages, the number of successfully recovered messages, the number of failed recovered messages, and the number of unique messages of the fault messages;
[0137] Calculating the message delivery rate, message recovery rate, and message uniqueness rate of the streaming data processing framework based on the total number of messages, the number of successfully delivered messages, the number of successfully recovered messages, the number of failed recovered messages, and the number of unique messages;
[0138] Combining the message delivery rate, message recovery rate, and message uniqueness rate, and calculating the fault integrity coefficient of the streaming data processing framework using the following formula:
[0139] F = (MDR * MRR * MUR) / (1 - (1 - MDR) * (-MRR) * (1 - MUR))
[0140] Among them, F represents the fault integrity coefficient of the streaming data processing framework, MDR represents the message delivery rate of the streaming data processing framework, MRR represents the message recovery rate of the streaming data processing framework, and MUR represents the message uniqueness rate of the streaming data processing framework.
[0141] Among them, the total number of messages refers to the number of all messages received by the streaming data processing framework during the fault time. The number of successfully recovered messages refers to the number of messages successfully recovered by the system after the fault. The number of failed recovered messages refers to the number of messages that the system fails to recover after the fault. The number of unique messages refers to the number of messages that maintain uniqueness during the processing. The message delivery rate refers to the proportion of messages successfully delivered by the system. The message recovery rate refers to the proportion of messages successfully recovered by the system after the fault occurs. The message uniqueness rate refers to the proportion of unique messages in the system. The fault integrity coefficient is a comprehensive index used to measure the overall performance of the streaming data processing framework in case of faults and the integrity of data.
[0142] S4. Based on the data processing metrics, optimize the streaming data processing framework to obtain an optimized streaming data processing framework. Based on the data type, data source, and data volume, construct a multi-modal data preprocessing component for the data to be processed, and establish a data buffer for the data to be processed and the data processing framework.
[0143] Based on the data processing metrics, the present invention optimizes the streaming data processing framework to obtain an optimized streaming data processing framework, which can better meet business requirements and provide higher performance and reliability. Among them, the optimized streaming data processing framework refers to the framework after optimizing the streaming data processing framework. Exemplarily, the optimization of the streaming data processing framework can be to increase or optimize hardware resources, such as improving CPU, memory, and storage performance, increasing the number of nodes, or adjusting the configuration parameters of the streaming data processing framework, such as cache size, number of threads, batch size, etc.
[0144] Based on the data type, data source, and data volume, the present invention constructs a multi-modal data preprocessing component for the data to be processed, which can construct an efficient multi-modal data preprocessing component and lay a solid foundation for subsequent data processing and analysis work. Among them, the multi-modal data preprocessing component refers to a system or software module specifically designed to process and prepare data sets containing multiple data types (such as text, images, audio, and video).
[0145] Optionally, the data buffer for the data to be processed and the data processing framework established by the present invention can establish an efficient and reliable data buffer, which will serve as a bridge between the data to be processed and the data processing framework to ensure the efficiency and stability of data processing.
[0146] Specifically, establishing the data buffer for the data to be processed and the data processing framework includes:
[0147] Determine the data buffering requirements of the data processing framework;
[0148] Define the data buffer type of the data processing framework according to the data buffering requirements;
[0149] Construct a distributed buffer architecture of the data buffer type;
[0150] Establish an initial data buffer for the data to be processed and the data processing framework through the distributed buffer architecture;
[0151] Simulate a high-load scenario of the initial data buffer;
[0152] Calculate the buffer performance of the initial data buffer based on the high-load scenario;
[0153] When the buffer performance meets the preset buffer performance threshold, use the initial data buffer as the data buffer for the data to be processed and the data processing framework.
[0154] Among them, the data buffering requirements refer to the performance and reliability requirements for the data processing framework when processing data. The data buffer type refers to the buffer implementation method selected according to the data buffering requirements. The distributed buffer architecture refers to a system structure that distributes the buffer on multiple nodes to improve performance, reliability, and scalability. The initial data buffer refers to the first buffer instance established and tested when constructing the distributed buffer architecture. The high-load scenario refers to a situation where a large amount of data flows in and out of the buffer is simulated or actually generated. The buffer performance refers to various performance indicators shown by the buffer in the high-load scenario. The buffer performance threshold refers to the upper or lower limit value of the buffer performance indicators preset to meet the requirements of the data processing framework. The data buffer refers to an intermediate storage area used to temporarily store the data to be processed in the actual data processing process to smooth the difference in the data inflow and outflow rates.
[0155] Optionally, the simulation of the high-load scenario of the initial data buffer can be implemented through Apache JMeter software.
[0156] S5. Use the multi-modal data preprocessing component to transfer the data to be processed into the data buffer to obtain buffer data, analyze the streaming processing logic of the buffer data, based on the streaming processing logic, construct a continuous processing algorithm set for the optimized streaming data processing framework, and continuously process the buffer data based on the continuous processing algorithm set to obtain the final data tuple.
[0157] It should be explained that the buffer data refers to the data that is transferred and stored in the data buffer after being processed by the data preprocessing component. The streaming processing logic refers to a series of rules and algorithms for real-time processing of the data in the buffer, including algorithms such as filtering, aggregation, and expansion.
[0158] Based on the streaming processing logic, the present invention can construct a continuous processing algorithm set for the optimized streaming data processing framework that is efficient, reliable, and easy to maintain.
[0159] Specifically, the construction of the continuous processing algorithm set for the optimized streaming data processing framework based on the streaming processing logic includes:
[0160] Define the data access algorithm, preprocessing algorithm, core processing algorithm, and data output algorithm for the optimized streaming data processing framework according to the streaming processing logic;
[0161] Analyze the algorithm continuity coefficients of the data access algorithm, preprocessing algorithm, core processing algorithm, and data output algorithm;
[0162] When the algorithm continuity coefficient meets the preset algorithm continuity threshold, integrate the data access algorithm, preprocessing algorithm, core processing algorithm, and data output algorithm into the optimized streaming data processing framework to obtain the continuous processing algorithm set for the optimized streaming data processing framework.
[0163] Among them, the data access algorithm is an algorithm responsible for reading the data stream from the data buffer and converting it into a format that can be processed by the streaming processing framework. The preprocessing algorithm is an algorithm for cleaning, formatting, filtering, and converting the original data. The core processing algorithm is an algorithm for implementing business logic, such as window functions, aggregation operations, complex event processing, etc. The data output algorithm is an algorithm responsible for writing the processed data into the target system. The algorithm continuity coefficient is an index for measuring the continuity and stability of the algorithm in the data stream. The algorithm continuity threshold is a threshold for judging whether the algorithm continuity coefficient meets the requirements for integration into the continuous processing algorithm set. The continuous processing algorithm set refers to integrating the data access algorithm, preprocessing algorithm, core processing algorithm, and data output algorithm in the order of the streaming processing logic to form a complete processing flow.
[0164] Based on the continuous processing algorithm set, the present invention continuously processes the buffer data to obtain the final data tuple, which can realize the intelligent information processing and data analysis fusion of the data. Among them, the final data tuple refers to the result data obtained after a series of operations such as conversion, calculation, and aggregation at the end of the data processing process, including key business fields, timestamps, calculation results, statistical information, status information, and metadata.
[0165] Compared with the problems described in the background art, by precisely defining the business scenarios, data types, data sources, and data volumes of the data to be processed, the present invention enables us to deeply analyze the business requirements of the scenarios, and accordingly establish an efficient server cluster, calculate the cluster load and identify security factors, enabling us to calculate the security coefficient of the server cluster, and then build a stable and secure cluster network environment. In this environment, we configure a streaming data processing framework and calculate its data processing metrics, including real-time performance, scalability, fault tolerance, and integration, to ensure the efficient operation of the framework. Through the optimization of the streaming data processing framework, we obtain an optimized streaming data processing framework that can better adapt to different data processing requirements. The constructed multimodal data preprocessing component can effectively process different types and sources of data, and the establishment of the data buffer ensures the fluency and stability of the data during the processing process. Therefore, the present invention can improve the timeliness and stability of information processing and data analysis.
[0166] Embodiment 2:
[0167] As Figure 2 shown, it is a system function module diagram of a system for constructing a platform based on the fusion of intelligent information processing and data analysis according to the present invention.
[0168] The system 200 for constructing a platform based on the fusion of intelligent information processing and data analysis according to the present invention can be installed in an electronic device. According to the functions to be realized, the system for constructing a platform based on the fusion of intelligent information processing and data analysis can include a server cluster establishment module 201, a cluster network environment construction module 202, a data processing metric determination module 203, a data buffer construction module 204, and a final data tuple generation module 205. The modules in the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by the processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.
[0169] In the embodiments of the present invention, the functions of each module / unit are as follows:
[0170] The server cluster establishment module 201 is used to determine the business scenario of the data to be processed, define the data type, data source, and data volume of the data to be processed, analyze the scenario business requirements of the business scenario through the data type, data source, and data volume, and establish a server cluster for the data to be processed based on the scenario business requirements;
[0171] The cluster network environment construction module 202 is used to calculate the cluster load of the server cluster, identify the security factors of the server cluster, calculate the security coefficient of the server cluster according to the security factors, and construct the cluster network environment of the server cluster by combining the cluster load and the security coefficient;
[0172] The data processing index determination module 203 is used to configure a streaming data processing framework for the server cluster in the cluster network environment and calculate the data processing indexes of the streaming data processing framework, where the data processing indexes include real-time performance, scalability, fault tolerance, and integration;
[0173] The data buffer construction module 204 is used to optimize the streaming data processing framework based on the data processing indexes to obtain an optimized streaming data processing framework, construct a multi-modal data preprocessing component for the data to be processed based on the data type, data source, and data volume, and establish a data buffer between the data to be processed and the data processing framework;
[0174] The final data tuple generation module 205 is used to use the multi-modal data preprocessing component to transfer the data to be processed into the data buffer to obtain buffer data, analyze the streaming processing logic of the buffer data, construct a continuous processing algorithm set for the optimized streaming data processing framework based on the streaming processing logic, and continuously process the buffer data based on the continuous processing algorithm set to obtain final data tuples.
[0175] Specifically, each module in the system 200 for constructing a platform based on the integration of intelligent information processing and data analysis in the embodiments of the present invention adopts the same technical means as those in the above-mentioned Figure 1 method for constructing a platform based on the integration of intelligent information processing and data analysis, and can produce the same technical effects, which will not be elaborated here.
[0176] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-mentioned exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.
[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for constructing a fusion platform based on intelligent information processing and data analysis, characterized in that: The method comprises: Determine the business scenario of the data to be processed, define the data type, data source and data volume of the data to be processed, analyze the scenario business requirements of the business scenario through the data type, data source and data volume, and establish a server cluster for the data to be processed based on the scenario business requirements; Calculating the cluster load of the server cluster, identifying the safety factor of the server cluster, calculating the safety factor of the server cluster according to the safety factor, and building a cluster network environment of the server cluster by combining the cluster load and the safety factor; In the cluster network environment, configuring a streaming data processing framework of the server cluster, and calculating data processing indicators of the streaming data processing framework, wherein the data processing indicators include real-time performance, scalability, fault tolerance, and integration; Based on the data processing index, the streaming data processing framework is optimized to obtain an optimized streaming data processing framework, based on the data type, data source and data volume, a multimodal data preprocessing component of the data to be processed is constructed, and a data buffer of the data to be processed and the data processing framework is established; The multimodal data preprocessing component is used to transfer the data to be processed to the data buffer to obtain buffer data, and the streaming processing logic of the buffer data is analyzed. Based on the streaming processing logic, a continuous processing algorithm set of the optimized streaming data processing framework is constructed, and the buffer data is continuously processed based on the continuous processing algorithm set to obtain a final data tuple.
2. The method for constructing a fusion platform based on intelligent information processing and data analysis according to claim 1, characterized in that: The analyzing the scenario business requirements of the business scenario by the data type, data source and data volume includes: Based on the data type, analyzing the data complexity of the business scenario; defining data source quality of said data source; Analyze the data growth status of the business scenario based on the data volume; Analyze the scenario business requirements of the business scenario in combination with the data complexity, the data source quality and the data growth status.
3. The method for constructing a fusion platform based on intelligent information processing and data analysis according to claim 2, characterized in that: The defining the data source quality of the data source includes: Defining data source quality indicators of the data source, wherein the data source quality indicators include data reliability, data accuracy, data integrity and data timeliness; Determining an indicator weight of the data source quality indicator; Based on the data reliability, data accuracy, data integrity, data timeliness and the indicator weights, the data source quality of the data source is calculated using the following formula: Q=1 / (e -(μR+σA+γI+θT) ) Among them, Q represents the data source quality of the data source, e represents the exponential function, R represents the data reliability of the data source, μ represents the indicator weight of data reliability, A represents the data accuracy of the data source, σ represents the indicator weight of data accuracy, I represents the data integrity of the data source, γ represents the indicator weight of data integrity, T represents the data timeliness of the data source, and θ represents the indicator weight of data timeliness.
4. The method for constructing a fusion platform based on intelligent information processing and data analysis according to claim 3, characterized in that: The establishing of the server cluster of the data to be processed based on the scenario business requirements includes: Determine the data processing performance index of the data to be processed according to the business requirements of the scenario; Based on the data processing performance indicator, define the server cluster type of the data to be processed; Determine the number of server nodes and server node configuration of the server cluster type according to the scenario business requirements; Establishing a cluster topology structure of the server cluster type according to the number of server nodes and the configuration of the server nodes; Based on the cluster topology, a server cluster for the data to be processed is constructed.
5. The method for constructing a fusion platform based on intelligent information processing and data analysis according to claim 4, characterized in that: The calculating the cluster load of the server cluster includes: Analyzing cluster load indicators of the server cluster; Configuring monitoring equipment for the server cluster according to the cluster load indicator; Based on the monitoring device, collecting load test data of the server cluster; Calculate the CPU load, memory load, disk I / O load and network bandwidth load of the server cluster through the load test data; The cluster load of the server cluster is determined based on the CPU load, memory load, disk I / O load, and network bandwidth load of the server cluster.
6. The method for constructing a fusion platform based on intelligent information processing and data analysis according to claim 5, characterized in that: The calculating of the data processing index of the streaming data processing framework includes: Obtaining benchmark test data of the streaming data processing framework; Extracting delay data of the benchmark test data, and calculating an average processing delay of the streaming data processing framework based on the delay data; Based on the average processing delay, calculating the real-time performance of the streaming data processing framework; Extracting node throughput data of the benchmark test data; Analyzing a node-throughput relationship of the streaming data processing framework according to the node throughput data; Calculating the scalability of the streaming data processing framework through the node-throughput relationship; Extracting node failure data of the benchmark test data, and analyzing a failure integrity coefficient of the streaming data processing framework based on the node failure data; Determining the fault tolerance of the streaming data processing framework through the fault integrity coefficient; Extracting integration data of the benchmark test data, and analyzing the integration complexity of the streaming data processing framework based on the integration data; The integration complexity is calculated using the following formula: Among them, I represents the integration of the streaming data processing framework, S c represents the cth dimension of integration complexity, ω c represents the dimension weight of the cth dimension of integration complexity, m represents the number of dimensions of integration complexity, and P represents the testability score of the streaming data processing framework; In combination with the real-time performance, scalability, fault tolerance, and integration, the data processing indicators of the streaming data processing framework are determined.
7. The method for constructing a fusion platform based on intelligent information processing and data analysis according to claim 6, characterized in that: The analyzing the failure integrity coefficient of the streaming data processing framework based on the node failure data includes: Based on the node failure data, analyzing the failure message of the streaming data processing framework after the failure occurs; Mark the total number of messages, the number of successful message recovery, the number of failed message recovery, and the number of unique messages of the fault message; Calculate the message delivery rate, message recovery rate and message uniqueness rate of the streaming data processing framework based on the total number of messages, the number of successful message delivery, the number of successful message recovery, the number of failed message recovery and the number of unique messages; Combining the message delivery rate, message recovery rate and message uniqueness rate, the failure integrity coefficient of the streaming data processing framework is calculated using the following formula: F=(MDR*MRR*MUR) / (1-(1-MDR)*(1-MRR)*(1-MUR)) Among them, F represents the fault integrity coefficient of the streaming data processing framework, MDR represents the message delivery rate of the streaming data processing framework, MRR represents the message recovery rate of the streaming data processing framework, and MUR represents the message unique rate of the streaming data processing framework.
8. The method for constructing a fusion platform based on intelligent information processing and data analysis according to claim 7, characterized in that: The step of establishing the data to be processed and the data buffer of the data processing framework comprises: determining data buffering requirements of the data processing framework; Defining a data buffer type of the data processing framework according to the data buffering requirement; Constructing a distributed buffer architecture of the data buffer type; Establishing the data to be processed and the initial data buffer of the data processing framework through the distributed buffer architecture; Simulating a high load scenario of the initial data buffer; Based on the high load scenario, calculating a buffer performance of the initial data buffer; When the buffer performance meets a preset buffer performance threshold, the initial data buffer is used as the data buffer for the data to be processed and the data processing framework.
9. The method for constructing a fusion platform based on intelligent information processing and data analysis according to claim 8, characterized in that: The continuous processing algorithm set for constructing the optimized streaming data processing framework based on the streaming processing logic includes: According to the stream processing logic, define the data access algorithm, preprocessing algorithm, core processing algorithm and data output algorithm of the optimized stream data processing framework; Analyze the algorithm continuity coefficients of the data access algorithm, preprocessing algorithm, core processing algorithm, and data output algorithm; When the algorithm continuity coefficient meets the preset algorithm continuity threshold, the data access algorithm, preprocessing algorithm, core processing algorithm and data output algorithm are integrated into the optimized streaming data processing framework to obtain the continuous processing algorithm set of the optimized streaming data processing framework.
10. A system for building a fusion platform based on intelligent information processing and data analysis, characterized in that: The system comprises: A server cluster establishment module is used to determine the business scenario of the data to be processed, define the data type, data source and data volume of the data to be processed, analyze the scenario business requirements of the business scenario through the data type, data source and data volume, and establish a server cluster for the data to be processed based on the scenario business requirements; A cluster network environment construction module, used to calculate the cluster load of the server cluster, identify the safety factor of the server cluster, calculate the safety factor of the server cluster according to the safety factor, and construct the cluster network environment of the server cluster by combining the cluster load and the safety factor; A data processing index determination module, used to configure the streaming data processing framework of the server cluster under the cluster network environment, and calculate the data processing index of the streaming data processing framework, wherein the data processing index includes real-time performance, scalability, fault tolerance, and integration; A data buffer construction module is used to optimize the streaming data processing framework based on the data processing index to obtain an optimized streaming data processing framework, construct a multimodal data preprocessing component of the data to be processed based on the data type, data source and data volume, and establish a data buffer for the data to be processed and the data processing framework; The final data tuple generation module is used to use the multimodal data preprocessing component to transfer the data to be processed to the data buffer to obtain buffer data, analyze the streaming processing logic of the buffer data, and build a continuous processing algorithm set for the optimized streaming data processing framework based on the streaming processing logic. Based on the continuous processing algorithm set, the buffer data is continuously processed to obtain the final data tuple.
Citation Information
Cited By
Service performance evaluation method, computer program product and electronic device
CN120723342A