Intelligent system based on multi-system data interaction
By designing a multi-system data interaction intelligent system and using data evaluation coefficients to judge data processing methods, the problem of unreasonable resource allocation in traditional data processing architectures is solved, and efficient computing resource utilization and data analysis accuracy is achieved.
Patent Information
- Application Number
- CN202510593252.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional data processing architecture lacks flexible and reasonable computing allocation strategies, resulting in unreasonable resource allocation, high real-time tasks are delayed due to insufficient resources, and low-priority tasks occupy too much resources, resulting in the overall inefficiency of the system.
Design a multi-system data interaction intelligent system, including data acquisition module, data calculation and allocation module, storage module, data analysis module and model training module. Scientifically judge the data processing method through data evaluation coefficients, select appropriate cloud computing or edge computing, and optimize resource allocation.
By accurately calculating the volume coefficient and aging coefficient, scientifically judging the data processing method has greatly improved the utilization efficiency of computing resources, reduced network transmission delay, and improved the accuracy and stability of data analysis.
Smart Images

Figure CN120104290A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to an intelligent system based on multi-system data interaction. Background Art
[0002] As the wave of digitalization sweeps the world, the operations of various organizations and enterprises are highly dependent on the effective use of data. Multi-system data interaction intelligent systems have emerged, and their importance is becoming increasingly prominent, but they are also facing many severe challenges.
[0003] The difference in data size and real-time requirements places complex demands on data processing; For example, high-frequency trading data in the financial sector may generate millions of records per second, and require microsecond-level processing responses to grasp the ever-changing market conditions; and equipment inspection data in the manufacturing industry, although the data volume is relatively small, needs to be regularly collected and analyzed according to the equipment operation cycle. Faced with such diverse data characteristics, the traditional data processing architecture lacks flexible and reasonable computing allocation strategies, which often leads to unreasonable resource allocation, high-real-time tasks are delayed due to insufficient resources, and low-priority tasks occupy too many resources, resulting in low overall system efficiency and failure to fully utilize the performance of hardware resources; Therefore, an intelligent system based on multi-system data interaction is needed to address the problem that the data processing architecture proposed above lacks a flexible and reasonable computing allocation strategy, which often leads to unreasonable resource allocation. Summary of the invention
[0004] The purpose of the present invention is to propose an intelligent system based on multi-system data interaction in order to solve the above problems.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions: An intelligent system based on multi-system data interaction, comprising: Data acquisition module: collects data from different data sources and pre-processes the collected data; Data computing allocation module: After making a preliminary judgment on the data of each data source, the data evaluation coefficient is obtained, and the data processing method of each data source is determined according to the data evaluation coefficient, including cloud computing and edge computing; Storage module: stores data in a distributed manner on nodes, and organizes and models data; Data analysis module: Use model fusion and integration to combine different models, integrate the prediction results of each model, and obtain a prediction model; Model training module: optimizes the model's hyperparameters and uses regularization techniques to prevent the model from overfitting to noise and details in the training data.
[0006] Preferably, the data acquisition module specifically includes: Use data collection tools to collect data from different data sources on a regular basis; Clean, transform and integrate the collected data.
[0007] Preferably, the data calculation and allocation module specifically includes the following contents: The volume coefficient is obtained by performing scale analysis on the data collected from each data source; After analyzing the real-time performance of the data, the timeliness coefficient is obtained; The data evaluation coefficient is obtained by comprehensive analysis of the volume coefficient and the time coefficient.
[0008] Preferably, the method for obtaining the volume coefficient includes the following parts: By parsing the task code, the number of floating-point operation instructions contained therein is analyzed; Establish a computing workload estimation model based on different task types to obtain the predicted computing intensity; Count the bytes of the statistical data source and estimate the peak memory requirements when the task is running; The volume coefficient is obtained by comprehensively analyzing the predicted computing intensity and peak memory demand.
[0009] Preferably, the method for obtaining the aging coefficient includes the following parts: Obtain the size of data received within a preset time period and the standard unit transmission time corresponding to the data, and divide the data size by the standard unit time to obtain the required bandwidth speed; Get the current bandwidth speed, and divide the current bandwidth speed by the required bandwidth speed to get the bandwidth satisfaction; After analyzing the network delay, the delay qualification coefficient is obtained; Obtain the number of data packets of any node within a preset time period and the data size of each data packet, calculate the average of each data packet to obtain the average size, and divide the product of the number of data packets and the average size of the data packets by the unit time to obtain the load rate; According to the data transmission speed and the length of the data transmission path, the unit estimated transmission time of the data is obtained; Obtaining the data transmission speed corresponding to each time interval within the preset time period, and dividing the data transmission speed in each time interval by the unit transmission time to obtain the unit actual time; Preset an allowable fluctuation range of the unit estimated transmission time, compare the unit actual time in each time interval with the allowable fluctuation range of the unit estimated transmission time in turn, and record the unit actual time within the allowable fluctuation range of the unit estimated transmission time as the effective unit time; Count all the valid unit time quantities, and divide all the unit valid time quantities by the unit actual time quantities to obtain the validity degree; The bandwidth satisfaction, load rate and effectiveness are comprehensively analyzed to obtain the time efficiency coefficient.
[0010] Preferably, the method of determining the data processing mode of each data source according to the data evaluation coefficient includes: Two sets of value ranges of data evaluation coefficient thresholds are preset, and the value ranges of the two sets of data evaluation coefficient thresholds correspond to cloud computing and edge computing respectively. The obtained data evaluation coefficient thresholds are compared with the value ranges of the two sets of data evaluation coefficient thresholds to obtain the calculation method corresponding to the data evaluation coefficient.
[0011] Preferably, the method further includes an interactive interface module: Use visualization technology to design visualization charts and graphs to display data analysis results and model prediction results; provide interactive functions; By optimizing natural language processing algorithms and models.
[0012] Preferably, the system also includes a system management module: Monitor and manage hardware resources, and dynamically allocate and schedule resources based on task priorities and resource requirements; Establish a system monitoring mechanism to monitor various performance indicators of the system in real time; Manage software versions, record the functional features, fixed vulnerabilities and improved performance information of different versions; and formulate version upgrade plans.
[0013] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: 1. The present invention obtains the data evaluation coefficient by accurately calculating the volume coefficient and the timeliness coefficient, and then scientifically determines the data processing method; this enables the system to reasonably select cloud computing or edge computing according to the scale and real-time requirements of the data, greatly improving the utilization efficiency of computing resources; edge computing is used after evaluation, which can quickly process and feedback the results and reduce network transmission delays; and for large-scale analysis tasks of historical production data, the powerful computing power of cloud computing is used to ensure efficient completion of the tasks.
[0014] 2. The present invention provides users with powerful intelligent analysis capabilities and excellent interactive experience through data analysis and interactive interfaces; the data analysis module adopts model fusion and integration technology to organically combine multiple models such as decision trees, support vector machines and neural networks; different models have their own strengths, and their prediction results are combined by weighted averaging, voting, etc., which can effectively overcome the limitations of a single model and significantly improve the accuracy and stability of the analysis results. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Further details, features and advantages of the present application are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which: Figure 1 is a flow chart of the present invention; DETAILED DESCRIPTION
[0016] Several embodiments of the present application will be described in more detail below with reference to the accompanying drawings so that those skilled in the art can implement the present application. The present application can be embodied in many different forms and purposes and should not be limited to the embodiments described herein. These embodiments are provided to make the present application comprehensive and complete, and to fully convey the scope of the present application to those skilled in the art. The embodiments do not limit the present application.
[0017] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art to which this application belongs. It will be further understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the relevant art and / or the context of this specification, and will not be interpreted in an idealized or overly formal sense unless explicitly defined as such herein.
[0018] See also Figure 1 As shown, the present invention provides a technical solution: An intelligent system based on multi-system data interaction, comprising: Data acquisition module: collects data from different data sources and pre-processes the collected data; The data acquisition module specifically includes: Use various data collection tools and technologies, such as ETL (Extract, Transform, Load) tools, web crawlers, sensor data collection interfaces, etc. to collect data from different data sources in a scheduled, real-time, or event-triggered manner; For example, data generated by IoT devices may need to be collected in real time to obtain the latest device status information; while data in some business databases can be collected regularly according to business needs; Clean, transform and integrate the collected data; Data cleaning: remove noise and duplicate data from the data, and handle missing values and outliers; For example, some data collected by sensors may produce abnormal values due to equipment failure or environmental interference, and these abnormal data need to be repaired or deleted through data cleaning; Data conversion: Perform operations such as format conversion, encoding conversion, and data standardization on data to make data from different sources have a unified format and specification, which is convenient for subsequent processing and analysis; For example, convert data in different time formats into a standard time format; Data integration: Integrate data from different data sources into a unified data storage to solve data consistency and data redundancy issues; Data format consistency: The data formats generated by different data sources vary greatly; For example, in an enterprise, the customer relationship management (CRM) system used by the sales department may use the "YYYY - MM - DD" format to record the customer's date of birth, while the accounting system of the finance department may use the "MM / DD / YYYY" format; when integrating data, these data in different formats need to be converted into a standard format, such as "YYYY - MM- DD", for subsequent analysis and processing. This is usually achieved by writing data conversion rules and using ETL tools; Data semantic consistency: The same concept may have different definitions and understandings in different systems; For example, in an e-commerce platform, "sales" in the sales report system only refers to the sales amount of goods, while in the financial accounting system, it may include adjustment items such as sales discounts, returns and refunds. In the data integration process, it is necessary to clarify the true meaning of the data in each data source, establish a unified data dictionary, give precise and consistent definitions for key data items such as "sales", and process them according to unified semantics during data extraction, conversion, and loading to ensure that the meaning of data remains unchanged when it flows between different systems. Data update consistency: When the data in the data source changes, the integrated data storage should also be updated promptly and accurately; For example, if an employee's position changes in the enterprise employee information system, the change should be quickly synchronized to other systems integrated with it, such as the human resources management system and office automation system, to ensure the real-time consistency of employee position information in each system. This is usually achieved through data synchronization mechanisms, such as synchronization technology based on database logs or message queues to achieve real-time push and update of data changes; Data computing allocation module: After making a preliminary judgment on the data of each data source, the data evaluation coefficient is obtained, and the data processing method of each data source is determined according to the data evaluation coefficient, including cloud computing and edge computing; The data calculation and allocation module specifically includes the following contents: The volume coefficient is obtained by performing scale analysis on the data collected from each data source; After analyzing the real-time performance of the data, the timeliness coefficient is obtained; The data evaluation coefficient is obtained by comprehensive analysis of the volume coefficient and the time efficiency coefficient; The method of obtaining the volume coefficient includes the following parts: By parsing the task code, the number of floating-point operation instructions contained in it is analyzed; or the FLOPs of the task are counted based on the number of floating-point operations generated when the task is executed as recorded in the historical log; For example, for an existing image filtering processing code, its floating-point operations can be counted through code analysis; Establish a computing workload estimation model based on different task types to obtain the predicted computing intensity; The formula is: ; in is the computational intensity of the prediction; and It is an empirical coefficient determined according to the task type (such as convolution operation For 3.2, sorting task is 1.5); Represents the data size, Indicates the task operation type; Count the bytes of the statistical data source and estimate the peak memory requirements when the task is running; Memory usage is calculated by the formula: ; in Represents the peak memory demand when the task is running, that is, how much memory resources the task will occupy at most during the entire running process; Represents all data objects ( is the size of the ith data object); for example, in a task involving multiple array processing, It can be individual arrays, adding up the memory sizes occupied by these arrays; Refers to the memory overhead of the runtime library. The runtime library is a collection of basic codes that the program depends on when it is running. They will occupy a certain amount of memory. For example, the standard library that the C language program depends on when it is running will generate corresponding memory overhead. This formula can be used to more accurately estimate the memory usage of the task when it is running, thereby quantifying the impact of the data size dimension on the task complexity. The volume coefficient is obtained by comprehensively analyzing the predicted computing intensity and peak memory requirements; The specific analysis is as follows: After normalizing the predicted computing intensity and peak memory demand, the predicted computing intensity and peak memory demand are used as the major and minor semi-axis of the ellipse, respectively, to establish an ellipse model, calculate the area of the ellipse model, and record it as the volume coefficient; The method of obtaining the time efficiency coefficient includes the following parts: Obtain the size of data received within a preset time period and the standard unit transmission time corresponding to the data, and divide the data size by the standard unit time to obtain the required bandwidth speed; Get the current bandwidth speed, and divide the current bandwidth speed by the required bandwidth speed to get the bandwidth satisfaction; After analyzing the network delay, the delay qualification coefficient is obtained; Obtain the number of data packets of any node within a preset time period and the data size of each data packet, calculate the average of each data packet to obtain the average size, and divide the product of the number of data packets and the average size of the data packets by the unit time to obtain the load rate; According to the data transmission speed and the length of the data transmission path, the unit estimated transmission time of the data is obtained; Obtaining the data transmission speed corresponding to each time interval within the preset time period, and dividing the data transmission speed in each time interval by the unit transmission time to obtain the unit actual time; Preset an allowable fluctuation range of the unit estimated transmission time, compare the unit actual time in each time interval with the allowable fluctuation range of the unit estimated transmission time in turn, and record the unit actual time within the allowable fluctuation range of the unit estimated transmission time as the effective unit time; Count all the valid unit time quantities, and divide all the unit valid time quantities by the unit actual time quantities to obtain the validity degree; The bandwidth satisfaction, load rate and effectiveness are comprehensively analyzed to obtain the time efficiency coefficient; After normalizing the bandwidth satisfaction, load rate and effectiveness, the product of the bandwidth satisfaction and effectiveness is used as one right-angled side of the right triangle, the load rate is used as the other right-angled side of the right triangle, and the remaining side is connected to form a complete right triangle. With the center of the side where the load rate is located as the center of the circle and half of the value of the load rate as the radius, a circle is drawn to cut the right triangle, and the remaining right triangle area is calculated and recorded as the time efficiency coefficient; The data processing method of each data source is determined based on the data evaluation coefficient, including: Preset the weight factors of the volume coefficient and the time-effect coefficient, respectively multiply the volume coefficient and the time-effect coefficient with their corresponding weight factors and then sum them up to obtain the data evaluation coefficient; Preset two sets of data evaluation coefficient threshold value ranges, the two sets of data evaluation coefficient threshold value value ranges correspond to cloud computing and edge computing respectively, compare the obtained data evaluation coefficient threshold value with the two sets of data evaluation coefficient threshold value value ranges, and obtain the calculation method corresponding to the data evaluation coefficient; Storage module: stores data in a distributed manner on nodes, and organizes and models data; Data analysis module: It uses model fusion and integration technology to combine multiple different models, such as decision tree, support vector machine and neural network models, and integrates the prediction results of each model through weighted average and voting to obtain a more accurate and stable prediction model; Model training module: optimizes the model's hyperparameters and uses regularization techniques to prevent the model from overfitting to noise and details in the training data; Use various hyperparameter tuning methods, such as grid search, random search, simulated annealing algorithm, genetic algorithm, etc. to optimize the model's hyperparameters; find the optimal hyperparameter settings by evaluating the model's performance under different hyperparameter combinations on the validation data set to give full play to the model's potential and improve the model's accuracy and generalization ability; use regularization techniques, such as L1 regularization, L2 regularization, and Dropout, to prevent the model from overfitting; regularization adds penalty terms to the loss function to limit the size of the model's parameters, making the model simpler and more generalizable, and preventing the model from overfitting to noise and details in the training data; Interactive interface module: Use advanced visualization technologies, such as D3.js and Echarts, to design a variety of visualization charts and graphics, such as bar charts, line charts, pie charts, maps, network diagrams, etc., to intuitively display data analysis results and model prediction results; at the same time, provide interactive functions, users can hover, click, zoom and other operations to view data details and analysis results in depth, and better understand data and models; Improve the system's ability to understand and answer natural language questions by optimizing natural language processing algorithms and models; use technologies such as word vector models, named entity recognition, and semantic role labeling to analyze and understand the natural language input by users, accurately extract key information from questions, and convert it into query statements or task instructions that the system can process, and then return the system's answers to users in a natural and fluent language format; System Management Module: Monitor and manage the system's hardware resources (such as CPU, memory, disk, network, etc.), dynamically allocate and schedule resources according to task priorities and resource requirements to ensure efficient operation of the system; For example, when multiple tasks are running at the same time, CPU time slices and memory space are reasonably allocated according to the importance and urgency of the tasks to avoid resource competition and system overload; Establish a system monitoring mechanism to monitor various system performance indicators in real time, such as the success rate of data collection, data processing delay, model training accuracy, system throughput, etc. When indicators exceed the normal range or abnormal conditions occur, issue alarm information in a timely manner to notify system administrators to handle them. Alarm methods can include email, SMS, instant messaging, etc., to ensure that administrators can respond and solve problems in a timely manner; Manage the software versions of the system, record the functional features of different versions, fixed vulnerabilities, improved performance and other information; formulate a reasonable version upgrade plan, ensure data security and system stability during the upgrade process, and conduct comprehensive testing and verification of the upgraded system to ensure the normal operation of new functions and the original functions are not affected; Take the software version management of the online advertising intelligent system as an example: Version records: Record detailed information about each version, such as: Version 1.0: Features include basic ad delivery functions and simple ad performance statistics. There are some minor bugs, such as occasional loss of ad delivery requests under high concurrency conditions. In terms of performance, the system responds slowly when processing large amounts of ad data; Version 2.0: Fixed the vulnerability of ad delivery request loss under high concurrency. Added the intelligent ad creative recommendation function to recommend appropriate ad creatives based on user portraits and ad history data. Optimized the data processing algorithm in terms of performance, and increased the system response speed by 30%; Version 3.0: Further optimized the advertising strategy algorithm and improved the click-through rate of ads. Fixed some interface display issues reported by users. Added support for new advertising platforms; Version upgrade plan: When preparing to upgrade from version 2.0 to 3.0, make a detailed upgrade plan: First, back up the data to ensure the security of advertising data and user data during the upgrade process; Perform the upgrade during a time period with low system traffic (such as 2 a.m. to 4 a.m.) to reduce the impact on business. After the upgrade is completed, comprehensive testing and verification will be carried out, including functional testing (checking whether the new functions are running normally and whether the original functions are affected), performance testing (verifying whether performance indicators such as system throughput and response time meet expectations), compatibility testing (ensuring that the system runs normally on different browsers and operating systems), etc.; only after the test is passed, the system will be officially put into use to ensure the normal operation of the new functions and the unaffected original functions.
[0019] The above formulas are obtained by collecting a large amount of data and performing software simulation, and a formula close to the actual value is selected. The influencing weight factor and specific coefficient value in the formula are set by technical personnel in this field according to actual conditions, and can be adjusted and modified later.
[0020] The above description of the embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An intelligent system based on multi-system data interaction, characterized in that: include: Data acquisition module: collects data from different data sources and pre-processes the collected data; Data calculation and allocation module: After making a preliminary judgment on the data of each data source, the data evaluation coefficient is obtained, including obtaining the volume coefficient after performing scale analysis on the data collected by each data source, obtaining the timeliness coefficient after analyzing the real-time performance of the data, and obtaining the data evaluation coefficient after comprehensive analysis of the volume coefficient and the timeliness coefficient; the data processing method of each data source is determined according to the data evaluation coefficient, including cloud computing and edge computing; Storage module: stores data in a distributed manner on nodes, and organizes and models data; Data analysis module: Use model fusion and integration to combine different models, integrate the prediction results of each model, and obtain a prediction model; Model training module: optimizes the model's hyperparameters and uses regularization techniques to prevent the model from overfitting to noise and details in the training data.
2. According to claim 1, a multi-system data interaction intelligent system is characterized in that: The data acquisition module specifically includes: Use data collection tools to collect data from different data sources on a regular basis; Clean, transform and integrate the collected data.
3. The multi-system data interaction intelligent system according to claim 1, characterized in that: The method of obtaining the volume coefficient includes the following parts: By parsing the task code, the number of floating-point operation instructions contained therein is analyzed; Establish a computing workload estimation model based on different task types to obtain the predicted computing intensity; Count the bytes of the statistical data source and estimate the peak memory requirements when the task is running; The volume coefficient is obtained by comprehensively analyzing the predicted computing intensity and peak memory demand.
4. The multi-system data interaction intelligent system according to claim 3 is characterized in that: The method of obtaining the time efficiency coefficient includes the following parts: Obtain the size of data received within a preset time period and the standard unit transmission time corresponding to the data, and divide the data size by the standard unit time to obtain the required bandwidth speed; Get the current bandwidth speed, and divide the current bandwidth speed by the required bandwidth speed to get the bandwidth satisfaction; After analyzing the network delay, the delay qualification coefficient is obtained; Obtain the number of data packets of any node within a preset time period and the data size of each data packet, calculate the average of each data packet to obtain the average size, and divide the product of the number of data packets and the average size of the data packets by the unit time to obtain the load rate; According to the data transmission speed and the length of the data transmission path, the unit estimated transmission time of the data is obtained; Obtaining the data transmission speed corresponding to each time interval within the preset time period, and dividing the data transmission speed in each time interval by the unit transmission time to obtain the unit actual time; Preset an allowable fluctuation range of the unit estimated transmission time, compare the unit actual time in each time interval with the allowable fluctuation range of the unit estimated transmission time in turn, and record the unit actual time within the allowable fluctuation range of the unit estimated transmission time as the effective unit time; Count all the valid unit time quantities, and divide all the unit valid time quantities by the unit actual time quantities to obtain the validity degree; The bandwidth satisfaction, load rate and effectiveness are comprehensively analyzed to obtain the time efficiency coefficient.
5. The multi-system data interaction intelligent system according to claim 4 is characterized in that: The data processing method of each data source is determined based on the data evaluation coefficient, including: Two sets of value ranges of data evaluation coefficient thresholds are preset, and the value ranges of the two sets of data evaluation coefficient thresholds correspond to cloud computing and edge computing respectively. The obtained data evaluation coefficient thresholds are compared with the value ranges of the two sets of data evaluation coefficient thresholds to obtain the calculation method corresponding to the data evaluation coefficient.
6. The multi-system data interaction intelligent system according to claim 1, characterized in that: Also includes interactive interface modules: Use visualization technology to design visualization charts and graphs to display data analysis results and model prediction results; provide interactive functions; By optimizing natural language processing algorithms and models.
7. The multi-system data interaction intelligent system according to claim 1, characterized in that: Also includes system management modules: Monitor and manage hardware resources, and dynamically allocate and schedule resources based on task priorities and resource requirements; Establish a system monitoring mechanism to monitor various performance indicators of the system in real time; Manage software versions, record the functional features, fixed vulnerabilities and improved performance information of different versions; and formulate version upgrade plans.
Citation Information
Patent Citations
Calculation method determining method and system
CN108920425A
Security integrated management method and system for industrial internet
CN116418603A
Data distribution method and system based on cloud edge collaboration
CN117997902A
Method and system for improving computing power efficiency
CN118550711A
Intelligent identification algorithm model edge-end fusion type deployment method
CN119865501A