Data acquisition method and system based on big data of intelligent operation and maintenance platform
By using a multi-level rule engine and target prediction model to screen and predict data acquisition methods on the intelligent operation and maintenance platform, and using a distributed computing framework for parallel acquisition, the problems of low accuracy and flexibility of data acquisition in the existing technology are solved, and more efficient and reliable data acquisition is achieved.
Patent Information
- Application Number
- CN202411819939.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-11
AI Technical Summary
Existing data acquisition methods are difficult to adapt to diverse data sources and complex operation and maintenance needs, resulting in low accuracy and flexibility in data acquisition.
The big data data acquisition method based on the intelligent operation and maintenance platform is adopted, and the configuration files of multiple data sources are obtained, the target data source is filtered using a multi-level rule engine, and the optimal data acquisition method is predicted in combination with the target prediction model, and tasks are allocated to multiple devices for parallel acquisition using a distributed computing framework.
It improves the accuracy and flexibility of data acquisition, and can automatically select the optimal data acquisition method according to different operation and maintenance needs and scenarios, reduce manual intervention, and improve acquisition efficiency and reliability.
Smart Images

Figure CN119988157A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of data collection technology, and in particular, to a data collection method and system based on big data of an intelligent operation and maintenance platform. Background Art
[0002] With the rapid development of information technology, the amount of data generated by enterprises and organizations has exploded. As an important tool for managing and maintaining large-scale systems, the intelligent operation and maintenance platform needs to efficiently collect and process data from multiple data sources. Data collection is one of the core functions of the intelligent operation and maintenance platform, which directly affects the performance, reliability and security of large-scale systems.
[0003] Most existing data collection methods are based on fixed rules or simple algorithms. Since different data sources and application scenarios have different requirements for data collection, existing data collection methods are difficult to adapt to diverse data sources and complex operation and maintenance requirements, resulting in low data collection accuracy and low flexibility. Summary of the invention
[0004] The embodiments of the present application provide a data collection method and system based on big data of an intelligent operation and maintenance platform, so as to solve the problems of low accuracy and low flexibility of data collection in the prior art.
[0005] In a first aspect, an embodiment of the present application provides a data collection method based on big data of an intelligent operation and maintenance platform, comprising:
[0006] Obtain configuration files of multiple data sources connected to the intelligent operation and maintenance platform, and use a multi-level rule engine to filter out a target data source from the multiple data sources based on the configuration files of the multiple data sources to form a data source set;
[0007] A target prediction model is used to predict the optimal data collection method corresponding to the data source set; wherein the target prediction model introduces a support vector machine, a target clustering algorithm and a long short-term memory network;
[0008] In combination with the distributed computing framework of the intelligent operation and maintenance platform, the data collection tasks of the data source set are allocated to multiple devices in the intelligent operation and maintenance platform according to the optimal data collection method, so that the multiple devices can collect the original data collected by all data sources in the data source set in parallel;
[0009] The original data is processed, and the operation and maintenance results are determined based on the processed data and the operation and maintenance business demand information.
[0010] Optionally, the target prediction model is used to predict the optimal data collection mode corresponding to the data source set; wherein the target prediction model introduces a support vector machine, a target clustering algorithm and a long short-term memory network, including:
[0011] Based on the configuration files of each data source in the data source set, the characteristics of each data source in the data source set are analyzed by using a support vector machine and a target clustering algorithm to obtain an analysis result;
[0012] According to the analysis results, combined with the enterprise operation and maintenance personalized demand information and historical data collection methods, a long short-term memory network is used to predict the optimal data collection method corresponding to the data source set.
[0013] Optionally, the analyzing the characteristics of each data source in the data source set by using a support vector machine and a target clustering algorithm based on the configuration file of each data source in the data source set to obtain the analysis result includes:
[0014] Based on the configuration files of each data source in the data source set, extracting first characteristic information of each data source in the data source set by using metadata extraction technology;
[0015] Based on the first characteristic information of each data source in the data source set, a support vector machine is used to classify the category attributes of all data sources in the data source set to obtain a data source classification result;
[0016] Based on the data source classification results, a target clustering algorithm is used to perform clustering processing on all data sources in each category to obtain a data source clustering result;
[0017] Based on the data source clustering result, a principal component analysis technique is used to select part of the characteristic information from the first characteristic information, and the part of the characteristic information is used as the second characteristic information; based on the second characteristic information, an association rule learning algorithm is used to generate an association relationship between the data sources in each category;
[0018] Based on the association relationship between the data sources, a genetic algorithm is used to optimize the selection of the initial center point of the target clustering algorithm during the clustering process to obtain an optimized data source clustering result;
[0019] Based on the optimized data source clustering results, a characteristic probability model of the data source is constructed using a Bayesian network to form a data source behavior prediction model;
[0020] Generate analysis results based on the configuration files of each data source in the data source set, the first characteristic information of each data source in the data source set, the data source classification results, the data source clustering results, the second characteristic information, the association relationships between the data sources, the optimized data source clustering results and the data source behavior prediction model.
[0021] Optionally, the target clustering algorithm includes a K-means clustering algorithm; based on the data source classification result, clustering all data sources in each category using the target clustering algorithm to obtain a data source clustering result includes:
[0022] Define a characteristic quantification indicator system of a data source, wherein the characteristic quantification indicator system includes the following characteristic quantification indicators: data volume, access delay, update frequency, and security;
[0023] Based on the data source classification result and the characteristic quantification index system, a standardization technology is used to perform standard quantification processing on the first characteristic information of all data sources in each category to obtain characteristic values of all data sources in each category;
[0024] For each category, a hierarchical clustering algorithm is introduced to determine the number of groups of the data source, and the number of groups is used as the value of the initial input parameter of the K-means clustering algorithm;
[0025] An initial center point is selected in a probability weighted manner, a K-means clustering process is performed using a K-means clustering algorithm based on the values of the initial input parameters and the initial center point, and a dynamic adjustment mechanism of the distance metric is introduced in the clustering process to obtain a data source clustering result; the dynamic adjustment mechanism of the distance metric means that different distance metrics are used for data sources of different characteristic types;
[0026] The silhouette coefficient method is applied to evaluate the data source clustering results to obtain a silhouette coefficient value. When the silhouette coefficient value is less than a preset threshold, the numerical value of the initial input parameter is adjusted or the distance metric is optimized, and the K-means clustering process and evaluation operation are repeated until the silhouette coefficient value is greater than or equal to the preset threshold.
[0027] Optionally, a dynamic adjustment mechanism of distance metric is introduced in the clustering process to obtain a data source clustering result; the dynamic adjustment mechanism of distance metric means using different distance metric standards for data sources of different characteristic types, including:
[0028] A distance metric library is constructed by using a plurality of distance metrics, wherein the distance metric library includes the following distance metrics: Euclidean distance, Manhattan distance, cosine similarity distance, and dynamic time warping distance;
[0029] Using a feature engineering method to classify all characteristic quantification indicators in the characteristic quantification indicator system to obtain an initial characteristic classification result; based on the initial characteristic classification result, using an adaptive distance metric selection algorithm to select a distance metric corresponding to each classification in the initial characteristic classification result from the distance metric standard library;
[0030] Assign corresponding weights to all characteristic quantification indicators in each category according to characteristic importance, and dynamically adjust the distance metric standard corresponding to each category according to the weights corresponding to all characteristic quantification indicators in each category to obtain an adjusted distance metric standard corresponding to each category;
[0031] The data source clustering results are determined based on the adjusted distance metrics corresponding to all classifications.
[0032] Optionally, based on the association relationship between the data sources, a genetic algorithm is used to optimize the selection of the initial center point of the target clustering algorithm in the clustering process to obtain an optimized data source clustering result, including:
[0033] Initializing a population to obtain an initial population, wherein each individual in the initial population represents a position of a to-be-determined center point of a group of data sources;
[0034] Calculating the fitness value of each individual in the population based on the fitness function; the population is the initial population in the first iteration process and is the updated population in other iteration processes;
[0035] Performing a selection operation based on the fitness value of each individual to obtain an optimizing population, and performing a crossover operation and a mutation operation on the individuals in the optimizing population to obtain an updated population;
[0036] Determine whether the preset stop iteration condition is reached. If not, re-execute the calculation step, selection operation, crossover operation, mutation operation and judgment step until the preset stop iteration condition is reached; the preset stop iteration condition is reaching the maximum number of iterations or the genetic algorithm converges.
[0037] Optionally, the method of predicting the optimal data collection method corresponding to the data source set using a long short-term memory network based on the analysis result, combined with the enterprise operation and maintenance personalized demand information and the historical data collection method, includes:
[0038] Based on the analysis results and combined with the enterprise's personalized operation and maintenance demand information, a comprehensive demand feature matrix is generated;
[0039] Based on the comprehensive demand feature matrix, the time series analysis method is used to generate the time series features of the historical collection mode;
[0040] Based on the comprehensive demand feature matrix and the time series characteristics of historical collection methods, a long short-term memory network is used to predict the optimal data collection method corresponding to the data source set.
[0041] In a second aspect, the embodiment of the present application provides a data collection system based on big data of an intelligent operation and maintenance platform, including:
[0042] An acquisition and screening module is used to acquire configuration files of multiple data sources connected to the intelligent operation and maintenance platform, and to screen out a target data source from the multiple data sources using a multi-level rule engine based on the configuration files of the multiple data sources to form a data source set;
[0043] A prediction module, used to predict the optimal data collection mode corresponding to the data source set by using a target prediction model; wherein the target prediction model introduces a support vector machine, a target clustering algorithm and a long short-term memory network;
[0044] A classification collection module is used to combine the distributed computing framework of the intelligent operation and maintenance platform to allocate the data collection tasks of the data source set to multiple devices in the intelligent operation and maintenance platform according to the optimal data collection method, so that the multiple devices can collect the original data collected by all data sources in the data source set in parallel;
[0045] The processing and determination module is used to process the original data and determine the operation and maintenance results based on the processed data and the operation and maintenance business demand information.
[0046] In a third aspect, an embodiment of the present application provides a computing device, comprising a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a data collection method based on big data of an intelligent operation and maintenance platform as described in any one of the first aspects.
[0047] In a fourth aspect, an embodiment of the present application provides a computer storage medium storing a computer program. When the computer program is executed by a computer, it implements a data collection method based on big data of an intelligent operation and maintenance platform as described in any one of the first aspects.
[0048] In an embodiment of the present application, a data collection method based on big data of an intelligent operation and maintenance platform is provided, the method comprising: obtaining configuration files of multiple data sources connected to the intelligent operation and maintenance platform, and using a multi-level rule engine to filter out target data sources from multiple data sources based on the configuration files of the multiple data sources to form a data source set; using a target prediction model to predict the optimal data collection method corresponding to the data source set; wherein the target prediction model introduces a support vector machine, a target clustering algorithm, and a long short-term memory network; combining the distributed computing framework of the intelligent operation and maintenance platform, allocating data collection tasks of the data source set to multiple devices in the intelligent operation and maintenance platform according to the optimal data collection method, so that the multiple devices can collect the original data collected by all data sources in the data source set in parallel; processing the original data, and determining the operation and maintenance results based on the processed data and operation and maintenance business demand information.
[0049] This embodiment can flexibly set multiple layers of rules according to different operation and maintenance requirements and scenarios through a multi-level rule engine to ensure the accuracy and applicability of the screening results. The target prediction model in this embodiment combines support vector machines, target clustering algorithms and long short-term memory networks to intelligently predict the optimal data collection method and improve the accuracy and flexibility of data collection. Specifically, this embodiment uses support vector machines to classify and analyze the characteristics of data sources, and can accurately distinguish different types of data sources, such as structured data and unstructured data, static data and dynamic data. This embodiment uses a target clustering algorithm to classify data sources into different groups according to the characteristics of the data source (such as data volume, access delay, etc.), which is convenient for more fine-grained feature analysis. This embodiment uses the time series analysis capability of the long short-term memory network, combined with historical data collection methods and enterprise operation and maintenance personalized needs, to predict the optimal data collection method, adapt to the needs of different data sources and application scenarios, and provide personalized data collection methods.
[0050] These and other aspects of the present application will become more clearly understood in the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0052] Figure 1 A flowchart of a data collection method based on big data of an intelligent operation and maintenance platform provided in an embodiment of the present application;
[0053] Figure 2 A structural diagram of a data acquisition system based on big data of an intelligent operation and maintenance platform provided in an embodiment of the present application;
[0054] Figure 3 A schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0055] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0056] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or executed in parallel. The sequence numbers of the operations, such as S11, S12, etc., are only used to distinguish between different operations, and the sequence numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit the "first" and "second" to be different types.
[0057] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0058] Figure 1 A flowchart of a data collection method based on big data of an intelligent operation and maintenance platform provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the method includes:
[0059] S11. Obtain configuration files of multiple data sources connected to the intelligent operation and maintenance platform, and use a multi-level rule engine to filter out a target data source from the multiple data sources based on the configuration files of the multiple data sources to form a data source set.
[0060] It should be understood that this embodiment uses the automation tools (such as network scanning tools) of the intelligent operation and maintenance platform to automatically identify and configure various data sources by scanning the network environment, reading configuration files, etc. Data sources may include databases, log files, network interfaces, etc.
[0061] Specifically, step S11 may include the following steps: Step 111, use a network scanning tool to scan the network environment of the intelligent operation and maintenance platform to obtain a scanning result including a data source list and its detailed information; the detailed information includes: Internet Protocol (IP) address, port number, user name, password, etc. Step 112, use a configuration file parser (such as regular expression, target parser, etc.) to read and parse the existing configuration file, and use a template engine and programming language (such as Python, Java, etc.) to dynamically generate a new configuration file based on the parsed existing configuration information and scanning results; genetic algorithms can be introduced in the generation process. The configuration file includes: data source type, IP address, port number, authentication information, etc. Among them, the target parser can be a JavaScript Object Notation (JSON) parser, an eXtensible Markup Language (XML) parser, etc. Step 113, use a configuration file verification tool to perform syntax and logic verification on the new configuration file to obtain a verified configuration file.
[0062] It should also be understood that the multi-level rule engine can be designed with a preset rule layer, a dynamic rule layer, and a comprehensive rule layer. Among them, the preset rule layer can have multiple preset rules built in, including screening conditions for common data sources. The dynamic rule layer can use machine learning algorithms to intelligently and dynamically generate appropriate screening rules based on historical data and user behavior data to improve the accuracy and efficiency of screening. The comprehensive rule layer can combine preset rules and screening rules to form comprehensive screening rules to improve the accuracy and efficiency of screening. Exemplarily, for the dynamic rule layer, the intelligent operation and maintenance platform can use machine learning algorithms (such as decision trees, random forests, etc.) to intelligently recommend the most appropriate screening rules based on historical data and user behavior to improve the accuracy and efficiency of screening. Among them, historical data includes screening records, screening results, user feedback, etc. in the past time period. User behavior data includes user screening habits, common rules, preference settings, etc.
[0063] Optionally, the intelligent operation and maintenance platform has a built-in powerful multi-level rule engine, which can filter data sources according to preset rules, comprehensive screening rules, etc. to obtain a data source set. These rules can be based on multiple conditions such as data type, data format, timestamp, keywords, etc. Exemplarily, the data type is a specified database or a specific type of log file. The data format is a log file in a preset format. The timestamp is the data source updated within the last week. The keyword is a specific word or word or symbol or a numerical value within a range.
[0064] S12. Use a target prediction model to predict the optimal data collection method corresponding to a data source set; wherein the target prediction model introduces a support vector machine, a target clustering algorithm, and a long short-term memory network.
[0065] Specifically, this embodiment can use support vector machines and K-means clustering algorithms to conduct in-depth analysis of the characteristics of data sources (such as data type, data volume, network environment, etc.) based on a data source set to obtain a detailed analysis report. Exemplarily, support vector machines are used to classify data source types to ensure that different types of data correspond to different processing methods. K-means clustering algorithms are used to group data sources to obtain data source clustering results. Long short-term memory networks are used to predict the optimal data collection method based on the data source clustering results.
[0066] S13. In combination with the distributed computing framework of the intelligent operation and maintenance platform, the data collection tasks of the data source set are allocated to multiple devices in the intelligent operation and maintenance platform according to the optimal data collection method, so that multiple devices can collect the original data collected by all data sources in the data source set in parallel.
[0067] It should be understood that the intelligent operation and maintenance platform supports multiple data access methods, including application programming interface (API) calls, file reading, message queues, etc. The intelligent operation and maintenance platform can automatically select the optimal data collection method to ensure the efficiency and reliability of data collection. The distributed computing framework of the intelligent operation and maintenance platform can be a distributed computing framework of the type of Apache Spark, Apache Hadoop or Spark, or it can be other types of frameworks. This embodiment does not specifically limit the structure of the framework. Exemplarily, this step can use the Apache Spark distributed computing framework to distribute the data collection tasks for the data in the data source in parallel to multiple devices of the intelligent operation and maintenance platform for execution based on the optimal data collection method predicted in the previous step. Accordingly, using the distributed computing framework, the intelligent operation and maintenance platform can process data collection tasks for multiple data sources in parallel, thereby improving the speed and concurrency of data collection.
[0068] Optionally, this embodiment can combine the adaptive task scheduling algorithm and reinforcement learning technology for the data collection tasks of the data source set, dynamically adjust the task allocation strategy according to the load and network conditions of each device, and obtain a task scheduling result, which can ensure that each data collection task can be efficiently executed on the most suitable device. In addition, based on the task scheduling results, this embodiment can combine the resource management algorithm and the linear programming method to further optimize the allocation of computing resources and obtain an optimized task scheduling result. This embodiment reduces resource waste while improving resource utilization, ensuring the best data collection effect under limited resources.
[0069] S14. Process the original data and determine the operation and maintenance results based on the processed data and operation and maintenance business demand information.
[0070] It should be understood that the purpose of data collection is to support various functions of intelligent operation and maintenance, such as fault detection, performance optimization, resource management, etc. Therefore, after executing the data collection task, this embodiment can transmit the collected raw data to the data processing module. Among them, the intelligent operation and maintenance platform can adopt efficient data transmission protocols (such as Kafka, Flume, etc.) to ensure the stability and speed of data transmission. At the same time, the intelligent operation and maintenance platform supports data compression and breakpoint continuation to ensure the integrity and reliability of data transmission. The data processing module in the intelligent operation and maintenance platform can process the raw data in real time. For example, the stream processing framework (such as Apache Flink, Storm) is used for data cleaning, formatting and preliminary analysis to ensure the immediate availability of data. Another exemplary embodiment can use data cleaning technology to remove invalid or erroneous data in the collected raw data, and then use feature selection algorithm to select the most valuable information. Finally, the anomaly detection algorithm is used to identify potential problems or risk points to provide accurate data support for operation and maintenance decisions.
[0071] Optionally, this embodiment can store the processed data in a designated data warehouse or database. Specifically, the intelligent operation and maintenance platform supports a variety of data storage methods, including relational databases (such as MySQL, PostgreSQL), NoSQL databases (such as MongoDB, Cassandra), data warehouses (such as Hive, Amazon Redshift), etc. The intelligent operation and maintenance platform can automatically select the most suitable storage method based on data characteristics and application scenarios. The intelligent operation and maintenance platform can automatically create data indexes and partitions to improve the performance of data query and analysis. At the same time, the intelligent operation and maintenance platform supports data lifecycle management, automatically archives and deletes expired data, and saves storage space.
[0072] This embodiment can flexibly set multiple layers of rules according to different operation and maintenance requirements and scenarios through a multi-level rule engine to ensure the accuracy and applicability of the screening results. The target prediction model in this embodiment combines support vector machines, target clustering algorithms, and long short-term memory networks to intelligently predict the optimal data collection method and improve the accuracy and flexibility of data collection. Specifically, this embodiment uses support vector machines to classify and analyze the characteristics of data sources, and can accurately distinguish different types of data sources, such as structured data and unstructured data, static data and dynamic data, etc.
[0073] In some possible embodiments, S12, using a target prediction model to predict the optimal data collection method corresponding to the data source set; wherein the target prediction model introduces a support vector machine, a target clustering algorithm and a long short-term memory network, including:
[0074] Step 121: Based on the configuration files of each data source in the data source set, the characteristics of each data source in the data source set are analyzed using a support vector machine and a target clustering algorithm to obtain an analysis result.
[0075] Step 122: Based on the analysis results, combined with the enterprise's personalized operation and maintenance demand information and historical data collection methods, the long short-term memory network is used to predict the optimal data collection method corresponding to the data source set.
[0076] This embodiment uses a target clustering algorithm to classify data sources into different groups (or groups) according to the characteristics of the data source (such as data volume, access latency, etc.), so as to facilitate more fine-grained characteristic analysis. This embodiment uses the time series analysis capability of the long short-term memory network, combines historical data collection methods and enterprise operation and maintenance personalized needs, predicts the optimal data collection method, adapts to the needs of different data sources and application scenarios, and provides personalized data collection methods.
[0077] In the above embodiment, as a possible implementation method, step 121, based on the configuration files of each data source in the data source set, uses a support vector machine and a target clustering algorithm to analyze the characteristics of each data source in the data source set to obtain an analysis result, including:
[0078] Step a1: Based on the configuration files of each data source in the data source set, metadata extraction technology is used to extract the first characteristic information of each data source in the data source set. The first characteristic information includes but is not limited to data type, data volume, update frequency, and network environment parameters, etc., to provide data support for subsequent characteristic analysis.
[0079] Step a2: Based on the first characteristic information of each data source in the data source set, a support vector machine is used to classify the category attributes of all data sources in the data source set to obtain a data source classification result. The classification process is used to evaluate the similarities and differences between different data sources, and the data source classification result is used to distinguish between structured data sources and unstructured data sources, or to distinguish between static data sources and dynamic data sources, etc.
[0080] Step a3: Based on the data source classification results, a target clustering algorithm is used to cluster all data sources in each category to obtain a data source clustering result. The target clustering algorithm includes but is not limited to the K-means clustering algorithm. In the clustering process, data sources can be classified into different groups based on factors such as data volume and access latency, so as to facilitate more detailed feature analysis.
[0081] Step a4: Based on the data source clustering results, principal component analysis technology is used to select part of the characteristic information from the first characteristic information, and part of the characteristic information is used as the second characteristic information; based on the second characteristic information, an association rule learning algorithm is used to generate the association relationship between the data sources in each category. It should be understood that principal component analysis technology removes redundant and irrelevant characteristic information, retains the key characteristics that have the greatest impact on the clustering results, and can reduce the characteristic dimension while keeping the key characteristics of the data source unchanged, so as to improve the analysis efficiency and optimize the performance of the support vector machine and the target clustering algorithm. The association rule learning algorithm can explore the implicit relationships that may exist between data sources, such as some data sources will be updated simultaneously under certain conditions, or certain types of data tend to appear in a specific network environment. These findings help to deeply understand the working mode of the data source.
[0082] Step a5: Based on the correlation between data sources, a genetic algorithm is used to optimize the selection of the initial center point of the target clustering algorithm during the clustering process to obtain the optimized data source clustering result. Using a genetic algorithm to optimize the selection of the initial center during the K-means clustering process can enhance the clustering effect and ensure that the data sources within each group are highly similar, while the data sources between different groups are significantly different.
[0083] Step a6: Based on the optimized data source clustering results, a Bayesian network is used to construct a characteristic probability model of the data source to form a data source behavior prediction model. The Bayesian network is used to construct a characteristic probability model of the data source, and by analyzing the conditional dependency relationship between data sources, the behavior pattern of a data source under specific conditions is predicted, providing a basis for the selection and priority sorting of data sources.
[0084] Step a7, based on the configuration files of each data source in the data source set, the first characteristic information of each data source in the data source set, the data source classification results, the data source clustering results, the second characteristic information, the association between the data sources, the optimized data source clustering results and the data source behavior prediction model, generate analysis results. It should be understood that the analysis results may refer to the characteristic analysis results of the data source. The analysis results include not only the quantitative analysis results of the characteristics of the data source, but also the qualitative evaluation and suggestions of the characteristics, such as which data sources are most suitable for real-time processing and which are more suitable for batch processing, etc., to provide comprehensive data support for decision makers.
[0085] Through the above detailed steps, this embodiment not only deepens the understanding of the characteristic analysis of the data source, but also ensures the scientificity and rationality of the analysis process by introducing multiple algorithms and technical means, and provides more accurate characteristic analysis results of the data source.
[0086] In the above embodiment, the target clustering algorithm includes a K-means clustering algorithm; accordingly, step a3, based on the data source classification result, clustering all data sources in each category using the target clustering algorithm to obtain a data source clustering result, includes:
[0087] Step a31, define a characteristic quantification index system of the data source, the characteristic quantification index system includes the following characteristic quantification indexes: data volume, access delay, update frequency and security. This embodiment can ensure that each characteristic quantification index can accurately reflect the characteristics of the data source and provide accurate data input for subsequent cluster analysis.
[0088] Step a32: Based on the data source classification results and the characteristic quantification index system, the first characteristic information of all data sources in each category is subjected to standard quantification processing by using standardization technology to obtain the characteristic values of all data sources in each category. The first characteristic information is processed by using standardization technology to eliminate the deviation caused by different dimensions, so as to ensure that the K-means clustering algorithm can fairly consider the influence of each characteristic quantification index and improve the accuracy of the clustering results.
[0089] Step a33: For each category, a hierarchical clustering algorithm is introduced to determine the number of groups of the data source, and the number of groups is used as the value of the initial input parameter of the K-means clustering algorithm. The number of groups is also called the number of groups. Specifically, this embodiment can introduce a hierarchical clustering algorithm as a preprocessing step of the K-means clustering algorithm, first preliminarily determine the number of groups of the data source through hierarchical clustering, and then use this number as the input parameter K of the K-means clustering algorithm to solve the problem that the K-means clustering algorithm needs to pre-specify the K value, thereby improving the effectiveness and rationality of clustering.
[0090] Step a34, select the initial center point in a probability weighted manner, use the K-means clustering algorithm to perform the K-means clustering process based on the values of the initial input parameters and the initial center point, and introduce a dynamic adjustment mechanism of the distance metric in the clustering process to obtain the data source clustering result; the dynamic adjustment mechanism of the distance metric means that different distance metrics are used for data sources with different characteristics. It should be understood that selecting the initial center point in a probability weighted manner can reduce the deviation caused by random selection, thereby improving the quality and stability of clustering. Introducing a dynamic adjustment mechanism of the distance metric means: dynamically adjusting the distance metric based on the actual differences between the data source characteristics. For example, the Euclidean distance can be used for high-dimensional characteristics, and the dynamic time warping distance can be used for time series characteristics to meet the needs of different types of data source characteristics and further improve clustering accuracy.
[0091] Specifically, in step a34, a dynamic adjustment mechanism of distance measurement is introduced in the clustering process to obtain the data source clustering result; the dynamic adjustment mechanism of distance measurement means that different distance measurement standards are used for data sources of different characteristic types, including:
[0092] Step a341, using multiple distance metrics to build a distance metric library, the distance metric library includes the following distance metrics: Euclidean distance, Manhattan distance, cosine similarity distance and dynamic time warping distance.
[0093] Step a342, using a feature engineering method to classify all characteristic quantification indicators in the characteristic quantification indicator system to obtain an initial characteristic classification result; based on the initial characteristic classification result, an adaptive distance metric selection algorithm is used to select a distance metric corresponding to each classification in the initial characteristic classification result from a distance metric library. This embodiment introduces multiple distance metrics and designs an adaptive distance metric selection algorithm to ensure that the most appropriate distance metric can be dynamically selected according to the actual characteristics of the data source during the data source clustering process, thereby improving the accuracy and reliability of the data source clustering results.
[0094] Step a343: assign corresponding weights to all characteristic quantitative indicators in each category according to characteristic importance, and dynamically adjust the distance metric corresponding to each category according to the weights corresponding to all characteristic quantitative indicators in each category to obtain the adjusted distance metric corresponding to each category.
[0095] Step a344: Determine the data source clustering result based on the adjusted distance metrics corresponding to all classifications.
[0096] By introducing a dynamic adjustment mechanism for distance metrics, this embodiment can dynamically select the most appropriate distance metric in the process of data source clustering according to the actual characteristics of the data source. Specifically, the mechanism constructs a library of multiple distance metrics (including Euclidean distance, Manhattan distance, cosine similarity distance, and dynamic time warping distance), and classifies the characteristic quantification indicators using feature engineering methods, and selects the most suitable distance metric from the standard library in combination with an adaptive distance metric selection algorithm. In addition, weights are assigned to the characteristic quantification indicators in each classification according to the importance of the characteristics, and the distance metric corresponding to each classification is dynamically adjusted to ensure the accuracy and reliability of the data source clustering results. This embodiment not only improves the accuracy of data source clustering, but also enhances the robustness and generalization ability of the model, so that the data source clustering results are more in line with the characteristics of the actual data, providing a reliable foundation for subsequent data analysis and application.
[0097] Step a35, apply the silhouette coefficient method to evaluate the data source clustering results to obtain the silhouette coefficient value. When the silhouette coefficient value is less than the preset threshold, adjust the value of the initial input parameter or optimize the distance metric, and repeat the K-means clustering process and evaluation operation until the silhouette coefficient value is greater than or equal to the preset threshold.
[0098] It should be understood that after clustering is completed, the silhouette coefficient method is used to evaluate the data source clustering results, and the quality of clustering is measured by calculating the distance ratio of each data source relative to its cluster and other clusters. If the silhouette coefficient is low, consider adjusting the K value or optimizing the distance metric, and re-execute the clustering process until a satisfactory data source clustering result is obtained.
[0099] Optionally, based on the optimized distance metric, this embodiment uses multiple runs, noise tests, and parameter sensitivity analysis methods to test the stability and robustness of the data source clustering results to obtain stable and robust data source clustering results.
[0100] Through the above steps, this embodiment can improve the applicability and accuracy of the K-means clustering algorithm by applying standardization technology, hierarchical clustering algorithm, K-means clustering algorithm and introducing a dynamic adjustment mechanism of distance metric, while ensuring the rationality of the clustering analysis process.
[0101] Specifically, step a5, based on the association relationship between data sources, adopting a genetic algorithm to optimize the selection of the initial center point of the target clustering algorithm during the clustering process to obtain an optimized data source clustering result, includes:
[0102] Step a51, initialize the population to obtain an initial population, each individual in the initial population represents the position of a to-be-determined center point of a group of data sources. The population size can be determined according to the number of data sources in the category and the characteristic dimension.
[0103] Step a52: Calculate the fitness value of each individual in the population based on the fitness function; the population is the initial population in the first iteration and the updated population in other iterations. The fitness function can select the individual with the highest fitness as the current optimal solution based on the compactness and separation of the clusters.
[0104] Step a53, based on the fitness value of each individual, a selection operation is performed to obtain the population being optimized, and the individuals in the population being optimized are subjected to crossover and mutation operations to obtain an updated population. Specifically, this embodiment adopts the crossover operation of the genetic algorithm, and generates new individuals by selecting two individuals with higher fitness to exchange genes, thereby increasing the diversity of the population and obtaining a population after crossover; based on the population after crossover, the mutation operation of the genetic algorithm is adopted to randomly change some genes of certain individuals to prevent the algorithm from converging prematurely, maintain the exploration ability of the population, and obtain an updated population.
[0105] Step a54, determine whether the preset stop iteration condition is reached. If not, re-execute the calculation steps, selection operations, crossover operations, mutation operations and judgment steps in steps a52 to a54 until the preset stop iteration condition is reached; the preset stop iteration condition is reaching the maximum number of iterations or the genetic algorithm converges.
[0106] Accordingly, by using a genetic algorithm to optimize the selection of the initial center point of the target clustering algorithm during the clustering process, this embodiment can improve the quality and stability of the data source clustering results. Specifically, this embodiment initializes the population and calculates the fitness value of each individual based on the fitness function, and selects the individual with the highest fitness as the current optimal solution. Then, the population is continuously optimized through selection, crossover and mutation operations to increase the diversity and exploration ability of the population and prevent the algorithm from converging prematurely. This embodiment not only improves the rationality of the initial center point selection, but also enhances the robustness and accuracy of the clustering algorithm, and finally obtains the optimized data source clustering results. This makes the data source clustering results more stable and reliable, can better reflect the actual correlation between data sources, and provides a solid foundation for subsequent data analysis and application.
[0107] In the above embodiment, as a possible implementation method, step 122, based on the analysis results, combined with the enterprise operation and maintenance personalized demand information and the historical data collection method, adopts the long short-term memory network to predict the optimal data collection method corresponding to the data source set, including:
[0108] Step b1: Based on the analysis results and combined with the enterprise's personalized operation and maintenance demand information, a comprehensive demand feature matrix is generated.
[0109] Step b2: Based on the comprehensive demand feature matrix, a time series analysis method is used to generate time series features of the historical collection method.
[0110] Step b3: Based on the comprehensive demand feature matrix and the time series characteristics of the historical collection method, a long short-term memory network is used to predict the optimal data collection method corresponding to the data source set.
[0111] Based on the analysis results and the enterprise's personalized operation and maintenance demand information, the generation of a comprehensive demand feature matrix ensures that the model can fully consider the specific needs and actual conditions of the enterprise, and improves the pertinence and practicality of the prediction. Based on the comprehensive demand feature matrix, the time series analysis method is used to generate the time series characteristics of the historical collection method, which makes full use of the time dependence and change trend of historical data and enhances the prediction ability of the model. The long short-term memory network improves the accuracy of the prediction data collection method and provides enterprises with more efficient and reliable operation and maintenance support.
[0112] In summary, this embodiment has the following advantages:
[0113] (1) Most existing data collection methods rely on manual configuration and selection of collection methods, and lack an intelligent automatic selection mechanism. The intelligent operation and maintenance platform in this embodiment uses automation tools, support vector machines, and K-means clustering algorithms to automatically select the optimal data collection method based on the characteristics of the data source (such as data type, data volume, network environment, etc.), combined with the enterprise's personalized operation and maintenance demand information and historical data collection methods, thereby reducing manual intervention and improving collection efficiency and reliability.
[0114] (2) Existing data collection methods usually use single-device or multi-device serial processing, which cannot fully utilize the computing resources of multiple devices, resulting in limited data collection speed. The intelligent operation and maintenance platform in this embodiment uses a distributed computing framework to distribute data collection tasks to multiple devices in parallel for execution, greatly improving the speed and concurrency of data collection. This distributed processing method can not only process large-scale data, but also ensure the real-time and high efficiency of data collection.
[0115] (3) Existing data collection methods often lack intelligent recommendation mechanisms. Users need to manually configure and adjust collection parameters, which is prone to errors and inefficient. The intelligent operation and maintenance platform in this embodiment has a built-in multi-level rule engine. The dynamic rule layer in the multi-level rule engine can use machine learning algorithms to intelligently recommend the most appropriate screening rules based on historical data and user behavior data, further improving the accuracy and efficiency of data collection.
[0116] Figure 2 A structural diagram of a data acquisition system based on big data of an intelligent operation and maintenance platform provided in an embodiment of the present application is shown in FIG. Figure 2As shown, the system includes:
[0117] The acquisition and screening module 21 is used to obtain configuration files of multiple data sources connected to the intelligent operation and maintenance platform, and use a multi-level rule engine to screen out a target data source from the multiple data sources based on the configuration files of the multiple data sources to form a data source set.
[0118] The prediction module 22 is used to use the target prediction model to predict the optimal data collection method corresponding to the data source set; wherein the target prediction model introduces a support vector machine, a target clustering algorithm and a long short-term memory network.
[0119] The classification collection module 23 is used to combine the distributed computing framework of the intelligent operation and maintenance platform to assign the data collection tasks of the data source set to multiple devices in the intelligent operation and maintenance platform according to the optimal data collection method, so that multiple devices can collect the original data collected by all data sources in the data source set in parallel.
[0120] The processing and determination module 24 is used to process the original data and determine the operation and maintenance results according to the processed data and the operation and maintenance business demand information.
[0121] Figure 2 The data acquisition system based on big data of intelligent operation and maintenance platform can execute Figure 1 The implementation principle and technical effect of the data collection method based on big data of the intelligent operation and maintenance platform described in the embodiment shown are not repeated here. The specific way in which each module and unit performs operations in the data collection system based on big data of the intelligent operation and maintenance platform in the above embodiment has been described in detail in the embodiment of the method, and will not be elaborated here.
[0122] In one possible design, Figure 2 The data collection system based on big data of the intelligent operation and maintenance platform of the embodiment shown can be implemented as a computing device, such as Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32 .
[0123] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32 .
[0124] The processing component 32 is used to: obtain configuration files of multiple data sources connected to the intelligent operation and maintenance platform, and use a multi-level rule engine to filter out target data sources from multiple data sources based on the configuration files of the multiple data sources to form a data source set; use a target prediction model to predict the optimal data collection method corresponding to the data source set; wherein the target prediction model introduces a support vector machine, a target clustering algorithm, and a long short-term memory network; in combination with the distributed computing framework of the intelligent operation and maintenance platform, the data collection tasks of the data source set are assigned to multiple devices in the intelligent operation and maintenance platform according to the optimal data collection method, so that multiple devices can collect the original data collected by all data sources in the data source set in parallel; process the original data, and determine the operation and maintenance results based on the processed data and operation and maintenance business demand information.
[0125] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented by one or more application-specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), controllers, microcontrollers, microprocessors or other electronic components to perform the above method.
[0126] The storage component 31 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as random access memory (RAM), static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0127] Of course, the computing device may also include other components, such as input / output interfaces, display components, communication components, etc.
[0128] The input / output interface provides an interface between the processing component and the peripheral interface module, which may be an output device, an input device, etc.
[0129] The communication component is configured to facilitate, among other things, wired or wireless communications between the computing device and other devices.
[0130] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.
[0131] The present application also provides a computer storage medium storing a computer program, wherein the computer program can achieve the above-mentioned Figure 1 The data collection method based on big data of the intelligent operation and maintenance platform of the embodiment shown.
[0132] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0133] The system embodiment described above is merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art may understand and implement it without creative work.
[0134] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0135] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A data collection method based on big data of intelligent operation and maintenance platform, characterized in that: include: Obtain configuration files of multiple data sources connected to the intelligent operation and maintenance platform, and use a multi-level rule engine to filter out a target data source from the multiple data sources based on the configuration files of the multiple data sources to form a data source set; A target prediction model is used to predict the optimal data collection method corresponding to the data source set; wherein the target prediction model introduces a support vector machine, a target clustering algorithm and a long short-term memory network; In combination with the distributed computing framework of the intelligent operation and maintenance platform, the data collection tasks of the data source set are allocated to multiple devices in the intelligent operation and maintenance platform according to the optimal data collection method, so that the multiple devices can collect the original data collected by all data sources in the data source set in parallel; The original data is processed, and the operation and maintenance results are determined based on the processed data and the operation and maintenance business demand information.
2. The method according to claim 1, characterized in that The target prediction model is used to predict the optimal data collection method corresponding to the data source set; wherein the target prediction model introduces a support vector machine, a target clustering algorithm and a long short-term memory network, including: Based on the configuration files of each data source in the data source set, the characteristics of each data source in the data source set are analyzed by using a support vector machine and a target clustering algorithm to obtain an analysis result; According to the analysis results, combined with the enterprise operation and maintenance personalized demand information and historical data collection methods, a long short-term memory network is used to predict the optimal data collection method corresponding to the data source set.
3. The method according to claim 2, characterized in that The configuration files of each data source in the data source set are analyzed by using a support vector machine and a target clustering algorithm to obtain analysis results, including: Based on the configuration files of each data source in the data source set, extracting first characteristic information of each data source in the data source set by using metadata extraction technology; Based on the first characteristic information of each data source in the data source set, a support vector machine is used to classify the category attributes of all data sources in the data source set to obtain a data source classification result; Based on the data source classification results, a target clustering algorithm is used to perform clustering processing on all data sources in each category to obtain a data source clustering result; Based on the data source clustering result, a principal component analysis technique is used to select part of the characteristic information from the first characteristic information, and the part of the characteristic information is used as the second characteristic information; based on the second characteristic information, an association rule learning algorithm is used to generate an association relationship between the data sources in each category; Based on the association relationship between the data sources, a genetic algorithm is used to optimize the selection of the initial center point of the target clustering algorithm during the clustering process to obtain an optimized data source clustering result; Based on the optimized data source clustering results, a characteristic probability model of the data source is constructed using a Bayesian network to form a data source behavior prediction model; Generate analysis results based on the configuration files of each data source in the data source set, the first characteristic information of each data source in the data source set, the data source classification results, the data source clustering results, the second characteristic information, the association relationships between the data sources, the optimized data source clustering results and the data source behavior prediction model.
4. The method according to claim 3, characterized in that The target clustering algorithm includes a K-means clustering algorithm; based on the data source classification result, the target clustering algorithm is used to perform clustering processing on all data sources in each category to obtain a data source clustering result, including: Define a characteristic quantification indicator system of a data source, wherein the characteristic quantification indicator system includes the following characteristic quantification indicators: data volume, access delay, update frequency, and security; Based on the data source classification result and the characteristic quantification index system, a standardization technology is used to perform standard quantification processing on the first characteristic information of all data sources in each category to obtain characteristic values of all data sources in each category; For each category, a hierarchical clustering algorithm is introduced to determine the number of groups of the data source, and the number of groups is used as the value of the initial input parameter of the K-means clustering algorithm; An initial center point is selected in a probability weighted manner, a K-means clustering process is performed using a K-means clustering algorithm based on the values of the initial input parameters and the initial center point, and a dynamic adjustment mechanism of the distance metric is introduced in the clustering process to obtain a data source clustering result; the dynamic adjustment mechanism of the distance metric means that different distance metrics are used for data sources of different characteristic types; The silhouette coefficient method is applied to evaluate the data source clustering results to obtain a silhouette coefficient value. When the silhouette coefficient value is less than a preset threshold, the numerical value of the initial input parameter is adjusted or the distance metric is optimized, and the K-means clustering process and evaluation operation are repeated until the silhouette coefficient value is greater than or equal to the preset threshold.
5. The method according to claim 4, characterized in that The dynamic adjustment mechanism of distance measurement is introduced in the clustering process to obtain the data source clustering result; The dynamic adjustment mechanism of the distance metric means that different distance metrics are used for data sources of different feature types, including: A distance metric library is constructed by using a plurality of distance metrics, wherein the distance metric library includes the following distance metrics: Euclidean distance, Manhattan distance, cosine similarity distance, and dynamic time warping distance; Using a feature engineering method to classify all characteristic quantification indicators in the characteristic quantification indicator system to obtain an initial characteristic classification result; based on the initial characteristic classification result, using an adaptive distance metric selection algorithm to select a distance metric corresponding to each classification in the initial characteristic classification result from the distance metric standard library; Assign corresponding weights to all characteristic quantification indicators in each category according to characteristic importance, and dynamically adjust the distance metric standard corresponding to each category according to the weights corresponding to all characteristic quantification indicators in each category to obtain an adjusted distance metric standard corresponding to each category; The data source clustering results are determined based on the adjusted distance metrics corresponding to all classifications.
6. The method according to claim 3, characterized in that Based on the association relationship between the data sources, a genetic algorithm is used to optimize the selection of the initial center point of the target clustering algorithm in the clustering process to obtain an optimized data source clustering result, including: Initializing a population to obtain an initial population, wherein each individual in the initial population represents a position of a to-be-determined center point of a group of data sources; Calculating the fitness value of each individual in the population based on the fitness function; the population is the initial population in the first iteration process and is the updated population in other iteration processes; Performing a selection operation based on the fitness value of each individual to obtain an optimizing population, and performing a crossover operation and a mutation operation on the individuals in the optimizing population to obtain an updated population; Determine whether the preset stop iteration condition is reached. If not, re-execute the calculation step, selection operation, crossover operation, mutation operation and judgment step until the preset stop iteration condition is reached; the preset stop iteration condition is reaching the maximum number of iterations or the genetic algorithm converges.
7. The method according to claim 2, characterized in that The method of using a long short-term memory network to predict the optimal data collection method corresponding to the data source set based on the analysis results, combined with the enterprise operation and maintenance personalized demand information and the historical data collection method, includes: Based on the analysis results and combined with the enterprise's personalized operation and maintenance demand information, a comprehensive demand feature matrix is generated; Based on the comprehensive demand feature matrix, the time series analysis method is used to generate the time series features of the historical collection mode; Based on the comprehensive demand feature matrix and the time series characteristics of historical collection methods, a long short-term memory network is used to predict the optimal data collection method corresponding to the data source set.
8. A data collection system based on big data of intelligent operation and maintenance platform, characterized in that: include: An acquisition and screening module is used to acquire configuration files of multiple data sources connected to the intelligent operation and maintenance platform, and to screen out a target data source from the multiple data sources using a multi-level rule engine based on the configuration files of the multiple data sources to form a data source set; A prediction module, used to predict the optimal data collection mode corresponding to the data source set by using a target prediction model; wherein the target prediction model introduces a support vector machine, a target clustering algorithm and a long short-term memory network; A classification collection module is used to combine the distributed computing framework of the intelligent operation and maintenance platform to allocate the data collection tasks of the data source set to multiple devices in the intelligent operation and maintenance platform according to the optimal data collection method, so that the multiple devices can collect the original data collected by all data sources in the data source set in parallel; The processing and determination module is used to process the original data and determine the operation and maintenance results according to the processed data and the operation and maintenance business demand information.
9. A computing device, characterized in that It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a data collection method based on big data of an intelligent operation and maintenance platform as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a computer, a data collection method based on big data of an intelligent operation and maintenance platform as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method and system for building medical insurance hospitalization fee prediction model
CN108197737A
Intelligent operation and maintenance multi-source data acquisition visual analysis system
CN112884452A
Big data intelligent acquisition method based on deep learning
CN119003849A
Data acquisition and processing platform for internet of things analysis and control
US10838836B1
Implementing and displaying digital transfer instruments
US20240193598A1
Cited By
File digital storage management method
CN120124594A
Algorithm discrimination identification method based on multi-dimensional data association analysis
CN120744584A