A data collection method and system based on big data of intelligent operation and maintenance platform
Through the big data acquisition method of the intelligent operation and maintenance platform, the multi-level rule engine and machine learning algorithm are used to screen data sources and collect data in parallel, solving the problem of insufficient accuracy and flexibility of data acquisition in the existing technology, and achieving efficient and personalized data acquisition.
Patent Information
- Application Number
- CN202411819939.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2044-12-11
AI Technical Summary
Existing data acquisition methods are difficult to adapt to diverse data sources and complex operation and maintenance needs, resulting in low accuracy and flexibility in data acquisition.
The big data acquisition method based on the intelligent operation and maintenance platform is adopted, and the target data source is screened using a multi-level rule engine, combining support vector machines, target clustering algorithms and long-term memory network prediction optimal data acquisition method, and data is collected in parallel through a distributed computing framework.
It improves the accuracy and flexibility of data acquisition, can provide personalized data acquisition methods according to different data sources and application scenarios, and enhances the efficiency and reliability of data acquisition.
Smart Images

Figure CN119988157B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of data acquisition technology, and in particular to a data acquisition method and system based on big data of an intelligent operation and maintenance platform. Background Art
[0002] With the rapid development of information technology, the amount of data generated by enterprises and organizations is exploding. As a crucial tool for managing and maintaining large-scale systems, intelligent operations platforms must efficiently collect and process data from multiple sources. Data collection is one of the core functions of intelligent operations platforms, directly impacting the performance, reliability, and security of large-scale systems.
[0003] Existing data collection methods are mostly based on fixed rules or simple algorithms. Because different data sources and application scenarios have different data collection requirements, existing data collection methods are difficult to adapt to diverse data sources and complex operation and maintenance requirements, resulting in low data collection accuracy and flexibility. Summary of the Invention
[0004] The embodiments of the present application provide a data collection method and system based on big data of an intelligent operation and maintenance platform, so as to solve the problems of low accuracy and low flexibility of data collection in the prior art.
[0005] In a first aspect, an embodiment of the present application provides a data collection method based on big data of an intelligent operation and maintenance platform, comprising:
[0006] Obtain configuration files of multiple data sources connected to the intelligent operation and maintenance platform, and use a multi-level rule engine to filter out a target data source from the multiple data sources based on the configuration files of the multiple data sources to form a data source set;
[0007] A target prediction model is used to predict the optimal data collection method corresponding to the data source set; wherein the target prediction model introduces a support vector machine, a target clustering algorithm, and a long short-term memory network;
[0008] In combination with the distributed computing framework of the intelligent operation and maintenance platform, data collection tasks of the data source set are assigned to multiple devices within the intelligent operation and maintenance platform according to the optimal data collection method, so that the multiple devices can collect the raw data collected by all data sources in the data source set in parallel;
[0009] The raw data is processed, and an operation and maintenance result is determined based on the processed data and the operation and maintenance business demand information.
[0010] Optionally, the target prediction model is used to predict the optimal data collection mode corresponding to the data source set; wherein the target prediction model introduces a support vector machine, a target clustering algorithm and a long short-term memory network, including:
[0011] Based on the configuration files of each data source in the data source set, using a support vector machine and a target clustering algorithm to analyze the characteristics of each data source in the data source set to obtain an analysis result;
[0012] According to the analysis results, combined with the enterprise operation and maintenance personalized demand information and historical data collection methods, a long short-term memory network is used to predict the optimal data collection method corresponding to the data source set.
[0013] Optionally, the analyzing the characteristics of each data source in the data source set using a support vector machine and a target clustering algorithm based on the configuration file of each data source in the data source set to obtain the analysis results includes:
[0014] Extracting first characteristic information of each data source in the data source set using metadata extraction technology based on a configuration file of each data source in the data source set;
[0015] Based on the first characteristic information of each data source in the data source set, a support vector machine is used to classify the category attributes of all data sources in the data source set to obtain a data source classification result;
[0016] Based on the data source classification results, a target clustering algorithm is used to cluster all data sources in each category to obtain a data source clustering result;
[0017] Based on the data source clustering result, a principal component analysis technique is used to select partial characteristic information from the first characteristic information, and the partial characteristic information is used as the second characteristic information; based on the second characteristic information, an association rule learning algorithm is used to generate an association relationship between data sources in each category;
[0018] Based on the association relationship between the data sources, a genetic algorithm is used to optimize the selection of the initial center point of the target clustering algorithm during the clustering process to obtain an optimized data source clustering result;
[0019] Based on the optimized data source clustering results, a characteristic probability model of the data source is constructed using a Bayesian network to form a data source behavior prediction model;
[0020] Generate analysis results based on the configuration files of each data source in the data source set, the first characteristic information of each data source in the data source set, the data source classification results, the data source clustering results, the second characteristic information, the association relationship between the data sources, the optimized data source clustering results and the data source behavior prediction model.
[0021] Optionally, the target clustering algorithm includes a K-means clustering algorithm; and based on the data source classification result, clustering all data sources in each category using the target clustering algorithm to obtain a data source clustering result includes:
[0022] Define a characteristic quantitative indicator system for data sources, which includes the following characteristic quantitative indicators: data volume, access latency, update frequency, and security;
[0023] Based on the data source classification result and the characteristic quantification index system, a standardization technology is used to perform standard quantization processing on the first characteristic information of all data sources in each category to obtain characteristic values of all data sources in each category;
[0024] For each category, a hierarchical clustering algorithm is introduced to determine the number of groups of the data source, and the number of groups is used as the value of the initial input parameter of the K-means clustering algorithm;
[0025] An initial center point is selected using a probability-weighted approach, and a K-means clustering process is performed using a K-means clustering algorithm based on the values of the initial input parameters and the initial center point. A dynamic adjustment mechanism for the distance metric is introduced into the clustering process to obtain a data source clustering result; the dynamic adjustment mechanism for the distance metric means that different distance metrics are used for data sources with different characteristic types;
[0026] The silhouette coefficient method is applied to evaluate the data source clustering results to obtain a silhouette coefficient value. When the silhouette coefficient value is less than a preset threshold, the values of the initial input parameters are adjusted or the distance metric is optimized, and the K-means clustering process and evaluation operation are repeated until the silhouette coefficient value is greater than or equal to the preset threshold.
[0027] Optionally, a dynamic adjustment mechanism of the distance metric is introduced in the clustering process to obtain a data source clustering result; the dynamic adjustment mechanism of the distance metric means using different distance metric standards for data sources of different characteristic types, including:
[0028] A distance metric library is constructed using a plurality of distance metrics, wherein the distance metric library includes the following distance metrics: Euclidean distance, Manhattan distance, cosine similarity distance, and dynamic time warping distance;
[0029] Using a feature engineering method to classify all characteristic quantitative indicators in the characteristic quantitative indicator system to obtain an initial characteristic classification result; based on the initial characteristic classification result, using an adaptive distance metric selection algorithm to select a distance metric corresponding to each category in the initial characteristic classification result from the distance metric standard library;
[0030] Assigning corresponding weights to all characteristic quantitative indicators in each category according to characteristic importance, and dynamically adjusting the distance metric corresponding to each category according to the corresponding weights of all characteristic quantitative indicators in each category to obtain an adjusted distance metric corresponding to each category;
[0031] The data source clustering results are determined based on the adjusted distance metrics corresponding to all categories.
[0032] Optionally, based on the association relationship between the data sources, a genetic algorithm is used to optimize the selection of the initial center point of the target clustering algorithm during the clustering process to obtain an optimized data source clustering result, including:
[0033] Initializing the population to obtain an initial population, wherein each individual in the initial population represents a position of a to-be-determined center point of a group of data sources;
[0034] Calculating the fitness value of each individual in the population based on the fitness function; the population is the initial population in the first iteration process and is the updated population in the subsequent iteration processes;
[0035] Performing a selection operation based on the fitness value of each individual to obtain an optimizing population, and performing a crossover operation and a mutation operation on the individuals in the optimizing population to obtain an updated population;
[0036] Determine whether the preset stop iteration condition is reached. If not, re-execute the calculation step, selection operation, crossover operation, mutation operation and judgment step until the preset stop iteration condition is reached; the preset stop iteration condition is reaching the maximum number of iterations or the genetic algorithm converges.
[0037] Optionally, the method of using a long short-term memory network to predict the optimal data collection method corresponding to the data source set based on the analysis results, combined with the enterprise operation and maintenance personalized demand information and historical data collection methods, includes:
[0038] Based on the analysis results and combined with the enterprise's personalized operation and maintenance demand information, a comprehensive demand feature matrix is generated;
[0039] Based on the comprehensive demand feature matrix, the time series analysis method is used to generate the time series features of the historical collection mode;
[0040] Based on the comprehensive demand feature matrix and the time series characteristics of historical collection methods, a long short-term memory network is used to predict the optimal data collection method corresponding to the data source set.
[0041] In a second aspect, an embodiment of the present application provides a data collection system based on big data of an intelligent operation and maintenance platform, comprising:
[0042] an acquisition and screening module, configured to acquire configuration files of multiple data sources connected to the intelligent operation and maintenance platform, and to screen out a target data source from the multiple data sources using a multi-level rule engine based on the configuration files of the multiple data sources to form a data source set;
[0043] A prediction module, configured to predict the optimal data collection method corresponding to the data source set using a target prediction model; wherein the target prediction model introduces a support vector machine, a target clustering algorithm, and a long short-term memory network;
[0044] A classification collection module is used to combine the distributed computing framework of the intelligent operation and maintenance platform to assign data collection tasks of the data source set to multiple devices in the intelligent operation and maintenance platform according to the optimal data collection method, so that the multiple devices can collect the raw data collected by all data sources in the data source set in parallel;
[0045] The processing and determination module is used to process the original data and determine the operation and maintenance results based on the processed data and the operation and maintenance business demand information.
[0046] In a third aspect, an embodiment of the present application provides a computing device comprising a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a data collection method based on big data of an intelligent operation and maintenance platform as described in any one of the first aspects.
[0047] In a fourth aspect, an embodiment of the present application provides a computer storage medium storing a computer program. When the computer program is executed by a computer, it implements a data collection method based on big data of an intelligent operation and maintenance platform as described in any one of the first aspects.
[0048] In an embodiment of the present application, a data collection method based on big data of an intelligent operation and maintenance platform is provided, the method comprising: obtaining configuration files of multiple data sources connected to the intelligent operation and maintenance platform, and using a multi-level rule engine to filter out target data sources from the multiple data sources based on the configuration files of the multiple data sources to form a data source set; using a target prediction model to predict the optimal data collection method corresponding to the data source set; wherein the target prediction model introduces a support vector machine, a target clustering algorithm and a long short-term memory network; combining the distributed computing framework of the intelligent operation and maintenance platform, allocating data collection tasks of the data source set to multiple devices within the intelligent operation and maintenance platform according to the optimal data collection method, so that the multiple devices can collect the original data collected by all data sources in the data source set in parallel; processing the original data, and determining the operation and maintenance results based on the processed data and operation and maintenance business demand information.
[0049] This embodiment can flexibly set multiple layers of rules according to different operation and maintenance needs and scenarios through a multi-level rule engine to ensure the accuracy and applicability of the screening results. The target prediction model in this embodiment combines support vector machines, target clustering algorithms and long short-term memory networks to intelligently predict the optimal data collection method and improve the accuracy and flexibility of data collection. Specifically, this embodiment uses support vector machines to classify and analyze the characteristics of data sources, and can accurately distinguish different types of data sources, such as structured data and unstructured data, static data and dynamic data, etc. This embodiment uses a target clustering algorithm to classify data sources into different groups according to the characteristics of the data source (such as data volume, access delay, etc.), which facilitates more fine-grained feature analysis. This embodiment uses the time series analysis capability of the long short-term memory network, combined with historical data collection methods and the personalized needs of enterprise operation and maintenance, to predict the optimal data collection method, adapt to the needs of different data sources and application scenarios, and provide personalized data collection methods.
[0050] These and other aspects of the present application will become more readily apparent from the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0052] Figure 1 A flowchart of a data collection method based on big data of an intelligent operation and maintenance platform provided in an embodiment of the present application;
[0053] Figure 2 A schematic diagram of the structure of a data acquisition system based on big data of an intelligent operation and maintenance platform provided in an embodiment of the present application;
[0054] Figure 3 A schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0055] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0056] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. The serial numbers of the operations, such as S11, S12, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this document are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do they limit "first" and "second" to be different types.
[0057] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0058] Figure 1 A flowchart of a data collection method based on big data of an intelligent operation and maintenance platform provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the method includes:
[0059] S11. Obtain configuration files of multiple data sources connected to the intelligent operation and maintenance platform, and use a multi-level rule engine to filter out a target data source from the multiple data sources based on the configuration files of the multiple data sources to form a data source set.
[0060] It should be understood that this embodiment utilizes automated tools (such as network scanning tools) of the intelligent operation and maintenance platform to automatically identify and configure various data sources by scanning the network environment, reading configuration files, etc. Data sources may include databases, log files, network interfaces, etc.
[0061] Specifically, step S11 may include the following steps: Step 111, use a network scanning tool to scan the network environment of the intelligent operation and maintenance platform to obtain a scan result including a data source list and its detailed information; the detailed information includes: Internet Protocol (IP) address, port number, user name, password, etc. Step 112, use a configuration file parser (such as regular expression, target parser, etc.) to read and parse the existing configuration file, and use a template engine and programming language (such as Python, Java, etc.) to dynamically generate a new configuration file based on the parsed existing configuration information and the scan results; a genetic algorithm may be introduced in the generation process. The configuration file includes: data source type, IP address, port number, authentication information, etc. Among them, the target parser can be a JavaScript Object Notation (JSON) parser, an eXtensible Markup Language (XML) parser, etc. Step 113, use a configuration file validation tool to perform syntax and logic validation on the new configuration file to obtain a validated configuration file.
[0062] It should also be understood that the multi-level rule engine can be designed with a preset rule layer, a dynamic rule layer, and a comprehensive rule layer. Among them, the preset rule layer can have multiple preset rules built in, including screening conditions for common data sources. The dynamic rule layer can use machine learning algorithms to intelligently and dynamically generate appropriate screening rules based on historical data and user behavior data to improve the accuracy and efficiency of screening. The comprehensive rule layer can combine preset rules and screening rules to form comprehensive screening rules to improve the accuracy and efficiency of screening. Exemplarily, for the dynamic rule layer, the intelligent operation and maintenance platform can use machine learning algorithms (such as decision trees, random forests, etc.) to intelligently recommend the most appropriate screening rules based on historical data and user behavior to improve the accuracy and efficiency of screening. Among them, historical data includes screening records, screening results, user feedback, etc. in the past time period. User behavior data includes users' screening habits, commonly used rules, preference settings, etc.
[0063] Optionally, the intelligent operation and maintenance platform includes a powerful, multi-layered rules engine that can filter data sources based on preset rules, comprehensive filtering rules, and other criteria to generate a data source set. These rules can be based on various conditions, such as data type, data format, timestamp, and keyword. For example, the data type can be a specified database or a specific type of log file. The data format can be a log file in a preset format. The timestamp can be for data sources updated within the past week. Keywords can be specific characters, words, symbols, or numerical values within a certain range.
[0064] S12. Use a target prediction model to predict the optimal data collection method corresponding to the data source set; wherein the target prediction model introduces a support vector machine, a target clustering algorithm, and a long short-term memory network.
[0065] Specifically, this embodiment can use a support vector machine and a K-means clustering algorithm to conduct an in-depth analysis of the characteristics of the data source (such as data type, data volume, network environment, etc.) based on a data source set to obtain a detailed analysis report. For example, a support vector machine is used to classify data source types to ensure that different types of data correspond to different processing methods. The K-means clustering algorithm is used to group the data sources to obtain data source clustering results. Based on the data source clustering results, a long short-term memory network is used to predict the optimal data collection method.
[0066] S13. In combination with the distributed computing framework of the intelligent operation and maintenance platform, the data collection tasks of the data source set are assigned to multiple devices in the intelligent operation and maintenance platform according to the optimal data collection method, so that the multiple devices can collect the original data collected by all data sources in the data source set in parallel.
[0067] It should be understood that the intelligent operation and maintenance platform supports a variety of data access methods, including Application Programming Interface (API) calls, file reading, message queues, etc. The intelligent operation and maintenance platform can automatically select the optimal data collection method to ensure the efficiency and reliability of data collection. The distributed computing framework of the intelligent operation and maintenance platform can be a distributed computing framework of the type of Apache Spark, Apache Hadoop or Spark, or it can be other types of frameworks. This embodiment does not specifically limit the structure of the framework. Exemplarily, this step can use the Apache Spark distributed computing framework to distribute the data collection tasks for the data in the data source in parallel to multiple devices of the intelligent operation and maintenance platform for execution based on the optimal data collection method predicted in the previous step. Accordingly, using the distributed computing framework, the intelligent operation and maintenance platform can process data collection tasks of multiple data sources in parallel, thereby improving the speed and concurrency of data collection.
[0068] Optionally, this embodiment can combine an adaptive task scheduling algorithm and reinforcement learning technology with data collection tasks for a set of data sources to dynamically adjust the task allocation strategy based on the load and network conditions of each device to obtain a task scheduling result. This task scheduling result can ensure that each data collection task can be efficiently executed on the most suitable device. In addition, based on the task scheduling results, this embodiment can combine resource management algorithms and linear programming methods to further optimize the allocation of computing resources and obtain an optimized task scheduling result. This embodiment improves resource utilization while reducing resource waste, ensuring the best data collection effect with limited resources.
[0069] S14. Process the original data and determine the operation and maintenance results based on the processed data and operation and maintenance business demand information.
[0070] It should be understood that the purpose of data collection is to support various functions of intelligent operation and maintenance, such as fault detection, performance optimization, resource management, etc. Therefore, after completing the data collection task, this embodiment can transmit the collected raw data to the data processing module. Among them, the intelligent operation and maintenance platform can adopt an efficient data transmission protocol (such as Kafka, Flume, etc.) to ensure the stability and speed of data transmission. At the same time, the intelligent operation and maintenance platform supports data compression and breakpoint resumption to ensure the integrity and reliability of data transmission. The data processing module in the intelligent operation and maintenance platform can process the raw data in real time. For example, a stream processing framework (such as Apache Flink, Storm) is used to clean, format and perform preliminary analysis of the data to ensure the immediate availability of the data. Another example is that this embodiment can use data cleaning technology to remove invalid or erroneous data from the collected raw data, and then use a feature selection algorithm to select the most valuable information. Finally, potential problems or risk points are identified through anomaly detection algorithms to provide accurate data support for operation and maintenance decisions.
[0071] Optionally, this embodiment can store the processed data in a designated data warehouse or database. Specifically, the intelligent operation and maintenance platform supports a variety of data storage methods, including relational databases (such as MySQL, PostgreSQL), NoSQL databases (such as MongoDB, Cassandra), data warehouses (such as Hive, Amazon Redshift), etc. The intelligent operation and maintenance platform can automatically select the most suitable storage method based on data characteristics and application scenarios. The intelligent operation and maintenance platform can automatically create data indexes and partitions to improve the performance of data query and analysis. At the same time, the intelligent operation and maintenance platform supports data lifecycle management, automatically archives and deletes expired data, and saves storage space.
[0072] This embodiment uses a multi-level rule engine to flexibly set multiple layers of rules based on different operation and maintenance needs and scenarios, ensuring the accuracy and applicability of the screening results. The target prediction model in this embodiment combines support vector machines, target clustering algorithms, and long-term short-term memory networks to intelligently predict the optimal data collection method and improve the accuracy and flexibility of data collection. Specifically, this embodiment uses support vector machines to classify and analyze the characteristics of data sources, and can accurately distinguish different types of data sources, such as structured data and unstructured data, static data and dynamic data, etc.
[0073] In some possible embodiments, S12, a target prediction model is used to predict an optimal data collection method corresponding to a data source set; wherein the target prediction model introduces a support vector machine, a target clustering algorithm, and a long short-term memory network, including:
[0074] Step 121 : Based on the configuration files of each data source in the data source set, the characteristics of each data source in the data source set are analyzed using a support vector machine and a target clustering algorithm to obtain analysis results.
[0075] Step 122: Based on the analysis results, combined with the enterprise's personalized operation and maintenance demand information and historical data collection methods, a long short-term memory network is used to predict the optimal data collection method corresponding to the data source set.
[0076] This embodiment uses a target clustering algorithm to categorize data sources into different groups (or groups) based on their characteristics (such as data volume and access latency), facilitating more fine-grained feature analysis. This embodiment leverages the time series analysis capabilities of long-short-term memory networks, combined with historical data collection methods and the personalized needs of enterprise operations and maintenance, to predict the optimal data collection method, adapting to the needs of different data sources and application scenarios, and providing personalized data collection methods.
[0077] In the above embodiment, as a possible implementation, step 121, based on the configuration files of each data source in the data source set, uses a support vector machine and a target clustering algorithm to analyze the characteristics of each data source in the data source set to obtain analysis results, including:
[0078] Step a1: Based on the configuration files of each data source in the data source set, metadata extraction technology is used to extract first characteristic information of each data source in the data source set. This first characteristic information includes, but is not limited to, data type, data size, update frequency, and network environment parameters, providing data support for subsequent characteristic analysis.
[0079] Step a2: Based on the first characteristic information of each data source in the data source set, a support vector machine is used to classify the category attributes of all data sources in the data source set to obtain a data source classification result. The classification process is used to evaluate the similarities and differences between different data sources. The data source classification result is used to distinguish structured data sources from unstructured data sources, or to distinguish static data sources from dynamic data sources.
[0080] In step a3, based on the data source classification results, a target clustering algorithm is used to cluster all data sources in each category to obtain data source clustering results. Target clustering algorithms include, but are not limited to, the K-means clustering algorithm. During the clustering process, data sources can be categorized into different groups based on factors such as data volume and access latency, facilitating more detailed feature analysis.
[0081] Step a4: Based on the data source clustering results, principal component analysis technology is used to select part of the feature information from the first feature information, and the part of the feature information is used as the second feature information; based on the second feature information, an association rule learning algorithm is used to generate the association relationship between the data sources in each category. It should be understood that principal component analysis technology eliminates redundant and irrelevant feature information and retains the key features that have the greatest impact on the clustering results. It can reduce the feature dimension while keeping the key features of the data source unchanged to improve analysis efficiency and optimize the performance of the support vector machine and target clustering algorithm. The association rule learning algorithm can explore the implicit relationships that may exist between data sources, such as certain data sources will be updated simultaneously under certain conditions, or certain types of data tend to appear in specific network environments. These findings help to gain a deeper understanding of the working mode of the data source.
[0082] Step a5: Based on the relationships between data sources, a genetic algorithm is used to optimize the selection of initial centers during the clustering process of the target clustering algorithm to obtain optimized data source clustering results. Using a genetic algorithm to optimize the selection of initial centers during the K-means clustering process can enhance the clustering effect, ensuring that the data sources within each group are highly similar, while the data sources between different groups are significantly different.
[0083] Step a6: Based on the optimized data source clustering results, a Bayesian network is used to construct a probability model of the data source characteristics to form a data source behavior prediction model. This Bayesian network is used to construct a probability model of the data source characteristics. By analyzing the conditional dependencies between data sources, the behavior pattern of a data source under specific conditions is predicted, providing a basis for data source selection and prioritization.
[0084] Step a7: Generate analysis results based on the configuration files of each data source in the data source set, the first characteristic information of each data source in the data source set, the data source classification results, the data source clustering results, the second characteristic information, the associations between data sources, the optimized data source clustering results, and the data source behavior prediction model. It should be understood that the analysis results may refer to the analysis results of the data source characteristics. These analysis results include not only quantitative analysis results of the data source characteristics, but also qualitative evaluations and recommendations, such as which data sources are most suitable for real-time processing and which are more suitable for batch processing, providing comprehensive data support for decision makers.
[0085] Through the above detailed steps, this embodiment not only deepens the understanding of the characteristic analysis of the data source, but also ensures the scientificity and rationality of the analysis process by introducing multiple algorithms and technical means, and provides more accurate characteristic analysis results of the data source.
[0086] In the above embodiment, the target clustering algorithm includes the K-means clustering algorithm; accordingly, step a3, based on the data source classification result, clustering all data sources in each category using the target clustering algorithm to obtain the data source clustering result, includes:
[0087] Step a31: Define a data source characteristic quantification indicator system. This characteristic quantification indicator system includes the following characteristic quantification indicators: data size, access latency, update frequency, and security. This embodiment ensures that each characteristic quantification indicator accurately reflects the characteristics of the data source, providing accurate data input for subsequent cluster analysis.
[0088] Step a32: Based on the data source classification results and the characteristic quantification index system, standardization techniques are used to perform standardized quantization processing on the first characteristic information of all data sources in each category, thereby obtaining characteristic values for all data sources in each category. Using standardization techniques to process each piece of first characteristic information eliminates bias caused by different dimensions, ensuring that the K-means clustering algorithm fairly considers the influence of each characteristic quantification index and improving the accuracy of the clustering results.
[0089] Step a33: For each category, a hierarchical clustering algorithm is introduced to determine the number of groups of the data source, and the number of groups is used as the value of the initial input parameter of the K-means clustering algorithm. The number of groups is also called the number of groups. Specifically, this embodiment can introduce a hierarchical clustering algorithm as a preprocessing step of the K-means clustering algorithm. The number of groups of the data source is first preliminarily determined through hierarchical clustering, and this number is then used as the input parameter K of the K-means clustering algorithm. This solves the problem that the K-means clustering algorithm requires pre-specified K values, thereby improving the effectiveness and rationality of clustering.
[0090] Step a34, select the initial center point in a probability weighted manner, use the K-means clustering algorithm to perform the K-means clustering process based on the values of the initial input parameters and the initial center point, and introduce a dynamic adjustment mechanism of the distance metric in the clustering process to obtain the data source clustering result; the dynamic adjustment mechanism of the distance metric means that different distance metrics are used for data sources with different characteristics. It should be understood that selecting the initial center point in a probability weighted manner can reduce the deviation caused by random selection, thereby improving the quality and stability of clustering. Introducing a dynamic adjustment mechanism of the distance metric means: dynamically adjusting the distance metric based on the actual differences between the data source characteristics. For example, Euclidean distance can be used for high-dimensional characteristics, while dynamic time warping distance can be used for time series characteristics to meet the needs of different types of data source characteristics and further improve clustering accuracy.
[0091] Specifically, in step a34, a dynamic adjustment mechanism of the distance metric is introduced into the clustering process to obtain the data source clustering result. The dynamic adjustment mechanism of the distance metric means that different distance metrics are used for data sources with different characteristic types, including:
[0092] Step a341: construct a distance metric standard library using multiple distance metrics, where the distance metric standard library includes the following distance metrics: Euclidean distance, Manhattan distance, cosine similarity distance, and dynamic time warping distance.
[0093] Step a342: Use feature engineering methods to classify all characteristic quantification indicators in the characteristic quantification indicator system to obtain initial characteristic classification results. Based on the initial characteristic classification results, use an adaptive distance metric selection algorithm to select a distance metric corresponding to each category in the initial characteristic classification results from a distance metric library. This embodiment introduces multiple distance metrics and designs an adaptive distance metric selection algorithm to ensure that the most appropriate distance metric can be dynamically selected during the data source clustering process based on the actual characteristics of the data source, thereby improving the accuracy and reliability of the data source clustering results.
[0094] Step a343: assign corresponding weights to all characteristic quantitative indicators in each category according to characteristic importance, and dynamically adjust the distance metric corresponding to each category according to the corresponding weights of all characteristic quantitative indicators in each category to obtain the adjusted distance metric corresponding to each category.
[0095] Step a344: Determine the data source clustering result based on the adjusted distance metrics corresponding to all categories.
[0096] By introducing a dynamic adjustment mechanism for distance metrics, this embodiment can dynamically select the most appropriate distance metric during the data source clustering process based on the actual characteristics of the data source. Specifically, the mechanism constructs a library of multiple distance metric standards (including Euclidean distance, Manhattan distance, cosine similarity distance, and dynamic time warping distance), and uses feature engineering methods to classify feature quantification indicators, and combines the adaptive distance metric selection algorithm to select the most suitable distance metric from the standard library. In addition, weights are assigned to feature quantification indicators within each category based on feature importance, and the distance metric corresponding to each category is dynamically adjusted to ensure the accuracy and reliability of the data source clustering results. This embodiment not only improves the accuracy of data source clustering, but also enhances the robustness and generalization ability of the model, so that the data source clustering results are more in line with the characteristics of actual data, providing a reliable foundation for subsequent data analysis and application.
[0097] Step a35: Apply the silhouette coefficient method to evaluate the data source clustering results to obtain a silhouette coefficient value. If the silhouette coefficient value is less than a preset threshold, adjust the values of the initial input parameters or optimize the distance metric, and repeat the K-means clustering process and evaluation operation until the silhouette coefficient value is greater than or equal to the preset threshold.
[0098] It should be understood that after clustering is complete, the Silhouette Coefficient method is used to evaluate the data source clustering results. The quality of the clustering is measured by calculating the distance ratio of each data source relative to its cluster and other clusters. If the Silhouette Coefficient is low, consider adjusting the K value or optimizing the distance metric and re-running the clustering process until satisfactory data source clustering results are obtained.
[0099] Optionally, based on the optimized distance metric, this embodiment uses multiple runs, noise testing, and parameter sensitivity analysis methods to test the stability and robustness of the data source clustering results to obtain stable and robust data source clustering results.
[0100] Through the above steps, this embodiment can improve the applicability and accuracy of the K-means clustering algorithm by applying standardization technology, hierarchical clustering algorithm, K-means clustering algorithm and introducing a dynamic adjustment mechanism of distance measurement, while ensuring the rationality of the clustering analysis process.
[0101] Specifically, step a5, based on the association relationship between data sources, uses a genetic algorithm to optimize the selection of the initial center point of the target clustering algorithm during the clustering process to obtain an optimized data source clustering result, including:
[0102] Step a51: Initialize the population to obtain an initial population. Each individual in the initial population represents the position of a to-be-determined center point of a group of data sources. The population size can be determined based on the number of data sources in the category and the characteristic dimension.
[0103] Step a52: Calculate the fitness value of each individual in the population based on the fitness function. The population is the initial population in the first iteration and the updated population in subsequent iterations. The fitness function can select the individual with the highest fitness as the current optimal solution based on factors such as cluster compactness and separation.
[0104] Step a53: Perform a selection operation based on the fitness value of each individual to obtain an optimizing population. Crossover and mutation operations are then performed on the individuals in the optimizing population to obtain an updated population. Specifically, this embodiment employs the crossover operation of a genetic algorithm to select two individuals with higher fitness and exchange genes to generate new individuals, thereby increasing the diversity of the population and obtaining a crossover population. Based on the crossover population, the mutation operation of the genetic algorithm is then employed to randomly alter some of the genes of certain individuals to prevent premature convergence of the algorithm and maintain the population's exploration capability, thereby obtaining an updated population.
[0105] Step a54, determine whether the preset stop iteration condition is reached. If not, re-execute the calculation steps, selection operations, crossover operations, mutation operations and judgment steps in steps a52 to a54 until the preset stop iteration condition is reached; the preset stop iteration condition is reaching the maximum number of iterations or the genetic algorithm converges.
[0106] Accordingly, by employing a genetic algorithm to optimize the target clustering algorithm's selection of initial center points during the clustering process, this embodiment can improve the quality and stability of data source clustering results. Specifically, this embodiment initializes the population and calculates the fitness value of each individual based on a fitness function, selecting the individual with the highest fitness as the current optimal solution. Then, through selection, crossover, and mutation operations, the population is continuously optimized, increasing its diversity and exploration capabilities and preventing the algorithm from converging prematurely. This embodiment not only improves the rationality of initial center point selection but also enhances the robustness and accuracy of the clustering algorithm, ultimately resulting in optimized data source clustering results. This makes the data source clustering results more stable and reliable, better reflecting the actual relationships between data sources and providing a solid foundation for subsequent data analysis and applications.
[0107] In the above embodiment, as a possible implementation method, step 122, based on the analysis results, combined with the enterprise's personalized operation and maintenance demand information and historical data collection methods, uses a long short-term memory network to predict the optimal data collection method corresponding to the data source set, including:
[0108] Step b1: Based on the analysis results and combined with the enterprise's personalized operation and maintenance demand information, a comprehensive demand feature matrix is generated.
[0109] Step b2: Based on the comprehensive demand feature matrix, a time series analysis method is used to generate time series features of the historical collection method.
[0110] Step b3: Based on the comprehensive demand feature matrix and the time series characteristics of the historical collection method, a long short-term memory network is used to predict the optimal data collection method corresponding to the data source set.
[0111] Based on the analysis results and the enterprise's personalized O&M requirements, a comprehensive demand feature matrix is generated, ensuring that the model fully considers the enterprise's specific needs and actual conditions, improving the pertinence and practicality of the forecast. Based on this comprehensive demand feature matrix, a time series analysis method is used to generate time series features of the historical data collection method, fully leveraging the temporal dependencies and changing trends of historical data to enhance the model's predictive capabilities. Long-short-term memory networks improve the accuracy of the forecast data collection method, providing enterprises with more efficient and reliable O&M support.
[0112] In summary, this embodiment has the following advantages:
[0113] (1) Most existing data collection methods rely on manual configuration and selection of collection methods and lack an intelligent automatic selection mechanism. However, the intelligent operation and maintenance platform in this embodiment uses automation tools, support vector machines, and K-means clustering algorithms to automatically select the optimal data collection method based on the characteristics of the data source (such as data type, data volume, network environment, etc.), combined with the enterprise's personalized operation and maintenance requirements and historical data collection methods, thereby reducing manual intervention and improving collection efficiency and reliability.
[0114] (2) Existing data collection methods typically use serial processing on a single device or multiple devices, which cannot fully utilize the computing resources of multiple devices, resulting in limited data collection speed. However, the intelligent operation and maintenance platform in this embodiment uses a distributed computing framework to distribute data collection tasks to multiple devices in parallel for execution, greatly improving the speed and concurrency of data collection. This distributed processing method can not only process large-scale data, but also ensure the real-time and efficient data collection.
[0115] (3) Existing data collection methods often lack intelligent recommendation mechanisms. Users need to manually configure and adjust collection parameters, which is error-prone and inefficient. However, the intelligent operation and maintenance platform in this embodiment has a built-in multi-level rule engine. The dynamic rule layer in this multi-level rule engine can use machine learning algorithms to intelligently recommend the most appropriate screening rules based on historical data and user behavior data, further improving the accuracy and efficiency of data collection.
[0116] Figure 2 A structural diagram of a data acquisition system based on big data of an intelligent operation and maintenance platform provided in an embodiment of the present application is shown as follows: Figure 2As shown, the system includes:
[0117] The acquisition and screening module 21 is used to obtain configuration files of multiple data sources connected to the intelligent operation and maintenance platform, and use a multi-level rule engine to screen out target data sources from the multiple data sources based on the configuration files of the multiple data sources to form a data source set.
[0118] The prediction module 22 is used to predict the optimal data collection method corresponding to the data source set using a target prediction model; wherein the target prediction model introduces a support vector machine, a target clustering algorithm and a long short-term memory network.
[0119] The classification collection module 23 is used to combine the distributed computing framework of the intelligent operation and maintenance platform to assign the data collection tasks of the data source set to multiple devices in the intelligent operation and maintenance platform according to the optimal data collection method, so that multiple devices can collect the original data collected by all data sources in the data source set in parallel.
[0120] The processing and determination module 24 is used to process the original data and determine the operation and maintenance results based on the processed data and the operation and maintenance business demand information.
[0121] Figure 2 The data acquisition system based on big data of intelligent operation and maintenance platform can perform Figure 1 The implementation principle and technical effects of the data collection method based on big data of the intelligent operation and maintenance platform described in the illustrated embodiment will not be repeated here. The specific manner in which each module and unit performs operations in the data collection system based on big data of the intelligent operation and maintenance platform in the above embodiment has been described in detail in the embodiment of the method and will not be elaborated on here.
[0122] In one possible design, Figure 2 The data acquisition system based on big data of intelligent operation and maintenance platform in the embodiment shown can be implemented as a computing device, such as Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32 .
[0123] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32 .
[0124] The processing component 32 is used to: obtain configuration files of multiple data sources connected to the intelligent operation and maintenance platform, and use a multi-level rule engine to filter out target data sources from multiple data sources based on the configuration files of the multiple data sources to form a data source set; use a target prediction model to predict the optimal data collection method corresponding to the data source set; wherein the target prediction model introduces a support vector machine, a target clustering algorithm and a long short-term memory network; in combination with the distributed computing framework of the intelligent operation and maintenance platform, the data collection tasks of the data source set are assigned to multiple devices in the intelligent operation and maintenance platform according to the optimal data collection method, so that multiple devices can collect the original data collected by all data sources in the data source set in parallel; process the original data, and determine the operation and maintenance results based on the processed data and operation and maintenance business demand information.
[0125] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above method.
[0126] The storage component 31 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as random access memory (RAM), static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0127] Of course, a computing device may also include other components, such as input / output interfaces, display components, communication components, etc.
[0128] The input / output interface provides an interface between the processing component and the peripheral interface module, which can be an output device, an input device, etc.
[0129] The communication component is configured to facilitate, among other things, wired or wireless communications between the computing device and other devices.
[0130] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.
[0131] The present application also provides a computer storage medium storing a computer program, wherein the computer program can achieve the above-mentioned Figure 1 The embodiment shown is a data collection method based on big data of an intelligent operation and maintenance platform.
[0132] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0133] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0134] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0135] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A data collection method based on big data of intelligent operation and maintenance platform, characterized in that: include: Obtain configuration files of multiple data sources connected to the intelligent operation and maintenance platform, and use a multi-level rule engine to filter out a target data source from the multiple data sources based on the configuration files of the multiple data sources to form a data source set; A target prediction model is used to predict the optimal data collection method corresponding to the data source set; wherein the target prediction model introduces a support vector machine, a target clustering algorithm, and a long short-term memory network; In combination with the distributed computing framework of the intelligent operation and maintenance platform, data collection tasks of the data source set are assigned to multiple devices within the intelligent operation and maintenance platform according to the optimal data collection method, so that the multiple devices can collect the raw data collected by all data sources in the data source set in parallel; Processing the raw data and determining the operation and maintenance results based on the processed data and the operation and maintenance business demand information; The target prediction model is used to predict the optimal data collection method corresponding to the data source set; wherein the target prediction model introduces a support vector machine, a target clustering algorithm and a long short-term memory network, including: Based on the configuration files of each data source in the data source set, using a support vector machine and a target clustering algorithm to analyze the characteristics of each data source in the data source set to obtain an analysis result; According to the analysis results, combined with the enterprise operation and maintenance personalized demand information and historical data collection methods, a long short-term memory network is used to predict the optimal data collection method corresponding to the data source set.
2. The method according to claim 1, characterized in that The configuration files of each data source in the data source set are analyzed using a support vector machine and a target clustering algorithm to obtain analysis results, including: Extracting first characteristic information of each data source in the data source set using metadata extraction technology based on a configuration file of each data source in the data source set; Based on the first characteristic information of each data source in the data source set, a support vector machine is used to classify the category attributes of all data sources in the data source set to obtain a data source classification result; Based on the data source classification results, a target clustering algorithm is used to cluster all data sources in each category to obtain a data source clustering result; Based on the data source clustering result, a principal component analysis technique is used to select partial characteristic information from the first characteristic information, and the partial characteristic information is used as the second characteristic information; based on the second characteristic information, an association rule learning algorithm is used to generate an association relationship between data sources in each category; Based on the association relationship between the data sources, a genetic algorithm is used to optimize the selection of the initial center point of the target clustering algorithm during the clustering process to obtain an optimized data source clustering result; Based on the optimized data source clustering results, a characteristic probability model of the data source is constructed using a Bayesian network to form a data source behavior prediction model; Generate analysis results based on the configuration files of each data source in the data source set, the first characteristic information of each data source in the data source set, the data source classification results, the data source clustering results, the second characteristic information, the association relationship between the data sources, the optimized data source clustering results and the data source behavior prediction model.
3. The method according to claim 2, characterized in that The target clustering algorithm includes a K-means clustering algorithm. Based on the data source classification result, the target clustering algorithm is used to cluster all data sources in each category to obtain a data source clustering result, including: Define a characteristic quantitative indicator system for data sources, which includes the following characteristic quantitative indicators: data volume, access latency, update frequency, and security; Based on the data source classification result and the characteristic quantification index system, a standardization technology is used to perform standard quantization processing on the first characteristic information of all data sources in each category to obtain characteristic values of all data sources in each category; For each category, a hierarchical clustering algorithm is introduced to determine the number of groups of the data source, and the number of groups is used as the value of the initial input parameter of the K-means clustering algorithm; An initial center point is selected using a probability-weighted approach, and a K-means clustering process is performed using a K-means clustering algorithm based on the values of the initial input parameters and the initial center point. A dynamic adjustment mechanism for the distance metric is introduced into the clustering process to obtain a data source clustering result; the dynamic adjustment mechanism for the distance metric means that different distance metrics are used for data sources with different characteristic types; The silhouette coefficient method is applied to evaluate the data source clustering results to obtain a silhouette coefficient value. When the silhouette coefficient value is less than a preset threshold, the values of the initial input parameters are adjusted or the distance metric is optimized, and the K-means clustering process and evaluation operation are repeated until the silhouette coefficient value is greater than or equal to the preset threshold.
4. The method according to claim 3, characterized in that The dynamic adjustment mechanism of distance metric is introduced into the clustering process to obtain the data source clustering result; The dynamic adjustment mechanism of the distance metric refers to using different distance metric standards for data sources with different feature types, including: A distance metric library is constructed using a plurality of distance metrics, wherein the distance metric library includes the following distance metrics: Euclidean distance, Manhattan distance, cosine similarity distance, and dynamic time warping distance; Using a feature engineering method to classify all characteristic quantitative indicators in the characteristic quantitative indicator system to obtain an initial characteristic classification result; based on the initial characteristic classification result, using an adaptive distance metric selection algorithm to select a distance metric corresponding to each category in the initial characteristic classification result from the distance metric standard library; Assigning corresponding weights to all characteristic quantitative indicators in each category according to characteristic importance, and dynamically adjusting the distance metric corresponding to each category according to the corresponding weights of all characteristic quantitative indicators in each category to obtain an adjusted distance metric corresponding to each category; The data source clustering results are determined based on the adjusted distance metrics corresponding to all categories.
5. The method according to claim 2, characterized in that The method of optimizing the selection of the initial center point of the target clustering algorithm in the clustering process based on the association relationship between the data sources to obtain the optimized data source clustering result includes: Initializing the population to obtain an initial population, wherein each individual in the initial population represents a position of a to-be-determined center point of a group of data sources; Calculating the fitness value of each individual in the population based on the fitness function; the population is the initial population in the first iteration process and is the updated population in the subsequent iteration processes; Performing a selection operation based on the fitness value of each individual to obtain an optimizing population, and performing a crossover operation and a mutation operation on the individuals in the optimizing population to obtain an updated population; Determine whether the preset stop iteration condition is reached. If not, re-execute the calculation step, selection operation, crossover operation, mutation operation and judgment step until the preset stop iteration condition is reached; the preset stop iteration condition is reaching the maximum number of iterations or the genetic algorithm converges.
6. The method according to claim 1, wherein The method of using a long short-term memory network to predict the optimal data collection method corresponding to the data source set based on the analysis results, combined with the enterprise operation and maintenance personalized demand information and historical data collection methods, includes: Based on the analysis results and combined with the enterprise's personalized operation and maintenance demand information, a comprehensive demand feature matrix is generated; Based on the comprehensive demand feature matrix, the time series analysis method is used to generate the time series features of the historical collection mode; Based on the comprehensive demand feature matrix and the time series characteristics of historical collection methods, a long short-term memory network is used to predict the optimal data collection method corresponding to the data source set.
7. A data collection system based on big data of intelligent operation and maintenance platform, characterized in that: include: An acquisition and screening module is used to obtain configuration files of multiple data sources connected to the intelligent operation and maintenance platform, and to screen out a target data source from the multiple data sources using a multi-level rule engine based on the configuration files of the multiple data sources to form a data source set; A prediction module, configured to predict the optimal data collection method corresponding to the data source set using a target prediction model; wherein the target prediction model introduces a support vector machine, a target clustering algorithm, and a long short-term memory network; A classification collection module is used to combine the distributed computing framework of the intelligent operation and maintenance platform to assign data collection tasks of the data source set to multiple devices in the intelligent operation and maintenance platform according to the optimal data collection method, so that the multiple devices can collect the raw data collected by all data sources in the data source set in parallel; A processing and determination module is used to process the raw data and determine the operation and maintenance results based on the processed data and the operation and maintenance business demand information; The target prediction model is used to predict the optimal data collection method corresponding to the data source set; wherein the target prediction model introduces a support vector machine, a target clustering algorithm and a long short-term memory network, including: Based on the configuration files of each data source in the data source set, using a support vector machine and a target clustering algorithm to analyze the characteristics of each data source in the data source set to obtain an analysis result; According to the analysis results, combined with the enterprise operation and maintenance personalized demand information and historical data collection methods, a long short-term memory network is used to predict the optimal data collection method corresponding to the data source set.
8. A computing device, characterized in that It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a data collection method based on big data of an intelligent operation and maintenance platform as described in any one of claims 1 to 6.
9. A computer storage medium, characterized in that A computer program is stored, and when the computer program is executed by a computer, the data collection method based on big data of an intelligent operation and maintenance platform as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Method and system for building medical insurance hospitalization fee prediction model
CN108197737A
Big data intelligent acquisition method based on deep learning
CN119003849A