A data cleaning system and method based on machine learning and data stream processing

Through an intelligent prediction model based on machine learning, the problem of improper server configuration in log data cleaning of the e-commerce cloud platform was solved, the accurate configuration of the cleaning server was achieved, and the efficient and stable operation of the cloud platform was ensured.

CN120508553BActive Publication Date: 2025-10-21GUANGZHOU SHANGHANG INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510950799.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-10-21
Estimated Expiration
2045-07-10

AI Technical Summary

Technical Problem

Existing technologies cannot accurately predict the number of duplicate, erroneous, and incomplete logs generated by e-commerce cloud platforms in future time periods, resulting in improper configuration of cleaning servers and problems such as insufficient or excessive efficiency.

Method used

By employing an AI model based on machine learning, and combining cloud product sales data streams, operation and maintenance data streams, platform-related information, and registered user-related information, the system can intelligently predict the number of duplicate, erroneous, and incomplete log entries in future time periods, and configure the number of cleaning servers based on the prediction results.

Benefits of technology

The accurate matching of the number of cleaning servers with actual demand is achieved, which avoids inefficiency or over-configuration and ensures the efficient and stable operation of the cloud platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508553B_ABST
    Figure CN120508553B_ABST
Patent Text Reader

Abstract

The application relates to a data cleaning system and method based on machine learning and data stream processing, and belongs to the field of electric digital data processing.The application adopts an artificial intelligence model based on machine learning to intelligently predict log repetition numbers, log error numbers and log incomplete numbers generated by a target cloud platform in a future time interval according to cloud product sales data streams and operation and maintenance data streams of the target cloud platform, a plurality of platform correlation information and a plurality of registered user correlation information of the target cloud platform, log repetition numbers, log error numbers and log incomplete numbers of each past time partition of the target cloud platform, and according to the intelligent prediction result, the number of cleaning servers allocated to the future time partition of the target cloud platform is determined, so that on the basis of fully utilizing limited cleaning server resources, the timeliness and reliability of cloud platform data cleaning are ensured, and strong support is provided for efficient and stable operation of the cloud platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of electronic digital data processing, and in particular relates to a data cleaning system and method based on machine learning and data stream processing. Background Art

[0002] Generally speaking, data cleaning refers to the processing of raw data to correct or delete duplicate, erroneous, or incomplete content, thereby improving data quality and usability. Data cleaning is a crucial step in data preprocessing and includes steps such as handling missing values, removing duplicate data, and correcting data errors. The goal of data cleaning is to remove redundant and invalid data, ensure the accuracy and consistency of retained data, and reduce redundancy and instability across the entire system. For example, cloud platforms designed for e-commerce, or cloud-based sales platforms, require frequent data cleaning of duplicate, erroneous, or incomplete log data generated during the cloud-based product sales process to ensure efficient and stable operation.

[0003] For example, Chinese invention patent publication CN107463639A proposes an artificial intelligence-based SMS data cleaning method, which includes the following steps: constructing an artificial intelligence-based SMS data cleaning system; collecting required SMS data, identifying it, and merging SMS data in the same field; presetting comparison rules, and reading SMS data in the same field one by one, comparing them with the pre-set comparison rules, thereby screening out problematic data; summarizing and cleaning the obtained problematic data; reconfirming the cleaned problematic data, and storing the confirmed cleaned data. The present invention constructs an artificial intelligence-based SMS data cleaning system, utilizing the system to complete the collection, identification, classification, and comparison of SMS data; summarizing and cleaning erroneous data, and storing the cleaned data; its cleaning process is simple, efficient, and accurate. Based on artificial intelligence, it avoids manual input and assistance, has low cleaning input and low cost, and is worthy of promotion.

[0004] For example, Chinese invention patent publication CN117891812A proposes an AI-based big data cleaning method and system, belonging to the field of data cleaning technology. The method includes: obtaining the data to be cleaned and dividing the data to be cleaned according to the data source to obtain several sub-datasets. At the same time, based on the AI-based cleaning task, path mining is performed on each data source to determine the source path of the corresponding data source, and then an interference mechanism is constructed to determine a first cleaning method for the corresponding source path; standard conditions are extracted for the cleaning task to obtain an initial cleaning method for each sub-dataset, and a cleaning optimization factor matching the corresponding sub-dataset is obtained to optimize the initial cleaning method to obtain a second cleaning method; cleaning rules are constructed based on all the first cleaning methods and the second cleaning methods, and the data to be cleaned is cleaned. To a certain extent, flexible data cleaning is achieved and data quality is improved.

[0005] It can be seen that the above-mentioned technical solutions either use artificial intelligence models to identify data that need to be cleaned, or use artificial intelligence models to trace the data that need to be cleaned, and cannot predict the quantity level of data that needs to be cleaned in the future. For example, it is impossible to predict the number of duplicate logs, log errors, and incomplete logs generated by the cloud platform designed for e-commerce in future time partitions, and thus it is impossible to configure the cloud platform in advance with a suitable number of cleaning servers suitable for future time partitions, resulting in a mismatch between the number of cleaning servers configured in advance and the actual number of error logs generated, which easily leads to the dilemma of insufficient cleaning efficiency of cleaning servers or excessive configuration of cleaning servers. Summary of the Invention

[0006] In order to solve the technical problems in the prior art, the present invention provides a data cleaning system and method based on machine learning and data stream processing. The system and method can, on the basis of designing a customized structure for a target cloud platform based on a machine learning artificial intelligence model and comprehensively and fully selecting various basic information, use the machine learning artificial intelligence model to intelligently predict the number of duplicate logs, error logs and incomplete logs generated by the target cloud platform in the future time interval according to the cloud product sales data flow and operation and maintenance data flow of the target cloud platform, multiple platform-related information and multiple registered user-related information of the target cloud platform, and the number of duplicate logs, error logs and incomplete logs in each past time partition of the target cloud platform, and determine the number of cleaning servers allocated to the target cloud platform in the future time partition based on the intelligent prediction results, thereby achieving an accurate match between the number of cleaning servers configured in advance and the actual number of error logs generated, avoiding the dilemma of insufficient cleaning efficiency of cleaning servers or excessive configuration of cleaning servers.

[0007] According to one aspect of the present invention, a data cleaning system based on machine learning and data stream processing is provided, the system comprising:

[0008] A data parsing component, used to parse the cloud product sales data stream and operation and maintenance data stream corresponding to the latest past time partition of the target cloud platform with the current time as the end time point. The target cloud platform is a cloud service-based sales platform used by the designated e-commerce company;

[0009] An information extraction device, used to extract multiple pieces of platform-related information and multiple pieces of registered user-related information of the target cloud platform;

[0010] Partition collection device, used to collect the number of duplicate log entries, the number of error entries, and the number of incomplete log entries corresponding to each past time partition of the target cloud platform;

[0011] An object building component is used to build a machine learning-based artificial intelligence model customized for the target cloud platform. The number of model learning times is positively correlated with the total number of registered users of the target cloud platform.

[0012] The content prediction device is connected to the data analysis device, the information extraction device, the partition collection device, and the object construction device respectively, and is used to use a machine learning-based artificial intelligence model to intelligently predict the number of duplicate log entries, the number of log errors, and the number of incomplete log entries corresponding to the future time partition of the target cloud platform starting at the current time point based on the duration of the time partition and the output of the data analysis device, the information extraction device, and the partition collection device;

[0013] The resource allocation device is connected to the content prediction device and is used to determine the number of cleaning servers allocated to the target cloud platform in the future time partition based on the cumulative values ​​of the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries predicted by the intelligent prediction.

[0014] According to another aspect of the present invention, a data cleaning method based on machine learning and data stream processing is provided, the method comprising:

[0015] Analyze the cloud product sales data stream and operation and maintenance data stream corresponding to the latest past time partition of the target cloud platform with the current time as the end time point. The target cloud platform is the cloud service-based sales platform used by the designated e-commerce company.

[0016] Extract multiple platform-related information and multiple registered user-related information of the target cloud platform;

[0017] Collect the number of duplicate log entries, error entries, and incomplete log entries for each past time partition of the target cloud platform.

[0018] Build a machine learning-based artificial intelligence model customized for the target cloud platform, with the number of model learning times positively correlated with the total number of registered users on the target cloud platform;

[0019] Adopting a machine learning-based artificial intelligence model, based on the duration of the time partition, the cloud product sales data flow and operation and maintenance data flow corresponding to the latest past time partition of the target cloud platform with the current moment as the end time point, multiple platform-related information and multiple registered user-related information of the target cloud platform, and the number of duplicate log entries, error entries, and incomplete log entries corresponding to each past time partition of the target cloud platform, intelligently predict the number of duplicate log entries, error entries, and incomplete log entries corresponding to the future time partition of the target cloud platform with the current moment as the start time point;

[0020] The number of cleaning servers allocated to the target cloud platform in future time partitions is determined based on the cumulative values ​​of the number of duplicate logs, the number of log errors, and the number of incomplete logs predicted by intelligent prediction.

[0021] It can be seen that the present invention has at least the following four key invention points:

[0022] First: Using a machine learning-based artificial intelligence model, we can intelligently predict the number of duplicate logs, error logs, and incomplete logs generated by the target cloud platform in the future time interval based on the target cloud platform's cloud product sales data stream and operation and maintenance data stream, multiple platform-related information and multiple registered user-related information, and the number of duplicate logs, error logs, and incomplete logs in each past time partition of the target cloud platform. Based on the intelligent prediction results, we can determine the number of cleaning servers allocated to the target cloud platform in the future time partition. This ensures the timeliness and reliability of cloud platform data cleaning while fully utilizing limited cleaning server resources, providing strong support for the efficient and stable operation of the cloud platform.

[0023] Second: In order to complete the intelligent prediction of the number of duplicate logs, log errors, and incomplete logs generated by the target cloud platform in the future time interval, a machine learning-based artificial intelligence model with a customized structure for the target cloud platform is designed. The model is a deep neural network that has completed multiple learning cycles. The number of model learning cycles is positively correlated with the total number of registered users of the target cloud platform, and the deep neural network includes multiple hidden layers, an input layer, and an output layer. In the deep neural network, the multiple hidden layers are located between the input layer and the output layer, and the number of hidden layers is positively correlated with the total number of product types on sale on the target cloud platform. The above-mentioned customized structural designs of the machine learning-based artificial intelligence model provide guarantees for the stability and effectiveness of the intelligent prediction results;

[0024] The third part: In order to complete the intelligent prediction of the number of duplicate logs, log errors and incomplete logs generated by the target cloud platform in the future time interval, various basic information are selected in a targeted manner, including the duration of the time partition, the cloud product sales data flow and operation and maintenance data flow corresponding to the latest past time partition of the target cloud platform with the current moment as the end time point, multiple platform-related information and multiple registered user-related information of the target cloud platform, the number of duplicate logs, the number of log errors and the number of incomplete logs corresponding to each past time partition of the target cloud platform. More specifically, the multiple platform-related information of the target cloud platform is the platform establishment time of the target cloud platform, the total number of product types on sale and the total number of platform sales regions, the multiple registered user-related information of the target cloud platform, The information is the total number of registered users of the target cloud platform, as well as the age information, gender identification and the total number of products purchased before the current moment of each registered user. The cloud product sales data stream corresponding to each time partition of the target cloud platform is the total amount of products sold by the target cloud platform in the time partition, the number of products corresponding to each product type, and the region numbers corresponding to each sales region involved. The operation and maintenance data stream corresponding to each time partition of the target platform is the number of product inquiries received by the target cloud platform in the time partition, the number of operation and maintenance personnel used, the duration of server downtime events, and the number of servers affected when the server downtime events occurred. The comprehensive and sufficient selection of the above basic information further guarantees the stability and effectiveness of the intelligent prediction results.

[0025] Fourth: In each learning of the machine learning-based artificial intelligence model customized for the target cloud platform, the known number of log duplicates, known number of log errors, and known number of incomplete logs corresponding to a certain past time partition of the target cloud platform are used as the output content of the machine learning-based artificial intelligence model, and the cloud product sales data flow and operation and maintenance data flow corresponding to the past time partition of the target cloud platform before and immediately adjacent to the said past time partition, multiple platform-related information and multiple registered user-related information of the target cloud platform, and the number of log duplicates, the number of log errors, and the number of incomplete logs corresponding to each past time partition of the target cloud platform before the said past time partition are used as the input content of the machine learning-based artificial intelligence model to complete this learning, thereby ensuring the learning effect of each learning of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The embodiments of the present invention will be described below with reference to the accompanying drawings, in which:

[0027] Figure 1 Schematic diagram of the working scenario of a data cleaning system and method based on machine learning and data stream processing according to the present invention.

[0028] Figure 2 This is a diagram showing the internal structure of a data cleaning system based on machine learning and data stream processing according to the first embodiment of the present invention.

[0029] Figure 3 This is a diagram showing the internal structure of a data cleaning system based on machine learning and data stream processing according to the second embodiment of the present invention.

[0030] Figure 4 This is a diagram showing the internal structure of a data cleaning system based on machine learning and data stream processing according to the third embodiment of the present invention.

[0031] Figure 5 This is a diagram showing the internal structure of a data cleaning system based on machine learning and data stream processing according to the fourth embodiment of the present invention.

[0032] Figure 6 This is a diagram showing the internal structure of a data cleaning system based on machine learning and data stream processing according to the fifth embodiment of the present invention.

[0033] Figure 7 The present invention is a flowchart showing the steps of a data cleaning method based on machine learning and data stream processing according to the sixth embodiment of the present invention. DETAILED DESCRIPTION

[0034] like Figure 1 As shown, a working scenario diagram of a data cleaning system and method based on machine learning and data stream processing according to the present invention is given. The present invention belongs to the field of electronic digital data processing.

[0035] The specific technical process of the present invention is as follows:

[0036] Technical Process A: For the target cloud platform for executing product cloud sales, that is, the cloud service-based sales platform used by e-commerce, a customized machine learning-based artificial intelligence model is designed to intelligently predict the number of duplicate logs, log errors, and log incompleteness generated by the target cloud platform in the future time period. Figure 1 As shown;

[0037] Specifically, the customized structure design of the machine learning-based artificial intelligence model designed for the target cloud platform is mainly reflected in the following aspects:

[0038] First, the machine learning-based artificial intelligence model designed for the target cloud platform is a deep neural network that has completed multiple learning cycles. The number of model learning cycles is positively correlated with the total number of registered users on the target cloud platform.

[0039] Second: The specific architecture of the deep neural network used is that the deep neural network includes multiple hidden layers, an input layer and an output layer, and in the deep neural network, the multiple hidden layers are located between the input layer and the output layer;

[0040] Third: In the deep neural network used, the number of hidden layers is positively correlated with the total number of product types sold on the target cloud platform. For example, when the total number of product types sold on the target cloud platform is 100,000, the number of hidden layers is 3; when the total number of product types sold on the target cloud platform is 500,000, the number of hidden layers is 5; when the total number of product types sold on the target cloud platform is 1 million, the number of hidden layers is 7; when the total number of product types sold on the target cloud platform is 5 million, the number of hidden layers is 9, and so on;

[0041] Fourth: In each learning of the machine learning-based artificial intelligence model customized for the target cloud platform, the number of known log duplicates, known log errors, and known log incompleteness corresponding to a certain past time partition of the target cloud platform is used as the output content of the machine learning-based artificial intelligence model, and the cloud product sales data stream and operation and maintenance data stream corresponding to the past time partition of the target cloud platform before and immediately adjacent to the said past time partition, multiple pieces of platform-related information and multiple pieces of registered user-related information of the target cloud platform, and the number of log duplicates, log errors, and incompleteness corresponding to each past time partition of the target cloud platform before the said past time partition are used as the input content of the machine learning-based artificial intelligence model to complete this learning, thereby ensuring the learning effect of each learning of the model;

[0042] In this way, through the customized structural design of the machine learning-based artificial intelligence model, the stability and effectiveness of the classification intelligent prediction results of the number of error logs in the future time-partitioned cloud platform are guaranteed;

[0043] Technical Process B: Targeting the cloud platform used by e-commerce companies for cloud-based product sales, the target cloud platform selects essential information to intelligently predict the number of duplicate, erroneous, and incomplete logs generated by the target cloud platform in future timeframes.

[0044] Specifically, the basic information includes the duration of the time partition, the cloud product sales data flow and operation and maintenance data flow corresponding to the latest past time partition of the target cloud platform with the current time as the end time point, multiple platform-related information and multiple registered user-related information of the target cloud platform, the number of duplicate logs corresponding to each past time partition of the target cloud platform, the number of errors in each log, and the number of incomplete logs. Figure 1As shown, the cloud product sales data flow and operation and maintenance data flow corresponding to the latest past time partition of the target cloud platform with the current time as the end time point are Figure 1 The two data streams in

[0045] More specifically, the target cloud platform's multiple platform-related information includes the target cloud platform's platform establishment time, the total number of product types on sale, and the total number of platform sales regions; the target cloud platform's multiple registered user-related information includes the target cloud platform's total number of registered users, the age information, gender identification, and the total number of products purchased by each registered user before the current moment; and the target cloud platform's corresponding cloud product sales data stream for each time partition includes the target cloud platform's total product amount sold within the time partition, the number of products corresponding to each product type, and the region numbers corresponding to each sales region involved; and the target platform's operation and maintenance data stream for each time partition includes the number of product inquiries received by the target cloud platform within the time partition, the number of operation and maintenance personnel used, the duration of the server downtime event, and the number of servers affected when the server downtime event occurs;

[0046] In this way, through the comprehensive and sufficient selection of the above basic information, the stability and effectiveness of the classification intelligent prediction results of the number of error logs in the future time partition cloud platform are further guaranteed;

[0047] Technical Process C: Using the machine learning-based artificial intelligence model customized by Technical Process A, based on the comprehensive and fully selected basic information from Technical Process B, we can complete the intelligent classification prediction of the number of error logs on the future time-partitioned cloud platform.

[0048] Specifically, the classification intelligent prediction results of the artificial intelligence model based on machine learning include: the number of duplicate logs, the number of log errors, and the number of incomplete logs generated by the target cloud platform in the future time interval, such as Figure 1 As shown in the figure, the intelligent prediction result is the future error log data, including the number of duplicate logs, log errors, and incomplete logs generated by the target cloud platform in the future time interval;

[0049] Technical process D: Determine the number of cleaning servers to be allocated to the target cloud platform in the future time partition based on the classification intelligent prediction results of technical process C;

[0050] For example, the number of cleaning servers allocated to the target cloud platform in the future time partition is determined based on the cumulative values ​​of the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries predicted by the intelligent prediction;

[0051] By way of further example, the number of cleaning servers determined to be allocated to the target cloud platform for the future time partition is proportional to the cumulative value. For example, when the cumulative value of the number of duplicate log entries, error log entries, and incomplete log entries predicted by the intelligent prediction is 100,000, the number of cleaning servers determined to be allocated to the target cloud platform for the future time partition is 1; when the cumulative value of the number of duplicate log entries, error log entries, and incomplete log entries predicted by the intelligent prediction is 200,000, the number of cleaning servers determined to be allocated to the target cloud platform for the future time partition is 2; when the cumulative value of the number of duplicate log entries, error log entries, and incomplete log entries predicted by the intelligent prediction is 300,000, the number of cleaning servers determined to be allocated to the target cloud platform for the future time partition is 3, and so on. The cleaning server not only needs to delete erroneous logs, but also needs to remove duplicate logs and restore incomplete logs to a complete state;

[0052] In this way, through the coordinated operation of the above four technical processes, an accurate match is achieved between the number of cleaning servers configured in advance and the actual number of error logs generated, avoiding the dilemma of insufficient cleaning efficiency or excessive configuration of cleaning servers. At the same time, by reducing data redundancy and erroneous data, the efficient and stable operation of the cloud platform is guaranteed.

[0053] The key points of the present invention are: accurate matching of the number of pre-configured cleaning servers based on intelligent prediction results and the actual number of error logs generated, customized structural design of machine learning-based artificial intelligence models for the target cloud platform, and comprehensive and sufficient selection of various basic information for intelligent prediction.

[0054] Below, a data cleaning system and method based on machine learning and data stream processing of the present invention will be specifically described in the form of an embodiment.

[0055] First embodiment

[0056] Figure 2 This is a diagram showing the internal structure of a data cleaning system based on machine learning and data stream processing according to the first embodiment of the present invention.

[0057] like Figure 2 As shown, the data cleaning system based on machine learning and data stream processing includes the following components:

[0058] A data parsing component, used to parse the cloud product sales data stream and operation and maintenance data stream corresponding to the latest past time partition of the target cloud platform with the current time as the end time point. The target cloud platform is a cloud service-based sales platform used by the designated e-commerce company;

[0059] For example, different cloud service-based sales platforms can be customized for different e-commerce companies. In fact, different cloud services are customized for different e-commerce companies, thereby providing a platform for the e-commerce companies to complete the sales of their products based on cloud services.

[0060] An information extraction device, used to extract multiple pieces of platform-related information and multiple pieces of registered user-related information of the target cloud platform;

[0061] For example, a first extraction component and a second extraction component may be used to respectively extract multiple pieces of platform association information of the target cloud platform and multiple pieces of registered user association information of the target cloud platform;

[0062] Partition collection device, used to collect the number of duplicate log entries, the number of error entries, and the number of incomplete log entries corresponding to each past time partition of the target cloud platform;

[0063] For example, when the duration of each time partition is 10 minutes, the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries corresponding to each past time partition of the target cloud platform are collected. If the current time is 5:00 PM, then the past time partitions are 4:50 PM to 5:00 PM, 4:40 PM to 4:50 PM, 4:30 PM to 4:40 PM, 4:20 PM to 4:30 PM, 4:10 PM to 4:20 PM, 4:00 PM to 4:10 PM, and so on.

[0064] For example, the number of time partitions of each selected past time partition is positively correlated with the total number of registered users of the target cloud platform. For example, if the total number of registered users of the target cloud platform is about 2 million, the number of time partitions of each selected past time partition is 6. If the total number of registered users of the target cloud platform is about 5 million, the number of time partitions of each selected past time partition is 8. If the total number of registered users of the target cloud platform is about 10 million, the number of time partitions of each selected past time partition is 10. If the total number of registered users of the target cloud platform is about 20 million, the number of time partitions of each selected past time partition is 12, and so on.

[0065] An object building component is used to build a machine learning-based artificial intelligence model customized for the target cloud platform. The number of model learning times is positively correlated with the total number of registered users of the target cloud platform.

[0066] For example, a machine learning-based artificial intelligence model customized for a target cloud platform is constructed, and the positive correlation between the number of model learning times and the total number of registered users of the target cloud platform includes: when the total number of registered users of the target cloud platform is around 2 million, the corresponding number of model learning times is 2,000; when the total number of registered users of the target cloud platform is around 5 million, the corresponding number of model learning times is 2,500; when the total number of registered users of the target cloud platform is around 10 million, the corresponding number of model learning times is 3,000; when the total number of registered users of the target cloud platform is around 20 million, the corresponding number of model learning times is 4,000, and so on;

[0067] The content prediction device is connected to the data analysis device, the information extraction device, the partition collection device, and the object construction device respectively, and is used to use a machine learning-based artificial intelligence model to intelligently predict the number of duplicate log entries, the number of log errors, and the number of incomplete log entries corresponding to the future time partition of the target cloud platform starting at the current time point based on the duration of the time partition and the output of the data analysis device, the information extraction device, and the partition collection device;

[0068] Here, the outputs of the data analysis device, the information extraction device, and the partition collection device are all the outputs of the three devices, namely, the duration of the time partition, the cloud product sales data flow and operation and maintenance data flow corresponding to the latest past time partition of the target cloud platform with the current time as the end time point, multiple platform-related information and multiple registered user-related information of the target cloud platform, and the number of duplicate log entries, the number of error entries in each log, and the number of incomplete log entries corresponding to each past time partition of the target cloud platform.

[0069] A resource allocation device, connected to the content prediction device, is used to determine the number of cleaning servers to be allocated to the target cloud platform in the future time partition based on the accumulated values ​​of the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries predicted by the intelligent prediction;

[0070] The number of cleaning servers allocated to the target cloud platform in the determined future time partition is proportional to the cumulative number of duplicate log entries, error log entries, and incomplete log entries based on intelligent prediction, and the cleaning servers allocated to the target cloud platform are servers that perform log data cleaning tasks for the target cloud platform;

[0071] For example, the number of cleaning servers allocated to the target cloud platform for the determined future time partition is proportional to the cumulative value of the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries based on intelligent prediction, and the cleaning servers allocated to the target cloud platform are servers that perform the target cloud platform log data cleaning task, including: when the cumulative value of the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries predicted by the intelligent prediction is 100,000, the number of cleaning servers allocated to the target cloud platform for the determined future time partition is 1, when the cumulative value of the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries predicted by the intelligent prediction is 200,000, the number of cleaning servers allocated to the target cloud platform for the determined future time partition is 2, when the cumulative value of the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries predicted by the intelligent prediction is 300,000, the number of cleaning servers allocated to the target cloud platform for the determined future time partition is 3, and so on. The cleaning server not only needs to delete erroneous logs, but also needs to remove duplicate logs and restore incomplete logs to a complete state;

[0072] The target cloud platform's multiple platform-related information includes the target cloud platform's establishment time, the total number of product types on sale, and the total number of platform sales regions; the target cloud platform's multiple registered user-related information includes the target cloud platform's total number of registered users, each registered user's age information, gender identifier, and the total number of products purchased before the current moment;

[0073] The cloud product sales data stream corresponding to each time partition of the target cloud platform is the total amount of products sold by the target cloud platform in the time partition, the number of products corresponding to each product type, and the region numbers corresponding to each sales region involved; the operation and maintenance data stream corresponding to each time partition of the target platform is the number of product inquiries received by the target cloud platform in the time partition, the number of operation and maintenance personnel used, the duration of the server downtime event, and the number of servers affected when the server downtime event occurs;

[0074] Among them, in each learning of the machine learning-based artificial intelligence model customized for the target cloud platform, the number of known log duplicates, the number of known log errors, and the number of known log incomplete entries corresponding to a certain past time partition of the target cloud platform are used as the output content of the machine learning-based artificial intelligence model, and the cloud product sales data flow and operation and maintenance data flow corresponding to the past time partition of the target cloud platform before and immediately adjacent to the certain past time partition, multiple platform-related information and multiple registered user-related information of the target cloud platform, and the number of log duplicates, the number of log errors, and the number of incomplete log entries corresponding to each past time partition of the target cloud platform before the certain past time partition are used as the input content of the machine learning-based artificial intelligence model to complete this learning;

[0075] The machine learning-based artificial intelligence model is a deep neural network that has completed multiple learning cycles. The deep neural network includes multiple hidden layers, an input layer, and an output layer. In the deep neural network, the multiple hidden layers are located between the input layer and the output layer, and the number of hidden layers is positively correlated with the total number of product types on sale on the target cloud platform.

[0076] For example, the artificial intelligence model based on machine learning is a deep neural network after completing multiple learning processes, wherein the deep neural network includes multiple hidden layers, an input layer, and an output layer. In the deep neural network, the multiple hidden layers are located between the input layer and the output layer, and the number of hidden layers is positively correlated with the total number of product types on sale on the target cloud platform, including: when the total number of product types on sale on the target cloud platform is 100,000, the number of hidden layers is 3; when the total number of product types on sale on the target cloud platform is 500,000, the number of hidden layers is 5; when the total number of product types on sale on the target cloud platform is 1 million, the number of hidden layers is 7; when the total number of product types on sale on the target cloud platform is 5 million, the number of hidden layers is 9, and so on.

[0077] Second embodiment

[0078] Figure 3 This is a diagram showing the internal structure of a data cleaning system based on machine learning and data stream processing according to the second embodiment of the present invention.

[0079] like Figure 3 As shown, compared with Figure 2 , the data cleaning system based on machine learning and data stream processing also includes:

[0080] A status warning device is connected to the resource allocation device and is used to execute a warning operation corresponding to insufficient number of cleaning servers when the number of existing cleaning servers is less than the number of cleaning servers allocated to the target cloud platform in the future time partition;

[0081] For example, the status warning device is connected to the resource allocation device and is used to perform a warning operation corresponding to insufficient number of cleaning servers when the number of existing cleaning servers is less than the number of cleaning servers allocated to the target cloud platform in the future time partition. The status warning device may be an acoustic warning device or an optical warning device.

[0082] The status warning device is further configured to temporarily suspend the warning operation corresponding to insufficient number of cleaning servers when the number of existing cleaning servers is greater than or equal to the number of cleaning servers allocated to the target cloud platform in the future time partition.

[0083] Third embodiment

[0084] Figure 4 This is a diagram showing the internal structure of a data cleaning system based on machine learning and data stream processing according to the third embodiment of the present invention.

[0085] like Figure 4 As shown, compared with Figure 3 , the data cleaning system based on machine learning and data stream processing also includes:

[0086] An on-site display device, connected to the resource allocation device, for receiving the number of cleaning servers allocated to the target cloud platform in the future time partition, and performing on-site display of the number of cleaning servers allocated to the target cloud platform in the future time partition;

[0087] For example, an on-site display device is connected to the resource allocation device, and is used to receive the number of cleaning servers allocated to the target cloud platform in the future time partition, and perform on-site display of the number of cleaning servers allocated to the target cloud platform in the future time partition. This includes: selecting an LED display array or a liquid crystal display screen to implement the on-site display device, connecting it to the resource allocation device, and receiving the number of cleaning servers allocated to the target cloud platform in the future time partition, and performing on-site display of the number of cleaning servers allocated to the target cloud platform in the future time partition.

[0088] Fourth embodiment

[0089] Figure 5 This is a diagram showing the internal structure of a data cleaning system based on machine learning and data stream processing according to the fourth embodiment of the present invention.

[0090] like Figure 5 As shown, compared with Figure 4 , the data cleaning system based on machine learning and data stream processing also includes:

[0091] a model storage device connected to the object construction device, configured to receive the machine learning-based artificial intelligence model customized for the target cloud platform and store the machine learning-based artificial intelligence model customized for the target cloud platform;

[0092] For example, a model storage device is connected to the object construction device, and is used to receive the machine learning-based artificial intelligence model customized for the target cloud platform, and store the machine learning-based artificial intelligence model customized for the target cloud platform. This includes: a TF storage chip or an MMC storage chip can be selected to implement the model storage device, which is connected to the object construction device, and is used to receive the machine learning-based artificial intelligence model customized for the target cloud platform, and store the machine learning-based artificial intelligence model customized for the target cloud platform;

[0093] Among them, receiving the machine learning-based artificial intelligence model customized for the target cloud platform and storing the machine learning-based artificial intelligence model customized for the target cloud platform includes: completing the model storage of the machine learning-based artificial intelligence model customized for the target cloud platform by storing various model parameters of the machine learning-based artificial intelligence model customized for the target cloud platform.

[0094] Fifth embodiment

[0095] Figure 6 This is a diagram showing the internal structure of a data cleaning system based on machine learning and data stream processing according to the fifth embodiment of the present invention.

[0096] like Figure 6 As shown, compared with Figure 5 , the data cleaning system based on machine learning and data stream processing also includes:

[0097] A timing server component is connected to the data analysis component and is used to provide timing services for the data analysis component to analyze the cloud product sales data stream and operation and maintenance data stream corresponding to the latest past time partition of the target cloud platform with the current time as the end time point;

[0098] Specifically, the timing server can have a built-in quartz oscillation unit for generating a reference clock signal that provides timing services for the data analysis device to analyze the cloud product sales data stream and operation and maintenance data stream corresponding to the latest past time partition of the target cloud platform with the current moment as the end time point.

[0099] Next, various embodiments of the present invention will be further described.

[0100] In each of the above embodiments, optionally, in the data cleaning system based on machine learning and data stream processing:

[0101] An artificial intelligence model based on machine learning is used to intelligently predict the number of duplicate log entries, the number of log errors, and the number of incomplete log entries corresponding to a future time partition of a target cloud platform starting at the current moment based on the duration of the time partition and the outputs of a data parsing device, an information extraction device, and a partition collection device. This includes: inputting the duration of the time partition, a cloud product sales data stream and an operation and maintenance data stream corresponding to the latest past time partition of the target cloud platform ending at the current moment, multiple pieces of platform-related information and multiple pieces of registered user-related information of the target cloud platform, and the number of duplicate log entries, the number of log errors, and the number of incomplete log entries corresponding to each past time partition of the target cloud platform into the artificial intelligence model based on machine learning in parallel;

[0102] For example, the duration of the time partition can be selected as 10 minutes. After the selection, the duration of each time partition is equal;

[0103] Among them, using a machine learning-based artificial intelligence model to intelligently predict the number of log duplicates, log errors, and log incompleteness corresponding to a future time partition of a target cloud platform starting at the current moment based on the duration of the time partition and the outputs of three devices: a data parsing device, an information extraction device, and a partition collection device also includes: executing the machine learning-based artificial intelligence model to obtain the number of log duplicates, log errors, and log incompleteness corresponding to a future time partition of the target cloud platform starting at the current moment output by the machine learning-based artificial intelligence model;

[0104] Among them, the duration of the time partition, the cloud product sales data flow and operation and maintenance data flow corresponding to the latest past time partition of the target cloud platform with the current moment as the end time point, multiple platform-related information and multiple registered user-related information of the target cloud platform, the number of duplicate log entries, the number of error entries in each log, and the number of incomplete log entries corresponding to each past time partition of the target cloud platform are input into the artificial intelligence model based on machine learning in parallel, including: the duration of the time partition, the cloud product sales data flow and operation and maintenance data flow corresponding to the latest past time partition of the target cloud platform with the current moment as the end time point, multiple platform-related information and multiple registered user-related information of the target cloud platform, the number of duplicate log entries, the number of error entries in each log, and the number of incomplete log entries corresponding to each past time partition of the target cloud platform are converted into binary values ​​and then input into the artificial intelligence model based on machine learning in parallel;

[0105] Wherein, executing the artificial intelligence model based on machine learning to obtain the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries corresponding to the future time partition of the target cloud platform output by the artificial intelligence model based on machine learning with the current moment as the starting time point includes: the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries corresponding to the future time partition of the target cloud platform output by the artificial intelligence model based on machine learning with the current moment as the starting time point are all expressed in binary numerical values;

[0106] Among them, the duration of the time partition, the cloud product sales data flow and operation and maintenance data flow corresponding to the latest past time partition of the target cloud platform with the current moment as the end time point, multiple platform-related information and multiple registered user-related information of the target cloud platform, the number of duplicate log entries, the number of error entries in each log, and the number of incomplete log entries corresponding to each past time partition of the target cloud platform are converted into binary values ​​and then input in parallel into the artificial intelligence model based on machine learning, including: using different programmable logic devices to respectively realize binary value conversion and parallel input;

[0107] For example, using different programmable logic devices to respectively implement binary value conversion and parallel input includes: the different programmable logic devices may be CPLD devices of different models.

[0108] And in each of the above embodiments, optionally, in the data cleaning system based on machine learning and data stream processing:

[0109] The artificial intelligence model based on machine learning is a deep neural network that has completed multiple learning cycles. The deep neural network includes multiple hidden layers, an input layer, and an output layer. In the deep neural network, the multiple hidden layers are located between the input layer and the output layer, and the number of hidden layers is positively correlated with the total number of product types on sale on the target cloud platform. This includes: using an FPGA chip designed in VHDL language to complete testing and simulation of a numerical conversion process in which the number of hidden layers is positively correlated with the total number of product types on sale on the target cloud platform;

[0110] For example, when the total number of product types sold on the target cloud platform is 100,000, the number of hidden layers is 3; when the total number of product types sold on the target cloud platform is 500,000, the number of hidden layers is 5; when the total number of product types sold on the target cloud platform is 1 million, the number of hidden layers is 7; when the total number of product types sold on the target cloud platform is 5 million, the number of hidden layers is 9, and so on;

[0111] The method includes constructing a machine learning-based artificial intelligence model customized for a target cloud platform, wherein the number of model learning times is positively correlated with the total number of registered users of the target cloud platform, and comprising: using an information conversion function to represent an information conversion relationship in which the number of learning times of the machine learning-based artificial intelligence model is positively correlated with the total number of registered users of the target cloud platform;

[0112] And wherein, the information conversion relationship that uses an information conversion function to represent the positive correlation between the number of learning times of the machine learning-based artificial intelligence model and the total number of registered users of the target cloud platform includes: in the information conversion function, the total number of registered users of the target cloud platform is the input content, and the number of learning times of the machine learning-based artificial intelligence model that is positively correlated with the total number of registered users of the target cloud platform is the output content.

[0113] Sixth embodiment

[0114] Figure 7 The present invention is a flowchart showing the steps of a data cleaning method based on machine learning and data stream processing according to the sixth embodiment of the present invention.

[0115] like Figure 7 As shown, the data cleaning method based on machine learning and data stream processing includes the following steps:

[0116] Step S701: parsing the cloud product sales data stream and operation and maintenance data stream corresponding to the latest past time partition of the target cloud platform with the current time as the end time point, where the target cloud platform is a cloud service-based sales platform used by the designated e-commerce company;

[0117] For example, different cloud service-based sales platforms can be customized for different e-commerce companies. In fact, different cloud services are customized for different e-commerce companies, thereby providing a platform for the e-commerce companies to complete the sales of their products based on cloud services.

[0118] Step S702: extracting multiple pieces of platform-related information and multiple pieces of registered user-related information of the target cloud platform;

[0119] For example, a first extraction component and a second extraction component may be used to respectively extract multiple pieces of platform association information of the target cloud platform and multiple pieces of registered user association information of the target cloud platform;

[0120] Step S703: Collect the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries corresponding to each past time partition of the target cloud platform;

[0121] For example, when the duration of each time partition is 10 minutes, the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries corresponding to each past time partition of the target cloud platform are collected. If the current time is 5:00 PM, then the past time partitions are 4:50 PM to 5:00 PM, 4:40 PM to 4:50 PM, 4:30 PM to 4:40 PM, 4:20 PM to 4:30 PM, 4:10 PM to 4:20 PM, 4:00 PM to 4:10 PM, and so on.

[0122] For example, the number of time partitions of each selected past time partition is positively correlated with the total number of registered users of the target cloud platform. For example, if the total number of registered users of the target cloud platform is about 2 million, the number of time partitions of each selected past time partition is 6. If the total number of registered users of the target cloud platform is about 5 million, the number of time partitions of each selected past time partition is 8. If the total number of registered users of the target cloud platform is about 10 million, the number of time partitions of each selected past time partition is 10. If the total number of registered users of the target cloud platform is about 20 million, the number of time partitions of each selected past time partition is 12, and so on.

[0123] Step S704: Build a machine learning-based artificial intelligence model customized for the target cloud platform, where the number of model learning times is positively correlated with the total number of registered users of the target cloud platform;

[0124] For example, a machine learning-based artificial intelligence model customized for a target cloud platform is constructed, and the positive correlation between the number of model learning times and the total number of registered users of the target cloud platform includes: when the total number of registered users of the target cloud platform is around 2 million, the corresponding number of model learning times is 2,000; when the total number of registered users of the target cloud platform is around 5 million, the corresponding number of model learning times is 2,500; when the total number of registered users of the target cloud platform is around 10 million, the corresponding number of model learning times is 3,000; when the total number of registered users of the target cloud platform is around 20 million, the corresponding number of model learning times is 4,000, and so on;

[0125] Step S705: Using a machine learning-based artificial intelligence model, based on the duration of the time partition, the cloud product sales data flow and operation and maintenance data flow corresponding to the latest past time partition of the target cloud platform with the current moment as the end time point, multiple pieces of platform-related information and multiple pieces of registered user-related information of the target cloud platform, and the number of duplicate log entries, the number of error entries, and the number of incomplete log entries corresponding to each past time partition of the target cloud platform, intelligently predict the number of duplicate log entries, the number of error entries, and the number of incomplete log entries corresponding to the future time partition of the target cloud platform with the current moment as the start time point;

[0126] Step S706: Determine the number of cleaning servers to be allocated to the target cloud platform in the future time partition based on the cumulative values ​​of the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries predicted by the intelligent method;

[0127] The number of cleaning servers allocated to the target cloud platform in the determined future time partition is proportional to the cumulative number of duplicate log entries, error log entries, and incomplete log entries based on intelligent prediction, and the cleaning servers allocated to the target cloud platform are servers that perform log data cleaning tasks for the target cloud platform;

[0128] For example, the number of cleaning servers allocated to the target cloud platform for the determined future time partition is proportional to the cumulative value of the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries based on intelligent prediction, and the cleaning servers allocated to the target cloud platform are servers that perform the target cloud platform log data cleaning task, including: when the cumulative value of the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries predicted by the intelligent prediction is 100,000, the number of cleaning servers allocated to the target cloud platform for the determined future time partition is 1, when the cumulative value of the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries predicted by the intelligent prediction is 200,000, the number of cleaning servers allocated to the target cloud platform for the determined future time partition is 2, when the cumulative value of the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries predicted by the intelligent prediction is 300,000, the number of cleaning servers allocated to the target cloud platform for the determined future time partition is 3, and so on. The cleaning server not only needs to delete erroneous logs, but also needs to remove duplicate logs and restore incomplete logs to a complete state;

[0129] The target cloud platform's multiple platform-related information includes the target cloud platform's establishment time, the total number of product types on sale, and the total number of platform sales regions; the target cloud platform's multiple registered user-related information includes the target cloud platform's total number of registered users, each registered user's age information, gender identifier, and the total number of products purchased before the current moment;

[0130] The cloud product sales data stream corresponding to each time partition of the target cloud platform is the total amount of products sold by the target cloud platform in the time partition, the number of products corresponding to each product type, and the region numbers corresponding to each sales region involved; the operation and maintenance data stream corresponding to each time partition of the target platform is the number of product inquiries received by the target cloud platform in the time partition, the number of operation and maintenance personnel used, the duration of the server downtime event, and the number of servers affected when the server downtime event occurs;

[0131] Among them, in each learning of the machine learning-based artificial intelligence model customized for the target cloud platform, the number of known log duplicates, the number of known log errors, and the number of known log incomplete entries corresponding to a certain past time partition of the target cloud platform are used as the output content of the machine learning-based artificial intelligence model, and the cloud product sales data flow and operation and maintenance data flow corresponding to the past time partition of the target cloud platform before and immediately adjacent to the certain past time partition, multiple platform-related information and multiple registered user-related information of the target cloud platform, and the number of log duplicates, the number of log errors, and the number of incomplete log entries corresponding to each past time partition of the target cloud platform before the certain past time partition are used as the input content of the machine learning-based artificial intelligence model to complete this learning;

[0132] The machine learning-based artificial intelligence model is a deep neural network that has completed multiple learning cycles. The deep neural network includes multiple hidden layers, an input layer, and an output layer. In the deep neural network, the multiple hidden layers are located between the input layer and the output layer, and the number of hidden layers is positively correlated with the total number of product types on sale on the target cloud platform.

[0133] For example, the artificial intelligence model based on machine learning is a deep neural network after completing multiple learning processes, wherein the deep neural network includes multiple hidden layers, an input layer, and an output layer. In the deep neural network, the multiple hidden layers are located between the input layer and the output layer, and the number of hidden layers is positively correlated with the total number of product types on sale on the target cloud platform, including: when the total number of product types on sale on the target cloud platform is 100,000, the number of hidden layers is 3; when the total number of product types on sale on the target cloud platform is 500,000, the number of hidden layers is 5; when the total number of product types on sale on the target cloud platform is 1 million, the number of hidden layers is 7; when the total number of product types on sale on the target cloud platform is 5 million, the number of hidden layers is 9, and so on.

[0134] In addition, in a data cleaning system and method based on machine learning and data stream processing according to the present invention:

[0135] In the deep neural network, the number of hidden layers is positively correlated with the total number of product types on sale on the target cloud platform, including: using a numerical mapping function to represent a numerical mapping relationship of the positive correlation between the number of hidden layers in the deep neural network and the total number of product types on sale on the target cloud platform;

[0136] For example, a numerical simulation mode may be selected to implement testing and simulation of a data processing process that uses a numerical mapping function to represent a numerical mapping relationship that is positively correlated between the number of hidden layers of the deep neural network and the total number of product types on sale on the target cloud platform.

[0137] The method of using a numerical mapping function to represent a numerical mapping relationship in which the number of hidden layers of the deep neural network and the total number of product types on sale on the target cloud platform are positively correlated includes: in the numerical mapping function, the total number of product types on sale on the target cloud platform is an input value of the numerical mapping function;

[0138] And wherein, the numerical mapping relationship of using a numerical mapping function to represent the positive correlation between the number of hidden layers of the deep neural network and the total number of product types on sale on the target cloud platform also includes: in the numerical mapping function, the number of hidden layers of the deep neural network that is positively correlated with the total number of product types on sale on the target cloud platform is the output value of the numerical mapping function.

[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and description of the present invention.

Claims

1. A data cleaning system based on machine learning and data stream processing, characterized in that: The system comprises: A data parsing component, used to parse the cloud product sales data stream and operation and maintenance data stream corresponding to the latest past time partition of the target cloud platform with the current time as the end time point. The target cloud platform is a cloud service-based sales platform used by the designated e-commerce company; An information extraction device, used to extract multiple pieces of platform-related information and multiple pieces of registered user-related information of the target cloud platform; Partition collection device, used to collect the number of duplicate log entries, the number of error entries, and the number of incomplete log entries corresponding to each past time partition of the target cloud platform; An object building component is used to build a machine learning-based artificial intelligence model customized for the target cloud platform. The number of model learning times is positively correlated with the total number of registered users of the target cloud platform. The content prediction device is connected to the data analysis device, the information extraction device, the partition collection device, and the object construction device respectively, and is used to use a machine learning-based artificial intelligence model to intelligently predict the number of duplicate log entries, the number of log errors, and the number of incomplete log entries corresponding to the future time partition of the target cloud platform starting at the current time point based on the duration of the time partition and the output of the data analysis device, the information extraction device, and the partition collection device; A resource allocation device, connected to the content prediction device, is used to determine the number of cleaning servers to be allocated to the target cloud platform in the future time partition based on the accumulated values ​​of the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries predicted by the intelligent prediction; Among them, in each learning of the machine learning-based artificial intelligence model customized for the target cloud platform, the number of known log duplicates, the number of known log errors, and the number of known log incomplete entries corresponding to a certain past time partition of the target cloud platform are used as the output content of the machine learning-based artificial intelligence model, and the cloud product sales data flow and operation and maintenance data flow corresponding to the past time partition of the target cloud platform before and immediately adjacent to the certain past time partition, multiple platform-related information and multiple registered user-related information of the target cloud platform, and the number of log duplicates, the number of log errors, and the number of incomplete log entries corresponding to each past time partition of the target cloud platform before the certain past time partition are used as the input content of the machine learning-based artificial intelligence model to complete this learning; Among them, the artificial intelligence model based on machine learning is a deep neural network after completing multiple learning processes. The deep neural network includes multiple hidden layers, an input layer and an output layer. In the deep neural network, the multiple hidden layers are located between the input layer and the output layer, and the number of hidden layers is positively correlated with the total number of product types on sale on the target cloud platform.

2. The data cleaning system based on machine learning and data stream processing according to claim 1, characterized in that: The number of cleaning servers allocated to the target cloud platform in the determined future time partition is proportional to the cumulative number of duplicate log entries, error log entries, and incomplete log entries based on intelligent prediction, and the cleaning servers allocated to the target cloud platform are servers that perform log data cleaning tasks for the target cloud platform; The target cloud platform's multiple platform-related information includes the target cloud platform's establishment time, the total number of product types on sale, and the total number of platform sales regions; the target cloud platform's multiple registered user-related information includes the target cloud platform's total number of registered users, each registered user's age information, gender identifier, and the total number of products purchased before the current moment; Among them, the cloud product sales data flow corresponding to the target cloud platform in each time partition is the total amount of products sold by the target cloud platform in the said time partition, the number of products corresponding to each product type, and the regional numbers corresponding to each sales region involved. The operation and maintenance data flow corresponding to the target platform in each time partition is the number of product inquiries received by the target cloud platform in the said time partition, the number of operation and maintenance personnel used, the duration of the server downtime event, and the number of servers affected when the server downtime event occurs.

3. The data cleaning system based on machine learning and data stream processing according to claim 2, characterized in that: The system further comprises: A status warning device is connected to the resource allocation device and is used to execute a warning operation corresponding to insufficient number of cleaning servers when the number of existing cleaning servers is less than the number of cleaning servers allocated to the target cloud platform in the future time partition; The status warning device is further configured to temporarily suspend the warning operation corresponding to insufficient number of cleaning servers when the number of existing cleaning servers is greater than or equal to the number of cleaning servers allocated to the target cloud platform in the future time partition.

4. The data cleaning system based on machine learning and data stream processing according to claim 2, characterized in that: The system further comprises: The on-site display device is connected to the resource allocation device, and is used to receive the number of cleaning servers allocated to the target cloud platform in the future time partition, and perform on-site display of the number of cleaning servers allocated to the target cloud platform in the future time partition.

5. The data cleaning system based on machine learning and data stream processing according to claim 2, characterized in that: The system further comprises: a model storage device connected to the object construction device, configured to receive the machine learning-based artificial intelligence model customized for the target cloud platform and store the machine learning-based artificial intelligence model customized for the target cloud platform; Among them, receiving the machine learning-based artificial intelligence model customized for the target cloud platform and storing the machine learning-based artificial intelligence model customized for the target cloud platform includes: completing the model storage of the machine learning-based artificial intelligence model customized for the target cloud platform by storing various model parameters of the machine learning-based artificial intelligence model customized for the target cloud platform.

6. The data cleaning system based on machine learning and data stream processing according to claim 2, characterized in that: The system further comprises: The timing server component is connected to the data analysis device and is used to provide timing services for the data analysis device to analyze the cloud product sales data flow and operation and maintenance data flow corresponding to the latest past time partition of the target cloud platform with the current time as the end time point.

7. The data cleaning system based on machine learning and data stream processing according to claim 2, characterized in that: The duration of the time partition, the cloud product sales data stream and operation and maintenance data stream corresponding to the latest past time partition of the target cloud platform with the current moment as the end time point, multiple platform-related information and multiple registered user-related information of the target cloud platform, the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries corresponding to each past time partition of the target cloud platform are input into the machine learning-based artificial intelligence model in parallel, and the machine learning-based artificial intelligence model is executed to obtain the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries corresponding to the future time partition of the target cloud platform with the current moment as the start time point output by the machine learning-based artificial intelligence model; Among them, the duration of the time partition, the cloud product sales data flow and operation and maintenance data flow corresponding to the latest past time partition of the target cloud platform with the current moment as the end time point, multiple platform-related information and multiple registered user-related information of the target cloud platform, the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries corresponding to each past time partition of the target cloud platform are converted into binary values ​​and then input in parallel into the machine learning-based artificial intelligence model. In addition, the number of duplicate log entries, the number of error log entries, and the number of incomplete log entries corresponding to the future time partition of the target cloud platform with the current moment as the start time point output by the machine learning-based artificial intelligence model are all expressed in binary values; Among them, different programmable logic devices are used to realize binary value conversion and parallel input respectively.

8. The data cleaning system based on machine learning and data stream processing according to claim 2, characterized in that: The artificial intelligence model based on machine learning is a deep neural network that has completed multiple learning cycles. The deep neural network includes multiple hidden layers, an input layer, and an output layer. In the deep neural network, the multiple hidden layers are located between the input layer and the output layer, and the number of hidden layers is positively correlated with the total number of product types on sale on the target cloud platform. This includes: using an FPGA chip designed in VHDL language to complete testing and simulation of a numerical conversion process in which the number of hidden layers is positively correlated with the total number of product types on sale on the target cloud platform; The method includes constructing a machine learning-based artificial intelligence model customized for a target cloud platform, wherein the number of model learning times is positively correlated with the total number of registered users of the target cloud platform, and comprising: using an information conversion function to represent an information conversion relationship in which the number of learning times of the machine learning-based artificial intelligence model is positively correlated with the total number of registered users of the target cloud platform; Among them, the information conversion relationship represented by the information conversion function, which shows that the number of learning times of the artificial intelligence model based on machine learning is positively correlated with the total number of registered users of the target cloud platform, includes: in the information conversion function, the total number of registered users of the target cloud platform is the input content, and the number of learning times of the artificial intelligence model based on machine learning that is positively correlated with the total number of registered users of the target cloud platform is the output content.

9. A data cleaning method based on machine learning and data stream processing, the method implementing the system according to claim 1, characterized in that: The method comprises: Analyze the cloud product sales data stream and operation and maintenance data stream corresponding to the latest past time partition of the target cloud platform with the current time as the end time point. The target cloud platform is the cloud service-based sales platform used by the designated e-commerce company. Extract multiple platform-related information and multiple registered user-related information of the target cloud platform; Collect the number of duplicate log entries, error entries, and incomplete log entries for each past time partition of the target cloud platform. Build a machine learning-based artificial intelligence model customized for the target cloud platform, with the number of model learning times positively correlated with the total number of registered users on the target cloud platform; Adopting a machine learning-based artificial intelligence model, based on the duration of the time partition, the cloud product sales data flow and operation and maintenance data flow corresponding to the latest past time partition of the target cloud platform with the current moment as the end time point, multiple platform-related information and multiple registered user-related information of the target cloud platform, and the number of duplicate log entries, error entries, and incomplete log entries corresponding to each past time partition of the target cloud platform, intelligently predict the number of duplicate log entries, error entries, and incomplete log entries corresponding to the future time partition of the target cloud platform with the current moment as the start time point; Determine the number of cleaning servers to be allocated to the target cloud platform in future time partitions based on the cumulative values ​​of the number of duplicate log entries, number of log errors, and number of incomplete log entries predicted by intelligent forecasting; Among them, in each learning of the machine learning-based artificial intelligence model customized for the target cloud platform, the number of known log duplicates, the number of known log errors, and the number of known log incomplete entries corresponding to a certain past time partition of the target cloud platform are used as the output content of the machine learning-based artificial intelligence model, and the cloud product sales data flow and operation and maintenance data flow corresponding to the past time partition of the target cloud platform before and immediately adjacent to the certain past time partition, multiple platform-related information and multiple registered user-related information of the target cloud platform, and the number of log duplicates, the number of log errors, and the number of incomplete log entries corresponding to each past time partition of the target cloud platform before the certain past time partition are used as the input content of the machine learning-based artificial intelligence model to complete this learning; Among them, the artificial intelligence model based on machine learning is a deep neural network after completing multiple learning processes. The deep neural network includes multiple hidden layers, an input layer and an output layer. In the deep neural network, the multiple hidden layers are located between the input layer and the output layer, and the number of hidden layers is positively correlated with the total number of product types on sale on the target cloud platform.

Citation Information

Patent Citations

  • Artificial intelligence-based short message data cleaning method

    CN107463639A

  • Big data cleaning method and system based on artificial intelligence

    CN117891812A

  • Data cleaning task processing method and device

    CN113360270A

  • Business website monitoring system based on natural language processing

    CN115438183A