Big data analysis and processing system based on deep learning
Through a big data analysis and processing system based on deep learning, the problems of slow data analysis and unintuitive results of enterprise are solved, and fast and accurate data collection and prediction analysis are achieved, supporting the rapid decision-making and management of enterprises.
Patent Information
- Application Number
- PCT/CN2024/123869
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-25
- Filing Date
- 2024-10-10
- Publication Date
- 2025-07-03
AI Technical Summary
The existing technology analyzes slowly when processing large amounts of enterprise data, which is prone to errors, and the analysis results are not intuitive, making it difficult to support the rapid decision-making of enterprises.
The big data analysis and processing system based on deep learning is adopted, including data acquisition module, deep learning model, prediction analysis module, result display module and data storage terminal. Through data acquisition, deep learning model analysis, prediction evaluation and histogram display, rapid data processing and intuitive analysis are achieved.
It realizes rapid collection, detailed analysis and prediction of enterprise operation data. Through histogram display, it supports enterprises to adjust and plan in a timely manner, and improves the efficiency and accuracy of data processing.
Smart Images

Figure CN2024123869_03072025_PF_FP_ABST
Abstract
Description
Big data analysis and processing system based on deep learning Technical Field
[0001] The present invention belongs to the field of data analysis technology, specifically to a big data analysis and processing system based on deep learning. Background Art
[0002] Enterprise big data analysis refers to the use of big data technology to analyze various types of enterprise data in order to extract valuable information and insights to support the decision-making and business development of the enterprise. The main purpose of enterprise big data analysis is to improve the operational efficiency of the enterprise, reduce costs, optimize business processes, discover new business opportunities, and enhance the market competitiveness of the enterprise. In enterprise big data analysis, the data sources are very wide, including internal production data, sales data, financial data, etc., as well as external market data, competitor data, customer data, etc. These data can be integrated, cleaned, analyzed and mined in big data analysis to discover patterns in the data and provide support for the decision-making of the enterprise. Enterprise big data analysis requires the help of professional Tools and technologies such as data mining, machine learning, natural language processing, etc. At the same time, a sound data governance system needs to be established to ensure the accuracy and completeness of the data. In short, enterprise big data analysis is an important part of enterprise digital transformation. It can help enterprises better understand and utilize data, improve their operational efficiency and competitiveness. The general data processing methods on the market can only process some simple data files and analyze the data. However, when there is too much data in the enterprise, it is easy to cause slow data analysis and errors in the analysis. It is inconvenient to call out the analysis results, and the results are not intuitive. It is only the collection and analysis of the current data. In this regard, we propose a big data analysis and processing system based on deep learning. Technical issues
[0003] In response to the shortcomings of the existing technology, the present invention provides a big data analysis and processing system based on deep learning to solve the above technical problems. Technical Solutions
[0004] To achieve the above objectives, the present invention provides the following technical solutions: a big data analysis and processing system based on deep learning, the analysis and processing system includes: a deep learning model, a data acquisition module, a prediction and analysis module, a result display module, a data storage terminal and a comparison module, and the analysis and processing system is built inside a computer; the analysis and processing steps are as follows:
[0005] S1. The data collection module collects the enterprise's operational data for the current month and uploads the collected data to the deep learning model;
[0006] S2. The deep learning model processes and analyzes the data to obtain detailed operational information within the current enterprise. The forecasting and analysis module then conducts a forecast and evaluation of the current month's operational data to obtain the operational data value for the next month.
[0007] S3. Upload the analysis result data of the current month to the data storage terminal for storage;
[0008] S4. Retrieve the historical data of the previous month from the data storage terminal and compare the data of the previous month with the data of the current month through the comparison module;
[0009] S5. The comparison results and the predicted operational data values for the next month are made into a histogram and uploaded to the result display module for enterprise personnel to review and analyze, supplement the analysis, strengthen management, and formulate future plans.
[0010] Preferably, the data collection module in step S1 collects log information recorded in the application by collecting logs from the computer operating the enterprise, and obtains the enterprise's operating data for the current month;
[0011] When operating computers process enterprise operating data, they generate various log information inside the computers. Flume collects the computer log information into the deep learning model.
[0012] When collecting data, standardize the log format, including timestamp, log level, and message content.
[0013] Prioritize, in step S2, the deep learning model processes data through a neural network model, learns complex feature representations, and performs classification and regression on the data;
[0014] The steps to build a neural network model are to import relevant modules, specify the input features and labels of the training set, build the network structure, describe each layer of the network layer by layer, build a sequential network structure where the upper layer output is the lower layer input, and create an initialization function;
[0015] The formula is expressed as determining the number of neural network layers, the number of neurons and the activation function, randomly initializing the neuron weights, passing the input data to the output layer, calculating the neural network loss function, calculating the gradient of the neurons based on the loss function, and repeatedly iterating the above process to reduce the value of the loss function.
[0016] First, the formula from input layer to hidden layer is:
[0017] Where z(1) is the output result of the first layer, represents the weight matrix of layer 1, represents the output result of the l-1 layer, Represented as the bias term of layer 1;
[0018] The formula from hidden layer to output layer is: in Represented as the output result of the first layer, Expressed as activation function, Represented as the output result of the z-1th layer.
[0019] Preferably, the data processed and analyzed in step S2 include sales data, financial data and customer data. The sales data includes sales volume and sales volume. The financial data includes income, expenditure, cost and profit. The customer data includes the number of customers, customer churn rate and customer praise rate.
[0020] Preferably, the method of saving data at the data storage end in step S3 is to save data through cloud storage. The comparison module in step S4 reveals the differences between the data by comparing multiple data, and understands the development and changes of things. The comparison step of the comparison module is to determine the comparison object, select two or more data as the comparison objects, which are the operation data of the previous month and the operation data of the current month, collect the operation data of the current month through the data collection module, obtain the comparison data by calling the historical operation data from the data storage end, clean the data, remove outliers and missing values, and perform comparative analysis on the data, wherein one set of data is coded as (1010111001) and the other set of data is coded as (0111000110), and the data codes are compared to obtain the differences.
[0021] Preferably, the calculation formula of the comparison module is:
[0022] Where x is the value of the indicator, y is the mean, and a is the standard deviation;
[0023] Preferably, the prediction analysis module in step S2 uses a logistic regression model for analysis. The calculation formula of the logistic regression model is:
[0024]
[0025] Where p represents the probability of the prediction result being 1, b0, b1…bn are model parameters, and q is the feature variable;
[0026] By learning sample data, we can find the optimal parameters and calculate the sum of the probability of the maximum sample data predicting the result to be 1 and the probability of the predicted result to be 0.
[0027] Prioritizing the steps for establishing a predictive analysis module is to collect relevant data as data for training the model, pre-process the data, select features that are closely related to the prediction target, train the logistic regression model through programming, continuously adjust the model parameters to continuously optimize the model, and deploy the model for actual use;
[0028] Data preprocessing includes data deduplication, data conversion, and data integration. The data deduplication step includes detecting and analyzing data to identify quality issues, defining deduplication rules based on quality issues found in data analysis, using clustering algorithms to deduplicate data, executing deduplication, and performing data cleaning operations;
[0029] The data conversion step converts the format of the deduplicated data and verifies it after conversion to ensure the correctness of the converted format.
[0030] The data integration step is the process of selecting and extracting a specific subset from the data source set. By relying on data extraction, only relevant data can be accurately copied from large amounts of data, and the extracted specific data subset can be sent to the destination location. By relying on data transmission, the circulation and sharing of data can be automatically maintained. For the data transmitted directly, data format, data encoding, and data consistency are cleaned to ensure the standardization of data in the central database. The cleaned data will be associated according to the new data organization logic to strengthen the internal connection of the data. According to the needs of the subject database layer, some data subsets in the central database will be regularly published to the subject database layer.
[0031] Prioritize, the histogram establishment step in step S5 is to collect and process the data, determine the number of groups and group intervals in the histogram based on the data analysis, make a frequency table based on the determined number of groups and group intervals, make a histogram, the height of the column represents the frequency of the group, the width of the column represents the group interval of the group, draw a cumulative frequency curve, and analyze and interpret the data. Beneficial effects
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] The present invention collects, uploads and analyzes the enterprise's operating data for the current month through the data acquisition module, and analyzes and predicts the operating data for the current month through the set prediction analysis module, and infers the operating data situation for the next month. The operating data of the current month is compared and analyzed with the operating data of the previous month through the comparison module, and is plotted into a histogram for enterprise personnel to view. It can quickly process and analyze the data, and compare and analyze the future data situation through historical data and analysis prediction, so that enterprise personnel can better understand the situation within the enterprise, adjust and plan strategies in a timely manner, and bring better usage prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] FIG1 is a framework diagram of the analysis and processing steps of the present invention. Modes for Carrying Out the Invention
[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0036] The present invention provides a technical solution: a big data analysis and processing system based on deep learning. The analysis and processing system includes: a deep learning model, a data acquisition module, a prediction analysis module, a result display module, a data storage terminal and a comparison module. The analysis and processing system is built inside a computer; the analysis and processing steps are:
[0037] S1. The data collection module collects the enterprise's operational data for the current month and uploads the collected data to the deep learning model;
[0038] S2. The deep learning model processes and analyzes the data to obtain detailed operational information within the current enterprise. The forecasting and analysis module then conducts a forecast and evaluation of the current month's operational data to obtain the operational data value for the next month.
[0039] S3. Upload the analysis result data of the current month to the data storage terminal for storage;
[0040] S4. Retrieve the historical data of the previous month from the data storage terminal and compare the data of the previous month with the data of the current month through the comparison module;
[0041] S5. The comparison results and the predicted operational data values for the next month are made into a histogram and uploaded to the result display module for enterprise personnel to review and analyze, supplement the analysis, strengthen management, and formulate future plans.
[0042] Furthermore, in step S1, the data collection module collects the log information recorded in the application by collecting the logs inside the computer operated by the enterprise, and obtains the operating data of the enterprise for the current month;
[0043] When operating computers process enterprise operating data, they generate various log information inside the computers. Flume collects the computer log information into the deep learning model.
[0044] Flume is a highly available, reliable, and distributed system provided by Cloudera for collecting, aggregating, and transmitting massive logs. It supports customizing various data senders within the log system to collect data, and also provides the ability to perform simple data processing and write data to various data receivers.
[0045] Flume uses a multi-master approach to ensure the consistency of configuration data. Flume introduces ZooKeeper to store configuration data. ZooKeeper itself can ensure the consistency and high availability of configuration data. In addition, when the configuration data changes, ZooKeeper can notify the Flume Master node. The Flume Master uses the gossip protocol to synchronize data.
[0046] The most obvious change in Flume is the removal of the centrally managed configuration of Master and Zookeeper, transforming it into a pure transmission tool. Another major difference in Flume is that reading and writing data are handled by different worker threads. In Flume, the reading thread also performs the writing work. If the writing is slow, it will block Flume's ability to receive data. This asynchronous design allows the reading thread to work smoothly without having to worry about any downstream problems.
[0047] When collecting data, standardize the log format, including timestamp, log level, and message content.
[0048] Furthermore, in step S2, the deep learning model processes the data through a neural network model, learns complex feature representations, and performs classification and regression on the data;
[0049] The steps to build a neural network model are to import relevant modules, specify the input features and labels of the training set, build the network structure, describe each layer of the network layer by layer, build a sequential network structure where the upper layer output is the lower layer input, and create an initialization function;
[0050] The formula is expressed as determining the number of neural network layers, the number of neurons and the activation function, randomly initializing the neuron weights, passing the input data to the output layer, calculating the neural network loss function, calculating the gradient of the neurons based on the loss function, and repeatedly iterating the above process to reduce the value of the loss function;
[0051] There are four basic characteristics of neural network models;
[0052] Nonlinear
[0053] Nonlinear relationships are a universal feature of nature. Brain intelligence is a nonlinear phenomenon. Artificial neurons are in two different states of activation or inhibition, which is mathematically manifested as a nonlinear relationship. Networks composed of neurons with thresholds have better performance to improve fault tolerance and storage capacity.
[0054] Unrestricted
[0055] Neural networks are typically composed of multiple neurons that are extensively connected. The overall behavior of the system is determined not only by the characteristics of individual neurons, but also primarily by the interactions and connections between units. The large number of connections between units is used to simulate the non-restrictive nature of the brain. Associative memory is a typical example of non-restrictiveness.
[0056] Very high quality
[0057] Artificial neural networks have the ability to adapt, self-organize, and self-learn. Neural networks not only process information that can cause various changes, but also the nonlinear dynamic system itself changes while processing information. Iterative processes are often used to describe the evolution of dynamic systems.
[0058] Non-convexity
[0059] The evolutionary direction of the system depends on a specific state function under specific conditions, such as the energy function, whose extreme value corresponds to the relatively stable state of the system. Non-convexity means that a function has multiple extreme values, resulting in the system having multiple stable equilibrium states, leading to the diversity of system evolution.
[0060] Furthermore, the formula from input layer to hidden layer is:
[0061]
[0062] Where z(1) is the output result of the first layer, represents the weight matrix of layer 1, represents the output result of the l-1 layer, Represented as the bias term of layer 1;
[0063] The formula from hidden layer to output layer is:
[0064] in Represented as the output result of the first layer, Expressed as activation function, Represented as the output result of the z-1th layer.
[0065] Furthermore, the data processed and analyzed in step S2 include sales data, financial data and customer data. The sales data includes sales volume and sales volume, the financial data includes income, expenditure, cost and profit, and the customer data includes the number of customers, customer churn rate and customer praise rate.
[0066] Furthermore, in step S3, the method for saving data at the data storage end is to save data through cloud storage. In step S4, the comparison module reveals the differences between multiple data by comparing them, and understands the development and changes of things. The comparison step of the comparison module is to determine the comparison object, select two or more data as the comparison objects, which are the operation data of the previous month and the operation data of the current month, collect the operation data of the current month through the data collection module, obtain the comparison data by calling the historical operation data from the data storage end, clean the data, remove outliers and missing values, and compare and analyze the data. Set one set of data encoding as (1010111001) and the other set of data encoding as (0111000110), and compare the data encodings to obtain the differences.
[0067] Furthermore, the calculation formula of the comparison module is:
[0068] Where x is the value of the indicator, y is the mean, and a is the standard deviation;
[0069] Furthermore, the prediction analysis module in step S2 uses a logistic regression model for analysis. The calculation formula of the logistic regression model is:
[0070] Where p represents the probability of the prediction result being 1, b0, b1…bn are model parameters, and q is the feature variable;
[0071] By learning sample data, we can find the optimal parameters and calculate the sum of the probability of the maximum sample data predicting the result to be 1 and the probability of the predicted result to be 0.
[0072] Furthermore, the steps for establishing the prediction analysis module include collecting relevant data as data for training the model, preprocessing the data, selecting features closely related to the prediction target, training the logistic regression model through programming, and continuously adjusting the model parameters to continuously optimize the model. The model is then deployed and put into practical use.
[0073] Data preprocessing includes data deduplication, data conversion, and data integration. The data deduplication step includes detecting and analyzing data to identify quality issues, defining deduplication rules based on quality issues found in data analysis, using clustering algorithms to deduplicate data, executing deduplication, and performing data cleaning operations;
[0074] The data conversion step converts the format of the deduplicated data and verifies it after conversion to ensure the correctness of the converted format.
[0075] The data integration step is the process of selecting and extracting a specific subset from the data source set. By relying on data extraction, only relevant data can be accurately copied from large amounts of data, and the extracted specific data subset can be sent to the destination location. By relying on data transmission, the circulation and sharing of data can be automatically maintained. For the data transmitted directly, data format, data encoding, and data consistency are cleaned to ensure the standardization of data in the central database. The cleaned data will be associated according to the new data organization logic to strengthen the internal connection of the data. According to the needs of the subject database layer, some data subsets in the central database will be regularly published to the subject database layer.
[0076] Furthermore, the histogram establishment step in step S5 is to collect and process the data. According to the data analysis, the number of groups and group intervals in the histogram are determined. A frequency table is made based on the determined number of groups and group intervals. A histogram is made. The height of the column represents the frequency of the group, and the width of the column represents the group interval of the group. A cumulative frequency curve is drawn to analyze and interpret the data.
[0077] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0078] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A big data analysis and processing system based on deep learning, characterized in that The analysis and processing system includes: a deep learning model, a data acquisition module, a prediction and analysis module, a result display module, a data storage terminal, and a comparison module. The analysis and processing system is built inside the computer. The analysis and processing steps are as follows: S1. The data acquisition module collects the operation data within the enterprise in the current month and uploads the collected data to the inside of the deep learning model. S2. The deep learning model processes and analyzes the data to obtain the detailed operation situation within the current enterprise, and predicts and evaluates the operation data in the current month through the prediction and analysis module to obtain the operation data value for the next month. S3. Upload the analysis result data in the current month to the data storage terminal for storage and preservation. S4. Retrieve the historical data of the previous month from the data storage terminal, and compare and process the data of the previous month and the current month through the comparison module. S5. Present the comparison result and the predicted operation data value for the next month in the form of a histogram, and upload it to the result display module for enterprise personnel to view, analyze, supplement, strengthen management, and formulate future plans.
2. The big data analysis and processing system based on deep learning according to claim 1, wherein: In step S1, the data acquisition module collects the log information recorded in the application program by collecting the logs inside the computer for enterprise operation, and obtains the operation situation data within the enterprise in the current month. When the operation computer processes the enterprise operation data, various log information is generated inside the computer, and the log information of the computer is collected into the deep learning model through Flume. When collecting, the log format is unified, including timestamp, log level, and message content.
3. The big data analysis and processing system based on deep learning according to claim 1, wherein: In step S2, the deep learning model processes the data through a neural network model, learns complex feature representations, and classifies and regresses the data. The steps for building the neural network model are to import relevant modules, specify the input features of the training set and the labels of the training set, build the network structure, describe each layer of the network layer by layer, build a sequential network structure where the output of the upper layer is the input of the lower layer, and create an initialization function. The formula for building is expressed as determining the number of neural network layers, the number of neurons, and the activation function, randomly initializing the weights of the neurons, transmitting the input data to the output layer, calculating the loss function of the neural network, calculating the gradient of the neurons according to the loss function, and repeatedly iterating the above process to reduce the value of the loss function.
4. The big data analysis and processing system based on deep learning according to claim 3, characterized in that, The formula from the input layer to the hidden layer is as follows: where z(1) is the output result of the first layer, represents the weight matrix of the first layer, represents the output result of the (l - 1)-th layer, represents the bias term of the first layer; The formula from the hidden layer to the output layer is as follows: Among them Indicates the output result of the first layer, Denoted as an activation function, It is expressed as the output result of the (z - 1)-th layer.
5. The big data analysis and processing system based on deep learning according to claim 1, wherein: The data processed and analyzed in step S2 includes sales data, financial data, and customer data. The sales data includes sales volume and sales amount. The financial data includes income, expenditure, cost, and profit. The customer data includes the number of customers, customer churn rate, and customer satisfaction rate.
6. The big data analysis and processing system based on deep learning according to claim 1, characterized in that: In the method for the data storage end to save data in step S3, data is saved through cloud storage. In step S4, the comparison module reveals the differences between multiple data to understand the development and changes of things. The comparison steps of the comparison module are to determine the comparison objects, select two or more data as comparison objects, which are the operation data of the previous month and the operation data of the current month. The operation data of the current month is collected through the data collection module, and the comparison data is obtained by retrieving the historical operation data from the data storage end. The data is cleaned to remove outliers and missing values, and then the data is analyzed by comparison. Among them, a set of data is encoded as (1010111001) and another set of data is encoded as (0111000110), and the data encodings are compared to obtain the difference situation.
7. The big data analysis and processing system based on deep learning according to claim 6, characterized in that, The calculation formula of the comparison module is as follows: Where x is the value of the index, y is the mean value, and a is the standard deviation.
8. The big data analysis and processing system based on deep learning according to claim 1, wherein: In step S2, the prediction analysis module uses a logistic regression model for analysis. The calculation formula of the logistic regression model is: where p represents the probability that the prediction result is 1, b0, b1... bn are model parameters, and q is a feature variable; By learning the sample data, the optimal parameters are found, and the sum of the probabilities of the prediction results of 1 and 0 is obtained through the maximum sample data.
9. The big data analysis and processing system based on deep learning according to claim 8, characterized in that: The steps for establishing the prediction analysis module are to collect relevant data as the data for training the model, preprocess the data, select the features closely related to the prediction target, program and train the logistic regression model, continuously adjust the parameters of the model to continuously optimize the model, and deploy the model for actual use. Data preprocessing includes data deduplication, data transformation, and data integration. The data deduplication step includes detecting and analyzing the data used to identify quality problems, defining deduplication rules according to the quality problems found in the data analysis, using clustering algorithms to deduplicate the data, and performing the deduplication and cleaning operations on the data. The data transformation step performs format conversion on the deduplicated data and validates it after conversion to ensure the correctness of the converted format. The data integration step is the process of selecting and extracting a specific subset from the data source set. Depending on data extraction, relevant data can be accurately copied only from a large amount of data, and the process of sending the extracted specific data subset to the destination location. Depending on data transmission, the data flow and sharing can be automatically maintained. For the directly transmitted data, data cleaning can ensure the standardization of the data in the central database in terms of data format, data encoding, and data consistency. The cleaned data is associated according to the new data organization logic to strengthen the internal connection of the data. According to the requirements of the theme database layer, some data subsets in the central database are regularly published to the theme database layer.
10. The big data analysis and processing system based on deep learning according to claim 1, characterized in that: In step S5, the steps for establishing the histogram are to collect and process the data. According to the data analysis situation, determine the number of groups and the group interval in the histogram, make a frequency table according to the determined number of groups and group interval, make a histogram, where the height of the column represents the frequency of the group and the width of the column represents the group interval, draw the cumulative frequency curve, and analyze and interpret the data.
Citation Information
Patent Citations
Intelligent chart generation method and device, computer system and readable storage medium
CN112597745A
Automatic replenishment method and replenishment system based on sales prediction of intelligent commodity system
CN114219412A
Sales prediction method, tool, system and device, and storage medium
CN114626898A
Hardware fitting supply prediction method and system based on deep learning
CN116843378A
Cigarette demand prediction method and system based on deep learning
CN117217788A
Cited By
Optimization method of low-carbon production data, equipment and medium
CN120633959A
Signal shielding device management method combining neural network and block chain
CN121000330A
Real-time acquisition and quality evaluation method and system based on multi-source heterogeneous data
CN121542253A