Enterprise data monitoring system based on big data analysis
By using an enterprise data monitoring system based on big data analytics, the problems of duplication, errors, and missing data in enterprise data processing have been solved. Through categorized processing and differentiated repair, the efficiency and accuracy of data analysis have been improved, ensuring the integrity and uniqueness of the data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN YICAI INFORMATION TECH CO LTD
- Filing Date
- 2025-04-02
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies, when processing enterprise data, suffer from data duplication, errors, and missing information due to the diverse sources and massive complexity of the data, which reduces the efficiency and accuracy of data analysis.
An enterprise data monitoring system based on big data analytics is adopted, including modules for data acquisition, diagnosis, analysis, repair, and processing. By determining the data missing rate and similarity, data categories are distinguished, anomaly characterization values are calculated, and different repair methods are adopted for different anomaly trends to eliminate duplicate data and ensure data integrity and accuracy.
It improves the efficiency and accuracy of enterprise data analysis and processing, ensures the integrity and uniqueness of data, and avoids wasted computing power and space due to duplicate data.
Smart Images

Figure CN120316200B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data monitoring technology, and in particular to an enterprise data monitoring system based on big data analysis. Background Technology
[0002] With the acceleration of enterprise digital transformation, the volume of data is exploding, requiring enterprises to manage and analyze this data more efficiently. Data monitoring systems help enterprises optimize operational efficiency, improve the scientific basis of decision-making, and support business innovation by collecting, analyzing, and visualizing massive amounts of data in real time. For example, by analyzing historical sales data, enterprises can predict future market trends; by monitoring customer behavior, enterprises can optimize marketing strategies; and through real-time analysis of business data, enterprises can promptly identify potential risks, such as abnormal transactions or equipment failures, and take measures to control them. Data monitoring technology is developing towards intelligence, efficiency, and comprehensiveness. In the future, with the development of IoT and big data technologies, data monitoring will increasingly combine artificial intelligence and machine learning technologies to achieve more accurate prediction, analysis, and repair, making data application scenarios more extensive.
[0003] Chinese Patent Publication No. CN113961770A discloses an enterprise big data monitoring system and method. The enterprise big data monitoring system includes an enterprise data acquisition module, a primary filtering module, a peripheral data judgment module, a filtering control module, a manual screening module, a system data module, and a data promotion module. The enterprise data acquisition module is used to collect the latest data information needed by the enterprise in real time. The primary filtering module is used to perform primary filtering of the latest data based on predetermined data from the surrounding industrial chain set by the enterprise. The peripheral data judgment module is used to perform similarity comparison judgment on the primary filtered data based on template data according to the actual needs of the enterprise. The filtering control module is used to directly store the comparison data that meets the storage requirements in the system data module. This invention has advantages such as high efficiency in enterprise big data monitoring and good data protection effects.
[0004] Chinese Patent Publication No. CN108173711A discloses a method for monitoring data exchange within an enterprise's internal systems, comprising the following steps: a data acquisition step, in which a business system sends requests and data to a service module through a data interface module; a service processing step, in which the service module receives the requests and data from the business system, processes the data according to the requests, and sends it to the corresponding business system; and a data monitoring step, which includes the following steps: an error detection step, in which an error detection submodule detects errors occurring in the data acquisition step, the service processing step, and the data monitoring step; and a data statistics step, in which a data statistics submodule statistically analyzes the usage of information resources during data transmission between various business systems. This invention provides a method for monitoring data exchange within an enterprise's internal systems, enabling enterprises to supervise and control the data transmitted between business systems.
[0005] However, the following problems still exist in the existing technology.
[0006] When processing enterprise data, due to the diverse sources of enterprise data and the large and complex content of the data, there may be situations where data is duplicated, incorrect, or missing due to the inability to accurately identify it. This results in incomplete and inaccurate data, reducing the efficiency and accuracy of the enterprise's data analysis and processing. Summary of the Invention
[0007] To address this issue, the present invention provides an enterprise data monitoring system based on big data analysis. This system aims to solve the problem that when processing enterprise data, due to the diverse sources and vast and complex content of enterprise data, there may be instances of duplicate, erroneous, or missing data due to the inability to accurately identify the data. This results in incomplete and inaccurate data, reducing the efficiency and accuracy of enterprise data analysis and processing.
[0008] To achieve the above objectives, the present invention provides an enterprise data monitoring system based on big data analysis, comprising:
[0009] The data acquisition module is used to acquire text-based data uploaded by enterprises in order to determine the data missing rate and data consistency.
[0010] A data diagnostic module, connected to the data acquisition module, is used to determine the data category based on the data missing rate and data identity, so as to distinguish the data analysis methods.
[0011] A data analysis module, connected to the data diagnostic module, is used to respond to the data category determination result of the data diagnostic module, calculate the data anomaly characterization value based on the data missing rate and data identity, classify the anomaly trend of the data, determine the data repair method, or mark the available data.
[0012] The data repair module, which is connected to the data analysis module, divides the data into several abnormal text data frames for data with a strong abnormal trend, determines the abnormal text data frame group and the adjacent text data frames adjacent to the abnormal text data frame group, predicts the abnormal text data frame group based on the adjacent text data frames, and verifies the predicted abnormal text data frame group.
[0013] For data with weak anomaly trends, a trained language model is used for data repair to confirm successful data replacement.
[0014] A data processing module, which is connected to the data diagnosis module and the data repair module, is used to determine the availability of data and mark usable data based on the data classification results of the data diagnosis module and the data repair results of the data repair module.
[0015] Among them, the abnormal text data frame group is a text data frame group composed of several abnormal text data frames.
[0016] Furthermore, the data acquisition module determines the data missing rate and data consistency, including,
[0017] The data missing rate is the ratio of the amount of missing data to the total amount of data.
[0018] Keywords used to extract the data;
[0019] The similarity between the keywords and the title / topic of the data is used to determine the data identity.
[0020] Furthermore, the data diagnostic module determines the data category based on the data missing rate and data similarity, including:
[0021] If the data meets the diagnostic criteria, then the data category is determined to be an abnormal data category;
[0022] If the data does not meet the diagnostic criteria, the data category is determined to be non-abnormal data.
[0023] The diagnostic criteria are that the data missing rate is greater than a preset data missing rate threshold and / or the data identity is less than a preset data identity threshold.
[0024] Furthermore, the data diagnostic module distinguishes data analysis methods, wherein,
[0025] If the category is an abnormal data category, then the abnormal data characterization value is calculated based on the data missing rate and data identity, and the abnormal trend of the data is divided to determine the data repair method.
[0026] If the category is a non-abnormal data category, then mark the data as available.
[0027] Furthermore, the data analysis module calculates data anomaly representation values based on the data missing rate and data identity, including:
[0028] The first influencing factor is used to determine the ratio of the data missing rate to the preset data missing rate;
[0029] The ratio of the preset data identity degree to the data identity degree is used to determine the second influence factor;
[0030] The weighted sum of the first impact factor and the second impact factor is used to determine the data anomaly characterization value.
[0031] Furthermore, the data analysis module identifies abnormal trends in the data to determine data repair methods, wherein...
[0032] If the data anomaly characterization value is greater than the preset data anomaly characterization value threshold, the anomaly tendency trend is classified as a strong anomaly tendency trend, the data is divided into several abnormal text data frames, the abnormal text data frame group and the adjacent text data frames adjacent to the abnormal text data frame group are determined, the abnormal text data frame group is predicted based on the adjacent text data frames, and the predicted abnormal text data frame group is verified.
[0033] If the data anomaly representation value is less than or equal to the preset data anomaly representation value threshold, the anomaly tendency trend is classified as a weak anomaly tendency trend, and the trained language model is used for data repair to determine that the data replacement was successful.
[0034] Furthermore, the data repair module predicts the abnormal text data frame group based on the adjacent text data frames, including:
[0035] Used to determine adjacent text data frames to the abnormal text data frame group, including the first adjacent text data frame and the last adjacent text data frame;
[0036] Used to predict and replace the first text data frame of an abnormal text data frame group based on the first adjacent text data frame;
[0037] Used to predict and replace the last text data frame of an abnormal text data frame group based on the last adjacent text data frames;
[0038] Used to identify unpredicted abnormal text data frame groups, and to use the predicted first-end text data frame as the first-end adjacent text data frame.
[0039] The predicted tail text data frame is used as the tail adjacent text data frame for repeated prediction until all abnormal text data frames in the abnormal text data frame group have been replaced.
[0040] Furthermore, the data repair module examines the predicted abnormal text data frame groups, including,
[0041] Used to determine the abnormal text data group corresponding to the abnormal text data frame group obtained by replacement;
[0042] This is used to determine several sample data in the enterprise database that maintain thematic consistency with the abnormal text data group;
[0043] Used to determine the similarity between the abnormal text data group and the plurality of sample data;
[0044] Used to detect abnormal text data frame groups based on the similarity.
[0045] Furthermore, the data repair module, based on the similarity check of the abnormal text data frame group, wherein,
[0046] If the similarity is greater than the similarity threshold, the predicted abnormal text data frame group is determined to be successfully replaced data.
[0047] Furthermore, the data processing module determines the availability of the data, including,
[0048] Used to determine the duplication status of the successfully replaced data and the available data;
[0049] If the successfully replaced data is duplicated with available data, the successfully replaced data will be removed.
[0050] If the successfully replaced data does not overlap with the available data, then the successfully replaced data is determined to be available data.
[0051] Compared with existing technologies, this invention sets up a data acquisition module, a data diagnosis module, a data analysis module, a data repair module, and a data processing module. The data acquisition module determines the data missing rate and data consistency. The data diagnosis module determines the data category based on the missing rate and data consistency to differentiate data analysis methods. The data analysis module calculates data anomaly characteristics, classifies the data's anomaly trends to determine data repair methods, or marks usable data. The data repair module, for data with strong anomaly trends, divides the data into several anomalous text data frames, identifies anomalous text data frame groups and adjacent text data frames, predicts the anomalous text data frame groups based on the adjacent text data frames, and verifies the predicted anomalous text data frame groups. For data with weak anomaly trends, a trained language model is used for data repair to determine successfully replaced data. The data processing module determines the usability of the data and marks usable data. This invention solves the problems of incomplete, inaccurate, and duplicate data from multiple channels in enterprises by classifying and processing data.
[0052] In particular, data categories are determined based on data missing rate and data identity to classify enterprise data and process it in a targeted manner. In reality, enterprise data is mostly obtained from multiple channels, which can lead to data ambiguity, inconsistent quality, and damage to storage media resulting in lost frames or fragments. Existing technologies mostly process the data directly, but using the same method to process data of different categories will result in wasted computing power and reduced efficiency. Based on this, this invention considers first classifying enterprise data, marking non-abnormal data categories, and further classifying abnormal data categories to repair abnormal data, ensuring the integrity and accuracy of enterprise data, and improving the efficiency and accuracy of enterprise data analysis and processing.
[0053] In particular, by calculating data anomaly representation values, the abnormal trend of data is divided to determine the repair methods for enterprise data with different abnormal trend trends. When processing abnormal data, the same repair method is usually selected. However, in practice, in order to save computing power, different repair methods can be determined according to the abnormal trend of data. Based on this, the present invention calculates data anomaly representation values to divide the abnormal trend of data. For data with strong abnormal trend trends, the method of predicting abnormal text data frame groups is used to repair enterprise data. For data with weak abnormal trend trends, an existing trained recurrent neural network is used for repair, which ensures the integrity and accuracy of enterprise data and improves the efficiency and accuracy of enterprise data analysis and processing.
[0054] In particular, the present invention performs usability verification on the repaired data. Preferably, most enterprise data comes from multiple channels and has redundancy. If the repaired data is directly marked for application, data duplication will cause space waste and reduce data analysis efficiency. Based on this, the present invention removes duplicate data by verifying the data to maintain data uniformity, ensure the integrity and accuracy of enterprise data, and improve the efficiency and accuracy of enterprise data analysis and processing. Attached Figure Description
[0055] Figure 1 This is a schematic diagram of the structure of an enterprise data monitoring system based on big data analysis, as described in an embodiment of the invention.
[0056] Figure 2 A logical block diagram of the data diagnosis module in an embodiment of the invention for determining data categories based on the data missing rate and data similarity;
[0057] Figure 3 A logic block diagram of the data diagnostic module differentiation data analysis method according to an embodiment of the invention;
[0058] Figure 4 The data analysis module of the invention is used to classify the abnormal trend of the data in order to determine the logical block diagram of the data repair method;
[0059] Figure 5 A logic block diagram for determining data availability in the data processing module of an embodiment of the invention. Detailed Implementation
[0060] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0061] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0062] It should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the term "connection" should be interpreted broadly. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be a connection within two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0063] Please see Figure 1 , Figure 1This is a schematic diagram of the structure of an enterprise data monitoring system based on big data analytics, according to an embodiment of the invention. The enterprise data monitoring system based on big data analytics according to this embodiment of the invention includes:
[0064] The data acquisition module is used to acquire data in text format uploaded by enterprises in order to determine the data missing rate and data consistency. It is understood that the data acquisition has been authorized.
[0065] A data diagnostic module, connected to the data acquisition module, is used to determine the data category based on the data missing rate and data identity, so as to distinguish the data analysis methods.
[0066] A data analysis module, connected to the data diagnostic module, is used to respond to the data category determination result of the data diagnostic module, calculate the data anomaly characterization value based on the data missing rate and data identity, classify the anomaly trend of the data, determine the data repair method, or mark the available data.
[0067] The data repair module, which is connected to the data analysis module, divides the data into several abnormal text data frames for data with a strong abnormal trend, determines the abnormal text data frame group and the adjacent text data frames adjacent to the abnormal text data frame group, predicts the abnormal text data frame group based on the adjacent text data frames, and verifies the predicted abnormal text data frame group.
[0068] For data with weak anomaly trends, a trained language model is used for data repair to confirm successful data replacement.
[0069] A data processing module, which is connected to the data diagnosis module and the data repair module, is used to determine the availability of data and mark usable data based on the data classification results of the data diagnosis module and the data repair results of the data repair module.
[0070] Among them, the abnormal text data frame group is a text data frame group composed of several abnormal text data frames.
[0071] It is understandable that the text data frame contains one or more keywords.
[0072] Specifically, when using a trained language model for data repair, there is no limitation on the specific model selected. It can be an open-source model that can identify parts of text data with low semantic relevance or missing parts, and then predict the content of that part based on contextual information. The predicted content replaces the original content to complete the repair. For example, the Transformer model can use a self-attention mechanism, can make predictions using contextual information at the same time, can better handle long-distance dependencies, and can also identify and predict parts with low semantic relevance to a certain extent. Of course, other forms can also be used, which will not be elaborated here.
[0073] Specifically, the data acquisition module determines the data missing rate and data consistency, including,
[0074] The data missing rate is the ratio of the amount of missing data to the total amount of data.
[0075] Keywords used to extract the data;
[0076] The similarity between the keywords and the title / topic of the data is used to determine the data identity.
[0077] In implementation, the data is in text form, and the total data volume is the total number of keywords in the data. Whether the data is missing can be determined based on the cosine similarity after the keyword vectorization. If the cosine similarity between a keyword and its adjacent keywords is lower than the preset missing cosine similarity threshold, it is determined that there is a missing keyword. The total number of missing keywords is counted to obtain the amount of missing data. The missing cosine similarity threshold is selected in the interval [0.25, 0.5].
[0078] Specifically, there are no restrictions on the methods for extracting keywords from the data. For example, word frequency statistics, TextRank algorithm, and TF-IDF method can be used. Of course, those skilled in the art can combine multiple methods, as long as the keywords in the data can be obtained. This is existing technology and will not be elaborated further.
[0079] Specifically, there is no limitation on the method for determining topic similarity. For example, in this embodiment, the cosine similarity method is used to vectorize each keyword in the data with the topic of the data, calculate the mean cosine similarity between the keywords and the topic words, and determine the mean cosine similarity as the topic similarity. Of course, the Jaccard similarity method, BERT and its variants can also be used. Those skilled in the art can choose the method according to the actual situation, as long as the topic similarity can be obtained. This is the prior art and will not be elaborated further.
[0080] Please see Figure 2 , Figure 2The present invention provides a logical block diagram of a data diagnostic module that determines data categories based on the data missing rate and data similarity. Specifically, the data diagnostic module determines data categories based on the data missing rate and data similarity, including:
[0081] If the data meets the diagnostic criteria, then the data category is determined to be an abnormal data category;
[0082] If the data does not meet the diagnostic criteria, the data category is determined to be non-abnormal data.
[0083] The diagnostic criteria are that the data missing rate is greater than a preset data missing rate threshold and / or the data identity is less than a preset data identity threshold.
[0084] Specifically, the preset data missing rate threshold is calculated in advance. The preset data missing rate threshold represents the lowest probability of identifying data anomalies when anomalies exist in the data. Therefore, a number of historical enterprise data are obtained in advance to determine a number of historical data missing rates, and the average of the historical data missing rates is determined as the preset data missing rate threshold.
[0085] Specifically, the preset data similarity threshold is calculated in advance. Several historical enterprise data are obtained in advance to determine the similarity of several historical data. The average value of the historical data similarity is determined as the preset data similarity threshold.
[0086] Please see Figure 3 , Figure 3 This is a logic block diagram of a data diagnostic module distinguishing data analysis methods according to an embodiment of the invention. Specifically, the data diagnostic module distinguishes data analysis methods, wherein...
[0087] If the category is an abnormal data category, then the abnormal data characterization value is calculated based on the data missing rate and data identity, and the abnormal trend of the data is divided to determine the data repair method.
[0088] If the category is a non-abnormal data category, then mark the data as available.
[0089] Specifically, data categories are determined based on data missing rate and data identity to classify enterprise data and process it in a targeted manner. In reality, enterprise data is mostly obtained from multiple channels, which can lead to data ambiguity, inconsistent quality, and damage to storage media resulting in lost frames or fragments. Existing technologies mostly process the data directly, but using the same method to process data of different categories will result in wasted computing power and reduced efficiency. Based on this, this invention considers first classifying enterprise data, marking non-abnormal data categories, and further classifying abnormal data categories to repair abnormal data, ensuring the integrity and accuracy of enterprise data, and improving the efficiency and accuracy of enterprise data analysis and processing.
[0090] Specifically, the data analysis module calculates data anomaly representation values based on the data missing rate and data identity, including:
[0091] The ratio of the data missing rate to the preset data missing rate threshold is used as the first influencing factor;
[0092] The ratio of the preset data similarity to the data similarity threshold is used to determine the second influence factor;
[0093] The weighted sum of the first impact factor and the second impact factor is used to determine the data anomaly characterization value.
[0094] Specifically, the sum of the weighting coefficients of the first impact factor and the second impact factor is 1, the weighting coefficient of the first impact factor is 0.53, and the weighting coefficient of the second impact factor is 0.47.
[0095] Please see Figure 4 , Figure 4 The data analysis module of this invention classifies data anomaly trends to determine the logical block diagram of a data repair method. Specifically, the data analysis module classifies data anomaly trends to determine the data repair method, wherein...
[0096] If the data anomaly characterization value is greater than the preset data anomaly characterization value threshold, the anomaly tendency trend is classified as a strong anomaly tendency trend, the data is divided into several abnormal text data frames, the abnormal text data frame group and the adjacent text data frames adjacent to the abnormal text data frame group are determined, the abnormal text data frame group is predicted based on the adjacent text data frames, and the predicted abnormal text data frame group is verified.
[0097] If the data anomaly representation value is less than or equal to the preset data anomaly representation value threshold, the anomaly tendency trend is classified as a weak anomaly tendency trend, and the trained language model is used for data repair to determine that the data replacement was successful.
[0098] Specifically, the preset data anomaly representation value threshold represents the minimum value at which data can be successfully replaced by the existing trained recurrent neural network. Therefore, the preset data anomaly representation value threshold is selected within the range [0.46, 0.65].
[0099] Specifically, the data repair module predicts the abnormal text data frame group based on the adjacent text data frames, including:
[0100] Used to determine adjacent text data frames to the abnormal text data frame group, including the first adjacent text data frame and the last adjacent text data frame;
[0101] Used to predict and replace the first text data frame of an abnormal text data frame group based on the first adjacent text data frame;
[0102] Used to predict and replace the last text data frame of an abnormal text data frame group based on the last adjacent text data frames;
[0103] Used to identify unpredicted abnormal text data frame groups, and to use the predicted first-end text data frame as the first-end adjacent text data frame.
[0104] The predicted tail text data frame is used as the tail adjacent text data frame for repeated prediction until all abnormal text data frames in the abnormal text data frame group have been replaced.
[0105] Understandably, in Example 1, the process of predicting abnormal text data frame groups is as follows:
[0106] In Example 1, there is an abnormal text data frame group, which contains abnormal text data frames A, B, C, D and E, a total of five abnormal text data frames.
[0107] At the same time, there are also first-end adjacent text data frames and last-end adjacent text data frames that are adjacent to the abnormal text data frame group.
[0108] By using the first adjacent text data frame, several text data frames that are related to the first adjacent text data frame can be predicted. It can be understood that the existence of a relationship means that the relationship between the text data frames is greater than a predetermined relationship threshold. The relationship threshold is selected in the range [0.5, 0.7], and the relationship is the cosine similarity between the text data frames.
[0109] The text data frame with the highest relevance is determined as the first text data frame, and A is replaced with this text data frame, denoted as A';
[0110] Similarly, the text data frame at the end can be obtained. Replace E with this text data frame and denote it as E'.
[0111] At this point, there are still three abnormal text data frames, B, C, and D, in the unpredicted abnormal text data frame group.
[0112] Use A' as the new first adjacent text data frame and E' as the new last adjacent text data frame.
[0113] By repeating the above steps, the data for B and D can be replaced;
[0114] Repeat the above operation to obtain the text data frame at the beginning and the text data frame at the end. Determine any text data frame and replace C with that text data frame.
[0115] At this point, the abnormal text data in the abnormal text data frame group has been completely replaced.
[0116] Specifically, the data repair module examines the predicted abnormal text data frame groups, including,
[0117] Used to determine the abnormal text data group corresponding to the abnormal text data frame group obtained by replacement;
[0118] This is used to determine several sample data in the enterprise database that maintain thematic consistency with the abnormal text data group;
[0119] Used to determine the similarity between the abnormal text data group and the plurality of sample data;
[0120] Used to detect abnormal text data frame groups based on the similarity.
[0121] Understandably, the title topic corresponding to the abnormal text data group is used as the topic word, and compared with the topic word corresponding to the sample data to determine the cosine similarity. If the cosine similarity is greater than the preset topic cosine similarity threshold, the preset topic cosine similarity threshold is the same as the preset data similarity threshold.
[0122] Specifically, the sample data can be keywords with pre-defined subject terms. This is understandable; for example, it can be extracted from data uploaded by the company in the past. First, the keywords in the data title are used as subject terms. The relevance between each keyword in the data and the subject terms is calculated and sorted in descending order. The top 30% of keywords are then selected, and the correspondence between each keyword and the subject terms is constructed and stored in the company database. The company database can be a database rebuilt by collecting the company's historical data, which will not be elaborated further.
[0123] It is understandable that the similarity between the abnormal text data group and the aforementioned sample data is the mean cosine similarity after keyword vectorization.
[0124] Specifically, the data repair module uses the similarity check to identify abnormal text data frames, wherein...
[0125] If the similarity is greater than the similarity threshold, the predicted abnormal text data frame group is determined to be successfully replaced data.
[0126] Specifically, the similarity threshold represents the probability that the abnormal text data group corresponding to the abnormal text data group exists in the enterprise's historical data. Therefore, the similarity threshold is set within the range of [0.72, 0.98].
[0127] Specifically, by calculating data anomaly representation values, the abnormal trend of the data is divided to determine the repair methods for enterprise data with different abnormal trend trends. When processing abnormal data, the same repair method is usually selected. However, in practice, to save computing power, different repair methods can be determined according to the abnormal trend of the data. Based on this, the present invention calculates data anomaly representation values to divide the abnormal trend of the data. For data with strong abnormal trend trends, the method of predicting abnormal text data frame groups is used to repair enterprise data. For data with weak abnormal trend trends, an existing trained recurrent neural network is used for repair, ensuring the integrity and accuracy of enterprise data and improving the efficiency and accuracy of enterprise data analysis and processing.
[0128] Please see Figure 5 , Figure 5 A logic block diagram illustrating how the data processing module in an embodiment of the invention determines data availability. Specifically, the data processing module determines data availability by including:
[0129] Used to determine the duplication status of the successfully replaced data and the available data;
[0130] If the successfully replaced data is duplicated with available data, the successfully replaced data will be removed.
[0131] If the successfully replaced data does not overlap with the available data, then the successfully replaced data is determined to be available data.
[0132] It is understood that the available data includes the available data marked by the data analysis module and the available data marked by the data processing module during the historical process. Once the available data is marked, it will be stored to facilitate the subsequent determination of the duplication status between the successfully replaced data and the available data. This will not be elaborated further.
[0133] Specifically, this invention performs usability verification on successfully replaced data. Preferably, most enterprise data comes from multiple channels and has redundancy. If successfully replaced data is directly marked for application, data duplication will cause space waste and reduce data analysis efficiency. Based on this, this invention verifies successfully replaced data and removes duplicate data to maintain data uniformity, ensure the integrity and accuracy of enterprise data, and improve the efficiency and accuracy of enterprise data analysis and processing.
[0134] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
[0135] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. An enterprise data monitoring system based on big data analytics, characterized in that, include: The data acquisition module is used to acquire text-based data uploaded by enterprises in order to determine the data missing rate and data consistency. A data diagnostic module, connected to the data acquisition module, is used to determine the data category based on the data missing rate and data identity, so as to distinguish the data analysis methods. A data analysis module, connected to the data diagnostic module, is used to respond to the data category determination result of the data diagnostic module, calculate the data anomaly characterization value based on the data missing rate and data identity, classify the anomaly trend of the data, determine the data repair method, or mark the available data. The data repair module, which is connected to the data analysis module, divides the data into several abnormal text data frames for data with a strong abnormal trend, determines the abnormal text data frame group and the adjacent text data frames adjacent to the abnormal text data frame group, predicts the abnormal text data frame group based on the adjacent text data frames, and verifies the predicted abnormal text data frame group. For data with weak anomaly trends, a trained language model is used for data repair to confirm successful data replacement. A data processing module, which is connected to the data diagnosis module and the data repair module, is used to determine the availability of data and mark usable data based on the data classification results of the data diagnosis module and the data repair results of the data repair module. Among them, the abnormal text data frame group is a text data frame group composed of several abnormal text data frames. The data acquisition module determines the data missing rate and data consistency, including: The data missing rate is the ratio of the amount of missing data to the total amount of data. Keywords used to extract the data; The similarity between the keywords and the title / topic of the data is used to determine the data identity degree. The data repair module predicts the abnormal text data frame group based on the adjacent text data frames, including: Used to determine adjacent text data frames to the abnormal text data frame group, including the first adjacent text data frame and the last adjacent text data frame; Used to predict and replace the first text data frame of an abnormal text data frame group based on the first adjacent text data frame; Used to predict and replace the last text data frame of an abnormal text data frame group based on the last adjacent text data frames; Used to identify unpredicted abnormal text data frame groups, and to use the predicted first-end text data frame as the first-end adjacent text data frame. The predicted tail text data frame is used as the tail adjacent text data frame for repeated prediction until all abnormal text data frames in the abnormal text data frame group have been replaced.
2. The enterprise data monitoring system based on big data analysis according to claim 1, characterized in that, The data diagnostic module determines the data category based on the data missing rate and data identity, including: If the data meets the diagnostic criteria, then the data category is determined to be an abnormal data category; If the data does not meet the diagnostic criteria, the data category is determined to be non-abnormal data. The diagnostic criteria are that the data missing rate is greater than a preset data missing rate threshold and / or the data identity is less than a preset data identity threshold.
3. The enterprise data monitoring system based on big data analysis according to claim 2, characterized in that, The data diagnostic module distinguishes data analysis methods, wherein, If the category is an abnormal data category, then the abnormal data characterization value is calculated based on the data missing rate and data identity, and the abnormal trend of the data is divided to determine the data repair method. If the category is a non-abnormal data category, then mark the data as available.
4. The enterprise data monitoring system based on big data analysis according to claim 1, characterized in that, The data analysis module calculates data anomaly representation values based on the data missing rate and data identity. include, The first influencing factor is used to determine the ratio of the data missing rate to the preset data missing rate; The ratio of the preset data identity degree to the data identity degree is used to determine the second influence factor; The weighted sum of the first impact factor and the second impact factor is used to determine the data anomaly characterization value.
5. The enterprise data monitoring system based on big data analysis according to claim 1, characterized in that, The data analysis module identifies abnormal trends in the data to determine data repair methods. If the data anomaly characterization value is greater than the preset data anomaly characterization value threshold, the anomaly tendency trend is classified as a strong anomaly tendency trend, the data is divided into several abnormal text data frames, the abnormal text data frame group and the adjacent text data frames adjacent to the abnormal text data frame group are determined, the abnormal text data frame group is predicted based on the adjacent text data frames, and the predicted abnormal text data frame group is verified. If the data anomaly representation value is less than or equal to the preset data anomaly representation value threshold, the anomaly tendency trend is classified as a weak anomaly tendency trend, and the trained language model is used for data repair to determine that the data replacement was successful.
6. The enterprise data monitoring system based on big data analysis according to claim 1, characterized in that, The data repair module examines the predicted abnormal text data frame groups, including, Used to determine the abnormal text data group corresponding to the abnormal text data frame group obtained by replacement; This is used to determine several sample data in the enterprise database that maintain thematic consistency with the abnormal text data group; Used to determine the similarity between the abnormal text data group and the plurality of sample data; Used to detect abnormal text data frame groups based on the similarity.
7. The enterprise data monitoring system based on big data analysis according to claim 6, characterized in that, The data repair module is based on the similarity check of the abnormal text data frame group, wherein, If the similarity is greater than the similarity threshold, the predicted abnormal text data frame group is determined to be successfully replaced data.
8. The enterprise data monitoring system based on big data analysis according to claim 1, characterized in that, The data processing module determines the availability of the data, including, Used to determine the duplication status of the successfully replaced data and the available data; If the successfully replaced data is duplicated with available data, the successfully replaced data will be removed. If the successfully replaced data does not overlap with the available data, then the successfully replaced data is determined to be available data.
Citation Information
Patent Citations
Data exchange monitoring method of enterprise internal systems
CN108173711A
Enterprise big data monitoring system and monitoring method
CN113961770A
Data verification classification processing system and method
CN112231310A
Abnormal monitoring method for business data
CN119377056A