Industrial Data Analysis Method and System Based on Dynamic Forms
By introducing multiple analysis modules into the industrial data analysis system, dynamic display, repetitive analysis, context vectorization and redundant identification of data are solved, and the data storage and analysis efficiency is optimized.
Patent Information
- Application Number
- CN202510077515.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The existing industrial data analysis system based on dynamic forms has redundant data problems in the data sharing process, resulting in waste of storage resources and reduced data analysis efficiency, and lack of effective redundant identification and data use analysis, resulting in data misuse and loss.
By introducing form construction modules, repeatability analysis modules, data use analysis modules, redundant identification modules and data processing modules in the industrial data analysis system, dynamic display, repetitive analysis, context vectorization and redundant identification of data can be realized, and non-redundant data can be screened out and compressed and stored.
Effectively identify and eliminate redundant data, optimize data storage, save storage resources, improve the efficiency and accuracy of data analysis, and ensure the simplicity and processing efficiency of data.
Smart Images

Figure CN119493813B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and specifically to an industrial data analysis method and system based on dynamic forms. Background Art
[0002] In the wide application of dynamic form technology, it has become a tool that can efficiently and flexibly collect and manage data, especially in the fields of data acquisition, information management, intelligent decision-making, etc.
[0003] Although the industrial data analysis system based on dynamic forms has played an important role in improving data display and acquisition efficiency, the current implementation still has several drawbacks and deficiencies. First, redundant data may be generated during the data sharing process. Especially when sharing data between multiple business systems, due to different data acquisition methods and requirements of each system, the same monitoring data may be repeatedly collected and stored. This redundant data not only occupies storage space but also may affect the efficiency and accuracy of subsequent data analysis. Second, the data items in the duplicate data groups may have different contexts, that is, the same data item may have different meanings in different systems or different scenarios. Currently, many systems lack effective redundant identification and data usage analysis during the data sharing process, often confusing duplicate data with redundant data, and it is difficult to accurately judge which data items are redundant and which data, even if it is duplicate data, is indispensable.
[0004] The occurrence of the above-mentioned drawbacks and deficiencies has led to several negative impacts. In the presence of redundant data, first, it will cause waste of storage resources. Especially in the scenario of large-scale industrial data acquisition, the storage cost of redundant data is quite huge. Second, due to data duplication and ineffective screening, data conflicts or deviations may be caused during subsequent data analysis, reducing the effectiveness of data and the accuracy of decision-making, and may even affect the timeliness of production decisions and equipment maintenance. In the case where the context differences are not accurately processed, the data between different business systems may be wrongly considered as duplicate or redundant, resulting in data loss or misuse, and ultimately affecting the overall performance of the system and the effect of business operations. For example, some data items may be key decision-making factors in certain scenarios, but due to the failure to correctly identify their context, they may ultimately be misjudged as redundant data because they belong to duplicate data and are finally eliminated, affecting subsequent production scheduling or quality control. Therefore, the correct identification, effective screening of redundant data, and clear analysis of the context are the keys to improving the performance and accuracy of industrial data analysis systems. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention provides an industrial data analysis method and system based on dynamic forms, which solves the problems in the above-mentioned background art.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: An industrial data analysis system based on dynamic forms, including a form construction module, a repeatability analysis module, a data usage analysis module, a redundancy identification module, and a data processing module;
[0007] The form construction module is used to obtain the monitoring data of each business system in the industrial park by using monitoring instruments, generate a dynamic form based on the monitoring data of each business system, and define conditions for the dynamic form to dynamically display the data item X under the corresponding conditions;
[0008] The repeatability analysis module is used to analyze the repeatability between business systems according to the data item X corresponding to the shared data when multiple groups of business systems share data, so as to calculate and obtain the similarity Xsd, and screen the duplicate data groups based on the similarity Xsd;
[0009] The data usage analysis module is used to identify the characteristics of each data item X in the duplicate data group in the context according to the duplicate data group, and perform data context vectorization to analyze the difference degree of the usage of each data item X in the duplicate data group in different contexts, so as to obtain the difference degree Cyd, and preliminarily screen out some non-redundant data groups from the duplicate data group based on the difference degree Cyd;
[0010] The redundancy identification module is used to obtain relevant numerical change data according to the monitoring interval on the basis of the data usage analysis module, so as to calculate and obtain the change value Bhz, combine the trained data identification model and the similarity Xsd, and after dimensionless processing, fit and output the determination index Pdzs;
[0011] The data processing module is used to identify and screen out non-redundant data groups from the duplicate data group again according to the determination index Pdzs, and perform compressed storage processing on the redundant data groups.
[0012] Preferably, the form construction module includes a data monitoring unit and a form definition unit;
[0013] The data monitoring unit is used to provide an interface for data exchange with the monitoring instruments in each business system to obtain the monitoring data of each business system in the industrial park, and perform data cleaning and convert the monitoring data of each business system into a unified format;
[0014] The form definition unit is used to generate a dynamic form according to the monitoring data of each business system obtained and processed by the data monitoring unit. The dynamic form includes the monitoring data of each business system, and conditions and rules are defined for the dynamic form to display relevant data item X required by the defined conditions and automatically hide the remaining data. The condition and rule definition includes status conditions, user selection conditions, time and date conditions, data item enable and disable conditions, data range verification, and data format verification. The data item X refers to the data values collected by each sensor under defined conditions.
[0015] Preferably, the repeatability analysis module includes a preparation unit, a repeat unit, and a judgment unit;
[0016] The preparation unit is used to determine the data item X involved in the sharing of multiple groups of business systems according to the sharing requirements of multiple groups of business systems when data sharing is performed among multiple groups of business systems. After summarization, shared data is generated. Based on the shared data, the data item X other than the current shared data is deleted from the monitoring data of each business system, and a shared data area is defined using condition and rule definitions;
[0017] The repeat unit is used to analyze the repeatability among business systems according to the data item X involved in the sharing of multiple groups of business systems displayed in the shared data area to calculate and obtain the similarity Xsd. The similarity Xsd is obtained through the following formula:
[0018] ;
[0019] In the formula, and are the vector representations of data items from two different business systems respectively, and are respectively the vectors and 's modulus.
[0020] Preferably, the judgment unit is used to set a preset threshold and compare the similarity Xsd with the preset threshold to screen out duplicate data groups. The specific screening content is as follows:
[0021] If the similarity Xsd exceeds the preset threshold, it means that and there is similarity between them. At this time, the corresponding data items are determined as duplicate data items and are not temporarily stored in the shared data area. After statistics, a duplicate data group is generated;
[0022] If the similarity Xsd does not exceed the preset threshold, it means that and There is no similarity yet, and at this time, the data items in each business system are stored in the shared data area.
[0023] Preferably, the data usage analysis module includes a vectorization unit, a difference unit, and a screening unit;
[0024] The vectorization unit is used to identify the specific business function roles of each data item X in the duplicate data group in the corresponding business system, the upstream and downstream relationships of the data items, and the changes and fluctuations of the data items in the time dimension. Among them, the specific business function roles include sensor data, control instructions, system status, and alarm signals, and are represented by classification labels; the upstream and downstream relationships of the data items refer to whether there are inputs or outputs of decisions or operations for the data item X; the changes and fluctuations of the data items in the time dimension refer to the time characteristics of the data item X, including the change rate and the fluctuation amplitude, and are represented as time series features; by vectorizing the specific business function roles of each data item X in the duplicate data group in the corresponding business system, the upstream and downstream relationships of the data items, and the changes and fluctuations of the data items in the time dimension, the feature vector Y of each data item in the duplicate data group is generated.
[0025] Preferably, the difference unit is used to analyze the difference degree of the uses of each data item X in the duplicate data group in different contexts according to the feature vector Y of each data item in the duplicate data group, so as to obtain the difference degree Cyd, and the difference degree Cyd is obtained by the following formula:
[0026] ;
[0027] In the formula, n represents the number of data items in the duplicate data group, i = 1, 2,..., n, and are the feature vectors of the i-th data item in the duplicate data groups from two different business systems respectively.
[0028] Preferably, the screening unit is used to preset a difference threshold, and compare and analyze the difference threshold with the difference degree Cyd, so as to preliminarily screen out some non-redundant data groups from the duplicate data group. The specific content is as follows:
[0029] If the difference degree Cyd exceeds the difference threshold, it indicates that the context difference of the corresponding data item is in an abnormal state. At this time, the corresponding data item is recorded as a non-redundant data item, and statistics are made to obtain a non-redundant data group;
[0030] If the difference degree Cyd does not exceed the difference threshold, it indicates that the context difference of the corresponding data item is in a normal state. At this time, the corresponding data item is recorded as a redundant data item.
[0031] Preferably, the redundancy identification module includes an obvious difference unit and a comprehensive analysis unit;
[0032] The obvious difference unit is used to re-update the duplicate data group according to the screening unit, and on the basis of re-updating the duplicate data group, obtain relevant numerical change data according to the monitoring interval, where the relevant numerical change data includes the numerical values within the data items obtained before and after the monitoring interval, so as to judge the numerical change situation of each data item X in the duplicate data group, and calculate to obtain the change value Bhz of each data item X between business systems;
[0033] The comprehensive analysis unit is used to construct a data recognition model by using deep learning technology, input the change value Bhz, the difference degree Cyd and the similarity Xsd into the data recognition model, and after dimensionless processing, fit and output the determination index Pdzs. The determination index Pdzs is obtained through the following formula:
[0034] ;
[0035] In the formula, , and are the weight values of the similarity Xsd, the change value Bhz and the difference degree Cyd respectively, , and The specific numerical values are set by the user according to the situation.
[0036] Preferably, the data processing module is used to preset an evaluation threshold M, and by comparing it with the determination index Pdzs, to identify and screen out the non-redundant data group from the duplicate data group again. If the determination index Pdzs exceeds the evaluation threshold M, at this time, the corresponding data item is recorded as a non-redundant data item, and counted to obtain the non-redundant data group, and combined with the re-updated duplicate data group in the obvious difference unit, determine the redundant data group, and perform compressed storage processing on the redundant data group to reduce the occupation of storage space.
[0037] The industrial data analysis method based on a dynamic form includes the following steps:
[0038] S1. Use monitoring instruments to obtain the monitoring data of each business system in the industrial park, generate a dynamic form based on the monitoring data of each business system, and define conditions for the dynamic form to dynamically display the data item X under the corresponding conditions;
[0039] S2. When multiple groups of business systems share data, analyze the repeatability between business systems according to the data item X corresponding to the shared data, calculate to obtain the similarity Xsd, and screen the duplicate data group based on the similarity Xsd;
[0040] S3. According to the duplicate data group, identify the characteristics of each data item X in the duplicate data group in the context, and perform data context vectorization to analyze the difference degree of the uses of each data item X in the duplicate data group in different contexts, so as to obtain the difference degree Cyd. Based on the difference degree Cyd, initially screen out some non-redundant data groups from the duplicate data group;
[0041] S4. On the basis of the data usage analysis module, according to the monitoring interval, obtain the relevant numerical change data to calculate and obtain the change value Bhz. Combine the trained data recognition model and the similarity Xsd, and after dimensionless processing, fit and output the determination index Pdzs;
[0042] S5. According to the determination index Pdzs, identify and screen out non-redundant data groups from the duplicate data group again, and perform compression storage processing on the redundant data groups.
[0043] The present invention provides an industrial data analysis method and system based on a dynamic form, having the following beneficial effects:
[0044] (1) Through the form construction module, the system integrates the monitoring data from different business systems into a dynamic form, enabling real-time collection and display of the monitoring data of each system. By defining conditions, the data items displayed in the form will be dynamically adjusted to meet different business needs, which enables the efficient sharing of data among multiple systems in the industrial park, reduces the existence of information silos, and improves the collaborative working efficiency among systems. When multiple systems share data, the system calculates the similarity Xsd of data items through the repetitive analysis module and filters out duplicate data groups based on this similarity, which enables the system to effectively identify and eliminate redundant data, avoid unnecessary repeated collection and storage, thereby optimizing data storage, saving storage resources, and reducing the burden of redundant data. Through the data usage analysis module, the system can identify the characteristics of each data item in the duplicate data group in different contexts. By vectorizing the data context and analyzing the usage difference degree Cyd in different contexts, the system can further determine which data items have different meanings in different systems or business scenarios, so as to screen out non-redundant data that should not be regarded as redundant. This kind of analysis ensures that the system makes judgments not only based on data similarity but also comprehensively considers the data context, improving the intelligence and accuracy of data processing. Through the redundancy identification module, the system not only identifies redundant data according to similarity analysis but also combines the change trend of monitoring data to make a numerical change judgment and fits the determination index Pdzs. The system can more accurately identify the redundant data group and compress and store it. This mechanism avoids excessive data storage, ensures the simplicity and processing efficiency of data. By comprehensively using the above modules, the system can accurately identify which data items are redundant and perform compression storage processing on redundant data. This not only improves the utilization efficiency of data storage, further avoids a large amount of useless data occupying storage space, but also can improve the efficiency of data processing and analysis through the streamlined data set and reduce the burden on the system. Through the flexibility of the dynamic form and the intelligence of the data analysis module, the system helps the managers and decision-makers in the industrial park better understand and analyze the real-time monitoring data. By reducing redundant data and enhancing the understanding of data context, the analysis results provided by the system are more accurate and operable, providing stronger support for business decisions. This industrial data analysis system based on dynamic forms can greatly improve data processing efficiency, reduce the burden brought by redundant data storage, and at the same time provide accurate and intelligent data analysis support for the industrial park through functions such as efficient data sharing, redundant data identification, context analysis, and compression storage, ultimately realizing the high efficiency and intelligence of data management and analysis.
[0045] (2) Through the preparation unit, the system can accurately determine the data items X to be shared according to the sharing requirements of multiple business systems, and screen out relevant data from the monitoring data of each business system. This function of data preprocessing and summarization ensures the accuracy and pertinence of the shared data, and avoids the mixing of irrelevant data. Through condition and rule definition, the system can clearly identify the shared data area, thereby optimizing the data sorting and display process and improving the efficiency of data sharing between different systems. In the shared data area, the repetition unit effectively analyzes the data repetition between business systems by calculating the similarity Xsd of data items. The similarity Xsd is calculated by a formula, which measures the similarity degree of data items between two different business systems. This calculation process uses vector representation and modulus value, and quantifies the similarity of data items in a numerical way, enabling the system to judge whether there are duplicate data items based on this value. This process not only improves the detection accuracy of duplicate data, but also avoids data redundancy caused by human factors or technical problems. By accurately screening duplicate data and performing intelligent processing, the system can maintain data consistency and interoperability between multiple business systems, ensuring the accuracy and real-time nature of shared data in different systems. This method provides a strong guarantee for cross-system collaboration within the industrial park and avoids data conflicts and inconsistencies caused by duplicate data. Through data optimization and standardization, it promotes the smooth exchange of data between various systems within the industrial park.
[0046] (3) Based on the feature vectors Y of data items within the duplicate data group, the difference unit calculates the context difference degree Cyd, and then analyzes the usage differences of data items in different contexts. The difference degree Cyd reflects the functional differences of data items in multiple systems by comparing the feature vectors of the same data item in different business systems. For example, a certain data item may represent the equipment status in one system and be used as a decision basis for production scheduling in another system. By calculating the difference degree, the application differences of these data items in different business systems can be accurately measured. This quantification process of the difference degree enables the system to intelligently distinguish which data items have significant functional differences, thus avoiding misjudging these data items as redundant data. Description of the Drawings
[0047] Figure 1 It is a block diagram of the industrial data analysis system based on dynamic forms according to the present invention;
[0048] Figure 2 It is a schematic flowchart of the industrial data analysis method based on dynamic forms according to the present invention. Detailed Embodiments
[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] Embodiment 1
[0051] Please refer to Figure 1 , the present invention provides an industrial data analysis system based on a dynamic form, including a form construction module, a repeatability analysis module, a data usage analysis module, a redundancy identification module, and a data processing module;
[0052] The form construction module is used to obtain the monitoring data of each business system in the industrial park by using monitoring instruments, generate a dynamic form based on the monitoring data of each business system, and define conditions for the dynamic form to dynamically display the data item X under the corresponding conditions;
[0053] The repeatability analysis module is used to analyze the repeatability between business systems according to the data item X corresponding to the shared data when multiple groups of business systems share data, so as to calculate and obtain the similarity Xsd, and screen the duplicate data groups based on the similarity Xsd;
[0054] The data usage analysis module is used to identify the characteristics of each data item X in the duplicate data group in the context according to the duplicate data group, and perform data context vectorization to analyze the difference degree of the uses of each data item X in the duplicate data group in different contexts, so as to obtain the difference degree Cyd, and preliminarily screen out some non-redundant data groups from the duplicate data group based on the difference degree Cyd;
[0055] The redundancy identification module is used to obtain relevant numerical change data according to the monitoring interval on the basis of the data usage analysis module, judge the numerical change situation of each data item X in the duplicate data group, so as to calculate and obtain the change value Bhz, and combine the trained data identification model and the similarity Xsd, and after dimensionless processing, fit and output the determination index Pdzs;
[0056] The data processing module is used to identify and screen out non-redundant data groups from the duplicate data group again according to the determination index Pdzs, and perform compression storage processing on the redundant data groups.
[0057] During the operation of this system, through the form construction module, the monitoring data of multiple business systems are aggregated and a dynamic form is generated, which can dynamically display corresponding data items according to actual needs. Through condition definition and adaptive display, data from different business systems can be efficiently displayed and processed in the same form, avoiding the problem of inflexible data display caused by static form design and improving the efficiency of data sharing and interaction. The repetitive analysis module can effectively screen out duplicate data groups by analyzing the similarity of data item X among business systems, avoiding the repeated collection and storage of the same monitoring data in multiple systems. The data usage analysis module can further judge which duplicate data is redundant and which is necessary by analyzing the data context differences, thus accurately screening out non-redundant data. This intelligent data redundancy identification can minimize the storage requirements of redundant data, saving storage resources and improving storage efficiency. The data usage analysis module can identify the usage differences of data in different systems and environments by performing vectorization analysis on the characteristics of data item X in different contexts in the duplicate data group, which provides more accurate and targeted input for subsequent data analysis and avoids data misunderstanding or incorrect analysis caused by unclear context. Through this refined analysis, the system can better support business decisions, improving the accuracy and efficiency of data analysis. Based on the data usage analysis, the redundancy identification module can judge the numerical change situation of duplicate data items through monitoring intervals, numerical changes, and model training, and output a determination index Pdzs in combination with similarity evaluation. Through this module, redundant data can not only be accurately identified, but also processed by compressed storage, reducing unnecessary data storage volume and optimizing data processing efficiency. By accurately identifying and screening out non-redundant data, the system can provide more efficient and accurate data support to help decision-makers in the industrial park conduct real-time monitoring, adjust production strategies, optimize equipment management, etc. The resulting decision-making data will be more reliable and can quickly respond to changes in the production process, thereby improving the operational efficiency of the entire industrial park. The design based on the dynamic form makes this system highly scalable and flexible. In summary, through the intelligent construction of the dynamic form and the accurate identification of redundant data, the present invention significantly improves the efficiency of data sharing, storage, and analysis, reduces the storage requirements of redundant data, and provides strong technical support for data management and business decision-making in the industrial park.
[0058] Embodiment 2
[0059] Please refer to Figure 1 , specifically: The form construction module includes a data monitoring unit and a form definition unit;
[0060] The data monitoring unit is used to provide an interface for data exchange with monitoring instruments in each business system, so as to obtain the monitoring data of each business system in the industrial park, and clean and convert the monitoring data of each business system into a unified format (such as JSON, XML, etc.). Among them, the monitoring data of each business system includes but is not limited to production status, equipment operation status, environmental parameters, etc.;
[0061] The form definition unit is used to generate a dynamic form according to the monitoring data of each business system obtained and processed by the data monitoring unit. Among them, the dynamic form includes the monitoring data of each business system, and defines conditions and rules for the dynamic form, so as to display the required relevant data item X through the defined conditions and automatically hide the rest of the data. Among them, the condition and rule definition includes status conditions, user selection conditions, time and date conditions, data item enable and disable conditions, data range verification, and data format verification. Among them, data item X refers to the data value collected by each sensor under the defined conditions.
[0062] Among them, the status condition determines the display of different fields according to the current status of the device or system (such as device operation status, production process status, etc.). The status condition usually involves the start / stop and maintenance status of the device. Example: If the production status is "paused", then hide the data items related to "production progress", that is, hide the relevant data items other than "paused";
[0063] The user selection condition dynamically displays relevant data items according to the user's selection in the form. When the user selects an option, the form may display specific data items or content according to the selection. Example: If the user selects the "equipment maintenance" option, then display the "maintenance reason" field.
[0064] Time and date conditions: The display of some data items is related to a specific time period or date. For example, some data is only displayed during working hours or during a specific production stage.
[0065] Data item enable and disable conditions: In addition to determining whether to display data items, the enable and disable conditions further control whether the data items can be edited or filled in by the user. This helps to control the user's input permissions according to the context and business requirements.
[0066] Data range verification: For example, a data item of a numerical type (such as temperature, pressure, etc.) must be within a specific range, and an error message will be given if it exceeds the range.
[0067] Data format verification: Verify whether the input format of the data item meets the requirements. For example, the input date must conform to a specific format, or the input temperature value must be within a reasonable range.
[0068] In this embodiment, the data monitoring unit exchanges data with the monitoring instruments in each business system through the provided interface, ensuring that data from different business systems can be smoothly collected into the system. By cleaning and converting the data format, the system can unify and standardize the monitoring data from different sources such as production, equipment operation status, and environmental parameters, thus eliminating the problems of data inconsistency and format differences, laying a solid foundation for subsequent data analysis and display. This data integration process can significantly improve the data collection efficiency and further avoid the errors and complexities brought by manual processing and multiple conversions. Through the form definition unit, the system dynamically generates forms based on the cleaned and formatted monitoring data. These dynamic forms can display relevant data item X according to specific requirements and automatically adjust the displayed content according to the set conditions, ensuring that only the data items related to the current task are displayed. This not only improves the flexibility of data display but also simplifies the user operation process and avoids unnecessary data interference. By defining conditions and rules, the form can intelligently display, hide, or enable / disable different data items, improving the user-friendliness of the system. The form definition unit supports the definition of multiple conditions and rules, such as data item display conditions, status conditions, user selection conditions, time and date conditions, etc., enhancing the configurability and adaptability of the system as much as possible.
[0069] Users can flexibly set the display method and enable / disable status of data items according to different business requirements to ensure the efficient operation of the form in various scenarios. The setting of these conditions and rules can not only prevent users from making mistakes in the data processing process but also ensure the accuracy and compliance of the data. For example, through the time and date conditions, the form only displays the data within the specified time range to ensure the timeliness and accuracy of the information. The system can monitor the data input by users in real time through functions such as integrated data range verification, logical relationship verification, and data format verification to ensure that it conforms to the set rules. The range verification of data items can prevent data from exceeding the preset range, the logical relationship verification can ensure the logical consistency between different data items, and the data format verification can ensure the correctness of the data. These verification functions effectively prevent data entry errors and improve the accuracy and reliability of the system data.
[0070] Embodiment 3
[0071] Please refer to Figure 1 , specifically: The repeatability analysis module includes a preparation unit, a repeat unit, and a judgment unit;
[0072] The preparation unit is used to determine the data item X involved in the sharing of multiple business systems according to the sharing requirements of multiple business systems when data sharing is carried out among multiple business systems. After summarization, shared data is generated. Based on the shared data, data item X other than the current shared data is deleted from the monitoring data of each business system, and a shared data area is defined using conditions and rule definitions.
[0073] The repetition unit is used to analyze the repeatability between business systems based on the data item X involved in the sharing of multiple business systems shown in the shared data area, so as to calculate and obtain the similarity Xsd. The similarity Xsd is obtained through the following formula:
[0074] ;
[0075] In the formula, and are the vector representations of data items from two different business systems respectively, and are respectively the and modulus of the vectors.
[0076] The judgment unit is used to set a preset threshold, and by comparing the similarity Xsd with the preset threshold, duplicate data groups are screened. The specific screening content is as follows:
[0077] If the similarity Xsd exceeds the preset threshold, it means that and there is similarity between them, that is, when two different business systems share data, there is a situation where the corresponding data items in them are similar; at this time, the corresponding data items are determined as duplicate data items and are not temporarily stored in the shared data area. After statistics, a duplicate data group is generated;
[0078] If the similarity Xsd does not exceed the preset threshold, it means that and there is no similarity between them for the time being, that is, when two different business systems share data, there is no situation where the corresponding data items in them are similar; at this time, the data items in each business system are stored in the shared data area.
[0079] In this embodiment, when the preparation unit performs data sharing among multiple groups of business systems, it can intelligently determine and summarize the data item X to be shared this time according to the sharing requirements of each system. By screening and processing the shared data, the system eliminates redundant or irrelevant data items from the monitoring data of the business system and only retains the data item X involved in this sharing. Through the management of the shared data area defined based on conditions and rules, the system can clearly define which data items are the core content of the sharing, avoiding the display and storage of redundant and irrelevant data. This not only ensures the accuracy of data sharing but also improves the efficiency of data exchange and avoids the waste of redundant data. Through the repetition unit, the system can analyze the data item X in the shared data area and accurately judge the data repeatability among business systems according to the similarity calculation (through vector representation and modulus operation). The similarity Xsd calculated by the system can reflect the similarity degree of the shared data items between two different business systems. If the similarity Xsd is relatively high, it indicates that the shared data items of the two systems have a high degree of similarity. The system can effectively identify and avoid storing duplicate data. This analysis process can significantly reduce data redundancy, optimize the data storage space, and improve the efficiency of subsequent data analysis and processing. The judgment unit can flexibly and accurately screen out duplicate data groups by comparing the similarity Xsd with a preset threshold. When the similarity exceeds the set threshold, the system determines that the data item is duplicate data and does not store it in the shared data area temporarily, thus avoiding the repeated storage of duplicate data. Through this precise screening mechanism, the system not only effectively reduces the generation of redundant data but also ensures the rationality and necessity of data storage. Through the cooperation of these modules, the system can reduce unnecessary data storage and sharing and improve the data sharing efficiency among multiple business systems in the industrial park. At the same time, the efficient identification and screening of duplicate data avoid the waste of storage resources, especially in the environment of large-scale data collection and processing, which can significantly reduce the storage cost. Finally, this mechanism will bring about data storage compression, resource savings, and system performance improvement, making data analysis and decision-making more efficient and accurate.
[0080] Embodiment 4
[0081] Please refer to Figure 1 , specifically: The data usage analysis module includes a vectorization unit, a difference unit, and a screening unit;
[0082] The vectorization unit is used to identify the specific business function roles of each data item X within the duplicate data group in the corresponding business system, the upstream and downstream relationships of the data items, and the changes and fluctuations of the data items in the time dimension. Among them, the specific business function roles include sensor data, control instructions, system status, and alarm signals, and are represented by classification tags; the upstream and downstream relationships of the data items refer to whether there are inputs or outputs of decisions or operations for the data item X, such as whether it is an input of a control signal or a basis for production scheduling decisions, which can be represented by a boolean value to reflect whether there is an impact; the changes and fluctuations of the data items in the time dimension refer to the time characteristics of the data item X, including the change rate and the fluctuation range, and are expressed as time series features; by vectorizing the specific business function roles of each data item X within the duplicate data group in the corresponding business system, the upstream and downstream relationships of the data items, and the changes and fluctuations of the data items in the time dimension, feature vectors Y of each data item within the duplicate data group are generated. Data context vectorization is the process of mapping these features into vectors.
[0083] The difference unit is used to analyze the difference degree of the uses of each data item X within the duplicate data group in different contexts according to the feature vectors Y of each data item within the duplicate data group, so as to obtain the difference degree Cyd. The difference degree Cyd is obtained through the following formula:
[0084] ;
[0085] In the formula, n represents the number of data items within the duplicate data group, i = 1, 2,..., n, and are the feature vectors of the i-th data item within the duplicate data groups from two different business systems respectively.
[0086] The screening unit is used to preset a difference threshold and compare and analyze the difference threshold with the difference degree Cyd, so as to preliminarily screen out some non-redundant data groups from the duplicate data group. The specific content is as follows:
[0087] If the difference degree Cyd exceeds the difference threshold, it indicates that the context difference of the corresponding data item is in an abnormal state. At this time, the corresponding data item is recorded as a non-redundant data item, and statistics are made to obtain a non-redundant data group, which means that the uses or functions of these data items in different systems are significantly different and cannot be regarded as redundant.
[0088] If the difference degree Cyd does not exceed the difference threshold, it indicates that the context difference of the corresponding data item is in a normal state. At this time, the corresponding data item is recorded as a redundant data item, which means that the uses or functions of these data items in different systems are not significantly different and can be regarded as redundant.
[0089] For the context analysis of data, how to identify the characteristics of data items in different contexts and how to measure their degree of difference are the keys to judging redundant data and non-redundant data. Context analysis pays more attention to the functional roles, uses, associated information, and usage scenarios of data in different systems. Therefore, the core of context analysis lies in understanding the business context of each data item, quantifying its degree of difference in some way, and finally judging whether the data items have the same use in different systems, and then judging whether they are redundant.
[0090] In this embodiment, the vectorization unit analyzes the specific business function roles, upstream and downstream relationships, and changes and fluctuations in the time dimension of each data item within the duplicate data group, and converts this complex context information into feature vectors. Through this process of data context vectorization, the system can not only identify the actual roles of each data item in the business (such as sensor data, control instructions, system status, etc.), but also reveal the change patterns of the data items within different time intervals. This vectorization processing enables the precise quantification of the business context and time characteristics of the data items, providing a solid foundation for subsequent difference analysis. Through the difference unit, the system analyzes the feature vectors Y of each data item within the duplicate data group, calculates and obtains the difference degree Cyd. The difference degree Cyd quantifies the difference in the usage of the same data item in two business systems within the context, helping the system identify which data items have significant functional differences in different contexts. This analysis process not only reveals the potential differences between the duplicate data items, but also clearly shows which data items perform different roles in different business systems and thus cannot be regarded as redundant. The calculation and analysis of the difference degree Cyd can ensure whether there are actual usage differences between the data items in different systems, thereby helping to accurately determine which data items belong to the non-redundant data group. The screening unit can screen out the non-redundant data group according to the difference degree Cyd by comparing it with a preset difference threshold. When the difference degree Cyd exceeds the set threshold, the system recognizes that the function or usage of this data item has a significant difference in different systems, indicating that this data item is non-redundant and must be retained. This screening mechanism effectively reduces the storage of redundant data, enabling only the data items with truly unique values and functions to be stored in the shared area, thereby optimizing data storage and improving data processing efficiency. By vectorizing the context information of the data items and combining difference analysis, the system can deeply understand the roles and changes of each data item in different systems. Compared with traditional data analysis methods, this analysis is not limited to the direct comparison of numerical values and attributes, but also considers the business functions, upstream and downstream relationships, and time characteristics of the data items, thus having higher intelligence and accuracy. This analysis ability enables the system to process complex multi-source data and provide more scientific and intelligent decision-making bases for data cleaning, storage, and sharing. By accurately screening non-redundant data, the system can effectively avoid unnecessary storage and transmission of redundant data, reduce system resource consumption, and improve the response speed and accuracy of data analysis. At the same time, this also reduces the interference of redundant information in subsequent data processing, making the data more concise and clear in practical applications, thereby optimizing the overall efficiency of industrial data analysis.
[0091] Embodiment 5
[0092] Please refer to Figure 1 , specifically: the redundant identification module includes an obvious difference unit and a comprehensive analysis unit;
[0093] The obvious difference unit is used to re-update the duplicate data group according to the screening unit, and on the basis of re-updating the duplicate data group, obtain relevant numerical change data according to the monitoring interval, where the relevant numerical change data includes the numerical values in the data items obtained before and after the monitoring interval, so as to judge the numerical change situation of each data item X in the duplicate data group, and calculate to obtain the change value Bhz of each data item X between business systems;
[0094] The comprehensive analysis unit is used to construct a data recognition model by using deep learning technology, input the change value Bhz, the difference degree Cyd and the similarity degree Xsd into the data recognition model, and after dimensionless processing, fit and output the determination index Pdzs. The determination index Pdzs is obtained through the following formula:
[0095] ;
[0096] In the formula, , and are the weight values of the similarity degree Xsd, the change value Bhz and the difference degree Cyd respectively. Among them, 0 < < 1, 0 < < 1, 0 < < 1, , and The specific numerical values are set by the user according to the situation.
[0097] The data processing module is used to preset an evaluation threshold M, and by comparing it with the determination index Pdzs, identify and screen out the non-redundant data group from the duplicate data group again. If the determination index Pdzs exceeds the evaluation threshold M, at this time, the corresponding data item is recorded as a non-redundant data item, and statistics are carried out to obtain the non-redundant data group. Combining with the re-updated duplicate data group in the obvious difference unit, the redundant data group is determined, and the redundant data group is compressed and stored to reduce the occupation of storage space.
[0098] Among them, the setting method of the evaluation threshold M is as follows: According to the data shared between other business systems, similarly, obtain the determination index Pdzs, and combine statistical algorithms to respectively obtain the average value and standard deviation of the determination index Pdzs. Based on the average value and standard deviation of the determination index Pdzs, set the evaluation threshold M: Evaluation threshold M = average value of the determination index Pdzs + k * standard deviation of the determination index Pdzs, where k is a constant, usually taking values from 1 to 3, corresponding to different confidence levels, and the specific numerical values are set by the user (according to the actual situation).
[0099] In this embodiment, the obvious difference unit, based on the re-updated duplicate data group, calculates the change value Bhz by obtaining the relevant numerical change data before and after the monitoring interval, and can accurately capture the changes of data items in the time dimension. This process can not only identify the fluctuations of data items at different time points, but also compare and analyze the change trends of the same data items in different business systems. This method can effectively eliminate the pseudo-redundant data caused by time differences or external factors, providing a strong guarantee for the accurate storage and subsequent processing of data. At the same time, the change value Bhz can be used as a basis for measuring the importance of data, further optimizing the data screening process and reducing unnecessary data storage. The comprehensive analysis unit takes the change value Bhz, the difference degree Cyd and the similarity Xsd as inputs, performs dimensionless processing through the constructed deep learning data recognition model, and calculates the determination index Pdzs. This process can comprehensively consider the change situations, functional differences and similarities of data items from different sources, and conduct intelligent evaluation and screening. By adjusting the weight value, the user can customize the sensitivity of the model according to specific business requirements, thereby optimizing the effect of data analysis. The output result of the model, that is, the determination index Pdzs, can provide an accurate basis for further identifying non-redundant data groups. On this basis, the redundant data group can be compressed and stored to reduce the occupancy rate of the storage space. The data processing module further screens out the non-redundant data group by setting the evaluation threshold M and combining the determination index Pdzs, and performs compression storage processing on the redundant data. This mechanism significantly reduces the occupancy of redundant data on storage resources, optimizes the data storage and management process, thereby improving the storage efficiency and data processing capacity of the entire system. At the same time, this redundant data identification and compression storage processing can effectively reduce the storage cost and calculation overhead, providing a more efficient storage environment for subsequent data analysis. Through deep learning technology and dimensionless processing, the system can quickly adapt to the changing data environment and requirements, providing efficient redundant data identification and compression storage services. This flexibility and efficiency enable the system to be widely applied to different industrial data analysis scenarios, improving the application scope and scalability of the system.
[0100] Embodiment 6
[0101] Please refer to Figure 2 , specifically: The industrial data analysis method based on a dynamic form includes the following steps:
[0102] S1. Use monitoring instruments to obtain the monitoring data of each business system in the industrial park. Based on the monitoring data of each business system, generate a dynamic form, and define conditions for the dynamic form to dynamically display the data item X under the corresponding conditions;
[0103] S2. When data is shared among multiple groups of business systems, analyze the repeatability among business systems based on the data item X corresponding to the shared data to calculate and obtain the similarity Xsd. Based on the similarity Xsd, screen the duplicate data groups;
[0104] S3. According to the duplicate data groups, identify the characteristics of each data item X in the duplicate data groups in the context and perform data context vectorization to analyze the difference degree of the uses of each data item X in different contexts in the duplicate data groups to obtain the difference degree Cyd. Based on the difference degree Cyd, initially screen out some non-redundant data groups from the duplicate data groups;
[0105] S4. On the basis of the data use analysis module, obtain the relevant numerical change data according to the monitoring interval to calculate and obtain the change value Bhz. Combine the trained data recognition model and the similarity Xsd, and after dimensionless processing, fit and output the determination index Pdzs;
[0106] S5. According to the determination index Pdzs, identify and screen out non-redundant data groups from the duplicate data groups again, and perform compression storage processing on the redundant data groups.
[0107] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. The industrial data analysis system based on dynamic forms is characterized by: It includes form construction module, repetitive analysis module, data usage analysis module, redundancy identification module and data processing module; The form construction module is used to obtain monitoring data of each business system in the industrial park using monitoring instruments, generate dynamic forms based on the monitoring data of each business system, and define conditions for the dynamic forms to dynamically display data items X under corresponding conditions; The repetitive analysis module is used to analyze the repetitiveness between business systems when multiple business systems share data, according to the data item X corresponding to the shared data, to calculate and obtain the similarity Xsd, and to filter the repeated data group based on the similarity Xsd; The repeatability analysis module includes a preparation unit, a repeating unit and a judgment unit; The preparation unit is used to determine the data items X involved in the sharing of the multiple business systems according to the sharing requirements of the multiple business systems when multiple business systems are sharing data, generate shared data after aggregation, delete the data items X other than the shared data from the monitoring data of each business system based on the shared data, and define the shared data area by using conditions and rules; The repeating unit is used to analyze the repetitiveness between the business systems according to the data item X involved in the sharing of multiple business systems displayed in the shared data area, so as to calculate and obtain the similarity Xsd. The similarity Xsd is obtained by the following formula: ; In the formula, and Vector representation of data items from two different business systems, and They are vectors and Model; The data usage analysis module is used to identify the features of each data item X in the repeated data group in the context according to the repeated data group, and perform data context vectorization to analyze the difference degree of usage of each data item X in the repeated data group in different contexts to obtain the difference degree Cyd, and preliminarily screen out some non-redundant data groups from the repeated data group based on the difference degree Cyd; The data usage analysis module includes a vectorization unit, a difference unit and a screening unit; The vectorization unit is used to identify, according to the repeated data group, the specific business function role of each data item X in the repeated data group in the corresponding business system, the upstream and downstream relationship of the data item, and the change and fluctuation of the data item in the time dimension, wherein the specific business function role includes sensor data, control instructions, system status and alarm signals, and is represented by classification labels; the upstream and downstream relationship of the data item refers to whether the data item X has input or output of a decision or operation; the change and fluctuation of the data item in the time dimension refers to the time characteristics of the data item X, including the change rate and the fluctuation amplitude, which are represented as time series features; by performing data context vectorization on the specific business function role of each data item X in the repeated data group in the corresponding business system, the upstream and downstream relationship of the data item, and the change and fluctuation of the data item in the time dimension, a feature vector Y of each data item in the repeated data group is generated; The redundant identification module is used to obtain relevant numerical change data based on the monitoring interval on the basis of the data usage analysis module to calculate the change value Bhz, and combine the trained data recognition model and similarity Xsd, and after dimensionless processing, fit the output judgment index Pdzs; The data processing module is used to identify and filter out non-redundant data groups from repeated data groups again according to the determination index Pdzs, and compress and store the redundant data groups.
2. The industrial data analysis system based on dynamic forms according to claim 1 is characterized in that: The form construction module includes a data monitoring unit and a form definition unit; The data monitoring unit is used to provide an interface to exchange data with monitoring instruments in each business system to obtain monitoring data of each business system in the industrial park, and to clean the monitoring data of each business system and convert it into a unified format; The form definition unit is used to generate a dynamic form based on the monitoring data of each business system acquired and processed by the data monitoring unit, wherein the dynamic form includes the monitoring data of each business system, and the conditions and rules of the dynamic form are defined to display the required related data items X through the defined conditions, and automatically hide the remaining data, wherein the condition and rule definitions include state conditions, user selection conditions, time and date conditions, data item enable and disable conditions, data range verification and data format verification, wherein the data item X refers to the data value collected by each sensor under the defined conditions.
3. The industrial data analysis system based on dynamic forms according to claim 2 is characterized in that: The judgment unit is used for a preset threshold, and compares the similarity Xsd with the preset threshold to screen the duplicate data group. The specific screening content is as follows: If the similarity Xsd exceeds a preset threshold, it means and If there is similarity between them, the corresponding data item is determined as a duplicate data item and is not stored in the shared data area temporarily. After statistics, a duplicate data group is generated; If the similarity Xsd does not exceed the preset threshold, it means and There is no similarity between them yet, so the data items in each business system are stored in the shared data area.
4. The industrial data analysis system based on dynamic forms according to claim 3 is characterized in that: The difference unit is used to analyze the difference degree of use of each data item X in the repeated data group in different contexts according to the characteristic vector Y of each data item in the repeated data group, so as to obtain the difference degree Cyd, and the difference degree Cyd is obtained by the following formula: ; Where n represents the number of data items in the repeated data group, i=1, 2, ..., n, and The feature vector of the i-th data item in the repeated data group comes from two different business systems.
5. The industrial data analysis system based on dynamic forms according to claim 4 is characterized in that: The screening unit is used to pre-set a difference threshold, and compare and analyze the difference threshold with the difference degree Cyd, so as to preliminarily screen out some non-redundant data groups from the repeated data groups. The specific contents are as follows: If the difference degree Cyd exceeds the difference threshold, it means that the context difference of the corresponding data item is in an abnormal state. At this time, the corresponding data item is recorded as a non-redundant data item and counted to obtain a non-redundant data group; If the difference degree Cyd does not exceed the difference threshold, it means that the context difference of the corresponding data item is in a normal state, and the corresponding data item is recorded as a redundant data item.
6. The industrial data analysis system based on dynamic forms according to claim 5 is characterized in that: The redundancy identification module includes a significant difference unit and a comprehensive analysis unit; The obvious difference unit is used to re-update the repeated data group according to the screening unit, and on the basis of the re-updated repeated data group, obtain relevant value change data according to the monitoring interval, wherein the relevant value change data includes the values in the data items obtained before and after the monitoring interval, so as to determine the value change of each data item X in the repeated data group, and obtain the change value Bhz of each data item X between the business systems through calculation; The comprehensive analysis unit is used to construct a data recognition model using deep learning technology, input the change value Bhz, the difference Cyd and the similarity Xsd into the data recognition model, and after dimensionless processing, fit and output the judgment index Pdzs, which is obtained by the following formula: ; In the formula, , and are the weight values of similarity Xsd, change value Bhz and difference Cyd, , and The specific value is set by the user according to the situation.
7. The industrial data analysis system based on dynamic forms according to claim 6 is characterized in that: The data processing module is used to pre-set an evaluation threshold M, and compare it with the determination index Pdzs to identify and screen out non-redundant data groups from the repeated data groups again. If the determination index Pdzs exceeds the evaluation threshold M, the corresponding data items are recorded as non-redundant data items, and statistics are performed to obtain non-redundant data groups. In combination with the repeated data groups updated again in the obvious difference units, redundant data groups are determined, and the redundant data groups are compressed and stored to reduce the storage space occupied.
8. An industrial data analysis method based on a dynamic form, used to implement the industrial data analysis system based on a dynamic form as described in any one of claims 1 to 7, characterized in that: The following steps are involved: S1. Use monitoring instruments to obtain monitoring data of various business systems in the industrial park, generate dynamic forms based on the monitoring data of various business systems, and define conditions for the dynamic forms to dynamically display data items X under corresponding conditions; S2. When multiple business systems share data, the duplication between business systems is analyzed according to the data item X corresponding to the shared data to calculate the similarity Xsd, and the duplicate data group is screened based on the similarity Xsd; S3, according to the repeated data group, identifying the characteristics of each data item X in the repeated data group in the context, and performing data context vectorization to analyze the difference in the use of each data item X in the repeated data group in different contexts to obtain the difference degree Cyd, and based on the difference degree Cyd, preliminarily screening out some non-redundant data groups from the repeated data group; S4. Based on the data usage analysis module, relevant numerical change data is obtained according to the monitoring interval to calculate the change value Bhz, and the output judgment index Pdzs is fitted after dimensionless processing in combination with the trained data recognition model and similarity Xsd; S5. According to the determination index Pdzs, non-redundant data groups are identified and screened out from the repeated data groups again, and the redundant data groups are compressed and stored.
Citation Information
Patent Citations
Method and device for carrying out repeated scheme detection on customer service session
CN115470771A
System and method for screening redundant data in power system
CN118227613A