An intelligence mining and analysis method in industrial scenarios
By determining the data state based on the data source correlation and mining difficulty coefficient and selecting an appropriate strategy analysis method, the problem of single data selection method in the existing technology is solved, and data mining efficiency and fault identification are improved.
Patent Information
- Application Number
- CN202510165241.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-14
AI Technical Summary
In the prior art, the data selection method is single, and adaptive selection cannot be made based on the actual data characteristics, resulting in poor abnormal data mining efficiency.
By acquiring each data source and determining the data state based on the data source correlation and mining difficulty coefficient, selecting appropriate strategy analysis methods, including extraction analysis for sub-data or selection analysis for related source groups.
It improves the pertinence and effectiveness of data analysis, avoids redundant processing, enhances data mining efficiency, and improves the efficiency and accuracy of fault identification.
Smart Images

Figure CN119621808B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligence mining and analysis, and in particular to an intelligence mining and analysis method in an industrial scenario. Background Art
[0002] With the rapid development of industry, industrial data has increased dramatically and data is updated frequently. Traditional data processing and analysis methods have great difficulty in extracting fault information and are unable to efficiently mine potential fault information hidden in the data, resulting in poor efficiency in fault identification. Therefore, how to improve the efficiency of data mining in industrial production processes is a technical problem that needs to be urgently solved by technical personnel in this field.
[0003] Chinese patent publication number CN118194204A discloses a method, system, equipment and medium for industrial data feature selection and outlier detection, including: it can identify some process variables that have a greater impact on industrial production results from a large number of industrial production process variables, which can effectively reduce the number of model input features and significantly improve work efficiency; at the same time, it can use filtering methods to detect outliers in the original data, which can replace manual work and quickly pick out abnormal data from the original data set, greatly improving the efficiency of data set production; in addition, it also toolizes the two parts of feature selection and outlier detection. It can be seen that the above technical solution has the following problems: data selection is only based on outliers, the data selection method is single, and it is impossible to make adaptive selections according to the actual data characteristics, resulting in poor efficiency in abnormal data mining. Summary of the invention
[0004] To this end, the present invention provides an intelligence mining and analysis method in an industrial scenario, so as to overcome the problem that the data selection method in the prior art is single and cannot be adaptively selected according to the characteristics of actual data, resulting in poor efficiency of abnormal data mining.
[0005] To achieve the above object, the present invention provides an intelligence mining and analysis method in an industrial scenario, comprising:
[0006] Obtain each data source, determine the data status according to the data source correlation and mining difficulty coefficient, and determine the strategic analysis method according to the data status. The strategic analysis method is to extract and analyze the sub-data or select and analyze the associated source group;
[0007] When extracting and analyzing the sub-data, a processing method is determined according to the distribution coefficient of the dangerous sub-data and the dangerous turbulence value to obtain a sub-data set. The processing method is to determine a first extraction method according to the radiation coefficient of the dangerous end and the correlation coefficient of the sub-data within the end, or to determine a second extraction method according to the category of the dangerous gathering area;
[0008] The first extraction method is to determine the number of selected dangerous sub-data according to the interaction coefficient of the associated terminal device and the environmental variation value, or to select the dangerous sub-data according to the influencing threshold within the terminal; the second extraction method is to select the dangerous sub-data according to the influencing factor, or to determine the number of selected dangerous sub-data in the dangerous concentration area according to the comprehensive evaluation value;
[0009] When selecting and analyzing the associated source group, the clustering method is determined according to the feature conversion coefficient and the number of data sources, and the data source in the associated source group is selected according to the data flow anomaly index, and the sub-data corresponding to each selected data source is used as the sub-data set.
[0010] Furthermore, when the data status is that the data source association degree is less than the preset data source association degree or the mining difficulty coefficient is greater than or equal to the preset mining difficulty coefficient, the strategic analysis method is to perform extraction analysis on the data source.
[0011] Furthermore, when the data status is that the data source association degree is greater than or equal to the preset data source association degree and the mining difficulty coefficient is less than the preset mining difficulty coefficient, the strategy analysis method is to perform selection analysis on the associated source group.
[0012] Furthermore, each sub-data category is provided with a corresponding dangerous sub-data determination method, wherein:
[0013] For a type of sub-data, the dangerous sub-data is determined by using a disorder threshold and a mutation coefficient;
[0014] For the second type of sub-data, the dangerous sub-data is determined based on the proportion of influencing keywords.
[0015] Furthermore, the processing method is determined according to the distribution coefficient of the dangerous sub-data and the dangerous turbulence value, wherein:
[0016] If the risk sub-data distribution coefficient is greater than or equal to the preset risk sub-data distribution coefficient or the risk turbulence value is greater than or equal to the preset risk turbulence value, the processing method is to determine the first extraction method according to the risk end radiation coefficient and the sub-data correlation coefficient within the end;
[0017] If the risk sub-data distribution coefficient is less than the preset risk sub-data distribution coefficient and the risk turbulence value is less than the preset risk turbulence value, the processing method is to determine the second extraction method according to the risk concentration area category.
[0018] Further, the first extraction method is determined according to the radiation coefficient of the dangerous end and the correlation coefficient of the sub-data within the end, wherein:
[0019] If the radiation coefficient of the dangerous end is greater than or equal to the preset radiation coefficient of the dangerous end or the correlation coefficient of the sub-data within the end is less than the preset correlation coefficient of the sub-data within the end, the first extraction method is to determine the number of selected dangerous sub-data according to the interaction coefficient of the associated end device and the environmental variation value;
[0020] If the dangerous end radiation coefficient is less than the preset dangerous end radiation coefficient and the intra-end sub-data correlation coefficient is greater than or equal to the preset intra-end sub-data correlation coefficient, the first extraction method is to select the dangerous sub-data according to the intra-end influence threshold.
[0021] Furthermore, the categories of dangerous gathering areas are determined according to the similarity of abnormality time and the impact value of key sub-data. The categories of dangerous gathering areas include:
[0022] A type of dangerous gathering area where the mutation time similarity is greater than or equal to the preset mutation time similarity and the key sub-data impact value is greater than or equal to the preset key sub-data impact value;
[0023] A second type of dangerous gathering area where the similarity of the mutation time is less than the preset similarity of the mutation time or the impact value of the key sub-data is less than the preset impact value of the key sub-data.
[0024] Furthermore, a second extraction method is determined according to the category of the dangerous concentration area, wherein:
[0025] For a type of dangerous clustering area, the second extraction method is to select dangerous sub-data according to the influencing factors;
[0026] For the second type of dangerous concentration areas, the second extraction method is to determine the number of dangerous sub-data in the selected dangerous concentration areas according to the comprehensive evaluation value.
[0027] Furthermore, the clustering method is determined according to the feature conversion coefficient and the number of data sources, where:
[0028] If the feature conversion coefficient is greater than or equal to the preset feature conversion coefficient or the number of data sources is greater than or equal to the preset number of data sources, the clustering method is to determine the associated source group according to the distance threshold;
[0029] If the feature conversion coefficient is less than the preset feature conversion coefficient and the number of data sources is less than the preset number of data sources, the clustering method is to take the set of data sources as an associated source group;
[0030] The distance threshold is positively correlated with the conversion reference value;
[0031] When determining the associated source group based on the distance threshold, the distance threshold is reduced and adjusted according to the clustering influence coefficient;
[0032] The reduction value of the distance threshold is positively correlated with the clustering influence coefficient.
[0033] Furthermore, the data flow anomaly index is determined according to the fluctuation difference, where:
[0034] If the fluctuation difference is greater than or equal to the preset fluctuation difference, the data flow anomaly index is determined according to the data flow fluctuation index and the mutation threshold;
[0035] If the fluctuation difference is less than the preset fluctuation difference, the data flow anomaly index is determined according to the environmental impact coefficient and the fluctuation trend value.
[0036] Compared with the prior art, the beneficial effect of the present invention lies in that, in the technical scheme of the present invention, the data status is determined according to the data source correlation and the mining difficulty coefficient, and the data source correlation and the mining difficulty coefficient effectively reflect the correlation degree of the data source and the analysis difficulty of the data source, and then the strategic analysis method is determined according to the data status, so that the selection of the strategic analysis method is more in line with the actual application scenario, which not only improves the pertinence and effectiveness of data analysis, but also avoids redundant processing when the data source correlation is high, thereby improving the data mining efficiency.
[0037] Furthermore, the present invention effectively reflects the distribution status of equipment corresponding to the dangerous sub-data in the data source and the degree of difference of the dangerous sub-data through the dangerous sub-data distribution coefficient and the dangerous turbulence value, and then adaptively selects different processing methods according to the dangerous sub-data distribution coefficient and the dangerous turbulence value, so that the sub-data selected by the processing method can better mine valuable information and can efficiently mine sub-data with potential fault information, thereby improving the efficiency and accuracy of fault identification.
[0038] Furthermore, the present invention effectively reflects the influence degree of the dangerous end and the correlation of the sub-data within the dangerous end according to the radiation coefficient of the dangerous end and the correlation coefficient of the sub-data within the end, and then adaptively selects different first extraction methods based on the radiation coefficient of the dangerous end and the correlation coefficient of the sub-data within the end, so that the first extraction method can more comprehensively evaluate the source and scope of the risk, thereby extracting sub-data closely related to the risk, thereby improving the accuracy and reliability of the sub-data selection results.
[0039] Furthermore, the present invention effectively reflects the degree of similarity in the time of abnormalities of the dangerous sub-data and the degree of influence of the key sub-data in the dangerous concentration area through the similarity in the time of abnormalities and the influence value of the key sub-data, and then determines the category of the dangerous concentration area according to the similarity in the time of abnormalities and the influence value of the key sub-data, and then adaptively selects different second extraction methods according to the category of the dangerous concentration area, while improving the pertinence and accuracy of the sub-data extraction, it also ensures the comprehensiveness of the sub-data selection, thereby improving the efficiency and accuracy of fault identification.
[0040] Furthermore, the present invention effectively reflects the degree of feature similarity of each data source through feature conversion coefficients and the number of data sources, and then adaptively selects different clustering methods according to the feature conversion coefficients and the number of data sources, so that the selection of clustering method is more in line with the actual data source status, avoiding the problem of poor accuracy of intelligence mining due to poor similarity of data sources in the determined associated source group, and thus being able to screen out representative data sources, thereby improving the efficiency of fault identification.
[0041] Furthermore, the present invention selects data sources in the associated source group based on the data flow anomaly index. The data flow anomaly index effectively reflects the degree of correlation of the abnormal situation of the data flow, and can effectively identify the data source with a larger degree of feature representativeness in the associated source group, which helps to avoid the use of erroneous or unreliable data sources for subsequent analysis, thereby improving the data quality in the data set. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic diagram of the intelligence mining and analysis method in the industrial scenario of the present invention;
[0043] Figure 2 A flow chart of the present invention for determining a strategy analysis method according to data status;
[0044] Figure 3 This is a flow chart of the present invention for determining a processing method according to a risk sub-data distribution coefficient and a risk turbulence value;
[0045] Figure 4 The present invention is a flow chart for determining the first extraction method according to the radiation coefficient of the dangerous end and the correlation coefficient of the sub-data within the end. DETAILED DESCRIPTION
[0046] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0047] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.
[0048] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the drawings. This is merely for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.
[0049] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0050] See also Figures 1 to 4 As shown, the present invention provides an intelligence mining and analysis method in an industrial scenario, comprising:
[0051] Obtain each data source, determine the data status according to the data source correlation and mining difficulty coefficient, and determine the strategic analysis method according to the data status. The strategic analysis method is to extract and analyze the sub-data or select and analyze the associated source group;
[0052] When extracting and analyzing the sub-data, a processing method is determined according to the distribution coefficient of the dangerous sub-data and the dangerous turbulence value to obtain a sub-data set. The processing method is to determine a first extraction method according to the radiation coefficient of the dangerous end and the correlation coefficient of the sub-data within the end, or to determine a second extraction method according to the category of the dangerous gathering area;
[0053] The first extraction method is to determine the number of selected dangerous sub-data according to the interaction coefficient of the associated terminal device and the environmental variation value, or to select the dangerous sub-data according to the influencing threshold within the terminal; the second extraction method is to select the dangerous sub-data according to the influencing factor, or to determine the number of selected dangerous sub-data in the dangerous concentration area according to the comprehensive evaluation value;
[0054] When selecting and analyzing the associated source group, the clustering method is determined according to the feature conversion coefficient and the number of data sources, and the data source in the associated source group is selected according to the data flow anomaly index, and the sub-data corresponding to each selected data source is used as the sub-data set.
[0055] The application scenario of the present invention is effective data screening before data intelligence mining. The present invention includes several data sources, a single data source contains several sub-data generated in the current monitoring cycle, and the single sub-data is a type of sub-data or a type of sub-data. The single type of sub-data is the monitoring parameters monitored in real time by a single industrial equipment at each time point during the use time, and the monitoring parameters include but are not limited to temperature, pressure, flow, speed and voltage. The time point is set by the user, and a time point setting method is provided. Every 1s is recorded as a time point, that is, the monitoring parameters are recorded once every 1s; the single type of sub-data is the text information corresponding to a single industrial equipment, and the text information is not limited to daily maintenance records, status assessment reports and fault reports; industrial equipment includes but is not limited to presses, assembly machines and compressors, and the details are not repeated. A continuous cyclic monitoring cycle is set in the present invention, and the data status is determined once at the end of each monitoring cycle. The duration of the monitoring cycle can be set according to the needs of the user. The greater the user's demand for monitoring accuracy, the shorter the duration of the monitoring cycle. A value of a monitoring cycle is provided, and the monitoring cycle is 24h.
[0056] In the present invention, several historical records are correspondingly set up, and any historical record records the data flow anomaly index, data source correlation, mining difficulty coefficient, disorder threshold, mutation coefficient, mutation reference value, dangerous sub-data distribution coefficient, dangerous turbulence value and dangerous end radiation coefficient in the historical process of at least one data screening, and each historical record corresponds to a qualified mark, which records whether the data screening process meets user requirements. The qualified mark can be recorded manually. It can be understood that the user can determine whether the data screening process meets the requirements based on self-set indicators. The self-set indicators can be but not limited to the wrong screening index, which will not be elaborated here. The wrong screening index is the number of sub-data in the sub-data set that is not fault data.
[0057] When selecting a data source in an associated source group according to a data flow anomaly index, a data source whose data flow anomaly index is greater than a preset data flow anomaly index is selected. The value of the preset data flow anomaly index can be determined by the user according to the actual application scenario. The greater the user's demand for improving the accuracy of fault information mining, the greater the value of the preset data flow anomaly index. A value of a preset data flow anomaly index is provided, and the historical records of selecting a data source in an associated source group according to the data flow anomaly index are detected, and the average value of the data flow anomaly index corresponding to the historical records that can meet the user's needs is recorded as the preset data flow anomaly index;
[0058] In the present invention, subsequent knowledge graph construction is performed through sub-data sets. When industrial equipment fails, users can query the knowledge graph by inputting the fault phenomenon or related parameters to quickly locate the fault information, thereby improving the efficiency of fault information mining. By screening the sub-data sets, redundant data can be reduced, so that the constructed knowledge graph can support more reliable knowledge representation and reasoning, and can also reduce the complexity and time cost of knowledge graph construction. When constructing the knowledge graph, key entities are identified from the sub-data sets using named entity recognition technology, and context information is analyzed through context analysis functions and predefined rule sets to extract the relationship between entities. The Neo4j graph database is used as a tool for storing and querying knowledge. According to the extracted key information, nodes (representing entities) and edges (representing relationships) are created in the graph database. For each identified entity, a corresponding node is created in the graph database. For each relationship between a pair of entities, an edge connecting the two nodes is created in the graph database to construct a knowledge graph containing entities and relationships. This is content that is easy for technicians in this field to understand and will not be described in detail.
[0059] Specifically, when the data status is that the data source correlation is less than the preset data source correlation or the mining difficulty coefficient is greater than or equal to the preset mining difficulty coefficient, the strategic analysis method is to perform extraction analysis on the data source.
[0060] The data state includes a first data state and a second data state, the first data state is that the data source association degree is less than the preset data source association degree or the mining difficulty coefficient is greater than or equal to the preset mining difficulty coefficient, and the second data state is that the data source association degree is greater than or equal to the preset data source association degree and the mining difficulty coefficient is less than the preset mining difficulty coefficient;
[0061] The data source correlation is the average value of the sub-correlations corresponding to each data source in the current monitoring period. For a single data source, the data source is recorded as the target data source, and other data sources other than the target data source are recorded as reference data sources. The sub-correlation corresponding to the target data source is the average value of the correlation reference values corresponding to each reference data source. The method for confirming the correlation reference value is that, for a single reference data source, the reference data source is recorded as the pre-analysis reference data source, and the correlation reference value corresponding to the pre-analysis reference data source = the first correlation + the second correlation. The confirmation method of the first correlation is that each type of sub-data in the pre-analysis reference data source in the current monitoring period is recorded as the first sub-data, and each type of sub-data in the target data source in the current monitoring period is recorded as the second sub-data. The first correlation = the average value of the maximum correlation value corresponding to each first sub-data + the number of second sub-data affected / the total amount of second sub-data. For a single first sub-data, the maximum value of the correlation reference values corresponding to the first sub-data and each second sub-data is recorded as the maximum correlation value; the second sub-data affected is the second sub-data corresponding to the maximum correlation value of each first sub-data.
[0062] For any two first sub-data, the calculation formula of the correlation reference value r corresponding to the two first sub-data is:
[0063]
[0064] Where n is the number of time points within the monitoring time corresponding to a single first sub-data; and are the values of the monitoring data corresponding to the i-th time point within the monitoring time corresponding to the two first sub-data, for The average value of the monitoring data corresponding to each time point within the monitoring time corresponding to the corresponding first sub-data, for The average value of the monitoring data corresponding to each time point within the monitoring time corresponding to the corresponding first sub-data, i=1, 2, 3, ..., n;
[0065] The second relevance is confirmed in the following manner: for two data sources, the keywords existing in both data sources in the current monitoring period are recorded as the same keywords, the second relevance = (the larger value of the number of the same keywords / the total number of keywords corresponding to the two data sources) + (the absolute value of the difference in the number of the same sub-data corresponding to the two data sources / the larger value of the number of the same sub-data corresponding to the two data sources), each data source corresponds to a second sub-data set, the second sub-data set is the set of the second sub-data in a single data source in the current monitoring period, the total number of keywords corresponding to a single data source is the number of keywords in the second sub-data set corresponding to the data source, and the same sub-data is the second sub-data containing the same keywords;
[0066] The mining difficulty coefficient is the average of the sub-difficulty coefficients corresponding to each data source in the current monitoring period. The sub-difficulty coefficient corresponding to a single data source = the number of second sub-data in the data source in the current monitoring period / (the number of first sub-data in the data source in the current monitoring period + the number of second sub-data);
[0067] The values of the preset data source association degree and the preset mining difficulty coefficient can be determined by the user according to the actual application scenario. The larger the value of the preset data source association degree and the smaller the value of the preset mining difficulty coefficient, the greater the user's demand for extraction and analysis of the data source. The values of the preset data source association degree and the preset mining difficulty coefficient are provided, and the historical records of extraction and analysis of the data source are detected. The average value of the data source association degree corresponding to the historical records that can meet the user's needs is recorded as the preset data source association degree, and the average value of the mining difficulty coefficient corresponding to the historical records that can meet the user's needs is recorded as the preset mining difficulty coefficient.
[0068] Specifically, when the data status is that the data source correlation is greater than or equal to the preset data source correlation and the mining difficulty coefficient is less than the preset mining difficulty coefficient, the strategy analysis method is to select and analyze the associated source group.
[0069] Specifically, each sub-data category is equipped with a corresponding dangerous sub-data determination method, among which:
[0070] For a type of sub-data, the dangerous sub-data is determined by using a disorder threshold and a mutation coefficient;
[0071] For the second type of sub-data, the dangerous sub-data is determined based on the proportion of influencing keywords.
[0072] Among them, for a type of sub-data, a type of sub-data whose disorder threshold is greater than a preset disorder threshold and whose mutation coefficient is greater than a preset mutation coefficient is regarded as dangerous sub-data;
[0073] For the second-category sub-data, the second-category sub-data whose impact keyword ratio is greater than the preset impact keyword ratio is regarded as dangerous sub-data;
[0074] The disorder threshold is the standard deviation of the values of the monitoring parameters corresponding to each time point in a single type of sub-data;
[0075] For a single type of sub-data, the mutation coefficient = the number of mutation monitoring values in the type of sub-data / the total value of the monitoring parameter corresponding to each time point in the type of sub-data; the method for confirming the mutation monitoring value is, for a single time point, the time point is recorded as the target time point, the time point adjacent to the target time point is recorded as the reference time point, the average value of the monitoring difference corresponding to each reference time point is recorded as the mutation reference value, and the monitoring data value corresponding to the time point when the mutation reference value is greater than the preset mutation reference value is recorded as the mutation monitoring value; the monitoring difference is the absolute value of the difference between the monitoring data value corresponding to a single reference time point and the monitoring data value corresponding to the target time point;
[0076] For a single second-category sub-data, the proportion of influencing keywords = the total number of influencing keywords in the second-category sub-data / the total number of keywords in the second-category sub-data, the influencing keywords are keywords with an influencing number greater than the preset influencing number, all the second-category sub-data in the sub-data set corresponding to the historical records that can meet the user's needs are recorded as the first text, and all the second-category sub-data in the sub-data set corresponding to the historical records that cannot meet the user's needs are recorded as the second text, for a single keyword, the difference between the number of times the keyword appears in the first text and the number of times it appears in the second text is recorded as the influencing number, it can be understood that the present invention reflects the degree of influence of the keyword on the fault information contained in the text through the influencing number, can accurately identify the influencing keywords that are highly related to the fault information, and is helpful to quickly locate and analyze the sub-data containing the fault information;
[0077] The user can determine the values of the preset disorder threshold, preset mutation coefficient, preset influencing keyword ratio, preset mutation reference value and preset number of influence times according to the actual application scenario. The higher the user's accurate demand for the dangerous sub-data judgment, the larger the values of the preset disorder threshold, preset mutation coefficient and preset influencing keyword ratio. Provide a preset disorder threshold, preset mutation coefficient, preset influencing keyword ratio, preset mutation reference value and preset number of influence values, detect the historical records of dangerous sub-data determined according to the disorder threshold and mutation coefficient, and record the average value of the disorder threshold corresponding to the historical records that can meet the user's needs as the preset disorder threshold, the average value of the mutation coefficient corresponding to the historical records that can meet the user's needs as the preset mutation coefficient, and the average value of the mutation reference value corresponding to the historical records that can meet the user's needs as the preset mutation reference value. The preset influencing keyword ratio is 70%, and the preset number of influences is 50 times.
[0078] Specifically, the processing method is determined according to the distribution coefficient of the dangerous sub-data and the dangerous turbulence value, where:
[0079] If the risk sub-data distribution coefficient is greater than or equal to the preset risk sub-data distribution coefficient or the risk turbulence value is greater than or equal to the preset risk turbulence value, the processing method is to determine the first extraction method according to the risk end radiation coefficient and the sub-data correlation coefficient within the end;
[0080] If the risk sub-data distribution coefficient is less than the preset risk sub-data distribution coefficient and the risk turbulence value is less than the preset risk turbulence value, the processing method is to determine the second extraction method according to the risk concentration area category.
[0081] Among them, the confirmation method of the dangerous sub-data distribution coefficient and the dangerous turbulence value is as follows: for a single data source, the data source is recorded as the first target data source, and each dangerous sub-data in the first target data source in the current monitoring period is recorded as analysis sub-data. The dangerous sub-data distribution coefficient is the average value of the sub-distribution coefficients corresponding to each analysis sub-data. For a single analysis sub-data, the analysis sub-data is recorded as the target analysis sub-data, and other analysis sub-data other than the target analysis sub-data are recorded as reference analysis sub-data. The average value of the shortest distance from the industrial equipment corresponding to the target analysis sub-data to the industrial equipment corresponding to each reference analysis sub-data is recorded as the sub-distribution coefficient; dangerous turbulence value = standard deviation of the dangerous coefficient corresponding to each first-class dangerous sub-data in the first target data source in the current monitoring period + standard deviation of the proportion of influencing keywords corresponding to each second-class dangerous sub-data in the first target data source in the current monitoring period. The first-class dangerous sub-data is the dangerous sub-data whose sub-data category is the first-class sub-data, and the second-class dangerous sub-data is the dangerous sub-data whose sub-data category is the second-class sub-data. The dangerous coefficient corresponding to a single first-class dangerous sub-data = the disorder threshold corresponding to the first-class dangerous sub-data + the mutation coefficient corresponding to the first-class dangerous sub-data;
[0082] The values of the preset danger sub-data distribution coefficient and the preset danger turbulence value can be determined by the user according to the actual application scenario. The larger the values of the preset danger sub-data distribution coefficient and the preset danger turbulence value are, the greater the user's need to determine the second extraction method according to the danger concentration area category. A value of the preset danger sub-data distribution coefficient and the preset danger turbulence value is provided, and the historical records of determining the second extraction method according to the danger concentration area category are detected. The average value of the danger sub-data distribution coefficient corresponding to the historical records that can meet the user's needs is recorded as the preset danger sub-data distribution coefficient, and the average value of the danger turbulence values corresponding to the historical records that can meet the user's needs is recorded as the preset danger turbulence value.
[0083] Specifically, the first extraction method is determined according to the radiation coefficient of the dangerous end and the correlation coefficient of the sub-data within the end, wherein:
[0084] If the radiation coefficient of the dangerous end is greater than or equal to the preset radiation coefficient of the dangerous end or the correlation coefficient of the sub-data within the end is less than the preset correlation coefficient of the sub-data within the end, the first extraction method is to determine the number of selected dangerous sub-data according to the interaction coefficient of the associated end device and the environmental variation value;
[0085] If the dangerous end radiation coefficient is less than the preset dangerous end radiation coefficient and the intra-end sub-data correlation coefficient is greater than or equal to the preset intra-end sub-data correlation coefficient, the first extraction method is to select the dangerous sub-data according to the intra-end influence threshold.
[0086] Among them, the confirmation method of the dangerous end is to perform association analysis on each analysis sub-data in the first target data source within the current monitoring period. When performing association analysis on a single analysis sub-data, the analysis sub-data is recorded as the first target analysis sub-data, and other analysis sub-data other than the first target analysis sub-data that are not recorded in the association combination are recorded as the first reference analysis sub-data. The first reference analysis sub-data and the first target analysis sub-data whose association reference value with the first target analysis sub-data is greater than the preset association reference value or whose proportion of similar keywords is greater than the preset proportion of similar keywords are recorded as the association combination, and the collection of the first reference analysis sub-data and the first target analysis sub-data are recorded as the association combination, and the association analysis is continued for the analysis sub-data that are not recorded in the association combination until all the analysis sub-data are recorded in the association combination, then the association analysis is stopped, and each association combination corresponds to a dangerous end, and the dangerous end corresponding to a single association combination is the smallest rectangular area that can contain the industrial equipment corresponding to each dangerous sub-data in the association combination;
[0087] The method for confirming the radiation coefficient of the dangerous end is to record a single dangerous end as the target dangerous end and record other dangerous ends other than the target dangerous end as reference dangerous ends. The radiation coefficient of the dangerous end = the number of reference dangerous ends with overlapping areas with the target dangerous end + the total amount of dangerous sub-data corresponding to each industrial equipment in the target dangerous end;
[0088] The intra-end sub-data correlation coefficient is the average value of the intra-end impact thresholds corresponding to each sub-data corresponding to each industrial equipment in a single dangerous end;
[0089] The confirmation methods of the in-terminal impact thresholds corresponding to the first type of sub-data and the second type of sub-data are different, among which:
[0090] For one type of sub-data, the intra-end impact threshold and the intra-end correlation are positively correlated;
[0091] For the second type of sub-data, the intra-end impact threshold and the intra-end impact value are positively correlated;
[0092] The confirmation method of the intra-end correlation and the intra-end influence value is to record each type of sub-data and each type of sub-data corresponding to each industrial equipment in a single dangerous end as type-one sub-data and type-two sub-data within the end respectively. For a single type-one sub-data within the end, the intra-end correlation corresponding to the type-one sub-data within the end is the average value of the correlation reference values corresponding to the type-one sub-data within the end and other type-one sub-data within the end, excluding the type-one sub-data within the end; for a single type-two sub-data within the end, the intra-end influence corresponding to the type-two sub-data within the end is the average value of the proportion of similar keywords corresponding to the type-two sub-data within the end and other type-two sub-data within the end, excluding the type-two sub-data within the end;
[0093] For the two sub-data, the proportion of similar keywords = the larger value of the number of identical associated words in the two sub-data / the number of keywords corresponding to the two sub-data;
[0094] The values of the preset correlation reference value, the preset proportion of similar keywords, the preset dangerous end radiation coefficient and the preset correlation coefficient of sub-data within the terminal can be determined by the user according to the actual application scenario. The higher the user's demand for improving the correlation degree of dangerous sub-data in a single dangerous terminal, the larger the values of the preset correlation reference value and the preset proportion of similar keywords. A preset correlation reference value and a preset proportion of similar keywords are provided. The preset correlation reference value is 0.7, and the preset proportion of similar keywords is 60%. The larger the value of the preset dangerous end radiation coefficient and the smaller the value of the preset correlation coefficient of sub-data within the terminal, the greater the user's demand for selecting dangerous sub-data according to the terminal impact threshold. A preset dangerous end radiation coefficient and a preset correlation coefficient of sub-data within the terminal are provided. The historical records of selecting sub-data according to the terminal impact threshold are detected, and the average value of the dangerous end radiation coefficient corresponding to the historical records that can meet the user's needs is recorded as the preset dangerous end radiation coefficient, and the average value of the correlation coefficient of sub-data within the terminal corresponding to the historical records that can meet the user's needs is recorded as the preset sub-data correlation coefficient within the terminal;
[0095] The first extraction method is to determine the number of selected dangerous sub-data according to the interaction coefficient of the associated terminal device and the environmental variation value, wherein the number of selected dangerous sub-data is positively correlated with the sum of the interaction coefficient of the associated terminal device and the environmental variation value. It should be noted that when selecting dangerous sub-data, selection is performed in descending order of the impact threshold within the terminal until the number of selected dangerous sub-data reaches the number of selected dangerous sub-data;
[0096] The first extraction method is to select dangerous sub-data according to the end-end impact threshold, and select dangerous sub-data whose end-end impact threshold is greater than the preset end-end impact threshold;
[0097] The confirmation method of the associated terminal device interaction coefficient is as follows: for a single dangerous terminal, the dangerous terminal is recorded as the first target dangerous terminal, and the dangerous terminal with the overlapping area with the first target dangerous terminal and the first target dangerous terminal are recorded as the analysis dangerous terminal. The associated terminal device interaction coefficient = (the maximum value of the fluctuation coefficients corresponding to each analysis dangerous terminal - the minimum value of the fluctuation coefficients corresponding to each analysis dangerous terminal) / the maximum value of the fluctuation coefficients corresponding to each analysis dangerous terminal. The fluctuation coefficient corresponding to a single analysis dangerous terminal is the standard deviation of the intra-terminal impact threshold corresponding to each dangerous sub-data in the analysis dangerous terminal;
[0098] For a single dangerous end, the environmental variation value = the area of the dangerous end / the number of industrial equipment in the dangerous end. The environmental variation value can reflect the degree of interference of a single industrial equipment by other equipment.
[0099] The value of the preset end-end impact threshold can be determined by the user according to the actual application scenario. The greater the user's demand for improving the efficiency of fault information mining, the larger the value of the preset end-end impact threshold. A value of the preset end-end impact threshold is provided, and the historical records of selecting dangerous sub-data according to the end-end impact threshold are detected. The average value of the end-end impact threshold corresponding to the historical records that can meet the user's needs is recorded as the preset end-end impact threshold.
[0100] Specifically, the categories of dangerous gathering areas are determined based on the similarity of mutation time and the impact value of key sub-data. The categories of dangerous gathering areas include:
[0101] A type of dangerous gathering area where the mutation time similarity is greater than or equal to the preset mutation time similarity and the key sub-data impact value is greater than or equal to the preset key sub-data impact value;
[0102] A second type of dangerous gathering area where the similarity of the mutation time is less than the preset similarity of the mutation time or the impact value of the key sub-data is less than the preset impact value of the key sub-data.
[0103] Among them, the danger gathering area is the smallest rectangular area that can contain the industrial equipment corresponding to each danger sub-data of a single data source in the current monitoring period. For a danger gathering area, the mutation time similarity = 1 / [(the maximum value of the mutation time length corresponding to each type of sub-data in the danger gathering area - the minimum value of the mutation time length corresponding to each type of sub-data in the danger gathering area) + (the maximum value of the mutation coefficient corresponding to each type of sub-data in the danger gathering area - the minimum value of the mutation coefficient corresponding to each type of sub-data in the danger gathering area)], and the method for confirming the mutation time length is that, for a single data source, the time point corresponding to the first mutation monitoring value that appears in the data source in order from early to late and the time point corresponding to the last mutation monitoring value that appears are recorded as the mutation time length corresponding to the data source;
[0104] The method for confirming the impact value of key sub-data is as follows: for a single data source, the sub-data of the type with the largest risk factor in the data source is recorded as key sub-data, and the sub-data of the type other than key sub-data in the data source is recorded as non-key sub-data. The impact value of key sub-data = the number of non-key sub-data with an associated reference value greater than 0.7 of the key sub-data / the total amount of a type of sub-data in the data source;
[0105] The values of the preset anomaly time similarity and the preset key sub-data influence value can be determined by the user according to the actual application scenario. The smaller the values of the preset anomaly time similarity and the preset key sub-data influence value are, the greater the user's need to select dangerous sub-data according to the influencing factor. A value of the preset anomaly time similarity and the preset key sub-data influence value is provided, and the historical records of selecting dangerous sub-data according to the influencing factor are detected. The average value of the anomaly time similarity corresponding to the historical records that can meet the user's needs is recorded as the preset anomaly time similarity, and the preset key sub-data influence value is 60%.
[0106] Specifically, the second extraction method is determined according to the category of the dangerous gathering area, wherein:
[0107] For a type of dangerous clustering area, the second extraction method is to select dangerous sub-data according to the influencing factors;
[0108] For the second type of dangerous concentration areas, the second extraction method is to determine the number of dangerous sub-data in the selected dangerous concentration areas according to the comprehensive evaluation value.
[0109] Among them, the second extraction method is to select the dangerous sub-data according to the influencing factor, and select the dangerous sub-data whose influencing factor is greater than the preset influencing factor. The confirmation method of the influencing factor is that for the first type of dangerous sub-data, the influencing factor is positively correlated with the risk coefficient, and for the second type of dangerous sub-data, the influencing factor is positively correlated with the proportion of influencing keywords;
[0110] The second extraction method is to determine the number of dangerous sub-data in the selected dangerous gathering area according to the comprehensive evaluation value. The number of dangerous sub-data in the selected dangerous gathering area is negatively correlated with the comprehensive evaluation value. It should be noted that when selecting dangerous sub-data in the dangerous gathering area, the selection is carried out in descending order of the influence factors until the number of dangerous sub-data in the dangerous gathering area to be selected is reached. The comprehensive evaluation value = the similarity of the abnormal time - the influence value of the key sub-data;
[0111] The value of the preset influence factor can be determined by the user according to the actual application scenario. The higher the user's demand for improving the efficiency of fault information mining, the larger the value of the preset influence factor. A value of the preset influence factor is provided, and the historical records of selecting dangerous sub-data according to the influence factor are detected. The average value of the influence factor corresponding to the historical records that can meet the user's needs is recorded as the preset influence factor.
[0112] Specifically, the clustering method is determined according to the feature conversion coefficient and the number of data sources, where:
[0113] If the feature conversion coefficient is greater than or equal to the preset feature conversion coefficient or the number of data sources is greater than or equal to the preset number of data sources, the clustering method is to determine the associated source group according to the distance threshold;
[0114] If the feature conversion coefficient is less than the preset feature conversion coefficient and the number of data sources is less than the preset number of data sources, the clustering method is to take the set of data sources as an associated source group;
[0115] The distance threshold is positively correlated with the conversion reference value;
[0116] When determining the associated source group based on the distance threshold, the distance threshold is reduced and adjusted according to the clustering influence coefficient;
[0117] The reduction value of the distance threshold is positively correlated with the clustering influence coefficient.
[0118] Among them, the feature conversion coefficient = the area of the minimum rectangle that can contain the coordinate points corresponding to each data source + the maximum value of the lengths of each side of the minimum rectangle that can contain the coordinate points corresponding to each data source, detect the first eigenvalue and the second eigenvalue corresponding to each data source, the first eigenvalue = the standard deviation of the risk coefficient corresponding to each type of risk sub-data in a single data source + the number of type one risk sub-data in a single data source, the second eigenvalue = the standard deviation of the proportion of influencing keywords corresponding to each type two risk sub-data in a single data source + the number of type two risk sub-data in a single data source, establish a rectangular coordinate system with the first eigenvalue as the horizontal coordinate and the second eigenvalue as the vertical coordinate, and obtain the coordinate points of each data source in the rectangular coordinate system, and then intuitively observe the feature correlation degree of each data source;
[0119] The values of the preset feature conversion coefficient and the preset number of data sources can be determined by the user according to the actual application scenario. The smaller the values of the preset feature conversion coefficient and the preset number of data sources are, the greater the user's need to determine the associated source group according to the distance threshold. A value of the preset feature conversion coefficient and the preset number of data sources is provided, and a historical record of a set of data sources is detected as an associated source group. The average value of the feature conversion coefficient corresponding to the historical records that can meet the user's needs is recorded as the preset feature conversion coefficient, and the average value of the number of data sources corresponding to the historical records that can meet the user's needs is recorded as the preset number of data sources;
[0120] When determining the associated source group according to the distance threshold, cluster analysis is performed on the coordinate points corresponding to each data source in the order of characteristic coefficients from large to small. When cluster analysis is performed on a single coordinate point, the coordinate point is recorded as the target coordinate point, and other coordinate points other than the target coordinate point are recorded as reference coordinate points. The data source corresponding to the reference coordinate point whose interval distance with the target coordinate point is less than the distance threshold and the data source corresponding to the target coordinate point are recorded as an associated source group, and cluster analysis is continued for the coordinate points corresponding to the data sources not recorded in the associated source group until all data sources are recorded in the associated source group, and then the cluster analysis is stopped;
[0121] The characteristic coefficient corresponding to a single data source = the first characteristic value corresponding to the data source + the second characteristic value corresponding to the data source; the interval distance is the shortest distance between the coordinate points corresponding to the two data sources; the conversion reference value = characteristic conversion coefficient + the number of data sources;
[0122] The clustering influence coefficient is the average value of the sub-clustering influence reference values corresponding to each associated source group. For a single associated source group, the smallest circle that can contain the coordinate points corresponding to each data source in the associated source group is recorded as the reference circle. The sub-clustering influence reference value = clustering distance average value - clustering distance difference. The maximum value of the shortest distance from the coordinate points corresponding to each data source in a single associated source group to the center of the reference circle corresponding to the associated source group is recorded as e1, and the minimum value is recorded as e2. The clustering difference = (e1-e2) / e1. The clustering distance average value is the average value of the shortest distance from the coordinate points corresponding to each data source in a single associated source group to the center of the reference circle corresponding to the associated source group.
[0123] Specifically, the data flow anomaly index is determined according to the fluctuation difference, where:
[0124] If the fluctuation difference is greater than or equal to the preset fluctuation difference, the data flow anomaly index is determined according to the data flow fluctuation index and the mutation threshold;
[0125] If the fluctuation difference is less than the preset fluctuation difference, the data flow anomaly index is determined according to the environmental impact coefficient and the fluctuation trend value.
[0126] Among them, for a single associated source group, the maximum value of the data flow fluctuation index corresponding to each data source in the associated source group is recorded as H1, and the minimum value is recorded as H2. The fluctuation difference = (H1-H2) / H1. The data flow fluctuation index is the standard deviation of the data flow corresponding to each time point of a single data source in the current monitoring period. The data flow corresponding to each data source is monitored by the SolarWinds NPM monitoring tool. This is easy for technicians in this field to understand and will not be described in detail.
[0127] The value of the preset fluctuation difference can be determined by the user according to the actual application scenario. The smaller the value of the preset fluctuation difference is, the greater the user's demand for determining the data flow anomaly index according to the data flow fluctuation index and the mutation threshold is. A value of the preset fluctuation difference is provided, and the historical records of determining the data flow anomaly index according to the data flow fluctuation index and the mutation threshold are detected. The average value of the fluctuation difference corresponding to the historical records that can meet the user's needs is recorded as the preset fluctuation difference;
[0128] If the fluctuation difference is greater than or equal to the preset fluctuation difference, the data flow anomaly index = data flow fluctuation index + mutation threshold;
[0129] If the fluctuation difference is less than the preset fluctuation difference, the data flow anomaly index = environmental impact coefficient + fluctuation trend value;
[0130] For a single data source, the mutation threshold is the absolute value of the difference between the maximum and minimum values of the data flow corresponding to each time point of the data source in the current monitoring period. The environmental impact coefficient = the total amount of sub-data of the data source in the current monitoring period + the number of industrial equipment corresponding to the sub-data of the data source in the current monitoring period;
[0131] Fluctuation trend value = number of incremental points / total number of time points in the current monitoring cycle. For a time point, this time point is recorded as the first target time point, and the time point adjacent to the target data point and monitored earlier than the target time point is recorded as the first reference time point. If the data flow corresponding to the first target time point is greater than the data flow corresponding to the first reference time point, then the first target time point is the incremental point, and the number of incremental points is the total number of incremental points of a single data source in the current monitoring cycle.
[0132] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
[0133] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An intelligence mining and analysis method in an industrial scenario, characterized in that: include: Obtain each data source, determine the data status according to the data source correlation and mining difficulty coefficient, and determine the strategic analysis method according to the data status. The strategic analysis method is to extract and analyze the sub-data or select and analyze the associated source group; When extracting and analyzing the sub-data, a processing method is determined according to the distribution coefficient of the dangerous sub-data and the dangerous turbulence value to obtain a sub-data set. The processing method is to determine a first extraction method according to the radiation coefficient of the dangerous end and the correlation coefficient of the sub-data within the end, or to determine a second extraction method according to the category of the dangerous gathering area; The first extraction method is to determine the number of selected dangerous sub-data according to the interaction coefficient of the associated terminal device and the environmental variation value, or to select the dangerous sub-data according to the influencing threshold within the terminal; the second extraction method is to select the dangerous sub-data according to the influencing factor, or to determine the number of selected dangerous sub-data in the dangerous concentration area according to the comprehensive evaluation value; When selecting and analyzing the associated source group, the clustering method is determined according to the feature conversion coefficient and the number of data sources, and the data source in the associated source group is selected according to the data flow anomaly index, and the sub-data corresponding to each selected data source is used as the sub-data set; The distribution coefficient of the dangerous sub-data is the average value of the sub-distribution coefficients corresponding to each analysis sub-data; Dangerous turbulence value = standard deviation of the risk coefficient corresponding to each type of first-class dangerous sub-data in the first target data source within the current monitoring period + standard deviation of the proportion of influencing keywords corresponding to each type of second-class dangerous sub-data in the first target data source within the current monitoring period; Dangerous end radiation coefficient = the number of reference dangerous ends with overlapping areas with the target dangerous end + the total amount of dangerous sub-data corresponding to each industrial equipment in the target dangerous end; The intra-end sub-data correlation coefficient is the average value of the intra-end impact thresholds corresponding to each sub-data corresponding to each industrial equipment in a single dangerous end; The confirmation method of the associated terminal device interaction coefficient is as follows: for a single dangerous terminal, the dangerous terminal is recorded as the first target dangerous terminal, and the dangerous terminal with an overlapping area with the first target dangerous terminal and the first target dangerous terminal are both recorded as the analysis dangerous terminal, and the associated terminal device interaction coefficient = (the maximum value of the fluctuation coefficients corresponding to each analysis dangerous terminal - the minimum value of the fluctuation coefficients corresponding to each analysis dangerous terminal) / the maximum value of the fluctuation coefficients corresponding to each analysis dangerous terminal; Environmental variation value = area of the dangerous end / number of industrial equipment in the dangerous end; Feature conversion coefficient=area of the smallest rectangle that can contain the coordinate points corresponding to each data source+maximum value of lengths corresponding to the sides of the smallest rectangle that can contain the coordinate points corresponding to each data source.
2. The intelligence mining and analysis method in industrial scenarios according to claim 1 is characterized in that: When the data status is that the data source correlation is less than the preset data source correlation or the mining difficulty coefficient is greater than or equal to the preset mining difficulty coefficient, the strategy analysis method is to perform extraction analysis on the data source.
3. The intelligence mining and analysis method in industrial scenarios according to claim 1 is characterized in that: When the data status is that the data source correlation is greater than or equal to the preset data source correlation and the mining difficulty coefficient is less than the preset mining difficulty coefficient, the strategy analysis method is to select and analyze the associated source group.
4. The intelligence mining and analysis method in industrial scenarios according to claim 3 is characterized in that: Each sub-data category has a corresponding dangerous sub-data determination method, among which: For a type of sub-data, the dangerous sub-data is determined by using a disorder threshold and a mutation coefficient; For the second type of sub-data, the dangerous sub-data is determined based on the proportion of influencing keywords.
5. The intelligence mining and analysis method in industrial scenarios according to claim 4 is characterized in that: The processing method is determined according to the distribution coefficient of the dangerous sub-data and the dangerous turbulence value, where: If the risk sub-data distribution coefficient is greater than or equal to the preset risk sub-data distribution coefficient or the risk turbulence value is greater than or equal to the preset risk turbulence value, the processing method is to determine the first extraction method according to the risk end radiation coefficient and the sub-data correlation coefficient within the end; If the risk sub-data distribution coefficient is less than the preset risk sub-data distribution coefficient and the risk turbulence value is less than the preset risk turbulence value, the processing method is to determine the second extraction method according to the risk concentration area category.
6. The intelligence mining and analysis method in industrial scenarios according to claim 5 is characterized in that: The first extraction method is determined according to the radiation coefficient of the dangerous end and the correlation coefficient of the sub-data in the end, wherein: If the dangerous end radiation coefficient is greater than or equal to the preset dangerous end radiation coefficient or the terminal intra-sub-data correlation coefficient is less than the preset terminal intra-sub-data correlation coefficient, the first extraction method is to determine the number of selected dangerous sub-data according to the associated terminal device interaction coefficient and the environmental variation value; If the dangerous end radiation coefficient is less than the preset dangerous end radiation coefficient and the intra-end sub-data correlation coefficient is greater than or equal to the preset intra-end sub-data correlation coefficient, the first extraction method is to select the dangerous sub-data according to the intra-end influence threshold.
7. The intelligence mining and analysis method in industrial scenarios according to claim 3 is characterized in that: The categories of dangerous gathering areas are determined based on the similarity of abnormal time and the impact value of key sub-data. The categories of dangerous gathering areas include: A type of dangerous gathering area where the mutation time similarity is greater than or equal to the preset mutation time similarity and the key sub-data impact value is greater than or equal to the preset key sub-data impact value; A second type of dangerous gathering area where the similarity of the mutation time is less than the preset similarity of the mutation time or the impact value of the key sub-data is less than the preset impact value of the key sub-data.
8. The intelligence mining and analysis method in industrial scenarios according to claim 7 is characterized in that: The second extraction method is determined according to the category of the dangerous gathering area, wherein: For a type of dangerous clustering area, the second extraction method is to select dangerous sub-data according to the influencing factors; For the second type of dangerous concentration areas, the second extraction method is to determine the number of dangerous sub-data in the selected dangerous concentration areas according to the comprehensive evaluation value.
9. The intelligence mining and analysis method in industrial scenarios according to claim 3 is characterized in that: Determine the clustering method based on the feature conversion coefficient and the number of data sources. in, If the feature conversion coefficient is greater than or equal to the preset feature conversion coefficient or the number of data sources is greater than or equal to the preset number of data sources, the clustering method is to determine the associated source group according to the distance threshold; If the feature conversion coefficient is less than the preset feature conversion coefficient and the number of data sources is less than the preset number of data sources, the clustering method is to take the set of data sources as an associated source group; The distance threshold is positively correlated with the conversion reference value; When determining the associated source group based on the distance threshold, the distance threshold is reduced and adjusted according to the clustering influence coefficient; The reduction value of the distance threshold is positively correlated with the clustering influence coefficient.
10. The intelligence mining and analysis method in industrial scenarios according to claim 9 is characterized in that: The data flow anomaly index is determined based on the fluctuation difference, where: If the fluctuation difference is greater than or equal to the preset fluctuation difference, the data flow anomaly index is determined according to the data flow fluctuation index and the mutation threshold; If the fluctuation difference is less than the preset fluctuation difference, the data flow anomaly index is determined according to the environmental impact coefficient and the fluctuation trend value.
Citation Information
Patent Citations
Industrial data feature selection and abnormal value detection method, system, equipment and medium
CN118194204A
Multi-source data fusion method and device, electronic equipment and storage medium
CN119089388A
Cited By
Intelligent identification and evaluation method and system for AI modeling sub-problem
CN122412551A