Data screening method and device, equipment, storage medium

By constructing a spatiotemporal graph and calculating attention weights and information entropy, the problem of inaccurate screening of chemical production data is solved, enabling efficient utilization and accurate screening of chemical production data, and supporting production optimization and fault diagnosis.

CN120371894BActive Publication Date: 2025-11-04TANGSHAN CAOFEIDIAN LIANCHENG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510501131.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-11-04
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently utilize chemical production data, particularly its dynamic evolution patterns over time and its inherent spatial relationships. This results in inaccurate data filtering, failing to meet the needs of industrial production optimization and fault diagnosis.

Method used

By constructing a spatiotemporal graph, calculating the attention weights of nodes in the spatiotemporal dimension, stitching together target features, and determining the target weights based on information entropy, efficient screening of industrial production data can be achieved.

Benefits of technology

It enables efficient utilization and precise screening of chemical production data, and can identify the causes of equipment malfunctions and sudden changes in production efficiency, supporting industrial production optimization and fault diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371894B_ABST
    Figure CN120371894B_ABST
Patent Text Reader

Abstract

The application provides a data screening method and device, equipment and storage medium, and belongs to the technical field of data screening. The method comprises the following steps: calculating the attention weight of each node in a space-time graph in a space-time dimension, the space-time graph being constructed by various industrial production data, each kind of industrial production data being taken as a node in the space-time graph, and the space-time dimension comprising a time dimension and a space dimension; splicing the attention weight of each node in the space-time graph in the space-time dimension to obtain a target feature corresponding to each node; obtaining a target weight of the variation degree of each node based on the information entropy of the target feature; and screening the various industrial production data based on the target weight of the variation degree of each node. The data screening method and device, equipment and storage medium provided by the application can realize efficient utilization and accurate screening of data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of data screening, and more particularly relates to a data screening method and device, equipment and storage medium. BACKGROUND

[0002] In the chemical production industry, with the expansion of production scale and the increase of process complexity, the amount of data is growing explosively, covering equipment operation, production efficiency, product quality and other aspects. However, most of the current data screening methods have limitations. Traditional methods mostly rely on simple statistical analysis or preset rules to screen data, which is difficult to mine the dynamic evolution law of chemical production data in the time dimension, such as the gradual trend of key parameters in long-term equipment operation, and the internal relationship of data in different production links in the space dimension, so as to realize efficient utilization and accurate screening of chemical production data. SUMMARY

[0003] The purpose of the present application is to provide a data screening method and device, equipment and storage medium to realize efficient utilization and accurate screening of data.

[0004] The first aspect of the embodiment of the present application provides a data screening method, comprising:

[0005] calculating the attention weight of each node in the time-space dimension in the time-space graph, the time-space graph being constructed by a plurality of industrial production data, each industrial production data being a node in the time-space graph, the time-space dimension including a time dimension and a space dimension;

[0006] splicing the attention weight of each node in the time-space dimension in the time-space graph to obtain the target feature corresponding to each node;

[0007] obtaining the target weight of the variation degree of each node based on the information entropy of the target feature;

[0008] screening the plurality of industrial production data based on the target weight of the variation degree of each node.

[0009] The second aspect of the embodiment of the present application provides a data screening device, comprising:

[0010] a first weight calculation module for calculating the attention weight of each node in the time-space dimension in the time-space graph, the time-space graph being constructed by a plurality of industrial production data, each industrial production data being a node in the time-space graph, the time-space dimension including a time dimension and a space dimension;

[0011] a weight splicing module for splicing the attention weight of each node in the time-space dimension in the time-space graph to obtain the target feature corresponding to each node;

[0012] The second weight calculation module is configured to obtain a target weight of the variation degree of each node based on the information entropy of the target feature.

[0013] The data screening module is configured to screen a plurality of industrial production data based on the target weight of the variation degree of each node.

[0014] In a third aspect, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor implements the steps of the above data screening method when executing the computer program.

[0015] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program implements the steps of the above data screening method when executed by a processor.

[0016] The data screening method and device, the equipment and the storage medium provided by the embodiments of the present application have the beneficial effects that: the present application can fully mine the correlation of data in time and space by preprocessing a plurality of industrial production data and constructing a space-time graph, calculating the node space-time dimension attention weight. The attention weight is spliced into a target feature to comprehensively describe the role of the node in the space-time graph. The target weight is determined based on the information entropy of the target feature, and the variation degree of the node data and the influence on the production process are reasonably reflected. The present application can preferentially focus on the information with large variation degree and key to production, such as equipment abnormality and efficiency mutation reason, to provide strong support for industrial production optimization and fault diagnosis, and realize efficient utilization and accurate screening of data. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0018] Figure 1 The flowchart of the data screening method provided by an embodiment of the present application is shown in the figure.

[0019] Figure 2 The structural block diagram of the data screening device provided by an embodiment of the present application is shown in the figure.

[0020] Figure 3 The schematic block diagram of the electronic device provided by an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0021] In the following description, for purposes of explanation and not limitation, specific details are set forth such as particular architectures, technologies, techniques, etc. in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present application with unnecessary detail.

[0022] In order to make the objects, technical solutions and advantages of the present application clearer, the following will be described by specific embodiments in conjunction with the accompanying drawings.

[0023] Reference will be made to Figure 1 , Figure 1 The flowchart of the data screening method provided by an embodiment of the present application is shown in the figure, and the method comprises the following steps.

[0024] S101: calculating the attention weight of each node in the space-time graph in the space-time dimension, the space-time graph being constructed by various industrial production data, each kind of industrial production data being a node in the space-time graph, and the space-time dimension comprising a time dimension and a space dimension.

[0025] In the embodiment, the data screening method can be applied to chemical production. In the chemical production process, a large amount of complex and diverse data will be generated, including but not limited to chemical equipment operation data (such as temperature, pressure, flow rate, etc. numerical data, and image data reflecting the equipment condition), production efficiency data (such as product output per unit time), and product quality data (such as product purity, component ratio).

[0026] Before calculating the attention weight of each node in the space-time graph in the space-time dimension, the following steps are further included.

[0027] The numerical data is normalized to obtain numerical data features, the image data is feature-extracted to obtain image data features, and the numerical data features and the image data features are fused to obtain equipment operation data.

[0028] For example, the numerical data such as equipment temperature and pressure is normalized to the interval [0, 1], and for image data, after feature extraction by a VGG16 model, the image data feature vector is spliced and fused with the normalized numerical data to form a more rich and comprehensive equipment operation state feature representation.

[0029] Considering the imbalance or sample deficiency problem of the original production efficiency data and the original product quality data, the original production efficiency data and the original product quality data are processed based on a generative adversarial network to obtain production efficiency data and product quality data.

[0030] For example, a generative adversarial network composed of a generator G and a discriminator D is constructed,

[0031] The generator G generates data similar to the original equipment production efficiency data and the original product quality data through a neural network with random noise z as input. The discriminator D is used to determine whether the input data is real data or generated data. Through adversarial training, the generator generates high-quality supplementary data.

[0032] The supplementary data is quantized to obtain production efficiency data and product quality data.

[0033] For example, for production efficiency data x, it is mapped into a set of limited quantum states. Assuming that the production efficiency data range is [a, b], it is divided into several intervals, and each interval corresponds to a quantum state.

[0034] Calculate which interval x falls into, and represent it as the quantum state corresponding to the interval. In this way, the data volume can be greatly reduced without losing too much information.

[0035] In this embodiment, various industrial production data is constructed into a space-time graph.

[0036] Chemical production has the characteristics of strong continuity and complex process. Production data at different time points is crucial for overall production condition judgment, and each production link is closely related, and data changes in one link may affect multiple other links. The space-time graph can effectively integrate these different types of data in the chemical field, use them as nodes, and use the data relationship between different time points and different chemical equipment and production links as edges to clearly present the association of data in the space-time dimension.

[0037] For each node in the space-time graph, this embodiment calculates its attention weight from the time dimension and the space dimension respectively. For example, in the time dimension, this embodiment determines the importance of each node at different time points in the time sequence according to factors such as the change of equipment operation data at different time points, the fluctuation of production efficiency data over time, and the performance of product quality data at different time points. In the spatial dimension, this embodiment can determine the close degree of the association of the node with other nodes in space by analyzing the association between different equipment and the influence between different production links, so as to obtain the attention weight of each node in the space-time dimension, so as to reflect the importance features of each node in the space-time dimension.

[0038] S102: Splice the attention weight of each node in the space-time graph in the space-time dimension to obtain the target feature corresponding to each node.

[0039] In the embodiment, the attention weight of each node in the time dimension and the space dimension is added to obtain a target feature corresponding to each node. The target feature fuses the information of the node in the time and space dimensions, and can more comprehensively describe the position and role of the node data in the whole space-time graph, and can contain the change of the node data over time, and can also reflect the associated characteristics of the node data with other nodes in the space.

[0040] In the embodiment, the information entropy is used to represent the uncertainty or disorder degree of data. The higher the information entropy of the target feature is, the greater the variation degree of the node data corresponding to the target feature is, the richer the information contained is or the more unstable the information is, and the greater the influence on the industrial production process is. Therefore, a higher target weight can be given. Conversely, the data of the node with lower information entropy is relatively stable, and the variation degree is small. Therefore, the target weight is relatively low.

[0041] In the embodiment, the information entropy is used to represent the uncertainty or disorder degree of data. The higher the information entropy of the target feature is, the greater the variation degree of the node data corresponding to the target feature is, the richer the information contained is or the more unstable the information is, and the greater the influence on the industrial production process is. Therefore, a higher target weight can be given. Conversely, the data of the node with lower information entropy is relatively stable, and the variation degree is small. Therefore, the target weight is relatively low.

[0042] In the embodiment, the information entropy is used to represent the uncertainty or disorder degree of data. The higher the information entropy of the target feature is, the greater the variation degree of the node data corresponding to the target feature is, the richer the information contained is or the more unstable the information is, and the greater the influence on the industrial production process is. Therefore, a higher target weight can be given. Conversely, the data of the node with lower information entropy is relatively stable, and the variation degree is small. Therefore, the target weight is relatively low.

[0043] In the embodiment, the information entropy is used to represent the uncertainty or disorder degree of data. The higher the information entropy of the target feature is, the greater the variation degree of the node data corresponding to the target feature is, the richer the information contained is or the more unstable the information is, and the greater the influence on the industrial production process is. Therefore, a higher target weight can be given. Conversely, the data of the node with lower information entropy is relatively stable, and the variation degree is small. Therefore, the target weight is relatively low.

[0044] As can be seen from the above, in the embodiment, the multiple industrial production data are preprocessed and a space-time graph is constructed to calculate the node space-time dimension attention weight, so that the correlation of the data in the time and space dimensions can be fully mined. Then, the attention weight is spliced into a target feature to comprehensively describe the role of the node in the space-time graph. The target weight is determined based on the information entropy of the target feature, and the variation degree of the node data and the influence on the production process are reasonably reflected. The data is screened according to the weight, the information with great variation degree and key to production can be preferentially focused on, such as the reasons for equipment abnormality and efficiency mutation, which can provide strong support for industrial production optimization and fault diagnosis, and realize efficient utilization and accurate screening of data.

[0045] In an embodiment of the present application, the multiple industrial production data include equipment operation data, production efficiency data and product quality data.

[0046] The data screening method further includes:

[0047] The device running data, the production efficiency data and the product quality data are taken as nodes respectively;

[0048] The feature relationship of different time points and different target devices is taken as an edge;

[0049] The space-time graph is constructed based on the nodes and the edges.

[0050] In the embodiment, because the industrial production data is an important data type in industrial production, various industrial production data reflects the state of the industrial production process from different angles. Different industrial production data is taken as a node, and the complex industrial production data can be abstracted as a basic element in the graph structure.

[0051] For example, the device running data can reflect the real-time state of the device, the production efficiency data reflects the output of the production process, and the product quality data relates to whether the final product meets the standard, and each node represents a type of key information.

[0052] In industrial production, industrial production data not only has a correlation at the same time, but also has certain change rules and causal relationships at different time points. At the same time, different target devices also influence each other, and there is a relationship between the data features of different target devices. In the embodiment, the feature relationship is taken as an edge to connect different nodes to form a complete network structure, thereby reflecting the correlation of industrial production data in the space-time dimension.

[0053] For example, the running data of a device at different time points can affect production efficiency and product quality, and the coordinated work between different devices can also have an impact on the overall production.

[0054] The space-time graph is a graph structure that can reflect the spatial relationship and temporal change of data. In the embodiment, the device running data, the production efficiency data and the product quality data are taken as nodes, and the feature relationship of different time points and different target devices is taken as an edge to connect, so that a space-time graph reflecting the complex relationship between industrial production data can be obtained. The space-time graph helps to mine important features of data in the space-time dimension, and then realizes more accurate data screening.

[0055] From the above, it can be concluded that the embodiment can clearly present the correlation of each data in the space-time dimension by integrating key production data, which facilitates accurate positioning of equipment failure, production efficiency bottleneck and quality problem root cause, and improves the overall efficiency of industrial production.

[0056] In an embodiment of the present application, the attention weight of each node in the space-time dimension includes a time attention weight and a space attention weight;

[0057] The attention weight of each node in the space-time graph in the space-time dimension is spliced to obtain the target feature corresponding to each node, including:

[0058] The time attention weight is weighted aggregated to obtain a time feature vector;

[0059] The space attention weight is weighted aggregated to obtain a space feature vector;

[0060] The time feature vector and the space feature vector are spliced to obtain the target feature corresponding to each node.

[0061] In the embodiment, the attention weight of each node in the space-time dimension includes the time attention weight and the space attention weight. The time attention weight is used to represent the importance of the node at different time points in the time sequence, such as the importance of the running state change of the device at different time points; and the space attention weight is used to represent the closeness of the node to other nodes in the space, such as the importance of the association between the feature nodes of different devices at the same time.

[0062] In the embodiment, based on the defined space-time graph G=(V, E), the attention weight of each node is represented by a feature vector .

[0063] The time attention weight is weighted aggregated to obtain a time feature vector, including:

[0064] The feature vectors of the target time points are spliced to obtain a time joint vector;

[0065] The time joint vector is input into a multi-layer perception machine to obtain a time attention score;

[0066] The time attention score is normalized to obtain the time attention weight;

[0067] Based on the time attention weight, the time feature vectors of the time-dimension adjacent nodes of the target device at the target time are weighted aggregated to obtain a time feature vector.

[0068] For example, the node of the device A at time t is taken as an example, the embodiment first collects the feature vectors , of the device A at time t-1 and t+1, then splices the two feature vectors into a joint vector , and inputs the joint vector into a multi-layer perception machine. The multi-layer perception machine calculates the joint vector through a weight matrix W1, W2, a bias vector b1, b2 and an activation function σ 1, σ 2, and outputs a time attention score , which measures the importance of the state change of device A at time t compared to t-1 and t+1. The time attention weights are obtained by applying a softmax function to normalize the time attention scores of all devices at each time point .

[0069] The feature vectors of the time-dimension adjacent nodes of device A at time t are weighted and aggregated based on the time attention weights to obtain a time feature vector and . .

[0070]

[0071] The time feature vector integrates the state change information of the node over the time sequence, highlighting the influence of different time points on the current node state.

[0072] The spatial feature vectors are obtained by weighted aggregation of the spatial attention weights, including:

[0073] The spatial joint vectors are obtained by splicing the feature vectors of other target device nodes based on the same target time point;

[0074] A plurality of spatial attention scores are obtained by inputting the spatial joint vectors into multiple perceptrons with the same structure but different parameters, respectively;

[0075] The spatial attention weights are obtained by normalizing the spatial attention scores of all target devices at each time point with other devices;

[0076] The spatial feature vectors are obtained by weighted aggregation of the spatial feature vectors of the spatial-dimension adjacent nodes of the target device at the target time point based on the spatial attention weights.

[0077] For example, for the node of device A at time t , this embodiment first obtains the feature vectors of other device nodes at the same time t , then splices and each into joint vectors , and finally inputs the obtained joint vectors into multiple perceptrons with the same structure but different parameters, respectively, to output spatial attention scores , reflecting the correlation importance between the feature nodes of device A at time t and other devices j at the same time. The spatial attention weights are obtained by applying a softmax function to normalize the spatial attention scores of all devices at each time point with other devices .

[0078] The spatial feature vector is obtained by weighted aggregation of eigenvectors of the spatial dimension adjacent nodes of device A at time t using spatial attention weights . .

[0079]

[0080] The spatial feature vector reflects the closeness of the node to other device nodes at the same time and contains mutual influence information between nodes in the spatial dimension.

[0081] The time feature vector and the spatial feature vector are added to obtain a target feature corresponding to each node.

[0082] From the above, it can be concluded that the target feature in the embodiment fuses information of nodes in time and space dimensions, reflects dynamic changes of the node with time, embodies the association of the node with other nodes in space, and comprehensively and deeply describes the characteristics of the node in the industrial production space-time system.

[0083] In an embodiment of the present application, a target weight of a variation degree of each node is obtained based on information entropy of the target feature, including:

[0084] The target weight of the variation degree of each node is obtained based on the first calculation formula;

[0085] The first calculation formula is:

[0086]

[0087] wherein, the target weight of the variation degree of the i-th node is represented by m, the total number of nodes is represented by m, and the information entropy corresponding to the target feature is represented by H.

[0088] In the embodiment, the information entropy H is a key indicator for measuring data uncertainty or variation degree. In the industrial data scenario, each node represents different industrial production feature data such as device running state, production efficiency, product quality, etc. After multi-modal feature fusion preprocessing and prediction model construction, these data are presented in the form of target features. The information entropy H corresponding to each target feature is calculated by the second calculation formula.

[0089]

[0090] wherein, , , ​​​​​is the value of the ith sample under the jth feature, which is used to measure the uncertainty of the data distribution; n represents the number of samples.

[0091] represents a balance parameter, which is used to adjust the influence degree of the added part on the calculation result of information entropy, The optimal value can be determined through experiments or cross-validation.

[0092] represents different time points in the time dimension, and S represents a set of other nodes that have a direct association with the current node in the space dimension (for example, in industrial production, different device nodes on the same production line or different component nodes on the same device, etc.).

[0093] represents a spatiotemporal correlation weight matrix, reflecting the influence weight of spatial node s on the current node at time t. It can be determined based on factors such as physical connection relationship between devices, upstream and downstream relationship in production process, etc. For example, on a production line, the influence weight of the device next to the current device is larger, while the influence weight of the device far away from the current device is smaller.

[0094] represents the difference in the influence of the data change of spatial node s on the ith sample data of the current node j at time t. For example, if the running state of device A changes, it will cause the production efficiency data of device B to change.

[0095] The information entropy in this embodiment can better reflect the spatiotemporal correlation and dynamic change of industrial production data by introducing the related parameters of time dimension T and space dimension S. The traditional information entropy formula only considers the distribution of data itself and cannot reflect the mutual influence of data at different time points and different devices in industrial production. For example, in industrial production, the running state of a device at a certain time point will be affected by the running state of other devices at the previous time point. The improved formula can capture this dynamic correlation in time; at the same time, the layout and interaction of different devices in space can affect the production process, and the space dimension part of the formula can reflect this spatial correlation.

[0096] In this embodiment, the first calculation formula is used to determine the target weight of the variation degree of each node. The numerator of the first calculation formula is a transformation of information entropy. Since the larger the information entropy is, the greater the data variation degree is, then the data variation degree is inversely proportional to , that is, the smaller the data variation degree is. Through this transformation, information entropy is connected with the target weight.

[0097] The denominator of the first calculation formula is the sum of all nodes , the effect is to normalize the molecule. The final calculation of is based on the relative size of all node entropy to determine the target weight reflecting the degree of variation of each node.

[0098] From the above, the target weight of the embodiment can objectively reflect the relative importance of each node in the entire data set. For the node with large variation degree (i.e. high information entropy), the target weight is relatively high, which means that the industrial production feature data represented by the node may contain more critical and valuable information, and should be given more attention in the subsequent data screening process; for the node with small variation degree (low information entropy), the target weight is relatively low, and the priority in data screening is relatively low.

[0099] In an embodiment of the application, the target weight based on the degree of variation of each node is used to screen a plurality of industrial production data, including:

[0100] determine the initial screening threshold corresponding to each node based on the historical data of the plurality of industrial production data;

[0101] obtain a comprehensive score corresponding to each industrial production data based on the initial screening threshold corresponding to each node and the target weight of the degree of variation;

[0102] screen the corresponding industrial production data based on the comprehensive score.

[0103] In this embodiment, the mean value of the historical production data features can be calculated. Then, combined with business experience, the initial screening threshold is obtained by adjusting the key feature statistics.

[0104] For example, for the product quality, an industrial production data node, if its historical mean value is , the adjustment amount is set according to the basic requirements and experience of the business for product quality, so as to obtain the initial screening threshold .

[0105] After obtaining the initial screening threshold corresponding to each node and the target weight of the variation degree of each node determined based on the entropy weight method, the embodiment calculates the comprehensive score corresponding to each kind of industrial production data. The comprehensive score calculation considers the initial screening threshold of the data and the target weight of the variation degree thereof. The target weight of the variation degree reflects the relative importance of the node data in the entire data set, and the initial screening threshold provides a reference standard related to historical data and business experience. The initial screening threshold and the target weight can be weighted and calculated to obtain the comprehensive score corresponding to each kind of industrial production data. The comprehensive score can more comprehensively reflect the value and priority of the industrial production data in the screening process, considering both the variation characteristics of the data itself (reflected by the target weight) and the preliminary screening standard based on history and business (reflected by the initial screening threshold).

[0106] Finally, the embodiment screens the corresponding industrial production data based on the obtained comprehensive score. The industrial production data with a higher comprehensive score means that it performs outstandingly in terms of data variation degree and screening standard based on history and business, and therefore contains information valuable for enterprise production decision-making, so it can be preferentially retained or focused on in the screening process; and the data with a lower comprehensive score can be assigned a lower priority or screened out in subsequent analysis. Thus, the screening of various industrial production data is realized, so that the screened data is more in line with the actual production needs of the enterprise, providing strong support for enterprise production decision-making.

[0107] In an embodiment of the present application, it further comprises:

[0108] determining a reward function based on production efficiency and defective rate;

[0109] adjusting the initial screening threshold based on the reward function.

[0110] In the embodiment, the proximal policy optimization algorithm can be used to dynamically adjust the initial screening threshold.

[0111] In the framework of the proximal policy optimization algorithm, the state of the embodiment not only contains the feature values of the current data sample after processing and the related statistical information of the screened data, such as the average comprehensive score and the contribution of different features , but also introduces the correlation strength information between the production indicators and the features.

[0112] For example, the embodiment calculates the Pearson correlation coefficient between the running state feature of the computing device and the production efficiency and the negative correlation coefficient between the product quality feature and the defective rate , etc., and takes these correlation coefficients as part of the state. Therefore, the state is expressed as: By combining these information into the state , the progress of current data screening and the relationship between data characteristics and production indicators can be comprehensively understood, so that more reasonable decisions can be made.

[0113] The action space is set to adjust the feature weight and the screening threshold . In each iteration process, the agent makes a decision according to the current state , and selects an action , where is the adjustment amount of the feature weight, which is used to change the relative importance of each feature in data screening; is the adjustment amount of the screening threshold, which determines the limit of data screening.

[0114] In this embodiment, the production efficiency P and the defective rate D are the key indicators in the production process of the enterprise, which directly reflect the benefit and quality of production. When determining the reward function, these two indicators are included, which can make the reward function closely around the core production target of the enterprise. The reward function calculation formula is:

[0115]

[0116] wherein, represents the initial production efficiency, represents the initial defective rate, represents the change amount of production efficiency before and after screening data, represents the change amount of defective rate before and after screening data, represents the relative change amount of production efficiency before and after screening data, which reflects the contribution of screening operation to the improvement of production efficiency; represents the relative change amount of defective rate before and after screening data, which reflects the effect of screening operation on product quality improvement (defective rate reduction); is the sum of the change amount of the correlation coefficient of each feature and production indicator, which represents the change of the correlation strength between the feature and the production indicator in the screening process; , and represent the weight coefficients, which are used to balance the contribution of these three parts to the reward, and can highlight the influence degree of different aspects on the reward according to the actual needs of the enterprise.

[0117] For example, if the enterprise pays more attention to the improvement of production efficiency, the value of can be appropriately increased. The reward function can comprehensively consider the influence of screening data on production efficiency, defective rate and the correlation between features and production indicators, and can accurately evaluate the contribution degree of each screening operation to the decision target of the enterprise.

[0118] In the reinforcement learning process, the agent continuously tries different actions (i.e., adjusts the initial screening threshold) by constantly interacting with the environment, and optimizes the decision-making according to the feedback of the reward function. When a positive reward is obtained, it indicates that the current threshold adjustment direction is helpful for screening data that is more valuable to the enterprise decision-making goal, and thus the threshold is further fine-tuned in this direction; if a negative reward is obtained, it indicates that the current adjustment direction is not conducive to screening data that meets the needs of the enterprise, and thus the threshold needs to be adjusted in the opposite direction. The specific adjustment method is that, assuming that the current screening threshold is , the adjustment amount is , and the new screening threshold is expressed as:

[0119]

[0120] wherein, as a learning rate, is used to accurately control the step size of threshold adjustment, preventing the adjustment amplitude from being too large and missing the optimal solution; is a sign function of the reward function value, which determines the direction of threshold adjustment according to the positive and negative of the reward.

[0121] From the above, it can be seen that the embodiment integrates the correlation strength of the features and the production indicators into the state representation and the reward function design, and can optimize the screening strategy in real time according to the dynamic mutual relationship of various factors in the production process, to ensure that the screened data always closely matches the production decision-making goal of the enterprise.

[0122] In an embodiment of the present application, the corresponding industrial production data is screened based on the comprehensive score, comprising:

[0123] screening the industrial production data that meets the first condition;

[0124] The first condition is that the comprehensive score corresponding to the industrial production data is greater than the comprehensive score threshold.

[0125] In this embodiment, a comprehensive score threshold can be pre-set, which is the baseline for screening industrial production data. The setting of the comprehensive score threshold can be based on the expectations of the enterprise on the value of the data, statistical analysis of historical production data, or determined by business experts according to experience. The comprehensive score threshold divides all industrial production data into two parts, and the data greater than the threshold is considered to contain more valuable information and meet the goal of the enterprise screening data.

[0126] The initial screening threshold corresponding to each node is determined based on the historical data of multiple industrial production data, and the target weight of the variation degree of each node is combined to obtain the comprehensive score corresponding to each kind of industrial production data. The comprehensive score considers the variation degree of the data itself and the quantitative indicators of the initial screening standard based on history and business experience, and can reflect the value and priority of the industrial production data in the screening process.

[0127] The comprehensive score of each industrial production data is compared with a set comprehensive score threshold, and only the industrial production data that meets the first condition of the comprehensive score being greater than the comprehensive score threshold can be screened out. The screening process of the embodiment focuses on data with higher comprehensive scores, because these data are more outstanding in terms of data variation degree and screening criteria based on history and business, and are more likely to contain information valuable to enterprise production decision-making. Thus, the screening of industrial production data is realized, so that the screened data is more in line with the actual production needs of the enterprise.

[0128] For example, if the enterprise finds in the production process that data with a comprehensive score higher than a certain threshold has a strong correlation with production efficiency improvement or product quality improvement, the threshold can be used as a screening criterion, so as to extract data that helps the enterprise achieve these goals from a large amount of industrial production data.

[0129] From the above, it can be concluded that the embodiment uses the comprehensive score being greater than the comprehensive score threshold as the screening condition, and can accurately focus on data with higher value. By determining the comprehensive score by combining the data variation degree and historical business experience, industrial production data with important value to enterprise production decision-making can be effectively screened out.

[0130] According to the data screening method of the above embodiment, Figure 2 A structural block diagram of a data screening device provided by an embodiment of the present application is shown. For ease of illustration, only parts related to the embodiment of the present application are shown. Reference is made to the above data screening method. Figure 2 The data screening device 20 includes:

[0131] The first weight calculation module 21 is configured to calculate the attention weight of each node in the spatio-temporal graph in the spatio-temporal dimension, the spatio-temporal graph being constructed by a plurality of industrial production data, each industrial production data being a node in the spatio-temporal graph, and the spatio-temporal dimension including a time dimension and a space dimension;

[0132] The weight splicing module 22 is configured to splice the attention weight of each node in the spatio-temporal graph in the spatio-temporal dimension to obtain a target feature corresponding to each node;

[0133] The second weight calculation module 23 is configured to obtain a target weight of the variation degree of each node based on the information entropy of the target feature;

[0134] The data screening module 24 is configured to screen the plurality of industrial production data based on the target weight of the variation degree of each node.

[0135] In an embodiment of the present application, the plurality of industrial production data includes equipment operation data, production efficiency data, and product quality data;

[0136] The data screening device 20 further comprises a space-time graph construction module, specifically configured to:

[0137] taking the equipment operation data, the production efficiency data and the product quality data as nodes respectively;

[0138] taking the feature relationship of different time points and different target equipment as edges;

[0139] constructing a space-time graph based on the nodes and the edges.

[0140] In an embodiment of the present application, the attention weight of each node in the space-time dimension comprises a time attention weight and a space attention weight;

[0141] The weight splicing module 22 is specifically configured to:

[0142] performing weighted aggregation on the time attention weight to obtain a time feature vector;

[0143] performing weighted aggregation on the space attention weight to obtain a space feature vector;

[0144] splicing the time feature vector and the space feature vector to obtain a target feature corresponding to each node.

[0145] In an embodiment of the present application, the second weight calculation module 23 is specifically configured to:

[0146] obtaining a target weight of the variation degree of each node based on a first calculation formula;

[0147] The first calculation formula is:

[0148]

[0149] wherein, denotes the target weight of the variation degree of the i-th node, m denotes the total number of nodes, the target feature corresponding to the information entropy. In an embodiment of the present application, the data screening module 24 is specifically configured to determine an initial screening threshold corresponding to each node based on historical data of a plurality of industrial production data;

[0150] obtaining a comprehensive score corresponding to each industrial production data based on the initial screening threshold corresponding to each node and the target weight of the variation degree;

[0151] screening the corresponding industrial production data based on the comprehensive score.

[0152]

[0153] In an embodiment of the present application, the data screening module 24 is specifically further configured to: ​​

[0154] The reward function is determined based on production efficiency and defect rate;

[0155] The initial screening threshold is adjusted based on the reward function.

[0156] In one embodiment of the present invention, the data filtering module 24 is further configured to:

[0157] The industrial production data that meets the first condition are screened;

[0158] The first condition is that the comprehensive score corresponding to the industrial production data is greater than the comprehensive score threshold.

[0159] See Figure 3 , Figure 3 This is a schematic block diagram of an electronic device provided according to an embodiment of the present invention. Figure 3 The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of each module / unit in the above-described device embodiments, for example... Figure 2 The functions of the first weight calculation module 21, the weight splicing module 22, the second weight calculation module 23, and the data filtering module 24 are shown.

[0160] It should be understood that, in this embodiment of the invention, the processor 301 may be a Central Processing Unit (CPU), but it may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0161] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.

[0162] The memory 304 can include read-only memory and random access memory, and provide instructions and data to the processor 301. A portion of the memory 304 can also include non-volatile random access memory. For example, the memory 304 can also store device type information.

[0163] In specific implementations, the processor 301, the input device 302, and the output device 303 described in the embodiments of the present application can perform the implementation manners described in the first embodiment and the second embodiment of the data filtering method provided by the embodiments of the present application, and can also perform the implementation manners of the electronic device described in the embodiments of the present application, which will not be described here.

[0164] In another embodiment of the present application, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program. The computer program includes program instructions, and the program instructions are executed by a processor to implement all or part of the processes of the above-mentioned embodiments. The computer program can also be used to instruct related hardware to complete, and the computer program can be stored in a computer readable storage medium. When the computer program is executed by the processor, the steps of each method embodiment described above can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0165] The computer readable storage medium can be an internal storage unit of the electronic device of any of the above-mentioned embodiments, such as a hard disk or a memory of the electronic device. The computer readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the computer readable storage medium can include both the internal storage unit and the external storage device of the electronic device. The computer readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer readable storage medium can also be used to temporarily store data that has been output or will be output.

[0166] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the electronic device and the units described above can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0167] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the electronic device and the units described above can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0168] In several embodiments provided in the present application, it should be understood that the disclosed electronic device and method can be implemented in other ways. For example, the above-described device embodiments are only schematic, for example, the division of units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface or unit, and can also be electrical, mechanical or other forms of connection.

[0169] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment of the present application.

[0170] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0171] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited to this. Any skilled person in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A data filtering method, characterized in that, include: Calculate the attention weight of each node in the spatiotemporal dimension of the spatiotemporal graph, which is constructed from various industrial production data, with each type of industrial production data serving as a node in the spatiotemporal graph. The spatiotemporal dimension includes a time dimension and a spatial dimension. The various industrial production data include equipment operation data, production efficiency data, and product quality data; the process of constructing the spatiotemporal map includes: Use equipment operation data, production efficiency data, and product quality data as nodes respectively; The characteristic relationships between different time points and different target devices are used as edges; Construct a spatiotemporal graph based on the nodes and edges; The attention weights of each node in the spatiotemporal dimension of the spatiotemporal graph are concatenated to obtain the target features corresponding to each node. The target weight of the degree of variation of each node is obtained based on the information entropy of the target features; The target weight of the degree of variation of each node is obtained based on the first calculation formula; The first calculation formula is: in, Indicates the first The target weights for the degree of variation of each node, where m represents the total number of nodes. Indicates the first Information entropy corresponding to the target feature; The information entropy corresponding to each target feature is calculated using the second calculation formula, which is: in, n represents the number of samples. , It is the value of the i-th sample under the j-th feature. Represents the balance parameters. Let S represent different moments in time, and let S represent the set of other nodes that are directly related to the current node in space. This represents the spatiotemporal correlation weight matrix, reflecting the influence weight of spatial node s on the current node at time t. This indicates that at time t, the data change at spatial node s affects the i-th sample data at current node j. The difference in impact; Multiple types of industrial production data are filtered based on the target weight of the degree of variation of each node.

2. The data filtering method as described in claim 1, characterized in that, The attention weight of each node in the spatiotemporal dimension includes: temporal attention weight and spatial attention weight; The process of concatenating the attention weights of each node in the spatiotemporal dimension in the spatiotemporal graph to obtain the target features corresponding to each node includes: The time attention weights are weighted and aggregated to obtain a time feature vector; The spatial attention weights are weighted and aggregated to obtain a spatial feature vector; The temporal feature vector and the spatial feature vector are concatenated to obtain the target feature corresponding to each node.

3. The data filtering method as described in claim 1, characterized in that, The target weights based on the degree of variation of each node are used to filter various industrial production data, including: The initial screening threshold for each node is determined based on historical data from various industrial production data. Based on the initial screening threshold and the target weight of the degree of variation corresponding to each node, a comprehensive score is obtained for each type of industrial production data. The corresponding industrial production data are filtered based on the comprehensive score.

4. The data filtering method as described in claim 3, characterized in that, Also includes: The reward function is determined based on production efficiency and defect rate; The initial screening threshold is adjusted based on the reward function.

5. The data filtering method as described in claim 3, characterized in that, The process of filtering the corresponding industrial production data based on the comprehensive score includes: The industrial production data that meets the first condition are screened; The first condition is that the comprehensive score corresponding to the industrial production data is greater than the comprehensive score threshold.

6. A data filtering device, characterized in that, include: The first weight calculation module is used to calculate the attention weight of each node in the spatiotemporal dimension of the spatiotemporal graph. The spatiotemporal graph is constructed from a variety of industrial production data, with each type of industrial production data serving as a node in the spatiotemporal graph. The spatiotemporal dimension includes a time dimension and a spatial dimension. The various industrial production data include equipment operation data, production efficiency data, and product quality data; The spatiotemporal graph construction module is used to: use equipment operation data, production efficiency data, and product quality data as nodes respectively; The characteristic relationships between different time points and different target devices are used as edges; A spatiotemporal graph is constructed based on the nodes and edges; a weight splicing module is used to splice the attention weights of each node in the spatiotemporal dimension to obtain the target features corresponding to each node; The second weight calculation module is used to obtain the target weight of the degree of variation of each node based on the information entropy of the target feature; The second weight calculation module is specifically used for: The target weight of the degree of variation of each node is obtained based on the first calculation formula; The first calculation formula is: in, Indicates the first The target weights for the degree of variation of each node, where m represents the total number of nodes. Indicates the first Information entropy corresponding to the target feature; The information entropy corresponding to each target feature is calculated using the second calculation formula, which is: in, n represents the number of samples. , It is the value of the i-th sample under the j-th feature. Represents the balance parameters. Let S represent different moments in time, and let S represent the set of other nodes that are directly related to the current node in space. This represents the spatiotemporal correlation weight matrix, reflecting the influence weight of spatial node s on the current node at time t. This indicates that at time t, the data change at spatial node s affects the i-th sample data at current node j. The difference in impact; The data filtering module is used to filter various industrial production data based on the target weight of the degree of variation of each node.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-dimensional time series data spatio-temporal feature extraction method based on attention mechanism

    CN118606682A

  • Method of malicious social activity prediction using spatial-temporal social network data

    US11195107B1