Data screening method and device, equipment and storage medium

By constructing a spatio-temporal map and calculating attention weights, and screening chemical production data with information entropy, the problem of inaccurate screening of chemical production data is solved, efficient use and accurate screening of data is achieved, and production optimization and fault diagnosis are supported.

CN120371894AActive Publication Date: 2025-07-25TANGSHAN CAOFEIDIAN LIANCHENG TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510501131.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-25
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

It is difficult for the existing technology to efficiently utilize chemical production data, especially the dynamic evolution laws in the time dimension and the internal connection in the spatial dimension, resulting in inaccurate data screening and unable to meet the optimization and fault diagnosis needs of industrial production.

Method used

Construct a spatiotemporal graph to calculate the attention weight of nodes in the spatiotemporal dimension, and realize accurate screening of industrial production data by splicing target features and using information entropy to determine the target weight.

Benefits of technology

By mining the correlation between data in time and space, we give priority to information that is highly varied and critical to production, providing strong support for industrial production optimization and fault diagnosis, and achieving efficient use and accurate screening of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371894A_ABST
    Figure CN120371894A_ABST
Patent Text Reader

Abstract

The invention provides a data screening method and device, equipment and a storage medium, and belongs to the technical field of data screening, and the method comprises the steps: calculating the attention weight of each node in a space-time diagram in a space-time dimension, the space-time diagram is constructed by multiple kinds of industrial production data, each kind of industrial production data is used as one node in the space-time diagram, and the attention weight of each node in the space-time diagram is calculated; the space-time dimension comprises a time dimension and a space dimension; splicing the attention weight of each node in the space-time diagram in the space-time dimension to obtain a target feature corresponding to each node; obtaining the target weight of each node variation degree based on the information entropy of the target features; and screening the various industrial production data based on the target weight of the variation degree of each node. According to the data screening method and device, the equipment and the storage medium provided by the invention, efficient utilization and accurate screening of the data can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data screening, and more specifically, relates to a data screening method and device, equipment, and storage medium. Background Art

[0002] In the chemical production industry, with the expansion of production scale and the increase of process complexity, the amount of data has exploded, covering equipment operation, production efficiency, product quality and other aspects of data. However, most current data screening methods have limitations. Traditional methods rely on simple statistical analysis or preset rules to screen data, which makes it difficult to explore the dynamic evolution of chemical production data in the time dimension, such as the gradual trend of key parameters in the long-term operation of equipment, and the inherent connection between data in different production links in the spatial dimension, making it impossible to achieve efficient use and accurate screening of chemical production data. Summary of the invention

[0003] The purpose of the present invention is to provide a data screening method and device, equipment, and storage medium to achieve efficient use and accurate screening of data.

[0004] A first aspect of an embodiment of the present invention provides a data screening method, comprising: Calculating the attention weight of each node in the space-time graph in the space-time dimension, wherein the space-time graph is constructed by a variety of industrial production data, each industrial production data is used as a node in the space-time graph, and the space-time dimension includes a time dimension and a space dimension; The attention weights of each node in the spatiotemporal dimension of the spatiotemporal graph are concatenated to obtain the target features corresponding to each node; Obtaining a target weight for the degree of variation of each node based on the information entropy of the target feature; A variety of industrial production data are screened based on the target weight of each node's variation degree.

[0005] A second aspect of an embodiment of the present invention provides a data screening device, comprising: A first weight calculation module is used to calculate the attention weight of each node in the space-time graph in the space-time dimension, wherein the space-time graph is constructed by a variety of industrial production data, each industrial production data is used as a node in the space-time graph, and the space-time dimension includes a time dimension and a space dimension; The weight splicing module is used to splice the attention weights of each node in the spatiotemporal dimension of the spatiotemporal graph to obtain the target features corresponding to each node; A second weight calculation module is used to obtain a target weight of each node variation degree based on the information entropy of the target feature; The data screening module is used to screen various industrial production data based on the target weight of each node's variation degree.

[0006] In the third aspect of the embodiments of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above data screening method are implemented.

[0007] In the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above data screening method are implemented.

[0008] The beneficial effects of the data screening method, device, equipment, and storage medium provided by the embodiments of the present invention are as follows: By preprocessing various industrial production data and constructing a spatio-temporal graph, calculating the spatio-temporal dimension attention weights of nodes, the present invention can fully explore the associations of data in time and space. The attention weights are concatenated into target features to comprehensively describe the roles of nodes in the spatio-temporal graph. Based on the information entropy of the target features, the target weights are determined to reasonably reflect the data variation degree of nodes and their impacts on the production process. The present invention screens data according to these weights, can give priority to focusing on information with large variation degrees and key to production, such as the reasons for equipment anomalies and efficiency mutations, provides strong support for industrial production optimization and fault diagnosis, and realizes the efficient utilization and accurate screening of data. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0010] Figure 1 It is a schematic flowchart of the data screening method provided by an embodiment of the present invention; Figure 2 It is a structural block diagram of the data screening device provided by an embodiment of the present invention; Figure 3 It is a schematic block diagram of the electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0011] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are set forth in order to provide a thorough understanding of the embodiments of the present invention. However, those skilled in the art should understand that the present invention can be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.

[0012] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will be described through specific embodiments with reference to the accompanying drawings.

[0013] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a data screening method provided by an embodiment of the present invention. The method includes: S101: Calculate the attention weights of each node in the spatio-temporal graph in the spatio-temporal dimension. The spatio-temporal graph is constructed from various industrial production data, and each industrial production data serves as a node in the spatio-temporal graph. The spatio-temporal dimension includes the time dimension and the space dimension.

[0014] In this embodiment, the data screening method can be applied to chemical production. During the chemical production process, a large amount of complex and diverse data will be generated, including but not limited to chemical equipment operation data (such as numerical data such as temperature, pressure, and flow rate, and image data reflecting the equipment status), production efficiency data (such as the product output per unit time), and product quality data (such as product purity, composition ratio).

[0015] Before calculating the attention weights of each node in the spatio-temporal graph in the spatio-temporal dimension, it further includes: Normalize the numerical data to obtain numerical data features; extract features from the image data to obtain image data features; fuse the numerical data features and the image data features to obtain equipment operation data.

[0016] For example, numerical data such as equipment temperature and pressure are first normalized to the interval [0, 1]. For image data, after extracting features through the VGG16 model, the image data feature vectors are concatenated and fused with the normalized numerical data to form a richer and more comprehensive representation of the equipment operation status.

[0017] Considering the problems of imbalance or insufficient samples in the original production efficiency data and the original product quality data, the original production efficiency data and the original product quality data are processed based on the generative adversarial network to obtain production efficiency data and product quality data.

[0018] For example, a generative adversarial network composed of a generator G and a discriminator D is constructed. The generator G takes a random noise z as input and generates data similar to the original equipment production efficiency data and the original product quality data through a neural network. The discriminator D is used to determine whether the input data is real data or generated data. Through adversarial training, the generator generates high-quality supplementary data. The supplementary data is quantized to obtain production efficiency data and product quality data.

[0019] For example, for the production efficiency data x, it is mapped into a finite set of quantum states. Suppose the range of production efficiency data is [a, b], which is divided into several intervals, and each interval corresponds to a quantum state.

[0020] Calculate which interval x falls into, and then quantize it to represent the quantum state corresponding to that interval. In this way, the amount of data can be significantly reduced without losing too much information.

[0021] In this embodiment, a spatio-temporal graph is constructed from various industrial production data.

[0022] Chemical production is characterized by strong continuity and complex processes. The production data at different time points is crucial for judging the overall production situation, and the data of each production link is closely related. The change of data in one link may affect multiple other links. The spatio-temporal graph can effectively integrate these different types of data in the chemical industry. Taking them as nodes and taking the data relationships between different time points and different chemical equipment and production links as edges, it clearly presents the association of data in the spatio-temporal dimension.

[0023] For each node in the spatio-temporal graph, in this embodiment, the attention weights are calculated separately from the time dimension and the space dimension. For example, in the time dimension, in this embodiment, according to factors such as the changes of equipment operation data at different times, the fluctuations of production efficiency data over time, and the performance of product quality data at different time points, etc., to determine the importance degree of each node at different moments in the time series. In the space dimension, in this embodiment, by analyzing the associations between different equipment and the influences between different production links, etc., to determine the closeness of the association between the node and other nodes in space, so as to obtain the attention weights of each node in the spatio-temporal dimension, in order to reflect the importance characteristics of each node in the spatio-temporal dimension.

[0024] S102: Concatenate the attention weights of each node in the spatio-temporal graph in the spatio-temporal dimension to obtain the target feature corresponding to each node.

[0025] In this embodiment, the attention weights of each node in the time dimension and the space dimension are added together to obtain the target feature corresponding to each node. The target feature integrates the information of the node in the two spatio-temporal dimensions, and can more comprehensively describe the status and role of the node data in the entire spatio-temporal graph. It can not only include the change situation of the node data over time, but also reflect its association characteristics with other nodes in space.

[0026] S103: Obtain the target weight of the variation degree of each node based on the information entropy of the target feature.

[0027] In this embodiment, information entropy is used to represent the uncertainty or degree of chaos of data. The higher the information entropy of the target feature, the greater the degree of variation of the node data corresponding to the target feature, the richer or more unstable the information it contains, and the greater the impact on the industrial production process. Therefore, a higher target weight can be assigned; conversely, for nodes with lower information entropy, their data is relatively more stable, with a smaller degree of variation, and the target weight is also relatively lower.

[0028] S104: Screen various industrial production data based on the target weights of the variation degree of each node.

[0029] In this embodiment, node data with a higher target weight, due to its greater degree of variation and more important role in the spatio-temporal graph, can be preferentially selected or focused on. These data contain information that has a key impact on the industrial production process, such as the abnormal state of equipment, the reasons for sudden changes in production efficiency, etc. While node data with a lower target weight can be given a lower priority in the analysis, thus realizing the effective screening of industrial production data and providing strong support for the optimization of industrial production, fault diagnosis, etc.

[0030] It can be concluded from the above that in this embodiment, first, by preprocessing various industrial production data and constructing a spatio-temporal graph, and calculating the spatio-temporal dimension attention weights of nodes, the association of data in time and space can be fully mined. Then, the attention weights are concatenated into target features to comprehensively describe the role of nodes in the spatio-temporal graph. Based on the information entropy of the target features, the target weights are determined to reasonably reflect the degree of variation of node data and its impact on the production process. According to this weight, this embodiment screens data, can preferentially focus on information with a large degree of variation and key to production, such as equipment anomalies and the reasons for sudden changes in efficiency, provides strong support for industrial production optimization and fault diagnosis, and realizes the efficient utilization and accurate screening of data.

[0031] In an embodiment of the present invention, various industrial production data include equipment operation data, production efficiency data, and product quality data; The data screening method further includes: Taking the equipment operation data, production efficiency data, and product quality data as nodes respectively; Taking the characteristic relationships at different time points and different target equipment as edges; Constructing a spatio-temporal graph based on the nodes and edges.

[0032] In this embodiment, since industrial production data is an important data type in industrial production, various industrial production data reflect the state of the industrial production process from different perspectives. Taking different industrial production data as nodes can abstract complex industrial production data into basic elements in the graph structure.

[0033] For example, equipment operation data can reflect the real-time status of the equipment, production efficiency data reflects the output of the production process, and product quality data is related to whether the final product meets the standards. Each node represents a type of key information.

[0034] In industrial production, industrial production data is not only related at the same time, but also has certain changing patterns and causal relationships at different time points. At the same time, different target devices will also affect each other, and there are connections between the data features of different target devices. This embodiment uses these feature relationships as edges to connect different nodes to form a complete network structure, thereby reflecting the correlation of industrial production data in the time and space dimensions.

[0035] For example, the operating data of a certain device at different time points can affect production efficiency and product quality, and the collaborative work between different devices can also have an impact on overall production.

[0036] A space-time graph is a graph structure that can simultaneously reflect the spatial relationship and temporal changes of data. This embodiment uses equipment operation data, production efficiency data, and product quality data as nodes, and connects the characteristic relationships of different time points and different target devices as edges to obtain a space-time graph that fully reflects the complex relationship between industrial production data. The space-time graph helps to mine the important features of data in the space-time dimension, thereby achieving more accurate data screening.

[0037] From the above, it can be concluded that this embodiment can clearly present the correlation between various data in the time and space dimensions by integrating key production data, facilitate the accurate positioning of equipment failures, production efficiency bottlenecks and root causes of quality problems, and improve the overall efficiency of industrial production.

[0038] In one embodiment of the present invention, the attention weight of each node in the spatiotemporal dimension includes: a temporal attention weight and a spatial attention weight; The attention weights of each node in the spatiotemporal graph in the spatiotemporal dimension are concatenated to obtain the target features corresponding to each node, including: Perform weighted aggregation on the temporal attention weights to obtain the temporal feature vector; Perform weighted aggregation on the spatial attention weights to obtain the spatial feature vector; The temporal feature vector and the spatial feature vector are concatenated to obtain the target feature corresponding to each node.

[0039] In this embodiment, the attention weights of each node in the spatio-temporal dimension include temporal attention weights and spatial attention weights. The temporal attention weights are used to represent the importance of nodes at different moments in the time series, such as the importance of the change in the operating state of a device at different time points; the spatial attention weights are used to represent the degree of closeness of the association between a node and other nodes in space, such as the importance of the association between feature nodes of different devices at the same time.

[0040] In this embodiment, based on the defined spatio-temporal graph G=(V, E), each node is represented by a feature vector .

[0041] Performing weighted aggregation on the temporal attention weights to obtain a temporal feature vector, including: Concatenating the feature vectors at the target time point to obtain a temporal joint vector; Inputting the temporal joint vector into a multi-layer perceptron to obtain a temporal attention score; Performing normalization processing on the temporal attention score to obtain the temporal attention weights; Based on the temporal attention weights, performing weighted aggregation on the temporal feature vectors of the temporal dimension adjacent nodes of the target device at the target time to obtain a temporal feature vector.

[0042] Exemplarily, taking the node of device A at time t as an example, this embodiment first collects its feature vectors at times t-1 and t+1 , , then concatenates the two feature vectors into a joint vector and inputs it into a multi-layer perceptron. The multi-layer perceptron calculates the joint vector through weight matrices W1, W2, bias vectors b1, b2, and activation functions σ 1, σ 2, and outputs a temporal attention score , which measures the importance of the change in the operating state of device A at time t compared to t-1 and t+1. Applying the softmax function to normalize the temporal attention scores of all devices at each time point to obtain the temporal attention weights .

[0043] Using the temporal attention weights to perform weighted aggregation on the feature vectors of the temporal dimension adjacent nodes and of device A at time t to obtain a temporal feature vector .

[0044]

[0045] The time feature vector integrates the state change information of the node in the time series, highlighting the influence degree of different time points on the current node state.

[0046] Weighted aggregation is performed on the spatial attention weights to obtain a spatial feature vector, including: Based on the feature vectors of other target device nodes at the same target time point, concatenation is performed to obtain a spatial joint vector; The spatial joint vector is respectively input into multi-layer perceptrons with the same structure but different parameters to obtain multiple spatial attention scores; Normalization processing is performed on the spatial attention scores of all target devices at each time point with other devices to obtain multiple spatial attention weights; Based on the spatial attention weights, weighted aggregation is performed on the spatial feature vectors of the spatial dimension adjacent nodes of the target device at the target time point to obtain a spatial feature vector.

[0047] Exemplarily, for the node of device A at time t , in this embodiment, first, the feature vectors of other device nodes at the same time t are obtained, and then is concatenated with each respectively to form a joint vector . Finally, the obtained joint vectors are respectively input into multi-layer perceptrons with the same structure but different parameters, and spatial attention scores are output, reflecting the association importance between the feature nodes of device A at time t and other devices j at the same time. After applying the softmax function for normalization to the spatial attention scores of all devices at each time point with other devices, spatial attention weights are obtained. .

[0048] Using the spatial attention weights, weighted aggregation is performed on the feature vectors of the spatial dimension adjacent nodes of device A at time t to obtain a spatial feature vector .

[0049]

[0050] The spatial feature vector reflects the association tightness degree between the node and other device nodes at the same time, and contains the mutual influence information among nodes in the spatial dimension.

[0051] Add the time feature vector and the spatial feature vector to obtain the target feature corresponding to each node.

[0052] It can be concluded from the above that the target feature in this embodiment integrates the information of the node in both the time and space dimensions, reflecting both the dynamic changes of the node over time and its association with other nodes in space, comprehensively and deeply depicting the characteristics of the node in the industrial production spatio-temporal system.

[0053] In one embodiment of the present invention, the target weight of the variation degree of each node is obtained based on the information entropy of the target feature, including: Obtaining the target weight of the variation degree of each node based on the first calculation formula; The first calculation formula is:

[0054] Wherein, represents the target weight of the variation degree of the th node, m represents the total number of nodes, the information entropy corresponding to the target feature.

[0055] In this embodiment, the information entropy is a key indicator for measuring the uncertainty or variation degree of data. In the industrial data scenario, each node represents different industrial production feature data such as equipment operation status, production efficiency, product quality, etc. After multi-modal feature fusion preprocessing and prediction model construction, these data are presented in the form of target features. Calculate the information entropy corresponding to each target feature through the second calculation formula .

[0056]

[0057] Wherein, , , is the value of the

[0058] th sample under the jth feature, used to measure the uncertainty of the data distribution; n represents the number of samples.

[0059] represents different moments in the time dimension, S represents the set of other nodes directly associated with the current node in the space dimension (for example, in industrial production, different equipment nodes on the same production line or different component nodes of the same equipment, etc.).

[0060] represents the spatio-temporal association weight matrix, reflecting the influence weight of the space node s on the current node at time t. It can be determined based on factors such as the physical connection relationship between devices and the upstream and downstream relationships in the production process. For example, on an assembly line, adjacent devices have a greater influence weight on the current device, while devices farther away have a smaller influence weight.

[0061] It represents the difference in the influence of the data change of the spatial node s at time t on the i-th sample data of the current node j. For example, if a change in the operating state of device A causes a change in the production efficiency data of device B. For example, if a change in the operating state of device A causes a change in the production efficiency data of device B.

[0062] In this embodiment, by introducing relevant parameters of the time dimension T and the space dimension S, the information entropy can better reflect the spatio-temporal correlation and dynamic changes of industrial production data. The traditional information entropy formula only considers the distribution of the data itself and cannot reflect the mutual influence of data at different time points and between different devices in industrial production. For example, in industrial production, the operating state of a device at a certain moment is affected by the operating conditions of other devices at the previous moment, and the improved formula can capture this dynamic correlation in time; at the same time, the layout and interaction of different devices in space can affect the production process, and the space dimension part in the formula can reflect this spatial correlation.

[0063] In this embodiment, the first calculation formula is used to determine the target weight of the variation degree of each node. The numerator of the first calculation formula is a transformation of the information entropy. Since the larger the information entropy, the greater the data variation degree, is inversely related to the data variation degree, that is the smaller it is, the greater the data variation degree. Through this transformation, a connection is established between the information entropy and the target weight.

[0064] The denominator of the first calculation formula is to sum up the of all nodes, and its function is to normalize the numerator. The finally calculated is the target weight that reflects the variation degree of each node and is determined based on the relative magnitudes of the information entropy of all nodes.

[0065] It can be concluded from the above that the target weight of this embodiment can objectively reflect the relative importance of each node in the entire dataset. For nodes with a large variation degree (i.e., high information entropy), the target weight is relatively high, which means that in the subsequent data screening process, the industrial production characteristic data represented by this node may contain more critical and valuable information and should be given higher attention; while for nodes with a small variation degree (low information entropy), the target weight is relatively low, and the priority in data screening is relatively low.

[0066] In an embodiment of the present invention, various industrial production data are screened based on the target weights of each node's mutation degree, including: Determine the initial screening threshold corresponding to each node based on the historical data of various industrial production data; Based on the initial screening threshold corresponding to each node and the target weight of the mutation degree, obtain the comprehensive score corresponding to each type of industrial production data; Screen the corresponding industrial production data based on the comprehensive score.

[0067] In this embodiment, the mean value of the historical production data features can be calculated. Then, combined with business experience, adjustments are made based on the key feature statistics to obtain the initial screening threshold.

[0068] For example, for the industrial production data node of product quality, if its historical mean is , and an adjustment amount is set according to the basic requirements and experience of the business for product quality, so as to obtain the initial screening threshold .

[0069] After obtaining the initial screening threshold corresponding to each node and the target weight of the mutation degree of each node determined based on the entropy weight method in this embodiment, calculate the comprehensive score corresponding to each type of industrial production data. The calculation of the comprehensive score takes into account the initial screening threshold of the data and its target weight of the mutation degree. The target weight of the mutation degree reflects the relative importance of the data of this node in the entire data set, while the initial screening threshold provides a reference standard related to historical data and business experience. The initial screening threshold and the target weight can be weighted and calculated to obtain the comprehensive score corresponding to each type of industrial production data. The comprehensive score can more comprehensively reflect the value and priority of this industrial production data in the screening process, taking into account both the mutation characteristics of the data itself (reflected by the target weight) and the preliminary screening criteria based on history and business (reflected by the initial screening threshold).

[0070] Finally, in this embodiment, the corresponding industrial production data are screened based on the obtained comprehensive score. Industrial production data with a higher comprehensive score means that it performs more prominently in terms of data mutation degree and screening criteria based on history and business, and contains information valuable for the enterprise's production decision-making. Therefore, it can be preferentially retained or focused on during the screening process; while data with a lower comprehensive score may be given a lower priority or screened out in subsequent analysis. Thus, the screening of various industrial production data is realized, making the screened data more in line with the actual production needs of the enterprise and providing strong support for the enterprise's production decision-making.

[0071] In an embodiment of the present invention, it further includes: Determine the reward function based on production efficiency and defective product rate; Adjust the initial screening threshold based on the reward function.

[0072] In this embodiment, the proximal policy optimization algorithm can be used to dynamically adjust the initial screening threshold.

[0073] In the framework of the proximal policy optimization algorithm in this embodiment, the state not only includes the feature values of the current data sample after processing , as well as the relevant statistical information of the screened data, such as the average comprehensive score and the contribution degrees of different features , but also introduces the correlation strength information between the production indicators and each feature.

[0074] For example, in this embodiment, the Pearson correlation coefficient between the device operation state feature and the production efficiency is calculated , as well as the negative correlation coefficient between the product quality feature and the defective rate etc., and these correlation coefficients are used as part of the state. Therefore, the state is expressed as: . By combining this information into the state , the progress of the current data screening and the relationship between the data features and the production indicators can be comprehensively understood, so as to make more reasonable decisions.

[0075] The action space is set as the adjustment of the weights of each feature and the screening threshold . In each iteration process, the agent makes a decision based on the current state and selects an action , where is the adjustment amount of the feature weight, which is used to change the relative importance of each feature in data screening; is the adjustment amount of the screening threshold, which determines the boundary of data screening.

[0076] In this embodiment, the production efficiency P and the defective rate D are the key indicators in the enterprise production process, which directly reflect the production efficiency and quality. When determining the reward function, these two indicators are incorporated, so that the reward function can closely revolve around the core production goals of the enterprise. The calculation formula of the reward function is:

[0077] Among them, represents the initial production efficiency, represents the initial defective rate, represents the change in production efficiency before and after screening the data, represents the change in defective rate before and after screening the data, It represents the relative change in production efficiency before and after screening data, reflecting the contribution of the screening operation to the improvement of production efficiency; It represents the relative change in the defective rate before and after screening data, reflecting the effect of the screening operation on product quality improvement (reduction of defective rate); It is the sum of the change amounts of the correlation coefficients between each feature and production indicators, indicating the change in the correlation strength between features and production indicators during the screening process; 、 and represent the weight coefficients, which are used to balance the contributions of these three parts to the reward, and can highlight the influence degree of different aspects on the reward according to the actual needs of the enterprise.

[0078] For example, if the enterprise pays more attention to the improvement of production efficiency, it can appropriately increase the value. The reward function can comprehensively consider the impact of screening data on production efficiency, defective rate, and the correlation between features and production indicators, and can accurately evaluate the contribution degree of each screening operation to the enterprise's decision-making goal.

[0079] In the process of reinforcement learning, the agent continuously interacts with the environment, continuously tries different actions (i.e., adjusts the initial screening threshold), and optimizes the decision according to the feedback of the reward function. When a positive reward is obtained, it indicates that the current threshold adjustment direction helps to screen out data that is more valuable to the enterprise's decision-making goal, and then fine-tunes the threshold in this direction; if a negative reward is obtained, it indicates that the current adjustment direction is not conducive to screening out data that meets the enterprise's needs, and at this time, the threshold needs to be adjusted in the opposite direction. The specific adjustment method is as follows: assuming that the current screening threshold is , the adjustment amount is , and the new screening threshold is expressed as:

[0080] Among them, as the learning rate, is used to precisely control the step size of the threshold adjustment to prevent the adjustment amplitude from being too large and missing the optimal solution; is the sign function of the reward function value, which determines the threshold adjustment direction according to the positive or negative of the reward.

[0081] It can be concluded from the above that in this embodiment, the correlation strength between features and production indicators is incorporated into the state representation and reward function design, and the screening strategy can be optimized in real time according to the dynamic mutual relationship of various factors in the production process to ensure that the screened data is always closely aligned with the enterprise's production decision-making goal.

[0082] In an embodiment of the present invention, screening the corresponding industrial production data based on the comprehensive score includes: Screening the industrial production data that meets the first condition; The first condition is that the comprehensive score corresponding to the industrial production data is greater than the comprehensive score threshold.

[0083] In this embodiment, a comprehensive score threshold can be set in advance. The comprehensive score threshold is the baseline for screening industrial production data. The setting of the comprehensive score threshold may be based on the enterprise's expectation of data value, statistical analysis of historical production data, or determined by business experts according to experience. The comprehensive score threshold divides all industrial production data into two parts. Data greater than this threshold is considered to contain more valuable information and better meets the enterprise's goal of screening data.

[0084] Based on the historical data of various industrial production data, the initial screening threshold corresponding to each node is determined, and combined with the target weight of the variation degree of each node, the comprehensive score corresponding to each type of industrial production data is obtained. The comprehensive score takes into account the variation degree of the data itself and the quantitative index of the initial screening criteria based on history and business experience, and can reflect the value and priority of the industrial production data in the screening process.

[0085] The comprehensive score of each type of industrial production data is compared with the set comprehensive score threshold. Only the industrial production data that meets the first condition that the comprehensive score is greater than the comprehensive score threshold can be screened out. The screening process of this embodiment focuses on the data with higher comprehensive scores because these data are more prominent in terms of data variation degree and screening criteria based on history and business, and are more likely to contain information valuable for the enterprise's production decision-making. Thus, the screening of industrial production data is realized, making the screened data more in line with the actual production needs of the enterprise.

[0086] For example, if an enterprise finds in the production process that the data with a comprehensive score higher than a certain threshold has a strong correlation with the improvement of production efficiency or product quality, this threshold can be used as the screening criterion, so as to extract data from a large amount of industrial production data that helps the enterprise achieve these goals.

[0087] It can be concluded from the above that this embodiment uses the comprehensive score being greater than the comprehensive score threshold as the screening condition, and can accurately focus on the data with higher value. By combining the data variation degree and historical business experience to determine the comprehensive score, it is possible to effectively screen out the industrial production data that is of important value to the enterprise's production decision-making.

[0088] Corresponding to the data screening method in the above embodiment, Figure 2 is the structural block diagram of the data screening device provided by an embodiment of the present invention. For the sake of convenience of description, only the parts related to the embodiment of the present invention are shown. Refer to Figure 2 wherein, the data screening device 20 includes: The first weight calculation module 21 is used to calculate the attention weights of each node in the spatio-temporal graph in the spatio-temporal dimension. The spatio-temporal graph is constructed from various industrial production data, and each industrial production data serves as a node in the spatio-temporal graph. The spatio-temporal dimension includes the time dimension and the space dimension; The weight splicing module 22 is used to splice the attention weights of each node in the spatio-temporal graph in the spatio-temporal dimension to obtain the target feature corresponding to each node; The second weight calculation module 23 is used to obtain the target weight of the variation degree of each node based on the information entropy of the target feature; The data screening module 24 is used to screen various industrial production data based on the target weight of the variation degree of each node.

[0089] In an embodiment of the present invention, the various industrial production data includes equipment operation data, production efficiency data, and product quality data; The data screening device 20 further includes: a spatio-temporal graph construction module, specifically used for: Taking the equipment operation data, production efficiency data, and product quality data as nodes respectively; Taking the feature relationships at different time points and different target devices as edges; Constructing a spatio-temporal graph based on the nodes and edges.

[0090] In an embodiment of the present invention, the attention weights of each node in the spatio-temporal dimension include: time attention weights and space attention weights; The weight splicing module 22 is specifically used for: Performing weighted aggregation on the time attention weights to obtain a time feature vector; Performing weighted aggregation on the space attention weights to obtain a space feature vector; Splicing the time feature vector and the space feature vector to obtain the target feature corresponding to each node.

[0091] In an embodiment of the present invention, the second weight calculation module 23 is specifically used for: Obtaining the target weight of the variation degree of each node based on the first calculation formula; The first calculation formula is:

[0092] Wherein, represents the target weight of the variation degree of the th node, m represents the total number of nodes, the information entropy corresponding to the target feature.

[0093] In an embodiment of the present invention, the data screening module 24 is specifically configured to: determine an initial screening threshold corresponding to each node based on historical data of various industrial production data; obtain a comprehensive score corresponding to each type of industrial production data based on the initial screening threshold corresponding to each node and the target weight of the degree of variation; screen the corresponding industrial production data based on the comprehensive score.

[0094] In an embodiment of the present invention, the data screening module 24 is further specifically configured to: determine a reward function based on production efficiency and defective product rate; adjust the initial screening threshold based on the reward function.

[0095] In an embodiment of the present invention, the data screening module 24 is further specifically configured to: screen the industrial production data that meets the first condition; The first condition is that the comprehensive score corresponding to the industrial production data is greater than the comprehensive score threshold.

[0096] See Figure 3 , Figure 3 is a schematic block diagram of an electronic device provided in an embodiment of the present invention. As Figure 3 shown, the electronic device 300 in this embodiment may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 complete mutual communication through a communication bus 305. The memory 304 is used to store a computer program, and the computer program includes program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module / unit in the above-mentioned device embodiments, such as Figure 2 the functions of the first weight calculation module 21, weight splicing module 22, second weight calculation module 23, and data screening module 24 shown.

[0097] It should be understood that in the embodiments of the present invention, the so-called processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0098] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.

[0099] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.

[0100] In specific implementation, the processor 301, the input device 302, and the output device 303 described in the embodiments of the present invention may implement the implementation manners described in the first embodiment and the second embodiment of the data screening method provided by the embodiments of the present invention, and may also implement the implementation manner of the electronic device described in the embodiments of the present invention, which will not be elaborated herein.

[0101] In another embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the method of the above embodiments are implemented. It can also be completed by instructing relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0102] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk equipped on the electronic device, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.

[0103] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0104] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described electronic devices and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described in detail here.

[0105] In several embodiments provided by this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection to each other can be an indirect coupling or communication connection through some interfaces or units, and can also be in the form of electrical, mechanical or other connections.

[0106] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention.

[0107] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0108] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A data screening method, characterized in that, Including: Calculating the attention weights of each node in the spatio-temporal graph in the spatio-temporal dimension, where the spatio-temporal graph is constructed from various industrial production data, and each industrial production data serves as a node in the spatio-temporal graph, and the spatio-temporal dimension includes a time dimension and a space dimension; Concatenating the attention weights of each node in the spatio-temporal graph in the spatio-temporal dimension to obtain the target feature corresponding to each node; Obtaining the target weight of the mutation degree of each node based on the information entropy of the target feature; Screening various industrial production data based on the target weight of the mutation degree of each node.

2. The data screening method according to claim 1, characterized in that The various industrial production data includes equipment operation data, production efficiency data, and product quality data; The data screening method further includes: Regarding the equipment operation data, production efficiency data, and product quality data as nodes respectively; Regarding the feature relationships at different time points and different target devices as edges; Constructing a spatio-temporal graph based on the nodes and the edges.

3. The data screening method according to claim 1, characterized in that, The attention weight of each node in the spatio-temporal dimension includes: a time attention weight and a space attention weight; The concatenating the attention weights of each node in the spatio-temporal graph in the spatio-temporal dimension to obtain the target feature corresponding to each node includes: Performing weighted aggregation on the time attention weights to obtain a time feature vector; Performing weighted aggregation on the space attention weights to obtain a space feature vector; Concatenating the time feature vector and the space feature vector to obtain the target feature corresponding to each node.

4. The data screening method according to claim 1, wherein The obtaining the target weight of the mutation degree of each node based on the information entropy of the target feature includes: Obtaining the target weight of the mutation degree of each node based on a first calculation formula; The first calculation formula is: Among them, represents the target weight of the mutation degree of the th node, m represents the total number of nodes, the information entropy corresponding to the target feature.

5. The data screening method according to claim 1, wherein The screening various industrial production data based on the target weight of the mutation degree of each node includes: Determining the initial screening threshold corresponding to each node based on the historical data of various industrial production data; Obtaining the comprehensive score corresponding to each industrial production data based on the initial screening threshold corresponding to each node and the target weight of the mutation degree; Screening the corresponding industrial production data based on the comprehensive score.

6. The data screening method according to claim 5, wherein It further includes: Determining a reward function based on production efficiency and defective product rate; Adjusting the initial screening threshold based on the reward function.

7. The data screening method according to claim 5, wherein The screening the corresponding industrial production data based on the comprehensive score includes: Screening the industrial production data that meets the first condition; The first condition is that the comprehensive score corresponding to the industrial production data is greater than the comprehensive score threshold.

8. A data screening device, characterized in that, Including: A first weight calculation module for calculating the attention weights of each node in the spatio-temporal graph in the spatio-temporal dimension, where the spatio-temporal graph is constructed from various industrial production data, and each industrial production data serves as a node in the spatio-temporal graph, and the spatio-temporal dimension includes a time dimension and a space dimension; A weight concatenating module for concatenating the attention weights of each node in the spatio-temporal graph in the spatio-temporal dimension to obtain the target feature corresponding to each node; A second weight calculation module for obtaining the target weight of the mutation degree of each node based on the information entropy of the target feature; A data screening module for screening various industrial production data based on the target weight of the mutation degree of each node.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Demand-oriented intelligent screening method for multi-source heterogeneous data of power transmission and transformation equipment

    CN116431623A

  • Node monitoring method and device, electronic equipment and storage medium

    CN116860556A

  • Industrial time series data anomaly detection method based on space-time diagram attention network

    CN117272196A

  • Network operation comprehensive evaluation method based on space-time perspective

    CN117875755A

  • Sensor data fusion method and device based on graph neural network, and storage medium

    CN118364432A