Method and system for mining and analyzing AMC standard exceeding reason of active learning framework
Through the active learning framework method, the analysis process of AMC reasons for exceeding the standard is simplified, the analysis efficiency and accuracy are improved, and the inefficiency problem caused by intricate data screening in the existing technology is solved.
Patent Information
- Application Number
- CN202510637911.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-19
AI Technical Summary
In the prior art, the automated data analysis scheme is not efficient in analyzing the causes of AMC exceeding the standard, mainly because the data screening process is not refined enough, resulting in less useful data in the analysis results.
Using the method of active learning framework, we use the method to obtain the monitoring point data of the AMC generated scene, create statistical matrix and data matrix, calculate the data information amount, extract sub-matrix sequence, delete data with low information amount, simplify the matrix sequence, and use it as a feature to build a sample set to determine the reason for exceeding the standard.
The efficiency of AMC reasons for exceeding the standard has been improved. By actively screening the sample set, the number of samples is reduced, the sample quality is improved, and the accuracy and efficiency of the analysis results are significantly improved.
Smart Images

Figure CN120162680A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of analysis of the causes of AMC exceeding the standard, and specifically, to a method and system for mining and analyzing the causes of AMC exceeding the standard based on an active learning framework. Background Art
[0002] AMC refers to airborne molecular contaminants, including organic and inorganic contaminants in the cleanroom air, existing in the form of gases, vapors, and airborne dust, such as acids, alkalis, polymer additives, organometallic compounds, etc. In microelectronics manufacturing, AMC can harm the production process, reduce the yield, and affect the product quality. Therefore, manual analysis is required.
[0003] Since there are many devices in the AMC generation scenario and the amount of data to be analyzed is large, the workload of the manual analysis process is extremely large. Therefore, many automated data analysis solutions have emerged in the prior art. However, these automated data analysis solutions mostly perform comprehensive analysis on the data, and there is not much really useful data, and the efficiency is still not very high. Therefore, how to provide an AMC exceeding the standard cause analysis solution containing a data screening process to improve the efficiency of analyzing the cause of exceeding the standard is the technical problem that the technical solution of the present invention wants to solve. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for mining and analyzing the causes of AMC exceeding the standard based on an active learning framework to solve the problems raised in the above background art.
[0005] To achieve the above purpose, the present invention provides the following technical solutions: A method and system for mining and analyzing the causes of AMC exceeding the standard based on an active learning framework, the method comprising: Obtain the monitoring points in the AMC generation scenario, create a statistical matrix according to the positions of the monitoring points, and obtain the data at the monitoring points at each moment based on the statistical matrix to obtain a data matrix with time tags; the monitoring points at least include the positions where sensors and signal transceivers are installed; Receive the analysis span input by the staff, obtain the data matrix sequence within the analysis span, and calculate the data information amount of any row and column position of the data matrices in the data matrix sequence; the information amount is jointly determined by the information entropy and periodicity; the analysis span is a time span; Query the AMC concentration containing time points obtained by the AMC monitor within the analysis span. When any AMC concentration reaches a preset threshold, extract the sub-matrix sequence corresponding to the AMC monitor in the data matrix sequence; wherein, the number of sub-matrices in the sub-matrix sequence is a preset value; Mark the row and column positions in the sub-matrix sequence where the data information amount is less than the preset information amount threshold, and delete the data at the marked row and column positions in the sub-matrix sequence to obtain a simplified matrix sequence; Use the obtained simplified matrix sequence as features and the AMC concentration as labels to construct a sample set, and determine the over-standard reason based on the sample set.
[0006] As a further solution of the present invention: the steps of receiving the analysis span input by the staff, obtaining the data matrix sequence within the analysis span, and calculating the data information amount of any row and column position in the data matrix in the data matrix sequence include: Receive the analysis span input by the staff and obtain the data matrix sequence within the analysis span; For any row and column position, sequentially read the data at this row and column position in the data matrices in the data matrix sequence, arrange the data based on the sequence order to obtain an array for each row and column position; Perform dimensionless processing on the data in the array and calculate the information entropy; Perform periodic identification on the dimensionless processed array and determine the probability of the existence of a period; Determine the data information amount according to the information entropy and the probability of the existence of a period.
[0007] As a further solution of the present invention: the process of performing periodic identification on the dimensionless processed array and determining the probability of the existence of a period includes: Perform discrete Fourier transform on the dimensionless processed array to obtain a frequency domain group; Calculate the standard deviation of the amplitudes of each frequency component in the frequency domain group; Fit the amplitudes of each frequency component in the frequency domain group into a curve and calculate the curvature of each point on the curve; When the curvature reaches the preset threshold, mark the corresponding frequency; When the interval length of the marked frequency is less than the preset length threshold and the standard deviation is greater than the preset standard deviation threshold, determine the probability of the existence of a period according to the marked interval length and the standard deviation.
[0008] As a further solution of the present invention: the steps of querying the AMC concentration containing time points obtained by the AMC monitor within the analysis span, and extracting the sub-matrix sequence corresponding to the AMC monitor in the data matrix sequence when any AMC concentration reaches the preset threshold include: Query the AMC concentration obtained in real time by the AMC monitor within the analysis span; the AMC concentration contains time tags; Compare the AMC concentration with the preset concentration threshold; When the concentration of any AMC reaches the preset threshold, query the location of the AMC monitor, query the monitoring point closest to the location of the AMC monitor, and query the row and column positions corresponding to the monitoring point; Determine the extraction side length according to the AMC concentration, and expand the queried row and column positions based on the extraction side length to obtain the position and size of the sub-matrix; the extraction side length is proportional to the AMC concentration; Determine the number of sub-matrices based on the preset backtracking duration, select the number of data matrices equal to the number of sub-matrices in the data matrix sequence based on the current time, intercept the sub-matrices in the data matrix according to the position and size of the sub-matrices, and arrange the sub-matrices in sequence order to obtain a sub-matrix sequence.
[0009] As a further solution of the present invention: the step of marking the row and column positions where the data information amount in the sub-matrix sequence is less than the preset information amount threshold, and deleting the data at the marked row and column positions in the sub-matrix sequence to obtain a simplified matrix sequence includes: Query the data information amount of each row and column position in the sub-matrix sequence; Compare the data information amount with the preset information amount threshold, and when the data information amount is less than the preset information amount threshold, mark the row and column position; For each sub-matrix in the sub-matrix sequence, delete the data at the marked row and column positions, and regularize the remaining data; When each sub-matrix is regularized, a simplified matrix sequence is obtained.
[0010] As a further solution of the present invention: the step of using the obtained simplified matrix sequence as a feature, using the AMC concentration as a label, constructing a sample set, and determining the cause of exceeding the standard based on the sample set includes: Use the obtained simplified matrix sequence as a feature, use the AMC concentration as a label, and construct a sample set; Train an AMC concentration prediction model according to the sample set; Read the parameters corresponding to each variable in the AMC concentration prediction model, and calculate the weights of the parameters; Select variables according to the weights of the parameters, and query the types of the monitoring points corresponding to the variables as the reasons for exceeding the standard.
[0011] The technical solution of the present invention also provides an AMC exceeding standard cause mining and analysis system for an active learning framework, and the system includes: A data statistics module, configured to obtain the monitoring points in the AMC generation scenario, create a statistical matrix according to the positions of the monitoring points, obtain the data at each monitoring point at each moment based on the statistical matrix, and obtain a data matrix with time tags; the monitoring points at least include the positions where sensors and signal transceivers are installed; An information quantity calculation module, configured to receive the analysis span input by the staff, obtain the data matrix sequence within the analysis span, and calculate the data information quantity at any row and column position of the data matrix in the data matrix sequence; the information quantity is jointly determined by information entropy and periodicity; the analysis span is a time span; A sub-matrix extraction module, configured to query the AMC concentration containing time points obtained by the AMC monitor within the analysis span, and when any AMC concentration reaches a preset threshold, extract the sub-matrix sequence corresponding to the AMC monitor in the data matrix sequence; wherein, the number of sub-matrices in the sub-matrix sequence is a preset value; A matrix simplification module, configured to mark the row and column positions where the data information quantity is less than the preset information quantity threshold in the sub-matrix sequence, and delete the data at the marked row and column positions in the sub-matrix sequence to obtain a simplified matrix sequence; An over-standard cause analysis module, configured to use the obtained simplified matrix sequence as a feature and the AMC concentration as a label to construct a sample set, and determine the over-standard cause based on the sample set.
[0012] As a further solution of the present invention: the information quantity calculation module includes: A matrix sequence acquisition unit, configured to receive the analysis span input by the staff and obtain the data matrix sequence within the analysis span; An array extraction unit, configured to, for any row and column position, sequentially read the data at this row and column position in the data matrices in the data matrix sequence, arrange the data based on the sequence order, and obtain an array for each row and column position; An information entropy calculation unit, configured to perform non-dimensionalization processing on the data in the array and calculate the information entropy; A period recognition unit, configured to perform periodicity recognition on the non-dimensionalized array and determine the probability of the existence of a period; A calculation execution unit, configured to determine the data information quantity according to the information entropy and the probability of the existence of a period.
[0013] As a further solution of the present invention: the sub-matrix extraction module includes: A concentration acquisition unit, configured to query the AMC concentration obtained in real time by the AMC monitor within the analysis span; the AMC concentration contains time tags; A concentration comparison unit, configured to compare the AMC concentration with a preset concentration threshold; A position query unit, configured to, when any AMC concentration reaches the preset threshold, query the position of the AMC monitor, query the monitoring point closest to the position of the AMC monitor, and query the row and column positions corresponding to the monitoring point; A position expansion unit, configured to determine an extraction side length according to the AMC concentration, expand the queried row and column positions based on the extraction side length, and obtain the position and size of the sub-matrix; the extraction side length is proportional to the AMC concentration; A matrix intercepting unit, configured to determine the number of sub-matrices based on a preset backtracking duration, select the number of data matrices in the data matrix sequence based on the current moment, intercept sub-matrices in the data matrix according to the position and size of the sub-matrices, and arrange the sub-matrices based on the sequence order to obtain a sub-matrix sequence.
[0014] As a further solution of the present invention: the matrix simplification module includes: An information amount query unit, configured to query the data information amounts of each row and column position in the sub-matrix sequence; A position marking unit, configured to compare the data information amount with a preset information amount threshold, and mark the row and column position when the data information amount is less than the preset information amount threshold; A data regularization unit, configured to, for each sub-matrix in the sub-matrix sequence, delete the data at the marked row and column positions and regularize the remaining data; A sequence output unit, configured to obtain a simplified matrix sequence when each sub-matrix has been regularized.
[0015] Compared with the prior art, the beneficial effects of the present invention are: The present invention provides a data screening solution based on the data volume, actively screens the sample set, reduces the number of samples in the sample set, improves the quality of the samples in the sample set. This preliminary work belongs to the active learning framework and is docked with the existing recognition process, greatly improving the cause analysis efficiency. Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention.
[0017] Figure 1 Shows the overall flow block diagram of the method for mining and analyzing the reasons for AMC exceeding the standard in the active learning framework.
[0018] Figure 2 Shows the structural diagram of the system for mining and analyzing the reasons for AMC exceeding the standard in the active learning framework. Detailed Embodiments
[0019] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer, the following further details the present invention with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0020] Figure 1 The overall process block diagram of the method and system for mining and analyzing the reasons for AMC exceeding the standard in the active learning framework. In an embodiment of the present invention, a method for mining and analyzing the reasons for AMC exceeding the standard in the active learning framework, the method includes: Step S100: Obtain the monitoring points in the AMC generation scenario, create a statistical matrix according to the positions of the monitoring points, and obtain the data at the monitoring points at each moment based on the statistical matrix to obtain a data matrix with time tags; the monitoring points at least include the positions where sensors and signal transceivers are installed; AMC refers to airborne molecular contaminants, including organic and inorganic contaminants in the cleanroom air, existing in the form of gas, vapor and airborne dust, such as acids, alkalis, polymer additives, organometallic compounds, etc. In microelectronics manufacturing, AMC will endanger the production process, resulting in a reduction in the yield rate, affecting the product quality. There are generally multiple devices in the scenario where AMC is generated, and various sensors are installed in these devices to obtain device data. In addition, some environmental monitoring devices, such as temperature measuring instruments, are also installed in the scenario. Any device with information collection function, its corresponding position is the monitoring point in the above content. Obtain the positions of all monitoring points, and a statistical matrix can be created according to the position relationship. The row and column positions in the statistical matrix correspond one by one to the monitoring points, and the corresponding relationship remains unchanged; the devices at these monitoring points obtain data in real time, obtain data with time tags, and insert them into the corresponding row and column positions in the statistical matrix to obtain a data matrix with time tags; it is worth mentioning that since the devices corresponding to the monitoring points may be installed at different heights in the scenario, therefore, the statistical matrix is actually a three-dimensional matrix. In addition, the number of monitoring points in each direction is not necessarily uniform, so there may be some empty positions in the statistical matrix, and default data can be directly inserted at the empty positions.
[0021] Step S200: Receive the analysis span input by the staff, obtain the data matrix sequence within the analysis span, and calculate the data information amount of the data at any row and column position of the data matrix in the data matrix sequence; the information amount is jointly determined by the information entropy and periodicity; the analysis span is a time span; The analysis span is a time span, obtained by the input of the staff, such as one day, three days or one week, etc.; query all the data matrices within the analysis span based on the time tags and sort them in chronological order to obtain a data matrix sequence (a set of data matrices with order). In the data matrix sequence, each data matrix is generated based on the statistical matrix and has the same size. For any row and column position, calculate the data information amount according to the data at this row and column position. The independent variables in the calculation process adopt the two parameters of information entropy and periodicity.
[0022] Step S300: query and analyze the AMC concentrations at the time points obtained by the AMC monitor within the span, and when any AMC concentration reaches a preset threshold, extract a submatrix sequence corresponding to the AMC monitor from the data matrix sequence; wherein the number of submatrices in the submatrix sequence is a preset value; At the same time, there will be multiple AMC monitors in the AMC generation scenario. The monitoring points do not include the locations where the AMC monitors are installed. The monitoring points and AMC monitors are two independent monitoring processes. The AMC concentrations at each moment obtained by the AMC monitors within the query and analysis span are analyzed. When any AMC concentration is large enough, the sub-matrix sequence corresponding to the AMC monitor is extracted from the data matrix sequence. The meaning of this process is that when there is a high concentration of AMC pollutants at a certain location, the monitoring data around the location is obtained with the location as the center. Under the premise that the monitoring data of the monitoring points are regarded as the cause, this process is equivalent to simplifying the monitoring data, which is essentially a restriction of the sample and belongs to the active learning framework.
[0023] Step S400: marking row and column positions in the submatrix sequence whose data information amount is less than a preset information amount threshold, and deleting the data at the marked row and column positions in the submatrix sequence to obtain a simplified matrix sequence; Based on the simplification in step S300, the above content deletes some data again according to the calculated data information amount, and further simplifies to obtain a simplified matrix sequence.
[0024] Step S500: Using the obtained simplified matrix sequence as a feature and the AMC concentration as a label, a sample set is constructed, and the cause of exceeding the standard is determined based on the sample set; The obtained sub-matrix sequence is the simplified monitoring data. It is used as a feature (independent variable), and the corresponding AMC concentration is used as a label (there is only one AMC concentration, the dependent variable). A sample set is constructed and analyzed to determine the cause of the exceeding of the standard. This process can be assisted by a supervised learning model. Since steps S100 to S400 provide a solution for simplifying the sample set, it is actually also an active learning framework.
[0025] Regarding step S200, the steps of receiving the analysis span input by the staff, obtaining the data matrix sequence within the analysis span, and calculating the data information amount of any row and column position of the data matrix in the data matrix sequence include: Receive the analysis span input by the staff, and obtain the data matrix sequence within the analysis span; For any row and column position, read the data at the row and column position in the data matrix in the data matrix sequence in turn, arrange the data based on the sequence order, and obtain an array for each row and column position; Dimensionalize the data in the array and calculate the information entropy; Perform periodic identification on the dimensionless array and determine the probability of the existence of a period; Determine the data information volume based on the information entropy and the probability of the existence of a period.
[0026] Receive the analysis span input by the staff, obtain the data matrix sequence within the analysis span. For any row and column position, each data matrix in the data matrix sequence has a data at this row and column position. Read these data and arrange them in sequence order to obtain an array. Since each row and column position corresponds to a monitoring device and there are many types of them, it is rather troublesome to process them together. Therefore, a dimensionalization process needs to be carried out first. The dimensionalization method is not complicated. Obtain the maximum value, calculate the difference between the maximum value and the minimum value as the denominator. For any data, calculate the difference between it and the minimum value as the numerator. The obtained ratio is the data after dimensionalization. Of course, a product constant can also be introduced to adjust the value range of the data after dimensionalization.
[0027] On this basis, perform periodic identification on the dimensionless array, determine the probability of the existence of a period, and determine the data information volume based on the information entropy and the probability of the existence of a period.
[0028] Furthermore, the calculation process of the information entropy of the array is not complicated and belongs to existing parameters. Its calculation method is public and will not be elaborated in this invention. However, the periodic identification process needs to be described as follows: The process of performing periodic identification on the dimensionless array and determining the probability of the existence of a period includes: Perform discrete Fourier transform on the dimensionless array to obtain a frequency domain group; Calculate the standard deviation of the amplitudes of each frequency component in the frequency domain group; Fit the amplitudes of each frequency component in the frequency domain group into a curve and calculate the curvature of each point on the curve; When the curvature reaches a preset threshold, mark the corresponding frequency; When the interval length of the marked frequency is less than the preset length threshold and the standard deviation is greater than the preset standard deviation threshold, determine the probability of the existence of a period based on the marked interval length and the standard deviation.
[0029] Performing a discrete Fourier transform operation can extract the frequency-domain features of an array, calculate the standard deviation of the amplitudes of each frequency component in the frequency-domain group, and then locate the peak data in the frequency-domain group. The fewer the peak data (the length of the marked position, corresponding to the frequency interval), the greater the standard deviation, indicating that the data in the frequency-domain group is more concentrated, the periodicity is more obvious, and the probability of the existence of a period is greater. Therefore, the probability of the existence of a period is inversely proportional to the length of the marked interval and directly proportional to the standard deviation. The relevant function is set by the staff according to the actual situation, as long as it conforms to this relationship. The present invention only defines the relevant relationship and does not limit the specific numerical relationship.
[0030] Furthermore, the data information amount is determined according to the information entropy and the probability of the existence of a period. The greater the information entropy, the more uncertain the corresponding data is. In the present invention, the more uncertain the data is considered to be, the higher its importance is, that is, the more it needs to be retained as a sample. Therefore, the greater the data information amount; the greater the probability of the existence of a period, the more obvious the periodicity is, the data is considered to be predictable, and its stability is higher, and it is considered to be less important and less in need of retention. Therefore, the smaller the data information amount; generally speaking, the data information amount is directly proportional to the information entropy and inversely proportional to the probability of the existence of a period. Similar to the above content, the present invention only defines the relevant relationship and does not limit the specific quantitative relationship, because those skilled in the technical field of the present invention can select many functions in the prior art to achieve this relationship.
[0031] Regarding step S300, when the AMC concentration containing time points obtained by the AMC monitor within the query analysis span reaches a preset threshold, the steps of extracting the sub-matrix sequence corresponding to the AMC monitor in the data matrix sequence include: The AMC concentration obtained in real time by the AMC monitor within the query analysis span; the AMC concentration contains time tags; Compare the AMC concentration with the preset concentration threshold; When any AMC concentration reaches the preset threshold, query the position of the AMC monitor, query the monitoring point closest to the position of the AMC monitor, and query the row and column positions corresponding to the monitoring point; Determine the extraction side length according to the AMC concentration, and expand the queried row and column positions based on the extraction side length to obtain the position and size of the sub-matrix; the extraction side length is directly proportional to the AMC concentration; Determine the number of sub-matrices based on the preset backtracking duration, select the number of data matrices as the number of sub-matrices in the data matrix sequence based on the current time, intercept the sub-matrices in the data matrix according to the position and size of the sub-matrices, and arrange the sub-matrices based on the sequence order to obtain the sub-matrix sequence.
[0032] Query the AMC concentration obtained in real time by the AMC monitor within the query analysis span. When any AMC concentration reaches a sufficiently large value, query the position of the AMC monitor, query the monitoring point position closest to the position of the AMC monitor, and query the corresponding row and column positions of the monitoring point. At this time, obtain the corresponding position of the AMC monitor in the data matrix. Determine the extraction side length according to the AMC concentration. The extraction side length is proportional to the AMC concentration. One calculation method is ; is the extraction side length, is a preset coefficient, is the concentration. The calculated extraction side length is an odd number, and the size of the obtained sub-matrix is 5*5*5 or 7*7*7. At this time, limit the position of the sub-matrix to the central position, and the central position can adopt the corresponding position of the AMC monitor in the data matrix. Then, starting from the current moment, read the data matrix within the preset retrospective duration forward, locate the sub-matrix in each read data matrix, intercept the data, and arrange the sub-matrices based on the sequence order to obtain the sub-matrix sequence.
[0033] Among them, the sequence order in the present invention is the order of the data matrix sequence, which is essentially the time order.
[0034] Regarding step S400, the steps of marking the row and column positions in the sub-matrix sequence where the data information amount is less than the preset information amount threshold and deleting the data at the marked row and column positions in the sub-matrix sequence to obtain the simplified matrix sequence include: Query the data information amount of each row and column position in the sub-matrix sequence; Compare the data information amount with the preset information amount threshold. When the data information amount is less than the preset information amount threshold, mark the row and column position; For each sub-matrix in the sub-matrix sequence, delete the data at the marked row and column positions and regularize the remaining data; When each sub-matrix is regularized, obtain the simplified matrix sequence.
[0035] Query the data information amount of each row and column position in the sub-matrix sequence. When the data information amount is small, it is considered that the data at the corresponding row and column positions is unimportant. At this time, for each sub-matrix in the sub-matrix sequence, delete the data at the marked row and column positions. After deleting the data, insert some preset default values at the row and column positions where the data is deleted to ensure that the size of the matrix remains unchanged. This belongs to the matrix regularization process. Of course, other data regularization processes can also be provided, which will not be elaborated in the present invention. When each sub-matrix is regularized, obtain the simplified matrix sequence.
[0036] As step S500, the step of using the obtained simplified matrix sequence as a feature and the AMC concentration as a label to construct a sample set and determining the cause of exceeding the standard based on the sample set includes: Using the obtained simplified matrix sequence as a feature and the AMC concentration as a label to construct a sample set; Training an AMC concentration prediction model according to the sample set; Reading the parameters corresponding to each variable in the AMC concentration prediction model and calculating the weights of the parameters; Selecting variables according to the weights of the parameters, querying the types of the monitoring points corresponding to the variables, and taking them as the reasons for exceeding the standard.
[0037] Using the obtained simplified matrix sequence as a feature and the AMC concentration as a label to construct a sample set, training an AMC concentration prediction model according to the sample set. The AMC concentration prediction model is essentially a function. The data at each row and column position is a variable (this variable is in array format, or can be said to be a vector), and each variable corresponds to a parameter. By comparing the parameters between different variables, the size of the parameters can be determined. The size of the parameters represents the degree of influence, that is, the weights in the above content; selecting the parameters with weights greater than the preset threshold, querying the types of the monitoring points corresponding to the variables, and taking them as the reasons for exceeding the standard; of course, the variables can also be sorted according to the parameters, and then the types of the monitoring points corresponding to the variables can be sorted to obtain the importance of the monitoring points, which is equivalent to determining the importance of various reasons.
[0038] It is worth mentioning that the trained AMC concentration prediction model in the above content can also be used for the prediction of AMC pollutants. The AMC concentration can be predicted based on the data obtained from the monitoring points, which is an additional function realized by the technical solution of the present invention.
[0039] Figure 2 The structure diagram of the AMC exceeding-standard cause mining and analysis system of the active learning framework is shown. In a preferred embodiment of the technical solution of the present invention, an AMC exceeding-standard cause mining and analysis system of the active learning framework is further provided. The system 10 includes: A data statistics module 11, configured to obtain the monitoring points in the AMC generation scenario, create a statistical matrix according to the positions of the monitoring points, and obtain the data at the monitoring points at each moment based on the statistical matrix to obtain a data matrix with time tags; the monitoring points at least include the positions where sensors and signal transceivers are installed; An information amount calculation module 12, configured to receive the analysis span input by the staff, obtain the data matrix sequence within the analysis span, and calculate the data information amount at any row and column position of the data matrices in the data matrix sequence; the information amount is jointly determined by the information entropy and periodicity; the analysis span is a time span; The sub - matrix extraction module 13 is used to query and analyze the AMC concentration with time points obtained by the AMC monitor within the analysis span. When any AMC concentration reaches the preset threshold, extract the sub - matrix sequence corresponding to the AMC monitor in the data matrix sequence; wherein, the number of sub - matrices in the sub - matrix sequence is a preset value; The matrix simplification module 14 is used to mark the row and column positions in the sub - matrix sequence where the data information amount is less than the preset information amount threshold, and delete the data at the marked row and column positions in the sub - matrix sequence to obtain a simplified matrix sequence; The over - standard cause analysis module 15 is used to use the obtained simplified matrix sequence as features and the AMC concentration as labels to construct a sample set, and determine the over - standard cause based on the sample set.
[0040] Furthermore, the information amount calculation module 12 includes: The matrix sequence acquisition unit is used to receive the analysis span input by the staff and acquire the data matrix sequence within the analysis span; The array extraction unit is used to, for any row and column position, sequentially read the data at this row and column position in the data matrices in the data matrix sequence, arrange the data based on the sequence order, and obtain an array for each row and column position; The information entropy calculation unit is used to perform non - dimensionalization processing on the data in the array and calculate the information entropy; The period recognition unit is used to perform periodic recognition on the non - dimensionalized array and determine the probability of the existence of a period; The calculation execution unit is used to determine the data information amount according to the information entropy and the probability of the existence of a period.
[0041] Specifically, the sub - matrix extraction module 13 includes: The concentration acquisition unit is used to query the AMC concentration obtained in real - time by the AMC monitor within the analysis span; the AMC concentration contains a time tag; The concentration comparison unit is used to compare the AMC concentration with the preset concentration threshold; The position query unit is used to, when any AMC concentration reaches the preset threshold, query the position of the AMC monitor, query the monitoring point closest to the position of the AMC monitor, and query the corresponding row and column positions of the monitoring point; The position expansion unit is used to determine the extraction side length according to the AMC concentration, and expand the queried row and column positions based on the extraction side length to obtain the position and size of the sub - matrix; the extraction side length is proportional to the AMC concentration; A matrix intercepting unit, configured to determine the number of sub-matrices based on a preset backtracking duration, select the number of data matrices as the number of sub-matrices in the data matrix sequence according to the current moment, intercept sub-matrices in the data matrix according to the positions and sizes of the sub-matrices, and arrange the sub-matrices based on the sequence order to obtain a sub-matrix sequence.
[0042] Further, the matrix simplification module 14 includes: An information quantity query unit, configured to query the data information quantity at each row and column position in the sub-matrix sequence; A position marking unit, configured to compare the data information quantity with a preset information quantity threshold, and mark the row and column position when the data information quantity is less than the preset information quantity threshold; A data regularization unit, configured to, for each sub-matrix in the sub-matrix sequence, delete the data at the marked row and column positions and regularize the remaining data; A sequence output unit, configured to obtain a simplified matrix sequence when each sub-matrix has been regularized.
[0043] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the description and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A method for mining and analyzing the causes of AMC exceeding the standard based on an active learning framework, characterized in that: The method comprises: Acquire monitoring points in the AMC generation scene, create a statistical matrix according to the locations of the monitoring points, acquire data at the monitoring points at each time based on the statistical matrix, and obtain a data matrix containing time tags; the monitoring points at least include locations where sensors and signal transceivers are installed; Receive the analysis span input by the staff, obtain the data matrix sequence within the analysis span, and calculate the amount of data information at any row and column position of the data matrix in the data matrix sequence; the amount of information is determined by information entropy and periodicity; the analysis span is a time span; The AMC concentrations at the time points obtained by the AMC monitor within the query analysis span are analyzed. When any AMC concentration reaches a preset threshold, a submatrix sequence corresponding to the AMC monitor is extracted from the data matrix sequence; wherein the number of submatrices in the submatrix sequence is a preset value; Marking row and column positions in the submatrix sequence where the amount of data information is less than a preset information amount threshold, and deleting the data at the marked row and column positions in the submatrix sequence to obtain a simplified matrix sequence; The obtained simplified matrix sequence was used as a feature, and the AMC concentration was used as a label to construct a sample set, and the cause of the exceeding of the standard was determined based on the sample set.
2. The AMC over-limit cause mining and analysis method based on the active learning framework according to claim 1 is characterized in that: The steps of receiving the analysis span input by the staff, obtaining the data matrix sequence within the analysis span, and calculating the data information amount of any row and column position of the data matrix in the data matrix sequence include: Receive the analysis span input by the staff, and obtain the data matrix sequence within the analysis span; For any row and column position, read the data at the row and column position in the data matrix in the data matrix sequence in turn, arrange the data based on the sequence order, and obtain an array for each row and column position; Non-dimensionalize the data in the array and calculate the information entropy; Perform periodicity identification on the dimensionless array and determine the probability of the existence of a period; The amount of data information is determined based on information entropy and the probability of existence of cycles.
3. The AMC over-standard cause mining and analysis method based on the active learning framework according to claim 2 is characterized in that: The process of identifying the periodicity of the dimensionless array and determining the probability of the existence of a period includes: Perform discrete Fourier transform on the dimensionless array to obtain a frequency domain group; Calculate the standard deviation of the amplitude of each frequency component in the frequency domain group; Fit the amplitude of each frequency component in the frequency domain group into a curve, and calculate the curvature of each point on the curve; When the curvature reaches a preset threshold, the corresponding frequency is marked; When the interval length of the marked frequency is less than a preset length threshold and the standard deviation is greater than a preset standard deviation threshold, the probability of the existence of a cycle is determined based on the marked interval length and the standard deviation.
4. The AMC over-limit cause mining and analysis method based on the active learning framework according to claim 1 is characterized in that: The query analysis span includes the AMC concentrations at the time points acquired by the AMC monitor, and when any AMC concentration reaches a preset threshold, the step of extracting a submatrix sequence corresponding to the AMC monitor in the data matrix sequence comprises: Query and analyze the AMC concentration obtained in real time by the AMC monitor within the span; the AMC concentration contains a time tag; Compare the AMC concentration with a preset concentration threshold; When any AMC concentration reaches a preset threshold, query the location of the AMC monitor, query the monitoring point closest to the AMC monitor, and query the row and column position corresponding to the monitoring point; Determine the extraction side length according to the AMC concentration, expand the queried row and column positions based on the extraction side length, and obtain the position and size of the submatrix; the extraction side length is proportional to the AMC concentration; The number of sub-matrices is determined based on the preset backtracking time, a data matrix with the number of sub-matrices is selected in the data matrix sequence based on the current moment, a sub-matrix is intercepted in the data matrix according to the position and size of the sub-matrix, and the sub-matrix is arranged based on the sequence order to obtain a sub-matrix sequence.
5. The AMC over-standard cause mining and analysis method based on the active learning framework according to claim 1 is characterized in that: The step of marking row and column positions in the submatrix sequence where the amount of data information is less than a preset information amount threshold, and deleting the data at the marked row and column positions in the submatrix sequence to obtain a simplified matrix sequence comprises: Query the amount of data information at each row and column position in the submatrix sequence; Compare the data information amount with a preset information amount threshold, and when the data information amount is less than the preset information amount threshold, mark the row and column position; For each submatrix in the submatrix sequence, delete the data at the marked row and column positions, and regularize the remaining data; When each submatrix is regularized, a simplified matrix sequence is obtained.
6. The AMC over-limit cause mining and analysis method based on the active learning framework according to claim 1 is characterized in that: The steps of using the obtained simplified matrix sequence as a feature, using the AMC concentration as a label, constructing a sample set, and determining the cause of exceeding the standard based on the sample set include: The obtained simplified matrix sequence is used as a feature and the AMC concentration is used as a label to construct a sample set; Training an AMC concentration prediction model based on the sample set; Read the parameters corresponding to each variable in the AMC concentration prediction model and calculate the weight of the parameters; Select variables according to the weight of the parameters, and query the type of monitoring point corresponding to the variable as the reason for exceeding the standard.
7. An AMC over-standard cause mining and analysis system based on an active learning framework, characterized in that: The system comprises: A data statistics module is used to obtain monitoring points in the AMC generation scene, create a statistical matrix according to the location of the monitoring points, obtain data at the monitoring points at each time based on the statistical matrix, and obtain a data matrix containing time tags; the monitoring points at least include locations where sensors and signal transceivers are installed; An information volume calculation module is used to receive the analysis span input by the staff, obtain the data matrix sequence within the analysis span, and calculate the data information volume of any row and column position of the data matrix in the data matrix sequence; the information volume is determined by information entropy and periodicity; the analysis span is a time span; A submatrix extraction module is used to query and analyze the AMC concentrations at the time points obtained by the AMC monitor within the span. When any AMC concentration reaches a preset threshold, a submatrix sequence corresponding to the AMC monitor is extracted from the data matrix sequence; wherein the number of submatrices in the submatrix sequence is a preset value; A matrix simplification module is used to mark the row and column positions in the submatrix sequence where the amount of data information is less than a preset information amount threshold, and delete the data at the marked row and column positions in the submatrix sequence to obtain a simplified matrix sequence; The module for analyzing the causes of exceeding the standard is used to construct a sample set by taking the obtained simplified matrix sequence as a feature and the AMC concentration as a label, and to determine the causes of exceeding the standard based on the sample set.
8. The AMC over-standard cause mining and analysis system based on the active learning framework according to claim 7 is characterized in that: The information volume calculation module includes: A matrix sequence acquisition unit is used to receive the analysis span input by the staff and acquire the data matrix sequence within the analysis span; An array extraction unit is used for, for any row and column position, sequentially reading the data at the row and column position in the data matrix in the data matrix sequence, arranging the data based on the sequence order, and obtaining an array of each row and column position; An information entropy calculation unit is used to perform dimensionless processing on the data in the array and calculate the information entropy; A period identification unit is used to identify the periodicity of the dimensionless array and determine the probability of the existence of a period; The computing execution unit is used to determine the amount of data information based on information entropy and the probability of existence of a cycle.
9. The AMC over-standard cause mining and analysis system based on the active learning framework according to claim 7 is characterized in that: The sub-matrix extraction module comprises: A concentration acquisition unit, used for querying and analyzing the AMC concentration acquired in real time by the AMC monitor within the span; the AMC concentration contains a time tag; A concentration comparison unit, used to compare the AMC concentration with a preset concentration threshold; A position query unit, used to query the position of the AMC monitor, query the monitoring point closest to the AMC monitor, and query the row and column position corresponding to the monitoring point when any AMC concentration reaches a preset threshold; A position expansion unit, used to determine the extraction side length according to the AMC concentration, and expand the queried row and column positions based on the extraction side length to obtain the position and size of the submatrix; the extraction side length is proportional to the AMC concentration; The matrix interception unit is used to determine the number of sub-matrices based on a preset backtracking time, select the number of data matrices of sub-matrices in the data matrix sequence based on the current moment, intercept the sub-matrix in the data matrix according to the position and size of the sub-matrix, and arrange the sub-matrix based on the sequence order to obtain a sub-matrix sequence.
10. The AMC over-standard cause mining and analysis system based on the active learning framework according to claim 7 is characterized in that: The matrix simplification module comprises: An information amount query unit, used to query the data information amount of each row and column position in the sub-matrix sequence; A position marking unit, used to compare the amount of data information with a preset information amount threshold, and mark the row and column position when the amount of data information is less than the preset information amount threshold; A data regularization unit is used to delete the data at the marked row and column positions of each submatrix in the submatrix sequence, and regularize the remaining data; The sequence output unit is used to obtain a simplified matrix sequence after each sub-matrix is regularized.
Citation Information
Patent Citations
AMC monitoring data clustering method based on space-time correlation
CN118378113A
Environmental monitoring data anomaly detection method, medium and system
CN119807728A
Distributed concentrator state monitoring method and system
CN119848707A
Optimised approximation architectures and forecasting systems
WO2021181107A1
Method and apparatus for classifying nodes of a graph
WO2023087303A1