Intelligent storage method for collected data of semiconductor equipment
By dynamically adjusting the dead zone threshold for each time period in the data collected by semiconductor equipment, combining the data change trend and abnormal characteristics of the period, the problem of low efficiency of dead zone compression method in the prior art is solved, and efficient data compression and guarantee of important data quality are achieved.
Patent Information
- Application Number
- CN202510429178.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-08
AI Technical Summary
In the data collected by existing semiconductor devices, the dead-band compression method uses a fixed dead-band threshold to cause loss of important data or low compression efficiency, affecting data analysis and storage efficiency.
By obtaining the ECG monitoring amplitude data collected by semiconductor devices, the data is divided into multiple time periods, the weight is determined based on the occurrence probability and distribution characteristics of the change trend of adjacent data, combined with the abnormal characteristics of the data period, the dead zone threshold of each time period is dynamically adjusted, and the dead zone compression method is used for efficient compression.
It realizes efficient compression of data collected by semiconductor equipment, while ensuring the quality of important data and improving data processing and storage efficiency.
Smart Images

Figure CN119937937A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a method for intelligently storing data collected by semiconductor equipment. Background Art
[0002] Semiconductor equipment refers to electronic components or devices made based on semiconductor materials. It is widely used in electronics, communications, computers, energy and other fields, and is a key component of the development of modern science and technology. Its development has made many innovations in modern science and technology possible, providing smaller, faster and more efficient electronic solutions, and promoting progress in the fields of information technology and communications. With the development of semiconductor technology, modern equipment will generate a large amount of data, so it is necessary to efficiently compress the data collected by semiconductor equipment to reduce the storage space occupied, and read and write data faster, thereby speeding up data processing and realizing intelligent storage of data collected by semiconductor equipment.
[0003] Dead zone compression is a technology commonly used in data compression, which can effectively reduce the storage space and transmission bandwidth requirements of data. The compression effect of this algorithm depends largely on the selection of an appropriate dead zone threshold. A larger dead zone threshold can improve compression efficiency but may cause compression distortion. A smaller dead zone threshold may result in a lower compression rate but maintain better data quality.
[0004] The existing problem is that the importance of data in different time periods in the data collected by semiconductor equipment to subsequent analysis and processing is different, and the dead zone compression method is a lossy compression. Therefore, when the conventional fixed dead zone threshold is large, it will lead to more loss of important data, affecting the accuracy of subsequent data analysis. When the dead zone threshold is small, it will lead to low compression efficiency, affecting the intelligent storage of data. Summary of the invention
[0005] The present invention provides a method for intelligently storing data collected by semiconductor equipment to solve the existing problems.
[0006] A semiconductor device data collection intelligent storage method of the present invention adopts the following technical solution: An embodiment of the present invention provides a method for intelligently storing data collected by a semiconductor device, the method comprising the following steps: Acquire the amplitude data of the electrocardiogram monitoring collected by the semiconductor device, record it as a time series data sequence, divide the time series data sequence into several time periods; take the difference between adjacent data in the time series data sequence as the change trend of adjacent data; Determine the corresponding weight of the changing trend of adjacent data according to the occurrence probability of the changing trend of adjacent data and the distribution characteristics of the changing trend of adjacent data in time; determine the trend abnormality characteristics of the data changing trend in the time period in the time series data sequence according to the corresponding weight of the changing trend of all adjacent data in the time period and the occurrence probability of the changing trend of all adjacent data in the time period; determine the initial abnormality degree of the data in the time period according to the trend abnormality characteristics of the data changing trend in the time period in the time series data sequence and the abnormal characteristics of the data cycle in the time period; Determine the precise abnormality level of the data in the time period based on the initial abnormality level of the data in the time period and the possibility of the existence of truly abnormal data in the time period; Determine the dead zone threshold corresponding to the time period according to the precise abnormality of the data in the time period and the difference of all adjacent data in the time period; obtain the dead zone range corresponding to the time period according to the dead zone threshold corresponding to the time period; According to the dead zone range corresponding to each time period, each time period is compressed using the dead zone compression method to obtain compressed semiconductor device acquisition data.
[0007] Furthermore, the step of determining the corresponding weight of the change trend of adjacent data according to the occurrence probability of the change trend of adjacent data and the distribution characteristics of the change trend of adjacent data in time includes the following specific steps: The change trends of all adjacent data in the time series data sequence constitute sequence B, and the change trends of all adjacent data in the time period constitute sequence C. According to the data value size of the change trend of adjacent data, the ordinal value corresponding to the data value of the change trend of each adjacent data in sequence C in sequence B is counted to obtain the ordinal value set; The difference between the previous data and the next data in the ordinal value set is subtracted to obtain the ordinal value difference set; If the number of data in the ordinal value set is not less than the preset number threshold, the temporal distribution characteristics of the change trend of adjacent data are determined according to the data variance in the ordinal value difference set; If the number of data in the ordinal value set is less than the preset number threshold, the temporal distribution characteristics of the change trend of adjacent data are set to the preset maximum temporal distribution characteristics of the change trend; The corresponding weight of the changing trend of the adjacent data is determined according to the distribution characteristics of the changing trend of the adjacent data in time and the occurrence probability of the changing trend of the adjacent data.
[0008] Furthermore, the abnormal characteristics of the data cycle within the time period include the following specific steps: Use the first-order derivative method to obtain the local extreme value points in the time period, and calculate the difference between the time point corresponding to the next local maximum value point in the time period and the time point corresponding to the previous local maximum value point in the time period in chronological order to obtain a first time difference value set; calculate the difference between the time point corresponding to the next local minimum value point in the time period and the time point corresponding to the previous local minimum value point in the time period in chronological order to obtain a second time difference value set; The abnormal characteristics of the data period in the time period are determined according to the number of local extreme value points in all time periods, the data variance in the first time difference value set, and the data variance in the second time difference value set.
[0009] Furthermore, the method of determining the abnormal characteristics of the data period in the time period according to the number of local extreme value points in all time periods and the data variance in the first time difference value set and the data variance in the second time difference value set includes the following specific steps: Calculate the data variance in the first time difference value set The variance of the data in the second time difference value set The mean of is taken as the first mean and the number of local extreme points in the time period is calculated. The mean of the number of local extreme points in all time periods The normalized value of the absolute value of the difference is taken as the first difference, and the product of the first mean and the first difference is recorded as the abnormal feature of the data period in the time period.
[0010] Furthermore, the initial abnormality degree of the data in the time period is determined according to the abnormal trend characteristics of the data change trend in the time period in the time series data sequence and the abnormal characteristics of the data cycle in the time period, and the specific steps include the following: Calculate 1 minus the probability of occurrence of the change trend of the i-th adjacent data in the time period The difference is taken as the first difference of the i-th adjacent data in the time period, the sum of the corresponding weights of the change trends of all adjacent data in the time period is calculated as the first sum, the ratio of the corresponding weight of the change trend of the i-th adjacent data in the time period to the first sum is recorded as the first ratio of the i-th adjacent data in the time period, the product of the first difference of the i-th adjacent data in the time period and the first ratio is taken as the first product of the i-th adjacent data in the time period, the sum of the first products of all adjacent data in the time period and the abnormal characteristics of the data cycle in the time period are The product of is recorded as the initial abnormality degree of the data in the time period.
[0011] Furthermore, the method of determining the precise abnormality degree of the data in the time period according to the initial abnormality degree of the data in the time period and the possibility of the existence of truly abnormal data in the time period includes the following specific steps: Divide the time period into small time periods according to the preset number of equal divisions, calculate the data mean in each small time period, and obtain the mean set; Use the first-order derivative method to obtain the local extreme value points in the mean set, and divide the mean set into several small mean sets according to the local extreme value points in the mean set; According to the weights corresponding to all small mean sets and the number of data in all small mean sets, the continuity of the same trend change of data in the mean set is determined; According to the continuity of the same trend change of the data in the mean set and the variance of the data in the mean set, the possibility of the existence of real abnormal data in the time period is determined; The precise abnormality degree of the data in the time period is determined according to the initial abnormality degree of the data in the time period and the possibility of the existence of real abnormal data in the time period, wherein the initial abnormality degree of the data in the time period and the possibility of the existence of real abnormal data in the time period are positively correlated.
[0012] Furthermore, the weight corresponding to the small mean value set includes the following specific steps: According to the absolute value of the difference between adjacent data in the small mean value set, a difference absolute value set is obtained; the number of local extreme value points in the difference absolute value set is obtained using the first-order derivative method; The weight corresponding to the small mean set is determined according to the number of data in the small mean set and the number of local extreme value points in the difference absolute value set corresponding to the small mean set.
[0013] Furthermore, the possibility of truly abnormal data existing in a time period is determined according to the continuity of the same trend change of the data in the mean set and the data variance in the mean set, and the specific steps include the following: Calculate the sum of the weights corresponding to all the small mean sets divided by the mean set as the second sum, take the ratio of the weight corresponding to the j-th small mean set divided by the mean set to the second sum as the second ratio of the j-th small mean set divided by the mean set, take the product of the second ratio of the j-th small mean set divided by the mean set and the number of data in the j-th small mean set divided by the mean set as the second product of the j-th small mean set divided by the mean set, take the ratio of the sum of the second products of all the small mean sets divided by the mean set to the preset number of equal divisions as the third ratio of the mean set, and calculate the difference between 1 and the third ratio of the mean set and the variance of the data in the mean set The normalized value of the product of is recorded as the third difference, and the sum of 1 and the third difference is taken as the possibility that there is real abnormal data in the time period.
[0014] Furthermore, the dead zone threshold corresponding to the time period is determined according to the precise abnormality degree of the data in the time period and the difference of all adjacent data in the time period, and the specific steps include the following: Determine the adaptive percentage threshold corresponding to the time period based on the preset percentage threshold interval and the precise abnormality of the data in the time period; The dead zone threshold corresponding to the time period is determined according to the adaptive percentage threshold corresponding to the time period and the average of the absolute values of all adjacent data differences in the time period.
[0015] Further, the dead zone threshold corresponding to the time period is determined according to the adaptive percentage threshold corresponding to the time period and the average of the absolute values of all adjacent data differences in the time period, and the specific steps include the following: Preset a percentage threshold interval as [ , ], and are the lower and upper boundaries of the preset percentage threshold interval respectively; calculate minus The product of the difference between the values of and the precise abnormality degree of the data in the time period is taken as the fourth product, and the fourth product is multiplied by The sum of the fourth sum and the absolute value of the difference between all adjacent data in the time period is calculated as The product of is recorded as the dead zone threshold corresponding to the time period.
[0016] The beneficial effects of the technical solution of the present invention are: In an embodiment of the present invention, the amplitude data of the electrocardiogram monitoring collected by the semiconductor device is obtained and recorded as a time series data sequence, the time series data sequence is divided into several time periods, the difference of adjacent data in the time series data sequence is recorded as the change trend of the adjacent data, the corresponding weight of the change trend of the adjacent data is determined according to the occurrence probability of the change trend of the adjacent data and the distribution characteristics of the change trend of the adjacent data in time, the trend abnormality characteristics of the data change trend in the time period in the time series data sequence are determined according to the corresponding weight of the change trend of all adjacent data in the time period and the occurrence probability of the change trend of all adjacent data in the time period, and the initial abnormality degree of the data in the time period is determined in combination with the abnormal characteristics of the data cycle in the time period. Considering that the abnormality of the data change trend may be caused by heart disease and physiological factors, it is necessary to further identify the real abnormal data caused by heart disease, so combined with the possibility of the existence of real abnormal data in the time period, the accurate abnormality degree of the data in the time period is determined, and then according to the difference of all adjacent data in the time period, the dead zone threshold corresponding to the time period is determined, and the dead zone range corresponding to the time period is obtained. According to the dead zone range corresponding to each time period, the dead zone compression method is used to efficiently compress each time period to obtain the compressed semiconductor device acquisition data. A smaller dead zone range is assigned to the time period with truly abnormal data in the time series data sequence to ensure the quality of important data after reconstruction, and a larger dead zone range is assigned to the time period with normal data in the time series data sequence to improve compression efficiency. In this way, efficient compression of semiconductor equipment acquisition data is achieved while ensuring the quality of important data after reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 The present invention is a flowchart of the steps of a method for intelligently storing data collected by a semiconductor device. DETAILED DESCRIPTION
[0019] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of the specific implementation method, structure, features and effects of a semiconductor device data acquisition intelligent storage method proposed by the present invention in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.
[0020] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0021] The specific scheme of the method for intelligently storing data collected by a semiconductor device provided by the present invention is described in detail below with reference to the accompanying drawings.
[0022] See also Figure 1 , which shows a flowchart of a method for intelligently storing data collected by a semiconductor device provided by an embodiment of the present invention, the method comprising the following steps: Step S001: Acquire amplitude data of ECG monitoring collected by semiconductor equipment, record it as a time series data sequence, divide the time series data sequence into several time periods; and use the difference between adjacent data in the time series data sequence as the change trend of adjacent data.
[0023] It is known that semiconductor devices are widely used in the medical and health fields for medical applications such as medical diagnosis, drug delivery, biosensing and monitoring. This embodiment takes the electrocardiogram data collected by semiconductor devices as an example to realize the intelligent storage of data collected by semiconductor devices.
[0024] The amplitude data of 24-hour ECG monitoring collected by semiconductor equipment is obtained and recorded as time series data sequence A. Since the dead zone compression method needs to segment the time series data first, this embodiment sets a fixed time step t, and divides the ECG time series data sequence A into time periods every t time periods in chronological order, and the fixed time step t is set to a value of 20 seconds.
[0025] Calculate the difference between the previous data and the next data in the time series data sequence A, record it as the change trend of adjacent data, and obtain the change trend sequence of adjacent data in the time series data sequence A. , where n represents the duration of the time series data sequence A, It represents the difference between the n-1th data and the nth data in the time series data sequence A. Take a time period divided by the time series data sequence A as an example to obtain the change trend sequence of adjacent data in this time period. , where t represents the duration of the time period, It represents the difference between the t-1th data and the tth data in this time period.
[0026] Step S002: Determine the corresponding weight of the changing trend of adjacent data according to the occurrence probability of the changing trend of adjacent data and the temporal distribution characteristics of the changing trend of adjacent data; determine the trend abnormality characteristics of the data changing trend within the time period in the time series data sequence according to the corresponding weight of the changing trend of all adjacent data within the time period and the occurrence probability of the changing trend of all adjacent data within the time period; determine the initial abnormality degree of the data within the time period according to the trend abnormality characteristics of the data changing trend within the time period in the time series data sequence and the abnormal characteristics of the data cycle within the time period.
[0027] Get the first data in sequence C For example, the data value size in the statistical sequence B is The ordinal value of each data, get the data The corresponding set of ordinal values , where q represents the size of the data value in sequence B The amount of data, Indicates that the data value size in sequence B is Then calculate the difference between the previous data and the next data in set D in turn to obtain the ordinal value difference set ,in Representing a collection The q-1th data Subtract the qth data The difference.
[0028] According to the above method, obtain the ordinal value set corresponding to each data in sequence C and the set of ordinal value differences .
[0029] It is known that ECG data presents obvious periodic characteristics in time. Under normal circumstances, the human heart rhythm is usually between 60 and 100 times per minute. Therefore, the duration of a complete cycle in the ECG data is generally between 0.6 and 1 second. Therefore, in this embodiment, there will be multiple complete cycles in each time period divided by the time series data sequence A.
[0030] Still taking the above-mentioned time period as an example, the local extreme value points in the time period are obtained using the method based on the first-order derivative, and the difference between the time point corresponding to the next local maximum value point and the time point corresponding to the previous local maximum value point is calculated in chronological order to obtain the first time difference value set, and then the data variance in the first time difference value set is calculated as , which represents the temporal distribution characteristics of the local maximum points within this time period.
[0031] Then, in chronological order, the difference between the time point corresponding to the next local minimum point and the time point corresponding to the previous local minimum point is calculated to obtain the second time difference value set, and then the data variance in the second time difference value set is calculated as , which represents the temporal distribution characteristics of the local minimum points within this time period.
[0032] Calculate the data variance in the first time difference value set The variance of the data in the second time difference value set The mean of is taken as the first mean and the number of local extreme points in the time period is calculated. The mean of the number of local extreme points in all time periods The normalized value of the absolute value of the difference is taken as the first difference, and the product of the first mean and the first difference is recorded as the abnormal feature of the data period in the time period.
[0033] Calculate 1 minus the probability of occurrence of the change trend of the i-th adjacent data in the time period The difference is taken as the first difference of the i-th adjacent data in the time period, the sum of the corresponding weights of the change trends of all adjacent data in the time period is calculated as the first sum, the ratio of the corresponding weight of the change trend of the i-th adjacent data in the time period to the first sum is recorded as the first ratio of the i-th adjacent data in the time period, the product of the first difference of the i-th adjacent data in the time period and the first ratio is taken as the first product of the i-th adjacent data in the time period, the sum of the first products of all adjacent data in the time period and the abnormal characteristics of the data cycle in the time period are The product of is recorded as the initial abnormality degree of the data in the time period.
[0034] It can be seen that the initial abnormality degree E of the data in this time period is: Where E is the initial abnormality of the data in this time period, F represents the abnormal characteristics of the data cycle in this time period, and t represents the length of this time period. Represents the i-th data in sequence C The corresponding weight of Represents the i-th data in sequence C The probability of occurrence, The process of obtaining is: the i-th data in sequence C The corresponding set of ordinal values The number of data in the sequence is divided by the number of data in the sequence B n-1. S represents the number of local extreme points in this time period (local extreme points include local maximum points and local minimum points), Represents the mean value of the number of local extreme points in all time periods divided by the time series data sequence A. is the data variance in the first time difference value set, is the data variance in the second time difference value set. Represents the i-th data in sequence C The corresponding set of ordinal values The number of data in is the set quantity threshold, Represents the i-th data in sequence C The corresponding set The variance of the data in It is the maximum distribution characteristic of the preset change trend in time. is a linear normalization function that normalizes the data value to the interval [0,1]. =1, This is described as an example, and other values may be set in other implementation modes, which are not limited in this embodiment.
[0035] It should be noted that: Under normal circumstances, the duration of the ECG data cycle and the trend of ECG data changes are similar for most of the day. Indicates the number of standard data cycles in time t. When the cycle length is disordered, it means that the ECG data is abnormal. Indicates the degree of change in the duration of each data cycle within this time period. The larger the value, the more drastic the change in the duration of each data cycle within this time period. It represents the difference between the number of data cycles in this time period and the number of standard data cycles, so it is normalized for The product of the two, F, represents the abnormal characteristics of the data cycle in this time period. The smaller it is, the smaller the number of times the change trend of the i-th adjacent data in the time period appears in the time series data sequence A, the greater the possibility of it being abnormal, and the more irregular the distribution of the change trend in time, the greater the degree of abnormality. Therefore, when the i-th data in sequence C The corresponding set of ordinal values The amount of data in Less than the set quantity threshold When , it means that the change trend occurs less frequently and its temporal distribution characteristics cannot be analyzed. Then let the temporal distribution characteristics of the change trend of the i-th adjacent data in this time period be the maximum temporal distribution characteristics of the preset change trend. . When the i-th data value in sequence C The corresponding set of ordinal values The amount of data in Not less than the set quantity threshold hour, The larger the value is, the more irregular the distribution of the change trend in time is. Let the distribution characteristics of the change trend of the i-th adjacent data in this time period in time be Therefore, the temporal distribution characteristics of the change trend of the i-th adjacent data in this time period are used as the inverse value of its occurrence probability The product of the two represents the corresponding weight of the change trend of the i-th adjacent data in this time period. , and then Take a weighted average Obtain the trend anomaly characteristics representing the change trend of the i-th adjacent data in the time period in the time series data sequence A. is the adjusted value of F, and obtains the initial abnormality of the data in this time period.
[0036] Step S003: Determine the precise abnormality level of the data in the time period according to the initial abnormality level of the data in the time period and the possibility of the existence of truly abnormal data in the time period.
[0037] Since the abnormalities in the ECG data cycle length and ECG data change trend may be caused by heart disease and physiological factors, it is necessary to further identify the real abnormal data caused by heart disease. Known heart diseases, such as arrhythmia, ischemic heart disease, myocarditis, etc., can cause sudden changes or fluctuations in heart rate, manifested as irregular heart rate rhythms, while physiological factors are heart rate changes caused by human activities. There is a rule that as the intensity of the activity increases, the heart rate also increases accordingly, and the heart rate changes gradually increase or slow down.
[0038] Still taking the above-mentioned time period as an example, it is known that the fixed time step t is set to 20 seconds, and the duration of a complete cycle in the ECG data is generally between 0.6 and 1 second, so the time period is divided into r small time periods. The value of the number of equal divisions r set in this embodiment is 20, and other implementations can be set to other values, which are not limited in this embodiment. Then calculate the mean of the data in each small time period to obtain the mean set ,in It represents the mean of the data in the rth small time period that is equally divided into the time period, and then uses the method based on the first-order derivative to obtain the local extreme points in the set G. The set G is divided into several small mean sets using these local extreme points as the dividing points, and the data trend changes in each small mean set are the same.
[0039] Calculate the sum of the weights corresponding to all the small mean sets divided by the mean set as the second sum, take the ratio of the weight corresponding to the j-th small mean set divided by the mean set to the second sum as the second ratio of the j-th small mean set divided by the mean set, take the product of the second ratio of the j-th small mean set divided by the mean set and the number of data in the j-th small mean set divided by the mean set as the second product of the j-th small mean set divided by the mean set, take the ratio of the sum of the second products of all the small mean sets divided by the mean set to the preset number of equal divisions as the third ratio of the mean set, and calculate the difference between 1 and the third ratio of the mean set and the variance of the data in the mean set The normalized value of the product of is recorded as the third difference, and the sum of 1 and the third difference is taken as the possibility that there is real abnormal data in the time period.
[0040] The precise abnormality of the data in this time period for: in is the precise abnormality of the data in this time period, E is the initial abnormality of the data in this time period, and H is the possibility of the existence of truly abnormal data in this time period. is the mean set The variance of the data in is the mean set The number of small mean sets divided, r is the set number of equal divisions, is the mean set The number of data in the jth small mean set divided, is the mean set The weight corresponding to the j-th small mean set divided, The process of obtaining is: calculate the mean set in sequence The absolute value of the difference between adjacent data in the jth small mean value set is divided, and the absolute value set of the difference is obtained. The number of local extreme points in the absolute value set of the difference is obtained using the method based on the first-order derivative. It should be noted that the definition of a local extreme point is relative to its adjacent data points. It is impossible to determine whether a data point is a local extreme point. There must be at least three consecutive data points. Therefore, in this embodiment, when the mean set The number of data in the partitioned small mean set If three are not satisfied, it means that there is no local extreme point in the small mean set, that is, the number of local extreme points corresponding to the small mean set is 0. is an exponential function with a natural constant as the base, and u is an adjustment value of the exponential function. In this embodiment, the value of u is set to 0.1. In other implementations, it can be set to other values, which are not limited in this embodiment. It is a linear normalization function that normalizes the data values to the interval [0,1].
[0041] What needs to be explained is: The larger the value is, the greater the overall trend of the data in this time period is, indicating that the data in this time period is more likely to be abnormal. At this time, it is necessary to further analyze whether the abnormal data in this time period is real abnormal data caused by heart disease or pseudo-abnormal data caused by physiological factors. It is known that heart disease can cause sudden fluctuations in ECG data, which appear as irregular fluctuations, while physiological factors can cause ECG data to fluctuate continuously or slow down. Therefore, when the data fluctuations in set G are more and more irregular, the possibility of real abnormal data in this time period is greater, that is, It indicates the number of data in the same trend change in set G. The larger its value is, the greater the possibility that the data abnormality is caused by physiological factors. The data in the same trend change may be in three states: uniform increase or decrease, increasing increase or decrease, and decreasing increase or decrease. It indicates the number of transitions of these three states in the same trend change in set G. The smaller its value is, the more regular the same trend change in set G is, and the greater the possibility that the data anomaly is caused by physiological factors. Therefore, the inverse normalization is used. for The product of the two represents the mean set The weight corresponding to the jth small set divided. Take a weighted average Indicates the continuity of the same trend change of the data in set G. The larger its value, the greater the possibility that the data abnormality is caused by physiological factors, that is, the smaller the possibility that the data abnormality is caused by disease factors. The normalized inverse value is calculated from this and The normalized value of the product plus 1, H, indicates the possibility of the existence of truly abnormal data in this time period. So far, H is used as The normalized value of the product of the two represents the precise abnormality of the data in that time period.
[0042] According to the above method, the precise abnormality degree of the data in each time period divided by the time series data sequence A is obtained.
[0043] Step S004: Determine the dead zone threshold corresponding to the time period according to the precise abnormality degree of the data in the time period and the difference of all adjacent data in the time period; obtain the dead zone range corresponding to the time period according to the dead zone threshold corresponding to the time period.
[0044] In the known dead zone compression method, the dead zone threshold is usually set according to the percentage of the mean of the absolute values of the differences between adjacent data in the time series data sequence. The percentage threshold interval set in this embodiment is [ , ],by , This is described as an example, and other values may be set in other implementation modes, which are not limited in this embodiment.
[0045] calculate minus The product of the difference between the values of and the precise abnormality degree of the data in the time period is taken as the fourth product, and the fourth product is multiplied by The sum of the fourth sum and the absolute value of the difference between all adjacent data in the time period is calculated as The product of is recorded as the dead zone threshold corresponding to the time period.
[0046] Still taking the above-mentioned time period as an example, the dead zone threshold Q corresponding to the time period is: Where Q is the dead zone threshold corresponding to this time period, is the precise abnormality of the data in this time period, then is the adaptive percentage threshold corresponding to the time period, is the mean of the absolute values of all adjacent data differences within this time period. and are the lower and upper boundaries of the preset percentage threshold interval respectively. Thus, the dead zone range [-Q,Q] corresponding to the time period is obtained.
[0047] According to the above method, the dead zone range corresponding to each time period divided by the time series data sequence A is obtained.
[0048] Step S005: according to the dead zone range corresponding to each time period, use the dead zone compression method to compress each time period to obtain compressed semiconductor device acquisition data.
[0049] The dead zone range corresponding to each time period divided by the known time series data sequence A is used, and the dead zone compression method is used to efficiently compress each time period according to the dead zone range corresponding to each time period, so as to complete the efficient compression processing of the data collected by the semiconductor device. Among them, the dead zone compression method is a well-known technology, and the specific method is not introduced here.
[0050] The efficiently compressed data collected by semiconductor devices is stored in appropriate storage media, such as semiconductor memory, hard disk, cloud storage, etc. Then data partitioning, index management, backup strategy, and data integrity verification are performed to improve data access efficiency and reliability, thereby completing the intelligent storage of semiconductor device collected data.
[0051] So far, the present invention is completed.
[0052] In summary, in an embodiment of the present invention, by segmenting the time series data sequence, the initial abnormal degree of the data in each time period is determined according to the abnormal trend characteristics of the data change trend in each time period in the time series data sequence and the abnormal characteristics of the data cycle in each time period. Then, each time period is divided equally, and the possibility of abnormal data in each time period and the continuity of the same trend change in the mean data are determined according to the difference and trend change between the data means in the equally divided small time periods, thereby obtaining the possibility of the existence of truly abnormal data in each time period, thereby obtaining the precise abnormal degree of the data in each time period. Then, the dead zone threshold corresponding to each time period is calculated, and the dead zone range corresponding to each time period is obtained, thereby using the dead zone compression method to perform efficient compression processing on each time period, and a smaller dead zone range is assigned to the time period with truly abnormal data in the time series data sequence to ensure the quality of important data after reconstruction, and a larger dead zone range is assigned to the time period of normal data in the time series data sequence to improve compression efficiency. The present invention can realize efficient compression of semiconductor device acquisition data while ensuring the quality of important data after reconstruction.
[0053] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for intelligently storing data collected by semiconductor equipment, characterized in that: The method comprises the following steps: Acquire the amplitude data of the electrocardiogram monitoring collected by the semiconductor device, record it as a time series data sequence, divide the time series data sequence into several time periods; take the difference between adjacent data in the time series data sequence as the change trend of adjacent data; Determine the corresponding weight of the changing trend of adjacent data according to the occurrence probability of the changing trend of adjacent data and the distribution characteristics of the changing trend of adjacent data in time; determine the trend abnormality characteristics of the data changing trend in the time period in the time series data sequence according to the corresponding weight of the changing trend of all adjacent data in the time period and the occurrence probability of the changing trend of all adjacent data in the time period; determine the initial abnormality degree of the data in the time period according to the trend abnormality characteristics of the data changing trend in the time period in the time series data sequence and the abnormal characteristics of the data cycle in the time period; Determine the precise abnormality level of the data in the time period based on the initial abnormality level of the data in the time period and the possibility of the existence of truly abnormal data in the time period; Determine the dead zone threshold corresponding to the time period according to the precise abnormality of the data in the time period and the difference of all adjacent data in the time period; obtain the dead zone range corresponding to the time period according to the dead zone threshold corresponding to the time period; According to the dead zone range corresponding to each time period, each time period is compressed using the dead zone compression method to obtain compressed semiconductor device acquisition data.
2. According to claim 1, a semiconductor device data collection intelligent storage method is characterized in that: The method of determining the corresponding weight of the change trend of adjacent data according to the occurrence probability of the change trend of adjacent data and the distribution characteristics of the change trend of adjacent data in time includes the following specific steps: The change trends of all adjacent data in the time series data sequence constitute sequence B, and the change trends of all adjacent data in the time period constitute sequence C. According to the data value size of the change trend of adjacent data, the ordinal value corresponding to the data value of the change trend of each adjacent data in sequence C in sequence B is counted to obtain the ordinal value set; The difference between the previous data and the next data in the ordinal value set is subtracted to obtain the ordinal value difference set; If the number of data in the ordinal value set is not less than the preset number threshold, the temporal distribution characteristics of the change trend of adjacent data are determined according to the data variance in the ordinal value difference set; If the number of data in the ordinal value set is less than the preset number threshold, the temporal distribution characteristics of the change trend of adjacent data are set to the preset maximum temporal distribution characteristics of the change trend; The corresponding weight of the changing trend of the adjacent data is determined according to the distribution characteristics of the changing trend of the adjacent data in time and the occurrence probability of the changing trend of the adjacent data.
3. According to claim 1, a semiconductor device data collection intelligent storage method is characterized in that: The abnormal characteristics of the data cycle within the time period include the following specific steps: Use the first-order derivative method to obtain the local extreme value points in the time period, and calculate the difference between the time point corresponding to the next local maximum value point in the time period and the time point corresponding to the previous local maximum value point in the time period in chronological order to obtain a first time difference value set; calculate the difference between the time point corresponding to the next local minimum value point in the time period and the time point corresponding to the previous local minimum value point in the time period in chronological order to obtain a second time difference value set; The abnormal characteristics of the data period in the time period are determined according to the number of local extreme value points in all time periods, the data variance in the first time difference value set, and the data variance in the second time difference value set.
4. A semiconductor device data collection intelligent storage method according to claim 3, characterized in that: The specific steps of determining the abnormal characteristics of the data cycle in the time period according to the number of local extreme value points in all time periods and the data variance in the first time difference value set and the data variance in the second time difference value set are as follows: Calculate the data variance in the first time difference value set The variance of the data in the second time difference value set The mean of is taken as the first mean and the number of local extreme points in the time period is calculated. The mean of the number of local extreme points in all time periods The normalized value of the absolute value of the difference is taken as the first difference, and the product of the first mean and the first difference is recorded as the abnormal feature of the data period in the time period.
5. The method for intelligently storing collected data of semiconductor equipment according to claim 1, characterized in that: The method of determining the initial abnormality degree of the data in the time period according to the abnormal trend characteristics of the data change trend in the time period in the time series data sequence and the abnormal characteristics of the data cycle in the time period includes the following specific steps: Calculate 1 minus the probability of occurrence of the change trend of the i-th adjacent data in the time period The difference is taken as the first difference of the i-th adjacent data in the time period, the sum of the corresponding weights of the change trends of all adjacent data in the time period is calculated as the first sum, the ratio of the corresponding weight of the change trend of the i-th adjacent data in the time period to the first sum is recorded as the first ratio of the i-th adjacent data in the time period, the product of the first difference of the i-th adjacent data in the time period and the first ratio is taken as the first product of the i-th adjacent data in the time period, the sum of the first products of all adjacent data in the time period and the abnormal characteristics of the data cycle in the time period are The product of is recorded as the initial abnormality degree of the data in the time period.
6. The method for intelligently storing collected data of semiconductor equipment according to claim 1, characterized in that: The specific steps of determining the precise abnormality degree of the data in the time period according to the initial abnormality degree of the data in the time period and the possibility of the existence of truly abnormal data in the time period are as follows: Divide the time period into small time periods according to the preset number of equal divisions, calculate the data mean in each small time period, and obtain the mean set; Use the first-order derivative method to obtain the local extreme value points in the mean set, and divide the mean set into several small mean sets according to the local extreme value points in the mean set; According to the weights corresponding to all small mean sets and the number of data in all small mean sets, the continuity of the same trend change of data in the mean set is determined; According to the continuity of the same trend change of the data in the mean set and the variance of the data in the mean set, the possibility of the existence of real abnormal data in the time period is determined; The precise abnormality degree of the data in the time period is determined according to the initial abnormality degree of the data in the time period and the possibility of the existence of real abnormal data in the time period, wherein the initial abnormality degree of the data in the time period and the possibility of the existence of real abnormal data in the time period are positively correlated.
7. A semiconductor device data collection intelligent storage method according to claim 6, characterized in that: The weight corresponding to the small mean value set includes the following specific steps: According to the absolute value of the difference between adjacent data in the small mean value set, a difference absolute value set is obtained; the number of local extreme value points in the difference absolute value set is obtained using the first-order derivative method; The weight corresponding to the small mean set is determined according to the number of data in the small mean set and the number of local extreme value points in the difference absolute value set corresponding to the small mean set.
8. A semiconductor device data collection intelligent storage method according to claim 6, characterized in that: The method of determining the possibility of the existence of truly abnormal data in a time period according to the continuity of the same trend change of the data in the mean set and the data variance in the mean set includes the following specific steps: Calculate the sum of the weights corresponding to all the small mean sets divided by the mean set as the second sum, take the ratio of the weight corresponding to the j-th small mean set divided by the mean set to the second sum as the second ratio of the j-th small mean set divided by the mean set, take the product of the second ratio of the j-th small mean set divided by the mean set and the number of data in the j-th small mean set divided by the mean set as the second product of the j-th small mean set divided by the mean set, take the ratio of the sum of the second products of all the small mean sets divided by the mean set to the preset number of equal divisions as the third ratio of the mean set, and calculate the difference between 1 and the third ratio of the mean set and the variance of the data in the mean set The normalized value of the product of is recorded as the third difference, and the sum of 1 and the third difference is taken as the possibility that there is real abnormal data in the time period.
9. The method for intelligently storing collected data of semiconductor equipment according to claim 1, characterized in that: The specific steps of determining the dead zone threshold corresponding to the time period according to the precise abnormality degree of the data in the time period and the difference of all adjacent data in the time period are as follows: Determine the adaptive percentage threshold corresponding to the time period based on the preset percentage threshold interval and the precise abnormality of the data in the time period; The dead zone threshold corresponding to the time period is determined according to the adaptive percentage threshold corresponding to the time period and the average of the absolute values of all adjacent data differences in the time period.
10. A semiconductor device data collection intelligent storage method according to claim 9, characterized in that: The dead zone threshold corresponding to the time period is determined according to the adaptive percentage threshold corresponding to the time period and the average of the absolute values of all adjacent data differences in the time period, and the specific steps include the following: Preset a percentage threshold interval as [ , ], and are the lower and upper boundaries of the preset percentage threshold interval respectively; calculate minus The product of the difference between the values of and the precise abnormality degree of the data in the time period is taken as the fourth product, and the fourth product is multiplied by The sum of the fourth sum and the absolute value of the difference between all adjacent data in the time period is calculated as The product of is recorded as the dead zone threshold corresponding to the time period.
Citation Information
Patent Citations
Data compressing method
CN104331495A
Method for testing wear resistance of self-leveling mortar
CN117992897A
Highly sensitive inertial mouse
EP1717669A1
Method and System of Using Inferential Measurements for Abnormal Event Detection in Continuous Industrial Processes
US20120330631A1
Data compression method, data restoration method and device
WO2020051895A1
Cited By
Reconnaissance data archiving and backup method and system for geotechnical engineering and medium
CN120179613A
Electromagnetic fingerprint real-time monitoring and abnormity identification system during operation of industrial equipment
CN120832623A
Electric energy meter data intelligent storage method based on cloud platform
CN120849366A
Machine data management method and device and storage medium
CN122019986A
A machine data management method, device and storage medium
CN122019986B