Compression method and device for operation and maintenance data of generic semiconductor industry, electronic equipment and medium

By identifying the temporal nature of pan-semiconductor industrial operation and maintenance data and splitting it into time and feature data groups, and adopting different compression strategies, the problem of insufficient adaptability of data types in existing technologies is solved, and more efficient data compression and storage optimization are achieved.

CN120811397APending Publication Date: 2025-10-17CLP JIUTIAN INTELLIGENT TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510385172.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

The compression methods in the prior art have limited adaptability to different types of data and are difficult to efficiently process multiple data forms simultaneously, which affects the compression rate.

Method used

By identifying the temporal nature of data, the pan-semiconductor industrial operation and maintenance data is split into time data groups and feature data groups, and different compression strategies are used to process them respectively, including Huffman coding and arithmetic coding.

Benefits of technology

It improves the balance between data compression efficiency and computing efficiency, and optimizes storage space requirements and data processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120811397A_ABST
    Figure CN120811397A_ABST
Patent Text Reader

Abstract

The invention discloses a compression method and device for generic semiconductor industry operation and maintenance data, electronic equipment and a medium, and relates to the technical field of data compression. The method comprises the following steps: acquiring to-be-compressed generic semiconductor industrial operation and maintenance data; identifying the time sequence of the generic semiconductor industrial operation and maintenance data to be compressed, and judging whether the generic semiconductor industrial operation and maintenance data is time sequence data or not according to the time sequence; and if the to-be-compressed generic semiconductor industrial operation and maintenance data is time series data, splitting the to-be-compressed generic semiconductor industrial operation and maintenance data into a time data set and a feature data set, and compressing the time data set and the feature data set respectively. According to the compression method, different compression strategies are adopted for different types of data, the compression efficiency of the data is improved, and the balance between the compression rate and the calculation efficiency is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data compression, and particularly relates to a compression method for generic semiconductor industry operation and maintenance data, a compression device for generic semiconductor industry operation and maintenance data, a computer readable storage medium and an electronic device. BACKGROUND

[0002] The compression method in the related art has limited adaptability to different types of data, and it is difficult to efficiently process multiple data forms at the same time, thereby affecting the compression rate. SUMMARY

[0003] The present application aims to at least partly solve one of the technical problems in the related art. To this end, the first object of the present application is to provide a compression method for generic semiconductor industry operation and maintenance data, which comprises: obtaining to-be-compressed generic semiconductor industry operation and maintenance data; identifying the time sequence of the to-be-compressed generic semiconductor industry operation and maintenance data, and determining whether it is time sequence data according to the time sequence; if the to-be-compressed generic semiconductor industry operation and maintenance data is time sequence data, then splitting the to-be-compressed generic semiconductor industry operation and maintenance data into a time data group and a feature data group, and compressing the time data group and the feature data group respectively. The compression method of the present application adopts different compression strategies for different types of data, improves the compression efficiency of the data, and optimizes the balance between the compression rate and the calculation efficiency.

[0004] The second object of the present application is to provide a compression device for generic semiconductor industry operation and maintenance data.

[0005] The third object of the present application is to provide a computer readable storage medium.

[0006] The fourth object of the present application is to provide an electronic device.

[0007] To achieve the above-mentioned objects, the first aspect of the present application provides a compression method for generic semiconductor industry operation and maintenance data, which comprises: obtaining to-be-compressed generic semiconductor industry operation and maintenance data; identifying the time sequence of the to-be-compressed generic semiconductor industry operation and maintenance data, and determining whether it is time sequence data according to the time sequence; if the to-be-compressed generic semiconductor industry operation and maintenance data is time sequence data, then splitting the to-be-compressed generic semiconductor industry operation and maintenance data into a time data group and a feature data group, and compressing the time data group and the feature data group respectively.

[0008] According to one embodiment of the present application, the above-mentioned method further comprises: if the to-be-compressed generic semiconductor industry operation and maintenance data is not time sequence data, then compressing the feature data group.

[0009] According to one embodiment of the present application, the feature data set comprises data and a feature vector corresponding to the data, and the compression of the feature data set comprises: inputting the feature vector into a preset feature prediction model to output probabilities of feature values in the feature vector; and compressing the feature vector according to the probabilities of the feature values and a first preset encoding strategy.

[0010] According to one embodiment of the present application, the compression of the feature vector according to the probabilities of the feature values and the first preset encoding strategy comprises: obtaining an encoding interval; determining a probability interval of each feature value according to the probabilities of the feature values; narrowing the encoding interval according to the probability intervals of the feature values to generate a target encoding interval; and generating a first compression result of the feature vector based on an upper limit value of the target encoding interval and a lower limit value of the target encoding interval.

[0011] According to one embodiment of the present application, the time data set comprises data and a time corresponding to the data, and the compression of the time data set comprises: obtaining a data sequence in a preset time period; calculating a frequency of occurrence of each data in the data sequence; and compressing the data sequence in the preset time period according to the frequency of occurrence of each data and a second preset encoding strategy.

[0012] According to one embodiment of the present application, the compression of the data sequence in the preset time period according to the frequency of occurrence of each data and the second preset encoding strategy comprises: constructing a Huffman tree according to each data and the frequency of occurrence of each data; determining an encoding of each data according to the Huffman tree; and splicing the encoding of each data in a data sequence of the data sequence to generate a second compression result of the data sequence in the preset time period.

[0013] According to one embodiment of the present application, before the feature vector is input into the preset feature prediction model, the preset feature prediction model is adjusted, comprising: constructing a preset external network layer based on the preset feature prediction model; training the preset external network layer based on a preset data set to determine a weight parameter of the preset external network layer; and adjusting the preset feature prediction model based on the weight parameter of the preset external network layer.

[0014] To achieve the above object, the second embodiment of the present application proposes a compression device for generic semiconductor industry operation and maintenance data, which comprises: an acquisition module configured to acquire to-be-compressed generic semiconductor industry operation and maintenance data; a judgment module configured to identify a time sequence of the to-be-compressed generic semiconductor industry operation and maintenance data, and determine whether the to-be-compressed generic semiconductor industry operation and maintenance data is time sequence data according to the time sequence; and a compression module configured to, if the to-be-compressed generic semiconductor industry operation and maintenance data is time sequence data, split the to-be-compressed generic semiconductor industry operation and maintenance data into a time data set and a feature data set, and compress the time data set and the feature data set respectively.

[0015] To achieve the above object, the third aspect of the present application provides a computer readable storage medium, which stores a compression program of the generic semiconductor industry operation and maintenance data, and the compression program of the generic semiconductor industry operation and maintenance data is executed by a processor to realize the compression method of the generic semiconductor industry operation and maintenance data.

[0016] To achieve the above object, the fourth aspect of the present application provides an electronic device, which comprises a memory, a processor and a compression program of the generic semiconductor industry operation and maintenance data stored in the memory and executable on the processor, and the processor executes the compression program of the generic semiconductor industry operation and maintenance data to realize the compression method of the generic semiconductor industry operation and maintenance data.

[0017] According to the compression method, device, electronic device and medium of the generic semiconductor industry operation and maintenance data, the compression method of the present application acquires the generic semiconductor industry operation and maintenance data to be compressed, identifies the time sequence of the generic semiconductor industry operation and maintenance data to be compressed, judges whether it is time sequence data according to the time sequence, splits the generic semiconductor industry operation and maintenance data to be compressed into time data groups and feature data groups if the generic semiconductor industry operation and maintenance data to be compressed is time sequence data, and compresses the time data groups and the feature data groups respectively. The compression method of the present application adopts different compression strategies for different types of data, improves the compression efficiency of the data, and optimizes the balance between the compression rate and the calculation efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 Flow chart of the compression method of the generic semiconductor industry operation and maintenance data according to some embodiments of the present application;

[0019] Figure 2 Data structure diagram according to some embodiments of the present application;

[0020] Figure 3 Large language model adjustment scheme diagram according to some embodiments of the present application;

[0021] Figure 4 Flow chart of the compression method of the generic semiconductor industry operation and maintenance data according to some other embodiments of the present application;

[0022] Figure 5 Data encoding diagram according to some embodiments of the present application;

[0023] Figure 6 Block diagram of the compression device of the generic semiconductor industry operation and maintenance data according to some embodiments of the present application;

[0024] Figure 7 Block diagram of the electronic device according to some embodiments of the present application. DETAILED DESCRIPTION

[0025] Embodiments of the present application are described below in detail, examples of which are shown in the accompanying drawings, wherein the same or similar notations represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by reference to the drawings are exemplary and are intended to explain the present application, and cannot be understood as a limitation of the present application.

[0026] A method, device, electronic equipment and medium for compressing generic semiconductor industry operation and maintenance data are described below in detail with reference to the accompanying drawings.

[0027] Figure 1 A flowchart of a method for compressing generic semiconductor industry operation and maintenance data according to some embodiments of the present application is shown in FIG. 1. Referring to FIG. 1, the method for compressing generic semiconductor industry operation and maintenance data according to some embodiments of the present application can include the following steps: Figure 1

[0028] S110, obtaining generic semiconductor industry operation and maintenance data to be compressed.

[0029] Specifically, the generic semiconductor industry operation and maintenance data to be compressed includes, but is not limited to, temperature, humidity, corpus, voice and device usage instructions of generic semiconductor industry equipment. Among them, the temperature and humidity of the generic semiconductor industry equipment can be collected by the temperature sensor and humidity sensor arranged on the equipment; the corpus and voice of the generic semiconductor industry equipment can be collected by the industrial Internet of Things gateway through the microphone array or voice sensor; the device usage instructions can be read through the cloud platform or local storage. It should be noted that the specific acquisition method is not limited here.

[0030] S120, identifying the time sequence of the generic semiconductor industry operation and maintenance data to be compressed, and determining whether it is time sequence data according to the time sequence.

[0031] Specifically, after obtaining the generic semiconductor industry operation and maintenance data to be compressed, feature engineering is used to process the generic semiconductor industry operation and maintenance data to be compressed to obtain the feature vector of the generic semiconductor industry operation and maintenance data to be compressed. For example, referring to Figure 2 , the generic semiconductor industry operation and maintenance data to be compressed includes temperature data and corpus data, and the temperature data and the corpus data are processed for feature extraction to determine the feature vector of the temperature data and the feature vector of the corpus data.

[0032] ​Further, after obtaining the to-be-compressed generic semiconductor industrial operation and maintenance data and the feature vectors thereof, it is necessary to consider the storage of the data rich in variability (such as data of changes in device temperature, humidity, etc. over time) and the feature vectors thereof, and other data (such as corpus, voice, device usage instructions, etc.) and the feature vectors thereof. Therefore, whether the to-be-compressed generic semiconductor industrial operation and maintenance data has a continuous timestamp field can be used to determine the timing of the to-be-compressed generic semiconductor industrial operation and maintenance data. For example, if the to-be-compressed generic semiconductor industrial operation and maintenance data has a continuous timestamp field, it is determined that the to-be-compressed generic semiconductor industrial operation and maintenance data is time-series data, such as temperature data. If the to-be-compressed generic semiconductor industrial operation and maintenance data does not have a continuous timestamp field, it is determined that the to-be-compressed generic semiconductor industrial operation and maintenance data is non-time-series data, such as corpus data.

[0033] In S130, if the to-be-compressed generic semiconductor industrial operation and maintenance data is time-series data, the to-be-compressed generic semiconductor industrial operation and maintenance data is split into a time data group and a feature data group, and the time data group and the feature data group are compressed respectively.

[0034] Specifically, referring to Figure 2 , the time-series data includes three parts of data, time corresponding to the data, and feature vectors corresponding to the data, and is rich in variability over time. If the table is directly stored, a large amount of repeated feature vector data (for example, the temperatures are all the same within 10 minutes, so their feature vectors are also the same) will be stored. Moreover, since the feature vectors usually learn small number arrays, such as 32-dimensional, 64-dimensional, and 128-dimensional double arrays, each double array occupies 64 bits, so each feature vector is a few kilobit array, and direct storage will consume a large amount of storage resources. In addition, the same coding strategy has limited adaptability to different types of data, and it is difficult to efficiently process multiple data forms (such as time-series data and feature vector data) at the same time.

[0035] Therefore, in the case that the to-be-compressed generic semiconductor industrial operation and maintenance data is time-series data, the to-be-compressed generic semiconductor industrial operation and maintenance data is split to obtain a time data group and a feature data group. The time data group includes data and time corresponding to the data, and the feature data group includes data and feature vectors corresponding to the data. The time data group and the feature data group are compressed respectively, for example, the feature data group is compressed using a first preset coding strategy, and the time data group is compressed using a second preset coding strategy.

[0036] The compression method of the present application uses different compression strategies for different types of data, improves the compression efficiency of the data, and optimizes the balance between compression rate and calculation efficiency.

[0037] In some embodiments, the method further comprises: if the to-be-compressed generic semiconductor industry operation and maintenance data is not time series data, compressing the feature data group.

[0038] Specifically, with continued reference to Figure 2 , the non-time series data includes two parts of data and a feature vector corresponding to the data, that is, the non-time series data only includes the feature data group. Therefore, in the case that the to-be-compressed generic semiconductor industry operation and maintenance data is not time series data, the feature data group can be directly compressed, for example, using a first preset encoding strategy to compress the feature data group.

[0039] In this way, by compressing the feature data group, the storage space requirement can be significantly reduced, the operation and maintenance cost can be reduced, and the data processing efficiency can be improved under the premise of ensuring data accuracy.

[0040] In some embodiments, the feature data group includes data and a feature vector corresponding to the data, and the compression of the feature data group includes: taking the feature vector as an input of a preset feature prediction model to output probabilities of feature values in the feature vector; and compressing the feature vector according to the probabilities of the feature values and a first preset encoding strategy.

[0041] In some embodiments, the compression of the feature vector according to the probabilities of the feature values and the first preset encoding strategy includes: obtaining an encoding interval; determining a probability interval of each feature value according to the probabilities of the feature values; narrowing the encoding interval according to the probability interval of each feature value to generate a target encoding interval; and generating a first compression result of the feature vector based on an upper limit value of the target encoding interval and a lower limit value of the target encoding interval. The preset feature prediction model can be a large language model, and the first preset encoding strategy can be arithmetic encoding, which is not limited here.

[0042] For example, assuming that the encoding interval is [0, 1), the feature data group is Y1[a, b, c], where the data is Y1 and the feature vector corresponding to the data is [a, b, c], the feature vector [a, b, c] is input into an LLM (Large Language Model) to predict the probabilities of the feature values a, b and c, and the output probabilities of the feature values a, b and c are P(a) = 0.5, P(b) = 0.3 and P(c) = 0.5 respectively. Then the probability interval of the feature value a is [0, 0.5), the probability interval of the feature value b is [0.5, 0.8), and the probability interval of the feature value b is [0.8, 1).

[0043] The feature value a is encoded as follows:

[0044] The probability interval corresponding to the feature value a is [0, 0.5), and the encoding interval [0, 1) is updated to [0, 0.5).

[0045] The feature value b is encoded as follows:

[0046] On the basis of the current encoding interval [0, 0.5), the probability interval is re-divided as follows:

[0047] The probability interval corresponding to the feature value a is [0, 0+0.5*0.5) = [0, 0.25);

[0048] The probability interval corresponding to the feature value b is [0.25, 0.25+0.5*0.3) = [0.25, 0.4);

[0049] The probability interval corresponding to the feature value c is [0.4, 0.4+0.5*0.2) = [0.4, 0.5);

[0050] The current feature value is b, and the probability interval corresponding to b is [0.25, 0.4), so the encoding interval [0, 0.5) is updated to [0.25, 0.4).

[0051] The feature value c is encoded as follows:

[0052] On the basis of the current encoding interval [0.25, 0.4), the probability interval is re-divided as follows:

[0053] The probability interval corresponding to the feature value a is [0.25, 0.25+0.15*0.5) = [0.25, 0.325);

[0054] The probability interval corresponding to the feature value b is [0.325, 0.325+0.15*0.3) = [0.325, 0.37);

[0055] The probability interval corresponding to the feature value c is [0.37, 0.37+0.15*0.2) = [0.37, 0.4);

[0056] The current feature value is c, and the probability interval corresponding to c is [0.37, 0.4), so the encoding interval [0.25, 0.4) is updated to [0.37, 0.4). That is, the target encoding interval is [0.37, 0.4).

[0057] Further, the upper limit value of the target encoding interval is 0.37, the lower limit value is 0.4, and the average of the upper limit value and the lower limit value is 0.385, so the arithmetic encoding result (the first compression result) of the feature vector [a, b, c] is 0.385. When storing, the binary result of 0.385 can be stored.

[0058] It should be noted that, in addition to taking the average of the upper limit value and the lower limit value as the encoding result, any value between the upper limit value and the lower limit value can be taken as the encoding result, such as 0.38.

[0059] Thus, using the preset feature prediction model and the first preset encoding strategy can effectively compress the feature data set, can effectively capture the complex relationship and information between the features, thereby reducing the occupation of the storage space and improving the compression rate; the first preset encoding strategy can be a lossless compression strategy, which can avoid information loss of the data to some extent, so as to restore the original data.

[0060] In some embodiments, the time data set includes data and time corresponding to the data, and compressing the time data set includes: obtaining a data sequence in a preset time period; calculating the occurrence frequency of each data in the data sequence; and compressing the data sequence in the preset time period according to the occurrence frequency of each data and a second preset encoding strategy.

[0061] In some embodiments, compressing the data sequence in the preset time period according to the occurrence frequency of each data and the second preset encoding strategy includes: constructing a Huffman tree according to each data and the occurrence frequency of each data; determining the encoding of each data according to the Huffman tree; and splicing the encoding of each data in the order of the data in the data sequence to generate a second compression result of the data sequence in the preset time period. The preset time period can be 1 minute, and the second preset encoding strategy can be Huffman encoding, which is not limited here.

[0062] For example, assume that the temperature data sequence in a preset time period (e.g., 1 minute) is y1 y2 y3 y1 y1 y2, where the occurrence frequency of y1 is 3, the occurrence frequency of y2 is 2, and the occurrence frequency of y3 is 1.

[0063] y3 and y2 are merged, and the occurrence frequency is 1+2=3;

[0064] The result of the previous step and y1 are merged, and the occurrence frequency is 3+3=6;

[0065] The Huffman tree is constructed according to each data and the occurrence frequency of each data as follows:

[0066]

[0067] According to the Huffman tree, the encoding of y1 is 0, the encoding of y2 is 10, and the encoding of y3 is 11, so the Huffman encoding result (second compression result) of y1 y2 y3 y1 y1 y2 is 01011100.

[0068] Thus, using the second preset encoding strategy can effectively compress the time data set, thereby reducing the occupation of the storage space; in addition, the second preset encoding strategy can be a lossless compression strategy, which can avoid information loss of the data to some extent, so as to restore the original data.

[0069] In some embodiments, the preset feature prediction model is adjusted before the feature vector is input into the preset feature prediction model, including: constructing a preset external network layer based on the preset feature prediction model; training the preset external network layer based on a preset data set to determine the weight parameters of the preset external network layer; and adjusting the preset feature prediction model based on the weight parameters of the preset external network layer.

[0070] Specifically, the preset feature prediction model can be adjusted before the feature vector is input into the preset feature prediction model to effectively improve the specialization level of the LLM in the field of generic semiconductor operation and maintenance, so that it can more accurately understand and process related languages and data, thereby improving operation and maintenance efficiency and prediction ability.

[0071] For example, referring to Figure 3 The preset feature prediction model (such as LLM) can be adjusted using LoRA (Learned Representation Augmentation, low-order adaptation of large language models). The weight parameters of the original LLM are fixed, a new preset external network layer is constructed externally, a preset data set is constructed using generic semiconductor industry operation and maintenance data, and the preset external network layer is trained based on the preset data set to determine the weight parameters of the preset external network layer, especially in the generic semiconductor operation and maintenance data and features.

[0072] Specifically, first, a data set containing generic semiconductor operation and maintenance related content is selected, which may include device running status, fault logs, operation manuals, technical documents, etc., as well as real-time data and historical data related to operation and maintenance, then the features related to operation and maintenance tasks can be extracted from the selected data through text processing, time series analysis, data cleaning and standardization preprocessing steps to construct a preset data set, to ensure data quality and model input adaptability.

[0073] Then, the preset external network layer is trained using the preset data set, and the weight parameters of the preset external network layer are adjusted according to the task requirements, so that it can better understand and process the language and content in the operation and maintenance field.

[0074] Then, the adjusted LLM is used to generate additional training samples. For example, the LLM is used to generate text fragments or fault reports similar to generic semiconductor operation and maintenance data; alternatively, new data samples are generated by interpolating or randomly disturbing existing data samples; alternatively, a part of the existing data is recombined or transformed to generate new data samples.

[0075] The additional training samples are then combined with the original pre-set dataset to form a larger, more diverse new dataset. This new dataset is then used to retrain the pre-set external network layer, adjusting its weight parameters and completing the adjustments to the pre-set feature prediction model. This allows the model to be exposed to a wider and more diverse range of operational data scenarios, thereby improving its adaptability and generalization capabilities in complex tasks and new domains.

[0076] Finally, the preset feature prediction model adjusted using the validation set is evaluated and optimized to verify the performance improvement of the model in pan-semiconductor operation and maintenance tasks, and the weight parameters or adjustment strategies of the preset external network layer are further adjusted based on the evaluation results.

[0077] In this way, the adjusted preset feature prediction model is used to capture the predicted distribution of data feature vectors, which improves the intelligence and adaptability of data compression, enabling the system to optimize compression according to the actual data characteristics, further reducing storage costs and improving data transmission efficiency.

[0078] It should be noted that this application can also implement efficient encryption and access control during data transmission and storage to ensure the security of critical operation and maintenance data, in line with the strict requirements of modern industry for data security and privacy protection.

[0079] As a specific example, see Figure 4 The method for compressing pan-semiconductor industrial operation and maintenance data in the embodiment of the present application may further include the following steps:

[0080] S401, pan-semiconductor industrial operation and maintenance data and its feature vector to be compressed.

[0081] S402: Determine whether the pan-semiconductor industry operation and maintenance data to be compressed is time series data. If yes, execute S403; otherwise, execute S407.

[0082] S403, splitting the compressed pan-semiconductor industrial operation and maintenance data.

[0083] S404, time data group.

[0084] Time data group such as Figure 5 Shown in the upper left box.

[0085] S405, feature data group.

[0086] Feature data sets such as Figure 5 Shown in the upper right box.

[0087] S406, Huffman coding.

[0088] The Huffman coding results of the time data group are as follows Figure 5 Shown in the lower left corner.

[0089] S407, large language model.

[0090] S408, predicted distribution of feature value.

[0091] S409, arithmetic coding.

[0092] The arithmetic coding result of the feature data set is as shown in the lower right corner. Figure 5

[0093] S410, memory storage.

[0094] Thus, the compression method of the present application combines the simple construction of Huffman coding and the high compression rate characteristics of arithmetic coding, and adopts appropriate compression methods for different data types, optimizing the balance between compression rate and computational efficiency. In particular, in the time data set, the application of Huffman coding can significantly improve the compression efficiency, while the combination of arithmetic coding and large language model can handle the complexity of the feature data set; through the combination of Huffman coding and arithmetic coding, efficient data compression can be achieved without sacrificing data quality, especially in the processing of time data sets and feature data sets, this method can effectively preserve important information and structure of the data, ensuring accurate restoration of the original data after decompression.

[0095] It should be noted that the results of Huffman coding and arithmetic coding can be stored separately and can be stored in different memories or storage units. When decoding, the Huffman coding result is decoded first for time series data, and then the arithmetic coding result of each data is matched and decoded to obtain the corresponding feature vector.

[0096] In summary, the compression method of the present application adopts different compression strategies for different types of data, such as using the second preset encoding strategy to compress the time data set, which takes advantage of the repetitive patterns and periodicity in time series data, and combines the preset feature prediction model and the first preset encoding strategy to compress the feature data set, which can effectively capture the complex relationships and information between features, improve the compression rate, and optimize the balance between compression rate and computational efficiency.

[0097] Corresponding to the above embodiments, the present application also proposes a compression device for general semiconductor industry operation and maintenance data.

[0098] Referring to Figure 6 , the compression device for general semiconductor industry operation and maintenance data 200 includes an acquisition module 210, a judgment module 220, and a compression module 230.

[0099] ​The acquisition module 210 is configured to acquire the to-be-compressed generic semiconductor industrial operation and maintenance data. The judgment module 220 is configured to identify the time sequence of the to-be-compressed generic semiconductor industrial operation and maintenance data, and determine whether the to-be-compressed generic semiconductor industrial operation and maintenance data is time sequence data according to the time sequence. The compression module 230 is configured to split the to-be-compressed generic semiconductor industrial operation and maintenance data into a time data group and a feature data group if the to-be-compressed generic semiconductor industrial operation and maintenance data is time sequence data, and compress the time data group and the feature data group respectively.

[0100] According to an embodiment of the present application, the compression module 230 is further configured to compress the feature data group if the to-be-compressed generic semiconductor industrial operation and maintenance data is not time sequence data.

[0101] According to an embodiment of the present application, the feature data group includes data and a feature vector corresponding to the data, and the compression module 230 is specifically configured to take the feature vector as an input of a preset feature prediction model to output probabilities of feature values in the feature vector, and compress the feature vector according to the probabilities of the feature values and a first preset encoding strategy.

[0102] According to an embodiment of the present application, the compression module 230 is specifically configured to acquire an encoding interval, determine a probability interval of each feature value according to the probabilities of the feature values, reduce the encoding interval according to the probability interval of each feature value to generate a target encoding interval, and generate a first compression result of the feature vector based on an upper limit value of the target encoding interval and a lower limit value of the target encoding interval.

[0103] According to an embodiment of the present application, the time data group includes data and a time corresponding to the data, and the compression module 230 is specifically configured to acquire a data sequence in a preset time period, calculate a frequency of occurrence of each data in the data sequence, and compress the data sequence in the preset time period according to the frequency of occurrence of each data and a second preset encoding strategy.

[0104] According to an embodiment of the present application, the compression module 230 is specifically configured to construct a Huffman tree according to each data and the frequency of occurrence of each data, determine an encoding of each data according to the Huffman tree, and splice the encoding of each data in a data sequence of the data sequence to generate a second compression result of the data sequence in the preset time period.

[0105] According to an embodiment of the present application, the preset feature prediction model is adjusted before the feature vector is input into the preset feature prediction model, including: constructing a preset external network layer based on the preset feature prediction model; training the preset external network layer based on a preset data set to determine a weight parameter of the preset external network layer; and adjusting the preset feature prediction model based on the weight parameter of the preset external network layer.

[0106] It should be noted that the above explanation of the embodiments and beneficial effects of the semiconductor industry operation and maintenance data compression method also applies to the semiconductor industry operation and maintenance data compression device of the embodiments of the present application. To avoid redundancy, it will not be expanded here.

[0107] Corresponding to the above embodiments, the present application also proposes a computer readable storage medium.

[0108] The computer readable storage medium of the present application has a semiconductor industry operation and maintenance data compression program stored thereon, which realizes the above-mentioned semiconductor industry operation and maintenance data compression method when executed by a processor.

[0109] It should be noted that the above explanation of the embodiments and beneficial effects of the semiconductor industry operation and maintenance data compression method also applies to the computer readable storage medium of the embodiments of the present application. To avoid redundancy, it will not be expanded here.

[0110] Corresponding to the above embodiments, the present application also proposes an electronic device.

[0111] Referring to Figure 7 The electronic device 300 of the present application includes a memory 310, a processor 320, and a semiconductor industry operation and maintenance data compression program stored on the memory 310 and executable on the processor 320. When the processor executes the semiconductor industry operation and maintenance data compression program, the above-mentioned semiconductor industry operation and maintenance data compression method is realized.

[0112] It should be noted that the above explanation of the embodiments and beneficial effects of the semiconductor industry operation and maintenance data compression method also applies to the electronic device of the embodiments of the present application. To avoid redundancy, it will not be expanded here.

[0113] It is to be appreciated that the above description and the examples that follow are intended to be illustrative only and that changes can be made to the description, either functionally or chronologically, as well as changes being made concerning the order of implementation. The logic and / or steps represented in the flow diagrams and / or described herein can be considered as a sequence of executable instructions, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, processor-containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions. For purposes of this specification, a "computer-readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-readable medium can be, for example, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system (or apparatus) or a propagation medium. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection (electronic) having one or more wires, a portable computer diskette (magnetic), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber (optical), and a portable compact disc read-only memory (CDROM). Note that the computer-readable medium can even be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example via the optical scanner of a device or device or via an intermediary, such as a facility bureau, then compiled, interpreted, or otherwise processed in a suitable manner, if necessary, and then stored in a computer storage medium.

[0114] It is to be understood that the various parts of the present application can be implemented in hardware, software, firmware or a combination thereof. In the above embodiments, various steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any of the following technologies, known in the art, or a combination thereof, can be used: discrete logic circuitry having logic gates for implementing logic functions upon an application of data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and so forth.

[0115] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Also, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in an appropriate manner.

[0116] In addition, the terms "first", "second", etc. are used only for the purpose of description, and should not be understood as indicating or implying relative importance or implying a number of the technical features indicated. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality of" is at least two, for example, two, three, etc., unless otherwise explicitly and specifically limited.

[0117] In the present application, unless otherwise explicitly specified and limited, the terms "mounting", "connecting", "connecting", "fixing" and the like should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or it can be integrated; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium; it can be the internal communication of two elements or the interaction relationship between two elements, unless otherwise explicitly limited. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0118] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application.

Claims

1. A method for compressing pan-semiconductor industrial operation and maintenance data, characterized in that: The method comprises: Obtain pan-semiconductor industrial operation and maintenance data to be compressed; Identifying the temporal nature of the pan-semiconductor industry operation and maintenance data to be compressed, and determining whether it is time series data based on the temporal nature; If the pan-semiconductor industry operation and maintenance data to be compressed is time series data, the pan-semiconductor industry operation and maintenance data to be compressed is split into a time data group and a feature data group, and the time data group and the feature data group are compressed respectively.

2. The method for compressing pan-semiconductor industrial operation and maintenance data according to claim 1, characterized in that: The method further comprises: If the pan-semiconductor industry operation and maintenance data to be compressed is not the time series data, the feature data group is compressed.

3. The method for compressing pan-semiconductor industrial operation and maintenance data according to claim 1 or 2, characterized in that: The feature data group includes data and feature vectors corresponding to the data, wherein compressing the feature data group includes: Using the feature vector as input to a preset feature prediction model to output the probability of each feature value in the feature vector; The feature vector is compressed according to the probability of each feature value and a first preset encoding strategy.

4. The method for compressing pan-semiconductor industrial operation and maintenance data according to claim 3, characterized in that: Compressing the feature vector according to the probabilities of the feature values ​​and a first preset encoding strategy includes: Get the encoding interval; Determine the probability interval of each eigenvalue according to the probability of each eigenvalue; Narrowing the coding interval according to the probability interval of each eigenvalue to generate a target coding interval; A first compression result of the feature vector is generated based on the upper limit value of the target encoding interval and the lower limit value of the target encoding interval.

5. The method for compressing pan-semiconductor industrial operation and maintenance data according to claim 1, characterized in that: The time data group includes data and time corresponding to the data, wherein compressing the time data group includes: Get the data sequence within the preset time period; Calculate the frequency of occurrence of each data in the data sequence; The data sequence within the preset time period is compressed according to the occurrence frequency of each data and a second preset coding strategy.

6. The method for compressing pan-semiconductor industrial operation and maintenance data according to claim 5, characterized in that: Compressing the data sequence within the preset time period according to the occurrence frequency of each data and a second preset coding strategy includes: Construct a Huffman tree according to each of the data and the occurrence frequency of each of the data; Determine the encoding of each data according to the Huffman tree; The codes of each data are spliced ​​according to the data order of the data sequence to generate a second compression result of the data sequence within the preset time period.

7. The method for compressing pan-semiconductor industrial operation and maintenance data according to claim 3, characterized in that: Before inputting the feature vector into the preset feature prediction model, adjusting the preset feature prediction model includes: Constructing a preset external network layer based on the preset feature prediction model; Training the preset external network layer based on a preset data set to determine weight parameters of the preset external network layer; The preset feature prediction model is adjusted based on the weight parameters of the preset external network layer.

8. A device for compressing pan-semiconductor industrial operation and maintenance data, characterized in that: The device comprises: An acquisition module is used to obtain the pan-semiconductor industrial operation and maintenance data to be compressed; a judgment module, configured to identify the temporal nature of the pan-semiconductor industry operation and maintenance data to be compressed, and determine whether the data is time series data based on the temporal nature; A compression module is used to split the pan-semiconductor industry operation and maintenance data to be compressed into a time data group and a feature data group if the pan-semiconductor industry operation and maintenance data to be compressed is time series data, and compress the time data group and the feature data group respectively.

9. A computer-readable storage medium, characterized in that A compression program for pan-semiconductor industry operation and maintenance data is stored thereon, and when the compression program for pan-semiconductor industry operation and maintenance data is executed by a processor, a method for compressing pan-semiconductor industry operation and maintenance data according to any one of claims 1-7 is implemented.

10. An electronic device, characterized in that: It includes a memory, a processor, and a compression program for pan-semiconductor industrial operation and maintenance data stored in the memory and runnable on the processor. When the processor executes the compression program for pan-semiconductor industrial operation and maintenance data, it implements the compression method for pan-semiconductor industrial operation and maintenance data according to any one of claims 1-7.

Citation Information

Cited By

  • Machine data management method and device and storage medium

    CN122019986A

  • A machine data management method, device and storage medium

    CN122019986B