Design data management method of novel server
By setting up multiple index tables and segmented control functions in the Hoffman encoding algorithm, the problem of slow decoding speed of traditional Hoffman encoding algorithm is solved, and a more efficient decoding process and lower storage cost are achieved.
Patent Information
- Application Number
- CN202510559346.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The traditional Hoffman encoding algorithm needs to match a large number of encodings during the decoding process, resulting in slow decoding speed and unable to effectively reduce storage costs.
By setting up multiple index tables, the number of encodings in each index table is small, which reduces the number of matches when positioning matching encodings, and searches and matches in the index table corresponding to each segment to improve decoding efficiency.
Effectively improve decoding efficiency, reduce the storage amount of duplicate data in the index table, reduce unnecessary storage waste, and reduce the number of data matching times.
Smart Images

Figure CN120074540A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing. More specifically, the present invention relates to a method for managing design data of a new type of server. Background Art
[0002] The design data of the new type of server contains a large number of configuration files, log files, virtual machine images, container data, etc. These data have a large volume and a high redundancy, and the storage cost required for directly storing these data is relatively large. Therefore, in order to reduce the storage cost, it is necessary to compress the design data of the new type of server.
[0003] As a commonly used coding compression algorithm, the Huffman coding compression algorithm has a good compression effect. The decoding process of this algorithm includes: first, obtaining the matching code by matching the code in the coding sequence with the code in the index table, and then performing decoding processing according to the index relationship between the matching code in the index table and the data. The traditional Huffman coding algorithm places all data and the corresponding codes in one index table. Therefore, there will be a large number of codes in one index table, and a large amount of matching work is required to locate the corresponding matching code among a large number of codes, which will greatly reduce the decoding speed. How to reduce the matching workload and thus improve the decoding efficiency has become the research focus of this solution.
[0004] The patent document with the authorization announcement number CN117811588B discloses a method, system, device and readable storage medium for log compression access based on Huffman coding and LZ77. The method in this patent document only compresses the template code using the Huffman coding algorithm and does not involve the content of how to improve the decoding efficiency of the Huffman coding algorithm. Therefore, the method in this patent document cannot solve the technical problems in this solution. Summary of the Invention
[0005] To solve the problem of how to reduce the matching workload and thus improve the decoding efficiency, the present invention proposes a method for managing design data of a new type of server, and this method includes the following steps: Obtain the sequence composed of all design data of the new type of server and denote it as the design data sequence; Use the Huffman coding algorithm to perform coding processing on the data in the design data sequence to obtain a coding sequence; Set a segmentation control function , divide the design data sequence into several feasible data segments, place the index relationship between the design data and the code in a feasible data segment in an independent index table to obtain a feasible local index table, and place the index relationship between all design data and the code in the design data sequence in an index table to obtain an overall index table. represents the data set composed of the data in the i-th feasible data segment. , respectively represent the data sets formed by the data in the left adjacent feasible data segment and the right adjacent feasible data segment of the i-th feasible data segment, represents the intersection symbol, represents the number of feasible data segments, represents the number of matches required to locate the index relationship between all design data and coding by using the feasible local index table, represents the number of matches required to locate the index relationship between all design data and coding by using the overall index table; The design data sequence is segmented and controlled by using a segmented control function to obtain a number of target data segments; the target local index table is set by using the target data segments to implement the decoding process.
[0006] In the present invention, by setting multiple index tables, the number of codings in each index table is small, thereby reducing the number of matches when locating and matching codings, and improving the decoding efficiency; further, by setting an independent index table for each segment of the design sequence, when decoding the coding corresponding to each segment, only search and match in the index table corresponding to the segment, effectively improving the decoding efficiency; further, when performing segmented control, the intersection data volume between segments and adjacent segments is considered, so that the intersection data between segments is small, effectively reducing the storage volume of duplicate data in the index table and reducing unnecessary storage waste; further, when performing segmented control, the number of data matches in the decoding process is considered, so that the number of data matches is small, effectively improving the decoding efficiency.
[0007] Preferably, the splitting the design data sequence into a number of feasible data segments includes: Obtaining the number of segments; Randomly select M - 1 segment positions in the design data sequence, and split the design data sequence at the segment positions to obtain a number of data segments, denoted as feasible data segments, where M represents the number of segments.
[0008] In the present invention, the design sequence is segmented by randomly selecting segment positions, so as to obtain segments more comprehensively and provide a basis for accurately obtaining target segments subsequently.
[0009] Preferably, the obtaining the number of segments includes: Analyze the data in the design data sequence by using the elbow method to obtain the number of segments.
[0010] In the present invention, the number of segments is obtained by using the elbow method, so as to more accurately grasp the data type situation in the design data sequence and provide a basis for realizing segment differentiation subsequently.
[0011] Preferably, the feasible local index table further includes an index table positioning flag value, and the index table positioning flag value is the starting position of the encoding of the first design data in the corresponding feasible data segment of the feasible local index table in the encoding sequence.
[0012] By setting the index table positioning flag value in the present invention, the encoding regions corresponding to each index table can be located more quickly, providing a basis for fast decoding.
[0013] Preferably, the number of matching times required to locate the index relationship between all design data and encoding by using the feasible local index table includes: Locate the starting encoding position corresponding to the feasible local index table in the encoding sequence according to the index table positioning flag value in the feasible local index table; take the encoding between the starting decoding position of the feasible local index table and the starting encoding position of the next feasible local index table as the analysis encoding of the feasible local index table; perform matching processing on the analysis encoding and the encoding in the feasible local index table in sequence until the matching encoding of the analysis encoding is located, obtain the number of matching times required to locate the matching encoding, take the sum of the matching times of all analysis encodings as the local matching times of the feasible local index table; take the sum of the local matching times of all feasible local index tables as the number of matching times required to locate the index relationship between all design data and encoding by using the feasible local index table.
[0014] Preferably, the number of matching times required to locate the index relationship between all design data and encoding by using the overall index table includes: Record the encoding in the encoding sequence as the research encoding, perform matching processing on the research encoding and the encoding in the overall index table in sequence until the matching encoding of the research encoding is located, obtain the number of matching times required to locate the matching encoding, take the sum of the matching times of all research encodings as the number of matching times required to locate the index relationship between all design data and encoding by using the overall index table.
[0015] Preferably, the process of using the segmentation control function to segment the design data sequence to obtain several target data segments includes: Take the feasible data segment obtained when the segmentation control function reaches the maximum value as the target data segment.
[0016] Preferably, the process of setting the target local index table by using the target data segment includes: Take the feasible local index table obtained from the target data segment as the target local index table.
[0017] The target local index table obtained by the present invention in this way can not only improve the decoding efficiency but also reduce the storage amount of duplicate data.
[0018] Preferably, the implementation of the decoding process includes: Based on the index relationship between the codes and the design data in the target local index table, obtain the design data corresponding to each code in the code sequence and perform decoding processing.
[0019] Preferably, the encoding the data in the design data sequence by using the Huffman coding algorithm to obtain a code sequence includes: Statistically analyze the data in the design data sequence to obtain the occurrence frequency of the design data with each value, and use the Huffman coding algorithm to encode the data in the design data sequence according to the occurrence frequency to obtain a code sequence and the index relationship between the design data and the codes.
[0020] The present invention has the following beneficial effects: By setting multiple index tables in the present invention, the number of codes in each index table is small, so as to reduce the number of matching times when locating and matching codes, and improve the decoding efficiency; Furthermore, by setting an independent index table for each segment of the design sequence, when decoding the codes corresponding to each segment, only search and match in the index table corresponding to that segment, effectively improving the decoding efficiency; Furthermore, when performing segment control, the amount of intersection data between segments and adjacent segments is considered, so that the intersection data between segments is small, effectively reducing the storage amount of duplicate data in the index table and reducing unnecessary storage waste; Furthermore, when performing segment control, the number of data matching times in the decoding process is considered, so that the number of data matching times is small, effectively improving the decoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flowchart of the steps of a method for managing design data of a new type of server according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0023] The following will describe in detail the specific embodiments of the present invention with reference to the accompanying drawings.
[0024] Please refer to Figure 1 , which shows a flowchart of the steps of a method for managing design data of a new type of server provided by an embodiment of the present invention. The method includes the following steps: S1: Obtain a sequence composed of all the design data of the new type of server and denote it as the design data sequence.
[0025] Specifically, all design data of the new server are obtained, and the text or symbols in the design data are converted into numbers through ASCII conversion. The sequence composed of the converted design data is denoted as the design data sequence.
[0026] S2: Use the Huffman coding algorithm to encode the data in the design data sequence to obtain an encoded sequence.
[0027] Preferably, as an example, using the Huffman coding algorithm to encode the data in the design data sequence to obtain an encoded sequence includes: Statistical analysis is performed on the data in the design data sequence to obtain the occurrence frequency of each value of the design data. According to the occurrence frequency, the Huffman coding algorithm is used to encode the data in the design data sequence to obtain an encoded sequence and the index relationship between the design data and the encoding.
[0028] S3: Set a segmentation control function.
[0029] It should be noted that there are many encodings in the index table, resulting in a large matching workload. Therefore, the matching workload can be reduced by splitting one index table into multiple ones to reduce the number of encodings in each index table. If the index table is simply split into multiple index tables, the above problem cannot be solved because simple splitting cannot determine which encoding in the encoded sequence corresponds to which index table, so it is impossible to perform matching search in a single index table, and thus the matching workload cannot be reduced. Therefore, to reduce the matching workload, the encoding in the encoded sequence needs to be able to locate the corresponding index table.
[0030] It should be further noted that to enable the encoding to locate the corresponding index table, the design data sequence can be segmented so that one segment corresponds to one index table. When decoding the design data within a segment, only matching search needs to be performed in its corresponding index table. To ensure that when decoding each segment, the index table corresponding to other segments is not used, the index table needs to contain the correspondence between all the design data and the encoding in that segment. Since there may be the same design data between segments, this will result in duplicate design data and corresponding encodings in the index tables of different segments, which will increase the storage capacity. At the same time, the data in the design sequence has local similarity. For example, in the log data of the design sequence, the user activity volume is recorded. The user activity volume during working hours is small, and the user activity volume during off-duty hours is large. Therefore, the data values within a local area of the design sequence are similar, and the value differences between different local areas are large. Therefore, the design data sequence can achieve differentiation between segments.
[0031] Preferably, as an example, setting a segmentation control function includes:
[0032] Among them, the design data sequence is segmented into several feasible data segments. The index relationship between the design data and the encoding in a feasible data segment is placed in an independent index table to obtain a feasible local index table, and the index relationship between all the design data and the encoding in the design data sequence is placed in an index table to obtain an overall index table. represents the data set composed of the data in the i-th feasible data segment represents the data set composed of the data in the left adjacent feasible data segment of the i-th feasible data segment represents the data set composed of the data in the right adjacent feasible data segment of the i-th feasible data segment represents the intersection symbol represents the number of feasible data segments represents the number of matching times required to locate all the index relationships between the design data and the encoding using the feasible local index table represents the number of matching times required to locate all the index relationships between the design data and the encoding using the overall index table represents the segmentation control function
[0033] It can be understood that reflects the data intersection of the design data in each feasible data segment with the data in the two adjacent feasible data segments on both sides. The smaller this value is, the smaller the number of duplicate data in the feasible data segment and the two adjacent feasible data segments on both sides, and the less the amount of duplicate stored data in the index tables corresponding to different segmentations. Therefore, the feasible data segments obtained by this segmentation method can be used to set up the index table. reflects the reduction of the matching amount. The larger this value is, the greater the reduction of the matching amount caused by the feasible data segments obtained by this segmentation method. Therefore, the feasible data segments obtained by this segmentation method can be used to set up the index table.
[0034] It should be added that the feasible local index table also contains an index table positioning flag value, and the index table positioning flag value is the starting position of the encoding of the first design data in the feasible data segment corresponding to the feasible local index table in the encoding sequence.
[0035] The above embodiments involve feasible data segments, the number of matching times required to locate all the index relationships between the design data and the encoding using the feasible local index table, and the number of matching times required to locate all the index relationships between the design data and the encoding using the overall index table. Next, the determination methods for the feasible data segments, the number of matching times required to locate all the index relationships between the design data and the encoding using the feasible local index table, and the number of matching times required to locate all the index relationships between the design data and the encoding using the overall index table will be described.
[0036] First, the method for obtaining feasible data segments will be introduced.
[0037] Preferably, as an example, the method for obtaining the feasible data segments includes: Analyze the data in the design data sequence using the elbow method to obtain the number of segments.
[0038] Randomly select M - 1 segment positions in the design data sequence, and split the design data sequence at these segment positions to obtain several data segments, denoted as feasible data segments, where M represents the number of segments.
[0039] Then introduce the number of matching times required to locate the index relationship between all design data and the encoding using the feasible local index table.
[0040] Preferably, as an example, the number of matching times required to locate the index relationship between all design data and the encoding using the feasible local index table includes: Locate the starting encoding position corresponding to the feasible local index table in the encoding sequence according to the index table positioning flag value in the feasible local index table; take the encoding between the starting decoding position of the feasible local index table and the starting encoding position of the next feasible local index table as the analysis encoding of the feasible local index table; sequentially match the analysis encoding with the encodings in the feasible local index table until a matching encoding with the analysis encoding is located, obtain the number of matching times required to locate the matching encoding, and take the cumulative sum of the matching times of all analysis encodings as the local matching times of the feasible local index table; take the cumulative sum of the local matching times of all feasible local index tables as the number of matching times required to locate the index relationship between all design data and the encoding using the feasible local index table.
[0041] Finally, introduce the number of matching times required to locate the index relationship between all design data and the encoding using the global index table.
[0042] Preferably, as an example, the number of matching times required to locate the index relationship between all design data and the encoding using the global index table includes: Denote the encoding in the encoding sequence as the research encoding, sequentially match the research encoding with the encodings in the global index table until a matching encoding with the research encoding is located, obtain the number of matching times required to locate the matching encoding, and take the cumulative sum of the matching times of all research encodings as the number of matching times required to locate the index relationship between all design data and the encoding using the global index table.
[0043] S4: Use the segmentation control function to perform segmentation control on the design data sequence to obtain several target data segments; use the target data segments to set up a target local index table to implement the decoding process.
[0044] S40: Use the segmentation control function to perform segmentation control on the design data sequence to obtain several target data segments.
[0045] Preferably, as an example, the design data sequence is segmented and controlled by using a segmented control function to obtain a number of target data segments, including: The feasible data segment obtained when the segmented control function takes the maximum value is used as the target data segment.
[0046] S41: Set a target local index table by using the target data segment.
[0047] Preferably, as an example, setting a target local index table by using the target data segment includes: The feasible local index table obtained from the target data segment is used as the target local index table.
[0048] S42: To implement decoding processing.
[0049] Preferably, as an example, to implement decoding processing includes: Locate the starting coding position corresponding to the target local index table in the coding sequence according to the index table positioning flag value in the target local index table; The coding between the starting decoding position of the target local index table and the starting coding position of the next target local index table is used as the analysis coding of the target local index table.
[0050] Perform matching processing on the analysis coding and the coding in the target local index table in sequence until the matching coding with the analysis coding is located, and obtain the design data corresponding to the matching coding according to the index relationship between the matching coding in the target local index table and the design data, thus completing the decoding processing.
[0051] So far, this embodiment is completed.
[0052] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A design data management method for a new type of server, characterized in that: include: A sequence consisting of all design data of the new server is obtained and recorded as a design data sequence; Using the Huffman coding algorithm to encode the data in the design data sequence to obtain a coding sequence; Set the segment control function , divide the design data sequence into several feasible data segments, place the index relationship between the design data and the code in a feasible data segment in an independent index table to obtain a feasible local index table, and place the index relationship between all the design data and the code in the design data sequence in an index table to obtain an overall index table, represents the data set composed of the data in the i-th feasible data segment, , They represent the data sets composed of the data in the left adjacent feasible data segment and the right adjacent feasible data segment of the i-th feasible data segment, respectively. represents the intersection symbol, Indicates the number of feasible data segments, It represents the number of matches required to locate the index relationship between all design data and codes using the feasible local index table. It indicates the number of matches required to locate the index relationship between all design data and codes using the overall index table; Using the segmentation control function to segment the design data sequence to obtain a number of target data segments; The target data segment is used to set the target local index table to implement the decoding process.
2. The design data management method of a new server according to claim 1, characterized in that: The step of dividing the design data sequence into a plurality of feasible data segments comprises: Get the number of segments; Randomly select M-1 segmentation positions in the design data sequence, and divide the design data sequence at the segmentation positions to obtain a number of data segments, which are recorded as feasible data segments, and M represents the number of segments.
3. The design data management method of a new server according to claim 2, characterized in that: The obtaining of the number of segments includes: The elbow method is used to analyze the data in the design data sequence to obtain the number of segments.
4. The design data management method of a new server according to claim 1, characterized in that: The feasible local index table also includes an index table positioning mark value, and the index table positioning mark value is the starting position of the encoding of the first design data in the feasible data segment corresponding to the feasible local index table in the encoding sequence.
5. The design data management method of a new server according to claim 4, characterized in that: The method of using the feasible local index table to locate the number of matches required for the index relationship between all design data and the code includes: According to the index table positioning mark value in the feasible local index table, the starting coding position corresponding to the feasible local index table is located in the coding sequence; the code between the starting decoding position of the feasible local index table and the starting coding position of the next feasible local index table is used as the analysis code of the feasible local index table; the analysis code and the code in the feasible local index table are matched in sequence until the matching code with the analysis code is located, and the number of matches required to locate the matching code is obtained, and the cumulative sum of the matching numbers of all analysis codes is used as the local matching number of the feasible local index table; the cumulative sum of the local matching numbers of all feasible local index tables is used as the number of matches required to locate the index relationship between all design data and codes using the feasible local index table.
6. The design data management method of a new server according to claim 1, characterized in that: The method of using the overall index table to locate the number of matches required for the index relationship between all design data and codes includes: The codes in the coding sequence are recorded as research codes, and the research codes are matched with the codes in the overall index table in sequence until the matching codes with the research codes are located, and the number of matches required to locate the matching codes is obtained, and the cumulative sum of the number of matches of all research codes is used as the number of matches required to locate the index relationship between all design data and codes using the overall index table.
7. The design data management method of a new server according to claim 1, characterized in that: The step of using the segmentation control function to segmentally control the design data sequence to obtain a plurality of target data segments includes: The feasible data segment obtained when the segment control function takes the maximum value is used as the target data segment.
8. The design data management method of a new server according to claim 1, characterized in that: The step of setting a target local index table using the target data segment includes: The feasible local index table obtained from the target data segment is used as the target local index table.
9. The design data management method of a new server according to claim 1, characterized in that: The decoding process is implemented by: Based on the index relationship between the codes and the design data in the target local index table, the design data corresponding to each code in the code sequence is obtained and decoded.
10. The design data management method of a new server according to claim 1, characterized in that: The method of using the Huffman coding algorithm to encode the data in the design data sequence to obtain the coding sequence includes: The data in the design data sequence are statistically analyzed to obtain the frequency of occurrence of each value of the design data. The data in the design data sequence are encoded using the Huffman coding algorithm according to the frequency of occurrence to obtain a coding sequence and an index relationship between the design data and the coding.
Citation Information
Patent Citations
A log compression access method, system, device and readable storage medium based on Huffman coding and LZ77
CN117811588B
Immediate data compressed encoding method and system
CN107463355A
Data query method and device
CN113448957A
Multi-model learning index construction method and system for time sequence database
CN116644069A
Intelligent logistics query method and system
CN118377756A